machine-bridge-mcp 3.0.0-beta.104 → 3.0.0-beta.115
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +77 -0
- package/README.md +2 -2
- package/browser-extension/manifest.json +1 -1
- package/docs/ARCHITECTURE.md +14 -13
- package/docs/AUDIT.md +134 -0
- package/docs/CLIENTS.md +4 -4
- package/docs/LOGGING.md +2 -2
- package/docs/MANAGED_JOBS.md +7 -5
- package/docs/OPERATIONS.md +7 -7
- package/docs/PRIVACY.md +10 -2
- package/docs/PROJECT_STANDARDS.md +22 -0
- package/docs/RELEASING.md +4 -1
- package/docs/TESTING.md +8 -5
- package/docs/TOOL_REFERENCE.md +15 -8
- package/docs/UPGRADING.md +11 -4
- package/package.json +1 -1
- package/src/local/browser-request-registry.mjs +1 -1
- package/src/local/cli-ready-output.mjs +6 -6
- package/src/local/cli.mjs +1 -1
- package/src/local/computer-use-observation-contract.mjs +89 -0
- package/src/local/computer-use.mjs +3 -96
- package/src/local/job-runner.mjs +3 -0
- package/src/local/log.mjs +4 -5
- package/src/local/managed-job-directory.mjs +16 -14
- package/src/local/managed-job-durable-process.mjs +1 -0
- package/src/local/managed-job-hosted-reconcile.mjs +16 -0
- package/src/local/managed-job-hosted-status.mjs +13 -6
- package/src/local/managed-job-idempotency.mjs +29 -0
- package/src/local/managed-job-plan.mjs +5 -1
- package/src/local/managed-job-read-wait.mjs +75 -0
- package/src/local/managed-job-retention.mjs +2 -1
- package/src/local/managed-job-runner-claim.mjs +20 -4
- package/src/local/managed-job-runner-liveness.mjs +35 -0
- package/src/local/managed-job-runner.mjs +2 -25
- package/src/local/managed-jobs.mjs +25 -29
- package/src/local/numbers.mjs +6 -0
- package/src/local/process-error-message.mjs +6 -0
- package/src/local/process-execution.mjs +3 -5
- package/src/local/process-identity.mjs +51 -0
- package/src/local/process-session-events.mjs +0 -5
- package/src/local/process-session-read.mjs +21 -1
- package/src/local/process-session-remote-poll.mjs +10 -4
- package/src/local/process-sessions.mjs +7 -4
- package/src/local/relay-call-recovery.mjs +33 -33
- package/src/local/relay-recovery-admission.mjs +19 -0
- package/src/local/relay-recovery-capacity.mjs +29 -0
- package/src/local/relay-recovery-diagnostics.mjs +28 -0
- package/src/local/relay-redelivery-safety.mjs +43 -0
- package/src/local/relay-result-retention.mjs +66 -0
- package/src/local/relay-tool-schema-extension.mjs +31 -0
- package/src/local/resource-admission.mjs +26 -28
- package/src/local/resource-host-cache.mjs +1 -1
- package/src/local/resource-host-darwin.mjs +1 -11
- package/src/local/resource-host-snapshot.mjs +1 -19
- package/src/local/resource-lease-liveness.mjs +28 -0
- package/src/local/resource-pressure.mjs +25 -6
- package/src/local/resource-probe-command.mjs +1 -14
- package/src/local/resource-process-ancestry-cache.mjs +3 -12
- package/src/local/resource-process-ancestry.mjs +1 -7
- package/src/local/resource-staging-recovery.mjs +12 -10
- package/src/local/resource-transaction-lock.mjs +37 -44
- package/src/local/resource-waiters.mjs +7 -9
- package/src/local/runtime-diagnostic-state.mjs +16 -0
- package/src/local/runtime-diagnostics.mjs +3 -3
- package/src/local/runtime-info-projection.mjs +5 -0
- package/src/local/runtime-relay.mjs +4 -1
- package/src/local/runtime-reporting.mjs +6 -0
- package/src/local/runtime-tool-handlers.mjs +9 -1
- package/src/local/runtime.mjs +16 -24
- package/src/local/short-identifiers.mjs +4 -0
- package/src/local/stdio.mjs +2 -2
- package/src/local/tool-executor.mjs +3 -20
- package/src/shared/log-redaction.d.mts +2 -0
- package/src/shared/log-redaction.mjs +8 -0
- package/src/shared/mcp-protocol.d.mts +3 -1
- package/src/shared/mcp-protocol.mjs +28 -3
- package/src/shared/relay-contract.json +8 -1
- package/src/shared/server-metadata.json +5 -3
- package/src/shared/tool-catalog.json +15 -8
- package/src/worker/daemon-http-controller.ts +1 -1
- package/src/worker/daemon-http-protocol.ts +2 -1
- package/src/worker/daemon-ready-dispatch.ts +9 -3
- package/src/worker/daemon-ready-waiter-policy.ts +55 -0
- package/src/worker/daemon-ready-waiter-state.ts +53 -0
- package/src/worker/daemon-ready-waiters.ts +27 -44
- package/src/worker/index.ts +90 -74
- package/src/worker/managed-job-read-timeout.ts +37 -0
- package/src/worker/mcp-controller.ts +42 -11
- package/src/worker/mcp-request-cancellation.ts +112 -0
- package/src/worker/mcp-response-cancel.ts +25 -0
- package/src/worker/mcp-response-proxy.ts +12 -17
- package/src/worker/mcp-response-stream.ts +0 -15
- package/src/worker/mcp-stale-schema-compat.ts +1 -1
- package/src/worker/mcp-streamed-tool-response.ts +19 -0
- package/src/worker/mcp-subscription-capacity.ts +47 -0
- package/src/worker/mcp-subscription-contract.ts +6 -0
- package/src/worker/mcp-subscription-registry.ts +122 -0
- package/src/worker/mcp-subscription-stream.ts +69 -0
- package/src/worker/observability.ts +2 -3
- package/src/worker/pending-call-capacity.ts +14 -0
- package/src/worker/pending-call-registration.ts +43 -0
- package/src/worker/pending-calls.ts +10 -15
- package/src/worker/server-info-tool-delivery.ts +24 -1
- package/src/worker/server-info.ts +3 -2
- package/src/worker/tool-call-recovery.ts +16 -0
- package/src/worker/tool-catalog.ts +22 -14
- package/src/worker/tool-timeout.ts +5 -5
- package/src/worker/worker-edge-log.ts +8 -7
- package/src/worker/worker-mcp-config.ts +7 -3
- package/src/worker/pending-admission.ts +0 -15
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,82 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 3.0.0-beta.115 - 2026-08-22
|
|
4
|
+
|
|
5
|
+
- Remove ChatGPT-internal `api_tool` / tool-loader cache inspection from release and upgrade acceptance entirely. Release evidence now stops at the governed Workspace Action control snapshot when that product layer applies, plus live MCP/runtime discovery, subscription behavior where the protocol contract changed, and harmless changed-generation invocation behavior. Historical loader-split observations remain audit history only; release automation must not query, compare, or grade those host-internal cache views.
|
|
6
|
+
- Correct the ChatGPT schema-freshness evidence model after live beta.115 review reproduced a loader split that invalidates the earlier “complete unfiltered catalog is authoritative” assumption. In the same connected session, an unfiltered `api_tool.list_resources` load retained generation 5 while a query-targeted resource load returned generation 6; live `server_info` reported beta.115/schema generation 6, Workspace Action control already showed the refreshed 54-action generation-6 snapshot, and a valid-but-unretained `read_job` reached generation-6 runtime behavior and returned typed `not_found`. Release/upgrade/client/testing guidance now treats both filtered and unfiltered `api_tool` resource bundles as diagnostic model/tool-loader caches rather than Workspace publication authority. Product-level freshness is established from the governed Workspace Action control snapshot when applicable plus live runtime/discovery and changed invocation behavior. A stale loader bundle alone must never trigger connector recreation or app republication.
|
|
7
|
+
- Remove the fixed >100-minute live managed-job soak from candidate acceptance. Long-running continuity remains release-blocking where affected, but acceptance now uses bounded evidence matched to the failure mode: hosted default long-poll survival, a second same-response `read_job`, durable same-`job_id` recovery across any exercised host/tool or daemon/relay boundary, representative helper churn, cancellation/cleanup, and replay-safety checks. A long soak may still be run as diagnostics, but elapsed duration alone is no longer a mandatory release gate.
|
|
8
|
+
- Supersede the activated beta.114 candidate after a fresh independent post-activation review deliberately reopened its interruption assumptions, architecture, failure branches, logging, tests, documentation, package inventory, and privacy surfaces. The review did **not** find another runtime cancellation defect: one-shot execution and process sessions retain a strict same-isolate final `AbortSignal` check immediately before synchronous `spawn`, while cancellation-aware managed-job main steps retain their final cross-process marker read immediately before `spawn` as the documented launch decision point. The distinction is now explicit: `cancel_job` acknowledges a cancellation request and post-decision cancellation terminates owned work; it does not claim a stronger wall-clock guarantee that a concurrently racing OS spawn syscall can never win after the runner's final marker read. Implementing that different guarantee would require a serialized cross-process launch/cancel protocol rather than another asynchronous check.
|
|
9
|
+
- Turn the review conclusions into maintenance guards instead of prose-only confidence. `src/local/job-runner.mjs` now has an architecture line budget near its current size so the most stateful managed-job module cannot continue expanding without an explicit extraction decision. The existing managed-job boundary assertion still pins the final cancellation-marker check directly to `spawn`, while integration tests cover cancellation/termination/cleanup and deterministic one-shot/session admission-time cancellation requires zero spawn plus exact lease release. Managed-job documentation now separates the launch decision point, post-dispatch cancellation, durable terminal authority, and idempotency/replay responsibilities.
|
|
10
|
+
- Make the cancellation safety valve diagnosable without weakening it or leaking request identity. If bounded pre-open MCP cancellation tombstones are exhausted, the Worker still fails closed rather than evicting evidence and potentially dispatching an orphaned side effect, but the first overflow in each fail-closed window now increments the existing low-cardinality `mcp_request_cancellation_fail_closed` observability counter. Repeated overflow during the same bounded window does not spam the counter, callback failure cannot change cancellation behavior, and no stream ID, account identifier, tool arguments, path, or response content is recorded.
|
|
11
|
+
- Tighten publication hygiene without discarding useful audit history. The beta.114 audit note that referred to beta.113 as a present-tense active runtime is rewritten as review-time history and linked to the later post-activation evidence. Tracked release/privacy/audit text no longer publishes a machine-specific ignored application-state directory name; it documents the reusable rule that ignored local databases remain archive-sensitive. The npm package intentionally continues shipping the engineering/security audit record, so its size remains an explicit transparency tradeoff rather than being reduced by silently removing historical security evidence.
|
|
12
|
+
- Advance package/runtime identity to **3.0.0-beta.115** while keeping the hosted tool contract at **schema generation 6**. Because the follow-up changes packaged documentation after beta.114 activation, the beta.114 tarball, OAuth canary, list-change subscription proof, and interrupted long-soak evidence remain diagnostic evidence only for those exact beta.114 bytes. Beta.115 requires a new frozen-tree verification receipt and candidate/install-only preparation; live activation and renewed acceptance remain separately owner-authorized release stages.
|
|
13
|
+
|
|
14
|
+
## 3.0.0-beta.114 - 2026-08-22
|
|
15
|
+
|
|
16
|
+
- Supersede beta.113 after its live >100-minute continuity attempt exposed an aggregate host-boundary/recovery defect that per-call 40-second pacing did not solve. The activated beta.113 Worker/daemon, packaged OAuth canary, complete unfiltered 54-tool generation-5 catalog, validator probes, `toolsListChanged` proof, and short unchanged-status fixture all passed. The long soak then produced roughly 125 consecutive structured 40-second `read_job` checkpoints before the host imposed a real response/tool boundary. A later explicit 60-second active read also succeeded, proving that per-call lifetime and aggregate assistant-response lifetime are independent constraints. Release guidance no longer treats `duration / wait interval` arithmetic as long-duration continuity evidence: an actual >100-minute job must either reach terminal state in one response or remain recoverable by the same `job_id` after a real host boundary.
|
|
17
|
+
- Protect explicit managed-job recovery evidence from ordinary helper-command churn without increasing retained private state. `exec_command`, `run_process`, and `run_local_command` one-step carriers now persist an internal `transient_process` retention class. Under the existing 50-state capacity bound, safely removable transient terminal history is reclaimed before explicit managed-job terminal history. The seven-day terminal and 24-hour staged expiry bounds remain unchanged, the marker is not part of public job projection, and it contains no argv/path/output/resource data. A valid job ID whose retained directory is gone now returns fixed non-retryable `not_found` instead of generic `execution_failed`; that absence is explicitly not proof that the underlying operation never executed and must not be converted into blind side-effect replay.
|
|
18
|
+
- Fail closed per call when relay result evidence is lost without turning one old incident into a daemon-lifetime outage. A completed result retained for Worker acknowledgement can expire after the bounded settlement lifetime or be discarded when reconnect grace expires. The affected call ID now enters a bounded private replay-safety tombstone set, so later `resume_calls` cannot misclassify that already-executed call as daemon-proven missing, while unrelated resumed IDs can still recover normally. Tombstones retire once the Worker no longer resumes that ID; only exhaustion of the bounded tombstone set escalates to global fail-closed redelivery disablement. `diagnose_runtime.runtime.relay_result_recovery` exposes aggregate `automatic_redelivery_safe`, `unsafe_call_tombstones`, and `global_redelivery_disabled` state without call IDs, tool names, arguments, or result content.
|
|
19
|
+
- Make hosted managed-job acceptance recoverable before side effects begin. The generation-6 Worker schema requires a caller-held `idempotency_key` for remote `start_job`, the actual invocation validator rejects a missing key before daemon dispatch, and timeout/disconnect errors carry the same idempotent-replay recovery descriptor already used by remote one-step process tools. The recovery descriptor identifies the key's source as the original request instead of echoing the key value into another host-visible payload. Retrying the same `start_job` arguments with the same key can therefore recover an ambiguous acceptance response instead of creating a second job. Local/stdio `start_job` keeps the underlying optional-key API; the stricter requirement is a hosted transport safety contract.
|
|
20
|
+
- Close the pre-dispatch stream-cancellation ordering race. A private stream cancel that reaches the Durable Object before the matching direct request is registered is now retained as a bounded short-lived cancellation tombstone; a direct request whose public `AbortSignal` was already aborted is likewise cancelled when ownership opens. The dispatch gate rechecks that signal before pending-call registration and daemon send. Tombstone overflow fails closed for the same bounded lifetime instead of evicting cancellation evidence and later starting an orphaned side effect, then automatically recovers when the retention window expires.
|
|
21
|
+
- Close the analogous local process-launch handoff window after resource admission. One-shot execution, interactive process sessions, and cancellation-aware managed-job main steps now perform their final synchronous cancellation check after launch arguments/resource ownership are ready and immediately before `spawn`, with no asynchronous boundary between the decision and dispatch. Cancellation visible at that decision point releases the admitted lease and remains definite non-execution; once dispatch wins, the existing process/session termination path keeps `side_effects_started=true`/pending-or-unknown settlement semantics instead of claiming safe replay. Deterministic one-shot/session regressions inject cancellation from resource admission and require zero spawns plus exact lease release; managed-job boundary coverage pins the durable-marker check directly to spawn and the full managed-job cancellation/cleanup integration remains green.
|
|
22
|
+
- Bound public MCP stream cancellation when its private Durable Object control request never settles. Normal request/reader cancellation still waits for the private cancel settlement before exposing public closure, but that internal control fetch now has its own two-second abortable deadline. A regression uses a never-resolving private fetch and proves the public SSE response cannot hang indefinitely. No request arguments, result bytes, paths, or credentials are added to logs or persistent state.
|
|
23
|
+
- Advance the hosted orchestration contract to **tool schema generation 6** and package/runtime identity to **3.0.0-beta.114**. Host-visible guidance now distinguishes per-call pacing from aggregate host-response lifetime, requires preserving the same durable identifier across a real host/tool boundary, and tells clients that `read_job` `not_found` is missing recovery evidence rather than a retry proof. Beta.113 live evidence remains diagnostic only; beta.114 requires fresh frozen verification, candidate preparation, explicit owner-authorized activation, packaged OAuth canary, generation-6 complete-catalog/validator proof, and renewed long-duration recovery acceptance before release acceptance.
|
|
24
|
+
|
|
25
|
+
## 3.0.0-beta.113 - 2026-08-22
|
|
26
|
+
|
|
27
|
+
- Correct the host-schema acceptance model after live beta.112 evidence proved that ChatGPT's filtered tool/resource search can retain a query-keyed stale schema snapshot independently of the complete Machine Bridge catalog. Repeated unfiltered discovery returned all 54 Machine Bridge tools at schema generation 5 with the 40,000 ms hosted `read_job` default, while an immediate filtered `read_job` search could still return a generation-4 bundle claiming the former 300,000 ms omitted default; the two result families alternated without replacing one another. Source tracing excludes a Machine Bridge query-dependent catalog path, and live state excluded old jobs, daemon sockets, and Worker isolates as the owner of the stale definition.
|
|
28
|
+
- Bind freshness to evidence the repository can actually interpret. Release, upgrade, testing, client, operations, audit, and project-standard guidance now treats one complete unfiltered host catalog snapshot as the schema-freshness authority and keeps filtered tool/resource search as routing/index evidence only. A stale filtered result that conflicts with the complete catalog is recorded as host query/search-index cache staleness and cannot by itself classify the live runtime as mixed-generation or block acceptance. Harmless previous-generation invocation-validator/behavior probes remain mandatory, because displayed schema freshness still does not prove the host invocation layer or daemon validation path accepted the new boundary.
|
|
29
|
+
- Keep the public hosted tool schema at **generation 5**. A control fixture loaded the stale generation-4 filtered definition, then an omitted-parameter `read_job` still executed with the live generation-5 40,000 ms default and returned generation-5 continuation metadata before a second same-response read reached terminal state. Machine Bridge cannot purge or determine the scope of the host's secondary filtered-search cache, so beta.113 removes the false release blocker rather than adding an unsupported cache-clear mechanism. The package identity advances because the corrected release/client/audit contracts are packaged bytes even though the public tool schema itself is unchanged.
|
|
30
|
+
|
|
31
|
+
## 3.0.0-beta.112 - 2026-08-21
|
|
32
|
+
|
|
33
|
+
- Supersede the activated beta.111 candidate after the required cross-generation invocation probe exposed a second schema split inside Machine Bridge. Beta.111 itself activated cleanly as Worker/login daemon, its activated-package OAuth canary passed authorization-code exchange, authenticated MCP, refresh rotation, refreshed MCP, and cleanup, and the 55-second live continuity fixture proved the beta.110 projection repair: the first omitted-parameter `read_job` remained inside one MCP call for 40 seconds and retained `host_turn_handoff_recommended=false`, `status_polling_mode=bounded_followup`, `tool_schema_generation=5`, and the other hosted metadata before a second same-response read reached terminal state. Host discovery also exposed generation-5 long-running tool definitions and a non-executing `stage_job(timeout_seconds=3601)` crossed the former one-hour boundary. However, a terminal `read_job(wait_ms=40001)` reached Machine Bridge and was rejected by the daemon with the local static schema maximum of 40,000 ms. The Worker correctly overlays hosted discovery to a 300,000 ms maximum, and the runtime wait implementation already permits that relay maximum, but `src/local/tool-executor.mjs` still validated every invocation against the static local catalog before dispatch. Beta.112 adds the missing relay-only validation extension, symmetric with the existing hosted durable-process timeout extension: only a relay `read_job` whose sole static-schema issue is `/wait_ms` `maximum` may use 40,001..300,000; malformed, negative, non-integer, or >300,000 values remain invalid, and local/stdio still rejects values above 40,000. Direct regression coverage binds all of those cases.
|
|
34
|
+
- Keep the public hosted tool schema at **generation 5** and advance package/runtime identity to **beta.112**. The public schema did not change; the implementation now makes the already-advertised generation-5 hosted `read_job` maximum executable through the daemon defense-in-depth validator. Beta.111 activation/canary/continuity evidence remains useful diagnosis but cannot authorize beta.112 acceptance because packaged runtime and documentation bytes changed after live activation.
|
|
35
|
+
|
|
36
|
+
## 3.0.0-beta.111 - 2026-08-21
|
|
37
|
+
|
|
38
|
+
- Supersede the activated beta.110 candidate after its required live continuity fixture exposed a result-projection defect. The exact beta.110 Worker/daemon converged on tool schema generation 5 and its activated-package OAuth canary passed authorization-code exchange, authenticated MCP, refresh rotation, refreshed MCP, and cleanup. A 55-second durable fixture then proved the first parameter-omitting `read_job` correctly stayed inside one MCP call for 40 seconds, but the unchanged timeout checkpoint lost `host_turn_handoff_recommended=false`, `status_polling_mode=bounded_followup`, `tool_schema_generation=5`, and related hosted metadata because each lightweight `readProgress()` probe replaced the richer `readHosted()` state. The long-poll now merges unchanged lightweight progress into the full state instead of replacing it, with a regression that gives hosted metadata only to full reads so the original live failure is reproduced by the old implementation. At that point the same ChatGPT conversation still rejected the harmless generation-5 `read_job(wait_ms=40001)` probe at 40,000 ms while live `server_info` reported beta.110/generation 5, so the evidence was conservatively treated as an external approved-action snapshot blocker. Beta.111 live retesting later refined that diagnosis: generation-5 host definitions and the 3,601-second staged-job boundary were accepted, while the 40,001 ms call reached Machine Bridge and failed in the daemon's static local catalog validator. The beta.112 entry records that corrected causal classification. Because the Machine Bridge projection repair changes packaged runtime bytes after live beta.110 activation, beta.110 activation/canary evidence remains diagnostic only and cannot authorize beta.111 acceptance.
|
|
39
|
+
|
|
40
|
+
- Independent 2026-08-21 follow-up reviews removed additional correctness/evidence hazards before release. Resource transaction reads no longer expand the fixed four-attempt/1 ms multiple-hard-link publication retry into transaction-deadline waiting; hosted managed-job liveness now reuses the secure runner-claim reader instead of bypassing path-identity/hard-link checks; `server_info` reports a server-opened list-change subscription rather than claiming client observation and explicitly says client receipt is unobservable; and an explicit five-minute `read_job` opt-in retains a 310-second daemon / 315-second settlement envelope so its non-current-runner fallback can cover both bounded process-generation probes. Live host probing then disproved the five-minute **default**: the model-facing schema advertised 300,000 ms, the same host's invocation validator still rejected values above 40,000 ms, an omitted wait entered the server-side five-minute default and the host terminated the tool call with `TimeoutError`, while an explicit 40,000 ms read completed normally. The hosted default is therefore restored to the empirically survivable 40 seconds while the 300-second explicit maximum remains available to clients that can actually carry it. Recovery now also refuses a `read_job` dispatch/redelivery when less than the dedicated ten-second reconciliation headroom remains instead of silently rewriting it to an impossible immediate read. Release acceptance includes harmless cross-generation argument probes rather than relying on discovery text alone. Current-tree/history privacy scanning and both npm audit modes remained clean; ignored application-local databases are documented as archive-sensitive local state without publishing machine-specific application directory names, and any pre-review beta.109 candidate artifact is historical evidence only because these source changes invalidate it for promotion.
|
|
41
|
+
- Close two additional interruption/authority gaps found while tracing the hosted call lifecycle. The managed-job long-poll deadline now begins before the initial full reconciliation and unchanged timeout return no longer appends another heavyweight reconciliation after the advertised wait, so a nominal 40-second hosted read cannot structurally become `initial reconcile + 40s + final reconcile`. Worker daemon dispatch also removes the redundant asynchronous `PendingAdmissionGate`: an already-ready daemon is selected synchronously and capacity check -> pending registration -> daemon send has no `await`, while daemon-down requests remain in the existing authority-bound ready-waiter registry. This removes an otherwise unowned queued state in which account/client/family revocation could occur after request authorization but before the request became visible to either waiter cancellation or pending-call cancellation.
|
|
42
|
+
- Tighten daemon handover and relay-result recovery under real transport interruptions. Ready-daemon waiters are released only after WebSocket/HTTPS handover reaches its final no-`await` point; while waiters exist, a newly visible ready socket remains handover-in-progress so fresh calls cannot leapfrog capacity or escape authority revocation. Waiter timeout no longer grabs a merely visible socket without formal readiness release. On the daemon, completed results awaiting Worker acknowledgement now consume the same 16-total / 14-ordinary / 2-reserved-control recovery ownership budget as active relay calls; ordinary overflow is rejected before execution with `side_effects_started=false`, the retained-result store has a separate 16-entry hard ceiling plus a monotonic 315-second acknowledgement lifetime, and diagnostics expose only aggregate recovery counts. This prevents lost acknowledgements from turning a network episode into unbounded memory/private-result retention while keeping diagnosis capacity available.
|
|
43
|
+
- Keep the hosted tool schema at **generation 5** and advance the package/runtime identity through **beta.111**. The activated-but-unpublished beta.109 runtime used generation 4 for the five-minute default and initial freshness repair; beta.110 was the first generation-5 candidate carrying the host-safe 40-second default and reconnect/subscription repairs, but its live unchanged-timeout projection defect requires new packaged bytes. Reusing either `3.0.0-beta.109` or `3.0.0-beta.110` would let activation convergence checks (`daemon.version` / `worker.version`) confuse earlier runtime bytes with the repaired candidate. Candidate/host acceptance must therefore converge on beta.111 generation 5 before these interruption fixes are treated as delivered.
|
|
44
|
+
- Bound `toolsListChanged` subscription ownership even when the Workers HTTP runtime does not surface client disconnect. Real Worker integration showed that aborting the public client closes its local response body without reliably aborting/cancelling the Durable Object stream—even after a public heartbeat interval—so relying only on request/stream cancellation can leave a ghost subscription consuming one of the 8-per-account / 32-global freshness slots. Current tool definitions are immutable within one Worker deployment and every listen receives an immediate level-trigger edge, so the stream now has a 10-second server lease in addition to explicit cancellation/authority revocation. `server_info.tool_delivery` advertises the lease; integration proves active capacity returns by that bound while server-opened history remains diagnostic only.
|
|
45
|
+
- Supersede beta.108 after live host evidence showed that the long-task correction was incomplete. Beta.108 removed the high-density immediate-status loop, but it still advertised `tools.listChanged=false`; the same ChatGPT connector subsequently exposed a mixed-generation catalog in which some tools carried generation 3 while other long-running tools still carried beta.104/beta.106 guidance about one-read handoff or a guessed host response/execution budget. Zero discovery TTL is not enough when the server simultaneously tells the client that the tool list never changes. Publication of beta.108 was therefore stopped before tag or GitHub Release.
|
|
46
|
+
- Implement the current MCP 2026-07-28 tool-list freshness contract instead of relying on eventual cache expiry. Current remote discovery advertises `tools.listChanged=true`; a client that opens `subscriptions/listen` with `toolsListChanged=true` receives a correlated `notifications/subscriptions/acknowledged`, an immediate level-trigger `notifications/tools/list_changed`, and a request-scoped SSE subscription bounded by explicit cancellation/authority revocation or the advertised 10-second server lease so it can re-fetch `tools/list`. The initialization-era 2025 compatibility surface continues to advertise `listChanged=false` because that protocol family uses different notification semantics. Subscription streaming lives in its own Worker module rather than expanding ordinary tool-response streaming responsibility. One Durable Object permits at most 32 simultaneously active subscriptions and one authenticated account at most 8; over-capacity requests fail before stream allocation, and abort/cancel/lease expiry returns the slot exactly once. This prevents both unbounded stream accumulation and one delegated account monopolizing the freshness channel for every other account. `server_info.tool_delivery` additionally exposes only the current account's active-subscription count, lease, and whether the server successfully opened a freshness subscription for that account during the current Durable Object instance; opened-stream history is bounded to the most recent 64 accounts. These are deliberately server-side facts only: `tools_list_change_subscription_client_receipt_observable=false` records that Machine Bridge cannot prove the external client read either SSE frame or refreshed its catalog. The fields do not leak another account's activity and are diagnostics, not an override for a host product's separate approval/cache policy.
|
|
47
|
+
- Correct the ChatGPT rollout premise after live beta.109 activation falsified the earlier universal-subscription assumption. The activated Worker/daemon immediately reported generation 4, the five-minute `read_job` contract, and the six-hour managed-job step ceiling, while the same ChatGPT host continued exposing its previously approved generation-3 action definitions and `server_info` correctly remained `tools_list_change_subscription_opened_for_account=false`. Current ChatGPT workspace-app behavior freezes approved tool/input definitions rather than automatically replacing them when an MCP server changes. Enterprise/Edu owners/admins refresh and review changed actions through Workspace Settings -> Apps -> Action control -> Refresh; a published Business custom app must instead be recreated and republished. A server-opened subscription is useful evidence that a client attempted the live freshness path, but it cannot prove frame receipt or catalog refresh and is not a universal ChatGPT freshness gate. Whole-catalog live host generation convergence remains release-blocking after the plan-appropriate ChatGPT action-snapshot update.
|
|
48
|
+
- Bound long-task interaction density without exceeding the hosted client's demonstrated per-call lifetime. Relay-origin `read_job` now defaults to a 40-second server-side wait and may explicitly request up to five minutes. A 100-minute unchanged job therefore needs at most 150 default hosted reads—far below the earlier rapid-loop incident class—while clients with verified longer request lifetimes may opt into fewer calls. The public SSE proxy emits a five-second heartbeat while the tool promise is pending. Internal lightweight status probes run every five seconds and full runner/recovery reconciliation at thirty-second intervals inside the same wait deadline. The deadline starts before the initial full read, and an unchanged timeout does not perform another heavy reconcile after the deadline; this prevents nominal wait duration from hiding extra reconciliation latency at both ends. The default call receives a 50-second daemon / 55-second settlement envelope; the explicit five-minute maximum receives 310/315 seconds. Pre-dispatch daemon recovery is deducted from those absolute envelopes and the requested wait is shortened accordingly. If fewer than ten seconds remain, the Worker fails retryably before dispatch/redelivery rather than pretending a recovery-sensitive read can safely execute as `wait_ms=0`. Ordinary remote tools retain the existing 50-second relay ceiling, and local/stdio reads remain immediate by default with their bounded local maximum.
|
|
49
|
+
- Remove the one-hour single-step ceiling that still contradicted the 100+ minute continuity target. `stage_job`/`start_job` main and finally steps retain their ten-minute default but may now explicitly request up to 21,600 seconds (six hours); the managed runner starts that execution timer only after cooperative resource admission. Remote one-step process tools remain capped at 600 seconds and continue to route legitimately longer continuous commands to `start_job`. Tests bind the local plan validator and both local/Worker schemas to the same six-hour ceiling, accept a 6,100-second staged step without executing it, and reject 21,601 seconds.
|
|
50
|
+
- Correct a false macOS host-pressure model that could block work even after the machine was idle. `iostat -Id` reports transfers and throughput activity, not SSD saturation, so the former 5,000-IOPS / 220-MB/s hard-red thresholds could produce `disk_iops_critical` / `host_pressure_red` from a short healthy NVMe burst. Fresh raw I/O throughput now contributes only Yellow concurrency pressure; Red requires independent critical evidence or a severe load backlog corroborated by current CPU/I/O evidence. Cached I/O hints expire after the same five-second freshness window instead of carrying one transient peak for thirty seconds. Live diagnostics during the reported incident class showed a Green 8-core host with roughly two busy cores, 41% free memory, no leases/waiters, and a direct one-second `iostat` sample of only 75 transfers / 0.62 MB, confirming that an already-idle machine must not remain blocked by the old peak model.
|
|
51
|
+
- Remove two smaller interruption/privacy hazards found by the independent follow-up review. Hosted `read_job` full reconciliation no longer performs a synchronous process-start `ps` probe every thirty seconds for a healthy runner: relay reads first verify runner ownership through the async process-identity path and enter the legacy synchronous recovery reconciliation only when the runner is non-current or ambiguous. Browser request dispatch also no longer exposes a raw lower-level `transport.send()` exception on read-only operations; the public failure is the fixed `browser extension send failed` message, while mutating operations retain their existing non-replayable unknown-outcome projection.
|
|
52
|
+
- The predecessor beta.109 activation established generation-4 live-host evidence but was never published to npm. Beta.110 invalidates that acceptance because its version identity, generation-5 tool contract, 40-second default continuation, handover ordering, relay-result ownership bounds, and subscription lease all differ. Beta.110 requires fresh frozen verification, candidate preparation, owner activation, activated-package OAuth canary, protocol-level subscription/lease verification, the plan-appropriate ChatGPT action-snapshot refresh or republication when ChatGPT is the hosted client, whole-catalog generation-5 verification, invocation-validator probes, and long-duration continuation proof before publication resumes.
|
|
53
|
+
|
|
54
|
+
## 3.0.0-beta.108 - 2026-08-20
|
|
55
|
+
|
|
56
|
+
- Correct the remaining autonomous-long-task orchestration defect after owner evidence disproved the earlier fixed-wall-clock explanation. The same hosted environment had repeatedly supported Machine Bridge work exceeding 100 minutes before the recent regressions, while beta.104's retained audit already demonstrated the relevant amplification class: 497 settled Machine Bridge calls in 25 minutes, including 91 immediate `read_job` calls. Beta.106 correctly removed the forced later-user-turn handoff, but it restored same-response `read_job` follow-up without moving durable-job waiting into the server, so an unchanged active job could again be followed through high-density host tool calls. The exact ChatGPT host quota/termination algorithm is not observable here and is not claimed; the supported causal conclusion is that Machine Bridge unnecessarily amplified host interactions and therefore could exhaust an external hosted-turn resource much sooner than the underlying long task required.
|
|
57
|
+
- Move waiting back inside Machine Bridge instead of consuming one hosted tool interaction per checkpoint. `read_job` now has a bounded `wait_ms` contract; local/stdio reads remain immediate by default, while relay-origin reads with omitted `wait_ms` default to a **40-second server-side long-poll**, return early on meaningful status/phase progress or terminal state, and accept `wait_ms=0` only as an intentional immediate checkpoint. The Worker gives that long-poll a 45-second daemon execution budget plus the existing settlement envelope, while its 5-second SSE heartbeat keeps the response stream active. At that default interval, a 100-minute unchanged durable job requires at most 150 hosted `read_job` calls instead of a rapid checkpoint loop. One-second internal probes read only the secure status record; full runner/recovery reconciliation is bounded to ten-second intervals plus the authoritative return path, so reducing host-call density does not create a synchronous process-identity loop inside the daemon. The same composition defect existed in interactive process-session pacing: beta.104 converted a repeated would-block read inside the fifteen-second cooldown into an immediate running result. That made sense only while beta.104 forced a user-turn handoff; under beta.106 same-turn autonomy it became another rapid-call path. The actual output/exit blocking wait remains capped at one second, but a repeated would-block request now stays inside the same MCP call until output/exit or the cooldown boundary, and the Worker reserves a 20-second execution / 25-second settlement envelope for that server-side pacing. For `wait_for_exit=true`, ordinary stdout/stderr changes do not release the cooldown stage early; the same call continues until process exit or the cooldown deadline before entering the bounded one-second exit wait. `server_info.tool_delivery` reports the managed-job read wait bounds, `host_turn_deadline_observable=false`, and `managed_jobs_detached_from_mcp_response=true`; these fields describe Machine Bridge's observable boundary without inventing a fixed host wall-clock limit.
|
|
58
|
+
- Make hosted schema freshness explicit instead of assuming a daemon/Worker replacement refreshes the client's cached tool contract. Both `server/discover` and `tools/list` now advertise `ttlMs=0`; hosted tool descriptions carry **tool schema generation 3**; `server_info.tool_delivery` reports the generation, live server version, discovery/tool-list TTLs, and `host_visible_schema_known_to_server=false`. The request-scoped server deliberately continues to advertise `tools.listChanged=false` because it has no persistent notification registry and does not claim to emit a notification it cannot deliver. Activation/acceptance must therefore compare the active client's host-visible generation with the live server generation whenever hosted tool semantics change.
|
|
59
|
+
- Remove the remaining resource-coordinator event-loop blocking path uncovered by the independent second review. Production host/Darwin/process-parent resource sampling is async-only; lease/waiter pruning now uses a cached lock-external async process-start snapshot and performs only `kill(0)` plus in-memory generation comparison while the transaction lock is held. Missing snapshot evidence fails closed by retaining a live owner; a proven PID-generation mismatch remains reclaimable and takes precedence over a numerically live isolated process group. Lease/waiter staging recovery uses the same snapshot evidence, while the beta.104 legacy transaction-owner migration reader now performs its process-generation check asynchronously before destructive recovery. The current complete-before-visible `transaction.lock` owner-state writer also uses the shared bounded transient-multiple-link retry on reads: the short `link(staging,target)` publication window may expose `nlink=2` to another process, but the reader never parses or reclaims that state until it settles; a persistent multiply-linked lock still fails closed after the bounded retry. Tests cover live-owner fail-closed behavior, PID reuse, process-group reuse, in-flight snapshot coalescing, stale staging, transient publication races, persistent hard-link rejection, and the absence of blocking resource-probe transport.
|
|
60
|
+
- These runtime, Worker, protocol, tests, and documentation changes supersede beta.107 source evidence and beta.106 live acceptance. Beta.108 requires a fresh frozen full verification, audits/signatures/SBOM, Worker/package dry runs, exact release candidate, persistent activation, host-visible schema-generation check where the client supports refresh, and live same-assistant-response multi-read proof before it can be accepted or published.
|
|
61
|
+
|
|
62
|
+
## 3.0.0-beta.107 - 2026-08-20
|
|
63
|
+
|
|
64
|
+
- Promote the beta.106 autonomous follow-up correction into a repository hard invariant. `AGENTS.md` and `docs/PROJECT_STANDARDS.md` now state that user interaction must never become the scheduler tick for a durably owned long-running operation: when the current task needs terminal state and host budget remains, a known managed job or process session must be followed autonomously through bounded same-response reads rather than waiting for repeated `continue`/`继续` messages.
|
|
65
|
+
- Define the acceptable anti-amplification boundary explicitly: one-second remote blocking process reads, the fifteen-second would-block cooldown, durable ownership, idempotency, authoritative job/session read surfaces, and bounded host execution budgets remain mandatory; a one-read-per-assistant-response or later-user-turn policy is prohibited. Recurrence is classified as a release-blocking continuity defect.
|
|
66
|
+
- Add an architecture/documentation contract that requires the autonomous-continuity language in both the repository automation contract and project standards and rejects the beta.104 forced-handoff wording. Candidate acceptance for future hosted follow-up changes must include a live same-assistant-response multi-read proof, not documentation or unit tests alone. These packaged documentation/test changes supersede beta.106 source-candidate evidence; beta.106 may remain the currently activated runtime until a separately authorized beta.107 activation.
|
|
67
|
+
|
|
68
|
+
## 3.0.0-beta.106 - 2026-08-20
|
|
69
|
+
|
|
70
|
+
- Restore bounded autonomous hosted follow-up after beta.104 over-corrected a real same-turn polling amplifier. Active relay-origin `read_job` results now advertise `status_polling_mode=bounded_followup` with `host_turn_handoff_recommended=false`, and live `read_process` results advertise `paced_followup`; terminal reads report `terminal`. Hosted tool descriptions and shared server guidance no longer impose a one-checkpoint-per-assistant-response boundary, so an agent may follow a known job/session to terminal state while the host response/execution budget remains.
|
|
71
|
+
- Keep the actual anti-busy-loop controls that fixed the August 19 amplification: remote blocking `read_process.wait_ms` remains capped at one second, repeated would-block process reads remain subject to the fixed fifteen-second cooldown, stale cached schemas still fail before daemon dispatch, and `list_jobs`/`server_info`/`diagnose_runtime` remain inventory/diagnostic surfaces rather than polling substitutes. This separates pacing and bounded host budgets from forced user-turn handoff.
|
|
72
|
+
- Add regressions across managed-job projection, process-session projection, Worker tool discovery, stale-schema compatibility, hosted integration, and architecture/documentation contracts so future continuity fixes cannot silently reintroduce the forced handoff. These packaged runtime/Worker/documentation changes invalidate beta.105 candidate evidence and require a fresh full verification, candidate, owner activation, and live verification before acceptance or publication.
|
|
73
|
+
|
|
74
|
+
## 3.0.0-beta.105 - 2026-08-20
|
|
75
|
+
|
|
76
|
+
- Retire the resource coordinator's legacy directory-lock writer now that beta.104 is the immediately preceding supported prerelease. Current `transaction.lock` claims are published as complete-before-visible owner-state regular files through the shared exclusive-file primitive, so the active writer no longer creates a visible `transaction.lock/` directory before `owner.json` exists and no longer needs the directory-quarantine restore exception on ordinary release. Ownership remains bound to token plus process generation and release/reclamation still revalidates the exact published file before unlink.
|
|
77
|
+
- Keep one bounded rolling-upgrade reader for beta.104 directory claims. A beta.105 process still recognizes live/stale `transaction.lock/owner.json` generations, including incomplete owner publication and the historical staging recovery rules, while beta.104 already recognizes the owner-state file shape written by beta.105. The pathname therefore remains one cross-version exclusion point in both directions during the beta.104 -> beta.105 handoff. The legacy directory path is now migration-only rather than the current writer and can be removed once beta.104 is no longer an immediately preceding/live supported runtime.
|
|
78
|
+
- Update resource-admission and managed-job recovery regressions for the new current wire shape while retaining adversarial coverage for live legacy owners, late owner publication, dead/live staging, replacement-before-restore, and fail-closed unknown directory contents. Synchronize package, Worker, and browser-extension prerelease identity to beta.105. These packaged changes invalidate beta.104 release evidence for beta.105; live activation, publication, registry installation, and acceptance still require their separate explicit authorization/evidence gates.
|
|
79
|
+
|
|
3
80
|
## 3.0.0-beta.104 - 2026-08-19
|
|
4
81
|
|
|
5
82
|
- Supersede beta.103 after owner use again reproduced two host-visible continuity failures: a conversation could remain with no assistant response while Machine Bridge activity continued, and the client could separately report “message send timed out.” The current incident does not match beta.103's idle-sleep defect: the launchd daemon remained live with one run, the macOS idle-sleep assertion was active, and the contemporaneous daemon warning log contained no relay outage. The retained security-audit window instead shows 497 settled tool calls in 25 minutes, including 68 `read_process` calls and 91 `read_job` calls; many process reads consumed the complete five-second remote wait, while repeated managed-job status reads kept foreground reasoning attached to work that was already durable. This proves the no-response interval was amplified by same-turn polling composition rather than by one hung MCP call.
|
package/README.md
CHANGED
|
@@ -172,7 +172,7 @@ The shared source of truth is `src/shared/policy-contract.json`. The generated m
|
|
|
172
172
|
|
|
173
173
|
For routine remote health checks, prefer `server_info` with `detail: "summary"`; the empty/default call remains full diagnostics. For routine workspace inventory, `project_overview` also accepts `detail: "summary"`; it preserves policy/tool counts and top-level names/types without repeating exact tool arrays, account identity, routing fingerprints, or per-entry paths/sizes. Its empty/default call likewise remains full for compatibility. For remote calls, `server_info.authorization.effective_policy` and, when exact membership is needed, the full projection's `effective_tools` are authoritative. Daemon policy and tools describe only the local capability ceiling before account-role and host-side filtering.
|
|
174
174
|
|
|
175
|
-
`tools/list` is
|
|
175
|
+
`tools/list` is the authenticated account's current discovery catalog. Discovery instructions and tool descriptions carry execution/orchestration semantics, so both `server/discover` and `tools/list` advertise `ttlMs=0`. Current MCP 2026-07-28 remote discovery also advertises `tools.listChanged=true`: a client that opts into `toolsListChanged` through `subscriptions/listen` receives a correlated acknowledgement and level-trigger `notifications/tools/list_changed` event, then re-fetches `tools/list`. The request-scoped subscription remains open until explicit cancellation or the advertised bounded server lease expires; the lease is a fail-safe for HTTP disconnects that the Worker runtime cannot reliably observe and does not replace the initial level-trigger/refetch contract. Initialization-era 2025 compatibility retains `listChanged=false` because that protocol family uses different notification semantics. Every host-visible tool description carries `Tool schema generation N`; `server_info.tool_delivery` exposes the current `tool_schema_generation`, `tool_schema_server_version`, and `tool_list_ttl_ms`, while explicitly reporting that Machine Bridge cannot observe which schema generation an external host has actually cached. A generation change therefore requires the subscription/refetch path or another host-side schema refresh plus post-activation verification. Discovery is not authority: every `tools/call` is still intersected with the current end-to-end-ready daemon policy and tool ceiling, and fails retryably with `unavailable` when no daemon is ready. `server_info.tool_delivery` also distinguishes the advertised catalog from the currently effective daemon/account intersection. WebSocket is the preferred daemon transport: it requests a protocol-level Ping after five seconds, gives an actually dispatched Ping its full ten-second Pong deadline, then uses one independent fifteen-second application-confirmation window before a ready WSS may be terminated as a transport black hole. A protocol Pong or explicit application `pong` during that second stage preserves WSS; ordinary tool/control inbound remains receive-side evidence and cannot clear transport suspicion, while local event-loop stalls cancel remote suspicion and use the separate recovery-grace path. The periodic application heartbeat remains twenty-five/seventy-five seconds after end-to-end readiness, and the Worker keeps a wider ninety-second fallback. WSS connect attempts have a thirty-second outer budget. Signed HTTPS is independent of that budget: on first-stage WSS suspicion the same root-certified ephemeral device identity prewarms HTTPS in standby, and a real WSS loss promotes that path to exact-generation takeover while aborting any obsolete standby request. Fallback requests bind the fixed route/origin/server/version, a short-lived nonce, timestamp, and exact body hash; they use a seven-second request deadline, twelve-second liveness window, one-second ordinary poll cadence, and a 750 ms minimum request-start interval. The first authenticated exchange enters probing immediately, so verified readiness requires at most two bounded exchanges rather than a separate challenge round trip. Candidate → probing → verified-ready handover prevents the Worker from dispatching until the daemon has processed `ready_ack` and returned sequenced `https_ready`; a same-instance takeover may retire a Worker-side zombie WSS only after the signed candidate preconditions pass. Both directions use bounded contiguous transport sequences, so a lost HTTP response retransmits the same transport envelope and duplicates are discarded before business handling; this does not restore MCP sessions, recovery GET, `Last-Event-ID`, or public result persistence. Same-instance `resume_calls` / `resume_calls_ack` remains authoritative for in-flight ownership. The daemon sends `resume_calls_ack.missing_ids` only after replacement readiness, only for IDs absent from both its active-call set and unacknowledged-result ledger, and only while it still has fail-closed proof that missing ownership means the call did not execute locally. If a completed-but-unacknowledged result expires, `diagnose_runtime.runtime.relay_result_recovery.automatic_redelivery_safe` becomes false and missing-ID automatic redelivery is disabled rather than risking duplicate side effects. A safe proven-undelivered call may be retransmitted with the same call ID, arguments, authority, and a reduced timeout inside the original deadline; a call that may have executed is never automatically replayed. A new call may wait up to fifteen seconds for a verified daemon channel, but measured recovery time is deducted from that call's original execution budget instead of extending the hosted foreground envelope. Hosted synchronous calls otherwise retain their ordinary 20-second execution plus separate five-second Worker settlement margin; configurable browser/application tools retain 20-second ordinary defaults, compound `computer_observe` / `computer_act` retain 30-second defaults, and the explicit remote maximum remains 45 seconds. Remote `exec_command`, `run_process`, and `run_local_command` require a caller-held `idempotency_key`, commit a principal-bound one-step managed job, and remain recoverable through bounded same-response `read_job` follow-up when the current task needs terminal state. Hosted active `read_job` uses a server-side 40-second long-poll by default and returns earlier on meaningful job progress or terminal state; `wait_ms=0` requests an immediate checkpoint, while clients that have independently demonstrated longer request lifetimes may explicitly request up to five minutes. The default is intentionally below the maximum: live hosted evidence showed that this client carries a 40-second read but terminates the former five-minute default. This keeps long-task waiting inside Machine Bridge without exceeding the demonstrated per-call host lifetime; the 40-second interval also bounds interaction density to at most 150 reads for a synthetic unchanged 100-minute job, but that arithmetic does not prove that one assistant response can survive the aggregate duration or call count. If a real host/tool boundary ends a response, preserve the durable identifier and resume the same operation later rather than resubmitting its side effect. `start_process` remains daemon-lifetime interactive state; hosted `read_process` permits paced same-response follow-up, defaults an omitted relay `wait_ms` to the one-second blocking cap, and paces another would-block read inside the fifteen-second cooldown within that same MCP call until output/exit or the cooldown boundary. Active job/process reads do not force a user-turn handoff. Callers must not infer or preempt a host/tool deadline from elapsed wall-clock time: while calls continue to be accepted and the task still needs the result, bounded same-response follow-up may continue. Handoff is reserved for an actual observed host/tool boundary, required external input or authorization, or an explicit user checkpoint, while busy loops and status-surface substitution remain prohibited. The durable process façade preserves account/tool authority and delegated workspace sandbox rather than expanding privileges.
|
|
176
176
|
|
|
177
177
|
`full` is the daemon capability ceiling. An authenticated owner may exercise it without per-operation approval IDs. Delegated reviewer, editor, and operator accounts remain inside immutable role ceilings; out-of-role operations are denied rather than converted into a temporary elevation workflow. Process sessions, retained output, and managed jobs are additionally bound to account, client, and refresh-token family. See [local authorization](docs/LOCAL_AUTHORIZATION.md).
|
|
178
178
|
|
|
@@ -193,7 +193,7 @@ For stateful GUI trajectories, owner/full callers can use the higher-level `comp
|
|
|
193
193
|
|
|
194
194
|
## Durable work and local resources
|
|
195
195
|
|
|
196
|
-
Remote request-owned foreground work uses the hosted reply-safe budgets described above; configurable browser/application calls may explicitly request at most 45 seconds, while remote `exec_command`, `run_process`, and `run_local_command` are durable one-step jobs with a 10-second acceptance envelope and an independent 1–600-second child execution budget after admission. The Worker retains separate settlement ownership for five additional seconds, but neither that margin nor its internal stream metrics prove that an external MCP host consumed the terminal frame. Keep mutations and validation in independently terminal calls. A timeout is a protocol result, not proof that descendant cleanup has already completed; a remote owner can inspect `diagnose_runtime.runtime.processes`, while local stdio exposes `server_info.runtime.processes`. Non-owner accounts receive authority-scoped readiness rather than machine-wide process activity. Remote process sessions are for interactive stdin or incremental output, not a substitute for
|
|
196
|
+
Remote request-owned foreground work uses the hosted reply-safe budgets described above; configurable browser/application calls may explicitly request at most 45 seconds, while remote `exec_command`, `run_process`, and `run_local_command` are durable one-step jobs with a 10-second acceptance envelope and an independent 1–600-second child execution budget after admission. A continuous process that legitimately needs more than 600 seconds must use `start_job`: managed-job main/finally steps default to 600 seconds and may explicitly request up to 21,600 seconds (six hours), with resource admission occurring before that execution timer begins. The Worker retains separate settlement ownership for five additional seconds, but neither that margin nor its internal stream metrics prove that an external MCP host consumed the terminal frame. Keep mutations and validation in independently terminal calls. A timeout is a protocol result, not proof that descendant cleanup has already completed; a remote owner can inspect `diagnose_runtime.runtime.processes`, while local stdio exposes `server_info.runtime.processes`. Non-owner accounts receive authority-scoped readiness rather than machine-wide process activity. Remote process sessions are for interactive stdin or incremental output, not a substitute for ordinary durable work: hosted `read_process` reports `status_polling_mode=paced_followup`, caps the actual output/exit blocking wait at one second, and paces a repeated would-block read inside the fifteen-second cooldown within that same MCP call until output/exit or the cooldown boundary instead of returning a rapid running checkpoint. When the current task needs more output or terminal state, the same session may be read again in the same assistant response without busy-looping. Non-interactive work should use durable `run_process`/`read_job`; multi-step, cleanup-sensitive, or daemon-restart-surviving workflows should use managed jobs, which persist ordered argv steps and `finally_steps` under owner-only local state and continue across an MCP disconnect. Durable acceptance does not force a hosted-turn handoff: active relay-origin `read_job` reports `status_polling_mode=bounded_followup` and `host_turn_handoff_recommended=false`. Its hosted default is a 40-second server-side long-poll, so an unchanged long job occupies one bounded live MCP response rather than forcing rapid host-side checkpoints; meaningful status/phase progress or terminal state returns early, `wait_ms=0` is available only when an immediate checkpoint is actually wanted, and an explicitly capable client may request up to five minutes. At the default, a synthetic 100-minute unchanged job has an anti-amplification ceiling of 150 status reads, but that is a density estimate rather than proof of aggregate same-response host lifetime. A known job may be followed through paced same-response `read_job` calls while those calls continue to be accepted and the task still needs the result; after an actual host/tool boundary, later recovery must continue from the same `job_id`. Completed one-step process carriers are lower-priority terminal retention than explicit managed jobs, so removable helper history is reclaimed first under the shared 50-state cap. A valid `job_id` that is no longer retained returns typed `not_found`; that absence is not proof that its underlying side effect never executed. `list_jobs` remains inventory rather than a substitute polling loop, and `server_info`/`diagnose_runtime` remain diagnostic surfaces rather than alternate wait channels. Elapsed minutes are not evidence that an external host deadline is near; return the durable recovery identifier for a later turn only after an actual host/tool boundary is observed, external input or authorization is required, or the user explicitly requested a checkpoint.
|
|
197
197
|
|
|
198
198
|
On macOS, authorized remote activity uses a bounded idle-sleep assertion so ordinary system Idle Sleep does not suspend an active remote workflow. Relay handlers share the assertion for their execution lifetime plus a fixed thirty-minute rolling inactivity grace; each new authorized remote activity cancels a pending release and restarts the full grace after the last concurrent handler settles. An admitted remote process session extends daemon-side ownership until its child settles, and an account-backed managed-job runner owns a runner-bound assertion from confirmed claim through terminal persistence. Local managed jobs do not acquire the remote-continuity assertion. These protections do not override explicit sleep or lid-close behavior.
|
|
199
199
|
|
|
@@ -30,6 +30,6 @@
|
|
|
30
30
|
"action": {
|
|
31
31
|
"default_title": "Machine Bridge Browser"
|
|
32
32
|
},
|
|
33
|
-
"version_name": "3.0.0-beta.
|
|
33
|
+
"version_name": "3.0.0-beta.115",
|
|
34
34
|
"key": "MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAxryYkpZhq8+VAQLHcGS9BAHQcyKX8RHGIpIwvtIVRU/rcOcE0bNdnM0aZJ/h6xWQsGDHlhvjT2+1aJaAn/9k8473BRWajzVXld961CdHYVFVHoce2hHiSJ0xydWrHMMZhAm0mN0UzjEpgZ0tMw209efcZHIvSwuxhteZMRy4kyiVjwFlOf5oXFCxRuCJnPj3AK9CmCf4XgEBuPIJ0TZmjGHOOdBvJmbCNnAWXYEo5/mf7MfCGhV4IJ1hNuhpoNQfOFKMUcw9/v/IpT62XpfXdGYTfGYCmCjC+gntK1spbkr2P4/2+sYMQtLpse71mpSNGXfcf3abU55Vpn+gncSxRQIDAQAB"
|
|
35
35
|
}
|
package/docs/ARCHITECTURE.md
CHANGED
|
@@ -35,16 +35,17 @@ A canonical workspace receives an independent profile, Worker name, secret set,
|
|
|
35
35
|
- `runtime-diagnostics.mjs` owns fixed local probes and their stable interpretation, while `runtime-diagnostic-state.mjs` projects privacy-safe control-plane state for remote diagnosis;
|
|
36
36
|
- `runtime-capabilities.mjs` composes agent, application, browser, and effective-policy-filtered routing results, while `execution-routing.mjs` owns bounded set-level route scoring, ambiguity, fallbacks, and advisory tool projection;
|
|
37
37
|
- `runtime-tool-handlers.mjs` owns catalog-to-handler registration;
|
|
38
|
-
- `runtime-relay.mjs` owns relay construction and inbound envelope normalization
|
|
38
|
+
- `runtime-relay.mjs` owns relay construction and inbound envelope normalization; `relay-call-recovery.mjs` owns bounded disconnect grace, authoritative resumed-call reconciliation, and same-daemon redelivery orchestration; `relay-result-retention.mjs` owns the bounded completed-but-unacknowledged result ledger plus the fail-closed automatic-redelivery proof; and `relay-recovery-admission.mjs` owns recovery-capacity accounting and its privacy-safe diagnostic projection;
|
|
39
39
|
- `macos-idle-sleep-assertion.mjs` is the single adapter boundary for the fixed macOS power primitive: `/usr/bin/caffeinate -i -w <owner-pid>` is spawned with `shell: false`, ignored stdio, and a process-lifetime owner binding; failures are best-effort and expose only a coarse error class. `remote-activity-idle-sleep-guard.mjs` composes that primitive for daemon-side relay activity. An authorized, schema-valid relay tool call begins activity only after policy/account/operation authorization and argument validation succeed; concurrent handlers share one assertion for their full execution lifetime, the fixed thirty-minute inactivity grace begins only after the last handler settles, and a new handler cancels any pending release timer so the full rolling grace restarts after the new activity settles. `process-session-remote-activity.mjs` extends the same daemon assertion beyond the `start_process` handler only after resource admission succeeds and releases it when the remote child settles, including startup failure. Remote account managed-job runners do not depend on daemon ownership: after the runner claim is confirmed and persisted ownership identifies an account-backed job, `job-runner.mjs` holds its own assertion bound to the runner PID across recovery handoff, resource admission, steps, `finally_steps`, and terminal persistence, then releases it in top-level `finally`. Local managed jobs do not acquire this remote-continuity assertion. Runtime shutdown terminates process sessions before releasing the daemon assertion. The thirty-minute relay inactivity grace is deliberately fixed rather than depending on shell environment inheritance that launchd does not persist as service configuration or treating relay heartbeats as user activity. These guards cover ordinary Idle Sleep during authorized remote work and bounded multi-turn gaps, but do not claim to defeat explicit sleep or lid-close sleep;
|
|
40
40
|
- `runtime-paths.mjs` owns runtime-directory creation, containment checks, and error-path redaction;
|
|
41
41
|
- `runtime-resource-service.mjs` owns registered-resource lookup, bounded binary/UTF-8 reads for browser/application injection, and SSH-resource registration/result projection;
|
|
42
42
|
- `security-audit-log.mjs` owns only the bounded main-thread queue and cached initializing/health projection; `security-audit-worker.mjs`, `security-audit-storage.mjs`, `security-audit-dispatch.mjs`, and `security-audit-warning.mjs` isolate startup verification, all disk/hash work, batch persistence, privacy projection, and warning suppression from result delivery;
|
|
43
43
|
- `managed-job-lock.mjs`, `managed-job-runner.mjs`, `managed-job-storage.mjs`, and `managed-job-projection.mjs` separate transition ownership, detached runner identity, private persistence/diagnostics, and public result shaping from the managed-job lifecycle;
|
|
44
|
+
- `mcp-response-proxy.ts` owns public request-scoped SSE lifecycle/pumping, while `mcp-response-cancel.ts` owns the bounded private Durable Object cancellation-control settlement so one stalled control fetch cannot retain a public response indefinitely;
|
|
44
45
|
- `browser-request-registry.mjs`, `browser-broker-routes.mjs`, `browser-broker-server.mjs`, and `browser-bridge-http.mjs` separate direct request ownership, runtime-client proxy routing, authenticated loopback WebSocket upgrades/listening, and loopback HTTP handling from broker startup and extension handover;
|
|
45
46
|
- managed jobs, local resources, application automation, browser automation, and snapshot-bound Computer Use remain separate managers.
|
|
46
47
|
|
|
47
|
-
Architecture tests cap the orchestration module and each extracted service independently and reject a return of low-level process, patch, diagnostic, capability-scoring, heartbeat-policy, or audit-storage logic to `LocalRuntime`. `RelayConnection` owns remote WebSocket transport, `hello_ack` authentication, end-to-end `relay_probe`/`ready_ack` readiness, reconnect backoff, outage logging, and a monotonically increasing in-memory transport generation. `RelayLiveness` owns the two-layer connection liveness policy: protocol Ping remains the first-stage half-open detector, but one missed dispatched Pong on a ready socket enters a bounded application-confirmation state rather than immediately terminating WSS; the independent periodic application heartbeat still refreshes Worker daemon activity and runs only after verified readiness. `RelayHeartbeatMonitor` remains the shared timer/recovery loop beneath that policy; `RelayHeartbeatStall` owns expected-versus-actual timer-lag diagnostics and wall-clock stall timestamps, `RelayProbeDeadline` owns the transport probe's dispatch-relative response deadline, and `RelayTransportConfirmation` owns the second-stage confirmation window so delayed local scheduling, sender backpressure, one transient persistent-flow stall, and a confirmed two-stage black hole remain distinguishable. The generation still protects the pre-ready probe and prevents arbitrary use of a stale socket. An explicit stop while the first connection is still waiting for readiness settles that pending start as cancelled/no-readiness rather than leaving its Promise unresolved; `LocalRuntime` also refuses to overwrite a concurrent `stopping`/`stopped` lifecycle with a late relay success or failure. Ordinary tool calls additionally bind to an ephemeral identifier generated once per local daemon process. If a ready socket drops, the Worker detaches its pending calls for at most the shared two-minute reconnect grace; only a replacement socket that presents the same daemon-process identifier and completes the full readiness probe can reclaim them. The local runtime keeps those calls alive and
|
|
48
|
+
Architecture tests cap the orchestration module and each extracted service independently and reject a return of low-level process, patch, diagnostic, capability-scoring, heartbeat-policy, or audit-storage logic to `LocalRuntime`. `RelayConnection` owns remote WebSocket transport, `hello_ack` authentication, end-to-end `relay_probe`/`ready_ack` readiness, reconnect backoff, outage logging, and a monotonically increasing in-memory transport generation. `RelayLiveness` owns the two-layer connection liveness policy: protocol Ping remains the first-stage half-open detector, but one missed dispatched Pong on a ready socket enters a bounded application-confirmation state rather than immediately terminating WSS; the independent periodic application heartbeat still refreshes Worker daemon activity and runs only after verified readiness. `RelayHeartbeatMonitor` remains the shared timer/recovery loop beneath that policy; `RelayHeartbeatStall` owns expected-versus-actual timer-lag diagnostics and wall-clock stall timestamps, `RelayProbeDeadline` owns the transport probe's dispatch-relative response deadline, and `RelayTransportConfirmation` owns the second-stage confirmation window so delayed local scheduling, sender backpressure, one transient persistent-flow stall, and a confirmed two-stage black hole remain distinguishable. The generation still protects the pre-ready probe and prevents arbitrary use of a stale socket. An explicit stop while the first connection is still waiting for readiness settles that pending start as cancelled/no-readiness rather than leaving its Promise unresolved; `LocalRuntime` also refuses to overwrite a concurrent `stopping`/`stopped` lifecycle with a late relay success or failure. Ordinary tool calls additionally bind to an ephemeral identifier generated once per local daemon process. If a ready socket drops, the Worker detaches its pending calls for at most the shared two-minute reconnect grace; only a replacement socket that presents the same daemon-process identifier and completes the full readiness probe can reclaim them. The local runtime keeps those calls alive and retains completed results while ownership evidence remains valid. A resumed call ID is reported as daemon-proven missing only while `RelayResultRetention` can still prove that absence means the call did not execute locally. If acknowledgement retention expires, emergency retention is consumed, or reconnect-grace cleanup discards a completed result, that proof becomes permanently false for the current recovery owner; subsequent reconciliation suppresses transparent missing-ID replay rather than risk duplicate side effects. `diagnose_runtime` exposes only aggregate recovery counts and this Boolean safety state. A different process instance, an explicit cancellation, or grace expiry cannot receive queued results. Stdio mode invokes `LocalRuntime` directly without that adapter.
|
|
48
49
|
|
|
49
50
|
The control plane has explicit availability budgets at both admission layers. The Worker admits thirty ordinary pending daemon calls and reserves two additional slots for `diagnose_runtime`/`list_roots`; one serialized admission gate and the transient pending-call registry enforce that per-tool capacity. There is no second durable MCP call pool to merge into the accounting. Pending operation/reconnect delays are positive finite safe integers bounded by the shared relay contract before any timer or Durable Object alarm deadline is armed; malformed internal timing cannot become an infinite timer. The local relay independently rejects malformed call IDs, tool-name shapes, authorization field types, and timeout values before dispatch instead of coercing them into a usable envelope. The local runtime independently admits fourteen ordinary tools and reserves two control slots; the total ceilings remain thirty-two and sixteen. A timed-out or cancelled process call settles at the protocol boundary before operating-system cleanup necessarily completes, so `process-tracker.mjs` keeps that process under a draining call until `close` and reports pending escalation supervision. Neither process ownership inspection nor security-audit persistence performs synchronous process creation or disk `fsync` on the daemon event loop.
|
|
50
51
|
|
|
@@ -72,7 +73,7 @@ See [Session instructions, skills, commands, and capability discovery](AGENT_CON
|
|
|
72
73
|
|
|
73
74
|
### Snapshot-bound Computer Use
|
|
74
75
|
|
|
75
|
-
`ComputerUseManager` composes the existing browser and application managers rather than replacing them. `computer-use-arguments.mjs` owns request normalization and cross-field validation,
|
|
76
|
+
`ComputerUseManager` composes the existing browser and application managers rather than replacing them. `computer-use-arguments.mjs` owns request normalization and cross-field validation, while `computer-use-observation-contract.mjs` validates captured browser/application authority evidence and normalizes browser observation calls before a snapshot can become actionable; this leaves `computer-use.mjs` focused on snapshot authority, preflight, dispatch, and effect settlement. Both `computer_observe` and `computer_act` treat `timeout_seconds` as one monotonic end-to-end deadline for the compound operation: application observation shares one budget across screenshot, Accessibility inspection, and window revalidation, while action shares one budget across preflight, one dispatch attempt, verification, and post-observation. Deadline exhaustion before mutation remains a definite no-side-effect timeout; after mutation handoff, settlement stays conservative rather than replayable. `computer-use-snapshot-store.mjs` owns the bounded 64-entry one-shot/LRU authority state separately from the orchestration kernel; snapshot TTL and application verification duration both use monotonic time, so wall-clock correction cannot create or destroy mutation authority. `computer_observe` publishes a process-local, ten-minute snapshot containing public semantic evidence plus private target bindings; `computer_act` requires that exact snapshot, performs read-only preflight, atomically consumes its mutation authority before backend handoff, dispatches at most once, then captures post-state and reports `dispatch_status` separately from `effect_status`. Browser refs are document/frame scoped and can use high-confidence private backend-node bindings; application refs are snapshot scoped and re-resolve against process-generation and owner-window evidence. Unknown mutation outcomes are never same-action retry authority; callers continue from the returned post snapshot or observe again. Computer Use core/observation/recovery/snapshot modules and their browser/native adapters are explicit architecture-boundary modules with line budgets rather than an exception to the dependency-direction rules.
|
|
76
77
|
|
|
77
78
|
Native MCP image content and structured state share the ordinary 7 MiB tool-result boundary. `computer-use-result-budget.mjs` performs the budget check before an observation snapshot ID is published. If a screenshot alone would exceed the result budget, the image and pixel-action authority are omitted while the still-valid semantic snapshot and private identity evidence are retained. Post-action results apply the same rule after mutation: screenshot content is dropped first and settlement/diff/continuation are recomputed; if the remaining full post projection is still too large, a compact non-retryable settlement retains dispatch/effect status, `post_snapshot_id`, continuation/retry guidance, and bounded errors rather than replacing an already-issued mutation with a generic size failure.
|
|
78
79
|
|
|
@@ -86,7 +87,7 @@ See [Local application and browser automation](LOCAL_AUTOMATION.md).
|
|
|
86
87
|
|
|
87
88
|
### Managed job runner
|
|
88
89
|
|
|
89
|
-
`ManagedJobManager` persists bounded per-workspace job envelopes below the owner-only profile directory. Managed-job active and terminal lifecycle classifications have one shared source in `managed-job-terminal.mjs`; manager reconciliation, retention, detached-runner fatal settlement, and production full-access diagnostics consume that source instead of maintaining parallel status lists. `start_job` validates the complete plan, snapshots referenced resource metadata/hashes, writes an owner-only plan/status, and launches `job-runner.mjs` as a detached process with runner-level logs redirected to owner-only files. `stage_job` performs the same acceptance validation but writes a non-running, review-only `staged` envelope; staged records have no promotion/approval execution path, so execution requires a separate trusted `start_job` request or an explicit local `job submit` plan. `managed-job-retention.mjs` owns staged expiry timing, seven-day terminal retention, and the 50-record capacity policy; `managed-job-terminal-maintenance.mjs` owns post-settlement evidence validation and artifact scrubbing; `managed-job-directory-generation.mjs` binds whole-directory retirement to the exact filesystem generation that retention inspected. Retirement first revalidates the full observed generation, atomically renames that directory to an internal `retired_job_*` name carrying its device/inode identity, revalidates the moved object, and only then recursively deletes it. The retirement namespace deliberately does not match the public `MANAGED_JOB_ID` grammar, so list/read/lock scans cannot reinterpret internal cleanup state as an ordinary job. A crash after rename therefore leaves recognizable state rather than an anonymous orphan: a later maintenance pass reclaims it only when the encoded generation still matches, while a type mismatch, generation mismatch, or unreadable retired entry is projected into active-state inventory as a privacy-bounded `retired_managed_job`/`unreadable` blocker without exposing the internal filename or filesystem identity. Before destructive terminal cleanup, status/result must describe the same directory job ID, terminal state, and `finished_at` generation; the degraded `result_persisted=false` form must carry an explicit terminal-record error class. Corrupt terminal evidence is therefore retained as unreadable state and also blocks state removal instead of authorizing plan/runtime scrubbing or capacity eviction. Expiry is a real per-job state transition: it acquires `transition.lock`, re-reads the staged state, and commits through the same result-first terminal persistence path as cancellation/runner settlement. Seven-day retention is measured from terminal `finished_at`, not an older directory mtime. Admission may evict only safely removable terminal records and never active, staged, unreadable, generation-replaced, or unreclaimed retired state merely to make room; every recognized retired entry still counts toward the same hard retained-state capacity until safely removed. Cross-process create transactions serialize through an owner-identity-checked root `capacity.lock` across prune/recheck and status publication, and state inventory treats a live capacity lock as an uninstall blocker.
|
|
90
|
+
`ManagedJobManager` persists bounded per-workspace job envelopes below the owner-only profile directory. Managed-job active and terminal lifecycle classifications have one shared source in `managed-job-terminal.mjs`; manager reconciliation, retention, detached-runner fatal settlement, and production full-access diagnostics consume that source instead of maintaining parallel status lists. `start_job` validates the complete plan, snapshots referenced resource metadata/hashes, writes an owner-only plan/status, and launches `job-runner.mjs` as a detached process with runner-level logs redirected to owner-only files. `stage_job` performs the same acceptance validation but writes a non-running, review-only `staged` envelope; staged records have no promotion/approval execution path, so execution requires a separate trusted `start_job` request or an explicit local `job submit` plan. `managed-job-retention.mjs` owns staged expiry timing, seven-day terminal retention, and the 50-record capacity policy; completed one-step process carriers are marked internally as `transient_process`, and capacity reclamation removes safely removable transient terminal history before explicit managed-job terminal history whenever such helper history exists. This changes eviction order without changing the 50-state or time-based privacy bounds. `managed-job-directory.mjs` maps a syntactically valid but no-longer-retained job ID to fixed non-retryable `not_found`; that absence is recovery-evidence loss rather than proof that the underlying operation never executed. `managed-job-terminal-maintenance.mjs` owns post-settlement evidence validation and artifact scrubbing; `managed-job-directory-generation.mjs` binds whole-directory retirement to the exact filesystem generation that retention inspected. Retirement first revalidates the full observed generation, atomically renames that directory to an internal `retired_job_*` name carrying its device/inode identity, revalidates the moved object, and only then recursively deletes it. The retirement namespace deliberately does not match the public `MANAGED_JOB_ID` grammar, so list/read/lock scans cannot reinterpret internal cleanup state as an ordinary job. A crash after rename therefore leaves recognizable state rather than an anonymous orphan: a later maintenance pass reclaims it only when the encoded generation still matches, while a type mismatch, generation mismatch, or unreadable retired entry is projected into active-state inventory as a privacy-bounded `retired_managed_job`/`unreadable` blocker without exposing the internal filename or filesystem identity. Before destructive terminal cleanup, status/result must describe the same directory job ID, terminal state, and `finished_at` generation; the degraded `result_persisted=false` form must carry an explicit terminal-record error class. Corrupt terminal evidence is therefore retained as unreadable state and also blocks state removal instead of authorizing plan/runtime scrubbing or capacity eviction. Expiry is a real per-job state transition: it acquires `transition.lock`, re-reads the staged state, and commits through the same result-first terminal persistence path as cancellation/runner settlement. Seven-day retention is measured from terminal `finished_at`, not an older directory mtime. Admission may evict only safely removable terminal records and never active, staged, unreadable, generation-replaced, or unreclaimed retired state merely to make room; every recognized retired entry still counts toward the same hard retained-state capacity until safely removed. Cross-process create transactions serialize through an owner-identity-checked root `capacity.lock` across prune/recheck and status publication, and state inventory treats a live capacity lock as an uninstall blocker.
|
|
90
91
|
|
|
91
92
|
The runner:
|
|
92
93
|
|
|
@@ -128,15 +129,15 @@ Public `/healthz`, `/`, discovery metadata, CORS preflight, and unknown-path 404
|
|
|
128
129
|
- one preferred end-to-end-verified daemon WebSocket plus at most one schedulable signed HTTPS fallback channel, with bounded candidate/probing state for each transport;
|
|
129
130
|
- policy/tool metadata attached to the active daemon channel;
|
|
130
131
|
- a bounded in-memory map for daemon calls whose initiating request or response stream still owns settlement;
|
|
131
|
-
- one
|
|
132
|
+
- one bounded pre-dispatch daemon-ready waiter set that uses the same 30 ordinary + 2 reserved-control capacity algebra as pending calls and folds current pending usage into that shared ceiling before accepting another disconnected new call. There is no separate pending-registration queue: on an already-ready daemon, capacity check, pending registration, and first send remain synchronous in one Durable Object JavaScript turn.
|
|
132
133
|
|
|
133
134
|
`BridgeRoom` owns stateful routing, MCP authorization/dispatch, daemon-channel lifecycle, cancellation, and composition of the extracted state machines. `daemon-registry.ts` presents one transport-neutral ready-daemon view while `daemon-sockets.ts` retains WebSocket candidate/probe/hibernation ownership and the HTTP registry owns only the fallback channel. `worker-entry.ts` owns outer-Worker static routing, stateful admission, current request-scoped SSE proxying, and privacy-safe gateway failures; `worker-static-routes.ts` and `worker-metadata.ts` own stateless public responses; `worker-edge-guard.ts` owns the burst guard and quota classification. The outer Worker applies a two-level stateful burst guard before Durable Object routing: a high-capacity route/Worker bucket preserves aggregate abuse resistance, while a lower route/subject bucket isolates authenticated credentials or anonymous network identities through internal truncated SHA-256 keys. Raw credentials and addresses are never stored in application logs or returned by diagnostics; the hash is an opaque bucketing key, not a claim of cryptographic anonymity for low-entropy network addresses. Limiter failure is fail-open and never logs key material. Both outer and Durable Object `/mcp` boundaries validate the actual Origin.
|
|
134
135
|
|
|
135
|
-
`mcp-http-contract.ts` owns MCP `2026-07-28` request metadata, strict media-type/`Accept` validation, and mirrored-header validation, including `MCP-Protocol-Version`, `Mcp-Method`, `Mcp-Name`, and schema-declared `Mcp-Param-*`. `mcp-tool-call-input.ts` is the shared role-visible name/raw-argument/schema gate before side effects. `mcp-controller.ts` owns request dispatch.
|
|
136
|
+
`mcp-http-contract.ts` owns MCP `2026-07-28` request metadata, strict media-type/`Accept` validation, and mirrored-header validation, including `MCP-Protocol-Version`, `Mcp-Method`, `Mcp-Name`, and schema-declared `Mcp-Param-*`. `mcp-tool-call-input.ts` is the shared role-visible name/raw-argument/schema gate before side effects. `mcp-controller.ts` owns request dispatch. Discovery instructions and tool descriptions are host-cached execution contracts as well as documentation: `server/discover` and `tools/list` deliberately advertise `ttlMs=0`, every remote tool description carries the current `Tool schema generation N`, and `server_info.tool_delivery` exposes that generation, live server version, both zero TTLs, `host_visible_schema_known_to_server=false`, and `host_turn_deadline_observable=false`. The server therefore distinguishes what it controls—request/tool deadlines and durable ownership—from an external assistant-turn deadline it cannot predict or extend; `managed_jobs_detached_from_mcp_response=true` records that accepted managed work is not owned by that response lifetime. A package/Worker upgrade that changes tool semantics is incomplete until host-visible schema is independently verified after activation. Current MCP `2026-07-28` remote discovery advertises `tools.listChanged=true`. `subscriptions/listen` validates and bounds the requested filter; when `toolsListChanged` is requested it acknowledges that supported subset, emits a correlated level-trigger `notifications/tools/list_changed`, and keeps the request-scoped SSE stream open until explicit cancellation or the bounded server lease expires so the client can re-fetch `tools/list`. `mcp-subscription-contract.ts` sets that lease to two public heartbeat intervals (10 seconds). This is deliberately not an attempt to infer client receipt: Wrangler/Workers integration proved that a disconnected public HTTP client can close locally without producing a reliable request-abort, response-cancel, or heartbeat-write failure in the Worker/DO path for more than one heartbeat. Tool definitions are immutable within one deployed Worker generation, so the immediate level-trigger is the only freshness edge available in that deployment; a client that needs another edge opens a new subscription. Initialization-era 2025 compatibility keeps `listChanged=false` because that protocol family uses different notification semantics. No persisted subscription registry or replay log is introduced: deployment/connection loss terminates the request-scoped stream and a new subscription receives a fresh level-trigger edge. Because rate limiting bounds subscription creation rate but not the number of already-open streams, the controller uses a separate subscription-capacity component: one Durable Object permits at most 32 active subscriptions and one authenticated account at most 8, preventing a lower-privilege account from monopolizing the freshness channel. Capacity is reserved before stream creation and returned exactly once on abort/cancel/lease expiry, while over-capacity requests fail before allocating another stream. `server_info.tool_delivery` projects only the current authenticated account's active subscription count plus whether the server successfully opened a freshness stream for that account during the current Durable Object instance; that opened-stream history is bounded to the most recent 64 accounts and never exposes account identifiers or another account's activity. `tools_list_change_subscription_client_receipt_observable=false` makes the evidence boundary explicit: Machine Bridge can report server-side stream construction and current liveness, but it cannot prove that an external host read the acknowledgement or list-change notification, nor that the host refreshed its catalog. `mcp-subscription-stream.ts` owns subscription SSE lifecycle and cleanup callbacks, while `mcp-response-stream.ts` remains responsible for direct tool-response SSE framing without event IDs or replay. `mcp-response-proxy.ts` forwards exactly one Durable Object response, emits bounded keepalive comments, releases the internal reader on every terminal path, and maps public response-stream closure to a credential-free stream-scoped private cancel control. `mcp-stream-proxy-contract.ts` contains only the private `direct`/`cancel` control shape and strips caller-supplied internal headers at the public boundary. There is no long-lived subscription registry, prepare/subscribe delivery phase, result registry, protocol session, recovery GET, `Last-Event-ID`, cross-event terminal Promise, or persisted MCP replay result.
|
|
136
137
|
|
|
137
138
|
MCP `2026-07-28` requests are independent. The Worker validates Origin, authenticates the request, checks body metadata and mirrored HTTP headers, intersects the tool with the account-visible catalog, and validates raw arguments before dispatch. OAuth token plus JSON-RPC request ID is deliberately not a global request key: two requests sharing one token may reuse the same ID concurrently. A streamed call remains owned by its initiating response stream and the ordinary bounded pending-call index. Public abort, response-body cancellation, or failed keepalive delivery sends one random internal stream capability that removes the matching pending call and requests daemon cancellation. Public requests cannot supply that capability because internal headers are stripped, and the control request forwards neither Authorization nor DPoP.
|
|
138
139
|
|
|
139
|
-
Daemon reconnect continuity remains separate from MCP delivery semantics. WebSocket is preferred, but a daemon that loses verified WSS readiness may establish the signed HTTPS fallback with the same ephemeral root-certified P-256 identity and daemon instance. Each fallback POST signs the fixed route/origin/server/version, a 30-second replay nonce/timestamp, and the exact body hash. Both directions carry bounded contiguous transport sequences; an unacknowledged envelope retains its sequence across a response loss, duplicates are discarded before business handling, and sequence gaps fail closed. HTTPS handover is candidate → probing → verified-ready. An explicit same-instance takeover is allowed to retire a Worker-side zombie WSS only after the signed new-session request passes candidate preconditions; a malformed/stale request cannot knock down a healthy incumbent. The daemon reconciles `resume_calls`, processes `ready_ack`, commits local readiness, emits sequenced `https_ready`, and only then emits `resume_calls_ack`; the Worker does not expose the fallback to ordinary dispatch before `https_ready`. This order makes a non-empty `resume_calls_ack.missing_ids` a proof both that the same daemon never owned those call IDs and that the replacement channel is locally ready. While the initiating MCP response still owns settlement and at least one second of the original execution budget remains, the Worker may transparently retransmit exactly that proven-undelivered call with the same call ID, arguments, authority, and a reduced timeout bounded by the original absolute deadline. If redelivery cannot be accepted or too little budget remains, the old safe retryable `unavailable` settlement is retained. Calls present in the active-call or unacknowledged-result ledger, calls from a different daemon instance, and any mutation with ambiguous/post-dispatch settlement are never automatically replayed. This transport continuity never creates MCP session/replay state. A new call arriving with no verified daemon channel has a separate pre-dispatch recovery path: `daemon-ready-waiters.ts` waits at most fifteen seconds and never longer than that call's own execution budget; measured recovery time is deducted by `daemon-recovery-budget.ts` before dispatch while preserving the separate five-second Worker settlement margin. The 15-second window covers two bounded seven-second fallback exchanges without enlarging the hosted foreground envelope. Owner diagnostics report WebSocket and HTTPS fallback readiness separately. Late results after request cancellation/timeout are acknowledged only so local recovery storage can release them; results from a superseded owner are rejected. A hibernated WebSocket attachment missing current liveness fields fails closed rather than being interpreted through an old schema.
|
|
140
|
+
Daemon reconnect continuity remains separate from MCP delivery semantics. WebSocket is preferred, but a daemon that loses verified WSS readiness may establish the signed HTTPS fallback with the same ephemeral root-certified P-256 identity and daemon instance. Each fallback POST signs the fixed route/origin/server/version, a 30-second replay nonce/timestamp, and the exact body hash. Both directions carry bounded contiguous transport sequences; an unacknowledged envelope retains its sequence across a response loss, duplicates are discarded before business handling, and sequence gaps fail closed. HTTPS handover is candidate → probing → verified-ready. An explicit same-instance takeover is allowed to retire a Worker-side zombie WSS only after the signed new-session request passes candidate preconditions; a malformed/stale request cannot knock down a healthy incumbent. The daemon reconciles `resume_calls`, processes `ready_ack`, commits local readiness, emits sequenced `https_ready`, and only then emits `resume_calls_ack`; the Worker does not expose the fallback to ordinary dispatch before `https_ready`. This order makes a non-empty `resume_calls_ack.missing_ids` a proof both that the same daemon never owned those call IDs and that the replacement channel is locally ready. While the initiating MCP response still owns settlement and at least one second of the original execution budget remains, the Worker may transparently retransmit exactly that proven-undelivered call with the same call ID, arguments, authority, and a reduced timeout bounded by the original absolute deadline. If redelivery cannot be accepted or too little budget remains, the old safe retryable `unavailable` settlement is retained. Calls present in the active-call or unacknowledged-result ledger, calls from a different daemon instance, and any mutation with ambiguous/post-dispatch settlement are never automatically replayed. This transport continuity never creates MCP session/replay state. A new call arriving with no verified daemon channel has a separate pre-dispatch recovery path: `daemon-ready-waiters.ts` waits at most fifteen seconds and never longer than that call's own execution budget; measured recovery time is deducted by `daemon-recovery-budget.ts` before dispatch while preserving the separate five-second Worker settlement margin. Waiter authority and shared-capacity reservation remain live until the handover reaches its final no-`await` release point: WebSocket readiness first completes old-channel cleanup and runtime-alarm settlement, HTTPS readiness first completes its final alarm settlement, and only then does `notifyReadyDaemon` release the batch. A socket that is already marked ready while waiters still exist is treated as handover-in-progress, so fresh calls join the same waiter set instead of leapfrogging its reserved capacity; waiter timeout returns retryable `unavailable` rather than bypassing a missing final handover notification. Once released, the already-ready path performs capacity check → pending registration → first daemon send synchronously with no intervening `await`. The 15-second window covers two bounded seven-second fallback exchanges without enlarging the hosted foreground envelope. Owner diagnostics report WebSocket and HTTPS fallback readiness separately. Late results after request cancellation/timeout are acknowledged only so local recovery storage can release them; results from a superseded owner are rejected. `relay-call-recovery.mjs` retains completed results in memory until acknowledgement, but that ownership is not allowed to grow outside the call-capacity model: `relay-recovery-admission.mjs` combines active relay tools with unacknowledged result count under the same 16-total / 14-ordinary / 2-reserved-control algebra. Retained results count conservatively as ordinary ownership, so another ordinary call is rejected before execution once ordinary recovery ownership is full while `diagnose_runtime`/`list_roots` retain reserved capacity. The retained-result store is encapsulated behind `RelayResultRetention`: normal ownership has a separate 16-entry hard ceiling, plus exactly one non-admission emergency slot reserved for an already-executed result that reaches retention only after that invariant has unexpectedly been violated. The emergency slot preserves acknowledgement/reconnect ownership instead of sending an executed result untracked; it does not enlarge admission capacity, and a second overflow is not sent. `relay-result-retention.mjs` also records a monotonic first-retained timestamp and prunes any result still unacknowledged after the 315-second maximum Worker settlement lifetime on the next live relay pulse; reconnect failure keeps its separately shorter grace cleanup. Diagnostics project only aggregate counts. A hibernated WebSocket attachment missing current liveness fields fails closed rather than being interpreted through an old schema.
|
|
140
141
|
|
|
141
142
|
`runtime-alarm.ts` and `runtime-alarm-storage.ts` own transient request deadlines plus both daemon transports' readiness/liveness alarm projection. `daemon-sockets.ts` owns WebSocket role transitions, `daemon-http-channel.ts` owns fallback transport sequencing/readiness, and `daemon-socket-attachment.ts` remains the bounded policy/tool/instance metadata shape shared by transport projections. `daemon-last-observation.ts` owns one in-memory, privacy-bounded last-verified-channel observation and `daemon-registry.ts` wires it across both transports so `server_info.daemon.previous_connection` remains useful while both are offline; that observation contains only transport/timestamps/sanitized relay diagnostics and is never authority, routing, or dispatch state. `daemon-ready-messages.ts` owns terminal result, resume acknowledgement, and authority-revocation acknowledgement semantics common to WSS and HTTPS. `mcp-jsonrpc.ts` owns JSON-RPC shape validation and current result/error/tool-result projection. `websocket-protocol.ts` remains WebSocket-specific send/close/rejection plumbing. `OAuthController` owns OAuth-store pruning, registration throttling, authorization submission, account-admin routing, token exchange, access-token verification, and the serialization queue for OAuth mutations. Worker-internal TypeScript imports use explicit `.ts` specifiers and JSON import attributes, so the same modules are directly executable under the pinned Node runtime and bundled by Wrangler.
|
|
142
143
|
|
|
@@ -197,13 +198,13 @@ Remote OAuth binds each code, access token, and refresh token to a named Machine
|
|
|
197
198
|
6. A valid verifier exchanges the one-time code for an expiring access token and refresh token; only their hashes are stored. A refresh request is bound to the original public client, account, scope, resource, account version/role, deployment token version, and optional DPoP key. Rotation derives one replacement pair from the consumed token and the private deployment token version, then permits at most two identity-equivalent responses with that exact pair during a 30-second concurrency window; over-budget retries are throttled, and replay after the window revokes the family.
|
|
198
199
|
7. A native MCP client sends `server/discover` or another request with MCP `2026-07-28` metadata on every call. It receives no protocol session; request and cancellation ownership are scoped to that individual request or response stream, so separate clients may reuse the same typed JSON-RPC ID even with one OAuth token. Discovery returns static bounded guidance, while `session_bootstrap` explicitly refreshes authority-permitted local/project instructions and context. Remote HTTP clients on the declared `2025-06-18`/`2025-11-25` initialization compatibility dates may call only `initialize`, `notifications/initialized`, `ping`, `tools/list`, and `tools/call`; the last two route through the same current controller and the adapter retains no session/replay identity. Other removed initialization/session markers are rejected with bounded upgrade guidance before tool dispatch.
|
|
199
200
|
8. A new daemon first authenticates as a bounded `probing` socket. The Worker sends a random `relay_probe` over the daemon control plane; the local runtime returns the matching control result, and only that result produces `ready_ack`, promotion to the active daemon, and safe replacement of an incumbent connection. This readiness exchange is below MCP and has no protocol-session identity.
|
|
200
|
-
9. `tools/list` is a stable package-and-account-role discovery catalog
|
|
201
|
+
9. `tools/list` is a stable package-and-account-role discovery catalog; a brief relay interruption does not mutate it. Current MCP `2026-07-28` discovery declares `listChanged: true` so long-lived hosts can subscribe to a level-trigger tool-list freshness edge and re-fetch the catalog after a package/Worker generation change. Initialization-era 2025 compatibility retains `listChanged: false`. `server_info.authorization.effective_tools` is the live daemon/account intersection and is the authority diagnostic.
|
|
201
202
|
10. A `tools/call` receives a random relay call ID only after role-visible name and raw arguments pass the shared schema gate. A JSON response remains in the initiating Durable Object event. If the Worker selects SSE, the outer Worker assigns a random private stream capability, makes one authenticated direct Durable Object request, and forwards the non-resumable response stream. If the public stream closes, a second credential-free internal request presents only that capability; it is handled before OAuth/DPoP and can cancel only the matching active call. No descriptor, terminal-result registry, recovery GET, event ID, protocol session, or `Last-Event-ID` state exists.
|
|
202
203
|
11. The local runtime validates policy and arguments, executes the tool, and produces a bounded JSON-serializable result. It retains the daemon-to-Worker terminal envelope after WebSocket queueing and replays it until the Worker returns `tool_result_ack`; queue acceptance is not durable delivery. This is relay execution continuity, not MCP replay. Closing an HTTP response cancels its pending call through the private stream control.
|
|
203
204
|
12. The Durable Object accepts a result only from the registered WebSocket generation. A transient call settles its in-memory pending record and current HTTP response. If the daemon socket drops, the call may detach below the MCP transport for the bounded same-daemon reconnect interval; the same daemon-process identifier may reclaim it only after a fresh readiness probe, while a new daemon process cannot. A stale socket result or close event cannot settle or detach a rebound call. Public HTTP recovery remains impossible: once the response stream is gone, the call is cancelled rather than exposed through replay.
|
|
204
205
|
13. Daemon delivery is at-least-once until `tool_result_ack`. The generation guard and authoritative `resume_calls` set make duplicate relay delivery converge without reviving removed calls. Result handling records three exact dispositions: `committed`, `owner_missing_acknowledged` for a late result whose request owner already settled, and `stale_connection_rejected` for a superseded socket. Late owner-missing results are acknowledged only so the daemon can release its recovery queue; stale-socket results are not acknowledged. On readiness handover, the runtime cancels active calls and queued results absent from `resume_calls` before accepting `ready_ack`.
|
|
205
206
|
14. A tool deadline cancels only that operation and never infers daemon death from tool duration. The independent daemon-liveness alarm owns socket invalidation. If same-instance readiness does not return before the grace deadline, the Worker rejects the detached request and the local runtime cancels ordinary calls, terminates their process trees, and discards queued results. A newly started daemon has a different instance identifier and cannot inherit prior calls.
|
|
206
|
-
15. `start_job` is different: after durable acceptance, the detached runner is no longer bound to an MCP response stream or daemon socket. Later cancellation uses `cancel_job` or the local CLI.
|
|
207
|
+
15. `start_job` is different: after durable acceptance, the detached runner is no longer bound to an MCP response stream or daemon socket. Later cancellation uses `cancel_job` or the local CLI. Cancellation-aware main steps re-read the owner-only cancellation marker after asynchronous resource admission/launch preparation and immediately before synchronous `spawn`; that marker read is the managed-job launch decision point. A marker already visible there prevents child creation, while cancellation appearing after the decision is post-dispatch and terminates the owned process tree. Because that marker is cross-process filesystem state, the cancellation API is a request/terminal-state protocol rather than a claim that its return instant is an atomic wall-clock barrier against a concurrently racing OS spawn syscall.
|
|
207
208
|
|
|
208
209
|
HTTP JSON-RPC IDs are scoped to each request or response stream and are not used as a token-wide duplicate key, so independent clients may reuse the same typed ID. Stdio has one process-local in-flight ID index because all requests share one explicit transport channel.
|
|
209
210
|
|
|
@@ -255,15 +256,15 @@ Fixed implementation-owned Git operations are not routed through that arbitrary-
|
|
|
255
256
|
|
|
256
257
|
The default `full` profile passes the complete parent environment. Isolated environment mode, used by the narrower named profiles unless overridden, creates private runtime HOME, temporary, and cache directories and passes only a small set of path/locale/platform variables. It reduces accidental credential inheritance but cannot prevent explicit access to known filesystem paths, credential stores, network services, or other user resources.
|
|
257
258
|
|
|
258
|
-
`execution-limits.mjs` is the shared source for local tool-call concurrency, one-shot process timeout/stdin/output limits, and process-session count/stdin/output/retention limits. Remote-owner `diagnose_runtime.runtime.execution_guardrails` reports those enforced limits, while local stdio `server_info.runtime.execution_guardrails` exposes the same contract together with explicit `not-enforced` values for CPU quota, memory quota, and network isolation. Public one-shot commands inline at most 32 KiB per stream. When either stream exceeds that preview, the runtime keeps up to 1 MiB per stream in a closed in-memory process session for thirty minutes and returns an `output_session_id`; `read_process` then reads monotonic byte-offset pages. Relay-origin process reads
|
|
259
|
+
`execution-limits.mjs` is the shared source for local tool-call concurrency, one-shot process timeout/stdin/output limits, and process-session count/stdin/output/retention limits. Remote-owner `diagnose_runtime.runtime.execution_guardrails` reports those enforced limits, while local stdio `server_info.runtime.execution_guardrails` exposes the same contract together with explicit `not-enforced` values for CPU quota, memory quota, and network isolation. Public one-shot commands inline at most 32 KiB per stream. When either stream exceeds that preview, the runtime keeps up to 1 MiB per stream in a closed in-memory process session for thirty minutes and returns an `output_session_id`; `read_process` then reads monotonic byte-offset pages. Relay-origin process reads use paced same-response follow-up. `process-session-read.mjs` owns read/wait/output orchestration while `process-session-remote-poll.mjs` owns only hosted projection/cooldown policy, leaving `process-sessions.mjs` focused on session lifecycle and ownership. The async read helper returns its internal plan/output snapshot before hosted projection is finalized; after the await the manager re-checks cancellation, then completes projection synchronously, so cancellation at the Promise boundary is not lost and a child exit there cannot produce `running=false` with stale cooldown metadata. The Worker and daemon cap the actual output/exit blocking wait at one second; a live session that has just consumed a remote blocking read enters a fifteen-second would-block cooldown. A repeated would-block request inside that cooldown is paced inside the same MCP call until output/exit or the cooldown boundary instead of returning an immediate running checkpoint, and the Worker execution budget covers that server-side delay. When `wait_for_exit=true`, ordinary output notifications remain readable but do not release the cooldown stage; only process exit or the monotonic cooldown deadline advances the call to its bounded one-second exit-wait stage. Every live relay-origin read identifies `status_polling_mode=paced_followup` and keeps `host_turn_handoff_recommended=false`; terminal reads identify `status_polling_mode=terminal`. The cooldown fields are explicitly blocking-only (`blocking_poll_throttled`, `next_blocking_poll_after_ms`) so a zero-wait status read is not misrepresented as globally rate-limited. Same-response follow-up is allowed when the task needs additional output or terminal state, but callers must not busy-loop and should respect the reported cooldown. Elapsed wall-clock time is not a host-deadline signal: callers continue bounded follow-up while calls are accepted and hand progress back only after an actual host/tool boundary is observed, external input/authorization is required, or the user explicitly requested a checkpoint. This state is per in-memory process session, does not alter local stdio/CLI wait behavior, and is intentionally a hosted pacing bound rather than durable scheduling. The oldest exited session is evicted before an active session is refused, so continuation retention is explicitly best effort rather than durable. The continuation stores command basename and cwd metadata but not argv or shell text. It is memory-only and disappears on runtime stop or daemon replacement.
|
|
259
260
|
|
|
260
|
-
Resource admission is a separate cooperative boundary, not an OS quota. `resource-foreground-wait.mjs` owns the ordinary process-start wait budget: a one-shot foreground call defaults to 20% of its execution timeout with a two-second floor, ten-second ceiling, and never more than the complete execution timeout. Owner-local process-session startup defaults to a ten-second cooperative wait, while relay-origin `start_process` performs one admission attempt without queueing by default so known host pressure is returned before the request-owned response budget is consumed. An explicit configured override remains authoritative for controlled diagnostics/tests. `LocalRuntime` leaves this default unconfigured in production so the services actually apply it, while explicit overrides are validated and capped at thirty minutes. These admission waits occur before process spawn and do not enlarge an outer relay deadline; cancellation propagates through the same coordinator wait. Detached managed-job steps use the shared durable-delivery admission ceiling instead: the runner may wait up to thirty minutes before spawn, child execution timeout starts only after admission succeeds, and persisted job status reports `current_phase=resource_admission` during that pre-spawn interval. Hosted status
|
|
261
|
+
Resource admission is a separate cooperative boundary, not an OS quota. `resource-foreground-wait.mjs` owns the ordinary process-start wait budget: a one-shot foreground call defaults to 20% of its execution timeout with a two-second floor, ten-second ceiling, and never more than the complete execution timeout. Owner-local process-session startup defaults to a ten-second cooperative wait, while relay-origin `start_process` performs one admission attempt without queueing by default so known host pressure is returned before the request-owned response budget is consumed. An explicit configured override remains authoritative for controlled diagnostics/tests. `LocalRuntime` leaves this default unconfigured in production so the services actually apply it, while explicit overrides are validated and capped at thirty minutes. These admission waits occur before process spawn and do not enlarge an outer relay deadline; cancellation propagates through the same coordinator wait. Detached managed-job steps use the shared durable-delivery admission ceiling instead: the runner may wait up to thirty minutes before spawn, child execution timeout starts only after admission succeeds, and persisted job status reports `current_phase=resource_admission` during that pre-spawn interval. Hosted managed-job status supports bounded autonomous follow-up: `managed-job-hosted-status.mjs` owns the relay-only projection, so relay-origin `read_job` reports `status_polling_mode=bounded_followup` plus `host_turn_handoff_recommended=false` for active jobs while terminal reads report `status_polling_mode=terminal`; local job reads retain their prior shape. Worker tool guidance allows a known active job to be followed again in the same assistant response when terminal state is required, while prohibiting busy loops, repeated `list_jobs` substitution, and speculative handoff based only on elapsed wall-clock time. `managed-job-read-wait.mjs` keeps that autonomy out of the host spin loop: the runtime `read_job` handler performs a relay-only 40-second default long-poll with five-second lightweight status-only progress probes, early return on meaningful progress, monotonic deadline accounting, and cancellation checks. The public maximum remains five minutes for clients that have independently demonstrated a longer request lifetime; the default is intentionally separate because live hosted evidence showed that the five-minute default outlived the target host tool invocation while a 40-second read survived. Full runner/recovery reconciliation is bounded to thirty-second intervals inside the advertised wait rather than executing on every lightweight probe. The monotonic wait deadline starts before the initial full read, so initial reconciliation consumes the same per-call budget, and an unchanged timeout returns the latest secure persisted progress state without adding a second heavyweight reconcile after the deadline. Hosted reads first verify a healthy runner with async process-start identity sampling; `managed-job-hosted-reconcile.mjs` owns that relay-only liveness/recovery decision, and only a non-current/ambiguous runner enters the existing recovery reconciliation path, so a normal long-running job no longer launches synchronous `ps` from the relay event loop every reconcile interval. `ManagedJobManager.read()` remains the local synchronous persistence/ownership projection while relay reads use `readHosted()`; `readProgress()` deliberately omits runner reconciliation and terminal-result projection during unchanged waits. Relay-origin `list_jobs` remains an `inventory` surface and does not recommend handoff. Ordinary managed-job result/status projection remains separate from this hosted pacing policy. Owner diagnostics separately expose the Boolean `waiters.drain_active`, computed by the same fairness selection state machine, so Green host/resource pressure is not confused with an aged protected waiter intentionally reserving a drain window. A fixed CPU request that exceeds the machine's best-case priority-specific launch window is structurally incapable of becoming admissible while the configured CPU headroom remains in force; it therefore returns non-retryable `cpu_request_exceeds_launch_window` before entering the retry sleep loop instead of masquerading as transient `cpu_pressure_window`. Explicit worker counts/argv are never rewritten and the headroom is not relaxed; elastic/unbounded requests retain their pressure-aware fitting behavior. `resource-command-profile.mjs` classifies known light operations, bounded adaptive unknowns, and known CPU/I/O/mixed build families. `resource-script-classification.mjs` classifies only an actually executed shell/Node/Python script operand or package-manager script name, and `resource-shell-analysis.mjs` owns conservative shell token/segment parsing. This lets direct orchestration roots reserve startup capacity before descendant fan-out without letting an unrelated argument such as a test filename or release note impersonate a heavy script. The arbitrary-process zero-resource allowlist is intentionally identity-bound: `resource-light-command.mjs` accepts only a small set of standard absolute executables for constant/output, process-table, uptime, and sleep probes. Bare PATH-resolved names and every arbitrary shell invocation remain adaptive even when their apparent command is cheap, because executable resolution, startup configuration, repository configuration, path behavior, options, helpers, or input size can change the actual work. Caller-controlled Git, filesystem metadata/query commands, lookup helpers, recursive/search/file-stream processors, `find`, application launchers, and script interpreters therefore never receive zero-resource admission from basename alone. Implementation-owned Git/diagnostic probes use their separate fixed-argv/internal boundary and do not depend on this allowlist. `resource-admission.mjs` persists per-user owner-only leases and waiters outside workspace state; ownership is bound to PID plus process-start identity and, where supported, the isolated process group. Releasing the caller-side lease does not delete a bound POSIX reservation while that isolated process group still exists: the persisted lease remains until ordinary pruning observes the group gone, so a detached descendant cannot become unaccounted merely because its direct caller returned. Independent nested process roots retain their own full durable leases for crash recovery, but live accounting does not blindly add an orchestration envelope and every child reservation. `resource-process-ancestry.mjs` samples the live parent graph through `resource-process-ancestry-cache.mjs`; the async cache coalesces concurrent requests and retains a completed snapshot for one second so admission retries do not repeatedly enumerate the complete process table. If ancestry cannot be sampled, accounting falls back to conservative full summation. `resource-lease-accounting.mjs` forms an ephemeral lease forest whose effective vector is the component-wise maximum of a node's own envelope and the sum of its direct lease children. A pending nested request is charged only for the additional vector it contributes to that forest. Same-key contention is exempted only for the request's actual ancestor lease chain; siblings and unrelated roots still serialize.
|
|
261
262
|
|
|
262
|
-
Live host-pressure observation is deliberately asynchronous. `resource-probe-command.mjs` owns bounded child-probe transport
|
|
263
|
+
Live host-pressure observation is deliberately asynchronous-only. `resource-probe-command.mjs` owns bounded `execFile` child-probe transport and intentionally exposes no synchronous `spawnSync` path; `resource-host-darwin.mjs` runs Darwin memory, VM, disk, and thermal probes concurrently, `resource-process-ancestry.mjs` samples the parent graph asynchronously, and `resource-host-snapshot.mjs` composes those results without blocking the daemon event loop on `ps`, `memory_pressure`, `vm_stat`, `iostat`, or `pmset`. The previous sync host/Darwin/process-parent samplers and sync ancestry cache are removed rather than retained as test-only APIs that a future runtime path could accidentally call. Process-start identity used when binding a spawned lease is also sampled asynchronously before the coordinator file lock is taken. Lease/waiter pruning uses a separate one-second cached async process-start snapshot collected outside the transaction lock; while holding the lock it performs only current PID liveness plus in-memory generation comparison. Missing snapshot evidence fails closed for a live PID, while a proven generation mismatch remains reclaimable before isolated process-group liveness can preserve the lease. Current lease/waiter staging recovery consumes the same snapshot, and the migration-only legacy transaction-owner staging path performs its process-generation check asynchronously before destructive recovery. `resource-host-cache.mjs` treats a same-project general host snapshot as fresh for 500 milliseconds while retaining successful I/O evidence for at most five seconds; a stale general sample can therefore refresh cheap CPU/memory/load data without waiting for another one-second `iostat` interval, but an older throughput peak is not copied into a current quick sample. Timestamps represent sample completion rather than sample start. `resource-admission-policy.mjs` evaluates effective CPU, memory, I/O, disk-reserve, startup-window, and host-pressure values without imposing global serialism. macOS `iostat -Id` exposes transfer and throughput activity, not device utilization/saturation; fresh high IOPS/MB-per-second evidence therefore tightens I/O capacity as Yellow pressure but cannot by itself mark the host Red. Red remains reserved for direct critical evidence such as thermal/memory/disk-headroom failure or a severe load backlog corroborated by current CPU/I/O evidence. Linux additionally uses `MemAvailable` and optional PSI avg10 observations through `resource-host-linux.mjs`; sustained `psi_io_full_avg10 >= 60` is direct critical I/O-stall evidence and therefore Red, while lower I/O PSI remains a Yellow capacity-throttling signal. The soft free-disk floor is `min(80 GiB, max(8 GiB, 15% of volume))`; the post-reservation hard floor is `min(50 GiB, max(5 GiB, 10% of volume))`. Disk-only hard-red pressure has one narrow self-recovery exception: only an internally classified direct standard absolute deletion executable with the exact small `disk-reclaim` envelope may enter using Yellow capacity limits, and only when `disk_free_headroom_critical` is the sole critical reason. PATH-resolved or shell-composed deletion cannot claim that class, and thermal, memory, PSI, CPU/load, or other Red evidence still blocks it. Diagnostics expose observed busy CPU next to reserved CPU and their ratio as a mismatch signal, not as process attribution. `resource-admission-diagnostics.mjs` owns the privacy-safe snapshot/check projection, while `resource-admission-diagnostic-error.mjs` separately classifies bounded transaction/staging-lock contention as retryable coordinator-busy snapshot unavailability instead of inflating the diagnostics module or mislabeling contention as process execution failure. `resource-waiters.mjs` normally skips an older request that does not fit, but a sufficiently aged rank-zero waiter blocked by coordinator-owned project/CPU/I/O/memory capacity can enter a protected drain phase if it is structurally feasible after current leases disappear. Fixed implementation-owned probes bypass this coordinator so diagnostics/recovery remain usable during resource pressure.
|
|
263
264
|
|
|
264
265
|
The same protocol is intentionally language-neutral: compatible workflow controllers may use the same per-user coordinator root, schema, lease ownership, contention keys, and waiter ordering. This permits Machine Bridge and workflow-bundle to coordinate independent process trees without a new daemon. `resource-project-key.mjs` canonicalizes the existing filesystem path before deriving the v1 project-contention hash, so path aliases such as macOS `/var` versus `/private/var` and symlinked workspace roots cannot create different mutex identities across implementations; Windows additionally normalizes separator form, path case, and extended-path prefixes before hashing. Lease/waiter records contain resource families, numeric reservations, hashes, timestamps, and process ownership but never argv, shell text, or raw project paths. The anonymous `project_hash` is derived from the same canonical project identity as contention and is diagnostic correlation only; it is not authorization or admission authority. The coordinator is fail-closed on malformed persistent authority.
|
|
265
266
|
|
|
266
|
-
`resource-transaction-lock.mjs`
|
|
267
|
+
`resource-transaction-lock.mjs` now publishes current `transaction.lock` ownership as one complete-before-visible owner-state regular file using the same exclusive-file primitive as other owner-state locks. The record carries a fixed `resource-coordinator` purpose, random ownership token, PID, and process-start identity; release and stale-file reclamation remove only a record whose token/purpose and secure-file identity still match. This removes the mkdir-before-owner publication state and directory quarantine/restore path from the current writer. The module retains a bounded beta.104 transition reader because beta.104 can still be the immediately preceding live daemon or durable runner during a beta.105 handoff and still writes the beta.60-origin schema-1 `transaction.lock/owner.json` directory shape. A live legacy directory is waited out; an incomplete directory is not reclaimable until the bounded owner-publication grace expires; dead owner-publication staging is removed only after publisher/process and exact-file identity checks; arbitrary, multiple, hard-linked, or live-publisher contents fail closed. Token-bearing stale legacy directories are quarantined and revalidated before recursive removal, and restore never overwrites an already-present replacement generation. The narrow same-user missing-check-to-restore race therefore remains only while consuming a beta.104 legacy directory generation, not in the beta.105 writer, and the directory reader is removable once beta.104 no longer falls within the supported rolling/live transition. Cross-version exclusion remains bidirectional because both versions contend on the identical final pathname: beta.105 reads beta.104 directories, while beta.104 already understands the owner-state regular-file shape beta.105 writes. Transaction-lock waiting remains operation-bounded rather than globally fixed: admission uses only the remaining caller admission budget, capped at 30 seconds, while spawned-process lease bind/release ownership settlement may wait up to 30 seconds and managed-job background admission retains its longer outer wait.
|
|
267
268
|
|
|
268
269
|
`resource-staging-recovery.mjs` is the narrow exception to the directory's final-name-only rule: it understands only the exact temporary naming emitted by the coordinator's own exclusive/atomic file helpers. Lease/waiter recovery handles a committed two-hard-link alias or a dead-publisher single-link uncommitted replacement after identity/link-count rechecks. A staging inode owned by a still-current publisher is classified as `MBM_RESOURCE_STAGING_BUSY`: `ResourceCoordinator.withLock()` releases the transaction generation and performs at most four five-millisecond retries so an ordinary concurrent complete-before-visible publication can settle, while a staging owner that remains live after that fixed budget still fails closed. Process admission maps the exhausted busy classification to the same retryable resource-unavailable surface as transaction-lock contention instead of exposing an internal error. Transaction-owner recovery separately recognizes only `.owner.json.<pid>.<nonce>.tmp` inside an aged ownerless lock directory, waits for a still-current publisher, and removes only a dead-publisher single-link staging generation before the canonical-directory `rmdir` recheck. It does not ignore or generically delete arbitrary temporary files.
|
|
269
270
|
|