@arnilo/prism 0.2.5 → 0.2.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/docs/migration.md CHANGED
@@ -1,5 +1,34 @@
1
1
  # Migration guide
2
2
 
3
+ ## 0.2.6 → 0.2.7 enterprise ERP production readiness (additive)
4
+
5
+ Release **0.2.7** (plan 027) adds the enterprise ERP production-readiness primitives behind optional host-activated seams: the transactional outbox/inbox + bounded dispatcher, the durable saga compensation/reconciliation engine, multi-party separation-of-duties approvals, signed hash-chained audit export with WORM/SIEM sinks, field-level classification + fail-closed redaction, and the deterministic ERP invariant evals. **Additive-only: no exported declaration removed or changed, no persisted 0.2.6 shape repurposed.**
6
+
7
+ New ERP tables use **separate forward-only migrations** (no down migrations exist; production rollback is roll-forward repair only):
8
+
9
+ - `prism_erp_outbox` / `prism_erp_inbox` (migration `004_erp_messaging`, version 4) — transactional outbox/inbox with `FOR UPDATE SKIP LOCKED` claim, `ON CONFLICT DO NOTHING` idempotent append, claim-token CAS, and three partial indexes. Outbox append must run in the caller-owned `PoolClient` transaction with the business mutation (atomicity is the host's responsibility).
10
+ - `prism_erp_approvals` (migration `005_erp_approvals`, version 5) — multi-party approval requests with decisions stored as JSONB, `FOR UPDATE` row locking for atomic quorum recomputation, rejection as any-party veto, expiry checked at every protected transition, and atomic grant consumption in the host transaction.
11
+
12
+ Saga state persists as a surrogate `WorkflowCheckpointRecord` through the existing `WorkflowCheckpointAdapter` (private workflow id `__prism_saga__/<key>`) — no saga-specific SQL or 0.2.6 shape is repurposed. Audit export, field policy, and ERP invariant evals are stateless or in-memory and add no persisted shape. Secret-manager adapters (Vault/AWS/Azure/GCP) stay **deferred** behind the demand gate; no adapter ships and no ambient credential discovery is added.
13
+
14
+ **Rollback notes.** Rollback = restore the 0.2.6 manifests/tag. The two new ERP migrations are forward-only; before downgrading, stop all 0.2.7 workers (outbox dispatcher, saga engine, audit exporter) and drop or ignore the `prism_erp_outbox`/`prism_erp_inbox`/`prism_erp_approvals` tables (they hold no 0.2.6 data). No 0.2.6 persisted shape changed, so an ordinary downgrade is store-safe; the added exports and ERP tables simply disappear. **"ERP production ready" remains blocked until the 0.3.0 live-service matrix is recorded** — this release adds the primitives and the protected journey evidence, not the live-service matrix.
15
+
16
+ ## 0.2.5 → 0.2.6 durable recovery, workspaces, and coding-agent readiness (additive)
17
+
18
+ Release **0.2.6** (plan 026) adds the coding-agent readiness capabilities behind optional host-activated seams: host-selected PTY backends, the indexed/semantic repository-search seam, the ownership-scoped multi-repository/worktree lifecycle, durable process/ACP recovery, and the patch-review/diagnostics workflow. **Additive-only: no exported declaration removed or changed, no persisted 0.2.5 shape repurposed.**
19
+
20
+ New durable records use **separate versioned checkpoint namespaces**, never the 0.2.5 shapes:
21
+
22
+ - `prism.coding-agent.process.v1` (schemaVersion 1) — managed-process recovery records. Readers reject unknown schema versions and corrupt/foreign records fail closed (dropped, never recovered).
23
+ - `prism.coding-agent.workspace.v1` (schemaVersion 1) — coding workspace lifecycle records.
24
+ - `prism.coding-agent.cancel.v1` (schemaVersion 1) — durable ACP run-cancel markers.
25
+
26
+ `CodingCheckpointMetadata` (schemaVersion 1, `prism.coding-agent`) is **never silently repurposed**; 0.2.5 readers reject unknown schema versions as before.
27
+
28
+ **ACP active-run references (Task 5 decision: additive optional field).** `PersistedAcpSession` gains an optional bounded `activeRun` ref (frozen 512-byte cap) recorded while a durable run is live. The decision recorded here: an additive optional field, not a separate recovery namespace, because the ref is advisory metadata — the authoritative run status is always re-queried from `AgentRunLifecycle.status` at restore time, and 0.2.5 hosts safely ignore the field. `PersistedAcpSession.activeRun` stays optional; 0.2.5 records remain readable and a 0.2.5 host reading a 0.2.6 record does not lose required recovery state (the run state itself lives in the existing `prism.agent-run` records, which are untouched).
29
+
30
+ **Downgrade to 0.2.5** is safe only after stopping 0.2.6 workers/replicas and marking any live 0.2.6 process/workspace records `unknown` (their leases expire within TTL); durable state never serializes a PTY fd, browser context, process object, controller, pending promise, raw terminal output, env, token, or credential, and no exact-process-survival claim is made (attach-if-attested, otherwise unknown).
31
+
3
32
  ## 0.2.4 → 0.2.5 maintainability and bounded performance (no migration)
4
33
 
5
34
  Release **0.2.5** (plan 025) is the maintainability-and-bounded-performance cut: the six remaining implementation god-modules split into cohesive internal family files behind preserved barrels (compat-preserving, no `exports`-map subpath), 21 pure persistence helpers moved into the dependency-free `session-store-codecs` package (ownership scope/assertion, checkpoint stale/encode/decode, branch cursors, lifecycle quota/reason/page-limit, search metadata/clipping, deepFreeze/string-array/throwIfAborted, feedback row mapping), the quadratic per-push `Buffer.concat` loops in language framing and tar parsing became chunk-array readers (linear; caps and fail-closed overflow byte-identical), two internal dead type aliases removed (`PostgresPersistenceCloseOptions`, `SqlitePersistenceCloseOptions` — never re-exported from their adapter indexes), and 76 behavior-backed coverage regressions closed the low-coverage core areas (core 91.43/84.80/91.60 lines/branches/functions). **No runtime contract change and no migration**: no exported declaration was removed or changed (the plain reviewed compat gate at 0.2.5 shows the version literal plus 105 additive internal-helper exports), no persisted shape/schema/default/behavior changed, no new runtime dependency. Store compatibility with 0.2.4: **compatible in both directions** — no migration step; rollback = restore the 0.2.4 manifests/tag (stores never change; the added exports disappear on downgrade). The 20 dead-but-compat-tracked exports deferred from Task 4 are the 0.3.0 breaking-cut removal list (see `docs/_evidence/phase25-dead-exports-triage.md`).
@@ -0,0 +1,104 @@
1
+ # Operations runbook: high availability, failover, and fencing
2
+
3
+ ## What it does
4
+
5
+ This page is the operator runbook for the high-availability story proven by plan 027 Task 6: two replicas can serve, inspect, cancel, resume, and reconcile durable ACP/workflow/saga/outbox/export operations after either replica dies. Correctness never depends on a dead process's in-memory registry — it comes from the durable `LeaseStore` (owner, token, fencing counter, expiry, renewal), the `CheckpointStore` (version CAS plus monotonic fencing token), and idempotent side-effect sinks such as the ERP outbox (`ON CONFLICT DO NOTHING` on a stable message id). The two-process proof is `scripts/phase27-ha-worker.mjs` orchestrated by `scripts/phase27-ha.test.mjs`, which records exact commands, process IDs, injected failures, timings, and durable final states in `docs/_evidence/phase27-ha-evidence.json`.
6
+
7
+ ## When to use it
8
+
9
+ - Before running any multi-replica deployment of the server, workflow coordinator, saga runner, ACP host, or enterprise dispatcher — read the local-registry limitations and the lease/fence model.
10
+ - When an operator or on-call engineer sees a hung lease, an uncertain commit, or a split-brain suspicion: follow "Failover procedure" and "Uncertain commits" below before touching anything.
11
+ - When sizing leases: the failover ceiling is lease TTL plus the peer's acquisition poll interval; the drill asserts `failoverMs <= ttlMs + 5000`.
12
+
13
+ ## Inputs / request
14
+
15
+ - A `LeaseStore` and `CheckpointStore` backed by the same durable store (PostgreSQL through `createPostgresPersistence`, or the corresponding production adapter). Lease keys carry an ownership scope (tenant/account/user) that is part of the trust boundary.
16
+ - Operations that need a leader: `acquireLease({ namespace, key, ownerId, ttlMs })` returns a lease with `token` and monotonic `fencingToken`, or `null` while another owner holds it.
17
+ - Durable progress: `saveCheckpoint({ namespace, key, version, expectedVersion, fencingToken, value })` — versions strictly increase, `expectedVersion` must match the current version, and a lower or absent fencing token can never replace a fenced record.
18
+ - Idempotent side effects: give every external effect a stable id (ERP outbox `messageId` is the designed carrier) so replay is safe.
19
+
20
+ ## Outputs / response / events
21
+
22
+ - A lease record: `{ namespace, key, ownerId, token, fencingToken, acquiredAt, expiresAt, updatedAt }`. Expired rows retain their fencing counter; the next owner inherits `fencingToken + 1`.
23
+ - A checkpoint record: `{ namespace, key, version, fencingToken?, value, createdAt, updatedAt }`. Cursor/value changes are CAS-committed; a peer can replay an unfinished step but can never skip ahead or move the cursor backward.
24
+ - Failover timing: the drill reports `failoverMs` (wall time between the owner's death and the peer's acquisition) and asserts it against the frozen ceiling.
25
+
26
+ ## Request/response example
27
+
28
+ ```ts
29
+ const lease = await stores.leases.tryAcquireLease({
30
+ namespace: "erp.ops", key: "invoice-42", ownerId: "worker-b", ttlMs: 30_000,
31
+ });
32
+ if (!lease) return "another replica owns invoice-42";
33
+ await stores.checkpoints.saveCheckpoint({
34
+ namespace: "erp.ops", key: "invoice-42",
35
+ version: 3, expectedVersion: 2, fencingToken: lease.fencingToken,
36
+ value: { cursor: 2, steps: ["reserve", "charge"] },
37
+ });
38
+ // Side effect with a stable id (idempotent replay):
39
+ await outbox.append(client, { tenantId, messageId: "pay-t/invoice-42/charge", topic: "erp.payment.requested", payload });
40
+ await stores.leases.releaseLease({ namespace: "erp.ops", key: "invoice-42", ownerId: "worker-b", token: lease.token });
41
+ ```
42
+
43
+ ## Implementation example
44
+
45
+ ```sh
46
+ # Protected two-replica drill (requires PRISM_TEST_POSTGRES_URL; Docker image
47
+ # postgres:16-alpine is the local stand-in). Recorded evidence lands in
48
+ # docs/_evidence/phase27-ha-evidence.json.
49
+ `PRISM_TEST_POSTGRES_URL` set to the protected connection string (locally a disposable `postgres:16-alpine` container): `node --test scripts/phase27-ha.test.mjs`
50
+ ```
51
+
52
+ The drill: worker A acquires, heartbeats, commits the charge effect into the
53
+ outbox, and is SIGKILLed inside the window between the effect commit and its
54
+ final cursor save. Worker B (separate process, own pool, no access to A's
55
+ registry) reads the durable state, waits out the lease expiry, acquires,
56
+ replays the uncertain commit idempotently (outbox count stays 1), finishes,
57
+ and releases. A stale-fence/stale-revision write and an old-token renewal are
58
+ rejected; two simultaneous acquisitions yield exactly one owner; a foreign
59
+ tenant's reads, writes, and lease takeover all fail closed.
60
+
61
+ ## Extension and configuration notes
62
+
63
+ - Lease TTL is the availability knob: too short spams contention, too long
64
+ delays failover. The frozen drill ceiling is `ttlMs + 5000` ms; renew at
65
+ `ttlMs / 3` (the workflow coordinator, saga engine, and this drill all
66
+ follow this pattern).
67
+ - Local in-memory registries (workflow active-run map, coordinator active
68
+ map, A2A live-task registry, ACP session registry, coding-agent sessions,
69
+ RPC active runs) are optional fast paths only. On restart, every component
70
+ reloads authority, status, and cursors from durable stores; a killed
71
+ replica's registry is never required to make progress.
72
+ - SIEM/alerting: emit owner/fence/lease-age/state metadata only — never tenant
73
+ payloads, tokens, or credentials. Metrics to watch: lease-acquire latency,
74
+ fence counter jumps (a jump means a takeover happened), and outbox pending
75
+ depth.
76
+
77
+ ## Security and performance notes
78
+
79
+ - Fencing is enforced at the store: a stale owner's write fails with
80
+ `ERR_PRISM_CHECKPOINT_CONFLICT` (stale version, failed CAS, or stale
81
+ fencing token), and a stale token's renewal returns `null`. There is
82
+ intentionally no "force unlock" operation — deleting a lease row manually
83
+ bypasses fencing and can cause split-brain writes; never do it.
84
+ - Tenant ownership is checked on every durable read/write; cross-tenant
85
+ reads, saves, and lease acquisitions fail closed with ownership-mismatch
86
+ errors.
87
+ - An uncertain commit (side effect landed, cursor not advanced) is resolved
88
+ by replaying the effect with its stable id — never by guessing. Reload
89
+ durable state before any retry; retries that skip the reload risk
90
+ overwriting a peer's fenced progress.
91
+ - Failover is bounded but not instantaneous: the peer cannot acquire until
92
+ the lease expires, so recovery time is at least the remaining TTL.
93
+ Performance is linear in record size; the drill records contention/latency
94
+ with bounded, jittered acquisition polls (no hot loops) and reports the
95
+ measured numbers in the evidence JSON.
96
+
97
+ ## Related APIs
98
+
99
+ - `LeaseStore` / `CheckpointStore` — the durable contracts this runbook relies on.
100
+ - `createPostgresPersistence` — the PostgreSQL adapter used by the drill.
101
+ - `ErpOutboxStore` — the idempotent side-effect carrier used to make replay safe.
102
+ - `createWorkflowCoordinator` and `defineSaga`/`runSaga`/`resumeSaga` — higher-level consumers with the same fencing/cursor semantics.
103
+ - `scripts/phase27-ha-worker.mjs` / `scripts/phase27-ha.test.mjs` — the reproducible drill; `docs/_evidence/phase27-ha-evidence.json` — the recorded run.
104
+ - [Signed, hash-chained audit export](audit-export.md) — the cursor/CAS pattern applied to audit exports.
@@ -118,6 +118,40 @@ Policy is optional. Hosts wire `record*` helpers or `evaluateAndAppend` at permi
118
118
  - The OPA decision fetch (0.2.1) is DNS-pinned through the core `pinnedFetch` primitive: one resolve per request, every resolved address SSRF-checked before the connect (rebinding defense), redirects rejected outright, timeouts/retries unchanged, and private-answer denials surface `MediaContentError` (`ssrf_denied`) rather than a transport error.
119
119
  - Export never full-scans: page size is capped; raise hard caps only with Phase 8 freeze + tests + docs updates.
120
120
 
121
+ ## Multi-party approvals
122
+
123
+ Immutable approval requests carry an action digest, requester, and required roles/quorum; verified host identities vote through an explicit `ApprovalAuthority`. Prism records decisions and enforces separation-of-duties, expiry, revocation, bounded delegation, rejection, and policy-revision pins — it never resolves identities or roles itself.
124
+
125
+ | API / field | Meaning |
126
+ | --- | --- |
127
+ | `ApprovalAuthority` | Host-owned `policyRevision` + `resolveRoles(identity, request)` returning role grants (and delegation chains) |
128
+ | `ApprovalRequest` | Immutable request: `action {kind, digest}`, `requirements [{role, quorum}]`, `separateFromRequester`, `expiresAt`, `delegationMaxDepth` |
129
+ | `ApprovalRecord` | Request + `status` (`pending/approved/rejected/revoked/consumed`) + `revision` + immutable `decisions` + `policyRevision` |
130
+ | `ApprovalStore.create/decide/revoke/consume/get/query` | Durable transitions; every transition records an `auditRef` |
131
+ | `createMemoryApprovalStore({ authority })` | Single-process reference adapter sharing the pure transition logic |
132
+ | `createPostgresApprovalStore({ pool, schema, authority })` | Cross-replica storage (migration `005_erp_approvals`, `@arnilo/prism-enterprise-postgres`) |
133
+
134
+ Quorum rules:
135
+
136
+ - Each requirement needs `quorum` **distinct approved principals** through that role. Same principal voting the same role twice is idempotent; changing a vote is a conflict.
137
+ - A rejection is terminal and vetoes the request (`rejected`), so `consume` always denies.
138
+ - `separateFromRequester: true` denies the requester (and any authority chain deriving from the requester) from deciding.
139
+ - `delegationMaxDepth` bounds the persisted delegation chain (delegator first); grants cannot outlive the request, leave the tenant, or widen the action.
140
+ - Pin: `ApprovalAuthority.policyRevision` must equal the record's revision; a policy bump invalidates outstanding approvals (release denied).
141
+
142
+ Revocation semantics: `revoke` is only valid for `pending`/`approved` grants, is taken by a verified tenant actor with an `auditRef`, and is terminal. After consumption, revocation cannot undo the effect — the record stays `consumed` as provenance, and reconciliation is host work.
143
+
144
+ Operator audit queries: `query({ tenantId, status })` pages by `(created_at, id)`; `get({ tenantId, requestId })` returns full immutable decisions with grants and delegation chains. Every decision and manual transition carries the host `auditRef`.
145
+
146
+ ```sql
147
+ -- pending approvals that never released (example host view)
148
+ SELECT id, policy_revision, expires_at, decisions
149
+ FROM prism.prism_erp_approvals
150
+ WHERE status = 'pending' AND expires_at < now();
151
+ ```
152
+
153
+ Hosts own identity verification and the role source; Prism does not certify NIST compliance. NIST SP 800-53 AC-5 (separation of duties) and AC-6 (least privilege) are control guidance only, not certification claims.
154
+
121
155
  ## OPA external policy adapter (`@arnilo/prism-policy/opa`, 0.0.28)
122
156
 
123
157
  Optional `createOpaPolicyEvaluator` evaluates `PolicyEvaluateRequest`s against a host-pinned OPA REST endpoint (`POST /v1/data/<path>` with `{"input": <document>}`) and returns a core `PolicyEvaluator` for `evaluateAndAppend`. Native `fetch` only; no OPA SDK dependency.
@@ -17,7 +17,7 @@
17
17
 
18
18
  ## When to use it
19
19
 
20
- Use when a host needs attachable long-running processes (watch modes, language servers, interactive CLIs) that one-shot `shell` cannot model. Do not use as a job-control language or PTY emulator — `pty: true` fails closed with `ERR_PRISM_PROCESS_PTY_UNSUPPORTED` until a platform capability is wired. Pass a sandbox with `startProcess` for contained long-running work; omit `sandbox` for native spawn.
20
+ Use when a host needs attachable long-running processes (watch modes, language servers, interactive CLIs) that one-shot `shell` cannot model. Do not use as a job-control language. `pty: true` (host-selected PTY) requires a `ptyBackend` passed to `createProcessSessions`; without one it fails closed before spawn with `ERR_PRISM_PROCESS_PTY_UNSUPPORTED`. Pass a sandbox with `startProcess` for contained long-running work; omit `sandbox` for native spawn.
21
21
 
22
22
  ```ts
23
23
  import { createProcessSessions } from "@arnilo/prism-coding-agent";
@@ -43,6 +43,7 @@ await sessions.dispose();
43
43
  | `ownership?` | `OwnershipScope` | Default owner key for sessions. |
44
44
  | `identity?` | `AgentIdentity` | When ownership omitted, owner key projects from identity. |
45
45
  | `sandbox?` | `ProcessSandboxBackend` | When set: require `startProcess` or fail closed; `status` loss → all running → `unknown`. |
46
+ | `ptyBackend?` | `ProcessPtyBackend` | Host-selected interactive-terminal backend. `pty: true` delegates only here; absent backend → `ERR_PRISM_PROCESS_PTY_UNSUPPORTED` before spawn. Host supplies the PTY engine (e.g. node-pty); Prism never depends on one. |
46
47
 
47
48
  `ProcessStartRequest`:
48
49
 
@@ -51,7 +52,8 @@ await sessions.dispose();
51
52
  | `command` / `args?` | Executable + argv (not a shell string). |
52
53
  | `cwd?` | Relative/absolute path contained under registry `cwd`. |
53
54
  | `env?` | Extra env merged onto `process.env` (never in fingerprint). |
54
- | `pty?` | Default false; unsupported → `ERR_PRISM_PROCESS_PTY_UNSUPPORTED`. |
55
+ | `pty?` | Default false. `true` requires the `ptyBackend` host option (delegated only to it); unsupported host → `ERR_PRISM_PROCESS_PTY_UNSUPPORTED` before spawn. |
56
+ | `terminal?` | `{ columns, rows, term? }` for `pty: true`; defaults `120 × 40`, `xterm-256color`. Bounds: columns 1–120 (hard 500), rows 1–40 (hard 200), TERM ≤ 64 bytes (hard 256). |
55
57
  | `lifetimeMs?` | Bounded by `maxLifetimeMs`. |
56
58
  | `owner?` | Override owner string. |
57
59
  | `releaseOnCancel?` | If true, `cancelOwned` releases instead of killing. |
@@ -68,10 +70,23 @@ await sessions.dispose();
68
70
  | `cancelOwned(owner)` | Kill (default) or release owned running sessions. |
69
71
  | `markUnknown` | Backend-loss terminal state; never fabricates `exitCode`. |
70
72
  | `reconcile()` | Host resume: mark every running/starting session `unknown` (O(sessions)). |
73
+ | `resize?` | Only when `ptyBackend.capabilities.resize` is true: bounded `{ columns, rows }` routed to the live terminal. |
71
74
 
72
75
  Events: `process_started`, `process_exited`, `process_killed`, `process_released`, `process_expired`, `process_unknown`.
73
76
 
74
- Errors: `ERR_PRISM_PROCESS_POLICY`, `ERR_PRISM_PROCESS_OWNERSHIP`, `ERR_PRISM_PROCESS_STATE`, `ERR_PRISM_PROCESS_LIMIT`, `ERR_PRISM_PROCESS_PTY_UNSUPPORTED`, `ERR_PRISM_PROCESS_UNSUPPORTED`.
77
+ Errors: `ERR_PRISM_PROCESS_POLICY`, `ERR_PRISM_PROCESS_OWNERSHIP`, `ERR_PRISM_PROCESS_STATE`, `ERR_PRISM_PROCESS_LIMIT`, `ERR_PRISM_PROCESS_PTY_UNSUPPORTED`, `ERR_PRISM_PROCESS_PTY_BACKEND`, `ERR_PRISM_PROCESS_PTY_LIMIT`, `ERR_PRISM_PROCESS_UNSUPPORTED`.
78
+
79
+ ## Host-selected PTY backend contract
80
+
81
+ `ptyBackend` is the host's interactive-terminal capability (plan 026 Task 1):
82
+
83
+ - **Delegation only.** `pty: true` starts a session exclusively through `ptyBackend.startPty`; the native spawn path never allocates a terminal. A missing backend or one without `startPty` fails closed before any spawn (`ERR_PRISM_PROCESS_PTY_UNSUPPORTED`). The non-PTY path is byte-compatible with the 0.2.5 baseline.
84
+ - **Contract.** `startPty({ file, args, cwd, env, columns, rows, term, onData })` returns `{ metadata?, write, signal, kill, release, wait, resize? }`. `capabilities.resize` is explicit — `resize` on the session exists only when declared; never duck-typed. `wait()` resolves on process exit (and rejects on backend loss); the host stops delivering `onData` once the session is terminal.
85
+ - **Terminal data is untrusted output.** Control sequences are never parsed or emulated; the host (or an attached terminal client) interprets them. Input is raw terminal bytes; NUL is rejected with a policy error, other control bytes pass through as terminal data.
86
+ - **Bounded attach.** `startPty` must settle within `maxPtyAttachTimeoutMs` (30 s default, 120 s hard); overflow removes the session record and fails with `ERR_PRISM_PROCESS_PTY_LIMIT`. Resize is rate-limited (60/min default, 600 hard) and fails with the same code. Backend `metadata` is bounded (`maxPtyBackendMetadataBytes`, 4 KiB default / 16 KiB hard).
87
+ - **Backend loss.** A throwing `startPty` or a lost `wait()` surfaces as `ERR_PRISM_PROCESS_PTY_BACKEND` with a generic message (backend error text is never embedded); the session becomes `unknown` with `exitCode: null` — never fabricated.
88
+ - **Parity.** Policy, cwd/ownership, cancel/expiry sweeps, input/lifetime/output caps, command fingerprint, events, and redaction behave exactly as non-PTY sessions; PTY sessions count against `maxSessions`.
89
+ - **Recovery caveat (phase 26, task 5).** PTY sessions are not durable across restart: no serialized terminal fd or raw output is ever persisted. Restart recovery reports such sessions `unknown` unless a host `recoveryBackend` re-attaches them; there is no exact-process-survival claim.
75
90
 
76
91
  ## Request/response example
77
92
 
@@ -123,6 +138,46 @@ await sessions.dispose();
123
138
  - Command fingerprint is SHA-256 of `[command, ...args]` only (no env).
124
139
  - Docker reference adapter does not implement `startProcess` yet — fail closed until a capable runtime is wired.
125
140
 
141
+ ## Durable process recovery (plan 026 Task 5)
142
+
143
+ Optional, host-activated: pass `checkpoints` + `leases` + `ownerId` (all three
144
+ together; a partial recovery configuration fails closed at construction) and
145
+ optionally `recoveryBackend` + `recoveryLimits`. With durability configured:
146
+
147
+ - Intent is persisted BEFORE spawn into the versioned namespace
148
+ `prism.coding-agent.process.v1` (schemaVersion 1, category `coding-process`),
149
+ and every lifecycle transition (running, exited, killed, released, expired,
150
+ unknown) is a CheckpointStore CAS write under a monotonic LeaseStore fencing
151
+ token. Transition writes are serialized per record so CAS order never inverts
152
+ on slow stores.
153
+ - `recover()` reconciles durable records against the live registry:
154
+ - records already live here report `attached` without mutation;
155
+ - terminal records report `terminal` with their exit code;
156
+ - `starting|running` records attach-if-attested: a `backendRef` (opaque
157
+ non-secret ref surfaced by a PTY/sandbox handle's optional `ref`) plus a
158
+ host `recoveryBackend.attach(ref)` (bounded attach timeout 30 s default /
159
+ 120 s hard) may reattach; otherwise the record atomically becomes `unknown`
160
+ with exit code null — no PID probing, no fabricated exit, no duplicate
161
+ spawn. `recover()` without durability configured throws
162
+ `ERR_PRISM_RECOVERY_UNSUPPORTED`.
163
+ - Recovered sessions are live registry sessions: input/signal/kill/release/
164
+ resize/wait reach the attached backend; output streaming is not re-established
165
+ after a restart (the host backend owns any buffered output behind its ref).
166
+ - Replica coordination: every mutation takes a per-record lease (30 s default /
167
+ 300 s hard); a crashed replica's lease lapses within TTL, a live one renews
168
+ on transitions and releases on terminal transitions. A held lease makes the
169
+ second replica report `unknown` without touching the record; CAS/fence
170
+ conflicts fail closed with `ERR_PRISM_RECOVERY_FENCE`.
171
+ - `cancelOwned` after recovery either reaches the attached backend or records
172
+ the durable record unknown — never a fabricated exit.
173
+ - Durable records are metadata only: no child/PTY handle, controller, promise,
174
+ raw output, env, token, or credential is ever serialized; forbidden fields,
175
+ corrupt, oversized, or cross-tenant records fail closed (dropped, never
176
+ recovered). Records are capped (32 default / 128 hard; oldest terminal
177
+ records evict beyond the cap; running records are never evicted).
178
+ - Errors: `ERR_PRISM_RECOVERY_UNSUPPORTED` / `_LIMIT` / `_OWNERSHIP` / `_FENCE`
179
+ / `_UNKNOWN` / `_UNTRUSTED` / `_TIMEOUT` (`ProcessRecoveryError`).
180
+
126
181
  ## Security and performance notes
127
182
 
128
183
  | Cap | Default | Hard |