@arnilo/prism 0.2.6 → 0.2.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +10 -2
- package/README.md +3 -2
- package/dist/agent-loops.js +4 -0
- package/dist/agent-run-lifecycle.js +2 -2
- package/dist/agent-session/helpers.d.ts +1 -1
- package/dist/agent-session/helpers.js +2 -1
- package/dist/agent-session/session.d.ts +1 -1
- package/dist/agent-session/session.js +28 -20
- package/dist/agent-session.d.ts +1 -1
- package/dist/agent-session.js +1 -1
- package/dist/agents.d.ts +1 -1
- package/dist/agents.js +1 -1
- package/dist/contracts-core/agent.d.ts +5 -5
- package/dist/contracts-core/agent.js +2 -0
- package/dist/contracts-core/extensions.d.ts +3 -3
- package/dist/contracts-core/extensions.js +2 -0
- package/dist/contracts-core/loop.d.ts +6 -1
- package/dist/contracts-core/session.d.ts +1 -1
- package/dist/contracts-core.d.ts +6 -6
- package/dist/contracts-core.js +6 -6
- package/dist/contracts-protocol.d.ts +8 -2
- package/dist/contracts.d.ts +1 -1
- package/dist/contracts.js +1 -1
- package/dist/field-policy.d.ts +119 -0
- package/dist/field-policy.js +418 -0
- package/dist/index.d.ts +6 -4
- package/dist/index.js +5 -4
- package/dist/input.js +1 -1
- package/dist/redaction.d.ts +6 -5
- package/dist/redaction.js +18 -10
- package/dist/tools.d.ts +1 -1
- package/dist/tools.js +1 -1
- package/docs/0.1.0-readiness.md +7 -7
- package/docs/acp-agent.md +78 -0
- package/docs/acp.md +21 -10
- package/docs/ag-ui.md +1 -1
- package/docs/agent-definitions.md +1 -1
- package/docs/agent-events.md +2 -2
- package/docs/agent-loops.md +2 -2
- package/docs/audit-export.md +151 -0
- package/docs/coding-agent-tools.md +3 -1
- package/docs/coding-security.md +6 -0
- package/docs/data-classification.md +82 -0
- package/docs/disaster-recovery.md +71 -0
- package/docs/enterprise-postgres-state.md +60 -7
- package/docs/evaluations.md +45 -0
- package/docs/host-security.md +4 -0
- package/docs/index.md +11 -5
- package/docs/migration.md +27 -0
- package/docs/operations.md +104 -0
- package/docs/policy-and-audit.md +34 -0
- package/docs/release-0.2.7-evidence.md +514 -0
- package/docs/release-and-install.md +55 -8
- package/docs/structured-output.md +10 -10
- package/docs/workflows.md +48 -3
- package/package.json +4 -3
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# Spawnable ACP agent (`@arnilo/prism-acp-agent`)
|
|
2
|
+
|
|
3
|
+
New in 0.2.8 (plan 028 Task 10 / adoption F3). A thin binary that serves [`createPrismAcpAgent`](acp.md) over stdio from a config file — the wiring you would otherwise copy out of [`examples/acp-coding-host.ts`](../examples/acp-coding-host.ts) into every host.
|
|
4
|
+
|
|
5
|
+
## Running
|
|
6
|
+
|
|
7
|
+
```sh
|
|
8
|
+
npx prism-acp-agent [--config prism-acp-agent.json]
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
The agent speaks ACP v1 as newline-delimited JSON on `stdin`/`stdout` (SDK `ndJsonStream` adapter over `Readable.toWeb(process.stdin)` / `Writable.toWeb(process.stdout)`). It serves until the client closes stdin; an `EPIPE` on stdout (client disconnected) is a normal shutdown.
|
|
12
|
+
|
|
13
|
+
```sh
|
|
14
|
+
# a config file must exist; missing/invalid config fails closed with a clear error and exit 1
|
|
15
|
+
printf '%s\n' '{"userId":"local","cwd":"/workspace"}' > prism-acp-agent.json
|
|
16
|
+
npx prism-acp-agent
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
## Config reference
|
|
20
|
+
|
|
21
|
+
The config file is the trust boundary: unknown keys are rejected (a typo cannot silently disable a security-relevant option), every value is shape-validated, and relative paths resolve against the config file's directory.
|
|
22
|
+
|
|
23
|
+
| Key | Required | Description |
|
|
24
|
+
| --- | --- | --- |
|
|
25
|
+
| `userId` | yes | Ownership user id for every session (single-local-user `authorize`). |
|
|
26
|
+
| `cwd` | yes | Workspace root the coding tools are bound to (must be an existing directory). Sessions always operate on this root — a client-supplied `cwd` never moves the tools. |
|
|
27
|
+
| `sessionStore` | no | `{ "type": "sqlite", "path": ".prism/sessions.db" }` or `{ "type": "memory" }` (default). SQLite persists sessions, runs, checkpoints, and leases (`createSqlitePersistence`). |
|
|
28
|
+
| `mcp.allow` | no | MCP allow-list. http/sse servers must have a `url` starting with an allow entry; stdio servers require the marker `"stdio"`. The UNSTABLE `acp` transport is never approved. |
|
|
29
|
+
| `modes` | no | Mode table `{ "modes": [{ "id", "name", "description?" }], "defaultModeId"? }`; ids unique, `defaultModeId` must name a mode. |
|
|
30
|
+
| `configOptions` | no | `{ "options": [{ "type": "boolean" \| "select", "id", "name", "defaultValue", ... }] }`; ids unique. Select options are advertised/settable per the B3 gate (see [acp.md](acp.md)). |
|
|
31
|
+
| `limits` | no | AG-UI/ACP caps passthrough (`AgUiLimitOptions`). |
|
|
32
|
+
|
|
33
|
+
Example:
|
|
34
|
+
|
|
35
|
+
```json
|
|
36
|
+
{
|
|
37
|
+
"userId": "local",
|
|
38
|
+
"cwd": ".",
|
|
39
|
+
"sessionStore": { "type": "sqlite", "path": ".prism/sessions.db" },
|
|
40
|
+
"mcp": { "allow": ["https://mcp.example.com"] },
|
|
41
|
+
"modes": { "modes": [{ "id": "edit", "name": "Edit" }], "defaultModeId": "edit" },
|
|
42
|
+
"configOptions": [{ "type": "boolean", "id": "verbose", "name": "Verbose", "defaultValue": false }]
|
|
43
|
+
}
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
## What it wires
|
|
47
|
+
|
|
48
|
+
The binary is pure wiring (~200 lines) — no protocol code lives here. It builds:
|
|
49
|
+
|
|
50
|
+
- `authorize` — single local user; every inbound call is scoped by session id.
|
|
51
|
+
- `sessionFactory` — real Prism sessions over `createAgent` with the nine coding tools (`createCodingTools(config.cwd)`), durable `runState` (`interruptBeforeTool`, checkpoints), ownership-scoped to `userId`.
|
|
52
|
+
- `lifecycle` — `createAgentRunLifecycle` over the same checkpoint store, so approvals suspend/resume durably.
|
|
53
|
+
- `mcp` — allow-list `select` gate with http/sse transports.
|
|
54
|
+
- `modes` / `configOptions` — from config.
|
|
55
|
+
- Provider — **mock by default** (full lifecycle, no tokens). Wire a real provider programmatically:
|
|
56
|
+
|
|
57
|
+
```ts
|
|
58
|
+
import { createSpawnableAgent, loadConfig } from "@arnilo/prism-acp-agent";
|
|
59
|
+
import { createOpenAIResponsesProvider } from "@arnilo/prism-provider-openai";
|
|
60
|
+
|
|
61
|
+
const agent = createSpawnableAgent({
|
|
62
|
+
config: loadConfig("prism-acp-agent.json"),
|
|
63
|
+
provider: createOpenAIResponsesProvider({ apiKey: process.env.OPENAI_API_KEY }),
|
|
64
|
+
});
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
## Library surface
|
|
68
|
+
|
|
69
|
+
- `loadConfig(path)` / `parseConfig(text, baseDir)` — read + validate; throw `ConfigError` (code `PRISM_ACP_AGENT_CONFIG`) with a clear message.
|
|
70
|
+
- `createSpawnableAgent({ config, provider? })` — build the ACP `AgentApp`.
|
|
71
|
+
- `selectMcpServers(allow, servers)` — the allow-list gate, exported for reuse in custom hosts.
|
|
72
|
+
|
|
73
|
+
## Security posture
|
|
74
|
+
|
|
75
|
+
- Config file = trust boundary: validated shape, no arbitrary code execution.
|
|
76
|
+
- MCP servers only from the allow-list; the UNSTABLE `acp` transport is never bridged.
|
|
77
|
+
- Coding tools are bound to `config.cwd` only; session ownership is fixed to `userId`.
|
|
78
|
+
- Session store paths are resolved against the config directory and fail closed on invalid config.
|
package/docs/acp.md
CHANGED
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
- `createPrismAcpAgent(options)` — serves ACP as an **agent**: an editor/AI client connects through the SDK transport and drives host-owned Prism sessions with `session/new`, `session/load`, `session/resume`, `session/prompt`, `session/set_mode`, `session/set_config_option`, `session/list`, `session/delete`, `session/close`, and `session/cancel`. The agent is a thin protocol adapter: every capability, decision, and byte cap is wired from host seams, and there is **no second policy engine** on the agent side.
|
|
8
8
|
- `createAcpEventMapper(options)` — maps a Prism `AgentEvent` stream (or `CoWorkEvent`) to ACP `SessionUpdate`s for hosts that stream through their own transport.
|
|
9
9
|
|
|
10
|
-
The adapter builds on the Phase 8/9 shared machinery: redacted event projection (`AgUiProjection`), the durable pending-decision batch model (`
|
|
10
|
+
The adapter builds on the Phase 8/9 shared machinery: redacted event projection (`AgUiProjection`), the durable pending-decision batch model (keyed by the permission optionIds `allow-once` / `allow-for-run` / `reject-once` / `reject-for-run`), `AgentRunLifecycle` resume, `CodingLifecycleEvent` emission, and the AG-UI/ACP package caps. It never ships experimental ACP v2 or UNSTABLE fields (`providers`, `nes`, `positionEncoding`, `sessionCapabilities.fork`, `mcpCapabilities.acp/auth`, `elicitation` is consumed client-side only and never advertised; the UNSTABLE `plan` surface is consumed client-side only — `plan_update`/`plan_removed` are emitted solely to clients that advertised `ClientCapabilities.plan`, F5).
|
|
11
11
|
|
|
12
12
|
## When to use it
|
|
13
13
|
|
|
@@ -24,11 +24,12 @@ Do **not** use it when the host needs a browser/TUI Web endpoint (use [AG-UI](ag
|
|
|
24
24
|
| `authorize` | `(input) => AcpAuthorization \| Promise` | **Required.** Ownership/identity gate for every inbound call, scoped by `sessionId`; unknown sessions fail. |
|
|
25
25
|
| `sessionFactory` | `(input) => AcpSessionBinding \| Promise` | **Required.** Builds the Prism `AgentSession` for `session/new`. Input carries `authorization`, `cwd`, `additionalDirectories`, `mcpServers` (policy-checked), `signal`, optional pre-generated `sessionId`, and `coding` (built client fs/terminal adapters when the client advertised them). |
|
|
26
26
|
| `lifecycle` | `AgentRunLifecycle` | **Required.** `status`/`resume`/`resumeStream` for durable `session/load` and `session/resume`. |
|
|
27
|
-
| `sessions?` | `AcpSessionStoreSeams` | `load` (advertises `sessionCapabilities.loadSession`), `list` (list), `delete` (delete), `resume` (resume), `additionalDirectories` (policy narrowing of `additionalDirectories`). `close` is always advertised. |
|
|
27
|
+
| `sessions?` | `AcpSessionStoreSeams` | `load` (advertises `sessionCapabilities.loadSession`), `list` (list), `delete` (delete), `resume` (resume), `additionalDirectories` (policy narrowing of `additionalDirectories`), `transcript` (F2: replay source for `session/load`/`session/resume`), `title` (F6: host-owned session titles — see below). `close` is always advertised. |
|
|
28
28
|
| `mcp?` | `AcpMcpSeams` | `transports: ("http" \| "sse")[]` and `select({ servers, signal })` — **required** for any client-supplied MCP server; select must approve before the bridge connects. Advertises `mcpCapabilities.http`/`sse` per transport. |
|
|
29
29
|
| `modes?` | `{ modes: AcpSessionMode[], defaultModeId? }` | `AcpSessionMode { id, name, description?, apply? }`; `apply({ sessionId?, fromModeId?, modeId, signal })` is the host hook run on switch. Advertised in `SessionModeState` on new/load/resume; enables `session/set_mode`. |
|
|
30
|
-
| `configOptions?` | `{ options: AcpConfigOption[], onChange? }` | `boolean`/`select` options with `defaultValue`; enables `session/set_config_option` (requires the client to advertise `session.configOptions.boolean`). |
|
|
31
|
-
| `capabilities?` | `AcpCapabilitiesOptions` | `prompt.media`/`prompt.embedded` policy seams, re-checked **live at prompt time**; presence advertises `promptCapabilities.image`/`audio`/`embeddedContext`. |
|
|
30
|
+
| `configOptions?` | `{ options: AcpConfigOption[], onChange? }` | `boolean`/`select` options with `defaultValue`; enables `session/set_config_option` (requires the client to advertise `session.configOptions.boolean`). **B3:** only `boolean` options are advertised in `session/new`/`load`/`resume` responses and `config_option_update`; `select` options are never settable — `set_config_option` on one fails with `ERR_PRISM_ACP_CAPABILITY` until the ACP spec defines a select capability. |
|
|
31
|
+
| `capabilities?` | `AcpCapabilitiesOptions` | `prompt.media`/`prompt.embedded` policy seams, re-checked **live at prompt time**; presence advertises `promptCapabilities.image`/`audio`/`embeddedContext`. `usage.contextWindow({ model, signal })` reports the model's context window in tokens; `usage_update` is emitted only when it returns a positive finite number — absent/undefined/throw ⇒ the update is omitted (never `size = used`). |
|
|
32
|
+
| `commands?` | `AcpCommandsSeam` | F9: `{ list({ sessionId, signal }) => AcpCommand[] }`. Presence emits `available_commands_update` on `session/new`, `session/load`, and `session/resume`. Absent seam ⇒ no update. |
|
|
32
33
|
| `coding?` | `AcpCodingSeams` | `filesystem(client, sessionId)` / `processes(client, sessionId)` factories building `AcpClientFilesystem` / `AcpClientTerminals` over client methods; `lifecycle?: CodingLifecycleEmitter` subscribes `CodingLifecycleEvent`s into ACP updates. |
|
|
33
34
|
| `name?` | `string` | `agentInfo.name` (default `"Prism"`). |
|
|
34
35
|
|
|
@@ -43,20 +44,26 @@ In-stream `SessionUpdate`s:
|
|
|
43
44
|
| Prism event | ACP update |
|
|
44
45
|
|---|---|
|
|
45
46
|
| Assistant text | `agent_message_chunk` |
|
|
46
|
-
|
|
|
47
|
-
| Tool
|
|
48
|
-
|
|
|
49
|
-
|
|
|
47
|
+
| Assistant thinking | `agent_thought_chunk` (same `messageId` scheme as text; through the shared redactor and byte caps) |
|
|
48
|
+
| Tool lifecycle | `tool_call` / `tool_call_update` (title/status/content) — `tool_call.kind` comes from the session's tool registry `kind` metadata when present (B4), else the name heuristic |
|
|
49
|
+
| Tool result, projected | `tool_call_update` with `locations` (≤ `acpLocationsPerUpdate`) and/or a `diff` block (≤ `acpDiffBytes`) — only from `AgUiProjection.toolLocations`/`toolDiff` allow-lists, at `finish()`. `toolResult` may return a string (text content) or `{ type: "image", data, mimeType }` (F8) — the mapper wraps the image as `{ type: "content", content: { type: "image", data, mimeType } }` and drops payloads over `acpImageBytes` (never truncated). Opt-in turnkey: `createCodingToolProjection()` (F7) recognizes first-party `edit` (path + unified patch as `newText`, `firstChangedLine` location) and `write` (path location only — result has no file body) results; default remains deny-by-default. |
|
|
50
|
+
| Provider usage | `usage_update` (only when the `capabilities.usage.contextWindow` seam reports a valid window — absent/undefined/throw ⇒ the update is omitted, never `size = used`) |
|
|
51
|
+
| Run-level failure | No transcript chunk — the `session/prompt` request rejects with `ERR_PRISM_ACP_RUN` (redacted, byte-capped message). Retryable provider-turn failures stay silent and may recover; only a terminal `error` event fails the request. |
|
|
52
|
+
| Run stop reason | `session/prompt` returns the SDK `StopReason` (F4): `cancelled` when the run was aborted, `max_turn_requests` for the tool-round ceiling (`finishReason: "turn_limit"`), `max_tokens` for `"token_limit"`, `refusal` for `"refusal"`, else `end_turn`. The generic `finishReason` field is set on `agent_finished` by loop strategies (single-shot records `turn_limit` at the `maxToolRounds` ceiling); `token_limit`/`refusal` have no core producer yet — the mapping is ready. |
|
|
53
|
+
| Durable suspension | `session/request_permission` with the four options `allow-once`→`allow_once`, `allow-for-run`→`allow_always`, `reject-once`→`reject_once`, `reject-for-run`→`reject_always` (optionId→SDK kind, as emitted by `permission-elicit.ts`); cancel, unknown options, and request failure deny. Sticky decisions expire at run end. |
|
|
50
54
|
| Elicitation suspension (all-elicitation batch + client advertised `elicitation`) | `elicitation/create` (form mode, bounded schema, redacted reason); `accept` → `allow_once` with the typed payload as `RunDecision.elicitation`, `decline`/`cancel` → `reject_once`. Otherwise falls back to the shared four-option permission path. |
|
|
51
55
|
| `file_changed` lifecycle | `tool_call_update` with `locations: [{ path }]` (needs a `toolCallId`; diff only from `fileDiff` allow-list, capped + redacted) |
|
|
52
56
|
| `worktree_changed` / process events | Projection-gated `agent_message_chunk` (deny-by-default: no `lifecycle` projection hook = no update) |
|
|
53
57
|
| `permission_denied` lifecycle | `tool_call_update` status `failed` (never raw args; synthesized id `prism:denied:<approvalId>` when no `toolCallId`) |
|
|
54
58
|
| `configuration_changed` lifecycle | `config_option_update` with the full current set, per streaming session |
|
|
59
|
+
| Plan lifecycle (F5, UNSTABLE-gated) | `plan_changed` → `plan_update` with `plan: { type: "items", planId = planPath, entries: [{ content, priority: "medium", status }] }` — the complete entry list per update (client replaces its plan wholesale); `plan_removed` → `plan_removed` with `planId = planPath`. Emitted only when the client advertised `ClientCapabilities.plan`; mapper stays capability-agnostic (gate in the agent wiring). Entries come from `writeCodingPlanFile`'s `onEvent` (parsed via `parseCodingPlanTodos`) or host-emitted through their `CodingLifecycleEmitter`; text passes the shared redactor and byte caps. |
|
|
60
|
+
| Session title (F6) | `sessions.title({ sessionId, prompt, signal })` resolves on `session/prompt`; a defined value differing from the last emitted title produces `session_info_update` with `{ sessionUpdate: "session_info_update", title }`. Best-effort: `undefined` or a throw means no title and no update (requests never fail on titles); the host owns title storage. Titles pass the shared redactor and are truncated at `maxTextBytes`/`maxEventBytes`. |
|
|
61
|
+
| Slash commands (F9) | `commands.list({ sessionId, signal })` on `session/new`/`load`/`resume` produces `available_commands_update` with `{ name, description, input?: { hint } }` (SDK `AvailableCommand`; description is required). Names/descriptions/hints pass the shared redactor and `maxTextBytes`; the list is sliced at `acpCommandsPerUpdate`. Best-effort: a throw or non-array omits the update (session start never fails on commands). |
|
|
55
62
|
| Session mode/config switch | `current_mode_update` / `config_option_update` |
|
|
56
63
|
|
|
57
|
-
Frozen caps (default / hard, from the Phase 10 freeze manifest): sessions 32/128, additional directories 8/32 (path 4 KiB/16 KiB), MCP servers 8/32 (config 16 KiB/256 KiB, header values 4 KiB/64 KiB), modes 16/64, config options 16/64, list page 20/100, diff bytes 64 KiB/1 MiB, locations per update 32/128, prompt media parts 16/64 and media bytes 64 KiB/1 MiB (shared AG-UI caps), terminal output chunks 51200 B/1 MiB (Phase 9 `process.outputChunkBytes`), stream events/bytes per AG-UI budgets. `session/load`/`session/resume` of a still-registered session rejects with `ERR_PRISM_ACP_INPUT` ("ACP session already exists"); model reconnect as resume of a pre-seeded stored session.
|
|
64
|
+
Frozen caps (default / hard, from the Phase 10 freeze manifest): sessions 32/128, additional directories 8/32 (path 4 KiB/16 KiB), MCP servers 8/32 (config 16 KiB/256 KiB, header values 4 KiB/64 KiB), modes 16/64, config options 16/64, list page 20/100, diff bytes 64 KiB/1 MiB, locations per update 32/128, projected tool-result images `acpImageBytes` 256 KiB/1 MiB (F8; oversize dropped), slash commands `acpCommandsPerUpdate` 32/128 (F9), prompt media parts 16/64 and media bytes 64 KiB/1 MiB (shared AG-UI caps), terminal output chunks 51200 B/1 MiB (Phase 9 `process.outputChunkBytes`), stream events/bytes per AG-UI budgets. `session/load`/`session/resume` of a still-registered session rejects with `ERR_PRISM_ACP_INPUT` ("ACP session already exists"); model reconnect as resume of a pre-seeded stored session.
|
|
58
65
|
|
|
59
|
-
Errors surface as `AcpError` with codes `ERR_PRISM_ACP_INPUT` (malformed), `ERR_PRISM_ACP_LIMIT` (caps), `ERR_PRISM_ACP_POLICY` (host denied), `ERR_PRISM_ACP_CAPABILITY` (not advertised), `ERR_PRISM_ACP_MCP` (MCP bridging). Over the wire they become JSON-RPC `-32603` with the message in `data.details` (SDK behavior).
|
|
66
|
+
Errors surface as `AcpError` with codes `ERR_PRISM_ACP_INPUT` (malformed), `ERR_PRISM_ACP_LIMIT` (caps), `ERR_PRISM_ACP_POLICY` (host denied), `ERR_PRISM_ACP_CAPABILITY` (not advertised), `ERR_PRISM_ACP_MCP` (MCP bridging), `ERR_PRISM_ACP_RUN` (run-level failure — the `session/prompt` request rejects instead of emitting a fake `Agent error:` chunk). Over the wire they become JSON-RPC `-32603` with the message in `data.details` (SDK behavior).
|
|
60
67
|
|
|
61
68
|
## Request/response example
|
|
62
69
|
|
|
@@ -104,6 +111,7 @@ const agent = createPrismAcpAgent({
|
|
|
104
111
|
## Extension and configuration notes
|
|
105
112
|
|
|
106
113
|
- **Seam = capability.** Wiring `sessions.load` advertises `loadSession`; removing it withdraws the method. There is no separate capability flag to keep in sync — the freeze manifest's advertise-when matrix is enforced by construction and asserted by `scripts/phase10-conformance.test.mjs`.
|
|
114
|
+
- **Transcript replay (F2).** When `sessions.transcript` is wired, `session/load` and `session/resume` replay `user_message_chunk`/`agent_message_chunk` text chunks (from `SessionEntry`s with `kind: "message"` and a user/assistant role, text blocks only) before returning `sessionState`. Each chunk passes the shared redactor and is truncated at `maxTextBytes`; replay stops at `maxReplayEvents` chunks and counts against the stream event/byte caps (an oversized transcript fails the load/resume request closed). Absent seam = no replay, behavior unchanged.
|
|
107
115
|
- **Client fs/terminal are adapters, not a second implementation.** `AcpClientFilesystem` / `AcpClientTerminals` wrap the client's `fs/*` and `terminal/*` methods behind the Phase 9 `ProcessSession`-flavored interfaces; the agent pre-generates the session id so terminal requests can carry it. Host repo operations remain default when the client fs is absent.
|
|
108
116
|
- **Modes and config options are a pure host overlay.** The agent stores only a thin per-session registry; `apply`/`onChange` hooks narrow the host's own behavior. Mode switches can narrow or host-authorized widen — never a parallel policy evaluator, never a client-enabled tool.
|
|
109
117
|
- **Lifecycle wiring.** Pass your `createCodingLifecycleEmitter()` as `coding.lifecycle`; `file_changed` etc. then flow to streaming sessions. `configuration_changed` broadcasts `config_option_update` (agent-message fallback if the SDK rejects the kind).
|
|
@@ -141,6 +149,9 @@ const agent = createPrismAcpAgent({
|
|
|
141
149
|
|
|
142
150
|
- **Untrusted client input.** Client-supplied paths, `additionalDirectories`, MCP server configs, terminal env/args, and media are validated at the boundary: count/byte caps, ownership-scoped sessions, path policy via the `sessions.additionalDirectories` seam, MCP servers only through host `select` (never auto-connected), UNSTABLE `acp` transport always rejected.
|
|
143
151
|
- **Deny-closed by default.** Unknown mode ids, unadvertised methods, unprojected lifecycle events, oversize diffs/locations/media, thrown projection hooks, and failed elicitation all fail closed. Raw tool arguments/results are never sent unless a projection allow-list says otherwise.
|
|
152
|
+
- **Slash commands (F9).** `commands.list` is a host-owned slash-command list (not derived from the tool registry). The agent emits `available_commands_update` on session start (`session/new`, `session/load`, `session/resume`). Mid-session refresh is not in this release — re-list by starting a session. Names, descriptions, and input hints pass the shared redactor; the list is sliced at `acpCommandsPerUpdate`. Absent seam or a thrown list ⇒ no update.
|
|
153
|
+
- **Projected images (F8).** `AgUiProjection.toolResult` may return `{ type: "image", data, mimeType }` (return-type widening — existing string returns stay valid). The mapper emits `{ type: "content", content: { type: "image", data, mimeType } }` (SDK v1 `ToolCallContent` has no top-level image variant). `data` is the host-supplied base64; it is not redacted and not truncated — payloads over `acpImageBytes` are dropped. Default (no hook / non-image return) emits no image.
|
|
154
|
+
- **Coding-tool projection (F7).** `createCodingToolProjection({ maxDiffBytes? })` is an opt-in `AgUiProjection` for first-party `@arnilo/prism-coding-agent` `edit`/`write` results: `edit` → `toolDiff` (`path` + unified `patch` as `newText`) and `toolLocations` (`path` + `firstChangedLine`); `write` → `toolLocations` (`path` only — write metadata has no file body, so no honest diff; use `file_changed` + `fileDiff` when bodies are needed). Pass as `projection: createCodingToolProjection()` on the agent/mapper. Mapper still redacts and enforces `acpDiffBytes` / `acpLocationsPerUpdate`; optional `maxDiffBytes` pre-truncates the patch so a slightly-oversize edit is shortened instead of dropped. Without the factory, behavior is unchanged (deny-by-default).
|
|
144
155
|
- **No secrets.** Updates carry no raw file bodies, terminal output is capped by the Phase 9 chunk budget, and the shared redactor is applied before anything leaves the host. `permission_denied` never includes raw args.
|
|
145
156
|
- **Performance.** The adapter is O(1) per update with no unbounded buffering; p95 targets (fs round trip 250 ms, mode switch 250 ms, terminal chunk ack 1000 ms, prompt first update 2000 ms, prompt end 30 s) are recorded by `scripts/benchmark-0.0.27.mjs` and gated in `scripts/budgets.json` `phase10`.
|
|
146
157
|
|
package/docs/ag-ui.md
CHANGED
|
@@ -202,7 +202,7 @@ const renderer = createA2UiRenderer({
|
|
|
202
202
|
const surface = await renderer.surface("chat"); // detached DOM node, kept in sync
|
|
203
203
|
```
|
|
204
204
|
|
|
205
|
-
The core is a DOM-free state machine (`reduceA2UiOps`): operations become a surface/component model (adjacency list with `id`/`component`/flat props, A2UI v0.9 JSON-Pointer data model, `deleteSurface`); a thin binding layer renders the model through catalog component renderers (framework-free `(props, ctx, dom) => node` functions). The core is also exported as values from the subpath entry (`A2UiSurfaceState`, `reduceA2UiOps`, `readA2UiBatch`, `resolvePointer`, `A2UI_VERSION`, 0.0.27,
|
|
205
|
+
The core is a DOM-free state machine (`reduceA2UiOps`): operations become a surface/component model (adjacency list with `id`/`component`/flat props, A2UI v0.9 JSON-Pointer data model, `deleteSurface`); a thin binding layer renders the model through catalog component renderers (framework-free `(props, ctx, dom) => node` functions). The core is also exported as values from the subpath entry (`A2UiSurfaceState`, `reduceA2UiOps`, `readA2UiBatch`, `resolvePointer`, `A2UI_VERSION`, 0.0.27, host FR) so framework hosts can drive the validated surface state machine and own the view layer; behavior and frozen caps are unchanged. Snapshots replace a surface's model (streaming mode sends cumulative ops); RFC 6902 deltas append. The same frozen caps as the server painter are enforced client-side: 64/512 ops per message, 64 KiB/1 MiB per op, 16/64 surfaces per run, depth 32/64. Invalid or oversized ops drop closed with one bounded `prism.a2ui.error` event (host logging via `onError`); unknown catalog components render an explicit placeholder. The renderer never executes remote HTML: only `createElement`/`createTextNode`/`appendChild`, no HTML-string assignment, no dynamic code evaluation. Data bindings `{"path": "/pointer"}` resolve against the per-surface data model; `deleteSurface` detaches content. The main `@arnilo/prism-ag-ui` entry stays runtime-agnostic — DOM code lives only behind the `renderer` subpath (the root entry re-exports renderer types only, no values). Hosts embedding it should follow the MCP Apps CSP/sandbox guidance (`docs/ag-ui-adoption.md`) for iframe/worker placement.
|
|
206
206
|
|
|
207
207
|
## Security and performance notes
|
|
208
208
|
|
|
@@ -15,7 +15,7 @@ A third helper, `discoverAgentBundles(options)` (same Node subpath), scans an ap
|
|
|
15
15
|
|
|
16
16
|
Use `resolveAgentDefinition` when an app already holds a `AgentDefinition` (from an extension, a manifest, or hand-written config) and wants to turn it into an `Agent` against its registries — without wiring every field by hand.
|
|
17
17
|
|
|
18
|
-
Use `discoverAgentBundles` + `resolveAgentBundle` when a host app keeps per-agent bundles on disk under an app-controlled config root (for example
|
|
18
|
+
Use `discoverAgentBundles` + `resolveAgentBundle` when a host app keeps per-agent bundles on disk under an app-controlled config root (for example `<appRoot>/extensions/prism/agents/<agentName>/AGENT.md`) and wants to honor them as first-class agents. The bundle layout is host-owned: Prism never picks the config root, never touches the user's home directory, and never auto-runs resolution — the host calls `discoverAgentBundles` and then `resolveAgentBundle` explicitly.
|
|
19
19
|
|
|
20
20
|
Do not use the bundle loader to discover providers — provider/model packages stay config/package-driven (Phase 24; see [Provider packages](provider-packages.md)). Do not use it to auto-activate undeclared tools or skills: omitted `tools` / `skills` means no active capabilities by default; named bundle entries are explicit activation, and runtime skill selection can narrow further with `RunOptions.activeSkills`.
|
|
21
21
|
|
package/docs/agent-events.md
CHANGED
|
@@ -88,7 +88,7 @@ Agent / turn / message events:
|
|
|
88
88
|
| Variant | Fields |
|
|
89
89
|
| --- | --- |
|
|
90
90
|
| `agent_started` | `sessionId`, `runId` |
|
|
91
|
-
| `agent_finished` | `sessionId`, `runId`, `usage?: Usage` (aggregate of all usage-bearing provider turns) |
|
|
91
|
+
| `agent_finished` | `sessionId`, `runId`, `usage?: Usage` (aggregate of all usage-bearing provider turns), `finishReason?: "turn_limit" \| "token_limit" \| "refusal"` (why a limit/ceiling ended the run cleanly — F4; absent = natural end) |
|
|
92
92
|
| `agent_suspended` | `sessionId`, `runId`, redacted `interruption`, checkpoint `version`; no tool side effect has started. |
|
|
93
93
|
| `agent_resumed` | `sessionId`, `runId`, checkpoint `version`. |
|
|
94
94
|
| `agent_denied` | `sessionId`, `runId`, redacted `interruption`, checkpoint `version`; no tool side effect runs. |
|
|
@@ -242,4 +242,4 @@ for await (const event of session.stream("draft", { loop: { strategy: "generate-
|
|
|
242
242
|
- [Observability](observability.md): `ProviderTurnMetadata`; optional adapter builds one parented GenAI span tree from metadata-only lifecycle events and ignores message/progress deltas.
|
|
243
243
|
- [Tools](tools.md): `tool_execution_*` variants.
|
|
244
244
|
- [Compaction and retry policies](compaction-and-retry.md): `compaction_*` and `retry_scheduled` variants.
|
|
245
|
-
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional redacted mapping of this stream; durable replay is ledger-backed and at-least-once, never a live-subscriber substitute. [ACP coding-host interop](acp.md) additionally maps `CodingLifecycleEvent`s from `@arnilo/prism-coding-agent` (`file_changed`, `worktree_changed`, `permission_denied`, `configuration_changed`; process events reuse `CodingProcessEvent`) into ACP session updates — locations/diff blocks only through projection allow-lists, terminal chunks under `process.outputChunkBytes
|
|
245
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional redacted mapping of this stream; durable replay is ledger-backed and at-least-once, never a live-subscriber substitute. [ACP coding-host interop](acp.md) additionally maps `CodingLifecycleEvent`s from `@arnilo/prism-coding-agent` (`file_changed`, `worktree_changed`, `permission_denied`, `configuration_changed`, `plan_changed`, `plan_removed`; process events reuse `CodingProcessEvent`) into ACP session updates — locations/diff blocks only through projection allow-lists, terminal chunks under `process.outputChunkBytes`, plan updates only to clients that advertised the UNSTABLE `plan` capability.
|
package/docs/agent-loops.md
CHANGED
|
@@ -232,7 +232,7 @@ await session.run(input, { loop: twoShotLoop });
|
|
|
232
232
|
- `generateValidateReviseLoop` makes at most `1 + maxRevisions + maxToolRounds` provider turns when bounded tools are enabled (otherwise `maxRevisions + 1`); it cannot loop forever. Each revision costs one provider turn plus one store append.
|
|
233
233
|
- Bounded artifact tool calls run sequentially through `dispatchToolCall` (permission + validation + execute); their assistant call and result are persisted before the next provider request. `singleShotLoop` retains its bounded parallel worker pool and original call-order transcript behavior.
|
|
234
234
|
- The loop is a plain object/factory; no class hierarchy, no background work, no extra dependencies. `LoopContext` is a single object literal of bound arrows built once per run.
|
|
235
|
-
- The
|
|
235
|
+
- The host-domain-free boundary is guarded by tests: `src/` imports no host-domain package, and the `Artifact*`/`AgentLoop*`/`LoopContext` contracts contain no `workflow`/`node`/`step` field names. Hosts supply their own schema; no host domain type is imported by `src/`.
|
|
236
236
|
|
|
237
237
|
## Guardrails
|
|
238
238
|
|
|
@@ -241,7 +241,7 @@ Built-in loops and custom loops that use `LoopContext.generate()` / `LoopContext
|
|
|
241
241
|
## Related APIs
|
|
242
242
|
- [Agent/session runtime](agent-session-runtime.md): `RuntimeAgentSession.run()` builds the `LoopContext` and delegates to the resolved loop.
|
|
243
243
|
- [Agent events](agent-events.md): the `artifact_*` event variants and ordering emitted by `generateValidateReviseLoop`.
|
|
244
|
-
- [Structured output](structured-output.md): the `ArtifactParser<T>`/`ArtifactValidator<T>`/`ArtifactRepairer<T>` seam (host-defined `T`, Prism never instantiates it) and a
|
|
244
|
+
- [Structured output](structured-output.md): the `ArtifactParser<T>`/`ArtifactValidator<T>`/`ArtifactRepairer<T>` seam (host-defined `T`, Prism never instantiates it) and a host schema→`ArtifactValidation` mapping example.
|
|
245
245
|
- [Public contracts](public-contracts.md): `AgentLoopStrategy`, `AgentLoopOptions`, `LoopContext`, `ProviderTurnResult`, and the `Artifact*` contracts.
|
|
246
246
|
- [Input and prompt assembly](input-and-prompt-assembly.md): `assembleProviderInput()`, the primitive behind `LoopContext.assemble`.
|
|
247
247
|
- [Tools](tools.md): `dispatchToolCall()`, the primitive behind `LoopContext.dispatchToolCall`.
|
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
# Signed, hash-chained audit export
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-policy` exports tenant-scoped audit records as signed, hash-chained
|
|
6
|
+
batches: each record envelope is canonicalized (RFC 8785 semantics), hashed with
|
|
7
|
+
SHA-256 including the prior digest, so records form a tamper-evident chain. A
|
|
8
|
+
batch of chained records is wrapped in a manifest that a host-provided
|
|
9
|
+
`AuditSigner` signs; the signed artifact is written to an immutable WORM sink
|
|
10
|
+
which must acknowledge the exact artifact digest before the export cursor
|
|
11
|
+
advances, and is mirrored to an optional SIEM sink with replayable status. An
|
|
12
|
+
independent verifier (`verifyAuditBatch` or `scripts/verify-audit-export.mjs`)
|
|
13
|
+
re-derives everything from the artifact bytes and a public key — no ledger,
|
|
14
|
+
sink, or private key needed.
|
|
15
|
+
|
|
16
|
+
## When to use it
|
|
17
|
+
|
|
18
|
+
- You keep an append-only policy/audit ledger and must prove to auditors that
|
|
19
|
+
exported records were not reordered, edited, inserted, or truncated after the
|
|
20
|
+
fact.
|
|
21
|
+
- You need a durable WORM copy with an integrity receipt and a replayable SIEM
|
|
22
|
+
mirror, without embedding cloud SDKs or key storage in Prism.
|
|
23
|
+
- You must export only the redacted bytes an external verifier will actually
|
|
24
|
+
see, with redaction provenance and legal-hold provenance preserved.
|
|
25
|
+
|
|
26
|
+
Do not use it for exactly-once delivery claims: like the messaging outbox,
|
|
27
|
+
exports are at-least-once, and the exporter never proves that a record exists —
|
|
28
|
+
it proves that what was exported is exactly what the verifier can reproduce.
|
|
29
|
+
|
|
30
|
+
## Inputs / request
|
|
31
|
+
|
|
32
|
+
- `createAuditExporter({ source, cursorStore, signer, wormSink, siemSink?, redact? })`
|
|
33
|
+
- `source` — tenant-scoped, stable-order `AuditPageSource`; page cursors are
|
|
34
|
+
one-shot tokens (re-reading a cursor that already served its final page
|
|
35
|
+
yields an empty page).
|
|
36
|
+
- `cursorStore` — CAS-versioned `AuditCursorStore`; the exporter advances the
|
|
37
|
+
cursor only on a matched version after the WORM acknowledgement.
|
|
38
|
+
- `signer` — host `AuditSigner` (`sign(bytes) -> Uint8Array`, optional
|
|
39
|
+
`keyId`/`algorithm`); Prism never accepts raw private keys.
|
|
40
|
+
- `wormSink` — required immutable sink returning `{ batchId, digest }`; the
|
|
41
|
+
exporter refuses to advance unless both match.
|
|
42
|
+
- `siemSink` — optional replayable mirror.
|
|
43
|
+
- `redact` — optional `AuditRedactionPolicy` applied before hashing.
|
|
44
|
+
- `exportNext({ tenantId, maxRecords?, maxBytes?, signal? })` — processes one
|
|
45
|
+
page; returns `{ batchId, firstSequence, lastSequence, recordCount,
|
|
46
|
+
wormAcked, siemStatus: "disabled" | "sent" | "pending", nextDigest,
|
|
47
|
+
artifactBytes }`.
|
|
48
|
+
- `verifyAuditBatch({ artifactBytes, publicKey, expectedTenantId,
|
|
49
|
+
previousDigest?, expectedFirstSequence?, expectedLastSequence? })`.
|
|
50
|
+
|
|
51
|
+
## Outputs / response / events
|
|
52
|
+
|
|
53
|
+
A batch artifact written to the WORM sink contains:
|
|
54
|
+
|
|
55
|
+
```json
|
|
56
|
+
{
|
|
57
|
+
"schemaVersion": 1,
|
|
58
|
+
"document": "{ ...canonical signed manifest as a single JSON string... }",
|
|
59
|
+
"signature": { "algorithm": "sha256", "keyId": "k1", "value": "<base64>" }
|
|
60
|
+
}
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
The embedded document is the manifest: `tenantId`, `batchId`, `algorithm`,
|
|
64
|
+
`firstSequence`, `lastSequence`, `previousDigest`, `nextDigest`, and
|
|
65
|
+
`records` — each record carrying `sequence`, `priorDigest`, `digest`,
|
|
66
|
+
optional `legalHold`, optional `redactions` (`{ path, reason }[]`), and the
|
|
67
|
+
canonical `record` payload. `verifyAuditBatch` returns `{ ok, errors, batch }`.
|
|
68
|
+
|
|
69
|
+
## Request/response example
|
|
70
|
+
|
|
71
|
+
```ts
|
|
72
|
+
import { createAuditExporter, createMemoryAuditCursorStore } from "@arnilo/prism-policy";
|
|
73
|
+
|
|
74
|
+
const exporter = createAuditExporter({
|
|
75
|
+
source, // host: tenant-scoped record pages
|
|
76
|
+
cursorStore: createMemoryAuditCursorStore(), // host durable store in production
|
|
77
|
+
signer, // host: HSM/KMS-backed signer
|
|
78
|
+
wormSink, // host: S3 object-lock / WORM bucket
|
|
79
|
+
siemSink, // host: SIEM or event stream
|
|
80
|
+
});
|
|
81
|
+
const result = await exporter.exportNext({ tenantId: "acme", maxRecords: 1000 });
|
|
82
|
+
// result.artifactBytes -> store on WORM; result.nextDigest -> pass as
|
|
83
|
+
// previousDigest when verifying the next batch.
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
## Implementation example
|
|
87
|
+
|
|
88
|
+
```ts
|
|
89
|
+
import { readFileSync } from "node:fs";
|
|
90
|
+
import { verifyAuditBatch } from "@arnilo/prism-policy";
|
|
91
|
+
|
|
92
|
+
const artifactBytes = new Uint8Array(readFileSync("./acme-000001.json"));
|
|
93
|
+
const verified = verifyAuditBatch({
|
|
94
|
+
artifactBytes,
|
|
95
|
+
publicKey: readFileSync("./audit-verification.pem", "utf8"),
|
|
96
|
+
expectedTenantId: "acme",
|
|
97
|
+
previousDigest: "0000000000000000000000000000000000000000000000000000000000000000",
|
|
98
|
+
});
|
|
99
|
+
if (!verified.ok) throw new Error(verified.errors.join("; "));
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
## Extension and configuration notes
|
|
103
|
+
|
|
104
|
+
- Key rotation: the artifact records the signer's `keyId`; verification fails
|
|
105
|
+
explicitly under a rotated key, so consumers select the key named by the
|
|
106
|
+
artifact's `signature.keyId`. Hosts own key lifecycle and the public-key
|
|
107
|
+
distribution path.
|
|
108
|
+
- Failed batches retain their one-page payload in exporter memory and replay
|
|
109
|
+
the same batch id on the next `exportNext`; WORM-failed or signer-failed
|
|
110
|
+
batches are never re-read from the source. A cursor CAS race after WORM
|
|
111
|
+
acknowledgment is surfaced as an explicit error and is not retried by this
|
|
112
|
+
exporter (the batch is already durable).
|
|
113
|
+
- SIEM failures do not fail the export: the batch reaches WORM, the cursor
|
|
114
|
+
advances, and a bounded `siemPending` list records what is not mirrored.
|
|
115
|
+
`retryPendingSiem` replays a pending batch when the host supplies its
|
|
116
|
+
artifact bytes (e.g. fetched from WORM) and verifies the digest matches.
|
|
117
|
+
- Legal holds: the exporter preserves a `legalHold` flag on envelope records
|
|
118
|
+
and never broadens tenant access; enforcing a hold is the host source's job.
|
|
119
|
+
Redaction (`redact`) strips values before hashing so the verifier sees
|
|
120
|
+
exactly the exported bytes; only `{ path, reason }` provenance survives.
|
|
121
|
+
- Caps are frozen: 1,000 records or 10 MiB per batch, at most 8 un-mirrored
|
|
122
|
+
SIEM batches retained.
|
|
123
|
+
|
|
124
|
+
## Security and performance notes
|
|
125
|
+
|
|
126
|
+
- Bytes are canonical JSON (RFC 8785 semantics: sorted keys, ECMAScript
|
|
127
|
+
shortest number round-trip, `-0` collapsed, lowercase control escapes);
|
|
128
|
+
non-finite numbers, BigInt, undefined, functions, symbols, and cyclic
|
|
129
|
+
values are rejected rather than coerced.
|
|
130
|
+
- Each record digest covers `schemaVersion`, `tenantId`, `sequence`,
|
|
131
|
+
`priorDigest`, `legalHold`, `redactions`, and the canonical record — a
|
|
132
|
+
cross-tenant record can never enter another tenant's chain, and the verifier
|
|
133
|
+
replays every envelope from the artifact's own bytes.
|
|
134
|
+
- The WORM acknowledgement must name the batch and match the artifact digest;
|
|
135
|
+
a lying or partial acknowledgement cannot falsely advance the cursor.
|
|
136
|
+
- Signer keys and raw values never enter logs, records, or artifacts.
|
|
137
|
+
- Performance: hashing and signing are linear in record bytes; batches are
|
|
138
|
+
bounded pages built without loading full history. Verification is also
|
|
139
|
+
linear and stateless. Prism does not certify compliance with NIST or any
|
|
140
|
+
SIEM/WORM vendor program; audit-export is the transport, hosts own the
|
|
141
|
+
custody chain and compliance posture.
|
|
142
|
+
|
|
143
|
+
## Related APIs
|
|
144
|
+
|
|
145
|
+
- `createPolicyDecisionStore` / `exportPolicyDecisions` — the ledger that
|
|
146
|
+
commonly backfills an `AuditPageSource`.
|
|
147
|
+
- `PersistenceLifecycleStore` — legal-hold and retention for lifecycle
|
|
148
|
+
records the source may consult.
|
|
149
|
+
- `createMemoryAuditCursorStore` — reference cursor store (replace with a
|
|
150
|
+
durable CAS store in production).
|
|
151
|
+
- `scripts/verify-audit-export.mjs` — standalone verifier CLI.
|
|
@@ -48,6 +48,8 @@ import { createCodingTools } from "@arnilo/prism-coding-agent";
|
|
|
48
48
|
const tools = createToolRegistry(createCodingTools(process.cwd()));
|
|
49
49
|
```
|
|
50
50
|
|
|
51
|
+
Every tool carries an explicit `kind` (`shell`→`execute`, `read`/`repo_list`→`read`, `write`/`edit`→`edit`, `repo_search`/`glob`→`search`, `delete`→`delete`, `move`→`move`) so ACP `tool_call` updates and other consumers can classify tools without name heuristics.
|
|
52
|
+
|
|
51
53
|
## When to use it
|
|
52
54
|
|
|
53
55
|
Use this package when a host wants ready-made coding tools for an agent, session, or run, registered explicitly into a `ToolRegistry` and dispatched through the normal Prism tool harness. The tools perform **real** shell and filesystem operations on the host — they are not mocked or sandboxed. Use the individual factories when you need per-tool options or custom operation backends; use the aggregators when you want the default set.
|
|
@@ -578,5 +580,5 @@ Every configurable value is a positive safe integer (context may be zero); Prism
|
|
|
578
580
|
- [Public contracts](public-contracts.md): `ToolDefinition`, `ToolResult`, `ToolExecutionContext`, `ContentBlock`, and `JsonObject` shapes.
|
|
579
581
|
- [Host security guide](host-security.md): fail-closed checklist for permission policies, tool validation, and trust boundaries that must gate these tools.
|
|
580
582
|
- [Tool conformance](tool-conformance.md): assertions for the tool-dispatch blocked-reason matrix these tools participate in.
|
|
581
|
-
- [ACP coding-host interop](acp.md): host editors drive these tools through stable ACP v1 — client fs/terminal adapters, `CodingLifecycleEvent` emission (`file_changed` etc. via the `onEvent` options), and permission/elicitation through the shared four-outcome decision model.
|
|
583
|
+
- [ACP coding-host interop](acp.md): host editors drive these tools through stable ACP v1 — client fs/terminal adapters, `CodingLifecycleEvent` emission (`file_changed` etc. via the `onEvent` options; `plan_changed` also fires from `writeCodingPlanFile`'s `onEvent`, F5), and permission/elicitation through the shared four-outcome decision model.
|
|
582
584
|
- [LLM compaction package](compaction-llm.md): optional `createCodingCompactionStrategy()` retains bounded paths, patch intent, checks, plan/todo state, blockers, and next verification—not complete diffs or raw command output.
|
package/docs/coding-security.md
CHANGED
|
@@ -210,6 +210,12 @@ The protected coding journey (0.2.6, plan 026 Task 7) exercises these boundaries
|
|
|
210
210
|
|
|
211
211
|
The egress proxy is a policy enforcer, not a firewall: it cannot stop a container whose Docker network reaches the internet directly. Egress attestation (`denyDirectEgress: true`) is a claim the host must make true by network topology; the adapter records it as evidence and fails closed when it is absent or malformed. The proxy performs no TLS interception, no DNS rebinding of its own beyond pinning, and no content filtering; audit records contain no secrets. Frozen caps: 32 concurrent connections (hard 256), 64 MiB request/response bytes (hard 1 GiB), 600 s transfer time (hard 1 h), 128 rules (hard 1,024), 5 redirect hops (hard 10).
|
|
212
212
|
|
|
213
|
+
## Windows hosts
|
|
214
|
+
|
|
215
|
+
`createNativeSandbox` is Linux-only. On any other platform (including Windows) it throws at creation and does **not** fall back to an unsandboxed process — egress denial cannot be enforced by construction without a network namespace. Do not catch that error and enable `shell` on the host; keep `shell` disabled, or run the agent inside a Docker container using `createDockerSandbox` and the documented [allow-list egress](#allow-list-egress-composition) policy (`network: none` or an attested custom network). Host-mode tools (`read`/`write`/`edit` under `workspaceMode: "host"`) remain available; they never claim containment.
|
|
216
|
+
|
|
217
|
+
A native Windows backend (Job objects / AppContainer) is tracked, not scheduled. Until one exists, Windows hosts that need isolation use Docker. Do not weaken the deny-by-default posture to compensate.
|
|
218
|
+
|
|
213
219
|
## Related APIs
|
|
214
220
|
|
|
215
221
|
- [Coding agent tools](coding-agent-tools.md): durable plan/todo Markdown helpers and `state.coding` checkpoint metadata for restart/resume without a second runtime
|
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
# Data classification and field-level redaction
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Field-level classification and fail-closed redaction at data boundaries (plan 027 Task 8). An explicit host policy classifies every field of a JSON-like value that crosses a boundary — provider prompt egress, tool dispatch/result persistence, artifact write/read/export, audit hashing, telemetry attributes/events, and export — and returns one of four decisions per field: `allow`, `redact`, `tokenize`, or `deny`. Unknown fields fail closed under the protected default; tenant and legal-hold context is carried through the walk. The policy walk is bounded (depth/keys/string budget, cycle detection, wall-clock budget when configured), rejects unsupported values instead of stringifying guesses, and preserves shape when redaction is required (only touched paths are allocated — untouched subtrees share the input reference).
|
|
6
|
+
|
|
7
|
+
Classification is explicit: labels come from a `labelFor` hint function supplied by the boundary owner. There is no automatic sensitive-data discovery, no global registry, no decorator framework, and no second policy language. Existing hardcoded secret redaction (`createSecretRedactor`) remains in place as defense in depth and runs before the policy pass at egress seams.
|
|
8
|
+
|
|
9
|
+
## When to use it
|
|
10
|
+
|
|
11
|
+
- When a boundary must guarantee that classified fields (secrets, financial data, personal data) never reach a sink unchanged, and unknown fields must be blocked rather than guessed at.
|
|
12
|
+
- When different destinations need different handling of the same payload — e.g. allow a field in the prompt but redact it in telemetry — the destination is part of every decision input.
|
|
13
|
+
- When the protected profile must run fail-closed: the provided `createProtectedFieldPolicy()` denies unknown labels on outbound/persisted boundaries by default.
|
|
14
|
+
- When deep-copy cost at a boundary matters: the sparse-copy walker avoids duplicate serialization and shares pristine subtrees.
|
|
15
|
+
|
|
16
|
+
Compatibility note: existing callers that do not supply a policy are untouched (identity fast path); the protected profile is what enables fail-closed behavior, and protected deployments should supply it at every boundary.
|
|
17
|
+
|
|
18
|
+
## Inputs / request
|
|
19
|
+
|
|
20
|
+
- `applyFieldPolicy(value, policy, options)` — `value` is any JSON-like structure (plain objects, arrays, primitives, `Date`/`RegExp`/buffers pass through; `Map`/`Set` normalize to object/array shapes; functions, bigints, symbols, class instances, and cycles are rejected).
|
|
21
|
+
- `policy` — `(input: FieldPolicyInput) => FieldPolicyDecision`, where the input carries `{path, destination, label?, kind, tenantId?, direction, purpose?}`.
|
|
22
|
+
- `options` — `destination` (required), `direction` (default `"outbound"`), `tenantId`, `purpose`, `labelFor` (explicit key→label hints; no auto-discovery), `onRedact` (provenance hook used by the audit adapter), `maxDepth` (32), `maxKeys` (10,000), `maxChars` (1,000,000), `maxPolicyMs` (only when set), `tokenPrefix` (`tok_`).
|
|
23
|
+
- The protected default is `createProtectedFieldPolicy({ publicLabels, deniedLabels, redactedLabels, tokenizedLabels })`: `public`/structural labels pass, `secret` and `financial` deny, `personal` redacts, `token` tokenizes, and anything unlabeled denies on outbound destinations (`prompt`, `tool`, `artifact`, `audit`, `telemetry`, `export`, `persistence`) while passing inbound.
|
|
24
|
+
|
|
25
|
+
## Outputs / response / events
|
|
26
|
+
|
|
27
|
+
- A tree of the same shape with decisions applied: `deny` replaces the value with `[DENIED]`, `redact` replaces string leaves with `[REDACTED]` while preserving containers, `tokenize` replaces string leaves with a deterministic `tok_<hash>` (stable across runs for the same path+value, safe for audit chains), `allow` keeps the value. Untouched branches share the input reference (sparse copy); the input is never mutated.
|
|
28
|
+
- `onRedact` fires once per transformed field with `{path, reason}`; values are never included.
|
|
29
|
+
- `FieldPolicyError` (code `ERR_PRISM_FIELD_POLICY`) on: policy throw, invalid decision, cyclic reference, unsupported value type, or depth/key/byte/time budget breach. Error messages contain the path and the policy error class — never the value.
|
|
30
|
+
|
|
31
|
+
## Request/response example
|
|
32
|
+
|
|
33
|
+
```ts
|
|
34
|
+
import { applyFieldPolicy, createProtectedFieldPolicy } from "@arnilo/prism";
|
|
35
|
+
|
|
36
|
+
const fieldPolicy = createProtectedFieldPolicy();
|
|
37
|
+
const labelFor = (key: string) =>
|
|
38
|
+
key === "apiKey" ? "secret" : key === "email" ? "personal" : key === "score" ? "public" : undefined;
|
|
39
|
+
|
|
40
|
+
const out = applyFieldPolicy(
|
|
41
|
+
{ score: 1, apiKey: "demo-secret-value", email: "ops@example.test", extra: "unknown" },
|
|
42
|
+
fieldPolicy,
|
|
43
|
+
{ destination: "prompt", direction: "outbound", labelFor },
|
|
44
|
+
);
|
|
45
|
+
// { score: 1, apiKey: "[DENIED]", email: "[REDACTED]", extra: "[DENIED]" }
|
|
46
|
+
|
|
47
|
+
// Audit seam: transformation precedes canonical hashing; only {path, reason} survives.
|
|
48
|
+
const redactor = createAuditFieldRedactor(fieldPolicy, { labelFor });
|
|
49
|
+
// pass redactor as the exporter's `redact` option: createAuditExporter({ redact: redactor, ... })
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
## Implementation example
|
|
53
|
+
|
|
54
|
+
- Root ownership: the contract lives in `src/field-policy.ts` of `@arnilo/prism` and is exported from the package index (`applyFieldPolicy`, `createProtectedFieldPolicy`, `hookFieldPolicy`-style adapter `createAuditFieldRedactor`, `ALLOW_FIELD_POLICY`, `FieldPolicyError`, `FIELD_POLICY_LIMITS`).
|
|
55
|
+
- Egress seams: `redactMessage`, `redactProviderRequest`, `redactAgentEvent`, `redactSessionEntry`, and `redactRunLedgerRecord` (from `./redaction.js`) take an optional `(fieldPolicy, destination, labelFor)` — secret redaction runs first, then the policy pass. Without a policy the functions are unchanged (identity).
|
|
56
|
+
- Audit export: `createAuditFieldRedactor(fieldPolicy, { tenantId, labelFor, purpose })` produces the structural `AuditRedactionPolicy` the exporter already applies before canonical hashing; the record's denied/redacted/tokenized bytes are exactly what gets hashed and verified.
|
|
57
|
+
- Telemetry: `createOpenTelemetryInstrumentation({ fieldPolicy })` filters or masks exported span attributes and events — `allow` keeps, `redact` masks with `[REDACTED]`, `deny` drops, `tokenize` hashes — and policy errors drop the attribute without ever echoing the value.
|
|
58
|
+
- Postgres persistence: stores persist exactly what the caller-appointed policy authorizes; the protected profile is a caller-supplied boundary control, not a store default (outbox payloads keep their canonical-collision semantics unchanged).
|
|
59
|
+
|
|
60
|
+
## Extension and configuration notes
|
|
61
|
+
|
|
62
|
+
- Label vocabulary is bounded by the boundary owner: `public`, `personal`, `secret`, `financial`, `token` in the protected default; custom profiles override `deniedLabels`/`redactedLabels`/`tokenizedLabels` maps with their own reasons.
|
|
63
|
+
- Legal hold does not broaden view/export permissions: holds are store-level (deletion prevention), and export-boundary policy denies classified/unknown fields regardless of hold flags.
|
|
64
|
+
- The `labelFor` hint function is the only discovery mechanism; it must be supplied per boundary by the owner who knows the shape. The walker never guesses labels from key names.
|
|
65
|
+
- Tokens are deterministic per (path, value) for a given run; they are not reversible by design and are not a pseudonymization system (no k-anonymity or re-identification risk model).
|
|
66
|
+
- Migration guidance: for existing callers, add the policy at the outermost seam (the redaction functions or the audit/telemetry options) and verify on canaries first; there is no global config key that enables classification, so adoption is per-boundary and explicit.
|
|
67
|
+
|
|
68
|
+
## Security and performance notes
|
|
69
|
+
|
|
70
|
+
- Fail-closed guarantees: unknown fields never cross outbound/persisted boundaries under the protected default; policy exceptions, invalid decisions, and budget breaches throw without echoing values; cycles and unsupported types throw instead of stringifying guesses; tenant mismatch can be enforced inside host policy via `tenantId` on every decision input.
|
|
71
|
+
- Secret canaries: the ERP-T9 matrix proves secret/personal/financial canaries never reach prompt, tool, artifact, audit, telemetry, persistence, or export sinks (denied = `[DENIED]`, redacted = `[REDACTED]`, tokenized = `tok_…`, and the canary string appears nowhere in transformed output or provenance lists).
|
|
72
|
+
- Bounds (frozen): depth 32, keys 10,000, string budget 1,000,000 chars, optional wall-clock budget; the walk is recursive with an active-path set — cycles terminate, diamond references are re-walked per branch.
|
|
73
|
+
- Overhead (frozen cap `classificationMaxOverheadPercent = 10`): measured against the pre-existing boundary walk (the secret-redaction walk boundaries already ran before classification existed) on the frozen representative payload sizes, interleaved A/B — prompt 4,164 B → 99.0%, toolArgs 2,114 B → 95.8%, toolResult 9,095 B → 97.1%, artifactMetadata 3,692 B → 99.0%, auditRecord 4,243 B → 99.8%, telemetry 1,760 B → 97.7%, exportPage 10,726 B → 98.0% of the redactor-walk baseline (peak 99.8%, all ≤ 110%). The raw ratio vs native `JSON.stringify` is recorded in the Task 8 evidence (≈1.0–1.3× across fixtures); a JS policy gateway cannot beat a native serializer, so the frozen cap is defined against the walk work the boundary already performed, and this stays the benchmark contract.
|
|
74
|
+
- The telemetry seam drops attributes on policy error rather than failing the whole span; the audit seam maintains redaction provenance `{path, reason}` only — values never enter the hash chain.
|
|
75
|
+
|
|
76
|
+
## Related APIs
|
|
77
|
+
|
|
78
|
+
- `redactMessage` / `redactProviderRequest` / `redactAgentEvent` / `redactSessionEntry` / `redactRunLedgerRecord` — the egress seams that take the optional policy (secret redaction first, then classification).
|
|
79
|
+
- `createAuditFieldRedactor` → the audit-export `redact` hook; see [Signed, hash-chained audit export](audit-export.md).
|
|
80
|
+
- `createOpenTelemetryInstrumentation` in `@arnilo/prism-observability-opentelemetry` — the telemetry `fieldPolicy` option.
|
|
81
|
+
- `createProtectedFieldPolicy`, `ALLOW_FIELD_POLICY`, `FieldPolicyError`, `FIELD_POLICY_LIMITS` — the protected default and limits.
|
|
82
|
+
- The ERP-T9 threat matrix (`src/__tests__/field-policy.test.ts`) and the boundary-drill scripts cover the enforcement evidence.
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# Disaster recovery and backup operations
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
This page is the operator runbook for backup, restore, migration rollback, point-in-time recovery (PITR), and disaster recovery (DR), proven by plan 027 Task 7 with a protected drill that uses only standard PostgreSQL tools (`pg_dump` custom format, `pg_restore`, `pg_basebackup`, `psql`) plus the existing migration runner. The drill seeds representative multi-tenant 0.2.7 state (sessions, workflow/saga/ACP/conversation checkpoints, leases, legal holds, tenant quotas, policy decisions, evaluations, work idempotency, tool effects, model-router budgets, ERP outbox/inbox, approvals) through the real store APIs, backs it up, restores it into an explicitly confirmed disposable database, verifies per-table row counts and content digests equal the source, rehearses the 0.2.6 → 0.2.7 migration forward with old rows preserved, rehearses rollback by restoring the pre-upgrade backup, and runs PITR against a WAL-archived cluster to a point between two known writes. The recorded run lives in `docs/_evidence/phase27-dr-evidence.json`.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
- Before running any migration or release in a production environment: decide the rollback path first (roll-forward repair preferred; backup-restore is the last resort and only in a disposable environment).
|
|
10
|
+
- When an operator must restore a database: read the guarded-command requirements below — this project has no "force restore" shortcut, and destructive commands always require explicit positive confirmation.
|
|
11
|
+
- When sizing backup windows or reviewing RPO/RTO: read the measured numbers in the evidence file and the ownership table (managed backup, encryption, retention scheduling, and cross-region replication are operator-owned and not claimed by Prism).
|
|
12
|
+
|
|
13
|
+
## Inputs / request
|
|
14
|
+
|
|
15
|
+
- A source instance URL (`PRISM_TEST_POSTGRES_URL` in the drill; any supported PostgreSQL ≥ 14 in practice).
|
|
16
|
+
- An explicitly named, disposable target database on a loopback non-production instance, supplied as `--target` plus the confirmation token `--confirm-target prism_dr_restore`. The target must not already exist; the drill refuses dirty state rather than clobbering it.
|
|
17
|
+
- A separate WAL-archived cluster for PITR (`PRISM_PITR_URL`) with `wal_level=replica`, `archive_mode=on`, and `archive_command='cp %p /wal_archive/%f'`; source and PITR containers must mount a shared host dir at `/dr` for artifact exchange.
|
|
18
|
+
- Sufficient free space (the drill asserts ≥ 512 MB headroom on the artifact dir before starting).
|
|
19
|
+
|
|
20
|
+
## Outputs / response / events
|
|
21
|
+
|
|
22
|
+
- A custom-format backup artifact (`.dump`) with its SHA-256 digest, byte size, duration, and table list count.
|
|
23
|
+
- A restore report: per-table count and content-digest equality against the source (application-level verification, not just exit codes), and duration.
|
|
24
|
+
- A migration report: the 0.2.6-era schema (migrations 001–003) with legacy rows, the upgraded 0.2.7 schema (all five migrations) with old rows preserved and new tables initialized empty, and the rollback rehearsal showing the pre-upgrade backup restores exactly and excludes 0.2.7 tables.
|
|
25
|
+
- A PITR report: recovery target time between two known writes, the earlier write present and the later write absent, recovery duration, and measured RPO/RTO.
|
|
26
|
+
|
|
27
|
+
## Request/response example
|
|
28
|
+
|
|
29
|
+
```sh
|
|
30
|
+
# Protected drill (standard tools only, orchestrated by the script):
|
|
31
|
+
PRISM_PITR_URL=postgresql://user:***@localhost:55436/postgres \
|
|
32
|
+
node scripts/phase27-dr.test.mjs \
|
|
33
|
+
--source "$PRISM_TEST_POSTGRES_URL" \
|
|
34
|
+
--target postgresql://user:***@localhost:55432/prism_dr_target \
|
|
35
|
+
--confirm-target prism_dr_restore
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
```ts
|
|
39
|
+
// Verifying a restore at the application level (from the drill):
|
|
40
|
+
const before = await digest(pool, "prism_erp_outbox"); // md5 of ordered rows
|
|
41
|
+
const after = await digest(restorePool, "prism_erp_outbox");
|
|
42
|
+
assert.equal(after, before); // content equality, not exit code
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
## Implementation example
|
|
46
|
+
|
|
47
|
+
The drill's legs: (1) seed multi-tenant state through `createPostgresPersistence` and `createPostgresEnterpriseState` plus `createPostgresApprovalStore` with a host authority; (2) `pg_dump -F c` into the shared `/dr` mount; (3) `createdb` the confirmed target and `pg_restore --no-owner --no-privileges`, then verify counts and digests per table; (4) build the 0.2.6-era schema from the raw DDL builders with the recorded migration registry rows, seed legacy rows through the real stores, take a backup, run `applyEnterpriseMigrations` (004/005 apply, old rows preserved, new tables empty), then restore the pre-upgrade backup into a fresh database and confirm 0.2.7 tables are absent; (5) on the WAL-archived cluster, take a `pg_basebackup`, insert two marker transactions a known distance apart, switch WAL and wait for archiving, then start a recovered instance with `recovery.signal`, a bounded `restore_command`, and `recovery_target_time` set between the two writes with `recovery_target_action=pause`, and verify the earlier marker exists and the later one does not.
|
|
48
|
+
|
|
49
|
+
## Extension and configuration notes
|
|
50
|
+
|
|
51
|
+
- Storage, encryption at rest, retention scheduling, and cross-region replication are operator-owned: the drill proves the command and verification path, not a managed backup service.
|
|
52
|
+
- Rollback decision tree: prefer roll-forward repair after a production migration. Restore-from-backup is the last resort, permitted only in a disposable evidence environment; the recorded loss window is "writes between the pre-upgrade backup and the rollback restore".
|
|
53
|
+
- Recovery parameters: `recovery_target_time` in the recovered instance needs `recovery.signal` (or `standby.signal`) present, a `restore_command`, and a full timestamp including the sub-second fraction and offset — truncating to whole seconds can land the recovery point before the intended write.
|
|
54
|
+
- Guards that must stay in place: source and target databases must differ; the target host must be loopback and the database name must not match production patterns; the target must not pre-exist; the confirmation token is mandatory; secret canaries seeded into the data must never appear in the manifest or console output.
|
|
55
|
+
- Quarterly re-run the drill and refresh `docs/_evidence/phase27-dr-evidence.json` as the schema evolves; treat any change to table counts/digests as needing a new recorded run.
|
|
56
|
+
|
|
57
|
+
## Security and performance notes
|
|
58
|
+
|
|
59
|
+
- Credentials are explicit and redacted: the manifest stores URLs with passwords masked, and the drill asserts the source/target/PITR passwords and the seeded secret canary never appear in the evidence, logs, or console. Connection strings are passed to the tools via the container environment, never printed.
|
|
60
|
+
- Destructive commands (dropping schemas/databases, restoring over an existing target) require explicit operator action; the drill fails closed on dirty state and never deletes anything itself.
|
|
61
|
+
- Legal-hold data is verified: legal-hold records and their referenced rows survive backup and restore with content digests intact; enforcement of holds stays host-owned (the stores preserve the records; the leases/quota/outbox lifecycle logic stays unchanged).
|
|
62
|
+
- Performance: measured in the disposable environment — backup 108,291 bytes in 122 ms, restore 382 ms, PITR recovery 1.2 s with the two markers a sub-second apart (RPO ≈ 0 s, RTO ≈ 1 s). These are environment-local measurements for sizing, not guarantees; the drill records sizes, timings, and digests in the evidence file and fails on breached frozen budgets.
|
|
63
|
+
- No exactly-once guarantee is claimed anywhere in the backup/restore path; the drill verifies observable equality (counts and digests) instead.
|
|
64
|
+
|
|
65
|
+
## Related APIs
|
|
66
|
+
|
|
67
|
+
- `createPostgresPersistence` / `createPostgresEnterpriseState` / `createPostgresApprovalStore` — the stores whose state the drill seeds and verifies.
|
|
68
|
+
- `applyEnterpriseMigrations` — the migration runner used for the forward upgrade rehearsal.
|
|
69
|
+
- `scripts/phase27-dr.test.mjs` — the protected drill; `docs/_evidence/phase27-dr-evidence.json` — the recorded run.
|
|
70
|
+
- [Operations runbook: high availability, failover, and fencing](operations.md) — the lease/fence model and failover ceiling the same state relies on.
|
|
71
|
+
- [Signed, hash-chained audit export](audit-export.md) — the append-only audit ledger preserved through backup/restore.
|