@arnilo/prism 0.0.23 → 0.0.25
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -3
- package/dist/agent-event-source.d.ts +11 -0
- package/dist/agent-event-source.js +512 -0
- package/dist/agent-loops.js +37 -5
- package/dist/agent-run-state.d.ts +27 -1
- package/dist/agent-run-state.js +86 -5
- package/dist/agents.js +890 -78
- package/dist/contracts.d.ts +338 -4
- package/dist/contracts.js +53 -0
- package/dist/index.d.ts +9 -4
- package/dist/index.js +5 -2
- package/dist/testing/agent-event-source-conformance.d.ts +4 -0
- package/dist/testing/agent-event-source-conformance.js +54 -0
- package/dist/testing/persistence-schema.d.ts +2 -2
- package/dist/testing/persistence-schema.js +58 -21
- package/dist/testing/tool-effect-store-conformance.d.ts +9 -0
- package/dist/testing/tool-effect-store-conformance.js +85 -0
- package/dist/tool-effects.d.ts +15 -0
- package/dist/tool-effects.js +338 -0
- package/dist/tools.d.ts +4 -1
- package/dist/tools.js +219 -9
- package/docs/0.1.0-readiness.md +10 -9
- package/docs/a2a.md +6 -2
- package/docs/ag-ui-adoption.md +77 -0
- package/docs/ag-ui.md +77 -42
- package/docs/agent-events.md +5 -1
- package/docs/agent-loops.md +9 -1
- package/docs/agent-session-runtime.md +9 -2
- package/docs/browser-automation.md +2 -0
- package/docs/coding-agent-tools.md +2 -0
- package/docs/coding-security.md +1 -1
- package/docs/database-persistence.md +2 -0
- package/docs/enterprise-postgres-state.md +5 -1
- package/docs/host-security.md +8 -1
- package/docs/index.md +13 -11
- package/docs/mcp-tools.md +19 -2
- package/docs/migration.md +46 -0
- package/docs/performance.md +25 -0
- package/docs/postgres-persistence.md +5 -2
- package/docs/public-contracts.md +2 -0
- package/docs/release-and-install.md +70 -690
- package/docs/server.md +10 -6
- package/docs/sqlite-persistence.md +10 -2
- package/docs/supervisors.md +6 -0
- package/docs/tool-effects.md +95 -0
- package/docs/tools.md +4 -0
- package/docs/work-tools.md +4 -0
- package/docs/workflows.md +1 -1
- package/package.json +11 -3
package/docs/ag-ui.md
CHANGED
|
@@ -4,12 +4,10 @@
|
|
|
4
4
|
|
|
5
5
|
`@arnilo/prism-ag-ui` is an optional, framework-free protocol adapter over Prism's existing redacted `AgentEvent`, session, durable-run, and persistence seams.
|
|
6
6
|
|
|
7
|
-
- Root export maps Prism events to AG-UI `@ag-ui/core` **0.0.57** events and offers `createAgUiHandler()` (`Request` → SSE `Response`)
|
|
7
|
+
- Root export maps Prism events to AG-UI `@ag-ui/core` **0.0.57** events and offers `createAgUiHandler()` (`Request` → SSE `Response`), compatible `createPersistenceAgUiReplay()` pages, distributed `createAgentEventSourceAgUiReplay()` follow, and explicit `createAgUiMcpAdapter()` / `createAgUiMcpAppHandler()` / `createAgUiA2AAdapter()` protocol handshakes.
|
|
8
8
|
- `@arnilo/prism-ag-ui/acp` uses stable `@agentclientprotocol/sdk` **1.3.0** root exports for `createAcpEventMapper()` and `createPrismAcpAgent()`.
|
|
9
9
|
- Core remains protocol-free. `resumeAgentRunStream()` / `AgentRunLifecycle.resumeStream()` are generic durable-resume streams shared by adapters.
|
|
10
10
|
|
|
11
|
-
This is not an app TUI, desktop shell, conversation database, terminal/filesystem bridge, A2A implementation, or frontend tool registry.
|
|
12
|
-
|
|
13
11
|
## When to use it
|
|
14
12
|
|
|
15
13
|
Use AG-UI when a host already authenticates users, owns sessions and durable run correlation, and needs a bounded Web endpoint for a browser/TUI/desktop client. Use ACP when an editor client already supplies an ACP transport and needs text, safe tool status, usage, and approval updates from a Prism session.
|
|
@@ -29,70 +27,69 @@ npm install @arnilo/prism @arnilo/prism-ag-ui
|
|
|
29
27
|
| Input | Purpose |
|
|
30
28
|
| --- | --- |
|
|
31
29
|
| `authorize` | Rebinds untrusted AG-UI thread/run selectors to host ownership on every request. `false` returns 403. |
|
|
32
|
-
| `sessionFactory` | Returns an authorized Prism `AgentSession`;
|
|
30
|
+
| `sessionFactory` | Returns an authorized Prism `AgentSession`; it receives only host-approved `AgUiPreparedInput`, never raw client tools/state. |
|
|
31
|
+
| `input.project` | Opts into full `RunAgentInput`; turns bounded, still-untrusted history, state, context, forwarded props, media, and lineage into host-selected Prism `Message` values. Omit for legacy final-text mode. |
|
|
32
|
+
| `input.frontendTools` | Explicitly selects client-side handoffs. Returned names must be request-tool subset; adapter never turns JSON tool declarations into Prism `ToolDefinition`s. |
|
|
33
|
+
| `mcp` | Optional `createAgUiMcpAdapter({ bridge, select })`; host selects reviewed `bridge.tools`, then `sessionFactory` receives them as `input.serverTools`. Normal Prism dispatch/loop remains sole executor. |
|
|
34
|
+
| `a2a` | Optional `createAgUiA2AAdapter({ client, select, correlate })`; verified remote A2A task stream replaces this handler's local session only. Host selection/correlation binds each remote task to ownership. |
|
|
33
35
|
| `lifecycle` + `resolveRun` | Optional durable status/resume path. Required only for a resumed interruption. |
|
|
34
|
-
| `
|
|
35
|
-
| `
|
|
36
|
+
| `interrupts.resume` | Optional host aggregate-policy callback for multiple AG-UI interrupts; it returns one current-version core approve/deny decision. |
|
|
37
|
+
| `replay` | Optional page adapter or `createAgentEventSourceAgUiReplay(source, options)` for gap-free distributed replay/live follow. |
|
|
38
|
+
| `projection` | Explicit safe tool/state/messages/activity/reasoning/raw/custom/interrupt projection. Omit each callback for default deny. Prefer `composeAgUiProjections(createMessagesFromSessionProjection(...), createStateFromStoreProjection(...), createActivityFromToolProgressProjection(), host)` for standard families. |
|
|
39
|
+
| `a2ui` | Opt-in A2UI painting middleware (`{ catalogId, mode, renderToolName?, allowedCatalogIds?, limits? }`). Detects `a2ui_operations` tool results and/or streams from `render_a2ui` args; paints `a2ui-surface` activity events. Absent = inert. |
|
|
40
|
+
| `capabilities` | Optional host declaration narrowed to implemented SSE/projector/lifecycle features; read `handler.capabilities`. |
|
|
36
41
|
| `redactor`, `limits` | Host redaction and narrowing-only finite caps. |
|
|
37
42
|
|
|
38
|
-
The handler accepts only `POST` JSON validated with AG-UI `RunAgentInputSchema`.
|
|
43
|
+
The handler accepts only `POST` JSON validated with official AG-UI `RunAgentInputSchema`. Every aggregate is bounded before a callback runs. With no `input.project`, it preserves compatibility: final text user message only; non-empty state or frontend tools fail before authorization/session lookup. With a projector, all current roles/history, context, state, forwarded props, multimodal parts, parent lineage, and tool-result continuations are available as untrusted input. The projector must apply Prism media URL/SSRF/MIME policy before forwarding media. Start a run with no `resume` and no `?cursor=`; replay supplies `?cursor=`.
|
|
39
44
|
|
|
40
45
|
## Outputs / response / events
|
|
41
46
|
|
|
42
|
-
The handler returns `text/event-stream`, one `data: <AG-UI event>\n\n` frame per output. Mapper lifecycle is ordered:
|
|
47
|
+
The handler returns `text/event-stream`, one `data: <AG-UI event>\n\n` frame per output. Mapper lifecycle is ordered: `RUN_*`, `STEP_*`, `TEXT_MESSAGE_*`, and `TOOL_CALL_*` are deterministic Prism mappings. Host projectors may additionally prove and emit `STATE_SNAPSHOT`/`STATE_DELTA`, `MESSAGES_SNAPSHOT`, `ACTIVITY_*`, current `REASONING_*`, `RAW`, and named `CUSTOM` values. All values revalidate against official `EventSchemas`; deprecated `THINKING_*` and convenience chunk events are not produced. Active message/tool/reasoning/step sequences close before error, interruption, or finish.
|
|
43
48
|
|
|
44
|
-
A Prism durable `agent_suspended` returns `RUN_FINISHED` with interrupt id `${runId}:${version}` and a strict `{ decision: "approve" | "deny" }` schema.
|
|
49
|
+
A Prism durable `agent_suspended` returns `RUN_FINISHED` with core interrupt id `${runId}:${version}` and a strict `{ decision: "approve" | "deny" }` schema. `projection.interrupt` may attach bounded expiry/metadata or additional host policy interrupts but must retain that core id. Without `interrupts.resume`, one exact entry is required; `cancelled` means deny. An aggregate policy may validate bounded multiple entries, then returns one current-version core decision. Payloads containing `editedArgs`/`args` always deny: Prism does not mutate persisted tool calls. The adapter checks host authorization, selected run, suspended status, and checkpoint version, then calls `AgentRunLifecycle.resumeStream()` once. Claimed/dispatched tools are never replayed.
|
|
45
50
|
|
|
46
|
-
`createPersistenceAgUiReplay()`
|
|
51
|
+
`createPersistenceAgUiReplay()` remains a compatible page adapter. `createAgentEventSourceAgUiReplay()` resolves exact ownership/run once per open, then consumes the shared durable source through terminal or live follow; it never attaches replica-local `session.subscribe()`. Every record must already be redacted. Mapped events carry stable `prismEventId` and bounded opaque `prismCursor`; records with no standard mapping emit `CUSTOM prism.replay_cursor`, so clients can persist progress. Terminal replay never creates a session or reruns a provider/tool.
|
|
47
52
|
|
|
48
53
|
ACP maps assistant text to `agent_message_chunk`, safe tool lifecycle to `tool_call`/`tool_call_update`, provider usage to `usage_update`, and durable suspension to `session/request_permission`. Only `allow_once` approves; reject, cancellation, unknown outcomes, and request failure deny. It advertises only close-session capability—no terminal, filesystem, MCP, editor state, location, diff, or raw input/output capability.
|
|
49
54
|
|
|
50
|
-
|
|
55
|
+
Selected MCP tools use normal `TOOL_CALL_*` core dispatch. Linked Apps add safe `mcp-apps` activity; app-only tools stay model-hidden. The separate reauthorizing Apps proxy allow-lists initialize/ping/logging/tool/resource calls for one bridge; its sandbox helper returns CSP/iframe config and never executes HTML.
|
|
51
56
|
|
|
52
|
-
|
|
57
|
+
`createAgUiA2AAdapter()` maps verified task text/activity. Non-text/tool/A2UI parts need host `projectPart`; non-streaming fallback accepts only a terminal task, otherwise host follows saved correlation.
|
|
53
58
|
|
|
54
|
-
|
|
55
|
-
{
|
|
56
|
-
"threadId": "thread-1",
|
|
57
|
-
"runId": "run-1",
|
|
58
|
-
"messages": [{ "id": "message-1", "role": "user", "content": "Summarize this" }],
|
|
59
|
-
"tools": [],
|
|
60
|
-
"state": {}
|
|
61
|
-
}
|
|
62
|
-
```
|
|
59
|
+
Co-work uses bounded, redacted `CUSTOM prism.cowork.*` events through `mapCoWork()` / `createCoWorkReplay()`; see [Work artifacts and review](work-artifacts-and-review.md).
|
|
63
60
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
```json
|
|
67
|
-
{
|
|
68
|
-
"type": "RUN_FINISHED",
|
|
69
|
-
"threadId": "thread-1",
|
|
70
|
-
"runId": "run-1",
|
|
71
|
-
"outcome": {
|
|
72
|
-
"type": "interrupt",
|
|
73
|
-
"interrupts": [{ "id": "run-1:4", "responseSchema": { "required": ["decision"] } }]
|
|
74
|
-
}
|
|
75
|
-
}
|
|
76
|
-
```
|
|
61
|
+
## Request/response example
|
|
77
62
|
|
|
78
|
-
Resume
|
|
63
|
+
Resume a default single interrupt with `resume: [{ "interruptId": "run-1:4", "status": "resolved", "payload": { "decision": "approve" } }]`. Full history, client tool results, and mutable state need authorized `input.project` selection. This adapter is not a conversation database.
|
|
79
64
|
|
|
80
65
|
## Implementation example
|
|
81
66
|
|
|
82
67
|
```ts
|
|
83
68
|
import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
|
|
84
|
-
import { createAgUiHandler } from "@arnilo/prism-ag-ui";
|
|
69
|
+
import { createAgentEventSourceAgUiReplay, createAgUiHandler } from "@arnilo/prism-ag-ui";
|
|
85
70
|
|
|
86
71
|
const agent = createAgent({
|
|
87
72
|
model: { provider: "mock", model: "offline" },
|
|
88
73
|
provider: createMockProvider([providerTextDelta("ready"), providerDone()]),
|
|
89
74
|
});
|
|
90
75
|
|
|
76
|
+
const replay = createAgentEventSourceAgUiReplay(persistence.events, {
|
|
77
|
+
resolveRun,
|
|
78
|
+
ownership: (authorization) => authorization.ownership,
|
|
79
|
+
});
|
|
80
|
+
|
|
91
81
|
const handle = createAgUiHandler({
|
|
92
82
|
authorize: ({ request }) => request.headers.get("authorization") === "Bearer host-checked"
|
|
93
83
|
? { ownership: { userId: "user-1" } }
|
|
94
84
|
: false,
|
|
85
|
+
// Create a run-scoped agent/session using input.serverTools when mcp is enabled.
|
|
95
86
|
sessionFactory: () => agent.createSession({ id: "host-owned-thread" }),
|
|
87
|
+
replay,
|
|
88
|
+
input: {
|
|
89
|
+
project: () => ({
|
|
90
|
+
messages: [{ id: "user", role: "user", content: [{ type: "text", text: "host-selected input" }] }],
|
|
91
|
+
}),
|
|
92
|
+
},
|
|
96
93
|
projection: { toolArguments: () => undefined, toolResult: () => undefined },
|
|
97
94
|
});
|
|
98
95
|
|
|
@@ -103,19 +100,54 @@ See runnable network-free [`examples/ag-ui-server.ts`](../examples/ag-ui-server.
|
|
|
103
100
|
|
|
104
101
|
## Extension and configuration notes
|
|
105
102
|
|
|
106
|
-
All identity, authorization, session/thread mapping, durable checkpoint lookup, persistence selection, replay cursor persistence, transport adaptation, and optional projection are host-owned. The adapter owns no listener, database, background reconnect loop, credential resolver, or UI state.
|
|
103
|
+
All identity, authorization, session/thread mapping, durable checkpoint lookup, persistence selection, replay cursor persistence, transport adaptation, MCP bridge/card configuration, app sandbox DOM, remote A2A task correlation, and optional projection are host-owned. The adapter owns no listener, database, background reconnect loop, credential resolver, or UI state.
|
|
107
104
|
|
|
108
|
-
`AgUiProjection` is an allow-list. Without a callback, raw tool arguments/results/progress,
|
|
105
|
+
`AgUiProjection` is an allow-list. Without a callback, raw tool arguments/results/progress, arbitrary state/patches/transcripts/activity/reasoning/raw events, paths, ACP locations/diffs/terminals/raw I/O, and frontend-supplied tools remain absent. Reasoning signatures do not become AG-UI encrypted values automatically: a host must explicitly provide an already client-encrypted opaque value. `input.project` is also an allow-list: do not merge client state/forwarded props into ownership, identity, tools, permissions, provider options, or media fetch policy.
|
|
109
106
|
|
|
110
107
|
Co-work projection reuses the same allow-list: `AgUiProjection.coWork(event)` may return a curated, JSON-serializable payload for a co-work event; absent it, the redacted event fields are exposed. Wire `coWorkContext` to derive thread/artifact/identity from the authorized request (never client JSON) and `coWork` to a `createCoWorkReplay()` over your durable artifact/draft/snapshot stores. The handler projects one bounded page after the run; mount a dedicated cursor-paged co-work endpoint when full pagination is needed.
|
|
111
108
|
|
|
112
|
-
|
|
109
|
+
Durable interrupts carry the shared decision batch: the fallback interrupt includes the redacted `pendingDecisions` under `metadata` and its `responseSchema` accepts either the legacy `{ decision: "approve" | "deny" }` or a `{ decisions: [{ approvalId, outcome, reason?, modifiedArguments?, elicitation? }] }` batch. All batch entries are shape- and cap-validated at the boundary (count ≤ 128, ids ≤ 128 chars, four outcomes, reason ≤ 8 KiB, payloads ≤ 64 KiB) and core re-validates each against the recorded pending set under the single CAS. `interrupts.resume` may return the batch form (`{ decisions, expectedVersion? }`); legacy `editedArgs` resume payloads still deny. ACP permission prompts offer the four outcomes (`allow_once` / `allow_always` / `reject_once` / `reject_always`) and map them onto the batch; a cancelled prompt stays deny-closed.
|
|
110
|
+
|
|
111
|
+
### Standard projectors (opt-in)
|
|
113
112
|
|
|
114
|
-
|
|
113
|
+
Three batteries-included factories return `AgUiProjection` fragments. Compose with host projectors via `composeAgUiProjections(...fragments)` — **first defined callback wins** (left to right); `undefined` fragments are skipped. Absent factories keep 0.0.24 default-deny.
|
|
114
|
+
|
|
115
|
+
```ts
|
|
116
|
+
createAgUiHandler({
|
|
117
|
+
projection: composeAgUiProjections(
|
|
118
|
+
createMessagesFromSessionProjection({ getMessages: () => authorizedAgUiMessages, redact }),
|
|
119
|
+
createStateFromStoreProjection(runStateStore),
|
|
120
|
+
createActivityFromToolProgressProjection(),
|
|
121
|
+
hostCustom,
|
|
122
|
+
),
|
|
123
|
+
});
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
| Factory | Emits | Notes |
|
|
127
|
+
| --- | --- | --- |
|
|
128
|
+
| `createMessagesFromSessionProjection` | `MESSAGES_SNAPSHOT` | Host `getMessages()` for authorized history, or live `message_finished` accumulation. Caps 128/1024. Redact drops closed. |
|
|
129
|
+
| `createStateFromStoreProjection(store)` | `STATE_SNAPSHOT` on `agent_started`; RFC 6902 `STATE_DELTA` (add/replace/remove) when `store.get()` changes | Host store; optional `subscribe` only marks dirty — no Prism watcher. Oversized/throw → drop closed. |
|
|
130
|
+
| `createActivityFromToolProgressProjection` | `ACTIVITY_SNAPSHOT` / `ACTIVITY_DELTA` from `tool_execution_progress` | Default `activityType: "tool-progress"`. Missing progress+metadata → drop closed. |
|
|
131
|
+
|
|
132
|
+
### A2UI painting middleware (opt-in)
|
|
133
|
+
|
|
134
|
+
`createAgUiHandler({ a2ui: { catalogId, mode } })` paints A2UI v0.9 surfaces without a host `projection.activity` callback:
|
|
135
|
+
|
|
136
|
+
| Mode | Source | Paint |
|
|
137
|
+
| --- | --- | --- |
|
|
138
|
+
| `fixed-schema` | Tool result `{ a2ui_operations: [...] }` | First batch with `createSurface` → `ACTIVITY_SNAPSHOT` (`activityType: "a2ui-surface"`); later batches → `ACTIVITY_DELTA` |
|
|
139
|
+
| `streaming` | `tool_call_delta` args of `renderToolName` (default `render_a2ui`) | Progressive `ACTIVITY_SNAPSHOT` with `replace: true` only when complete ops extractable — never partial JSON |
|
|
140
|
+
| `both` | Both paths; streamed surfaces are not re-painted from the final envelope |
|
|
141
|
+
|
|
142
|
+
Host `catalogId` is stamped when absent; model-supplied ids outside `allowedCatalogIds` (default `[catalogId]`) are overwritten. Invalid ops emit one bounded `CUSTOM` event `prism.a2ui.error` and paint nothing. Caps: 64/512 ops per message, 64 KiB/1 MiB per op, 16/64 surfaces per run, depth 32/64.
|
|
143
|
+
|
|
144
|
+
User actions arrive as untrusted `AgUiA2UiAction` values on `input.project({ a2uiActions })` (from `forwardedProps.a2uiAction` or activity/tool-result shapes). Without `input.project` they stay default-deny — Prism never synthesizes a `log_a2ui_event` tool call (documented divergence from official `@ag-ui/a2ui-middleware`). Example: `examples/ag-ui-a2ui.ts`.
|
|
145
|
+
|
|
146
|
+
## Security and performance notes
|
|
115
147
|
|
|
116
|
-
|
|
148
|
+
Authorize every start/replay/resume/proxy/follow. Treat protocol fields, MCP metadata/HTML, and A2A cards/parts as untrusted; persist run/task correlation before output and redact streams. MCP Apps requires extension acknowledgement, exact proxy origin, same-bridge visibility, approval, `ui://` HTML/MIME bounds, and sandbox CSP. It never retries UI mutations; Task 4 adds recovery.
|
|
117
149
|
|
|
118
|
-
|
|
150
|
+
Defaults / hard caps: request 64 KiB / 1 MiB; input 128 / 1024 messages, 32 / 256 tools/contexts, 8 / 64 interrupts, 16 / 64 media parts, and 64 KiB / 1 MiB text/state/media; frontend tool/context payloads 16 KiB / 256 KiB; projected event/state/activity/reasoning/raw values 64 KiB / 1 MiB; patches 128 / 4096 operations; JSON depth 16 / 64, properties 128 / 4096, arrays 512 / 8192; cursor 4 / 16 KiB; replay page 100 / 500 records; queue 128 / 4096 events; stream 10,000 / 100,000 events and 10 / 64 MiB; wall time 120 seconds / 30 minutes. Overflow yields a bounded error/closed stream, not an unbounded queue. SSE is declared; WebSocket/protobuf/push are not. Reconnect is at-least-once, so clients de-duplicate stable event/message/tool IDs.
|
|
119
151
|
|
|
120
152
|
## Related APIs
|
|
121
153
|
|
|
@@ -124,5 +156,8 @@ Benchmark command/result placeholder: Task 8 adds `node scripts/benchmark-0.0.12
|
|
|
124
156
|
- [Runs and usage ledger](runs-and-usage.md): durable `AgentEventRecord` query source.
|
|
125
157
|
- [Web-standard server handler](server.md): generic Prism HTTP API, separate from AG-UI.
|
|
126
158
|
- [A2A interoperability](a2a.md): remote agent-to-agent tasks, not frontend protocol mapping.
|
|
159
|
+
- [AG-UI adoption evaluation](ag-ui-adoption.md): official 0.0.57 event/input matrix and shipped explicit MCP/MCP Apps/A2A handshakes.
|
|
160
|
+
- [MCP bridge/server](mcp-tools.md): `mcpApps` negotiation, bounded resources, and remote tool trust.
|
|
161
|
+
- [A2A interoperability](a2a.md): verified rich task client and remote task lifecycle.
|
|
127
162
|
- [Host security guide](host-security.md): authorization, ownership, redaction, and credential boundaries.
|
|
128
163
|
- [Work artifacts and review](work-artifacts-and-review.md): durable artifact service that produces the co-work approval/progress/download-link events projected here.
|
package/docs/agent-events.md
CHANGED
|
@@ -10,7 +10,7 @@ Events are emitted by the runtime and by loops through `LoopContext.emit`, both
|
|
|
10
10
|
|
|
11
11
|
Subscribe via `session.stream()` for a single owned run, or `session.subscribe()` when a host needs a long-lived observer across runs: render streamed assistant text in a UI, react to tool execution, drive observability/telemetry, or audit artifact validation outcomes. Do not parse provider stream events directly for these — `AgentEvent` is the stable, normalized surface across providers and loops.
|
|
12
12
|
|
|
13
|
-
Do not use `
|
|
13
|
+
Do not use live `session.subscribe()` for cross-replica reconnect — use durable `AgentEventSource` below. Live subscribe remains process-local.
|
|
14
14
|
|
|
15
15
|
## Durable event ledger
|
|
16
16
|
|
|
@@ -18,6 +18,10 @@ When `AgentConfig.runLedger` or `RunOptions.runLedger` is configured, every emit
|
|
|
18
18
|
|
|
19
19
|
Event records preserve emission order within a run because the runtime drains pending event appends before writing the final `RunRecord`. Subscribers still see the live, in-memory stream; the ledger is the durable copy.
|
|
20
20
|
|
|
21
|
+
## Durable AgentEventSource
|
|
22
|
+
|
|
23
|
+
`AgentEventSource` (`createMemoryAgentEventSource` / `persistence.events` on PostgreSQL) appends, pages, and subscribes with opaque ownership-bound cursors. `subscribe` registers wake interest before replaying history so replay-to-live handoff has no gap. Delivery is at-least-once; consumers dedupe `record.id`. PostgreSQL uses transactional sequence allocation plus `LISTEN`/`NOTIFY` wakeups with polling fallback. Transport adapters (server SSE `Last-Event-ID`, AG-UI, A2A `afterEventId`) map source envelopes only — they do not invent private replay loops. This is not exactly-once.
|
|
24
|
+
|
|
21
25
|
## Inputs / request
|
|
22
26
|
|
|
23
27
|
```ts
|
package/docs/agent-loops.md
CHANGED
|
@@ -118,7 +118,15 @@ Optional steer hooks on `LoopContext` (0.0.11): `hasPendingSteers?()` / `applyPe
|
|
|
118
118
|
|
|
119
119
|
## Durable runs
|
|
120
120
|
|
|
121
|
-
`RunOptions.runState` supports
|
|
121
|
+
`RunOptions.runState` supports the built-in loop options (`single-shot` and `generate-validate-revise`) and custom strategies that opt into durable state. `single-shot` is durable via the runtime's pending-call mechanism and carries no loop-local state. `generate-validate-revise` snapshots `{ attempts, artifactPhase, savedSchema, pendingHistory }` at `revision: "1"`. A custom `AgentLoopStrategy` must declare both snapshot hooks or durable configuration rejects it with `AgentLoopStateError` (`ERR_PRISM_LOOP_NOT_DURABLE`) before any provider call:
|
|
122
|
+
|
|
123
|
+
| Member | Purpose |
|
|
124
|
+
| --- | --- |
|
|
125
|
+
| `revision?: string` | Host-authored loop revision. Joins the durable-run fingerprint, so a loop change without a `definitionRevision` bump fails closed on resume. |
|
|
126
|
+
| `snapshot?(): JsonValue` | Capture loop-local resumable state at suspension. Must be JSON-compatible; core redacts it and bounds it inside the durable run-state envelope (`maxStateBytes`, depth 32). A non-JSON value fails the run with `ERR_PRISM_LOOP_SNAPSHOT`. |
|
|
127
|
+
| `restore?(snapshot): void` | Rehydrate from the captured snapshot; must throw on drift. Called once before `run(ctx)` on resume. Also available as `ctx.restoredLoopState`. |
|
|
128
|
+
|
|
129
|
+
The snapshot is stored as `loopState: { name, revision, snapshot }` on the durable run state and cleared when the run reaches a terminal status. On resume, a name/revision mismatch between the stored `loopState` and the resolved strategy fails closed (`ERR_PRISM_LOOP_REVISION`), and the fingerprint check independently rejects any loop drift. Suspension occurs only before an input provider call or immediately before a tool side effect; completed provider turns remain in `SessionStore` history and are not repeated after `resumeAgentRun()`.
|
|
122
130
|
|
|
123
131
|
## Outputs / response / events
|
|
124
132
|
|
|
@@ -176,7 +176,14 @@ await agent.createSession().run("Hi", { model: overrideModel });
|
|
|
176
176
|
|
|
177
177
|
## Durable interruption
|
|
178
178
|
|
|
179
|
-
Set `runState` with a host-owned `CheckpointStore`, stable `definitionRevision`, and `interruptBeforeTool: true` to suspend at a persisted pre-side-effect boundary. A suspended result has `status: "suspended"`, a redacted `interruption`, and `runState.version`; it releases session resources before returning.
|
|
179
|
+
Set `runState` with a host-owned `CheckpointStore`, stable `definitionRevision`, and `interruptBeforeTool: true` to suspend at a persisted pre-side-effect boundary. A suspended result has `status: "suspended"`, a redacted `interruption`, and `runState.version`; it releases session resources before returning. When a provider turn requests several tools, the round is collected into **one** suspension whose `interruption.pendingDecisions` holds one redacted `PendingDecision` per gated call (`approvalId`, kind, scope with tool name/effect kind/identity/arguments hash — never raw arguments); ungated calls still dispatch.
|
|
180
|
+
|
|
181
|
+
`resumeAgentRun` accepts exactly one of:
|
|
182
|
+
|
|
183
|
+
- `decision: "approve" | "deny"` — legacy single-approval path. `approve` allows every pending decision once; `deny` terminates the run as `denied`.
|
|
184
|
+
- `decisions: readonly RunDecision[]` — one atomic batch. Every entry validates against the recorded pending set (unknown/foreign `approvalId`, duplicates, stale `expectedVersion`, invalid outcomes fail the whole batch closed with `AgentDecisionError` and leave state and version untouched). Outcomes: `allow_once`, `allow_for_run`, `reject_once`, `reject_for_run`. `reject_*` continues the run with a blocked tool result carrying the bounded (2 KB) `reason`. `modifiedArguments` are revalidated (schema, then input guardrails; permission/trust re-run at dispatch) and produce a new arguments hash. `elicitation` payloads are validated against the pending decision's `elicitationSchema` (required keys plus the configured host validator) and resolve the suspended call without executing it. A batch deciding a strict subset persists the decided entries and re-suspends with the remainder pending at the bumped version.
|
|
185
|
+
|
|
186
|
+
`*_for_run` outcomes append a `StickyDecision` to the durable run state: later calls in the same run matching the scope exactly (all recorded fields) proceed or are blocked without a new suspension, policy still enforced at dispatch. Sticky decisions expire when the run reaches any terminal status. Caps: 32 pending decisions per run (hard 128), 64 sticky decisions (hard 256), 2 KB decision reasons, 16 KB elicitation payloads.
|
|
180
187
|
|
|
181
188
|
```ts
|
|
182
189
|
const result = await session.run("Publish draft", {
|
|
@@ -189,7 +196,7 @@ if (result.status === "suspended") {
|
|
|
189
196
|
}
|
|
190
197
|
```
|
|
191
198
|
|
|
192
|
-
Resume requires exact checkpoint ownership, version, agent fingerprint, and revision. The fingerprint hashes the agent id/name, `definitionRevision`, model, instructions, system-prompt contributions, skills (name/instructions/tool names), tool definitions (name/parameters/exclusive), guardrail definitions (name/stage/revision), and loop strategy — changing any of them without bumping `definitionRevision` fails resume closed instead of silently continuing with different agent semantics. Prism CAS-claims approval before work, rechecks normal guardrail/permission/validation/limit paths, and marks a pending tool dispatched before its side effect. `createAgentRunLifecycle()` wraps the same core path for server/MCP hosts: adapters pass only authorized ownership, status returns only `{ state, version }`, and `resolveAgent()` supplies current agent/revision. `resumeStream()` uses that same claim path and bounded subscriber, so adapters do not poll or duplicate resume logic. Remote restart requires both checkpoint and session stores to be durable. A crash after that mark is ambiguous and is never replayed automatically; use host tool idempotency keyed by `runId`/`toolCallId` or resolve it manually. Checkpoints contain bounded redacted state plus session/leaf references, never provider objects, callbacks, signals, credentials, or raw secrets. State is bounded at save by `runState.maxStateBytes` (default 256 KB, at most the 1 MB hard cap); load bounds against the 1 MB hard cap only, so state saved with a raised limit stays resumable while oversized records are still rejected.
|
|
199
|
+
Resume requires exact checkpoint ownership, version, agent fingerprint, and revision. The fingerprint hashes the agent id/name, `definitionRevision`, model, instructions, system-prompt contributions, skills (name/instructions/tool names), tool definitions (name/parameters/exclusive), guardrail definitions (name/stage/revision), and loop strategy — changing any of them without bumping `definitionRevision` fails resume closed instead of silently continuing with different agent semantics. Prism CAS-claims approval before work, rechecks normal guardrail/permission/validation/limit paths, and marks a pending tool dispatched before its side effect. `createAgentRunLifecycle()` wraps the same core path for server/MCP hosts: adapters pass only authorized ownership, status returns only `{ state, version }`, and `resolveAgent()` supplies current agent/revision. `resumeStream()` uses that same claim path and bounded subscriber, so adapters do not poll or duplicate resume logic. Remote restart requires both checkpoint and session stores to be durable. A crash after that mark is ambiguous and is never replayed automatically; use host tool idempotency keyed by `runId`/`toolCallId` or resolve it manually. Checkpoints contain bounded redacted state plus session/leaf references, never provider objects, callbacks, signals, credentials, or raw secrets. State is bounded at save by `runState.maxStateBytes` (default 256 KB, at most the 1 MB hard cap); load bounds against the 1 MB hard cap only, so state saved with a raised limit stays resumable while oversized records are still rejected. Built-in loop options are durable; custom `AgentLoopStrategy` instances are durable when they declare `snapshot`/`restore` hooks (see [Agent loops § Durable runs](agent-loops.md#durable-runs)) and reject before provider work otherwise.
|
|
193
200
|
|
|
194
201
|
## Secure composition
|
|
195
202
|
|
|
@@ -108,6 +108,8 @@ await browser.close();
|
|
|
108
108
|
- Raw CSS is absent from production defaults. Ref resolution uses Playwright’s built-in `aria-ref=` selector with a package-owned snapshot ref table for staleness checks.
|
|
109
109
|
- Verified-state checkpoints (0.0.14): `createBrowserCheckpointLedger()` records navigation state — URL, a domain-state hash, and host-owned data refs — never serialized browser internals (cookies/storage/contexts), which are fragile and secret-bearing. Frozen caps: URL 8 KiB/16 KiB, domain-state hash 256 B/1 KiB, host-data ref 2 KiB/8 KiB (refs only, never bodies), 16/64 checkpoints per run (oldest evicted). After any resume/interruption `markResumed(runId)` marks state stale; `assertVerifiedBeforeSideEffect(runId)` fails closed until the host reloads + `verify()`s, so side effects never replay on stale state. Checkpoints are run-scoped: a conversation thread composes through the run it owns, reusing the manager's sandbox/egress/approval/limit policy above.
|
|
110
110
|
|
|
111
|
+
Observation tools declare `kind: none`; mutations are `external_mutation`/`unsupported` and fail closed on stale checkpoint state. See [tool effects](tool-effects.md).
|
|
112
|
+
|
|
111
113
|
## Security and performance notes
|
|
112
114
|
|
|
113
115
|
Import is inert. Construction fails clearly when neither `browser` nor `manager` is supplied. Browser installation, launch, version, and control endpoint are host-owned. Prism never exposes `page.evaluate`, init scripts, CDP, extensions, persistent profiles, or model-supplied Playwright launch options. Secrets and storage state must not appear in snapshots, tool results, logs, or checkpoints. Finite caps charge before context/page/action/queue/snapshot/network/artifact retention; snapshots retain no unbounded DOM, console, request, response, or trace history. Unreleased downloads are deleted on context close.
|
|
@@ -491,6 +491,8 @@ Packed capability demo: `examples/coding-tools-capability-gaps.ts` (search modes
|
|
|
491
491
|
- **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
|
|
492
492
|
- No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
|
|
493
493
|
|
|
494
|
+
Read/list/search/glob are observation effects; write/edit/delete/move are optional local mutations; shell/check are unsupported external mutations. `reconcileCodingToolEffect` proves local postconditions or returns `unknown`. See [tool effects](tool-effects.md).
|
|
495
|
+
|
|
494
496
|
## Security and performance notes
|
|
495
497
|
|
|
496
498
|
- **Host shell/filesystem access.** These tools run real commands and read/write/list/search/glob/delete/move real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
|
package/docs/coding-security.md
CHANGED
|
@@ -143,7 +143,7 @@ Policies are ordinary host values: attach one globally through `createCodingTool
|
|
|
143
143
|
|
|
144
144
|
The Docker reference adapter starts by recorded container ID/label, uses argument arrays only, mounts source read-only, populates a size-bounded tmpfs `/workspace`, drops all capabilities, enables `no-new-privileges`, runs with `--init`, and never exposes the Docker socket, privileged mode, or host PID/IPC namespaces. Image pull/build/update stays outside Prism. Protected real-Docker checks are opt-in via `PRISM_TEST_DOCKER_SANDBOX=1` with host-supplied `PRISM_TEST_DOCKER_BIN` and digest-pinned `PRISM_TEST_DOCKER_IMAGE`.
|
|
145
145
|
|
|
146
|
-
Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
|
|
146
|
+
Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. `ask_user_decision` also maps onto the shared decision model: inside a durable gated agent run its call suspends as a kind-`elicitation` pending decision whose schema carries the choice contract (option-id enums) plus the full question/options UX payload on a Prism-owned schema extension property; the resume decision's `elicitation` payload (a `selectedId`/`selectedIds`/`customText` answer) is validated against the schema and the tool-level answer-shape rules, then resolves the call without invoking the blocking `ask()` callback. The process-local `ask()` path and the workflow suspend/resume path are unchanged. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
|
|
147
147
|
|
|
148
148
|
## Security and performance notes
|
|
149
149
|
|
|
@@ -451,6 +451,8 @@ const dbStore: ProductionPersistenceStore = {
|
|
|
451
451
|
- Cursor values and idempotency keys are host-defined and opaque to Prism.
|
|
452
452
|
- First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume, human suspension, multi-process coordination, Phase 11 schedule records/fire leases, shared state, and replay lineage; workflow code owns no SQL table. `suspended`/`denied`, schedules, state history, and replay lineage remain namespaces/categories plus bounded checkpoint JSON values, so Phases 8 and 11 need no database migration.
|
|
453
453
|
|
|
454
|
+
Schema version **7** adds the exact-owner durable event retention index (`prism_agent_events_owner_timestamp_sequence_idx`). Distributed subscribe/LISTEN remains PostgreSQL-only via `persistence.events`.
|
|
455
|
+
|
|
454
456
|
## Security and performance notes
|
|
455
457
|
|
|
456
458
|
- **No credentials in storage.** The contracts never include `CredentialResolver`, `AIProvider`, `ProviderResolver`, provider API keys, or credential values.
|
|
@@ -10,8 +10,10 @@
|
|
|
10
10
|
| Evaluations | `evaluations` | `EvaluationStore` append/query records with exact owner pages. |
|
|
11
11
|
| Work mutations | `workIdempotency` | Atomic claim/CAS lifecycle for connector effects. |
|
|
12
12
|
| Model routing | `modelRouter` | Shared rate, budget, and circuit state for router replicas. |
|
|
13
|
+
| Tool effects | `toolEffects` | Durable `ToolEffectStore` claim/CAS for recoverable tool side effects (migration 002). |
|
|
14
|
+
| Tool effects | `toolEffects` | Durable `ToolEffectStore` claim/CAS for recoverable tool side effects (migration 002). |
|
|
13
15
|
|
|
14
|
-
`createPostgresEnterpriseState()` opens a host-supplied or adapter-owned `pg` pool, verifies/applies
|
|
16
|
+
`createPostgresEnterpriseState()` opens a host-supplied or adapter-owned `pg` pool, verifies/applies checksum-protected enterprise migrations (`001_enterprise_state`, `002_tool_effects`), and returns those stores plus explicit cleanup and close operations. Importing it performs no I/O. It is separate from session/run persistence in [`@arnilo/prism-session-store-postgres`](postgres-persistence.md).
|
|
15
17
|
|
|
16
18
|
## When to use it
|
|
17
19
|
|
|
@@ -55,6 +57,8 @@ interface PostgresEnterpriseState {
|
|
|
55
57
|
readonly evaluations: EvaluationStore;
|
|
56
58
|
readonly workIdempotency: IdempotencyStore;
|
|
57
59
|
readonly modelRouter: ModelRouterStateStore;
|
|
60
|
+
readonly toolEffects: ToolEffectStore;
|
|
61
|
+
readonly toolEffects: ToolEffectStore;
|
|
58
62
|
cleanup(input: EnterpriseStateCleanupInput): Promise<EnterpriseStateCleanupResult>;
|
|
59
63
|
close(): Promise<void>;
|
|
60
64
|
}
|
package/docs/host-security.md
CHANGED
|
@@ -143,6 +143,9 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
143
143
|
- Evaluation trace readers require exact supplied ownership plus session/run identity, reject cursor/identity drift, and redact before bounded scorer/judge input. Model-judge callbacks receive no credential resolver, tools, or workspace; keep live judges outside default CI and redact report artifacts.
|
|
144
144
|
- Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
|
|
145
145
|
- MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands/resources/prompts, requires per-operation `authorize`, and retains core gates. Sampling, roots, model/credential selection, and elicitation consent stay host-owned; URL elicitation is never opened automatically. Stateful web mode requires host `resolveAuthInfo` plus `resolveIdentity`, exact origin policy, and binds every POST/GET/DELETE/SSE request to one non-secret principal; mismatches return 404. Handler still needs TLS and edge rate limiting. See [MCP client/server exposure](mcp-tools.md).
|
|
146
|
+
- AG-UI fields stay untrusted after schema validation. `input.project` returns host-selected messages/handoffs only; never merge state/tools/context/props into ownership, identity, permissions, provider options, or media policy. Apply Prism media SSRF/MIME bounds before resolution; output projectors are bounded allow-lists; interrupt edits deny rather than mutate persisted calls.
|
|
147
|
+
- AG-UI MCP Apps requires negotiated `mcpApps`, exact proxy origin/auth, owned-run context, approval, one bridge, separate-origin sandbox (`allow-scripts allow-same-origin`), and no-wider CSP. Never execute HTML in host origin or retry a UI mutation; Task 4 adds recovery.
|
|
148
|
+
- AG-UI A2A requires exact-origin verified client, host-owned task selection/correlation, explicit data/tool/A2UI projection, and reauthorized follow/cancel.
|
|
146
149
|
- `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. The artifact review service (`createArtifactService`) requires authenticated identity + thread ownership on every attach/revise/compare/approve/reject/download, resolves concurrent reviewers via checkpoint CAS (no lost approvals), rejects local filesystem paths in `uri`/citations, redacts records before persist and on response, and serves downloads only through signed expiring links that are reauthorized against the token's ownership per request. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
|
|
147
150
|
- Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, repository list/search depth/entry/match/scan/time caps, structured Git path/ref/message/output/patch/worktree caps, named-check concurrency/output caps, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Opt-in `createGitTools()` uses argument arrays with hooks/credential prompts/external diff disabled, requires host `commitIdentity` for commits, and never pushes or opens PRs. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell/repository backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, identity-scoped approval caching, required `workspaceMode` on `createSandboxCodingComposition()` / `createSandboxCodingTools()`, and the optional `createDockerSandbox()` reference adapter. **Host mode is never contained execution** (`containmentClaim: false`). Sandbox mode claims containment only when FS backends target the disposable tree; mixed wiring requires `allowMixedWorkspaceWiring` and still does not claim containment. Limits alone are not containment: construct the Docker adapter (absolute CLI, digest-pinned image, network none by default) or an equivalent host sandbox before treating coding execution as production-safe. Docker daemon/image trust, egress firewall/proxy, and artifact retention remain host-owned.
|
|
148
151
|
- Optional `@arnilo/prism-browser` requires a host-supplied Playwright Browser (`playwright-core@1.61.0` peer). Import is inert. One non-persistent context belongs to one run; actions serialize; refs are snapshot-scoped; CSS/evaluate/CDP/persistent profiles are denied. Context routing + `serviceWorkers: "block"` deny file/data/blob/devtools/private/loopback by default and require contained-proxy attestation for external egress (Playwright routing is defense in depth, not DNS containment). Uploads are realpath-rooted; downloads quarantine with hash/MIME until host `approveRelease`; screenshots return bounded `ImageContent`. Observation vs mutation/high-impact actions map to `ExecutionPolicy`. Treat snapshot/page text as untrusted external content. Close contexts with `browser_close` or `manager.closeRun(runId)` on abort/terminal. Browser control endpoint, binary/image pin, and real egress firewall/proxy remain host-owned. Shared sandbox: `createSharedSandboxBrowserOptions()` + `assertBrowserSandboxNetwork()`.
|
|
@@ -199,10 +202,14 @@ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy
|
|
|
199
202
|
- Scheduled/manual coding/browser containment checks run in protected `sandbox-browser` environment (`.github/workflows/sandbox-browser.yml`). They receive no provider/npm/OIDC secrets; Docker/Playwright enablement is variable-gated with host-preloaded digest-pinned images/binaries; uploads are redacted aggregate status only.
|
|
200
203
|
- Live endpoint operators own TLS, egress allow-lists, account-dollar budget, cleanup beyond MCP session DELETE, and revocation. Failed canaries log only operation kind plus status/timeout; inspect provider-side audit logs for details.
|
|
201
204
|
|
|
205
|
+
## Distributed events and tool effects
|
|
206
|
+
|
|
207
|
+
Every durable `AgentEventSource` page/subscribe and tool-effect claim rechecks exact ownership. Opaque cursors never select tenants. Oversized events/effects fail closed. Ambiguous tool outcomes stay `unknown` until operator reconciliation — never silent replay. Request-path SQL stays DML-only; migration/listener principals stay separate. See [tool effects](tool-effects.md) and [agent events](agent-events.md).
|
|
208
|
+
|
|
202
209
|
## Related APIs
|
|
203
210
|
|
|
204
211
|
- [Web-standard server handler](server.md): remote agent/workflow route, ownership, limits, abort, and deployment boundary.
|
|
205
|
-
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): authorize every protocol selector/operation;
|
|
212
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): authorize every protocol selector/operation; full AG-UI input/output only through bounded host allow-lists; exact interrupt/version resume; redacted, ownership-scoped replay.
|
|
206
213
|
- [Supervisor delegation](supervisors.md): local child permission/memory/budget boundary.
|
|
207
214
|
- [A2A interoperability](a2a.md): remote card/auth/origin/signature boundary.
|
|
208
215
|
- [Settings, auth, trust, and security controls](settings-auth-trust-security.md): low-level helpers and boundary hardening table.
|
package/docs/index.md
CHANGED
|
@@ -11,15 +11,15 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
11
11
|
- [Model routing](model-routing.md): optional `@arnilo/prism-model-router` allow-list/residency/budget/rate/circuit/fallback governance with redacted diagnostics; durable state requires awaited identity-scoped calls.
|
|
12
12
|
|
|
13
13
|
## Agent/session runtime
|
|
14
|
-
- [Agent/session runtime](agent-session-runtime.md): create explicit or opt-in secure agents/sessions, get direct `AgentRunResult` values from `run`/`prompt`, mid-run `steer` (turn-boundary or softInterrupt), use integrated `stream()`/`resumeAgentRunStream()`, subscribe to normalized events, and expose opted-in durable lifecycle capabilities.
|
|
14
|
+
- [Agent/session runtime](agent-session-runtime.md): create explicit or opt-in secure agents/sessions, get direct `AgentRunResult` values from `run`/`prompt`, mid-run `steer` (turn-boundary or softInterrupt), use integrated `stream()`/`resumeAgentRunStream()`, shared batch pending-decisions / sticky run-scope approvals, subscribe to normalized events, and expose opted-in durable lifecycle capabilities.
|
|
15
15
|
- [Agent definitions](agent-definitions.md): resolve declarative `AgentDefinition` values via `resolveAgentDefinition`, and turn app-config `<configRoot>/agents/<name>/AGENT.md` bundles into runnable agents via `discoverAgentBundles` / `resolveAgentBundle` (explicit tool/skill activation by name, fail-closed omitted capabilities, migration-only `activateAllCapabilities`, strict duplicate scope checks, configurable prompt layers, no auto-discovery).
|
|
16
|
-
- [Agent loops](agent-loops.md): replaceable per-run control loops — `singleShotLoop` default
|
|
16
|
+
- [Agent loops](agent-loops.md): replaceable per-run control loops — `singleShotLoop` default, opt-in bounded artifact-loop tool rounds, and durable custom-loop `revision`/`snapshot`/`restore` hooks with fail-closed resume.
|
|
17
17
|
- [Guardrails](guardrails.md): typed fail-closed input/output/tool checks with buffered provider output and redacted decision records.
|
|
18
|
-
- [Agent events](agent-events.md):
|
|
18
|
+
- [Agent events](agent-events.md): live `session.subscribe` plus durable `AgentEventSource` page/subscribe/resume for cross-replica reconnect; message/progress deltas never create spans.
|
|
19
19
|
- [Observability](observability.md): OTel GenAI agent/provider/tool hierarchy, host context parenting, bounded trace linkage, safe evaluation events, controlled metrics, and exporter isolation.
|
|
20
20
|
- [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, coding/browser adversarial fixtures, ID-only linkage to immutable owned run feedback, and optional durable PostgreSQL records.
|
|
21
21
|
- [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence, optional bounded FIFO durability policies, session snapshot caching, and immutable run/trace feedback.
|
|
22
|
-
- [Performance limits](performance.md): 0.0.
|
|
22
|
+
- [Performance limits](performance.md): 0.0.25 durable-loop/HITL/A2UI network-free evidence, 0.0.24 distributed event/effect PostgreSQL evidence, 0.0.23 enterprise state evidence, 0.0.15 network-free provider/RAG/memory benchmark evidence and frozen caps, bounded evaluation traces/judges/reports, and production sizing assumptions.
|
|
23
23
|
- [Structured output](structured-output.md): the `Artifact*` seam plus provider-native `StructuredOutputOptions` / `structuredOutputMode` for capable models.
|
|
24
24
|
|
|
25
25
|
## Compaction/session memory
|
|
@@ -34,8 +34,8 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
34
34
|
- [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention/legal-hold/quota lifecycle (`lifecycle`), and NoSQL mapping.
|
|
35
35
|
- [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, FTS `searchSessions` (migration-v4), and transactionally verified/backfilled migration metadata.
|
|
36
36
|
- [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, FTS `searchSessions` (migration-v4), advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
|
|
37
|
-
- [Enterprise PostgreSQL state](enterprise-postgres-state.md): optional `@arnilo/prism-enterprise-postgres` composition for durable policy/evaluation/work-idempotency/model-router state, exact ownership, checksummed
|
|
38
|
-
- [Migration guide](migration.md): **0.0.23** enterprise PostgreSQL state adapters (async router and work reconciliation); **0.0.22** third-party behavior integrations (Caveman, Ponytail); **0.0.21** coding-tool capability gaps (`outputMode`, glob, read-before-write, delete/move, aggregator 9/4); **0.0.20** progressive skill disclosure, empty registry default, `load_skill`, priority budget demotion, optional tool-result fold; **0.0.19** observational memory lifecycle; **0.0.15** OpenAI hosted tools/continuation/Realtime, exact AI SDK v4 matrix, RAG lifecycle/reranking/trust/status, and memory export/rebuild; **0.0.14** conversations, memory consent/lifecycle, artifact co-work review, AG-UI co-work events, scoped M365/GWS OAuth connectors, browser checkpoints, device contracts, and Alibaba/Ollama providers; plus prior release migrations.
|
|
37
|
+
- [Enterprise PostgreSQL state](enterprise-postgres-state.md): optional `@arnilo/prism-enterprise-postgres` composition for durable policy/evaluation/work-idempotency/model-router/`toolEffects` state, exact ownership, checksummed migrations, and explicit cleanup.
|
|
38
|
+
- [Migration guide](migration.md): **0.0.25** durable custom loops and shared human-in-the-loop decisions; **0.0.24** distributed events and recoverable tool effects; **0.0.23** enterprise PostgreSQL state adapters (async router and work reconciliation); **0.0.22** third-party behavior integrations (Caveman, Ponytail); **0.0.21** coding-tool capability gaps (`outputMode`, glob, read-before-write, delete/move, aggregator 9/4); **0.0.20** progressive skill disclosure, empty registry default, `load_skill`, priority budget demotion, optional tool-result fold; **0.0.19** observational memory lifecycle; **0.0.15** OpenAI hosted tools/continuation/Realtime, exact AI SDK v4 matrix, RAG lifecycle/reranking/trust/status, and memory export/rebuild; **0.0.14** conversations, memory consent/lifecycle, artifact co-work review, AG-UI co-work events, scoped M365/GWS OAuth connectors, browser checkpoints, device contracts, and Alibaba/Ollama providers; plus prior release migrations.
|
|
39
39
|
- [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety; `searchSessions` throws `SessionSearchUnsupportedError`.
|
|
40
40
|
- [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
|
|
41
41
|
|
|
@@ -63,6 +63,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
63
63
|
- [Retrieval-augmented generation](rag.md): optional bounded source lifecycle, document adapters, host reranking, ingestion status, attributable citations, and inert context injection.
|
|
64
64
|
|
|
65
65
|
## Tools
|
|
66
|
+
- [Recoverable tool effects](tool-effects.md): optional effect declarations, `ToolEffectStore` claim/CAS, unknown reconciliation (not exactly-once), and adapter classifications.
|
|
66
67
|
- [Tools](tools.md): register host-owned active tools with replace-or-error duplicate policy, apply exact allow/deny filtering, dispatch normal or opt-in bounded artifact-loop calls, and optionally bound untrusted JSON Schema compilation.
|
|
67
68
|
- [Tool execution primitives](tool-execution-primitives.md): finite JSON Schema LRU validation, exclusive-aware bounded parallel dispatch, MCP bridge mapping, coding execution policy, and image-read bounds.
|
|
68
69
|
- [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
|
|
@@ -88,12 +89,13 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
88
89
|
- [Resource loading](resource-loading.md): decode text, JSON, binary, and manifests through caller-provided loaders; bridge host-authorized artifacts to bounded RAG document loading.
|
|
89
90
|
|
|
90
91
|
## Server/API
|
|
91
|
-
- [Web-standard server handler](server.md): optional framework-free authorized direct/SSE agent, durable
|
|
92
|
+
- [Web-standard server handler](server.md): optional framework-free authorized direct/SSE agent, cross-replica durable event reconnect via `Last-Event-ID`, durable agent lifecycle/workflow routes, plus health/drain/rate-limit/replay/deployment-lease seams; explicit bounds and zero default exposure.
|
|
92
93
|
|
|
93
94
|
## Multi-agent and interoperability
|
|
94
95
|
- [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, finite budgets, host-projected delegation telemetry, and separate A2A durable adapter boundary.
|
|
95
|
-
- [A2A interoperability](a2a.md): A2A 1.0 JSON-RPC/HTTPS cards plus host-owned durable task get/list/cancel/subscribe, bounded rich parts/replay, principal-scoped push configs,
|
|
96
|
-
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional `@arnilo/prism-ag-ui` AG-UI mapper
|
|
96
|
+
- [A2A interoperability](a2a.md): A2A 1.0 JSON-RPC/HTTPS cards plus host-owned durable task get/list/cancel/subscribe, shared `AgentEventSource` task adapter, bounded rich parts/replay, principal-scoped push configs, exact-origin verified client, and rich stream seam for explicit AG-UI fronting.
|
|
97
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional `@arnilo/prism-ag-ui` full AG-UI 0.0.57 input/event/capability mapper, authorized Web handler/distributed source follow, opt-in A2UI painting middleware, explicit hardened MCP/MCP Apps/remote A2A adapters, and stable ACP sibling over shared redacted event and durable-approval seams; 0.0.14 adds reconnectable co-work events.
|
|
98
|
+
- [AG-UI adoption evaluation](ag-ui-adoption.md): official 0.0.57 input/event/capability matrix and shipped hardened MCP/MCP Apps/A2A handshake boundaries.
|
|
97
99
|
|
|
98
100
|
## CLI/RPC
|
|
99
101
|
- [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including mid-run `steer`, branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
|
|
@@ -115,14 +117,14 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
115
117
|
- [Compaction conformance](compaction-conformance.md): assert any `CompactionStrategy` returns a non-empty redacted summary and observes abort from `@arnilo/prism/testing/compaction-conformance`.
|
|
116
118
|
- [Tool conformance](tool-conformance.md): assert the tool-dispatch blocked-reason matrix (unknown/denied/invalid/permission/validator) and success path from `@arnilo/prism/testing/tool-conformance`.
|
|
117
119
|
- [Extension conformance](extension-conformance.md): assert an `Extension` setup runs, contributions stay inert, and setup errors are redacted or rethrown from `@arnilo/prism/testing/extension-conformance`.
|
|
118
|
-
- `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts), [`examples/enterprise-identity.ts`](../examples/enterprise-identity.ts), [`examples/enterprise-policy-audit.ts`](../examples/enterprise-policy-audit.ts), [`examples/enterprise-work-connectors.ts`](../examples/enterprise-work-connectors.ts), [`examples/enterprise-postgres-state.ts`](../examples/enterprise-postgres-state.ts), [`examples/conversation-durable-replay.ts`](../examples/conversation-durable-replay.ts), [`examples/artifact-review-delivery.ts`](../examples/artifact-review-delivery.ts), [`examples/server-deployment-seams.ts`](../examples/server-deployment-seams.ts), cache-aware prompt assembly, NeuralWatt agent run ([`examples/neuralwatt-agent-run.ts`](../examples/neuralwatt-agent-run.ts)), [`examples/coding-compaction.ts`](../examples/coding-compaction.ts), [`examples/caveman-ponytail.ts`](../examples/caveman-ponytail.ts), stores/branching, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
|
|
120
|
+
- `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts), [`examples/ag-ui-a2ui.ts`](../examples/ag-ui-a2ui.ts), [`examples/ag-ui-mcp-apps.ts`](../examples/ag-ui-mcp-apps.ts), [`examples/enterprise-identity.ts`](../examples/enterprise-identity.ts), [`examples/enterprise-policy-audit.ts`](../examples/enterprise-policy-audit.ts), [`examples/enterprise-work-connectors.ts`](../examples/enterprise-work-connectors.ts), [`examples/enterprise-postgres-state.ts`](../examples/enterprise-postgres-state.ts), [`examples/conversation-durable-replay.ts`](../examples/conversation-durable-replay.ts), [`examples/artifact-review-delivery.ts`](../examples/artifact-review-delivery.ts), [`examples/server-deployment-seams.ts`](../examples/server-deployment-seams.ts), cache-aware prompt assembly, NeuralWatt agent run ([`examples/neuralwatt-agent-run.ts`](../examples/neuralwatt-agent-run.ts)), [`examples/coding-compaction.ts`](../examples/coding-compaction.ts), [`examples/caveman-ponytail.ts`](../examples/caveman-ponytail.ts), stores/branching, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
|
|
119
121
|
|
|
120
122
|
## Third-party integrations
|
|
121
123
|
- [Caveman behavior integration](caveman.md): optional `@arnilo/prism-caveman` — upstream Caveman skills/commands, `caveman-mode` injector, session `caveman-level` persistence, progressive catalog + `load_skill`; requires host `upstreamPath` and session attach callbacks; inert until `kernel.load`.
|
|
122
124
|
- [Ponytail behavior integration](ponytail.md): optional `@arnilo/prism-ponytail` — upstream Ponytail skills/commands, `ponytail-mode` injector, session `ponytail-mode` persistence; resolves peer `@dietrichgebert/ponytail` or `upstreamPath`; opt-in (not in code/sdk profiles).
|
|
123
125
|
|
|
124
126
|
## Release and install
|
|
125
|
-
- [Release and install](release-and-install.md): current **0.0.
|
|
127
|
+
- [Release and install](release-and-install.md): current **0.0.24** 47-package graph (Phase 7 distributed events and tool effects; plan 007), exact-peer/install/tarball rules, deterministic resumable publication and publish dry-run, protected PostgreSQL gate, pinned supply-chain gates, offline tests, the 0.0.15 provider/AI-SDK/RAG/memory protected live-canary matrix, and sandbox-browser Docker/Playwright gates.
|
|
126
128
|
- [0.1.0 / 1.0 readiness gates](0.1.0-readiness.md): command-per-gate 1.0 readiness table — frozen API surface + compat gate, migration/docs tripwires, budget table, live-suite matrix, security matrix, current-line status (**0.0.23** published target), signed-publication/live-canary prerequisites for 1.0, and Phase 12 demand-evidence entry criteria.
|
|
127
129
|
- [Review coverage (2026-07-26 Phase 11)](review-coverage-2026-07-26-phase-11.md): Plan 079 evidence freeze — baseline size/startup/benchmark budgets, hotspot domain extraction table, confirmed duplication survivors (redactor/cleanJson/row-codecs/checkpoints/exec-runner/approval/ownership), profile adoption recommendations, and tarball artifact-diet findings for 0.0.16.
|
|
128
130
|
- [Review coverage (2026-07-26 Phase 10)](review-coverage-2026-07-26-phase-10.md): Plan 078 evidence freeze — OpenAI hosted tools/continuation/realtime, AI SDK version matrix, remaining provider metadata parity, RAG replaceSource/loaders/parsers/reranker/provenance/ingestion-status, memory export/rebuild/conformance, and 0.0.15 (43 → 43 manifests; no new package) release gates.
|
package/docs/mcp-tools.md
CHANGED
|
@@ -21,6 +21,17 @@ await bridge.close(); // close client + transport
|
|
|
21
21
|
|
|
22
22
|
Advanced hosts that manage their own `Client` + `Transport` can call `attachMcpToolBridge()` or `attachMcpCapabilities()` after connect. `connectMcpCapabilities()` keeps resources/prompts as host-facing facades rather than converting them into model tools, and declares roots/sampling/elicitation only when callbacks are supplied.
|
|
23
23
|
|
|
24
|
+
MCP Apps is an explicit opt-in on the normal bridge:
|
|
25
|
+
|
|
26
|
+
```ts
|
|
27
|
+
const bridge = await connectMcpTools({ serverId: "weather", transport, mcpApps: true });
|
|
28
|
+
// Fails unless the server acknowledges io.modelcontextprotocol/ui.
|
|
29
|
+
const app = await bridge.apps!.readResource("ui://weather/card");
|
|
30
|
+
// bridge.tools excludes _meta.ui.visibility: ["app"] tools.
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
`bridge.apps` exposes reviewed UI metadata, linked bounded `ui://` HTML, and same-server app tools for a host renderer/proxy; it never creates an iframe or executes HTML.
|
|
34
|
+
|
|
24
35
|
```ts
|
|
25
36
|
const bridge = await connectMcpCapabilities({
|
|
26
37
|
serverId: "research",
|
|
@@ -78,7 +89,7 @@ Do **not** use this package as a sandbox, permission engine, or auto-discovery l
|
|
|
78
89
|
|
|
79
90
|
## Outputs / response / events
|
|
80
91
|
|
|
81
|
-
|
|
92
|
+
`McpToolBridge` exposes `tools`, optional `apps`, `refresh()`, and `close()`. Normal tools are Prism `ToolDefinition`s. Apps requires server acknowledgement; nested resource metadata wins over flat/deprecated and app-only tools stay outside `tools`. Resource reads require linked bounded `ui://` HTML5 with exact MIME; content metadata wins over list defaults.
|
|
82
93
|
|
|
83
94
|
`createPrismMcpServer()` returns the SDK `McpServer`. It lists only passed tools/commands and explicitly selected `agentRuns` lifecycle tools; JSON Schema parameters are converted through installed Zod v4 for SDK validation, then Prism tool calls still pass through `dispatchToolCall` permission/validator/redactor gates. Command definitions support explicitly selected direct/background/replay workflow operations and optional ownership-scoped schedule operations from `createWorkflowCommands()`; none are registered unless the host passes those command definitions. Calls return bounded MCP text content and `isError` on denial/failure. `createPrismMcpWebHandler()` returns `(Request) => Promise<Response>`.
|
|
84
95
|
|
|
@@ -126,6 +137,7 @@ Duplicate prefixed names throw `McpToolNameCollisionError` at refresh time.
|
|
|
126
137
|
| `serverId` | required | Stable identifier used in default name prefix |
|
|
127
138
|
| `transport` | required | `stdio` or `streamable-http` config |
|
|
128
139
|
| `namePrefix` | `mcp:<serverId>:` | Registry namespace for remote tools |
|
|
140
|
+
| `mcpApps` | `false` | Explicitly negotiate `io.modelcontextprotocol/ui`; exposes `bridge.apps` only after server acknowledgement |
|
|
129
141
|
| `listCacheTtlMs` | 30 s (24 h hard) | Skip re-listing until TTL expires (invalidated on list-changed) |
|
|
130
142
|
| `callTimeoutMs` | 60 s (30 min hard) | Connect, list-page, and tool-call SDK request timeout/abort |
|
|
131
143
|
| `maxListPages` / `maxTools` | 20 / 500 (hard 100 / 5,000) | Stop pagination before another request/append |
|
|
@@ -138,6 +150,8 @@ Duplicate prefixed names throw `McpToolNameCollisionError` at refresh time.
|
|
|
138
150
|
| `maxJsonDepth` / `maxJsonProperties` | 64 / 10,000 (hard 128 / 100,000) | Bound schema and result JSON walks |
|
|
139
151
|
| `signal` | none | Abort connect/list and trigger close on connect abort |
|
|
140
152
|
|
|
153
|
+
MCP elicitation maps onto the shared decision model: `mcpElicitationDecision(approvalId, params)` converts an untrusted `ElicitRequest` (message ≤ 2 KiB, schema ≤ 16 KiB) into a kind-`elicitation` pending decision, and `mcpElicitationResultFromDecision(decision, { humanInteraction })` maps a decision back to a protocol result — `reject_*` declines, `allow_*` accepts with the payload and fails closed unless the host proved explicit human interaction. Wire behavior is unchanged; the marker never reaches protocol output.
|
|
154
|
+
|
|
141
155
|
### Stdio transport
|
|
142
156
|
|
|
143
157
|
```ts
|
|
@@ -187,6 +201,8 @@ Plaintext is accepted only when `allowLoopbackHttp: true`, the URL hostname is l
|
|
|
187
201
|
|
|
188
202
|
Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard), 32 concurrent requests (512 hard), 60 s timeout (30 min hard), and 32 sessions (512 hard). Stateful mode is intentionally one official SDK transport/session lineage per handler; use one handler/server instance per independently hosted endpoint when multi-tenant transport isolation is required. It parses bounded JSON before passing `parsedBody` to the SDK transport. `allowedHosts`/`allowedOrigins` activate SDK DNS-rebinding checks only when explicitly configured. Authentication data comes only from host `resolveAuthInfo()`.
|
|
189
203
|
|
|
204
|
+
Remote MCP tools default to `external_mutation`/`unsupported` unless the host `effect` policy classifies them. MCP Apps (`io.modelcontextprotocol/ui`) stay behind host CSP/origin/visibility gates. See [tool effects](tool-effects.md) and [AG-UI adoption](ag-ui-adoption.md).
|
|
205
|
+
|
|
190
206
|
## Security and performance notes
|
|
191
207
|
|
|
192
208
|
| Risk | Mitigation |
|
|
@@ -195,6 +211,7 @@ Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard),
|
|
|
195
211
|
| SSRF / DNS rebinding / redirects (HTTP) | Exact HTTPS origins; credentials/fragments/redirects denied; every DNS answer public; one address pinned per request; explicit loopback-only HTTP escape hatch |
|
|
196
212
|
| Hostile discovery / schema compilation | Raw SDK `tools/list` requests avoid SDK Ajv output-schema compilation; finite pages/tools/cursors/metadata/schema totals; failed refresh leaves previous tools unchanged |
|
|
197
213
|
| Tool-name shadowing | Prefixed names + `createToolRegistry({ duplicate: "error" })` |
|
|
214
|
+
| MCP Apps metadata/HTML | Explicit extension acknowledgement; bounded nested metadata; app-only tools absent from model list; only linked `ui://` HTML/MIME resource body reaches the host renderer |
|
|
198
215
|
| Oversized/deep/wide server output | One aggregate byte/depth/property walk covers content, structured content, compatibility `toolResult`, and bounded remote errors before `ToolResult` |
|
|
199
216
|
| Unvalidated arguments | Register tools with `createJsonSchemaToolArgumentValidator()` at dispatch |
|
|
200
217
|
| Missing permission gate | Client direction: `PermissionPolicy` on `tool:mcp:<serverId>:<name>:execute`; server direction: required MCP `authorize` plus optional core `PermissionPolicy` |
|
|
@@ -207,7 +224,7 @@ Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard),
|
|
|
207
224
|
|
|
208
225
|
For durable lifecycle exposure, construct `createAgentRunLifecycle({ checkpoints, resolveAgent })` in core, then pass selected entries as `agentRuns: { support: { lifecycle } }`. MCP registers two tools: `agent.support.status` accepts `{ runId, sessionId? }`; `agent.support.resume` accepts `{ runId, sessionId?, decision, expectedVersion }`. Do not expose an agent without durable checkpoints and a restart-safe `SessionStore`; no lifecycle tool appears by default.
|
|
209
226
|
|
|
210
|
-
MCP output is untrusted. Register bridge tools through core dispatch with a `SecretRedactor`
|
|
227
|
+
MCP output is untrusted. Register bridge tools through core dispatch with a `SecretRedactor`. Apps renderer needs separate origin, `allow-scripts allow-same-origin`, no-wider CSP, and authenticated proxy approval; no app mutation retry before Task 4 recovery. `CreatePrismMcpServerOptions.guardrails` applies shared tool-input/output stages to registered Prism tools; commands remain host callbacks. See [Guardrails](guardrails.md). Prism does not infer unknown secrets. MCP server authorization does not replace tool `PermissionPolicy`, argument validation, coding `ExecutionPolicy`, workflow ownership checks, TLS, rate limiting, or sandboxing. A timed-out tool must cooperate with `AbortSignal` to stop side effects; protocol retention and HTTP responses remain bounded when remote work ignores abort.
|
|
211
228
|
|
|
212
229
|
Discovery validation is atomic: cursor/page/tool/name/description/schema failures reject `refresh()` and preserve the previous immutable tool-array reference. The bridge intentionally uses raw SDK `request()` for `tools/list` and `tools/call`; this avoids eager Ajv compilation/validation of untrusted remote output schemas. Host `ToolValidator` remains the argument-validation owner.
|
|
213
230
|
|