@arnilo/prism 0.0.96 → 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +285 -2
- package/README.md +17 -3
- package/dist/agent-definitions.js +2 -3
- package/dist/agent-event-source.d.ts +11 -0
- package/dist/agent-event-source.js +512 -0
- package/dist/agent-loops.d.ts +5 -0
- package/dist/agent-loops.js +99 -14
- package/dist/agent-run-lifecycle.d.ts +5 -2
- package/dist/agent-run-lifecycle.js +18 -2
- package/dist/agent-run-state.d.ts +27 -1
- package/dist/agent-run-state.js +113 -7
- package/dist/agents.d.ts +3 -1
- package/dist/agents.js +1255 -129
- package/dist/artifacts.d.ts +132 -0
- package/dist/artifacts.js +44 -0
- package/dist/cache-helpers.js +18 -9
- package/dist/checkpoints.d.ts +4 -0
- package/dist/checkpoints.js +17 -9
- package/dist/cli-init.js +3 -7
- package/dist/cli-runner.d.ts +2 -6
- package/dist/cli-runner.js +71 -33
- package/dist/compaction.js +5 -4
- package/dist/config.js +7 -4
- package/dist/content.js +26 -24
- package/dist/context-budget.d.ts +67 -0
- package/dist/context-budget.js +288 -0
- package/dist/contracts.d.ts +590 -8
- package/dist/contracts.js +142 -1
- package/dist/contribution-parsing.js +6 -2
- package/dist/contributions.d.ts +2 -0
- package/dist/contributions.js +3 -0
- package/dist/conversations.d.ts +50 -0
- package/dist/conversations.js +98 -0
- package/dist/credentials.d.ts +22 -2
- package/dist/credentials.js +18 -3
- package/dist/devices.d.ts +94 -0
- package/dist/devices.js +138 -0
- package/dist/event-multiplexer.js +18 -4
- package/dist/extensions.d.ts +18 -1
- package/dist/extensions.js +79 -6
- package/dist/feedback.js +12 -10
- package/dist/guardrails.d.ts +1 -1
- package/dist/guardrails.js +26 -17
- package/dist/identity.d.ts +92 -0
- package/dist/identity.js +265 -0
- package/dist/index.d.ts +94 -72
- package/dist/index.js +48 -36
- package/dist/input.d.ts +10 -1
- package/dist/input.js +152 -52
- package/dist/instruction-injection.d.ts +1 -1
- package/dist/middleware.js +9 -1
- package/dist/models.d.ts +2 -0
- package/dist/models.js +3 -0
- package/dist/node/agent-definitions.js +16 -8
- package/dist/node/contribution-discovery.d.ts +1 -2
- package/dist/node/contribution-discovery.js +3 -3
- package/dist/node/session-store-jsonl.js +13 -7
- package/dist/node/settings.d.ts +1 -1
- package/dist/node/settings.js +1 -1
- package/dist/node/system-project-prompts.js +2 -4
- package/dist/node/trust.js +1 -1
- package/dist/persistence-lifecycle.d.ts +103 -0
- package/dist/persistence-lifecycle.js +202 -0
- package/dist/provider-events.d.ts +1 -0
- package/dist/provider-events.js +6 -1
- package/dist/provider-request-policy.js +3 -4
- package/dist/providers/media.d.ts +1 -1
- package/dist/providers/openai-compatible.d.ts +46 -1
- package/dist/providers/openai-compatible.js +123 -53
- package/dist/providers/openai-primitives.js +10 -7
- package/dist/providers/transport.d.ts +6 -0
- package/dist/providers/transport.js +21 -0
- package/dist/providers.d.ts +2 -0
- package/dist/providers.js +3 -0
- package/dist/redaction.d.ts +1 -0
- package/dist/redaction.js +26 -9
- package/dist/resources.d.ts +2 -2
- package/dist/resources.js +2 -2
- package/dist/retry.d.ts +5 -0
- package/dist/retry.js +8 -1
- package/dist/rpc.js +55 -11
- package/dist/run-ledger.d.ts +6 -0
- package/dist/run-ledger.js +16 -13
- package/dist/run-limits.js +49 -10
- package/dist/secure-agent.js +8 -2
- package/dist/security.js +7 -2
- package/dist/session-stores.d.ts +7 -2
- package/dist/session-stores.js +195 -21
- package/dist/skill-disclosure.d.ts +35 -0
- package/dist/skill-disclosure.js +101 -0
- package/dist/skill-load.d.ts +25 -0
- package/dist/skill-load.js +112 -0
- package/dist/structured-output.d.ts +5 -1
- package/dist/structured-output.js +20 -2
- package/dist/system-prompts.js +7 -2
- package/dist/testing/agent-event-source-conformance.d.ts +4 -0
- package/dist/testing/agent-event-source-conformance.js +54 -0
- package/dist/testing/compaction-conformance.js +5 -1
- package/dist/testing/extension-conformance.js +15 -3
- package/dist/testing/feedback.d.ts +1 -3
- package/dist/testing/feedback.js +1 -1
- package/dist/testing/persistence-schema.d.ts +2 -2
- package/dist/testing/persistence-schema.js +280 -35
- package/dist/testing/provider-conformance.js +3 -3
- package/dist/testing/run-ledger-conformance.js +1 -1
- package/dist/testing/session-store-conformance.d.ts +6 -0
- package/dist/testing/session-store-conformance.js +37 -2
- package/dist/testing/tool-conformance.js +30 -5
- package/dist/testing/tool-effect-store-conformance.d.ts +9 -0
- package/dist/testing/tool-effect-store-conformance.js +85 -0
- package/dist/thinking.js +4 -1
- package/dist/tool-effects.d.ts +15 -0
- package/dist/tool-effects.js +352 -0
- package/dist/tool-result-fold.d.ts +40 -0
- package/dist/tool-result-fold.js +176 -0
- package/dist/tools.d.ts +8 -3
- package/dist/tools.js +248 -13
- package/docs/0.1.0-readiness.md +202 -0
- package/docs/a2a.md +33 -2
- package/docs/acp.md +126 -0
- package/docs/ag-ui-adoption.md +77 -0
- package/docs/ag-ui.md +225 -0
- package/docs/agent-events.md +34 -3
- package/docs/agent-identity.md +144 -0
- package/docs/agent-loops.md +17 -2
- package/docs/agent-session-runtime.md +21 -4
- package/docs/browser-automation.md +5 -0
- package/docs/caveman.md +129 -0
- package/docs/cli-rpc.md +3 -6
- package/docs/coding-agent-tools.md +229 -25
- package/docs/coding-security.md +77 -11
- package/docs/compaction-and-retry.md +5 -2
- package/docs/compaction-llm.md +20 -1
- package/docs/compaction-observational-memory.md +52 -8
- package/docs/context-and-skills.md +94 -7
- package/docs/contribution-registries.md +1 -0
- package/docs/conversations.md +135 -0
- package/docs/credential-storage.md +34 -1
- package/docs/credentials-and-redaction.md +11 -1
- package/docs/database-persistence.md +27 -7
- package/docs/device-adapters.md +97 -0
- package/docs/enterprise-postgres-state.md +178 -0
- package/docs/evaluations.md +14 -1
- package/docs/extensions.md +4 -1
- package/docs/forge-integration.md +113 -0
- package/docs/guardrails.md +16 -2
- package/docs/host-security.md +35 -4
- package/docs/index.md +69 -37
- package/docs/input-and-prompt-assembly.md +8 -7
- package/docs/language-intelligence.md +162 -0
- package/docs/mcp-tools.md +62 -5
- package/docs/middleware-hooks.md +2 -2
- package/docs/migration.md +423 -2
- package/docs/model-routing.md +111 -0
- package/docs/multimodal-content.md +8 -5
- package/docs/node-jsonl-session-store.md +1 -1
- package/docs/observability.md +2 -0
- package/docs/openapi-tools.md +56 -0
- package/docs/performance.md +282 -0
- package/docs/policy-and-audit.md +171 -0
- package/docs/ponytail.md +127 -0
- package/docs/postgres-persistence.md +8 -4
- package/docs/process-sessions.md +147 -0
- package/docs/provider-caching.md +13 -1
- package/docs/provider-conformance.md +29 -5
- package/docs/provider-packages.md +43 -2
- package/docs/provider-request-policies.md +2 -0
- package/docs/providers/ai-sdk.md +24 -7
- package/docs/providers/alibaba.md +179 -0
- package/docs/providers/anthropic.md +93 -0
- package/docs/providers/azure.md +74 -0
- package/docs/providers/bedrock.md +72 -0
- package/docs/providers/google.md +89 -0
- package/docs/providers/ollama.md +166 -0
- package/docs/providers/openai-compatible.md +31 -2
- package/docs/providers/openai.md +24 -5
- package/docs/providers/openrouter.md +2 -0
- package/docs/providers/vertex.md +71 -0
- package/docs/public-contracts.md +61 -4
- package/docs/rag.md +41 -12
- package/docs/release-and-install.md +323 -206
- package/docs/resource-loading.md +3 -0
- package/docs/runs-and-usage.md +3 -0
- package/docs/server.md +44 -6
- package/docs/session-store-conformance.md +2 -0
- package/docs/session-stores.md +41 -2
- package/docs/sqlite-persistence.md +11 -3
- package/docs/structured-output.md +7 -1
- package/docs/supervisors.md +8 -0
- package/docs/tool-effects.md +95 -0
- package/docs/tools.md +5 -0
- package/docs/work-artifacts-and-review.md +102 -0
- package/docs/work-connectors.md +32 -0
- package/docs/work-tools.md +137 -0
- package/docs/workflows.md +6 -0
- package/docs/working-and-semantic-memory.md +40 -7
- package/package.json +30 -8
- package/templates/init/providers.json +22 -0
- package/docs/review-coverage-2026-07-14.md +0 -260
- package/docs/review-coverage-2026-07-15.md +0 -193
- package/docs/review-coverage-2026-07-17-provider-validation.md +0 -192
- package/docs/review-coverage-2026-07-19-phase-3.md +0 -174
- package/docs/review-coverage-2026-07-20-phase-4.md +0 -175
package/docs/ag-ui.md
ADDED
|
@@ -0,0 +1,225 @@
|
|
|
1
|
+
# Frontend interoperability (AG-UI and ACP)
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-ag-ui` is an optional, framework-free protocol adapter over Prism's existing redacted `AgentEvent`, session, durable-run, and persistence seams.
|
|
6
|
+
|
|
7
|
+
- Root export maps Prism events to AG-UI `@ag-ui/core` **0.0.57** events and offers `createAgUiHandler()` (`Request` → SSE `Response`), compatible `createPersistenceAgUiReplay()` pages, distributed `createAgentEventSourceAgUiReplay()` follow, and explicit `createAgUiMcpAdapter()` / `createAgUiMcpAppHandler()` / `createAgUiA2AAdapter()` protocol handshakes.
|
|
8
|
+
- `@arnilo/prism-ag-ui/acp` is the stable ACP **v1** sibling: `createAcpEventMapper()` and `createPrismAcpAgent()` over `@agentclientprotocol/sdk` **1.3.0** root exports. ACP is a protocol adapter — sessions, modes, MCP, fs/terminal, lifecycle mapping, and caps live on the host seams. See [ACP coding-host interop](acp.md) for the full reference; this page covers AG-UI only.
|
|
9
|
+
- Core remains protocol-free. `resumeAgentRunStream()` / `AgentRunLifecycle.resumeStream()` are generic durable-resume streams shared by adapters.
|
|
10
|
+
|
|
11
|
+
## When to use it
|
|
12
|
+
|
|
13
|
+
Use AG-UI when a host already authenticates users, owns sessions and durable run correlation, and needs a bounded Web endpoint for a browser/TUI/desktop client. Use ACP when an editor client already supplies an ACP transport and needs text, safe tool status, usage, and approval updates from a Prism session.
|
|
14
|
+
|
|
15
|
+
Use [A2A interoperability](a2a.md) for remote agent-to-agent JSON-RPC/HTTPS tasks. AG-UI/ACP are frontend/client protocol adapters; neither replaces A2A task lifecycle or storage.
|
|
16
|
+
|
|
17
|
+
## Inputs / request
|
|
18
|
+
|
|
19
|
+
Install the optional package beside the core runtime (it becomes publishable with the 0.0.12 release graph):
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
npm install @arnilo/prism @arnilo/prism-ag-ui
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
`createAgUiHandler()` takes host-owned callbacks:
|
|
26
|
+
|
|
27
|
+
| Input | Purpose |
|
|
28
|
+
| --- | --- |
|
|
29
|
+
| `authorize` | Rebinds untrusted AG-UI thread/run selectors to host ownership on every request. `false` returns 403. |
|
|
30
|
+
| `sessionFactory` | Returns an authorized Prism `AgentSession`; it receives only host-approved `AgUiPreparedInput`, never raw client tools/state. |
|
|
31
|
+
| `input.project` | Opts into full `RunAgentInput`; turns bounded, still-untrusted history, state, context, forwarded props, media, and lineage into host-selected Prism `Message` values. Omit for legacy final-text mode. |
|
|
32
|
+
| `input.frontendTools` | Explicitly selects client-side handoffs. Returned names must be request-tool subset; adapter never turns JSON tool declarations into Prism `ToolDefinition`s. |
|
|
33
|
+
| `mcp` | Optional `createAgUiMcpAdapter({ bridge, select })`; host selects reviewed `bridge.tools`, then `sessionFactory` receives them as `input.serverTools`. Normal Prism dispatch/loop remains sole executor. |
|
|
34
|
+
| `a2a` | Optional `createAgUiA2AAdapter({ client, select, correlate })`; verified remote A2A task stream replaces this handler's local session only. Host selection/correlation binds each remote task to ownership. |
|
|
35
|
+
| `lifecycle` + `resolveRun` | Optional durable status/resume path. Required only for a resumed interruption. |
|
|
36
|
+
| `interrupts.resume` | Optional host aggregate-policy callback for multiple AG-UI interrupts; it returns one current-version core approve/deny decision. |
|
|
37
|
+
| `replay` | Optional page adapter or `createAgentEventSourceAgUiReplay(source, options)` for gap-free distributed replay/live follow. |
|
|
38
|
+
| `projection` | Explicit safe tool/state/messages/activity/reasoning/raw/custom/interrupt projection. Omit each callback for default deny. Prefer `composeAgUiProjections(createMessagesFromSessionProjection(...), createStateFromStoreProjection(...), createActivityFromToolProgressProjection(), host)` for standard families. |
|
|
39
|
+
| `a2ui` | Opt-in A2UI painting middleware (`{ catalogId, mode, renderToolName?, allowedCatalogIds?, limits? }`). Detects `a2ui_operations` tool results and/or streams from `render_a2ui` args; paints `a2ui-surface` activity events. Absent = inert. |
|
|
40
|
+
| `capabilities` | Optional host declaration narrowed to implemented SSE/projector/lifecycle features; read `handler.capabilities`. |
|
|
41
|
+
| `redactor`, `limits` | Host redaction and narrowing-only finite caps. |
|
|
42
|
+
|
|
43
|
+
The handler accepts only `POST` JSON validated with official AG-UI `RunAgentInputSchema`. Every aggregate is bounded before a callback runs. With no `input.project`, it preserves compatibility: final text user message only; non-empty state or frontend tools fail before authorization/session lookup. With a projector, all current roles/history, context, state, forwarded props, multimodal parts, parent lineage, and tool-result continuations are available as untrusted input. The projector must apply Prism media URL/SSRF/MIME policy before forwarding media. Start a run with no `resume` and no `?cursor=`; replay supplies `?cursor=`.
|
|
44
|
+
|
|
45
|
+
## Outputs / response / events
|
|
46
|
+
|
|
47
|
+
The handler returns `text/event-stream`, one `data: <AG-UI event>\n\n` frame per output. Mapper lifecycle is ordered: `RUN_*`, `STEP_*`, `TEXT_MESSAGE_*`, and `TOOL_CALL_*` are deterministic Prism mappings. Host projectors may additionally prove and emit `STATE_SNAPSHOT`/`STATE_DELTA`, `MESSAGES_SNAPSHOT`, `ACTIVITY_*`, current `REASONING_*`, `RAW`, and named `CUSTOM` values. All values revalidate against official `EventSchemas`; deprecated `THINKING_*` and convenience chunk events are not produced. Active message/tool/reasoning/step sequences close before error, interruption, or finish.
|
|
48
|
+
|
|
49
|
+
A Prism durable `agent_suspended` returns `RUN_FINISHED` with core interrupt id `${runId}:${version}` and a strict `{ decision: "approve" | "deny" }` schema. `projection.interrupt` may attach bounded expiry/metadata or additional host policy interrupts but must retain that core id. Without `interrupts.resume`, one exact entry is required; `cancelled` means deny. An aggregate policy may validate bounded multiple entries, then returns one current-version core decision. Payloads containing `editedArgs`/`args` always deny: Prism does not mutate persisted tool calls. The adapter checks host authorization, selected run, suspended status, and checkpoint version, then calls `AgentRunLifecycle.resumeStream()` once. Claimed/dispatched tools are never replayed.
|
|
50
|
+
|
|
51
|
+
`createPersistenceAgUiReplay()` remains a compatible page adapter. `createAgentEventSourceAgUiReplay()` resolves exact ownership/run once per open, then consumes the shared durable source through terminal or live follow; it never attaches replica-local `session.subscribe()`. Every record must already be redacted. Mapped events carry stable `prismEventId` and bounded opaque `prismCursor`; records with no standard mapping emit `CUSTOM prism.replay_cursor`, so clients can persist progress. Terminal replay never creates a session or reruns a provider/tool.
|
|
52
|
+
|
|
53
|
+
ACP maps assistant text to `agent_message_chunk`, safe tool lifecycle to `tool_call`/`tool_call_update` (locations/diffs only from projection allow-lists), provider usage to `usage_update`, and durable suspension to `session/request_permission` with the four shared outcomes — plus, per the client's advertised capabilities, editor-buffer fs, terminal, prompt media, modes/config options, lifecycle events, and elicitation. Advertisement is a pure function of the wired seams: no seam, no capability, no method. Only `allow_once` approves; reject, cancellation, unknown outcomes, and request failure deny. See [ACP coding-host interop](acp.md).
|
|
54
|
+
|
|
55
|
+
Selected MCP tools use normal `TOOL_CALL_*` core dispatch. Linked Apps add safe `mcp-apps` activity; app-only tools stay model-hidden. The separate reauthorizing Apps proxy allow-lists initialize/ping/logging/tool/resource calls for one bridge; its sandbox helper returns CSP/iframe config and never executes HTML.
|
|
56
|
+
|
|
57
|
+
### UI-initiated mutation retry through `ToolEffectStore` (FR-4)
|
|
58
|
+
|
|
59
|
+
`createAgUiMcpAppHandler` accepts an optional `effectStore` (Phase 7 `ToolEffectStore`) plus `effectContext` (identity/ownership; falls back to `authorization.ownership` + `context.identity`). Every approved `tools/call` then records `begin` → `markDispatched` → `complete`/`fail`/`markUnknown` in the store; effect keys derive from identity + ownership + tool name + arguments hash (`deriveAppEffectKey`). The proxy **never auto-retries** — the host decides:
|
|
60
|
+
|
|
61
|
+
```ts
|
|
62
|
+
import { createAgUiMcpAppHandler, reconcileAppEffect } from "@arnilo/prism-ag-ui";
|
|
63
|
+
|
|
64
|
+
const handler = createAgUiMcpAppHandler({
|
|
65
|
+
apps, authorize, context, approveToolCall, allowedOrigins,
|
|
66
|
+
effectStore, // optional: records UI mutations for idempotent retry
|
|
67
|
+
effectContext: ({ authorization, context }) => ({ identity: context.identity, ownership: authorization.ownership }),
|
|
68
|
+
});
|
|
69
|
+
|
|
70
|
+
// After transport/abort loss the record is `unknown`; the host verifies the
|
|
71
|
+
// actual outcome and resolves it (claim/CAS), then the UI can retry idempotently:
|
|
72
|
+
await reconcileAppEffect({ effectStore, identity, ownership, sessionId, runId, toolName, arguments: args, outcome: "completed", result });
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
A retried call whose record is `completed` replays the recorded result without re-dispatching; `failed_retryable`/`failed_terminal`/`dispatched`/`unknown` records fail closed with a `409` until the host reconciles. Wrong-owner or unresolvable identity/ownership fails closed; absent `effectStore` keeps 0.0.25 behavior exactly.
|
|
76
|
+
|
|
77
|
+
`createAgUiA2AAdapter()` maps verified task text/activity. Non-text/tool/A2UI parts need host `projectPart`; non-streaming fallback accepts only a terminal task, otherwise host follows saved correlation. `createAgUiA2AServer()` fronts a local AG-UI agent as an A2A 1.0 server for remote A2A clients (reverse direction; see [A2A interoperability](a2a.md)).
|
|
78
|
+
|
|
79
|
+
Co-work uses bounded, redacted `CUSTOM prism.cowork.*` events through `mapCoWork()` / `createCoWorkReplay()`; see [Work artifacts and review](work-artifacts-and-review.md).
|
|
80
|
+
|
|
81
|
+
## Request/response example
|
|
82
|
+
|
|
83
|
+
Resume a default single interrupt with `resume: [{ "interruptId": "run-1:4", "status": "resolved", "payload": { "decision": "approve" } }]`. Full history, client tool results, and mutable state need authorized `input.project` selection. This adapter is not a conversation database.
|
|
84
|
+
|
|
85
|
+
## Implementation example
|
|
86
|
+
|
|
87
|
+
```ts
|
|
88
|
+
import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
|
|
89
|
+
import { createAgentEventSourceAgUiReplay, createAgUiHandler } from "@arnilo/prism-ag-ui";
|
|
90
|
+
|
|
91
|
+
const agent = createAgent({
|
|
92
|
+
model: { provider: "mock", model: "offline" },
|
|
93
|
+
provider: createMockProvider([providerTextDelta("ready"), providerDone()]),
|
|
94
|
+
});
|
|
95
|
+
|
|
96
|
+
const replay = createAgentEventSourceAgUiReplay(persistence.events, {
|
|
97
|
+
resolveRun,
|
|
98
|
+
ownership: (authorization) => authorization.ownership,
|
|
99
|
+
});
|
|
100
|
+
|
|
101
|
+
const handle = createAgUiHandler({
|
|
102
|
+
authorize: ({ request }) => request.headers.get("authorization") === "Bearer host-checked"
|
|
103
|
+
? { ownership: { userId: "user-1" } }
|
|
104
|
+
: false,
|
|
105
|
+
// Create a run-scoped agent/session using input.serverTools when mcp is enabled.
|
|
106
|
+
sessionFactory: () => agent.createSession({ id: "host-owned-thread" }),
|
|
107
|
+
replay,
|
|
108
|
+
input: {
|
|
109
|
+
project: () => ({
|
|
110
|
+
messages: [{ id: "user", role: "user", content: [{ type: "text", text: "host-selected input" }] }],
|
|
111
|
+
}),
|
|
112
|
+
},
|
|
113
|
+
projection: { toolArguments: () => undefined, toolResult: () => undefined },
|
|
114
|
+
});
|
|
115
|
+
|
|
116
|
+
const response = await handle(request); // adapt this Web Response in host framework
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
See runnable network-free [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts). For ACP, construct `createPrismAcpAgent({ authorize, sessionFactory, lifecycle, ...seams })` and connect the returned stable SDK agent through the host's ACP transport — see [`examples/acp-coding-host.ts`](../examples/acp-coding-host.ts) and [ACP coding-host interop](acp.md).
|
|
120
|
+
|
|
121
|
+
## Extension and configuration notes
|
|
122
|
+
|
|
123
|
+
All identity, authorization, session/thread mapping, durable checkpoint lookup, persistence selection, replay cursor persistence, transport adaptation, MCP bridge/card configuration, app sandbox DOM, remote A2A task correlation, and optional projection are host-owned. The adapter owns no listener, database, background reconnect loop, credential resolver, or UI state.
|
|
124
|
+
|
|
125
|
+
`AgUiProjection` is an allow-list. Without a callback, raw tool arguments/results/progress, arbitrary state/patches/transcripts/activity/reasoning/raw events, paths, ACP locations/diffs/terminals/raw I/O, and frontend-supplied tools remain absent. Reasoning signatures do not become AG-UI encrypted values automatically: a host must explicitly provide an already client-encrypted opaque value. `input.project` is also an allow-list: do not merge client state/forwarded props into ownership, identity, tools, permissions, provider options, or media fetch policy.
|
|
126
|
+
|
|
127
|
+
### Reasoning encrypted-value helper (FR-3)
|
|
128
|
+
|
|
129
|
+
`createReasoningEncryptedValue({ encrypt, content, event, maxBytes? })` produces the `encryptedValue` fragment for the `reasoning` projection callback (AG-UI `REASONING_ENCRYPTED_VALUE`):
|
|
130
|
+
|
|
131
|
+
```ts
|
|
132
|
+
import { createReasoningEncryptedValue } from "@arnilo/prism-ag-ui";
|
|
133
|
+
|
|
134
|
+
const mapper = createAgUiEventMapper({
|
|
135
|
+
projection: {
|
|
136
|
+
reasoning: (content, event) =>
|
|
137
|
+
createReasoningEncryptedValue({ encrypt: hostEncryptForClient, content, event }),
|
|
138
|
+
},
|
|
139
|
+
});
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
`encrypt` is host-owned (client key) and receives the redacted `ThinkingContent` and the Prism event; return `undefined` to decline. The helper is synchronous and pure like the other projection callbacks: it never infers an encrypted value from a Prism reasoning signature, fails closed (returns `undefined`) when `encrypt` is missing, throws, or returns a non-string, and truncates output to `maxBytes` (default `DEFAULT_MAX_REASONING_BYTES`, clamped to `HARD_MAX_REASONING_BYTES`). The mapper additionally caps the emitted value at the resolved `maxReasoningBytes` limit.
|
|
143
|
+
|
|
144
|
+
Co-work projection reuses the same allow-list: `AgUiProjection.coWork(event)` may return a curated, JSON-serializable payload for a co-work event; absent it, the redacted event fields are exposed. Wire `coWorkContext` to derive thread/artifact/identity from the authorized request (never client JSON) and `coWork` to a `createCoWorkReplay()` over your durable artifact/draft/snapshot stores. The handler projects one bounded page after the run; mount a dedicated cursor-paged co-work endpoint when full pagination is needed.
|
|
145
|
+
|
|
146
|
+
Durable interrupts carry the shared decision batch: the fallback interrupt includes the redacted `pendingDecisions` under `metadata` and its `responseSchema` accepts either the legacy `{ decision: "approve" | "deny" }` or a `{ decisions: [{ approvalId, outcome, reason?, modifiedArguments?, elicitation? }] }` batch. All batch entries are shape- and cap-validated at the boundary (count ≤ 128, ids ≤ 128 chars, four outcomes, reason ≤ 8 KiB, payloads ≤ 64 KiB) and core re-validates each against the recorded pending set under the single CAS. `interrupts.resume` may return the batch form (`{ decisions, expectedVersion? }`); legacy `editedArgs` resume payloads still deny. ACP permission prompts offer the four outcomes (`allow_once` / `allow_always` / `reject_once` / `reject_always`) and map them onto the batch; a cancelled prompt stays deny-closed.
|
|
147
|
+
|
|
148
|
+
### Standard projectors (opt-in)
|
|
149
|
+
|
|
150
|
+
Three batteries-included factories return `AgUiProjection` fragments. Compose with host projectors via `composeAgUiProjections(...fragments)` — **first defined callback wins** (left to right); `undefined` fragments are skipped. Absent factories keep 0.0.24 default-deny.
|
|
151
|
+
|
|
152
|
+
```ts
|
|
153
|
+
createAgUiHandler({
|
|
154
|
+
projection: composeAgUiProjections(
|
|
155
|
+
// async transcript source: AgentSession.entries() is async
|
|
156
|
+
createMessagesFromSessionProjection({
|
|
157
|
+
getMessages: async () => (await session.entries()).map(entryToAgUiMessage),
|
|
158
|
+
redact,
|
|
159
|
+
}),
|
|
160
|
+
createStateFromStoreProjection(runStateStore),
|
|
161
|
+
createActivityFromToolProgressProjection(),
|
|
162
|
+
hostCustom,
|
|
163
|
+
),
|
|
164
|
+
});
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
Every `AgUiProjection` callback may return a promise (types are `Awaitable<T>` — Task 15, 0.0.26), so projectors can call async host APIs like `session.entries()` directly. Sync-only hosts keep exact prior behavior: sync return values short-circuit, hooks are awaited strictly in event order (never `Promise.all`), and a rejected hook fails closed per event (omitted value, stream continues) exactly like today's sync throw handling. `createMessagesFromSessionProjection({ getMessages })` accepts an async transcript source and emits `MESSAGES_SNAPSHOT` from it at `agent_started` and `message_finished` (no sync `getMessages` needed for full session history); `agent_finished` is terminal and the mapper projects nothing after it, so the final snapshot arrives at the last `message_finished`.
|
|
168
|
+
|
|
169
|
+
| Factory | Emits | Notes |
|
|
170
|
+
| --- | --- | --- |
|
|
171
|
+
| `createMessagesFromSessionProjection` | `MESSAGES_SNAPSHOT` | Host `getMessages()` for authorized history (sync or async), or live `message_finished` accumulation. Caps 128/1024. Redact drops closed. |
|
|
172
|
+
| `createStateFromStoreProjection(store)` | `STATE_SNAPSHOT` on `agent_started`; RFC 6902 `STATE_DELTA` (add/replace/remove) when `store.get()` changes | Host store; optional `subscribe` only marks dirty — no Prism watcher. Oversized/throw → drop closed. |
|
|
173
|
+
| `createActivityFromToolProgressProjection` | `ACTIVITY_SNAPSHOT` / `ACTIVITY_DELTA` from `tool_execution_progress` | Default `activityType: "tool-progress"`. Missing progress+metadata → drop closed. |
|
|
174
|
+
|
|
175
|
+
### A2UI painting middleware (opt-in)
|
|
176
|
+
|
|
177
|
+
`createAgUiHandler({ a2ui: { catalogId, mode } })` paints A2UI v0.9 surfaces without a host `projection.activity` callback:
|
|
178
|
+
|
|
179
|
+
| Mode | Source | Paint |
|
|
180
|
+
| --- | --- | --- |
|
|
181
|
+
| `fixed-schema` | Tool result `{ a2ui_operations: [...] }` | First batch with `createSurface` → `ACTIVITY_SNAPSHOT` (`activityType: "a2ui-surface"`); later batches → `ACTIVITY_DELTA` |
|
|
182
|
+
| `streaming` | `tool_call_delta` args of `renderToolName` (default `render_a2ui`) | Progressive `ACTIVITY_SNAPSHOT` with `replace: true` only when complete ops extractable — never partial JSON |
|
|
183
|
+
| `both` | Both paths; streamed surfaces are not re-painted from the final envelope |
|
|
184
|
+
|
|
185
|
+
Host `catalogId` is stamped when absent; model-supplied ids outside `allowedCatalogIds` (default `[catalogId]`) are overwritten. Invalid ops emit one bounded `CUSTOM` event `prism.a2ui.error` and paint nothing. Caps: 64/512 ops per message, 64 KiB/1 MiB per op, 16/64 surfaces per run, depth 32/64.
|
|
186
|
+
|
|
187
|
+
User actions arrive as untrusted `AgUiA2UiAction` values on `input.project({ a2uiActions })` (from `forwardedProps.a2uiAction` or activity/tool-result shapes). Without `input.project` they stay default-deny — Prism never synthesizes a `log_a2ui_event` tool call (documented divergence from official `@ag-ui/a2ui-middleware`). Example: `examples/ag-ui-a2ui.ts`.
|
|
188
|
+
|
|
189
|
+
### Reference frontend renderer (Task 14, 0.0.26)
|
|
190
|
+
|
|
191
|
+
`@arnilo/prism-ag-ui/renderer` ships a framework-free client renderer for AG-UI/A2UI surfaces: it consumes an AG-UI event stream (SSE via `@ag-ui/client`, or any `AsyncIterable`) and renders `a2ui-surface` activity snapshots/deltas into DOM surfaces from a host component catalog. No framework dependency, no host build step, no jsdom (tests use an in-memory DOM stub).
|
|
192
|
+
|
|
193
|
+
```ts
|
|
194
|
+
import { createA2UiRenderer } from "@arnilo/prism-ag-ui/renderer";
|
|
195
|
+
|
|
196
|
+
const renderer = createA2UiRenderer({
|
|
197
|
+
stream: agUiEventStream, // SSE or AsyncIterable of AGUIEvent
|
|
198
|
+
catalog: myComponents, // optional; defaults to Text/Container/Column/Row/Button
|
|
199
|
+
onAction: (action) => sendA2UiAction(action), // optional: Button clicks etc.
|
|
200
|
+
onError: (error) => console.warn(error.code, error.message),
|
|
201
|
+
});
|
|
202
|
+
const surface = await renderer.surface("chat"); // detached DOM node, kept in sync
|
|
203
|
+
```
|
|
204
|
+
|
|
205
|
+
The core is a DOM-free state machine (`reduceA2UiOps`): operations become a surface/component model (adjacency list with `id`/`component`/flat props, A2UI v0.9 JSON-Pointer data model, `deleteSurface`); a thin binding layer renders the model through catalog component renderers (framework-free `(props, ctx, dom) => node` functions). The core is also exported as values from the subpath entry (`A2UiSurfaceState`, `reduceA2UiOps`, `readA2UiBatch`, `resolvePointer`, `A2UI_VERSION`, 0.0.27, Synapta FR) so framework hosts can drive the validated surface state machine and own the view layer; behavior and frozen caps are unchanged. Snapshots replace a surface's model (streaming mode sends cumulative ops); RFC 6902 deltas append. The same frozen caps as the server painter are enforced client-side: 64/512 ops per message, 64 KiB/1 MiB per op, 16/64 surfaces per run, depth 32/64. Invalid or oversized ops drop closed with one bounded `prism.a2ui.error` event (host logging via `onError`); unknown catalog components render an explicit placeholder. The renderer never executes remote HTML: only `createElement`/`createTextNode`/`appendChild`, no HTML-string assignment, no dynamic code evaluation. Data bindings `{"path": "/pointer"}` resolve against the per-surface data model; `deleteSurface` detaches content. The main `@arnilo/prism-ag-ui` entry stays runtime-agnostic — DOM code lives only behind the `renderer` subpath (the root entry re-exports renderer types only, no values). Hosts embedding it should follow the MCP Apps CSP/sandbox guidance (`docs/ag-ui-adoption.md`) for iframe/worker placement.
|
|
206
|
+
|
|
207
|
+
## Security and performance notes
|
|
208
|
+
|
|
209
|
+
Authorize every start/replay/resume/proxy/follow. Treat protocol fields, MCP metadata/HTML, and A2A cards/parts as untrusted; persist run/task correlation before output and redact streams. MCP Apps requires extension acknowledgement, exact proxy origin, same-bridge visibility, approval, `ui://` HTML/MIME bounds, and sandbox CSP. It never retries UI mutations; with an `effectStore` it records them for host-driven idempotent retry and unknown-outcome reconciliation (FR-4).
|
|
210
|
+
|
|
211
|
+
Defaults / hard caps: request 64 KiB / 1 MiB; input 128 / 1024 messages, 32 / 256 tools/contexts, 8 / 64 interrupts, 16 / 64 media parts, and 64 KiB / 1 MiB text/state/media; frontend tool/context payloads 16 KiB / 256 KiB; projected event/state/activity/reasoning/raw values 64 KiB / 1 MiB; patches 128 / 4096 operations; JSON depth 16 / 64, properties 128 / 4096, arrays 512 / 8192; cursor 4 / 16 KiB; replay page 100 / 500 records; queue 128 / 4096 events; stream 10,000 / 100,000 events and 10 / 64 MiB; wall time 120 seconds / 30 minutes. Overflow yields a bounded error/closed stream, not an unbounded queue. SSE is declared; WebSocket/protobuf/push are not. Reconnect is at-least-once, so clients de-duplicate stable event/message/tool IDs.
|
|
212
|
+
|
|
213
|
+
## Related APIs
|
|
214
|
+
|
|
215
|
+
- [Agent/session runtime](agent-session-runtime.md): `session.stream()`, `resumeAgentRunStream()`, and durable lifecycle.
|
|
216
|
+
- [Agent events](agent-events.md): normalized source events and ledger redaction.
|
|
217
|
+
- [Runs and usage ledger](runs-and-usage.md): durable `AgentEventRecord` query source.
|
|
218
|
+
- [Web-standard server handler](server.md): generic Prism HTTP API, separate from AG-UI.
|
|
219
|
+
- [A2A interoperability](a2a.md): remote agent-to-agent tasks, not frontend protocol mapping.
|
|
220
|
+
- [AG-UI adoption evaluation](ag-ui-adoption.md): official 0.0.57 event/input matrix and shipped explicit MCP/MCP Apps/A2A handshakes.
|
|
221
|
+
- [ACP coding-host interop](acp.md): the full ACP reference — seam-based capability advertisement, session modes/config, MCP select, fs/terminal adapters, lifecycle mapping, elicitation, and caps.
|
|
222
|
+
- [MCP bridge/server](mcp-tools.md): `mcpApps` negotiation, bounded resources, and remote tool trust.
|
|
223
|
+
- [A2A interoperability](a2a.md): verified rich task client and remote task lifecycle.
|
|
224
|
+
- [Host security guide](host-security.md): authorization, ownership, redaction, and credential boundaries.
|
|
225
|
+
- [Work artifacts and review](work-artifacts-and-review.md): durable artifact service that produces the co-work approval/progress/download-link events projected here.
|
package/docs/agent-events.md
CHANGED
|
@@ -10,7 +10,7 @@ Events are emitted by the runtime and by loops through `LoopContext.emit`, both
|
|
|
10
10
|
|
|
11
11
|
Subscribe via `session.stream()` for a single owned run, or `session.subscribe()` when a host needs a long-lived observer across runs: render streamed assistant text in a UI, react to tool execution, drive observability/telemetry, or audit artifact validation outcomes. Do not parse provider stream events directly for these — `AgentEvent` is the stable, normalized surface across providers and loops.
|
|
12
12
|
|
|
13
|
-
Do not use `
|
|
13
|
+
Do not use live `session.subscribe()` for cross-replica reconnect — use durable `AgentEventSource` below. Live subscribe remains process-local.
|
|
14
14
|
|
|
15
15
|
## Durable event ledger
|
|
16
16
|
|
|
@@ -18,6 +18,35 @@ When `AgentConfig.runLedger` or `RunOptions.runLedger` is configured, every emit
|
|
|
18
18
|
|
|
19
19
|
Event records preserve emission order within a run because the runtime drains pending event appends before writing the final `RunRecord`. Subscribers still see the live, in-memory stream; the ledger is the durable copy.
|
|
20
20
|
|
|
21
|
+
## Durable AgentEventSource
|
|
22
|
+
|
|
23
|
+
`AgentEventSource` (`createMemoryAgentEventSource` / `persistence.events` on PostgreSQL) appends, pages, and subscribes with opaque ownership-bound cursors. `subscribe` registers wake interest before replaying history so replay-to-live handoff has no gap. Delivery is at-least-once; consumers dedupe `record.id`. PostgreSQL uses transactional sequence allocation plus `LISTEN`/`NOTIFY` wakeups with polling fallback. Transport adapters (server SSE `Last-Event-ID`, AG-UI, A2A `afterEventId`) map source envelopes only — they do not invent private replay loops. This is not exactly-once.
|
|
24
|
+
|
|
25
|
+
### Placement (FR-7 answer, 0.0.26)
|
|
26
|
+
|
|
27
|
+
The durable `AgentEventSource` **stays in `@arnilo/prism-session-store-postgres`** for the 0.0.26 line and is importable from the package root (FR-6):
|
|
28
|
+
|
|
29
|
+
```ts
|
|
30
|
+
import { createPostgresAgentEventSource } from "@arnilo/prism-session-store-postgres";
|
|
31
|
+
const source = createPostgresAgentEventSource({ pool, schema: "prism", cursorSecret });
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
PostgreSQL `LISTEN`/`NOTIFY` remains the **reference durable implementation**; `createPostgresPersistence` still bundles the same source as `persistence.events` (the canonical path — no behavior change). The standalone root export exists for consumers that want a durable source without full persistence. The migration path from the 0.0.24/0.0.25 API is: `persistence.events` and `createPostgresAgentEventSource` both keep working unchanged; a future relocation (if any) ships a replacement export with a deprecation note before removing the old one. See [migration](migration.md) `0.0.25 → 0.0.26` and the FR-6/FR-7 record `prism-agent-event-source-export-and-location.md`.
|
|
35
|
+
|
|
36
|
+
### NATS JetStream adapter (FR-5)
|
|
37
|
+
|
|
38
|
+
`@arnilo/prism-session-store-nats` ships a sibling durable `AgentEventSource` over NATS JetStream for JetStream backbones (Postgres remains the reference implementation):
|
|
39
|
+
|
|
40
|
+
```ts
|
|
41
|
+
import { connect } from "@nats-io/transport-node";
|
|
42
|
+
import { createNatsAgentEventSource, createNatsJetStream } from "@arnilo/prism-session-store-nats";
|
|
43
|
+
|
|
44
|
+
const nc = await connect({ servers: process.env.NATS_URL });
|
|
45
|
+
const source = createNatsAgentEventSource({ connection: await createNatsJetStream(nc), stream: "prism_agent_events" });
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
One subject per run (`prism.agent-events.<tenant>.<session>.<run>`); the JetStream per-subject sequence is the per-run event sequence. `append` is idempotent by `record.id` within the stream's dedupe window; `page`/`subscribe` replay per subject from HMAC-signed cursors; `subscribe` uses a durable pull consumer with explicit acks (at-least-once, 30s redelivery, dedupe by `record.id`); `cleanup` deletes ownership-scoped messages older than `before`. The host provisions the stream (subjects `prism.agent-events.>`, retention limits, dedupe window). Inert on import; network-free tests use an in-memory fake of the narrow `NatsJetStream` seam.
|
|
49
|
+
|
|
21
50
|
## Inputs / request
|
|
22
51
|
|
|
23
52
|
```ts
|
|
@@ -44,7 +73,7 @@ The `AgentEvent` union (grouped by concern):
|
|
|
44
73
|
| Assistant messages | `message_started`, `message_delta`, `message_finished` |
|
|
45
74
|
| Tool execution | `tool_execution_started`, `tool_execution_progress`, `tool_execution_finished`, `tool_execution_error`, `tool_execution_blocked` |
|
|
46
75
|
| Guardrails | `guardrail_decision` |
|
|
47
|
-
| Queue/subscribers | `queue_updated`, `event_subscriber_overflow` |
|
|
76
|
+
| Queue/subscribers | `queue_updated`, `event_subscriber_overflow`, `steer_rejected` |
|
|
48
77
|
| Compaction | `compaction_started`, `compaction_finished` |
|
|
49
78
|
| Retry | `retry_scheduled` |
|
|
50
79
|
| Artifacts | `artifact_validation_started`, `artifact_validation_finished`, `artifact_revision_started`, `artifact_finished`, `artifact_failed` |
|
|
@@ -92,6 +121,7 @@ Queue / subscriber / compaction / retry / provider events:
|
|
|
92
121
|
| Variant | Fields |
|
|
93
122
|
| --- | --- |
|
|
94
123
|
| `queue_updated` | `sessionId`, `runId`, `size: number` |
|
|
124
|
+
| `steer_rejected` | `sessionId`, `runId`, `message: Message` (redacted), `record: GuardrailRecord` — a steered message dropped by a terminal input guardrail; the run continues without it |
|
|
95
125
|
| `event_subscriber_overflow` | `sessionId`, `droppedEvents: number`, `maxQueuedEvents: number`, `overflow: "close" \| "drop_oldest" \| "drop_newest"` |
|
|
96
126
|
| `compaction_started` | `sessionId`, `runId?` |
|
|
97
127
|
| `compaction_finished` | `sessionId`, `runId?`, `summary: string` |
|
|
@@ -191,7 +221,7 @@ for await (const event of session.stream("draft", { loop: { strategy: "generate-
|
|
|
191
221
|
|
|
192
222
|
- All events flow through `redactAgentEvent(event, activeRedactor)` before subscribers observe them. Configure `AgentConfig.redactor` / `RunOptions.redactor` via `createSecretRedactor([...knownSecretStrings])` so secret values are redacted in `message` content, `errors[].message`, `metadata`, and artifact `result`/`failure` payloads.
|
|
193
223
|
- The artifact variants are emitted only by `generateValidateReviseLoop`. `singleShotLoop` (the default when no `AgentConfig.loop` / `RunOptions.loop` is set) emits zero artifact events. See [Agent loops](agent-loops.md).
|
|
194
|
-
- Subscribers are in-process; the broadcaster is in-memory and live-only. Multiple `subscribe()` calls receive the same stream.
|
|
224
|
+
- Subscribers are in-process; the broadcaster is in-memory and live-only. Multiple `subscribe()` calls receive the same stream. `resumeAgentRunStream()` and `AgentRunLifecycle.resumeStream()` subscribe before resumed execution and yield only their selected durable `runId`; approval emits the normal `agent_started` then `agent_resumed` envelope, denial emits only `agent_denied`.
|
|
195
225
|
- `session.subscribe(options)` accepts `maxQueuedEvents` (default `1024`, minimum `1`) and `overflow` (default `"close"`). The `close` policy clears queued payload events, queues one `event_subscriber_overflow` notice for that subscriber, then closes it. `drop_oldest` keeps the newest queued events; `drop_newest` ignores new events while full.
|
|
196
226
|
- The union is additive: new variants are appended without renumbering; subscribers should handle unknown `event.type` gracefully.
|
|
197
227
|
|
|
@@ -212,3 +242,4 @@ for await (const event of session.stream("draft", { loop: { strategy: "generate-
|
|
|
212
242
|
- [Observability](observability.md): `ProviderTurnMetadata`; optional adapter builds one parented GenAI span tree from metadata-only lifecycle events and ignores message/progress deltas.
|
|
213
243
|
- [Tools](tools.md): `tool_execution_*` variants.
|
|
214
244
|
- [Compaction and retry policies](compaction-and-retry.md): `compaction_*` and `retry_scheduled` variants.
|
|
245
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional redacted mapping of this stream; durable replay is ledger-backed and at-least-once, never a live-subscriber substitute. [ACP coding-host interop](acp.md) additionally maps `CodingLifecycleEvent`s from `@arnilo/prism-coding-agent` (`file_changed`, `worktree_changed`, `permission_denied`, `configuration_changed`; process events reuse `CodingProcessEvent`) into ACP session updates — locations/diff blocks only through projection allow-lists, terminal chunks under `process.outputChunkBytes`.
|
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
# Agent identity
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Authenticated `Principal` / `AgentIdentity` contracts let hosts attach verified tenant, sponsor/owner, delegated actor, scopes, credential references, issued/expiry, and revocation metadata to runs and tools. Core helpers assert activity, narrow scopes for delegation, project onto `OwnershipScope`, refuse silent widening, and emit redacted telemetry attributes. Prism does not store identities or verify tokens itself — hosts supply an `IdentityVerifier`.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
Use these APIs when embedding Prism in multi-tenant or enterprise hosts that already authenticate callers (for example Microsoft Entra Agent ID governance). Use them before tools, providers, MCP, A2A, workflows, or persistence that must carry attributable identity.
|
|
10
|
+
|
|
11
|
+
Do not treat optional `ownership` strings as identity provenance. Do not accept caller-asserted identity headers without a host verifier. Do not put JWTs or secret credential material on identity records.
|
|
12
|
+
|
|
13
|
+
## Inputs / request
|
|
14
|
+
|
|
15
|
+
| Field | Meaning |
|
|
16
|
+
| --- | --- |
|
|
17
|
+
| `Principal` | Actor id/kind (user, service, agent) plus optional display name |
|
|
18
|
+
| `AgentIdentity` | Verified context: required `tenantId`, optional account/user, principal, sponsor/owner, scopes, credential refs, issued/expiry/revocation, `verified: true` |
|
|
19
|
+
| `IdentityVerifier.verify(input)` | Host-owned authentication → `AgentIdentity` |
|
|
20
|
+
| `RunOptions.identity` / `AgentConfig.identity` | Optional verified identity for a run or agent default |
|
|
21
|
+
| Server/MCP/A2A `authorization.identity` | Optional verified identity on authorize results |
|
|
22
|
+
|
|
23
|
+
Frozen caps (defaults / hard): scopes `64 / 256`, scope bytes `128 / 512`, metadata `4 KiB / 16 KiB`, credential ref / principal id `256 B / 2 KiB`.
|
|
24
|
+
|
|
25
|
+
## Outputs / response / events
|
|
26
|
+
|
|
27
|
+
- `assertIdentityActive` — fail closed on unverified/expired/revoked/wrong-tenant/over-limit shapes (sync, no network).
|
|
28
|
+
- `narrowIdentity` — child scopes ⊆ parent; tenant and ownership ids immutable; expiry cannot extend.
|
|
29
|
+
- `ownershipFromIdentity` — projects tenant/account/user onto existing ownership seams.
|
|
30
|
+
- `assertIdentityMatchesOwnership` / `assertIdentityPropagation` — refuse widen across ownership or boundary hop.
|
|
31
|
+
- `identityTelemetryAttributes` — redacted refs for metadata/OTel (`prism.identity.*`); never includes credential secrets or raw tokens.
|
|
32
|
+
- Tool `ToolExecutionContext.identity` — set when a run carries verified identity.
|
|
33
|
+
|
|
34
|
+
## Request/response example
|
|
35
|
+
|
|
36
|
+
```json
|
|
37
|
+
{
|
|
38
|
+
"tenantId": "tenant-1",
|
|
39
|
+
"userId": "user-1",
|
|
40
|
+
"principal": { "kind": "agent", "id": "agent-42" },
|
|
41
|
+
"sponsor": { "kind": "user", "id": "sponsor-7" },
|
|
42
|
+
"scopes": ["mail.read", "mail.draft"],
|
|
43
|
+
"credentialRefs": ["m365:tenant-1:user-1"],
|
|
44
|
+
"issuedAt": "2026-07-23T00:00:00.000Z",
|
|
45
|
+
"expiresAt": "2026-07-23T01:00:00.000Z",
|
|
46
|
+
"verified": true
|
|
47
|
+
}
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
## Implementation example
|
|
51
|
+
|
|
52
|
+
```ts
|
|
53
|
+
import {
|
|
54
|
+
assertIdentityActive,
|
|
55
|
+
createAgent,
|
|
56
|
+
identityTelemetryAttributes,
|
|
57
|
+
narrowIdentity,
|
|
58
|
+
ownershipFromIdentity,
|
|
59
|
+
type AgentIdentity,
|
|
60
|
+
type IdentityVerifier,
|
|
61
|
+
} from "@arnilo/prism";
|
|
62
|
+
|
|
63
|
+
const verifier: IdentityVerifier = {
|
|
64
|
+
async verify(request) {
|
|
65
|
+
// Host validates JWT/session, then returns AgentIdentity with verified: true
|
|
66
|
+
return hostVerifiedIdentityFrom(request);
|
|
67
|
+
},
|
|
68
|
+
};
|
|
69
|
+
|
|
70
|
+
const identity = await verifier.verify(incomingRequest);
|
|
71
|
+
assertIdentityActive(identity);
|
|
72
|
+
const child = narrowIdentity(identity, { scopes: ["mail.read"] });
|
|
73
|
+
|
|
74
|
+
const agent = createAgent({
|
|
75
|
+
model,
|
|
76
|
+
provider,
|
|
77
|
+
ownership: ownershipFromIdentity(identity),
|
|
78
|
+
identity,
|
|
79
|
+
});
|
|
80
|
+
|
|
81
|
+
await agent.createSession().run("Summarize inbox", {
|
|
82
|
+
identity: child,
|
|
83
|
+
ownership: ownershipFromIdentity(child),
|
|
84
|
+
metadata: identityTelemetryAttributes(child),
|
|
85
|
+
});
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
Server / MCP / A2A authorize callbacks may include the same `identity` beside `ownership`. Handlers assert activity and ownership match before admitting work.
|
|
89
|
+
|
|
90
|
+
## OIDC/JWKS verifier adapter (`@arnilo/prism-credentials-node/oidc`)
|
|
91
|
+
|
|
92
|
+
Optional `createOidcIdentityVerifier` turns a pinned issuer/audience and pinned JWKS URL into a core `IdentityVerifier` — one bounded reference adapter for hosts that already authenticate callers with OIDC JWTs (Entra, Keycloak, Auth0, …). Native `fetch` + WebCrypto only; no JOSE dependency.
|
|
93
|
+
|
|
94
|
+
| Option | Meaning |
|
|
95
|
+
| --- | --- |
|
|
96
|
+
| `issuer` / `audience` | Exact `iss` and accepted `aud` value(s); anything else fails closed |
|
|
97
|
+
| `jwksUrl` | Host-pinned JWKS URL; SSRF-checked (`assertSsrfAllowedUrl`), never discovered at runtime, never followed through redirects |
|
|
98
|
+
| `mapClaims` | Bounded claims → `tenantId`, `principal`, `scopes`, optional account/user/sponsor/owner/refs/metadata |
|
|
99
|
+
| `algorithms` | `RS256`/`ES256` default; hosts may only narrow |
|
|
100
|
+
| `clockSkewMs` | Bounded `exp`/`nbf` slack (default 30 s) |
|
|
101
|
+
| `isRevoked` | Optional revocation callback; `true` or a thrown error fails closed |
|
|
102
|
+
| `limits` | Bounded JWKS/claims knobs; `identity` reuses core identity caps |
|
|
103
|
+
|
|
104
|
+
```ts
|
|
105
|
+
import { createOidcIdentityVerifier } from "@arnilo/prism-credentials-node/oidc";
|
|
106
|
+
|
|
107
|
+
const verifier = createOidcIdentityVerifier({
|
|
108
|
+
issuer: "https://id.example.com/tenant",
|
|
109
|
+
audience: "prism-api",
|
|
110
|
+
jwksUrl: "https://id.example.com/tenant/.well-known/jwks.json",
|
|
111
|
+
mapClaims: (claims) => ({
|
|
112
|
+
tenantId: String(claims.tid),
|
|
113
|
+
principal: { kind: "user", id: String(claims.sub) },
|
|
114
|
+
scopes: Array.isArray(claims.scp) ? claims.scp.map(String) : [],
|
|
115
|
+
}),
|
|
116
|
+
});
|
|
117
|
+
|
|
118
|
+
const identity = await verifier.verify({ token }); // -> AgentIdentity (verified: true)
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Fail-closed reasons (`IdentityError.reason`): `ERR_PRISM_OIDC_ISSUER_MISMATCH`, `AUDIENCE_MISMATCH`, `ALGORITHM`, `SIGNATURE`, `EXPIRED`, `NOT_YET_VALID`, `JWKS_FETCH`, `JWKS_KEY_MISSING`, `JWKS_PARSE`, `CLAIMS_BOUNDS`, `REVOKED`, `TENANT_MAPPING`. JWKS is cached (bounded entries/TTL, single-flight refetch, one bounded refetch on unknown `kid`); a parse/bounds failure on refresh fails closed while a transport failure keeps serving the last valid keys. SSRF denials surface the core `MediaContentError` (`ssrf_denied`). Tokens/claims never appear in errors, logs, or telemetry.
|
|
122
|
+
|
|
123
|
+
## Extension and configuration notes
|
|
124
|
+
|
|
125
|
+
Identity is optional. Hosts that only set `ownership` keep prior behavior. When identity is present, run start and tool dispatch assert it before side effects. Workflows forward `RunWorkflowOptions.identity` into agent nodes. Credential values stay behind `CredentialResolver` keys listed in `credentialRefs`.
|
|
126
|
+
|
|
127
|
+
## Security and performance notes
|
|
128
|
+
|
|
129
|
+
- Caller-asserted identity without `IdentityVerifier` is unsupported at trust boundaries.
|
|
130
|
+
- Delegation only narrows scopes; tenant/account/user cannot widen on propagation.
|
|
131
|
+
- Credential refs never expand to secrets in events, ledgers, or telemetry attributes.
|
|
132
|
+
- Checks are O(fields) and network-free in core; remote auth stays in the host verifier.
|
|
133
|
+
- Raising hard caps requires updating `docs/review-coverage-2026-07-23-phase-8.md`, tests, and docs.
|
|
134
|
+
|
|
135
|
+
## Related APIs
|
|
136
|
+
|
|
137
|
+
- [Policy and audit](policy-and-audit.md)
|
|
138
|
+
- [Host security guide](host-security.md)
|
|
139
|
+
- [Public contracts](public-contracts.md)
|
|
140
|
+
- [Server](server.md)
|
|
141
|
+
- [Supervisors](supervisors.md) / [A2A](a2a.md)
|
|
142
|
+
- [MCP tools](mcp-tools.md)
|
|
143
|
+
- [Observability](observability.md)
|
|
144
|
+
- [Runs and usage ledger](runs-and-usage.md)
|
package/docs/agent-loops.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
Agent loops make the agent's per-run turn-control flow a replaceable strategy without forking the runtime. The runtime owns provider calls, retry, abort, store appends, redaction, and event emission; a loop only orchestrates those shared primitives through a `LoopContext`. The default `singleShotLoop` is the former inline turn loop extracted verbatim — assemble → generate → append assistant message → optional tool dispatch → next turn. `generateValidateReviseLoop` is the first alternative loop: generate → parse → validate → revise up to a budget.
|
|
5
|
+
Agent loops make the agent's per-run turn-control flow a replaceable strategy without forking the runtime. The runtime owns provider calls, retry, abort, store appends, redaction, and event emission; a loop only orchestrates those shared primitives through a `LoopContext`. The default `singleShotLoop` is the former inline turn loop extracted verbatim — optional drain pending steers → assemble → generate → append assistant message → optional tool dispatch → next turn (continue while steers remain even if the provider returned no tool calls). `generateValidateReviseLoop` is the first alternative loop: generate → parse → validate → revise up to a budget.
|
|
6
6
|
|
|
7
7
|
Loops are opt-in. When no `loop` is configured, the runtime runs `singleShotLoop` and behavior is bit-for-bit with the pre-loop runtime.
|
|
8
8
|
|
|
@@ -58,6 +58,7 @@ await session.run(input, {
|
|
|
58
58
|
repairer: hostRepairer, // optional; default stringifies validation.errors[].message
|
|
59
59
|
maxRevisions: 3, // optional; default 3
|
|
60
60
|
toolCalls: "bounded", // optional; default "disabled"; uses limits.maxToolRounds
|
|
61
|
+
structuredOutputTiming: "final-turn-only", // optional; default "every-turn"
|
|
61
62
|
},
|
|
62
63
|
});
|
|
63
64
|
|
|
@@ -82,6 +83,10 @@ type AgentLoopOptions =
|
|
|
82
83
|
readonly maxRevisions?: number;
|
|
83
84
|
/** Default "disabled". "bounded" dispatches sequentially up to limits.maxToolRounds. */
|
|
84
85
|
readonly toolCalls?: "disabled" | "bounded";
|
|
86
|
+
readonly structuredOutput?: StructuredOutputOptions;
|
|
87
|
+
readonly structuredOutputMode?: "native" | "artifact-loop";
|
|
88
|
+
/** Default "every-turn". "final-turn-only" omits schema while tools may run. */
|
|
89
|
+
readonly structuredOutputTiming?: "every-turn" | "final-turn-only";
|
|
85
90
|
};
|
|
86
91
|
```
|
|
87
92
|
|
|
@@ -96,6 +101,8 @@ Host callback contracts (all generic over host `T`):
|
|
|
96
101
|
| `ArtifactContext` | `{ sessionId, runId, turn, signal, metadata }` — passed to every callback. |
|
|
97
102
|
| `ArtifactParseResult<T>` | `{ ok: boolean; value?: T; error?: string }`. |
|
|
98
103
|
|
|
104
|
+
Optional steer hooks on `LoopContext` (0.0.11): `hasPendingSteers?()` / `applyPendingSteers?()`. Hosts/custom loops that omit them keep pre-steer behavior; built-in loops drain at turn start.
|
|
105
|
+
|
|
99
106
|
`LoopContext` (what the runtime builds for the loop each run):
|
|
100
107
|
|
|
101
108
|
| Field | Purpose |
|
|
@@ -111,7 +118,15 @@ Host callback contracts (all generic over host `T`):
|
|
|
111
118
|
|
|
112
119
|
## Durable runs
|
|
113
120
|
|
|
114
|
-
`RunOptions.runState` supports
|
|
121
|
+
`RunOptions.runState` supports the built-in loop options (`single-shot` and `generate-validate-revise`) and custom strategies that opt into durable state. `single-shot` is durable via the runtime's pending-call mechanism and carries no loop-local state. `generate-validate-revise` snapshots `{ attempts, artifactPhase, savedSchema, pendingHistory }` at `revision: "1"`. A custom `AgentLoopStrategy` must declare both snapshot hooks or durable configuration rejects it with `AgentLoopStateError` (`ERR_PRISM_LOOP_NOT_DURABLE`) before any provider call:
|
|
122
|
+
|
|
123
|
+
| Member | Purpose |
|
|
124
|
+
| --- | --- |
|
|
125
|
+
| `revision?: string` | Host-authored loop revision. Joins the durable-run fingerprint, so a loop change without a `definitionRevision` bump fails closed on resume. |
|
|
126
|
+
| `snapshot?(): JsonValue` | Capture loop-local resumable state at suspension. Must be JSON-compatible; core redacts it and bounds it inside the durable run-state envelope (`maxStateBytes`, depth 32). A non-JSON value fails the run with `ERR_PRISM_LOOP_SNAPSHOT`. |
|
|
127
|
+
| `restore?(snapshot): void` | Rehydrate from the captured snapshot; must throw on drift. Called once before `run(ctx)` on resume. Also available as `ctx.restoredLoopState`. |
|
|
128
|
+
|
|
129
|
+
The snapshot is stored as `loopState: { name, revision, snapshot }` on the durable run state and cleared when the run reaches a terminal status. On resume, a name/revision mismatch between the stored `loopState` and the resolved strategy fails closed (`ERR_PRISM_LOOP_REVISION`), and the fingerprint check independently rejects any loop drift. Suspension occurs only before an input provider call or immediately before a tool side effect; completed provider turns remain in `SessionStore` history and are not repeated after `resumeAgentRun()`.
|
|
115
130
|
|
|
116
131
|
## Outputs / response / events
|
|
117
132
|
|
|
@@ -19,6 +19,7 @@ The agent/session runtime adds the minimal shared SDK surface for running provid
|
|
|
19
19
|
- `session.fork(options?)`
|
|
20
20
|
- `session.clone(options?)`
|
|
21
21
|
- `resumeAgentRun(agent, ref, decision, options)`
|
|
22
|
+
- `resumeAgentRunStream(agent, ref, decision, options)` → owned durable-resume `AsyncIterable<AgentEvent>`
|
|
22
23
|
- `createAgentRunLifecycle({ checkpoints, resolveAgent })` for host-selected remote status/resume adapters
|
|
23
24
|
|
|
24
25
|
The runtime streams provider text/tool-call content into `AgentEvent` values. Complete `tool_call` events are dispatched through the active host `ToolRegistry`, then returned as tool-result messages on the next provider turn. When a store is supplied, user, assistant, tool-result, and model-change entries are appended under the current branch leaf. Abort propagation and run exclusivity use native `AbortController`.
|
|
@@ -48,7 +49,9 @@ string | Message | readonly Message[]
|
|
|
48
49
|
|
|
49
50
|
`AgentConfig.limits` sets run ceilings; `RunOptions.limits` may only narrow configured agent values. Limits cover turns, provider attempts, tool rounds/calls, wall time, request/response bytes, tokens, and optional single-currency cost. A breach emits one `run_limit_exceeded` event and throws `AgentRunError` with `result.limit`; see [Runs and usage ledger](runs-and-usage.md#run-limits).
|
|
50
51
|
|
|
51
|
-
`RunOptions.model` can override the request model for a run. Model overrides append a `model_change` entry. `AgentConfig.inputLayout` selects the default input assembly layout (`"
|
|
52
|
+
`RunOptions.model` can override the request model for a run. Model overrides append a `model_change` entry. `AgentConfig.inputLayout` selects the default input assembly layout (`"cache_aware"` by default, or opt-in `"legacy"`); `RunOptions.inputLayout` wins for one run. `AgentConfig.providerOptions`/`RunOptions.providerOptions` supply generic provider request options; `timeoutMs`, `maxRetries`, and `maxRetryDelayMs` are deprecated inert provider-level hints in first-party providers. Use `RunOptions.signal`/host abort controllers for timeouts and `AgentConfig.retry`/`RunOptions.retry` for retry. `AgentConfig.providerRequestPolicies`/`RunOptions.providerRequestPolicies` run before `AIProvider.generate()` and before `provider_request` middleware. `AgentConfig.systemPrompt` and `RunOptions.systemPrompt` add explicit layered system prompt contributions; `RunOptions.systemPrompt: false` disables configured prompt layers for that run while keeping `AgentConfig.instructions` as the base path. `RunOptions.compaction` can enable auto-compaction for that run or use `false` to disable configured auto-compaction. `RunOptions.retry` can enable provider-turn retry for that run or use `false` to disable configured retry. `RunOptions.metadata` is merged with agent/session metadata for assembly, provider requests, and tool contexts. Deprecated `RunOptions.maxToolRounds` narrows `limits.maxToolRounds`. `RunOptions.signal` is bridged into the per-run abort signal passed to assembly, providers, tools, auto-compaction, and retry backoff.
|
|
53
|
+
|
|
54
|
+
`RunOptions.activeSkills` selects named skills from a configured `SkillRegistry`; `RunOptions.skills` replaces a plain `Skill[]` config for one run. When `AgentConfig.skills` is a registry and neither is set, **no skills activate** unless `activateAllSkills: true` (run or agent). `skillsDisclosure` (`"progressive"` default, `"eager"` opt-in; run wins) controls catalog vs full instruction bodies; the session-owned `LoadedSkillSet` is populated by `load_skill` when the host registers `createLoadSkillTool`. `toolResultFold` (off unless the host supplies `summarize`) optionally folds aged large tool results in provider input only. See [Context and skills](context-and-skills.md).
|
|
52
55
|
|
|
53
56
|
## Outputs / response / events
|
|
54
57
|
|
|
@@ -56,6 +59,8 @@ string | Message | readonly Message[]
|
|
|
56
59
|
|
|
57
60
|
`session.stream(input, options?)` subscribes first, starts exactly one run, yields only that run's events, and terminates when the run succeeds, fails, or aborts. Early consumer return aborts the owned run and releases the session. `SubscribeOptions.maxQueuedEvents` / `overflow` may be passed alongside `RunOptions`.
|
|
58
61
|
|
|
62
|
+
`resumeAgentRunStream(agent, ref, resume, options?)` does the same for one existing suspended durable run. It validates checkpoint ownership, revision/fingerprint, and `expectedVersion`, then subscribes before emitting `agent_started` / `agent_resumed` and resumed message/tool/terminal events. `AgentRunResumeStreamOptions` combines existing resume options with `signal`, `maxQueuedEvents`, and `overflow`; early return aborts only resumed execution. It does not replay a claimed/dispatched tool, poll a ledger, or retain a worker. `createAgentRunLifecycle().resumeStream(ref, resume, request?)` adds the same behavior after host agent-capability resolution.
|
|
63
|
+
|
|
59
64
|
`session.subscribe(options?)` remains available for hosts that want a long-lived subscriber across runs. Subscribe before `run()` to observe that run's events. The consumer loop and `session.run()` must run concurrently (e.g. start the `for await` consumer, then `await Promise.all([consumer, session.run("Hi")])`): events are only emitted during a live run, so awaiting the subscribe loop before calling `run()` deadlocks. Prefer `session.stream()` when you only need one run's events. `SubscribeOptions.maxQueuedEvents` defaults to `1024` (minimum `1`) and caps events queued while the consumer is not awaiting `next()`. `SubscribeOptions.overflow` defaults to `"close"`; it clears queued payload events, delivers one `event_subscriber_overflow` notice to that subscriber, then closes it. `"drop_oldest"` keeps newest events; `"drop_newest"` ignores new events while full.
|
|
60
65
|
|
|
61
66
|
For a text-only provider turn, the runtime emits:
|
|
@@ -78,7 +83,11 @@ Provider `thinking`/`reasoning` content emitted during a turn is preserved as `t
|
|
|
78
83
|
|
|
79
84
|
Missing providers fail closed: `run()` emits `error` and rejects before calling any provider. Provider `error` events emit session `error` and reject unless configured retry handles a transient provider-turn failure before output. Unknown tools fail closed through the tool harness and do not execute. Tool exceptions emit `tool_execution_error`, return an error `ToolResult`, and may still continue to the next provider turn.
|
|
80
85
|
|
|
81
|
-
Only one `run()` may be active per session. Concurrent `run()` calls emit `error` and reject immediately; Prism does not queue
|
|
86
|
+
Only one `run()` may be active per session. Concurrent `run()` / `prompt` / `followUp` calls emit `error` and reject immediately; Prism does not queue second prompts. Manual `compact()` also rejects while a run is active.
|
|
87
|
+
|
|
88
|
+
### Mid-run steer (0.0.11)
|
|
89
|
+
|
|
90
|
+
`session.steer(input, options?)` enqueues user text into the **same** active run (fail closed when no run). Default: inject at the next turn boundary (after tool rounds / before next provider assemble). `options.softInterrupt: true` aborts only the current provider stream, then continues the same `runId` with steered text. Pending queue caps: **8** messages / **64 KiB** UTF-8 total (`DEFAULT_MAX_PENDING_STEERS` / `DEFAULT_MAX_PENDING_STEER_BYTES`); overflow throws. Steered messages pass input guardrails + normal session append/redaction. A `block`/`tripwire` on a steered message drops just that message: Prism emits `guardrail_decision` plus a `steer_rejected` event (redacted message + `GuardrailRecord`) and the run continues; the message never enters history or the session store. Run-start input blocking still fails the run. `interrupt` on a steered message fails closed (durable suspension is only for run-start input). Loops drain via optional `LoopContext.hasPendingSteers` / `applyPendingSteers`.
|
|
82
91
|
|
|
83
92
|
`session.abort(reason)` aborts the active run. The abort signal is passed to input assembly, provider requests, and tool execution; if a tool/provider path aborts after a tool call, Prism does not start another provider turn.
|
|
84
93
|
|
|
@@ -167,7 +176,14 @@ await agent.createSession().run("Hi", { model: overrideModel });
|
|
|
167
176
|
|
|
168
177
|
## Durable interruption
|
|
169
178
|
|
|
170
|
-
Set `runState` with a host-owned `CheckpointStore`, stable `definitionRevision`, and `interruptBeforeTool: true` to suspend at a persisted pre-side-effect boundary. A suspended result has `status: "suspended"`, a redacted `interruption`, and `runState.version`; it releases session resources before returning.
|
|
179
|
+
Set `runState` with a host-owned `CheckpointStore`, stable `definitionRevision`, and `interruptBeforeTool: true` to suspend at a persisted pre-side-effect boundary. A suspended result has `status: "suspended"`, a redacted `interruption`, and `runState.version`; it releases session resources before returning. When a provider turn requests several tools, the round is collected into **one** suspension whose `interruption.pendingDecisions` holds one redacted `PendingDecision` per gated call (`approvalId`, kind, scope with tool name/effect kind/identity/arguments hash — never raw arguments); ungated calls still dispatch.
|
|
180
|
+
|
|
181
|
+
`resumeAgentRun` accepts exactly one of:
|
|
182
|
+
|
|
183
|
+
- `decision: "approve" | "deny"` — legacy single-approval path. `approve` allows every pending decision once; `deny` terminates the run as `denied`.
|
|
184
|
+
- `decisions: readonly RunDecision[]` — one atomic batch. Every entry validates against the recorded pending set (unknown/foreign `approvalId`, duplicates, stale `expectedVersion`, invalid outcomes fail the whole batch closed with `AgentDecisionError` and leave state and version untouched). Outcomes: `allow_once`, `allow_for_run`, `reject_once`, `reject_for_run`. `reject_*` continues the run with a blocked tool result carrying the bounded (2 KB) `reason`. `modifiedArguments` are revalidated (schema, then input guardrails; permission/trust re-run at dispatch) and produce a new arguments hash. `elicitation` payloads are validated against the pending decision's `elicitationSchema` (required keys plus the configured host validator) and resolve the suspended call without executing it. A batch deciding a strict subset persists the decided entries and re-suspends with the remainder pending at the bumped version.
|
|
185
|
+
|
|
186
|
+
`*_for_run` outcomes append a `StickyDecision` to the durable run state: later calls in the same run matching the scope exactly (all recorded fields) proceed or are blocked without a new suspension, policy still enforced at dispatch. Sticky decisions expire when the run reaches any terminal status. Caps: 32 pending decisions per run (hard 128), 64 sticky decisions (hard 256), 2 KB decision reasons, 16 KB elicitation payloads.
|
|
171
187
|
|
|
172
188
|
```ts
|
|
173
189
|
const result = await session.run("Publish draft", {
|
|
@@ -180,7 +196,7 @@ if (result.status === "suspended") {
|
|
|
180
196
|
}
|
|
181
197
|
```
|
|
182
198
|
|
|
183
|
-
Resume requires exact checkpoint ownership, version, agent fingerprint, and revision. Prism CAS-claims approval before work, rechecks normal guardrail/permission/validation/limit paths, and marks a pending tool dispatched before its side effect. `createAgentRunLifecycle()` wraps the same core path for server/MCP hosts: adapters pass only authorized ownership, status returns only `{ state, version }`, and `resolveAgent()` supplies current agent/revision. Remote restart requires both checkpoint and session stores to be durable. A crash after that mark is ambiguous and is never replayed automatically; use host tool idempotency keyed by `runId`/`toolCallId` or resolve it manually. Checkpoints contain bounded redacted state plus session/leaf references, never provider objects, callbacks, signals, credentials, or raw secrets.
|
|
199
|
+
Resume requires exact checkpoint ownership, version, agent fingerprint, and revision. The fingerprint hashes the agent id/name, `definitionRevision`, model, instructions, system-prompt contributions, skills (name/instructions/tool names), tool definitions (name/parameters/exclusive), guardrail definitions (name/stage/revision), and loop strategy — changing any of them without bumping `definitionRevision` fails resume closed instead of silently continuing with different agent semantics. Prism CAS-claims approval before work, rechecks normal guardrail/permission/validation/limit paths, and marks a pending tool dispatched before its side effect. `createAgentRunLifecycle()` wraps the same core path for server/MCP hosts: adapters pass only authorized ownership, status returns only `{ state, version }`, and `resolveAgent()` supplies current agent/revision. `resumeStream()` uses that same claim path and bounded subscriber, so adapters do not poll or duplicate resume logic. Remote restart requires both checkpoint and session stores to be durable. A crash after that mark is ambiguous and is never replayed automatically; use host tool idempotency keyed by `runId`/`toolCallId` or resolve it manually. Checkpoints contain bounded redacted state plus session/leaf references, never provider objects, callbacks, signals, credentials, or raw secrets. State is bounded at save by `runState.maxStateBytes` (default 256 KB, at most the 1 MB hard cap); load bounds against the 1 MB hard cap only, so state saved with a raised limit stays resumable while oversized records are still rejected. Built-in loop options are durable; custom `AgentLoopStrategy` instances are durable when they declare `snapshot`/`restore` hooks (see [Agent loops § Durable runs](agent-loops.md#durable-runs)) and reject before provider work otherwise.
|
|
184
200
|
|
|
185
201
|
## Secure composition
|
|
186
202
|
|
|
@@ -205,6 +221,7 @@ Per-run options may narrow `limits` and append `guardrails`; they cannot replace
|
|
|
205
221
|
- [CLI/RPC](cli-rpc.md): terminal and JSONL adapters over this runtime.
|
|
206
222
|
- [Workflows](workflows.md): optional DAG orchestration that calls `AgentSession.run()` for agent nodes.
|
|
207
223
|
- [A2A interoperability](a2a.md): direct text exposure calls `AgentSession.run()`; durable/rich/reconnect behavior uses host `A2ATaskLifecycle` over existing checkpoints/persistence, never an in-memory runtime cache.
|
|
224
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional adapters use `session.stream()` and `resumeAgentRunStream()` / `AgentRunLifecycle.resumeStream()`; protocol/UI state remains outside core.
|
|
208
225
|
|
|
209
226
|
`AgentConfig.loop` and `RunOptions.loop` select a replaceable per-run control loop (`singleShotLoop` default, or `generate-validate-revise` with host callbacks); see [Agent loops](agent-loops.md). `RunOptions.loop` wins over `AgentConfig.loop`. Built-in loops emit the same normal turn/message envelope around provider turns, and both add the first run input to live history once after the first provider turn so later turns see the same transcript shape.
|
|
210
227
|
|