@arnilo/prism 0.0.4 → 0.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (70) hide show
  1. package/CHANGELOG.md +18 -0
  2. package/README.md +34 -10
  3. package/dist/agents.js +146 -19
  4. package/dist/cli-init.d.ts +41 -0
  5. package/dist/cli-init.js +390 -0
  6. package/dist/cli-runner.d.ts +7 -1
  7. package/dist/cli-runner.js +13 -1
  8. package/dist/content.d.ts +19 -0
  9. package/dist/content.js +197 -69
  10. package/dist/contracts.d.ts +94 -9
  11. package/dist/contracts.js +8 -0
  12. package/dist/feedback.d.ts +48 -0
  13. package/dist/feedback.js +230 -0
  14. package/dist/index.d.ts +6 -4
  15. package/dist/index.js +4 -3
  16. package/dist/providers/media.d.ts +3 -1
  17. package/dist/providers/media.js +11 -1
  18. package/dist/testing/feedback.d.ts +6 -0
  19. package/dist/testing/feedback.js +37 -0
  20. package/dist/testing/persistence-schema.d.ts +3 -3
  21. package/dist/testing/persistence-schema.js +32 -2
  22. package/dist/testing/run-ledger-conformance.js +7 -1
  23. package/docs/a2a.md +73 -0
  24. package/docs/agent-events.md +4 -6
  25. package/docs/agent-loops.md +1 -1
  26. package/docs/agent-session-runtime.md +14 -16
  27. package/docs/cli-rpc.md +35 -7
  28. package/docs/coding-agent-tools.md +2 -2
  29. package/docs/coding-security.md +7 -3
  30. package/docs/compaction-observational-memory.md +2 -0
  31. package/docs/context-and-skills.md +1 -0
  32. package/docs/credentials-and-redaction.md +2 -2
  33. package/docs/database-persistence.md +9 -6
  34. package/docs/evaluations.md +122 -0
  35. package/docs/extensions.md +2 -2
  36. package/docs/host-security.md +20 -3
  37. package/docs/index.md +29 -17
  38. package/docs/mcp-tools.md +49 -4
  39. package/docs/migration.md +33 -3
  40. package/docs/multimodal-content.md +14 -6
  41. package/docs/observability.md +14 -6
  42. package/docs/performance.md +209 -0
  43. package/docs/postgres-persistence.md +6 -4
  44. package/docs/provider-conformance.md +1 -0
  45. package/docs/provider-packages.md +2 -0
  46. package/docs/providers/ai-sdk.md +113 -0
  47. package/docs/public-contracts.md +6 -5
  48. package/docs/rag.md +113 -0
  49. package/docs/release-and-install.md +100 -77
  50. package/docs/review-coverage-2026-07-15.md +193 -0
  51. package/docs/runs-and-usage.md +41 -4
  52. package/docs/server.md +139 -0
  53. package/docs/settings-auth-trust-security.md +5 -5
  54. package/docs/sqlite-persistence.md +4 -3
  55. package/docs/supervisors.md +71 -0
  56. package/docs/workflow-orchestration-primitives.md +19 -3
  57. package/docs/workflows.md +97 -23
  58. package/docs/working-and-semantic-memory.md +169 -0
  59. package/package.json +12 -2
  60. package/templates/init/README.md.tmpl +28 -0
  61. package/templates/init/env.example.tmpl +1 -0
  62. package/templates/init/gitignore.tmpl +11 -0
  63. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  64. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  65. package/templates/init/package.json.tmpl +22 -0
  66. package/templates/init/providers.json +76 -0
  67. package/templates/init/src/agent.ts.tmpl +10 -0
  68. package/templates/init/src/index.ts.tmpl +12 -0
  69. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  70. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -7,8 +7,9 @@ The agent/session runtime adds the minimal shared SDK surface for running provid
7
7
  - `createAgent(config)`
8
8
  - `createAgentSession(config)`
9
9
  - `agent.createSession(config)`
10
- - `session.run(input, options)`
11
- - `session.prompt(input, options)`
10
+ - `session.run(input, options)` → `AgentRunResult`
11
+ - `session.prompt(input, options)` → `AgentRunResult`
12
+ - `session.stream(input, options)` → owned-run `AsyncIterable<AgentEvent>`
12
13
  - `session.compact(options?)`
13
14
  - `session.subscribe(options?)`
14
15
  - `session.abort()`
@@ -46,7 +47,11 @@ string | Message | readonly Message[]
46
47
 
47
48
  ## Outputs / response / events
48
49
 
49
- `session.subscribe(options?)` returns a live `AsyncIterable<AgentEvent>`. Subscribe before `run()` to observe that run's events. The consumer loop and `session.run()` must run concurrently (e.g. start the `for await` consumer, then `await Promise.all([consumer, session.run("Hi")])`): events are only emitted during a live run, so awaiting the subscribe loop before calling `run()` deadlocks. `SubscribeOptions.maxQueuedEvents` defaults to `1024` (minimum `1`) and caps events queued while the consumer is not awaiting `next()`. `SubscribeOptions.overflow` defaults to `"close"`; it clears queued payload events, delivers one `event_subscriber_overflow` notice to that subscriber, then closes it. `"drop_oldest"` keeps newest events; `"drop_newest"` ignores new events while full.
50
+ `session.run()` / `session.prompt()` resolve to an `AgentRunResult` with `sessionId`, `runId`, `status`, `text`, `content`, optional `message`/`usage`/`leafId`, and terminal `error`/`abortReason` when applicable. Callers may ignore the return value. Failed and aborted runs still emit their terminal events, then reject with `AgentRunError` whose `.result` carries the same shape.
51
+
52
+ `session.stream(input, options?)` subscribes first, starts exactly one run, yields only that run's events, and terminates when the run succeeds, fails, or aborts. Early consumer return aborts the owned run and releases the session. `SubscribeOptions.maxQueuedEvents` / `overflow` may be passed alongside `RunOptions`.
53
+
54
+ `session.subscribe(options?)` remains available for hosts that want a long-lived subscriber across runs. Subscribe before `run()` to observe that run's events. The consumer loop and `session.run()` must run concurrently (e.g. start the `for await` consumer, then `await Promise.all([consumer, session.run("Hi")])`): events are only emitted during a live run, so awaiting the subscribe loop before calling `run()` deadlocks. Prefer `session.stream()` when you only need one run's events. `SubscribeOptions.maxQueuedEvents` defaults to `1024` (minimum `1`) and caps events queued while the consumer is not awaiting `next()`. `SubscribeOptions.overflow` defaults to `"close"`; it clears queued payload events, delivers one `event_subscriber_overflow` notice to that subscriber, then closes it. `"drop_oldest"` keeps newest events; `"drop_newest"` ignores new events while full.
50
55
 
51
56
  For a text-only provider turn, the runtime emits:
52
57
 
@@ -101,29 +106,22 @@ const agent = createAgent({
101
106
  });
102
107
 
103
108
  const session = agent.createSession({ id: "s1" });
104
- const reader = (async () => {
105
- for await (const event of session.subscribe()) console.log(event.type);
106
- })();
109
+ const result = await session.run("Hi", { maxToolRounds: 1, compaction: { thresholdEntries: 20, keepRecentEntries: 6 }, retry: { maxAttempts: 3, baseDelayMs: 50 } });
110
+ console.log(result.text, result.usage?.totalTokens);
111
+
112
+ for await (const event of session.stream("Follow up")) console.log(event.type);
107
113
 
108
- await session.run("Hi", { maxToolRounds: 1, compaction: { thresholdEntries: 20, keepRecentEntries: 6 }, retry: { maxAttempts: 3, baseDelayMs: 50 } });
109
114
  await session.compact({ keepRecentEntries: 4 });
110
115
  const branch = await session.entries();
111
116
  await session.checkout(branch.at(-1)?.id);
112
117
  const clone = await session.clone({ id: "s2" });
113
- await reader;
114
118
  ```
115
119
 
116
120
  ## Extension and configuration notes
117
121
 
118
122
  The runtime calls `assembleProviderInput()` on every turn and uses only runtime-consumed values supplied on `AgentConfig`: `instructions`, `systemPrompt`, `inputBuilder`, `promptBuilder`, `inputLayout`, `context`, selected `skills`, active `tools`, `middleware`, `resourceLoader`, metadata, `compaction`, `retry`, and `RunOptions.model`/`systemPrompt`/`inputLayout`/`compaction`/`retry`. Contributions remain inert until a host passes selected values into the agent config.
119
123
 
120
- `AgentConfig` fields that are host-owned metadata, not runtime work:
121
-
122
- | Field | Runtime behavior |
123
- | --- | --- |
124
- | `extensions` | Preserved on `agent.config` only. `createAgent()` / `session.run()` do not call `setup()`, load packages, or auto-register contributions. Load extensions with `createExtensionKernel()` before building config. |
125
- | `settings` | Preserved on `agent.config` only. The runtime does not call `settings.get()`; hosts or provider packages read settings before passing concrete runtime options. |
126
- | `credentials` | Preserved on `agent.config` only. The runtime does not call `credentials.resolve()`; provider adapters/request policies resolve credentials at the provider edge and pass exact secret values to redaction when needed. |
124
+ `AgentConfig` no longer accepts inert `extensions`, `settings`, or `credentials` fields. Load extensions with `createExtensionKernel()` before building config; read settings in the host before passing concrete runtime options; resolve credentials at the provider edge and pass exact secret values to redaction when needed.
127
125
 
128
126
  The runtime calls `middleware.run("compaction", { context, result })` after a compaction strategy returns and before appending the standard compaction entry. Middleware can adjust the result summary/data, but the runtime still owns store append ordering and branch parent ids.
129
127
 
@@ -131,7 +129,7 @@ Provider request policy application is one ordered in-memory pass per provider t
131
129
 
132
130
  The runtime calls `middleware.run("retry", { context, decision })` after the retry policy decision and before emitting `retry_scheduled`. Middleware can stop retrying or adjust the delay. Retry wraps only the current provider turn, reuses the same assembled request, and never retries after assistant output has been emitted.
133
131
 
134
- `createAgent()` is a thin wrapper over explicit config. It does not load `AgentConfig.extensions`, scan packages, resolve credentials, read settings, call `Extension.setup()`, or consult hidden registries. External `AgentDefinition` implementations can call it from their own `create()` method:
132
+ `createAgent()` is a thin wrapper over explicit config. It does not scan packages, resolve credentials, read settings, call `Extension.setup()`, or consult hidden registries. External `AgentDefinition` implementations can call it from their own `create()` method:
135
133
 
136
134
  ```ts
137
135
  import { createAgent, createContributionRegistries } from "@arnilo/prism";
package/docs/cli-rpc.md CHANGED
@@ -2,23 +2,41 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- The `prism` bin is a thin adapter over `AgentSession`:
5
+ The `prism` bin is a thin adapter over `AgentSession` plus a tiny project scaffold:
6
6
 
7
7
  - `prism -p "prompt"`: print assistant text deltas.
8
8
  - `prism --mode json -p "prompt"`: write one normalized event envelope per line.
9
9
  - `prism --mode rpc`: read LF-delimited JSON requests from stdin and write correlated JSON responses/events to stdout.
10
+ - `prism init <dir>`: create a minimal TypeScript project with one selected provider, `.env.example`, and one offline mock test.
10
11
 
11
- It does not add a TUI, app tools, provider globals, extension discovery, resource discovery, or credential storage.
12
+ It does not add a TUI, app tools, provider globals, extension discovery, resource discovery, or credential storage. `init` uses Node standard-library filesystem APIs and checked-in templates only — no interactive prompts or template-engine dependency.
12
13
 
13
14
  ## When to use it
14
15
 
15
- Use the CLI for terminal smoke tests, scriptable JSON event streams, and simple non-Node clients that can speak newline-delimited JSON.
16
+ Use the CLI for terminal smoke tests, scriptable JSON event streams, simple non-Node clients that can speak newline-delimited JSON, and bootstrapping a tiny host project with `prism init`.
16
17
 
17
18
  Use the SDK directly when an app needs custom providers, tools, resources, credentials, trust prompts, or UI behavior.
18
19
 
19
20
  ## Inputs / request
20
21
 
21
- CLI flags:
22
+ ### `prism init`
23
+
24
+ ```bash
25
+ prism init <dir> [--provider <name>] [--with-workflows] [--with-evals] [--force]
26
+ ```
27
+
28
+ | Flag / arg | Purpose |
29
+ | --- | --- |
30
+ | `<dir>` | Destination directory (created if missing). |
31
+ | `--provider <name>` | `mock` (default), `openai`, `openrouter`, `kimi`, `zai`, `opencode-go`, or `neuralwatt`. |
32
+ | `--with-workflows` | Add `@arnilo/prism-workflows` and `src/workflows-example.ts`. |
33
+ | `--with-evals` | Add `@arnilo/prism-evals` and `src/evals-example.ts`. |
34
+ | `--force` | Overwrite generated files when the destination already exists. |
35
+ | `-h`, `--help` | Print init usage. |
36
+
37
+ Default generation installs only `@arnilo/prism` (mock provider). Selecting a real provider adds exactly one `@arnilo/prism-provider-*` package. Storage, telemetry, memory, and server packages are never added unless a later phase introduces an explicit flag for them. Rerunning without `--force` refuses non-empty destinations and existing generated files. `.env.example` contains placeholders only; `.gitignore` excludes `.env` and local stores.
38
+
39
+ ### Run/RPC CLI flags
22
40
 
23
41
  | Flag | Purpose |
24
42
  | --- | --- |
@@ -127,6 +145,11 @@ Events streamed during a run keep the original prompt request id, even when an `
127
145
  prism --provider mock --model demo -p "Hi"
128
146
  prism --provider mock --mode json -p "Hi"
129
147
  printf '{"id":"1","command":"prompt","params":{"input":"Hi"}}\n' | prism --provider mock --mode rpc
148
+
149
+ prism init my-agent
150
+ prism init my-agent --provider openai
151
+ prism init my-agent --provider openrouter --with-workflows --with-evals
152
+ cd my-agent && npm install && npm test
130
153
  ```
131
154
 
132
155
  Programmatic hosts should use the public runtime directly:
@@ -147,7 +170,9 @@ CLI/RPC are adapters over `AgentSession`. They do not scan packages, import exte
147
170
 
148
171
  RPC `command` executes only explicitly registered `CommandDefinition` values. `setModel` stores a model override for later prompt/follow-up calls. `compact`, `switchSession`, `forkSession`, `cloneSession`, and `checkout` call the existing session APIs.
149
172
 
150
- Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.start`, `workflow.status`, `workflow.list`, `workflow.cancel`, and `workflow.resume` via `createWorkflowCommands({ workflows, checkpoints, runOptions? })`. Pass the returned `CommandDefinition[]` into `runRpcServer({ commands })` the same way as observational-memory commands. Cancel aborts in-process runs through the package active-run registry; orphaned durable checkpoints still marked `running` are fail-closed to `aborted`.
173
+ Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.start`, `workflow.enqueue`, `workflow.replay`, `workflow.status`, `workflow.list`, `workflow.cancel`, and `workflow.resume` via `createWorkflowCommands({ workflows, checkpoints, runOptions? })`. Supplying an ownership-scoped `schedules` service additionally registers `schedule.create`, `schedule.list`, `schedule.pause`, `schedule.resume`, `schedule.trigger`, and `schedule.delete`. Pass the returned `CommandDefinition[]` into `runRpcServer({ commands })` the same way as observational-memory commands. Cancel aborts in-process runs through the package active-run registry; orphaned durable checkpoints still marked `running` are fail-closed to `aborted`.
174
+
175
+ Suspended workflow resume parameters are `{ workflowId, runId, decision: "approve" | "deny", input?, expectedVersion, ownership? }`. Read `expectedVersion` from `workflow.status`/`workflow.list`; stale or duplicate decisions fail checkpoint CAS before node execution. Ordinary recovery resume for failed/aborted runs remains backward-compatible without decision fields.
151
176
 
152
177
  `forkSession` creates another handle for the same `sessionId` and selected `leafId`; it no longer overwrites the parent handle in the RPC map. Keep the returned `handleId` when a UI needs to switch among sibling branches. `switchSession` accepts `handleId` (preferred), `sessionId`, or `id`; with multiple branch handles, use `handleId` to avoid ambiguity. `checkout` requires `params.leafId`, calls `AgentSession.checkout(leafId)`, and keeps the active handle id unchanged while moving that handle to the existing leaf. `messages` returns entries for the active branch path.
153
178
 
@@ -157,7 +182,10 @@ Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.s
157
182
  - No hidden provider, credential, extension, resource, config, settings, or tool globals are created.
158
183
  - No full TUI or sandbox is provided or implied.
159
184
  - JSONL is processed line by line with Node stdlib; no parser dependency, worker, watcher, or queue is added.
160
- - Unknown or malformed CLI/RPC input fails closed.
185
+ - Unknown or malformed CLI/RPC input fails closed. Workflow resume validates decision and positive `expectedVersion`; ownership remains host-selected and checkpoint-enforced.
186
+ - `prism init` refuses non-empty destinations without `--force`, keeps writes inside the destination root, and never executes downloaded code beyond the user's later `npm install`.
187
+ - Generated `.env.example` values are placeholders only; `.gitignore` excludes `.env` and local store files.
188
+ - Default generated install stays small (~27 MB with TypeScript tooling in a clean consumer install versus Mastra's measured 439 MB scaffold); unselected storage/telemetry/eval/workflow packages are omitted.
161
189
  - Branch handles (`handleId`, `sessionId`, `leafId`) are identifiers only; do not encode credentials, tokens, provider objects, or secrets into them.
162
190
  - Do not put resolved credential values, tokens, headers, or secrets in prompts, CLI flags, config, events, or docs examples.
163
191
 
@@ -171,7 +199,7 @@ Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.s
171
199
  - [Resource loading](resource-loading.md): explicit resource loading primitives.
172
200
  - [Credentials and redaction](credentials-and-redaction.md): secret redaction helpers and credential boundaries.
173
201
  - [Observational memory compaction package](compaction-observational-memory.md): optional `om:status` and `om:view` command factories for explicitly wired hosts.
174
- - [Workflows](workflows.md): optional `createWorkflowCommands()` for start/status/list/cancel/resume over the same RPC `command` seam.
202
+ - [Workflows](workflows.md): optional `createWorkflowCommands()` for direct/background/replay/status/cancel/resume and selected schedule control over the same RPC `command` seam.
175
203
 
176
204
  The CLI records flags but does not auto-load project-local resources, extensions, tools, or config. The two system/project prompt files are the exception: in print/json modes the CLI auto-loads `<workspaceRoot>/AGENTS.md` (trust-gated) and an app-supplied `SYSTEM.md` layer as `AgentConfig.systemPrompt` layers composed with `--system` (base); `--no-agents-md` / `--no-system-md` skip them and `--agents-md-file` / `--system-md-file` override the paths. The CLI does not default `globalRoot` to the user's home directory — pass it from a host adapter or use `--agents-config <path>` for the app-config bundle layout. RPC mode does not auto-read these files (the host owns the session factory). Hosts must make explicit trust and permission decisions before wiring any other local loading.
177
205
 
@@ -222,13 +222,13 @@ const remoteWrite = createWriteTool("/repo", {
222
222
 
223
223
  - **Pluggable operation backends.** Every tool accepts an `operations` seam so a host can delegate to a remote system (e.g. SSH) while keeping the tool's matching/serialization behavior: `BashOperations` (`shell`), `ReadOperations` (`read`), `WriteOperations` (`write`), `EditOperations` (`edit`).
224
224
  - **Per-tool options.** `ShellToolOptions` (`shellPath`, `commandPrefix`, `maxLines`, `maxBytes`, `tempFilePrefix`, `operations`, `spawnHook`, `executionPolicy`); `ReadToolOptions` (`operations`, `maxImageBytes`, `transformImage`, `maxLines`, `maxBytes`, `executionPolicy`; `autoResizeImages` deprecated); `WriteToolOptions` (`operations`, `executionPolicy`); `EditToolOptions` (`operations`, `executionPolicy`).
225
- - **Aggregator options.** `ToolsOptions` (`{ shell?, read?, write?, edit? }`) threads each sub-object to the matching tool.
225
+ - **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override.
226
226
  - **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
227
227
  - No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
228
228
 
229
229
  ## Security and performance notes
230
230
 
231
- - **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
231
+ - **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
232
232
  - **Non-zero exit is not an error.** A failing command is a normal `shell` result (exit code in metadata); only timeout/abort/spawn failures are error results. Do not assume `error == undefined` means the command succeeded.
233
233
  - **Bounded output.** `shell`/`read` accumulate output into a rolling tail bounded by `maxLines`/`maxBytes`; oversized output spills to a temp file (`fullOutputPath`), so memory use is bounded regardless of command output size.
234
234
  - **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
@@ -34,9 +34,11 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
34
34
  | `approvalCacheScope` | `"none"` | Optional `run` or `session` decision cache scope. |
35
35
  | `approvalTimeoutMs` | `30000` | Bound approval wait; caller abort also cancels it. |
36
36
 
37
+ `run` caching keys decisions by the tool execution context's `runId`; `session` uses `sessionId`. Coding tools pass both identities to the policy. A missing/empty identity disables caching for that check rather than creating a global bucket. Identical actions in different runs/sessions never share approvals or denials.
38
+
37
39
  ## Outputs / response / events
38
40
 
39
- `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations` and never grant policy approval themselves.
41
+ `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations`, receive `onData(Buffer)` for ordered stdout/stderr forwarding through the shell tool's existing bounded accumulator, and never grant policy approval themselves.
40
42
 
41
43
  ## Request/response example
42
44
 
@@ -70,11 +72,13 @@ const tools = createCodingTools(workspaceRoot, {
70
72
 
71
73
  ## Extension and configuration notes
72
74
 
73
- Policies are ordinary host values: attach one globally through `createCodingTools()` or per tool. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers. Use run-scoped approval caching unless a wider host identity/lifecycle is explicit.
75
+ Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers.
76
+
77
+ Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
74
78
 
75
79
  ## Security and performance notes
76
80
 
77
- Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
81
+ Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
78
82
 
79
83
  ## Related APIs
80
84
 
@@ -6,6 +6,8 @@
6
6
 
7
7
  Current status: ledger/projection/render/recall utilities, explicit worker runtime, fast compaction strategy, inert extension helper, recall tool, and status/view command factories are available.
8
8
 
9
+ This package is distinct from `@arnilo/prism-memory` working/semantic memory: observational memory compresses and recalls source-backed observations/reflections; semantic memory retrieves embeddings; working memory stores the current structured profile/state. Hosts may compose both.
10
+
9
11
  ## When to use it
10
12
 
11
13
  Use it when a host wants to opt in to long-session memory that records observations/reflections as session custom entries, renders prepared memory during compaction, and supports exact-id recall.
@@ -179,6 +179,7 @@ Use `activateAllCapabilities: true` only as a temporary all-skills/all-tools com
179
179
  - [Agent/session runtime](agent-session-runtime.md): consumes host-selected context providers and skills from explicit agent config.
180
180
  - [Input and prompt assembly](input-and-prompt-assembly.md): default prompt builder and provider-input assembly helper.
181
181
  - [Instruction injection](instruction-injection.md): package injectors contribute `contextBlocks` that merge after host+skill provider blocks.
182
+ - [Retrieval-augmented generation](rag.md): optional retrieved citations contribute through the same explicit inert context seam.
182
183
  - [Public contracts](public-contracts.md): `ContextProvider`, `ContextResolutionContext`, `ContextBlock`, `Skill`, `SkillRegistry`, `PromptBuilder`, and `PromptBuildRequest`.
183
184
  - [Middleware hooks](middleware-hooks.md): `context` and `prompt_build` hooks.
184
185
  - [Contribution registries](contribution-registries.md): inert context provider and skill contributions.
@@ -95,7 +95,7 @@ console.log(error.message);
95
95
  ## Extension and configuration notes
96
96
 
97
97
  - Hosts and extension packages can implement `CredentialResolver` and pass it explicitly to code that needs credentials.
98
- - `AgentConfig.credentials` is host-owned metadata for compatibility; `createAgent()` / `session.run()` do not call `credentials.resolve()`. Provider adapters, compaction workers, or request policies should receive and resolve credentials at the provider edge.
98
+ - Credentials stay host-owned outside `AgentConfig`. `createAgent()` / `session.run()` do not call `credentials.resolve()`. Provider adapters, compaction workers, or request policies should receive and resolve credentials at the provider edge.
99
99
  - Use `createExplicitCredentialResolver()` when documenting a fixed order such as runtime override, stored credential, caller-provided env object, then fallback resolver.
100
100
  - Use `createEnvCredentialResolver()` only with an object supplied by the host; Prism does not read `process.env` for you.
101
101
  - Provider adapters should resolve credentials as late as possible, per request.
@@ -110,7 +110,7 @@ console.log(error.message);
110
110
  - Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via a `WeakSet` visited-set. Self-referential or mutually referenced objects render `"[Circular]"` at the back-reference instead of throwing. `Date` and `RegExp` values are passed through unchanged; `ArrayBuffer` and typed arrays are passed through unchanged; `Map` is normalized to a plain object and `Set` to an array so the output stays JSON-compatible. `errorToErrorInfo()` tolerates a cyclic `error.cause` (rendered via `String()`).
111
111
  - Use placeholders in tests and docs. Never commit real tokens.
112
112
  - Live provider/worker tests are gated behind explicit environment variables and skipped by default: `PRISM_LIVE_PROVIDER_TESTS`, `PRISM_LIVE_COMPACTION_TESTS`, `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS`. Default `npm test` is network-free; do not add ungated network calls to default tests.
113
- - `AgentConfig.credentials` is not eagerly resolved, serialized into provider requests/events/stores, or passed to loops/compaction by the core runtime.
113
+ - Credentials are not eagerly resolved by the core runtime, serialized into provider requests/events/stores, or passed to loops/compaction.
114
114
  - `resolveCredentialValue()` and `createExplicitCredentialResolver()` do not cache values. Add host-side caching only if a real credential source needs it.
115
115
  - `refreshOAuthCredential()` only calls the supplied OAuth provider and optional store; it has no built-in persistence or retry loop.
116
116
  - OpenAI Codex device-code OAuth polls inside `createOpenAICodexOAuthProvider().login()` with bounded delays and abort support via `OAuthLoginCallbacks.signal`. Token-endpoint failures redact authorization codes, PKCE verifiers, device/user codes, and access/refresh tokens when those values are known.
@@ -67,7 +67,7 @@ Important shapes:
67
67
  | `RunRecord` | Stored run with `sessionId`, `branchId`, status (`queued` \| `running` \| `succeeded` \| `failed` \| `aborted`), `model`, `provider`, `idempotencyKey`, `abortReason`, and `error`. |
68
68
  | `AgentEventRecord` | Event ledger row with `event: AgentEvent` and a `redacted` flag. Hosts redact before storage. |
69
69
  | `ToolCallRecord` | Tool-call row with `arguments`, optional `result: ToolResult`, `reason`, `progress` snapshots, status, and a `redacted` flag. |
70
- | `UsageRecord` | Usage row wrapping `Usage` with session/run/entry linkage. |
70
+ | `UsageRecord` | Scoped provider-turn or aggregate run usage with session/run/entry and turn/attempt linkage. |
71
71
  | `AgentDefinitionRecord` | Versioned agent definition snapshot. Only stores `AgentDefinition` data; never provider credentials/resolvers/instances. |
72
72
  | `RetentionPolicy` | Policy with `maxAgeDays`, `maxEntriesPerSession`, `maxTotalBytes`, `archiveStore`, and `appliedKinds`. |
73
73
  | `MigrationRecord` | Applied migration with name, version, timestamp, checksum, and applied-by. |
@@ -166,7 +166,7 @@ The `event` JSONB stores a redacted `AgentEvent`. The `sequence` column is an im
166
166
 
167
167
  | Table | Key columns |
168
168
  | --- | --- |
169
- | `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
169
+ | `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `scope`, `turn`, `attempt`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
170
170
 
171
171
  The `usage` JSONB stores the `Usage` shape: input/output/total/cache tokens, cost, and currency.
172
172
 
@@ -274,7 +274,8 @@ Recommended indexes for the reference schema. Hosts should add DB-specific parti
274
274
  | `prism_tool_calls` | `(session_id, name, started_at)` | tool usage by name |
275
275
  | `prism_tool_calls` | `(run_id, started_at)` | run tool-call listing |
276
276
  | `prism_tool_calls` | `(tool_call_id)` | deduplication / replay |
277
- | `prism_usage` | `(session_id, recorded_at)` | usage aggregation |
277
+ | `prism_usage` | `(session_id, recorded_at)` | usage pagination |
278
+ | `prism_usage` | `(session_id, scope, recorded_at)` | scope-safe billing/aggregate queries |
278
279
  | `prism_usage` | `(run_id, recorded_at)` | run usage |
279
280
  | `prism_agent_definitions` | `(name, version)` | definition lookup |
280
281
  | `prism_retention_policies` | `(tenant_id, account_id, user_id)` | policy listing |
@@ -422,18 +423,20 @@ const dbStore: ProductionPersistenceStore = {
422
423
  - `ProductionPersistenceStore` is an optional extension point. The runtime does not require it.
423
424
  - `ProductionPersistenceStore.checkpoints?: CheckpointStore` exposes generic versioned save/load/bounded-list/delete with compare-and-swap and fencing tokens, without workflow vocabulary.
424
425
  - `ProductionPersistenceStore.leases?: LeaseStore` exposes atomic acquire/renew/release/get with opaque claim tokens, expiries, ownership scope, and monotonic fencing tokens.
426
+ - `ProductionPersistenceStore.feedback?: RunFeedbackStore` exposes immutable append, bounded owned query, and owned deletion. First-party adapters store migration-003 rows in `prism_run_feedback`, FK-link `run_id`, and index owner/run/trace creation cursors.
425
427
  - Hosts choose the database, schema, transaction, and indexing strategy. The contracts specify query and checkpoint capability shapes.
426
428
  - `SessionStore` (`append`/`list`/`get`/optional `readBranchPath`) can be implemented on top of `ProductionPersistenceStore` or kept separate.
427
429
  - Cursor values and idempotency keys are host-defined and opaque to Prism.
428
- - First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume and multi-process coordination; workflow code owns no SQL table.
430
+ - First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume, human suspension, multi-process coordination, Phase 11 schedule records/fire leases, shared state, and replay lineage; workflow code owns no SQL table. `suspended`/`denied`, schedules, state history, and replay lineage remain namespaces/categories plus bounded checkpoint JSON values, so Phases 8 and 11 need no database migration.
429
431
 
430
432
  ## Security and performance notes
431
433
 
432
434
  - **No credentials in storage.** The contracts never include `CredentialResolver`, `AIProvider`, `ProviderResolver`, provider API keys, or credential values.
433
435
  - **Redact before storage.** Runtime session entries are redacted before `SessionStore.append`; `AgentEventRecord.event` and `ToolCallRecord.result` may contain secrets, so hosts must redact them (for example with `redactAgentEvent()` and a `SecretRedactor`) before writing to durable storage and set `redacted: true`.
434
- - **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility.
436
+ - **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility. First-party feedback stores require tenant plus account/user and use exact run ownership on append, query, and delete.
437
+ - **Feedback retention/deletion.** Comments, tags, and metadata are redacted and bounded before insert. `RunFeedbackStore.delete()` provides explicit erasure; deleting a run cascades its feedback in first-party SQL schemas. Apply host retention policy to `created_at`.
435
438
  - **Pagination and branch reads.** Every query supports `cursor`/`limit`/`order` so hosts can avoid full-table or full-session scans. Loading an entire large session into memory to serve a provider context is an anti-pattern; implement `readBranchPath` and use branch-relevant filters / recursive ancestor queries.
436
- - **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, and entry kind. See the reference indexes above.
439
+ - **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, entry kind, and feedback owner/run/trace creation cursors. See the reference indexes above.
437
440
 
438
441
  ## Related APIs
439
442
 
@@ -0,0 +1,122 @@
1
+ # Evaluations
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, live post-run scoring, and bounded batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
6
+
7
+ ## When to use it
8
+
9
+ Use this package when a host needs offline quality checks or sampled live scoring without coupling scorers into core agent execution. Install it directly or through `@arnilo/prism-all`; installation does not attach scorers to runs.
10
+
11
+ ## Inputs / request
12
+
13
+ | API | Key inputs |
14
+ | --- | --- |
15
+ | `defineScorer` | `id`, `score({ result, item?, expected?, signal? })` |
16
+ | `defineDataset` | `id`, `version?`, immutable `items[]` with unique ids |
17
+ | `scoreRun` / `scoreRunLive` | `AgentRunResult`, scorers, optional `sampleRate`, store, ownership, redactor |
18
+ | `runExperiment` | `agent`, dataset, scorers, bounded `concurrency`, optional store/ownership |
19
+ | `createMemoryEvaluationStore` | optional seed records |
20
+ | `appendEvaluationFeedback` | `RunFeedbackStore`, `EvaluationStore`, feedback fields, and 1–64 known evaluation IDs |
21
+
22
+ ## Outputs / response / events
23
+
24
+ | API | Output |
25
+ | --- | --- |
26
+ | `scoreRun` | `EvaluationRecord[]` with `scored` / `skipped` / `failed` |
27
+ | `scoreRunLive` | same records; never mutates the agent result; host may ignore the promise |
28
+ | `runExperiment` | `ExperimentReport` with stable item order, evaluations, and aggregates |
29
+ | `EvaluationStore.query` | cursor-paginated, ownership-filtered page |
30
+ | `appendEvaluationFeedback` | immutable `RunFeedbackRecord` containing only evaluation/scorer IDs |
31
+
32
+ ## Request/response example
33
+
34
+ ```json
35
+ {
36
+ "scorerId": "contains-citation",
37
+ "status": "scored",
38
+ "score": 1,
39
+ "runId": "run_1",
40
+ "sessionId": "session_1",
41
+ "experimentId": "exp_1",
42
+ "sampled": true
43
+ }
44
+ ```
45
+
46
+ ## Implementation example
47
+
48
+ ```ts
49
+ import { createAgent, createMemoryRunFeedbackStore, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
50
+ import {
51
+ appendEvaluationFeedback,
52
+ createMemoryEvaluationStore,
53
+ defineDataset,
54
+ defineScorer,
55
+ runExperiment,
56
+ scoreRunLive,
57
+ } from "@arnilo/prism-evals";
58
+
59
+ const scorer = defineScorer({
60
+ id: "contains-citation",
61
+ score: ({ result }) => ({ score: result.text.includes("[") ? 1 : 0 }),
62
+ });
63
+
64
+ const dataset = defineDataset({
65
+ id: "citations",
66
+ version: "1",
67
+ items: [{ id: "1", input: "Summarize with a citation" }],
68
+ });
69
+
70
+ const agent = createAgent({
71
+ model: { provider: "mock", model: "demo" },
72
+ provider: createMockProvider([providerTextDelta("ok [1]"), providerDone()]),
73
+ });
74
+
75
+ const store = createMemoryEvaluationStore();
76
+ const report = await runExperiment({
77
+ agent,
78
+ dataset,
79
+ scorers: [scorer],
80
+ concurrency: 2,
81
+ store,
82
+ ownership: { tenantId: "t1", userId: "u1" },
83
+ });
84
+
85
+ const result = await agent.createSession().run("Follow up");
86
+ void scoreRunLive(result, { scorers: [scorer], store });
87
+ const evaluation = report.evaluations[0]!;
88
+ const feedbackStore = createMemoryRunFeedbackStore({
89
+ resolveRun: ({ runId }) => runId === evaluation.runId
90
+ ? { runId, sessionId: evaluation.sessionId!, tenantId: "t1", userId: "u1" }
91
+ : false,
92
+ });
93
+ const linked = await appendEvaluationFeedback({
94
+ feedbackStore,
95
+ evaluationStore: store,
96
+ evaluationIds: [evaluation.id],
97
+ feedback: { id: "fb_1", runId: evaluation.runId!, rating: 1, tenantId: "t1", userId: "u1" },
98
+ });
99
+ console.log(report.aggregate.meanScore, linked.evaluationIds);
100
+ ```
101
+
102
+ ## Extension and configuration notes
103
+
104
+ - Function scorers are the base primitive. No mandatory LLM judge, dashboard, or schema library is included.
105
+ - Evaluation-result persistence remains package-local (`EvaluationStore`) and in-memory by default. Linked feedback is separately durable through optional `ProductionPersistenceStore.feedback`; evaluation score/reason payloads are not copied there.
106
+ - `sampleRate` is explicit (`0`–`1`). Inject `random` for deterministic tests.
107
+ - Dataset snapshots are frozen; duplicate item ids fail closed.
108
+ - `appendEvaluationFeedback()` resolves every supplied ID from `EvaluationStore`, rejects missing IDs, verifies each evaluation has the same run, optional trace, and exact ownership as feedback, then copies only deduplicated `evaluationIds`/`scorerIds`. Evaluation scores, reasons, errors, and metadata are not duplicated.
109
+
110
+ ## Security and performance notes
111
+
112
+ - Scorers receive result/item data only. Credentials, tools, and workspace access are not provided unless the host deliberately closes over them.
113
+ - Records pass through `SecretRedactor` / `secrets` before store append.
114
+ - Queries filter by ownership scope. Feedback linkage additionally requires tenant plus account/user and the feedback store re-verifies the run.
115
+ - Experiment concurrency defaults to `1` and is capped at `32`. Scoring can reference run IDs without duplicating unbounded event payloads.
116
+
117
+ ## Related APIs
118
+
119
+ - [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
120
+ - [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
121
+ - [Observability](observability.md): trace/run metadata hosts may copy into `traceId`
122
+ - [Release and install](release-and-install.md): optional package install
@@ -101,7 +101,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
101
101
  ## Extension and configuration notes
102
102
 
103
103
  - Extension loading is explicit. Prism does not discover packages, read manifests, or load filesystem config in the kernel.
104
- - `AgentConfig.extensions` is host-owned metadata for compatibility; `createAgent()` and `session.run()` do not load it or call `Extension.setup()`. Load extensions with `createExtensionKernel().load(...)`, then pass selected contributions (`tools`, `context`, `skills`, middleware, etc.) into `createAgent()`.
104
+ - Extensions stay host-owned outside `AgentConfig`. `createAgent()` and `session.run()` do not load extension lists or call `Extension.setup()`. Load extensions with `createExtensionKernel().load(...)`, then pass selected contributions (`tools`, `context`, `skills`, middleware, etc.) into `createAgent()`.
105
105
  - Setup order is the order provided by the host.
106
106
  - The kernel writes only to explicit registries returned by `createContributionRegistries()` or provided by the host.
107
107
  - `api.registerTool()` contributes an inert `ToolDefinition` to `registries.tools`; it does not add the tool to an active tool registry, allow list, or dispatch loop.
@@ -117,7 +117,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
117
117
  ## Security and performance notes
118
118
 
119
119
  - No hidden global extension kernel, provider registry, credential resolver, settings provider, store, or resource loader is created.
120
- - `AgentConfig.extensions` does not auto-execute, so constructing or running an agent cannot unexpectedly run extension code.
120
+ - Extension packages do not auto-execute from `createAgent()`, so constructing or running an agent cannot unexpectedly run extension code.
121
121
  - Error events use `ErrorInfo` and redact only known secret values passed in `secrets`.
122
122
  - Do not put resolved credential values in extension events, registry metadata, docs, logs, prompts, or session stores.
123
123
  - Event and middleware dispatch are ordered and dependency-free. They use no timers, background workers, filesystem discovery, network calls, provider calls, or tool execution.
@@ -26,9 +26,12 @@ Start from explicit host inputs. Do not let runtime code discover security state
26
26
  | Tool allow-list | active tools for this agent/session/run | `createToolRegistry`, `filterTools()`, `dispatchToolCall()` |
27
27
  | Tool argument rules | host validator | `AgentConfig.validator`, `RunOptions.validate`, `ToolValidator` |
28
28
  | Coding execution policy | path/command approval adapter | `ExecutionPolicy`, `@arnilo/prism-coding-security` |
29
+ | Remote media policy | public/default pinned DNS or explicit trusted transport | `SsrfPolicy`, `resolveMediaContentBlock()` |
29
30
  | Durable history | host database adapter | `SessionStore`, `assertSessionStoreConforms()` |
30
31
  | Durable audit | host ledger adapter | `RunLedger`, `redactRunLedgerRecord()` |
31
32
  | Extensions | explicit package imports only | `createExtensionKernel`, `ExtensionAPI` |
33
+ | Remote agent/workflow API | host authentication + ownership mapping | `@arnilo/prism-server`, `createPrismHandler()` |
34
+ | MCP server exposure | host MCP auth + selected capability list | `createPrismMcpServer()`, `createPrismMcpWebHandler()` |
32
35
 
33
36
  ## Outputs / response / events
34
37
 
@@ -105,7 +108,7 @@ Wire those values where they matter: provider adapters receive the resolved cred
105
108
 
106
109
  ## Extension and configuration notes
107
110
 
108
- - Keep security state explicit. `AgentConfig.settings` and `AgentConfig.credentials` are host-owned metadata; `createAgent()` and `session.run()` do not automatically call `settings.get()` or `credentials.resolve()`.
111
+ - Keep security state explicit. Settings and credentials are host-owned outside `AgentConfig`; `createAgent()` and `session.run()` do not automatically call `settings.get()` or `credentials.resolve()`.
109
112
  - Resolve credentials at the provider/request edge, as late as possible. Do not put resolved credentials in configs, manifests, registries, prompts, messages, events, session entries, run ledgers, idempotency keys, cache keys, or logs.
110
113
  - Use `createExplicitCredentialResolver()` to document source order such as runtime override → stored credential → caller-supplied env object → fallback.
111
114
  - Use `createEnvCredentialResolver()` only with an object the host passes in. Prism does not read `process.env` for credentials.
@@ -122,8 +125,10 @@ Wire those values where they matter: provider adapters receive the resolved cred
122
125
  - Redaction is exact known-secret replacement only. It is not arbitrary secret detection, entropy scanning, or DLP.
123
126
  - Known secrets must be passed into redactors before data is emitted or persisted. Redact again in host adapters if they transform records after Prism redaction.
124
127
  - Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects.
125
- - MCP tools from `@arnilo/prism-mcp` are untrusted remote servers. Configure stdio commands and HTTP URLs explicitly; bound output with `maxResultBytes`; register prefixed tools only after trust review. See [MCP client bridge](mcp-tools.md).
126
- - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects. Use `@arnilo/prism-coding-security` for path roots, command rules, and approval caching. Prism does not provide OS sandboxing unless the host supplies a sandbox adapter.
128
+ - MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Configure stdio commands and HTTP URLs explicitly; bound output with `maxResultBytes`; register prefixed tools only after trust review. MCP server direction exposes only passed tools/commands, requires per-call `authorize`, and retains `PermissionPolicy`/`ToolValidator` gates for tools. Its web handler needs host `resolveAuthInfo`, TLS, rate limiting, and exact host/origin policy. See [MCP client/server exposure](mcp-tools.md).
129
+ - `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive non-empty ownership from validated host identity, never request JSON. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
130
+ - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Caching defaults to none; missing run/session identity never falls back to a global cache. Prism does not provide OS sandboxing unless the host supplies a sandbox adapter.
131
+ - Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
127
132
  - Permission checks happen before tool validation and before `tool.execute()`. Middleware cannot grant permission by renaming a tool.
128
133
  - Session stores and ledgers receive redacted values when a redactor is active, but durable storage remains host-owned. Enforce tenant/account/user ownership and retention in the database layer.
129
134
  - Provider-owned auth/content/session/cache/security headers win over caller headers in adapters that merge headers.
@@ -140,8 +145,20 @@ Wire those values where they matter: provider adapters receive the resolved cred
140
145
 
141
146
  PostgreSQL TLS/network policy, MCP endpoint allow-listing, provider base URLs, OS keychain availability, process sandboxing, workflow tenant identity, and ANSI/control-sequence sanitization in any host terminal renderer remain host boundaries. Prism 0.0.4 ships JSON-line RPC, not an interactive TUI; hosts must render untrusted model/tool text safely. Credential-gated PostgreSQL/provider/keychain tests are separate operator/CI gates, not silently replaced by mocks.
142
147
 
148
+ ## Supervisor and A2A boundaries
149
+
150
+ - Register children explicitly. AND-compose parent/child/hook permissions; never let a delegation hook replace broader parent policy.
151
+ - Build each child's context/memory with supervisor-provided `resourceId`/`threadId`; resolve provider credentials inside that child factory.
152
+ - Keep depth, active children, input, turn/tool/token, timeout, and queue ceilings finite; propagate abort through nested calls.
153
+ - Expose A2A only behind per-request authentication/authorization, TLS, edge rate limits, and replay policy. Public card discovery grants no invoke access.
154
+ - Remote A2A endpoints/card URLs require exact HTTPS origin allow-lists and redirect rejection. Pin ES256 card keys/expiry; never auto-fetch untrusted `jku`.
155
+ - Treat cards, task status, errors, artifacts, and SSE frames as untrusted bounded input and redact before logs/hooks/events.
156
+
143
157
  ## Related APIs
144
158
 
159
+ - [Web-standard server handler](server.md): remote agent/workflow route, ownership, limits, abort, and deployment boundary.
160
+ - [Supervisor delegation](supervisors.md): local child permission/memory/budget boundary.
161
+ - [A2A interoperability](a2a.md): remote card/auth/origin/signature boundary.
145
162
  - [Settings, auth, trust, and security controls](settings-auth-trust-security.md): low-level helpers and boundary hardening table.
146
163
  - [Credentials and redaction](credentials-and-redaction.md): credential resolver order, caller-supplied env objects, OAuth refresh, exact redaction, and no persistent secret store.
147
164
  - [Tools](tools.md): active tool registry, allow/deny filters, permission order, validator order, blocked events, and no sandbox.