@arnilo/prism 0.0.5 → 0.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/CHANGELOG.md +39 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +27 -16
  4. package/dist/agent-run-lifecycle.d.ts +28 -0
  5. package/dist/agent-run-lifecycle.js +33 -0
  6. package/dist/agent-run-state.d.ts +53 -0
  7. package/dist/agent-run-state.js +127 -0
  8. package/dist/agents.d.ts +3 -1
  9. package/dist/agents.js +337 -46
  10. package/dist/contracts.d.ts +205 -3
  11. package/dist/contracts.js +4 -0
  12. package/dist/guardrails.d.ts +25 -0
  13. package/dist/guardrails.js +133 -0
  14. package/dist/ids.d.ts +2 -0
  15. package/dist/ids.js +6 -0
  16. package/dist/index.d.ts +17 -3
  17. package/dist/index.js +10 -3
  18. package/dist/input.js +2 -0
  19. package/dist/resources.js +2 -1
  20. package/dist/run-limits.d.ts +34 -0
  21. package/dist/run-limits.js +163 -0
  22. package/dist/secure-agent.d.ts +3 -0
  23. package/dist/secure-agent.js +63 -0
  24. package/dist/session-stores.js +2 -3
  25. package/dist/testing/persistence-schema.d.ts +45 -7
  26. package/dist/testing/persistence-schema.js +138 -24
  27. package/dist/thinking.d.ts +42 -0
  28. package/dist/thinking.js +92 -0
  29. package/dist/tools.d.ts +10 -2
  30. package/dist/tools.js +56 -7
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +4 -2
  34. package/docs/agent-events.md +23 -16
  35. package/docs/agent-loops.md +19 -8
  36. package/docs/agent-session-runtime.md +33 -1
  37. package/docs/coding-agent-tools.md +33 -12
  38. package/docs/coding-security.md +2 -2
  39. package/docs/compaction-llm.md +17 -7
  40. package/docs/compaction-observational-memory.md +28 -4
  41. package/docs/credential-storage.md +58 -9
  42. package/docs/credentials-and-redaction.md +1 -1
  43. package/docs/database-persistence.md +8 -3
  44. package/docs/guardrails.md +75 -0
  45. package/docs/host-security.md +16 -8
  46. package/docs/index.md +26 -22
  47. package/docs/mcp-tools.md +32 -12
  48. package/docs/migration.md +164 -2
  49. package/docs/node-filesystem-config.md +1 -0
  50. package/docs/node-jsonl-session-store.md +5 -4
  51. package/docs/postgres-persistence.md +3 -3
  52. package/docs/provider-caching.md +16 -4
  53. package/docs/provider-conformance.md +39 -1
  54. package/docs/provider-packages.md +60 -3
  55. package/docs/providers/ai-sdk.md +36 -0
  56. package/docs/providers/kimi.md +124 -61
  57. package/docs/providers/neuralwatt.md +19 -13
  58. package/docs/providers/openai.md +56 -13
  59. package/docs/providers/opencode-go.md +118 -30
  60. package/docs/providers/openrouter.md +105 -35
  61. package/docs/providers/zai.md +94 -45
  62. package/docs/release-and-install.md +47 -49
  63. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  64. package/docs/runs-and-usage.md +30 -3
  65. package/docs/server.md +5 -2
  66. package/docs/sqlite-persistence.md +2 -2
  67. package/docs/structured-output.md +1 -1
  68. package/docs/thinking-and-reasoning.md +98 -0
  69. package/docs/tool-execution-primitives.md +3 -3
  70. package/docs/tools.md +21 -1
  71. package/docs/use-case-model-selection.md +109 -0
  72. package/docs/workflow-orchestration-primitives.md +1 -0
  73. package/docs/workflows.md +18 -10
  74. package/docs/working-and-semantic-memory.md +1 -0
  75. package/package.json +2 -2
@@ -15,6 +15,8 @@
15
15
  | `createAllTools(cwd, options?)` | Every tool the package provides (currently identical to `createCodingTools`). |
16
16
  | `detectSupportedImageMimeType(buf)` / `detectSupportedImageMimeTypeFromFile(path)` | Magic-byte image MIME detection (PNG/JPEG/GIF/WebP/BMP) used by `read`. |
17
17
  | `DEFAULT_MAX_IMAGE_BYTES` | Default `read` image size ceiling (10 MB). |
18
+ | `DEFAULT_*` / `HARD_*` coding limit constants | Published text-scan, image, write/edit, shell timeout, display, and total-output ceilings. |
19
+ | `ReadTextOptions` / `ReadTextResult` | Bounded text-page contract required by custom `ReadOperations`. |
18
20
  | `TransformImage` / `TransformImageInput` | Types for the optional `read` `transformImage` callback. |
19
21
  | `withFileMutationQueue(path, fn)` | Per-path serialization primitive re-exported for hosts. |
20
22
 
@@ -65,7 +67,7 @@ Run a shell command and return combined stdout+stderr.
65
67
  | Field | Type | Purpose |
66
68
  | --- | --- | --- |
67
69
  | `command` | `string` | Shell command to execute (required). |
68
- | `timeout` | `number` | Timeout in **seconds** (optional; no default). |
70
+ | `timeout` | `number` | Timeout in **seconds** (optional; defaults to 600, hard maximum 3600). |
69
71
 
70
72
  **Outputs:** a `ToolResult` whose `content[0]` is a `TextContent` with the combined output. Non-zero exit is **not** a tool error: it is returned as a normal result with `[Command exited with code N]` appended to the content and `exitCode` in metadata. Timeout and abort are error results that still carry the partial output captured so far.
71
73
 
@@ -75,9 +77,11 @@ Run a shell command and return combined stdout+stderr.
75
77
  | --- | --- | --- |
76
78
  | `exitCode` | always | Process exit code, or `null` when the process was killed by timeout/abort. |
77
79
  | `truncation` | always | `TruncationResult` from the bounded output accumulator. |
78
- | `fullOutputPath?` | truncated only | Path to the spilled temp file holding the full output. |
80
+ | `fullOutputPath?` | successful and truncated only | Host-owned path to retained output. Failed/aborted/timed-out/output-limited calls remove unpublished spills. |
81
+ | `totalOutputBytes` | shell executed | Raw bytes retained, never above `maxTotalOutputBytes`. |
82
+ | `outputLimitExceeded` / `outputStorageFailed` | shell executed | Attributable resource failure flags. |
79
83
 
80
- Shell resolution honors `options.shellPath` → `SHELL` env → `/bin/bash` → `sh`, and the process group is killed on timeout/abort (`process.kill(-pid)` on Unix, `taskkill /F /T` on Windows).
84
+ Shell resolution honors `options.shellPath` → `SHELL` env → `/bin/bash` → `sh`. The process group is killed on timeout, caller abort, spill failure, or total-output overflow (`process.kill(-pid)` on Unix, `taskkill /F /T` on Windows). Combined output defaults to a 64 MiB total cap (1 GiB hard cap). Spill files use random exclusive creation and Unix mode `0600`; hosts own and must delete a successful result's `fullOutputPath` after consumption.
81
85
 
82
86
  ### `read`
83
87
 
@@ -91,7 +95,7 @@ Read a text or image file.
91
95
  | `offset` | `number` | Line to start reading from (1-indexed). |
92
96
  | `limit` | `number` | Maximum number of lines to read. |
93
97
 
94
- **Outputs:** text files become a single `TextContent`, truncated to `maxLines`/`maxBytes` (defaults 2000 lines / 50 KB) with a `Use offset=N to continue` footer when more remains. Image files (PNG/JPEG/GIF/WebP/BMP by **magic bytes**, not extension) become `[TextContent note, ImageContent]` with base64 `data` and `mimeType`. Oversize images are rejected by `stat` (when available) or `buffer.length` against `maxImageBytes` (default 10 MB) before base64 encoding. An optional `transformImage` callback lets hosts resize or re-encode images without adding image-processing dependencies to the base package. Read failures (missing file, offset beyond end, oversize image, abort) are error results.
98
+ **Outputs:** text files are scanned incrementally until one requested page, `maxLines`/`maxBytes`, EOF, or `maxScanBytes` (default 64 MiB scanned per call; 1 GiB hard cap). The default path never loads the complete file and returns a `Use offset=N to continue` footer when more remains. Exact total line count is reported only when EOF was already reached in the bounded scan. Image files (PNG/JPEG/GIF/WebP/BMP by **magic bytes**, not extension) become `[TextContent note, ImageContent]` with base64 `data` and `mimeType`. Oversize images are rejected by `stat` (when available) or `buffer.length` against `maxImageBytes` (default 10 MB) before base64 encoding. An optional `transformImage` callback lets hosts resize or re-encode images without adding image-processing dependencies to the base package. Read failures (missing file, offset beyond end, oversize image, abort) are error results.
95
99
 
96
100
  `read` tool options (via `createReadTool(cwd, options)` or `ToolsOptions.read`):
97
101
 
@@ -100,8 +104,9 @@ Read a text or image file.
100
104
  | `maxImageBytes` | `DEFAULT_MAX_IMAGE_BYTES` (10 MB) | Reject image reads larger than this many bytes. |
101
105
  | `transformImage` | — | Host callback `( { buffer, mimeType } ) => Promise<Buffer>` run after read, before base64. |
102
106
  | `autoResizeImages` | — | **Deprecated.** Ignored unless `transformImage` is also set (use `transformImage` instead). |
103
- | `maxLines` / `maxBytes` | 2000 / 50 KB | Text head truncation limits. |
104
- | `operations` | local fs | Pluggable `ReadOperations` backend. |
107
+ | `maxLines` / `maxBytes` | 2000 / 50 KiB | Text page display limits (hard: 100,000 / 1 MiB). |
108
+ | `maxScanBytes` | 64 MiB | Raw bytes scanned to reach one page (hard: 1 GiB). |
109
+ | `operations` | local fs | Pluggable bounded `ReadOperations` backend. |
105
110
  | `executionPolicy` | — | Structured pre-execution policy (see [Coding security](coding-security.md)). |
106
111
 
107
112
  ```ts
@@ -133,7 +138,7 @@ Create or overwrite a file, creating parent directories as needed.
133
138
  | `path` | `string` | Path to the file to write (relative or absolute). Required. |
134
139
  | `content` | `string` | Content to write (empty string creates an empty file). Required. |
135
140
 
136
- **Outputs:** a `TextContent` confirmation naming the **absolute path** with UTF-8 byte and line counts (e.g. `Successfully wrote 42 bytes (3 lines) to /abs/path.txt`). Write failures and abort are error results. Empty `content` is valid.
141
+ **Outputs:** a `TextContent` confirmation naming the **absolute path** with UTF-8 byte and line counts (e.g. `Successfully wrote 42 bytes (3 lines) to /abs/path.txt`). `maxInputBytes` defaults to 8 MiB (64 MiB hard cap); oversized UTF-8 input fails before policy evaluation, directory creation, or write. Write failures and abort are error results. Empty `content` is valid.
137
142
 
138
143
  `write` result `metadata`: `{ bytes, lines, path }` (absolute path). Concurrent writes to the same path serialize through `withFileMutationQueue`; writes to different paths run in parallel.
139
144
 
@@ -148,7 +153,7 @@ Precise text replacement in an existing file via exact-then-fuzzy matching.
148
153
  | `path` | `string` | Path to the file to edit. Required. |
149
154
  | `edits` | `Array<{ oldText: string, newText: string }>` | Targeted replacements, each matched against the **original** file (not incrementally). No overlapping/nested edits. Required, non-empty. |
150
155
 
151
- Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse). A BOM is stripped before matching and re-prepended on write; original line endings are restored.
156
+ Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse). A BOM is stripped before matching and re-prepended on write; original line endings are restored. Defaults reject targets over 8 MiB, aggregate old/new UTF-8 input over 2 MiB, or more than 100 edits (hard caps: 64 MiB, 16 MiB, and 1,000). Stat and bounded read checks run before matching or mutation.
152
157
 
153
158
  **Outputs:** a `TextContent` confirmation (`Successfully replaced N block(s) in {path}.`) plus `metadata`. Any failure — missing/unreadable file, no match, duplicate (non-unique) match, overlap, empty `oldText`, no-op edit, or abort — is an error result, and the file is left **unchanged** (the match runs before the write).
154
159
 
@@ -156,7 +161,7 @@ Each `edits[].oldText` must match a unique, non-overlapping region of the origin
156
161
 
157
162
  ## Outputs / response / events
158
163
 
159
- Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`. Mutating tools (`shell` with same cwd, `write`, `edit`) serialize per realpath through `withFileMutationQueue` so concurrent calls targeting one file do not interleave. The package emits no events of its own; hosts observe tool execution through the normal Prism `AgentEvent` stream via `dispatchToolCall`.
164
+ Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`. `write` and `edit` serialize per realpath through `withFileMutationQueue` so concurrent calls targeting one file do not interleave. `shell` is marked `exclusive`; tool dispatch serializes it at the turn level. The package emits no events of its own; hosts observe tool execution through the normal Prism `AgentEvent` stream via `dispatchToolCall`.
160
165
 
161
166
  ## Request/response example
162
167
 
@@ -208,6 +213,8 @@ const shell = createShellTool("/repo", {
208
213
  shellPath: "/bin/bash",
209
214
  commandPrefix: "set -euo pipefail",
210
215
  maxLines: 500,
216
+ timeout: 600,
217
+ maxTotalOutputBytes: 64 * 1024 * 1024,
211
218
  });
212
219
 
213
220
  const remoteWrite = createWriteTool("/repo", {
@@ -220,8 +227,8 @@ const remoteWrite = createWriteTool("/repo", {
220
227
 
221
228
  ## Extension and configuration notes
222
229
 
223
- - **Pluggable operation backends.** Every tool accepts an `operations` seam so a host can delegate to a remote system (e.g. SSH) while keeping the tool's matching/serialization behavior: `BashOperations` (`shell`), `ReadOperations` (`read`), `WriteOperations` (`write`), `EditOperations` (`edit`).
224
- - **Per-tool options.** `ShellToolOptions` (`shellPath`, `commandPrefix`, `maxLines`, `maxBytes`, `tempFilePrefix`, `operations`, `spawnHook`, `executionPolicy`); `ReadToolOptions` (`operations`, `maxImageBytes`, `transformImage`, `maxLines`, `maxBytes`, `executionPolicy`; `autoResizeImages` deprecated); `WriteToolOptions` (`operations`, `executionPolicy`); `EditToolOptions` (`operations`, `executionPolicy`).
230
+ - **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
231
+ - **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`. Invalid/non-finite/unsafe/above-hard-cap values throw during tool construction; request `timeout` errors before spawn.
225
232
  - **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override.
226
233
  - **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
227
234
  - No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
@@ -230,10 +237,24 @@ const remoteWrite = createWriteTool("/repo", {
230
237
 
231
238
  - **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
232
239
  - **Non-zero exit is not an error.** A failing command is a normal `shell` result (exit code in metadata); only timeout/abort/spawn failures are error results. Do not assume `error == undefined` means the command succeeded.
233
- - **Bounded output.** `shell`/`read` accumulate output into a rolling tail bounded by `maxLines`/`maxBytes`; oversized output spills to a temp file (`fullOutputPath`), so memory use is bounded regardless of command output size.
240
+ - **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
234
241
  - **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
235
242
  - **Bounded image reads.** `read` rejects images over `maxImageBytes` (default 10 MB) by `stat` before read when possible; MIME is detected from magic bytes only. Optional `transformImage` is host-owned — the base package has no image-processing dependency.
236
243
 
244
+ ### Resource-limit defaults and hard caps
245
+
246
+ | Boundary | Default | Hard cap | Failure point |
247
+ | --- | ---: | ---: | --- |
248
+ | Display lines / bytes | 2,000 / 50 KiB | 100,000 / 1 MiB | tool construction |
249
+ | Text scan per read | 64 MiB | 1 GiB | bounded scan before more input is retained |
250
+ | Image | 10,000,000 bytes | 32 MiB | stat and bounded read before base64/transform result use |
251
+ | Write UTF-8 input | 8 MiB | 64 MiB | before policy/filesystem mutation |
252
+ | Edit target / input / count | 8 MiB / 2 MiB / 100 | 64 MiB / 16 MiB / 1,000 | before target read/matching/write |
253
+ | Shell wall time | 600 seconds | 3,600 seconds | process-tree kill |
254
+ | Shell total stdout+stderr | 64 MiB | 1 GiB | process-tree kill; spill removal |
255
+
256
+ Every configurable value is a positive safe integer; Prism rejects rather than clamps invalid values. Limits control resources, not authority: they do not replace root containment, approval, validation, or a sandbox.
257
+
237
258
  ## Related APIs
238
259
 
239
260
  - [Tools](tools.md): the host-owned tool harness — `createToolRegistry`, `dispatchToolCall`, filtering, and the `ToolDefinition` contract these factories satisfy.
@@ -74,11 +74,11 @@ const tools = createCodingTools(workspaceRoot, {
74
74
 
75
75
  Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers.
76
76
 
77
- Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
77
+ Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
78
78
 
79
79
  ## Security and performance notes
80
80
 
81
- Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
81
+ Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
82
82
 
83
83
  ## Related APIs
84
84
 
@@ -25,15 +25,16 @@ Key exports:
25
25
  | Field | Purpose |
26
26
  | --- | --- |
27
27
  | `provider` / `summaryProvider` | Explicit `AIProvider`, or factory receiving a resolved credential. |
28
- | `model` / `summaryModel` | Explicit summary `ModelConfig`. |
28
+ | `model` / `summaryModel` | Summary `ModelConfig`. `summaryModel` wins; `model` is the session/host fallback via `resolveUseCaseModel`. See [Use-case model selection](use-case-model-selection.md). |
29
29
  | `credential`, `credentialRequest` | Optional per-call credential resolution for provider factories. |
30
30
  | `providerOptions` | Generic `ProviderRequest.options`, including cache fields. |
31
31
  | `providerRequestPolicies` | Optional Prism provider request policies applied before the summary call. |
32
32
  | `customInstructions` | Additional summary focus appended to prompts. |
33
- | `thinkingLevel` | Passed as `ProviderRequest.options.extra.thinkingLevel`. |
34
- | `reserveTokens` | Output budget basis; defaults to `16384`. |
33
+ | `thinkingLevel` | Mapped into `ProviderRequest.options.compat` via `applyThinkingLevel` / `thinkingFamilyForModel` (not inert `extra.thinkingLevel`). See [Thinking and reasoning](thinking-and-reasoning.md). |
34
+ | `reserveTokens` | Output budget basis; defaults to `16384`, hard cap `131072`. |
35
35
  | `keepRecentTokens` | Approximate recent-token budget; defaults to `20000`. |
36
- | `maxSummaryTokens` / `maxOutputTokens` | Computes the summary output budget, writes it to `summaryModel/model.parameters.maxTokens`, and truncates oversized collected summaries. First-party providers serialize that generic field to their real request field (`max_output_tokens` for OpenAI Responses, `max_tokens` for OpenAI-compatible/Anthropic-style providers). |
36
+ | `maxSummaryTokens` / `maxOutputTokens` | Summary retention/request ceiling; default `16384`, hard cap `131072`. `maxSummaryTokens` wins over the compatibility alias. The finite value is written to `model.parameters.maxTokens`; first-party providers map it to their wire field. |
37
+ | `maxErrorBytes` | Retained provider/factory/policy error detail; default `1024`, hard cap `8192`, UTF-8-safe and known-secret redacted. |
37
38
  | `maxToolResultChars` | Tool-result JSON truncation limit; defaults to `2000`. |
38
39
  | `trackFileOperations`, `includeFileOperations` | Control file path extraction and final summary blocks. |
39
40
  | `secrets` | Exact strings to redact from serialized prompts and final summaries. |
@@ -64,7 +65,8 @@ const strategy = createLlmCompactionStrategy({
64
65
  model: { provider: "openai", model: "gpt-4.1-mini" },
65
66
  keepRecentTokens: 20_000,
66
67
  reserveTokens: 16_384,
67
- maxOutputTokens: 800,
68
+ maxSummaryTokens: 800,
69
+ maxErrorBytes: 1_024,
68
70
  providerOptions: { cacheRetention: "short" },
69
71
  customInstructions: "Focus on current files and failing tests.",
70
72
  });
@@ -102,10 +104,18 @@ const agent = createAgent({ model, provider, compaction: { strategy, thresholdEn
102
104
  Registration only contributes an inert strategy. The host must resolve and pass it to runtime config.
103
105
 
104
106
  ## Security and performance notes
105
- Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization. Output-budget calculation is O(1) and does not add an extra summarization call. The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. It does not discover credentials, read files, start background jobs, or add provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history or provider output.
107
+ Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization. Limit options must be positive safe integers at or below their hard caps and reject during strategy creation. Missing output options use a 16,384-token summary ceiling; reserve ratio/model metadata may narrow the provider request, never remove its finite `maxTokens`. A request policy that replaces `maxTokens` with NaN, Infinity, zero, an unsafe integer, or above-hard-cap input fails before provider generation.
108
+
109
+ Provider deltas are redacted while retained and stop at `maxSummaryTokens * 4` UTF-16 code units without splitting a surrogate pair. Provider iteration is closed/aborted on overflow. A derived finite event ceiling also stops endless empty/non-text deltas. Final history/turn/file composition receives the same cap. Provider error events, generator throws, provider-factory failures, and policy failures expose only bounded redacted detail; host abort remains authoritative.
110
+
111
+ The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. It does not discover credentials, read files, start background jobs, or add provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history or provider output.
106
112
 
107
113
  ## Related APIs
108
- - [Compaction and retry policies](compaction-and-retry.md): core compaction strategy surface.
114
+
115
+ - [Use-case model selection](use-case-model-selection.md): `summaryModel` vs session `model` fallback.
116
+ - [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → `compat`.
117
+ - [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary and core compaction strategy surface.
118
+ - [Observational memory compaction package](compaction-observational-memory.md): source-backed memory workers with the same use-case binding pattern.
109
119
  - [Agent/session runtime](agent-session-runtime.md): `AgentSession.compact()` and opt-in auto-compaction.
110
120
  - [Provider layer](provider-layer.md): mock providers and provider request contracts.
111
121
  - [Credentials and redaction](credentials-and-redaction.md): exact known-secret redaction behavior.
@@ -27,6 +27,20 @@ Memory records use `SessionEntry.kind: "custom"` with `entry.data.type` markers:
27
27
 
28
28
  Ids are known, source-backed 12-character lowercase hex strings matching `^[a-f0-9]{12}$`.
29
29
 
30
+ Worker limits are finite positive safe integers:
31
+
32
+ | Runtime option | Default | Hard cap | Scope |
33
+ | --- | ---: | ---: | --- |
34
+ | `maxWorkerTurns` | 16 | 64 | Provider turns per observer/reflector/dropper run; overrides settings `agentMaxTurns` |
35
+ | `maxWorkerToolCallsPerTurn` | 32 | 256 | Calls retained from one provider response |
36
+ | `maxWorkerToolCalls` | 128 | 1,024 | Calls across all turns in one worker run |
37
+ | `maxWorkerArgumentBytes` | 64 KiB | 1 MiB | Each raw and redacted JSON argument object |
38
+ | `maxWorkerResultBytes` | 64 KiB | 1 MiB | Full tool result and replayed value/error payload |
39
+ | `maxWorkerMessageBytes` | 1 MiB | 8 MiB | System/prompt plus assistant-call/tool-result transcript |
40
+ | `maxWorkerErrorBytes` | 1 KiB | 8 KiB | Provider/tool/runtime error text after exact known-secret redaction |
41
+
42
+ Direct `runObserver()` / `runReflector()` / `runDropper()` calls retain required `maxTurns` and accept the corresponding shorter worker fields (`maxToolCalls`, `maxResultBytes`, etc.). Named default/hard constants and `resolveMemoryWorkerLimits()` are exported.
43
+
30
44
  ## Outputs / response / events
31
45
 
32
46
  Key exports:
@@ -78,7 +92,12 @@ const memory = createObservationalMemoryRuntime({
78
92
  session,
79
93
  appendEntry: (entry) => store.append(entry),
80
94
  workerProvider,
81
- workerModel: { provider: "mock", model: "memory" },
95
+ sessionModel: agent.config.model, // fallback when workerModel unset
96
+ // workerModel: { provider: "mock", model: "memory" }, // optional override
97
+ maxWorkerTurns: 8,
98
+ maxWorkerToolCalls: 64,
99
+ maxWorkerResultBytes: 64 * 1024,
100
+ overrides: { thinkingLevel: "low" },
82
101
  });
83
102
  await memory.flush();
84
103
  await session.compact({ strategy: createObservationalMemoryCompactionStrategy({ keepRecentEntries: 8 }) });
@@ -92,9 +111,9 @@ await kernel.load([createObservationalMemoryExtension({ recallTool: { getEntries
92
111
 
93
112
  ## Extension and configuration notes
94
113
 
95
- Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`.
114
+ Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`. `agentMaxTurns` now rejects non-integer/non-finite/out-of-range input (hard 64) instead of flooring/falling back. Runtime `maxWorkerTurns` takes precedence.
96
115
 
97
- The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, `workerProvider`, and `workerModel`. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution.
116
+ The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, and `workerProvider`. Model selection uses [use-case model selection](use-case-model-selection.md): pass optional `workerModel` (or settings `workerModel`) to override, and `sessionModel: agent.config.model` so workers fall back to the session model when no worker model is configured. `requireExplicitModel: true` restores the historical `missing_model` skip when no explicit worker model is set. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution. Default credential requests use the **resolved** model's provider id.
98
117
 
99
118
  `createObservationalMemoryCompactionStrategy()` keeps recent message entries like the default compaction strategy, renders existing observations/reflections as the summary, and returns a standard Prism compaction entry. Its `data` includes `throughEntryId`, `keepEntryIds`, `strategy`, `trigger`, and `memory: { type: "om.folded", version: 1, fullFold, observations, reflections, droppedObservationIds }`. When active observations exceed `observationsPoolMaxTokens`, it performs a full fold into `data.memory`.
100
119
 
@@ -110,13 +129,18 @@ The runtime requires host-supplied `session`, an `appendEntry` callback bound to
110
129
  - Recall tool and commands only see current-branch entries supplied by the host callback.
111
130
  - Invalid or missing ids fail closed; invalid recall tool ids skip entry lookup.
112
131
  - Utilities and fast compaction are O(n) over supplied entries and use no provider, network, filesystem, timer, worker, credential, or settings access.
113
- - Workers serialize only supplied branch entries, enforce `agentMaxTurns`, and run one consolidation pipeline at a time per runtime. Worker transcripts replay assistant `tool_call` messages before matching role `tool` `tool_result` messages so provider requests stay valid for call/result-pairing providers.
132
+ - Workers serialize only supplied branch entries within `maxWorkerMessageBytes`, enforce finite turns/calls/arguments/results/messages/errors, and run one consolidation pipeline at a time per runtime. Source serialization and reflection/drop prompts fail before joining beyond the transcript cap.
133
+ - Every provider call must name a registered worker tool. Unknown calls, call overflow, oversized/deep/cyclic/non-JSON arguments/results, and transcript overflow fail deterministically; no excess call enters the assistant transcript or executes.
134
+ - Raw arguments are measured before tool execution. Full results are measured before redaction/replay; the bounded redacted value/error is then measured again because replacement text can grow. Replayed call arguments, tool values/errors, runtime `lastError`, and debug error data contain exact known-secret redaction. Host tools may already have caused side effects before returning an invalid oversized result; keep worker tools small/idempotent.
135
+ - Worker transcripts replay assistant `tool_call` messages before matching role `tool` `tool_result` messages so provider requests stay valid for call/result-pairing providers. Calls produced on the final allowed turn execute and persist, but no additional provider turn starts.
114
136
  - Compaction preserves raw history; Prism appends one standard compaction entry and rebuilds provider context from its summary plus kept recent messages.
115
137
  - Pass known secrets to render/recall/runtime/tool/command helpers to redact exact values from prompts, records, structured results, and text output.
116
138
  - Live tests are opt-in with `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS=1`.
117
139
 
118
140
  ## Related APIs
119
141
 
142
+ - [Use-case model selection](use-case-model-selection.md): session vs worker model binding and `resolveUseCaseModel`.
143
+ - [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → provider `compat`.
120
144
  - [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary.
121
145
  - [LLM compaction package](compaction-llm.md): existing optional compaction-package pattern.
122
146
  - [Session stores and branching](session-stores-and-branching.md): branch entries that observational memory reads and appends to.
@@ -44,8 +44,11 @@ import {
44
44
  | --- | --- | --- |
45
45
  | `path` | `string` | Vault file path. Parent directories are created as needed. |
46
46
  | `getPassphrase` | `() => string \| Promise<string>` | Host-owned passphrase retrieval. Never logged by the adapter. |
47
- | `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32`. Minimum `N=16384`. |
48
- | `fileMode` | `number` | Unix mode for newly written files. Defaults to `0o600`. |
47
+ | `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32`; limits are listed below. |
48
+ | `fileMode` | `number` | Unix mode for files. Defaults to `0o600`; group/other permissions are rejected. |
49
+ | `limits.maxFileBytes` | `number` | Encrypted envelope file: 4 MiB default, 16 MiB hard cap. |
50
+ | `limits.maxVaultBytes` | `number` | Decrypted vault/plaintext: 3 MiB default, 12 MiB hard cap. |
51
+ | `limits.maxScryptMemoryBytes` | `number` | `128*N*r` memory estimate: 256 MiB default and hard cap. |
49
52
 
50
53
  ### System keychain
51
54
 
@@ -53,7 +56,8 @@ import {
53
56
  | --- | --- | --- |
54
57
  | `service` | `string` | Keychain service name (application identifier). |
55
58
  | `namespace` | `string` | Optional prefix separating environments or tenants within one service. |
56
- | `timeoutMs` | `number` | Operation timeout. Defaults to `5000`. |
59
+ | `timeoutMs` | `number` | Operation timeout. Defaults to 5,000 ms; hard cap 60,000 ms. |
60
+ | `maxPayloadBytes` | `number` | Decrypted keychain payload: 3 MiB default, 12 MiB hard cap. |
57
61
 
58
62
  ## Outputs / response / events
59
63
 
@@ -71,6 +75,8 @@ Encrypted file stores also expose:
71
75
  - `reload()` — re-read and decrypt from disk
72
76
  - `flush()` — force rewrite of the encrypted envelope
73
77
 
78
+ `encryptBytes()` and `decryptBytes()` are Promise-based because they use asynchronous `node:crypto.scrypt`.
79
+
74
80
  Errors are explicit and fail closed:
75
81
 
76
82
  | Error | Code | When |
@@ -123,6 +129,7 @@ import {
123
129
  const store = await openEncryptedCredentialStore({
124
130
  path: "./credentials.vault",
125
131
  getPassphrase: () => process.env.MY_APP_CREDENTIAL_PASSPHRASE!,
132
+ limits: { maxFileBytes: 4 * 1024 * 1024, maxVaultBytes: 3 * 1024 * 1024 },
126
133
  });
127
134
 
128
135
  const resolver = createExplicitCredentialResolver([
@@ -151,6 +158,47 @@ await rotateEncryptedCredentialStorePassphrase({
151
158
  });
152
159
  ```
153
160
 
161
+ ### Desktop keychain and explicit overrides
162
+
163
+ ```ts
164
+ import {
165
+ createEnvCredentialResolver,
166
+ createExplicitCredentialResolver,
167
+ createMemoryCredentialStore,
168
+ } from "@arnilo/prism";
169
+ import {
170
+ createKeychainCredentialStore,
171
+ createStoredCredentialResolver,
172
+ } from "@arnilo/prism-credentials-node";
173
+ import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
174
+
175
+ const keychain = createKeychainCredentialStore({
176
+ service: "com.example.my-app",
177
+ namespace: "production",
178
+ });
179
+ await keychain.set({
180
+ name: "apiKey",
181
+ provider: "openai",
182
+ credential: { type: "api_key", value: userSuppliedKey },
183
+ });
184
+
185
+ const runtimeOverrides = createMemoryCredentialStore();
186
+ // Set only for this user/agent instance; it wins over stored and env values.
187
+ runtimeOverrides.set({
188
+ name: "apiKey",
189
+ provider: "openai",
190
+ credential: { type: "api_key", value: temporaryOverride },
191
+ });
192
+
193
+ const apiKey = createExplicitCredentialResolver([
194
+ { name: "runtime", resolver: runtimeOverrides },
195
+ { name: "keychain", resolver: createStoredCredentialResolver(keychain) },
196
+ { name: "env", resolver: createEnvCredentialResolver(process.env, { openai: "OPENAI_API_KEY" }) },
197
+ ]);
198
+
199
+ const providers = createOpenAIProviderPackage({ apiKey });
200
+ ```
201
+
154
202
  ## Extension and configuration notes
155
203
 
156
204
  - Passphrase retrieval, TLS, and OS permission prompts remain host-owned.
@@ -161,12 +209,13 @@ await rotateEncryptedCredentialStorePassphrase({
161
209
 
162
210
  ## Security and performance notes
163
211
 
164
- - Authenticated encryption uses Node built-in `aes-256-gcm` and `scrypt`; no extra crypto dependencies for the file backend.
165
- - Atomic writes use temp file + rename; partial writes cannot replace a valid vault.
166
- - Derived keys are zeroed after encrypt/decrypt operations where practical.
167
- - Default scrypt `N=32768` targets interactive CLI unlock; raise `N` for higher security at the cost of unlock latency.
168
- - Keychain operations honor `timeoutMs` and surface `CredentialStoreTimeoutError` instead of blocking indefinitely.
169
- - Never log passphrases, derived keys, or decrypted credential payloads. Error messages do not echo secret values.
212
+ - Authenticated encryption uses Node built-in `aes-256-gcm` and asynchronous `scrypt`; no extra crypto dependency is added for the file backend.
213
+ - Envelope parsing rejects unknown shape, non-canonical/oversized base64, wrong salt/IV/tag size, unsupported algorithms/version, and excessive KDF work before scrypt. `N` must be a power of two from 16,384–262,144; `r≤32`, `p≤16`, `keyLength=32`, `N*r*p≤2,097,152`, and `128*N*r` must fit `maxScryptMemoryBytes`.
214
+ - Existing Unix vaults are checked before content read and must deny group/other access. Atomic writes create a random exclusive temp file at the requested restrictive mode, then rename; Windows skips Unix mode checks.
215
+ - Derived keys and package-owned plaintext buffers are zeroed after use. JavaScript passphrase strings and returned credentials remain host-owned.
216
+ - Keychain operations use `@napi-rs/keyring`'s abort-aware `AsyncEntry`, so native work runs outside the JavaScript event loop. A main-loop timer aborts and rejects at `timeoutMs`; native cancellation remains OS/backend-dependent and may briefly retain one libuv worker after rejection.
217
+ - Keychain payloads are bytes rather than password strings and are zeroed after parse/write. Unknown native errors are mapped to sanitized typed errors; no native message or secret value is echoed.
218
+ - Never log passphrases, derived keys, or decrypted credential payloads.
170
219
  - Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
171
220
 
172
221
  ## Related APIs
@@ -107,7 +107,7 @@ console.log(error.message);
107
107
  - Redaction only removes exact known secret values passed to the helper. It is not a general-purpose secret detector.
108
108
  - Do not pass empty strings as secrets; they are ignored.
109
109
  - `redactSecrets()` recursively walks arrays and object entries, so avoid using it on huge objects unless needed.
110
- - Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via a `WeakSet` visited-set. Self-referential or mutually referenced objects render `"[Circular]"` at the back-reference instead of throwing. `Date` and `RegExp` values are passed through unchanged; `ArrayBuffer` and typed arrays are passed through unchanged; `Map` is normalized to a plain object and `Set` to an array so the output stays JSON-compatible. `errorToErrorInfo()` tolerates a cyclic `error.cause` (rendered via `String()`).
110
+ - Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via an active-path `WeakSet`. Ancestor cycles render `"[Circular]"` at the back-reference instead of throwing; shared diamond references on separate branches stay structured (they are not collapsed to `"[Circular]"`). String object keys and `Map` keys are redacted like values, with deterministic `__2`/`__3`/… suffixes on collisions. `Date` and `RegExp` values are passed through unchanged; `ArrayBuffer` and typed arrays are passed through unchanged; `Map` is normalized to a plain object and `Set` to an array so the output stays JSON-compatible. `errorToErrorInfo()` tolerates a cyclic `error.cause` (rendered via `String()`). Symbols and non-enumerable properties remain outside JSON-shaped redaction.
111
111
  - Use placeholders in tests and docs. Never commit real tokens.
112
112
  - Live provider/worker tests are gated behind explicit environment variables and skipped by default: `PRISM_LIVE_PROVIDER_TESTS`, `PRISM_LIVE_COMPACTION_TESTS`, `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS`. Default `npm test` is network-free; do not add ungated network calls to default tests.
113
113
  - Credentials are not eagerly resolved by the core runtime, serialized into provider requests/events/stores, or passed to loops/compaction.
@@ -128,7 +128,7 @@ A branch leaf (`leaf_entry_id`) is the current entry id for that branch. Rebuild
128
128
  | --- | --- |
129
129
  | `prism_session_entries` | `id` PK, `session_id` FK, `parent_id`, `run_id`, `timestamp`, `kind`, `schema_version`, `message` JSONB, `event` JSONB, `model` JSONB, `previous_model` JSONB, `label`, `summary`, `data` JSONB, `metadata` JSONB |
130
130
 
131
- Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version` defaults to `1`. `parent_id` may be null for the root entry of a session.
131
+ Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version` is nullable in schema version 3 for existing entry compatibility; hosts writing new rows should persist the entry's current schema version. `parent_id` may be null for the root entry of a session.
132
132
 
133
133
  `SessionAppendOptions.idempotencyKey` is not part of `SessionEntry`; store it in an adapter-owned side table when you need durable retry detection:
134
134
 
@@ -214,7 +214,8 @@ await runRunLedgerConformance(() => createLedger(testDatabase), { exerciseReopen
214
214
  | Primitive | Purpose |
215
215
  | --- | --- |
216
216
  | `PersistenceSchemaModel` | Versioned table/column/index model covering sessions, entries, parent chain, idempotency side table, runs, events, tool calls, usage, tenant columns, and `prism_migrations` |
217
- | `createPersistenceMigrationContract()` | Strictly increasing migration steps, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
217
+ | `createPersistenceMigrationContract()` | Strictly increasing checked-in migration steps with deterministic SHA-256 checksums, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
218
+ | `assertAppliedPersistenceMigrations()` / `assertPersistenceSchemaShape()` | Testable fail-closed history and normalized SQLite/PostgreSQL catalog checks for every required table, column/type/null/default, PK/unique/FK, and named index. |
218
219
  | `getPersistencePaginationCursors()` | Indexed `(session_id, timestamp, id)`, `(run_id, sequence)`, `(run_id, recorded_at, id)` cursor shapes that avoid offset scans |
219
220
  | `assertPersistenceQueryPaginationConforms()` | Generic cursor pagination fixture for `queryEntries` |
220
221
  | `assertTenantScopedQueryIsolation()` | Tenant-filtered reads must not leak rows or primary-id collisions across tenants |
@@ -345,7 +346,11 @@ Retention jobs should not run inside the agent/session runtime. They are a host
345
346
 
346
347
  ## Migrations
347
348
 
348
- Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library. Recommended migration practices:
349
+ Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library. First-party SQLite/PostgreSQL adapters automatically verify their checked-in schema-v3 history and catalog at open, before runtime writes. Their catalog reads are bounded metadata queries/PRAGMAs, not table-data scans.
350
+
351
+ Each new adapter-owned migration row records the contract SHA-256 checksum. A complete known v0.0.5 history whose checksum values are all `NULL` is a one-time compatibility case: under the SQLite transaction or PostgreSQL advisory transaction lock, the adapter verifies the full current shape, backfills all checksums, and continues. Unknown/duplicate/out-of-order/name-version/checksum mismatch, mixed/partial legacy values, or any missing/renamed/wrong-type/null/default/key/index artifact fails closed. Restore or apply a reviewed host migration; never edit checksums to silence drift.
352
+
353
+ Recommended migration practices:
349
354
 
350
355
  - Use a sequential or timestamped migration naming convention.
351
356
  - Store applied migrations in `prism_migrations` with `name`, `version`, `applied_at`, `applied_by`, and `checksum`.
@@ -0,0 +1,75 @@
1
+ # Guardrails
2
+
3
+ ## What it does
4
+
5
+ Guardrails are typed, fail-closed checks at input, completed provider output, tool input, and raw tool output boundaries. `session.run()` evaluates configured stages through one core runner; `dispatchToolCall()` uses same runner for direct, MCP-server, and workflow tool calls.
6
+
7
+ ## When to use it
8
+
9
+ Use guardrails to block unsafe prompts, model responses, tool arguments, or tool results before their next boundary. Use a redactor for known secrets. Do not treat guardrails as a sandbox, secret detector, permission policy, or validation replacement.
10
+
11
+ ## Inputs / request
12
+
13
+ ```ts
14
+ import type { Guardrail, Guardrails } from "@arnilo/prism";
15
+
16
+ const pii: Guardrail<"input"> = {
17
+ name: "pii",
18
+ stage: "input",
19
+ evaluate: ({ value }) => JSON.stringify(value).includes("SSN")
20
+ ? { action: "tripwire", reason: "pii" }
21
+ : { action: "allow" },
22
+ };
23
+
24
+ const guardrails: Guardrails = { input: [pii], maxConcurrency: 1 };
25
+ ```
26
+
27
+ Set `AgentConfig.guardrails` for every session run or `RunOptions.guardrails` to append checks for one run. `DispatchToolCallOptions.guardrails`, workflow `RunWorkflowOptions.guardrails`, and MCP server `CreatePrismMcpServerOptions.guardrails` apply tool stages to direct calls. A stage has `Guardrail<"input" | "output" | "tool_input" | "tool_output">`, a name, optional revision, and `evaluate(context)` result.
28
+
29
+ Decisions are `allow`, `block`, `tripwire`, or `interrupt`. Evaluation defaults to declaration-order sequential. `maxConcurrency` may be 1–16; records are emitted in declaration order. Thrown or malformed decisions become a fail-closed tripwire. Decision reasons are capped at 4 KiB and metadata at 16 KiB after JSON normalization and optional redaction.
30
+
31
+ ## Outputs / response / events
32
+
33
+ Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
34
+
35
+ Ordering is fixed:
36
+
37
+ 1. input before session append, compaction, or provider work;
38
+ 2. provider output is privately collected, then output checks run before any assistant message event or persistence;
39
+ 3. tool input runs after tool-call middleware normalization and before lookup, permission, validation, execution policy, and side effect;
40
+ 4. tool output runs after the side effect but before redaction, tool events, ledger rows, transcript append, or next turn.
41
+
42
+ With no output guardrails, provider streaming retains existing behavior. With output guardrails, message events are buffered until the completed provider turn is allowed.
43
+
44
+ ## Request/response example
45
+
46
+ ```json
47
+ {
48
+ "event": {
49
+ "type": "guardrail_decision",
50
+ "record": { "guardrail": "pii", "stage": "input", "action": "tripwire", "reason": "pii" }
51
+ }
52
+ }
53
+ ```
54
+
55
+ ## Implementation example
56
+
57
+ ```ts
58
+ const agent = createAgent({ model, provider, guardrails: { input: [pii], output: [responseGuard] } });
59
+ await agent.createSession().run("Draft reply", { guardrails: { toolInput: [commandGuard] } });
60
+ ```
61
+
62
+ ## Extension and configuration notes
63
+
64
+ Guardrails are callbacks supplied by the host. Prism does not discover, load, retry, or persist callback code. `createSecureAgent()` keeps configured guardrails and only appends run-level checks; it never lets a run remove secure defaults. Custom loops receive guarded `LoopContext.generate()` and `LoopContext.dispatchToolCall()`; host code that directly calls a provider or `ToolDefinition.execute()` is outside the runtime boundary.
65
+
66
+ ## Security and performance notes
67
+
68
+ Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work.
69
+
70
+ ## Related APIs
71
+
72
+ - [Agent/session runtime](agent-session-runtime.md)
73
+ - [Tools](tools.md)
74
+ - [Agent events](agent-events.md)
75
+ - [Host security](host-security.md)