@arnilo/prism 0.0.5 → 0.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -1
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +27 -16
- package/dist/agent-run-lifecycle.d.ts +28 -0
- package/dist/agent-run-lifecycle.js +33 -0
- package/dist/agent-run-state.d.ts +53 -0
- package/dist/agent-run-state.js +127 -0
- package/dist/agents.d.ts +3 -1
- package/dist/agents.js +337 -46
- package/dist/contracts.d.ts +205 -3
- package/dist/contracts.js +4 -0
- package/dist/guardrails.d.ts +25 -0
- package/dist/guardrails.js +133 -0
- package/dist/ids.d.ts +2 -0
- package/dist/ids.js +6 -0
- package/dist/index.d.ts +17 -3
- package/dist/index.js +10 -3
- package/dist/input.js +2 -0
- package/dist/resources.js +2 -1
- package/dist/run-limits.d.ts +34 -0
- package/dist/run-limits.js +163 -0
- package/dist/secure-agent.d.ts +3 -0
- package/dist/secure-agent.js +63 -0
- package/dist/session-stores.js +2 -3
- package/dist/testing/persistence-schema.d.ts +45 -7
- package/dist/testing/persistence-schema.js +138 -24
- package/dist/thinking.d.ts +42 -0
- package/dist/thinking.js +92 -0
- package/dist/tools.d.ts +10 -2
- package/dist/tools.js +56 -7
- package/dist/use-case-model.d.ts +63 -0
- package/dist/use-case-model.js +52 -0
- package/docs/a2a.md +4 -2
- package/docs/agent-events.md +23 -16
- package/docs/agent-loops.md +19 -8
- package/docs/agent-session-runtime.md +33 -1
- package/docs/coding-agent-tools.md +33 -12
- package/docs/coding-security.md +2 -2
- package/docs/compaction-llm.md +17 -7
- package/docs/compaction-observational-memory.md +28 -4
- package/docs/credential-storage.md +58 -9
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/database-persistence.md +8 -3
- package/docs/guardrails.md +75 -0
- package/docs/host-security.md +16 -8
- package/docs/index.md +26 -22
- package/docs/mcp-tools.md +32 -12
- package/docs/migration.md +164 -2
- package/docs/node-filesystem-config.md +1 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/postgres-persistence.md +3 -3
- package/docs/provider-caching.md +16 -4
- package/docs/provider-conformance.md +39 -1
- package/docs/provider-packages.md +60 -3
- package/docs/providers/ai-sdk.md +36 -0
- package/docs/providers/kimi.md +124 -61
- package/docs/providers/neuralwatt.md +19 -13
- package/docs/providers/openai.md +56 -13
- package/docs/providers/opencode-go.md +118 -30
- package/docs/providers/openrouter.md +105 -35
- package/docs/providers/zai.md +94 -45
- package/docs/release-and-install.md +47 -49
- package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
- package/docs/runs-and-usage.md +30 -3
- package/docs/server.md +5 -2
- package/docs/sqlite-persistence.md +2 -2
- package/docs/structured-output.md +1 -1
- package/docs/thinking-and-reasoning.md +98 -0
- package/docs/tool-execution-primitives.md +3 -3
- package/docs/tools.md +21 -1
- package/docs/use-case-model-selection.md +109 -0
- package/docs/workflow-orchestration-primitives.md +1 -0
- package/docs/workflows.md +18 -10
- package/docs/working-and-semantic-memory.md +1 -0
- package/package.json +2 -2
|
@@ -15,6 +15,8 @@
|
|
|
15
15
|
| `createAllTools(cwd, options?)` | Every tool the package provides (currently identical to `createCodingTools`). |
|
|
16
16
|
| `detectSupportedImageMimeType(buf)` / `detectSupportedImageMimeTypeFromFile(path)` | Magic-byte image MIME detection (PNG/JPEG/GIF/WebP/BMP) used by `read`. |
|
|
17
17
|
| `DEFAULT_MAX_IMAGE_BYTES` | Default `read` image size ceiling (10 MB). |
|
|
18
|
+
| `DEFAULT_*` / `HARD_*` coding limit constants | Published text-scan, image, write/edit, shell timeout, display, and total-output ceilings. |
|
|
19
|
+
| `ReadTextOptions` / `ReadTextResult` | Bounded text-page contract required by custom `ReadOperations`. |
|
|
18
20
|
| `TransformImage` / `TransformImageInput` | Types for the optional `read` `transformImage` callback. |
|
|
19
21
|
| `withFileMutationQueue(path, fn)` | Per-path serialization primitive re-exported for hosts. |
|
|
20
22
|
|
|
@@ -65,7 +67,7 @@ Run a shell command and return combined stdout+stderr.
|
|
|
65
67
|
| Field | Type | Purpose |
|
|
66
68
|
| --- | --- | --- |
|
|
67
69
|
| `command` | `string` | Shell command to execute (required). |
|
|
68
|
-
| `timeout` | `number` | Timeout in **seconds** (optional;
|
|
70
|
+
| `timeout` | `number` | Timeout in **seconds** (optional; defaults to 600, hard maximum 3600). |
|
|
69
71
|
|
|
70
72
|
**Outputs:** a `ToolResult` whose `content[0]` is a `TextContent` with the combined output. Non-zero exit is **not** a tool error: it is returned as a normal result with `[Command exited with code N]` appended to the content and `exitCode` in metadata. Timeout and abort are error results that still carry the partial output captured so far.
|
|
71
73
|
|
|
@@ -75,9 +77,11 @@ Run a shell command and return combined stdout+stderr.
|
|
|
75
77
|
| --- | --- | --- |
|
|
76
78
|
| `exitCode` | always | Process exit code, or `null` when the process was killed by timeout/abort. |
|
|
77
79
|
| `truncation` | always | `TruncationResult` from the bounded output accumulator. |
|
|
78
|
-
| `fullOutputPath?` | truncated only |
|
|
80
|
+
| `fullOutputPath?` | successful and truncated only | Host-owned path to retained output. Failed/aborted/timed-out/output-limited calls remove unpublished spills. |
|
|
81
|
+
| `totalOutputBytes` | shell executed | Raw bytes retained, never above `maxTotalOutputBytes`. |
|
|
82
|
+
| `outputLimitExceeded` / `outputStorageFailed` | shell executed | Attributable resource failure flags. |
|
|
79
83
|
|
|
80
|
-
Shell resolution honors `options.shellPath` → `SHELL` env → `/bin/bash` → `sh
|
|
84
|
+
Shell resolution honors `options.shellPath` → `SHELL` env → `/bin/bash` → `sh`. The process group is killed on timeout, caller abort, spill failure, or total-output overflow (`process.kill(-pid)` on Unix, `taskkill /F /T` on Windows). Combined output defaults to a 64 MiB total cap (1 GiB hard cap). Spill files use random exclusive creation and Unix mode `0600`; hosts own and must delete a successful result's `fullOutputPath` after consumption.
|
|
81
85
|
|
|
82
86
|
### `read`
|
|
83
87
|
|
|
@@ -91,7 +95,7 @@ Read a text or image file.
|
|
|
91
95
|
| `offset` | `number` | Line to start reading from (1-indexed). |
|
|
92
96
|
| `limit` | `number` | Maximum number of lines to read. |
|
|
93
97
|
|
|
94
|
-
**Outputs:** text files
|
|
98
|
+
**Outputs:** text files are scanned incrementally until one requested page, `maxLines`/`maxBytes`, EOF, or `maxScanBytes` (default 64 MiB scanned per call; 1 GiB hard cap). The default path never loads the complete file and returns a `Use offset=N to continue` footer when more remains. Exact total line count is reported only when EOF was already reached in the bounded scan. Image files (PNG/JPEG/GIF/WebP/BMP by **magic bytes**, not extension) become `[TextContent note, ImageContent]` with base64 `data` and `mimeType`. Oversize images are rejected by `stat` (when available) or `buffer.length` against `maxImageBytes` (default 10 MB) before base64 encoding. An optional `transformImage` callback lets hosts resize or re-encode images without adding image-processing dependencies to the base package. Read failures (missing file, offset beyond end, oversize image, abort) are error results.
|
|
95
99
|
|
|
96
100
|
`read` tool options (via `createReadTool(cwd, options)` or `ToolsOptions.read`):
|
|
97
101
|
|
|
@@ -100,8 +104,9 @@ Read a text or image file.
|
|
|
100
104
|
| `maxImageBytes` | `DEFAULT_MAX_IMAGE_BYTES` (10 MB) | Reject image reads larger than this many bytes. |
|
|
101
105
|
| `transformImage` | — | Host callback `( { buffer, mimeType } ) => Promise<Buffer>` run after read, before base64. |
|
|
102
106
|
| `autoResizeImages` | — | **Deprecated.** Ignored unless `transformImage` is also set (use `transformImage` instead). |
|
|
103
|
-
| `maxLines` / `maxBytes` | 2000 / 50
|
|
104
|
-
| `
|
|
107
|
+
| `maxLines` / `maxBytes` | 2000 / 50 KiB | Text page display limits (hard: 100,000 / 1 MiB). |
|
|
108
|
+
| `maxScanBytes` | 64 MiB | Raw bytes scanned to reach one page (hard: 1 GiB). |
|
|
109
|
+
| `operations` | local fs | Pluggable bounded `ReadOperations` backend. |
|
|
105
110
|
| `executionPolicy` | — | Structured pre-execution policy (see [Coding security](coding-security.md)). |
|
|
106
111
|
|
|
107
112
|
```ts
|
|
@@ -133,7 +138,7 @@ Create or overwrite a file, creating parent directories as needed.
|
|
|
133
138
|
| `path` | `string` | Path to the file to write (relative or absolute). Required. |
|
|
134
139
|
| `content` | `string` | Content to write (empty string creates an empty file). Required. |
|
|
135
140
|
|
|
136
|
-
**Outputs:** a `TextContent` confirmation naming the **absolute path** with UTF-8 byte and line counts (e.g. `Successfully wrote 42 bytes (3 lines) to /abs/path.txt`). Write failures and abort are error results. Empty `content` is valid.
|
|
141
|
+
**Outputs:** a `TextContent` confirmation naming the **absolute path** with UTF-8 byte and line counts (e.g. `Successfully wrote 42 bytes (3 lines) to /abs/path.txt`). `maxInputBytes` defaults to 8 MiB (64 MiB hard cap); oversized UTF-8 input fails before policy evaluation, directory creation, or write. Write failures and abort are error results. Empty `content` is valid.
|
|
137
142
|
|
|
138
143
|
`write` result `metadata`: `{ bytes, lines, path }` (absolute path). Concurrent writes to the same path serialize through `withFileMutationQueue`; writes to different paths run in parallel.
|
|
139
144
|
|
|
@@ -148,7 +153,7 @@ Precise text replacement in an existing file via exact-then-fuzzy matching.
|
|
|
148
153
|
| `path` | `string` | Path to the file to edit. Required. |
|
|
149
154
|
| `edits` | `Array<{ oldText: string, newText: string }>` | Targeted replacements, each matched against the **original** file (not incrementally). No overlapping/nested edits. Required, non-empty. |
|
|
150
155
|
|
|
151
|
-
Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse). A BOM is stripped before matching and re-prepended on write; original line endings are restored.
|
|
156
|
+
Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse). A BOM is stripped before matching and re-prepended on write; original line endings are restored. Defaults reject targets over 8 MiB, aggregate old/new UTF-8 input over 2 MiB, or more than 100 edits (hard caps: 64 MiB, 16 MiB, and 1,000). Stat and bounded read checks run before matching or mutation.
|
|
152
157
|
|
|
153
158
|
**Outputs:** a `TextContent` confirmation (`Successfully replaced N block(s) in {path}.`) plus `metadata`. Any failure — missing/unreadable file, no match, duplicate (non-unique) match, overlap, empty `oldText`, no-op edit, or abort — is an error result, and the file is left **unchanged** (the match runs before the write).
|
|
154
159
|
|
|
@@ -156,7 +161,7 @@ Each `edits[].oldText` must match a unique, non-overlapping region of the origin
|
|
|
156
161
|
|
|
157
162
|
## Outputs / response / events
|
|
158
163
|
|
|
159
|
-
Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`.
|
|
164
|
+
Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`. `write` and `edit` serialize per realpath through `withFileMutationQueue` so concurrent calls targeting one file do not interleave. `shell` is marked `exclusive`; tool dispatch serializes it at the turn level. The package emits no events of its own; hosts observe tool execution through the normal Prism `AgentEvent` stream via `dispatchToolCall`.
|
|
160
165
|
|
|
161
166
|
## Request/response example
|
|
162
167
|
|
|
@@ -208,6 +213,8 @@ const shell = createShellTool("/repo", {
|
|
|
208
213
|
shellPath: "/bin/bash",
|
|
209
214
|
commandPrefix: "set -euo pipefail",
|
|
210
215
|
maxLines: 500,
|
|
216
|
+
timeout: 600,
|
|
217
|
+
maxTotalOutputBytes: 64 * 1024 * 1024,
|
|
211
218
|
});
|
|
212
219
|
|
|
213
220
|
const remoteWrite = createWriteTool("/repo", {
|
|
@@ -220,8 +227,8 @@ const remoteWrite = createWriteTool("/repo", {
|
|
|
220
227
|
|
|
221
228
|
## Extension and configuration notes
|
|
222
229
|
|
|
223
|
-
- **Pluggable operation backends.** Every tool accepts an `operations` seam
|
|
224
|
-
- **Per-tool options.** `ShellToolOptions`
|
|
230
|
+
- **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
|
|
231
|
+
- **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`. Invalid/non-finite/unsafe/above-hard-cap values throw during tool construction; request `timeout` errors before spawn.
|
|
225
232
|
- **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override.
|
|
226
233
|
- **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
|
|
227
234
|
- No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
|
|
@@ -230,10 +237,24 @@ const remoteWrite = createWriteTool("/repo", {
|
|
|
230
237
|
|
|
231
238
|
- **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
|
|
232
239
|
- **Non-zero exit is not an error.** A failing command is a normal `shell` result (exit code in metadata); only timeout/abort/spawn failures are error results. Do not assume `error == undefined` means the command succeeded.
|
|
233
|
-
- **Bounded
|
|
240
|
+
- **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
|
|
234
241
|
- **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
|
|
235
242
|
- **Bounded image reads.** `read` rejects images over `maxImageBytes` (default 10 MB) by `stat` before read when possible; MIME is detected from magic bytes only. Optional `transformImage` is host-owned — the base package has no image-processing dependency.
|
|
236
243
|
|
|
244
|
+
### Resource-limit defaults and hard caps
|
|
245
|
+
|
|
246
|
+
| Boundary | Default | Hard cap | Failure point |
|
|
247
|
+
| --- | ---: | ---: | --- |
|
|
248
|
+
| Display lines / bytes | 2,000 / 50 KiB | 100,000 / 1 MiB | tool construction |
|
|
249
|
+
| Text scan per read | 64 MiB | 1 GiB | bounded scan before more input is retained |
|
|
250
|
+
| Image | 10,000,000 bytes | 32 MiB | stat and bounded read before base64/transform result use |
|
|
251
|
+
| Write UTF-8 input | 8 MiB | 64 MiB | before policy/filesystem mutation |
|
|
252
|
+
| Edit target / input / count | 8 MiB / 2 MiB / 100 | 64 MiB / 16 MiB / 1,000 | before target read/matching/write |
|
|
253
|
+
| Shell wall time | 600 seconds | 3,600 seconds | process-tree kill |
|
|
254
|
+
| Shell total stdout+stderr | 64 MiB | 1 GiB | process-tree kill; spill removal |
|
|
255
|
+
|
|
256
|
+
Every configurable value is a positive safe integer; Prism rejects rather than clamps invalid values. Limits control resources, not authority: they do not replace root containment, approval, validation, or a sandbox.
|
|
257
|
+
|
|
237
258
|
## Related APIs
|
|
238
259
|
|
|
239
260
|
- [Tools](tools.md): the host-owned tool harness — `createToolRegistry`, `dispatchToolCall`, filtering, and the `ToolDefinition` contract these factories satisfy.
|
package/docs/coding-security.md
CHANGED
|
@@ -74,11 +74,11 @@ const tools = createCodingTools(workspaceRoot, {
|
|
|
74
74
|
|
|
75
75
|
Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers.
|
|
76
76
|
|
|
77
|
-
Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
|
|
77
|
+
Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
|
|
78
78
|
|
|
79
79
|
## Security and performance notes
|
|
80
80
|
|
|
81
|
-
Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
|
|
81
|
+
Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
|
|
82
82
|
|
|
83
83
|
## Related APIs
|
|
84
84
|
|
package/docs/compaction-llm.md
CHANGED
|
@@ -25,15 +25,16 @@ Key exports:
|
|
|
25
25
|
| Field | Purpose |
|
|
26
26
|
| --- | --- |
|
|
27
27
|
| `provider` / `summaryProvider` | Explicit `AIProvider`, or factory receiving a resolved credential. |
|
|
28
|
-
| `model` / `summaryModel` |
|
|
28
|
+
| `model` / `summaryModel` | Summary `ModelConfig`. `summaryModel` wins; `model` is the session/host fallback via `resolveUseCaseModel`. See [Use-case model selection](use-case-model-selection.md). |
|
|
29
29
|
| `credential`, `credentialRequest` | Optional per-call credential resolution for provider factories. |
|
|
30
30
|
| `providerOptions` | Generic `ProviderRequest.options`, including cache fields. |
|
|
31
31
|
| `providerRequestPolicies` | Optional Prism provider request policies applied before the summary call. |
|
|
32
32
|
| `customInstructions` | Additional summary focus appended to prompts. |
|
|
33
|
-
| `thinkingLevel` |
|
|
34
|
-
| `reserveTokens` | Output budget basis; defaults to `16384`. |
|
|
33
|
+
| `thinkingLevel` | Mapped into `ProviderRequest.options.compat` via `applyThinkingLevel` / `thinkingFamilyForModel` (not inert `extra.thinkingLevel`). See [Thinking and reasoning](thinking-and-reasoning.md). |
|
|
34
|
+
| `reserveTokens` | Output budget basis; defaults to `16384`, hard cap `131072`. |
|
|
35
35
|
| `keepRecentTokens` | Approximate recent-token budget; defaults to `20000`. |
|
|
36
|
-
| `maxSummaryTokens` / `maxOutputTokens` |
|
|
36
|
+
| `maxSummaryTokens` / `maxOutputTokens` | Summary retention/request ceiling; default `16384`, hard cap `131072`. `maxSummaryTokens` wins over the compatibility alias. The finite value is written to `model.parameters.maxTokens`; first-party providers map it to their wire field. |
|
|
37
|
+
| `maxErrorBytes` | Retained provider/factory/policy error detail; default `1024`, hard cap `8192`, UTF-8-safe and known-secret redacted. |
|
|
37
38
|
| `maxToolResultChars` | Tool-result JSON truncation limit; defaults to `2000`. |
|
|
38
39
|
| `trackFileOperations`, `includeFileOperations` | Control file path extraction and final summary blocks. |
|
|
39
40
|
| `secrets` | Exact strings to redact from serialized prompts and final summaries. |
|
|
@@ -64,7 +65,8 @@ const strategy = createLlmCompactionStrategy({
|
|
|
64
65
|
model: { provider: "openai", model: "gpt-4.1-mini" },
|
|
65
66
|
keepRecentTokens: 20_000,
|
|
66
67
|
reserveTokens: 16_384,
|
|
67
|
-
|
|
68
|
+
maxSummaryTokens: 800,
|
|
69
|
+
maxErrorBytes: 1_024,
|
|
68
70
|
providerOptions: { cacheRetention: "short" },
|
|
69
71
|
customInstructions: "Focus on current files and failing tests.",
|
|
70
72
|
});
|
|
@@ -102,10 +104,18 @@ const agent = createAgent({ model, provider, compaction: { strategy, thresholdEn
|
|
|
102
104
|
Registration only contributes an inert strategy. The host must resolve and pass it to runtime config.
|
|
103
105
|
|
|
104
106
|
## Security and performance notes
|
|
105
|
-
Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization.
|
|
107
|
+
Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization. Limit options must be positive safe integers at or below their hard caps and reject during strategy creation. Missing output options use a 16,384-token summary ceiling; reserve ratio/model metadata may narrow the provider request, never remove its finite `maxTokens`. A request policy that replaces `maxTokens` with NaN, Infinity, zero, an unsafe integer, or above-hard-cap input fails before provider generation.
|
|
108
|
+
|
|
109
|
+
Provider deltas are redacted while retained and stop at `maxSummaryTokens * 4` UTF-16 code units without splitting a surrogate pair. Provider iteration is closed/aborted on overflow. A derived finite event ceiling also stops endless empty/non-text deltas. Final history/turn/file composition receives the same cap. Provider error events, generator throws, provider-factory failures, and policy failures expose only bounded redacted detail; host abort remains authoritative.
|
|
110
|
+
|
|
111
|
+
The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. It does not discover credentials, read files, start background jobs, or add provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history or provider output.
|
|
106
112
|
|
|
107
113
|
## Related APIs
|
|
108
|
-
|
|
114
|
+
|
|
115
|
+
- [Use-case model selection](use-case-model-selection.md): `summaryModel` vs session `model` fallback.
|
|
116
|
+
- [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → `compat`.
|
|
117
|
+
- [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary and core compaction strategy surface.
|
|
118
|
+
- [Observational memory compaction package](compaction-observational-memory.md): source-backed memory workers with the same use-case binding pattern.
|
|
109
119
|
- [Agent/session runtime](agent-session-runtime.md): `AgentSession.compact()` and opt-in auto-compaction.
|
|
110
120
|
- [Provider layer](provider-layer.md): mock providers and provider request contracts.
|
|
111
121
|
- [Credentials and redaction](credentials-and-redaction.md): exact known-secret redaction behavior.
|
|
@@ -27,6 +27,20 @@ Memory records use `SessionEntry.kind: "custom"` with `entry.data.type` markers:
|
|
|
27
27
|
|
|
28
28
|
Ids are known, source-backed 12-character lowercase hex strings matching `^[a-f0-9]{12}$`.
|
|
29
29
|
|
|
30
|
+
Worker limits are finite positive safe integers:
|
|
31
|
+
|
|
32
|
+
| Runtime option | Default | Hard cap | Scope |
|
|
33
|
+
| --- | ---: | ---: | --- |
|
|
34
|
+
| `maxWorkerTurns` | 16 | 64 | Provider turns per observer/reflector/dropper run; overrides settings `agentMaxTurns` |
|
|
35
|
+
| `maxWorkerToolCallsPerTurn` | 32 | 256 | Calls retained from one provider response |
|
|
36
|
+
| `maxWorkerToolCalls` | 128 | 1,024 | Calls across all turns in one worker run |
|
|
37
|
+
| `maxWorkerArgumentBytes` | 64 KiB | 1 MiB | Each raw and redacted JSON argument object |
|
|
38
|
+
| `maxWorkerResultBytes` | 64 KiB | 1 MiB | Full tool result and replayed value/error payload |
|
|
39
|
+
| `maxWorkerMessageBytes` | 1 MiB | 8 MiB | System/prompt plus assistant-call/tool-result transcript |
|
|
40
|
+
| `maxWorkerErrorBytes` | 1 KiB | 8 KiB | Provider/tool/runtime error text after exact known-secret redaction |
|
|
41
|
+
|
|
42
|
+
Direct `runObserver()` / `runReflector()` / `runDropper()` calls retain required `maxTurns` and accept the corresponding shorter worker fields (`maxToolCalls`, `maxResultBytes`, etc.). Named default/hard constants and `resolveMemoryWorkerLimits()` are exported.
|
|
43
|
+
|
|
30
44
|
## Outputs / response / events
|
|
31
45
|
|
|
32
46
|
Key exports:
|
|
@@ -78,7 +92,12 @@ const memory = createObservationalMemoryRuntime({
|
|
|
78
92
|
session,
|
|
79
93
|
appendEntry: (entry) => store.append(entry),
|
|
80
94
|
workerProvider,
|
|
81
|
-
|
|
95
|
+
sessionModel: agent.config.model, // fallback when workerModel unset
|
|
96
|
+
// workerModel: { provider: "mock", model: "memory" }, // optional override
|
|
97
|
+
maxWorkerTurns: 8,
|
|
98
|
+
maxWorkerToolCalls: 64,
|
|
99
|
+
maxWorkerResultBytes: 64 * 1024,
|
|
100
|
+
overrides: { thinkingLevel: "low" },
|
|
82
101
|
});
|
|
83
102
|
await memory.flush();
|
|
84
103
|
await session.compact({ strategy: createObservationalMemoryCompactionStrategy({ keepRecentEntries: 8 }) });
|
|
@@ -92,9 +111,9 @@ await kernel.load([createObservationalMemoryExtension({ recallTool: { getEntries
|
|
|
92
111
|
|
|
93
112
|
## Extension and configuration notes
|
|
94
113
|
|
|
95
|
-
Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`.
|
|
114
|
+
Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`. `agentMaxTurns` now rejects non-integer/non-finite/out-of-range input (hard 64) instead of flooring/falling back. Runtime `maxWorkerTurns` takes precedence.
|
|
96
115
|
|
|
97
|
-
The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, `workerProvider
|
|
116
|
+
The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, and `workerProvider`. Model selection uses [use-case model selection](use-case-model-selection.md): pass optional `workerModel` (or settings `workerModel`) to override, and `sessionModel: agent.config.model` so workers fall back to the session model when no worker model is configured. `requireExplicitModel: true` restores the historical `missing_model` skip when no explicit worker model is set. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution. Default credential requests use the **resolved** model's provider id.
|
|
98
117
|
|
|
99
118
|
`createObservationalMemoryCompactionStrategy()` keeps recent message entries like the default compaction strategy, renders existing observations/reflections as the summary, and returns a standard Prism compaction entry. Its `data` includes `throughEntryId`, `keepEntryIds`, `strategy`, `trigger`, and `memory: { type: "om.folded", version: 1, fullFold, observations, reflections, droppedObservationIds }`. When active observations exceed `observationsPoolMaxTokens`, it performs a full fold into `data.memory`.
|
|
100
119
|
|
|
@@ -110,13 +129,18 @@ The runtime requires host-supplied `session`, an `appendEntry` callback bound to
|
|
|
110
129
|
- Recall tool and commands only see current-branch entries supplied by the host callback.
|
|
111
130
|
- Invalid or missing ids fail closed; invalid recall tool ids skip entry lookup.
|
|
112
131
|
- Utilities and fast compaction are O(n) over supplied entries and use no provider, network, filesystem, timer, worker, credential, or settings access.
|
|
113
|
-
- Workers serialize only supplied branch entries
|
|
132
|
+
- Workers serialize only supplied branch entries within `maxWorkerMessageBytes`, enforce finite turns/calls/arguments/results/messages/errors, and run one consolidation pipeline at a time per runtime. Source serialization and reflection/drop prompts fail before joining beyond the transcript cap.
|
|
133
|
+
- Every provider call must name a registered worker tool. Unknown calls, call overflow, oversized/deep/cyclic/non-JSON arguments/results, and transcript overflow fail deterministically; no excess call enters the assistant transcript or executes.
|
|
134
|
+
- Raw arguments are measured before tool execution. Full results are measured before redaction/replay; the bounded redacted value/error is then measured again because replacement text can grow. Replayed call arguments, tool values/errors, runtime `lastError`, and debug error data contain exact known-secret redaction. Host tools may already have caused side effects before returning an invalid oversized result; keep worker tools small/idempotent.
|
|
135
|
+
- Worker transcripts replay assistant `tool_call` messages before matching role `tool` `tool_result` messages so provider requests stay valid for call/result-pairing providers. Calls produced on the final allowed turn execute and persist, but no additional provider turn starts.
|
|
114
136
|
- Compaction preserves raw history; Prism appends one standard compaction entry and rebuilds provider context from its summary plus kept recent messages.
|
|
115
137
|
- Pass known secrets to render/recall/runtime/tool/command helpers to redact exact values from prompts, records, structured results, and text output.
|
|
116
138
|
- Live tests are opt-in with `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS=1`.
|
|
117
139
|
|
|
118
140
|
## Related APIs
|
|
119
141
|
|
|
142
|
+
- [Use-case model selection](use-case-model-selection.md): session vs worker model binding and `resolveUseCaseModel`.
|
|
143
|
+
- [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → provider `compat`.
|
|
120
144
|
- [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary.
|
|
121
145
|
- [LLM compaction package](compaction-llm.md): existing optional compaction-package pattern.
|
|
122
146
|
- [Session stores and branching](session-stores-and-branching.md): branch entries that observational memory reads and appends to.
|
|
@@ -44,8 +44,11 @@ import {
|
|
|
44
44
|
| --- | --- | --- |
|
|
45
45
|
| `path` | `string` | Vault file path. Parent directories are created as needed. |
|
|
46
46
|
| `getPassphrase` | `() => string \| Promise<string>` | Host-owned passphrase retrieval. Never logged by the adapter. |
|
|
47
|
-
| `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32
|
|
48
|
-
| `fileMode` | `number` | Unix mode for
|
|
47
|
+
| `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32`; limits are listed below. |
|
|
48
|
+
| `fileMode` | `number` | Unix mode for files. Defaults to `0o600`; group/other permissions are rejected. |
|
|
49
|
+
| `limits.maxFileBytes` | `number` | Encrypted envelope file: 4 MiB default, 16 MiB hard cap. |
|
|
50
|
+
| `limits.maxVaultBytes` | `number` | Decrypted vault/plaintext: 3 MiB default, 12 MiB hard cap. |
|
|
51
|
+
| `limits.maxScryptMemoryBytes` | `number` | `128*N*r` memory estimate: 256 MiB default and hard cap. |
|
|
49
52
|
|
|
50
53
|
### System keychain
|
|
51
54
|
|
|
@@ -53,7 +56,8 @@ import {
|
|
|
53
56
|
| --- | --- | --- |
|
|
54
57
|
| `service` | `string` | Keychain service name (application identifier). |
|
|
55
58
|
| `namespace` | `string` | Optional prefix separating environments or tenants within one service. |
|
|
56
|
-
| `timeoutMs` | `number` | Operation timeout. Defaults to
|
|
59
|
+
| `timeoutMs` | `number` | Operation timeout. Defaults to 5,000 ms; hard cap 60,000 ms. |
|
|
60
|
+
| `maxPayloadBytes` | `number` | Decrypted keychain payload: 3 MiB default, 12 MiB hard cap. |
|
|
57
61
|
|
|
58
62
|
## Outputs / response / events
|
|
59
63
|
|
|
@@ -71,6 +75,8 @@ Encrypted file stores also expose:
|
|
|
71
75
|
- `reload()` — re-read and decrypt from disk
|
|
72
76
|
- `flush()` — force rewrite of the encrypted envelope
|
|
73
77
|
|
|
78
|
+
`encryptBytes()` and `decryptBytes()` are Promise-based because they use asynchronous `node:crypto.scrypt`.
|
|
79
|
+
|
|
74
80
|
Errors are explicit and fail closed:
|
|
75
81
|
|
|
76
82
|
| Error | Code | When |
|
|
@@ -123,6 +129,7 @@ import {
|
|
|
123
129
|
const store = await openEncryptedCredentialStore({
|
|
124
130
|
path: "./credentials.vault",
|
|
125
131
|
getPassphrase: () => process.env.MY_APP_CREDENTIAL_PASSPHRASE!,
|
|
132
|
+
limits: { maxFileBytes: 4 * 1024 * 1024, maxVaultBytes: 3 * 1024 * 1024 },
|
|
126
133
|
});
|
|
127
134
|
|
|
128
135
|
const resolver = createExplicitCredentialResolver([
|
|
@@ -151,6 +158,47 @@ await rotateEncryptedCredentialStorePassphrase({
|
|
|
151
158
|
});
|
|
152
159
|
```
|
|
153
160
|
|
|
161
|
+
### Desktop keychain and explicit overrides
|
|
162
|
+
|
|
163
|
+
```ts
|
|
164
|
+
import {
|
|
165
|
+
createEnvCredentialResolver,
|
|
166
|
+
createExplicitCredentialResolver,
|
|
167
|
+
createMemoryCredentialStore,
|
|
168
|
+
} from "@arnilo/prism";
|
|
169
|
+
import {
|
|
170
|
+
createKeychainCredentialStore,
|
|
171
|
+
createStoredCredentialResolver,
|
|
172
|
+
} from "@arnilo/prism-credentials-node";
|
|
173
|
+
import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
|
|
174
|
+
|
|
175
|
+
const keychain = createKeychainCredentialStore({
|
|
176
|
+
service: "com.example.my-app",
|
|
177
|
+
namespace: "production",
|
|
178
|
+
});
|
|
179
|
+
await keychain.set({
|
|
180
|
+
name: "apiKey",
|
|
181
|
+
provider: "openai",
|
|
182
|
+
credential: { type: "api_key", value: userSuppliedKey },
|
|
183
|
+
});
|
|
184
|
+
|
|
185
|
+
const runtimeOverrides = createMemoryCredentialStore();
|
|
186
|
+
// Set only for this user/agent instance; it wins over stored and env values.
|
|
187
|
+
runtimeOverrides.set({
|
|
188
|
+
name: "apiKey",
|
|
189
|
+
provider: "openai",
|
|
190
|
+
credential: { type: "api_key", value: temporaryOverride },
|
|
191
|
+
});
|
|
192
|
+
|
|
193
|
+
const apiKey = createExplicitCredentialResolver([
|
|
194
|
+
{ name: "runtime", resolver: runtimeOverrides },
|
|
195
|
+
{ name: "keychain", resolver: createStoredCredentialResolver(keychain) },
|
|
196
|
+
{ name: "env", resolver: createEnvCredentialResolver(process.env, { openai: "OPENAI_API_KEY" }) },
|
|
197
|
+
]);
|
|
198
|
+
|
|
199
|
+
const providers = createOpenAIProviderPackage({ apiKey });
|
|
200
|
+
```
|
|
201
|
+
|
|
154
202
|
## Extension and configuration notes
|
|
155
203
|
|
|
156
204
|
- Passphrase retrieval, TLS, and OS permission prompts remain host-owned.
|
|
@@ -161,12 +209,13 @@ await rotateEncryptedCredentialStorePassphrase({
|
|
|
161
209
|
|
|
162
210
|
## Security and performance notes
|
|
163
211
|
|
|
164
|
-
- Authenticated encryption uses Node built-in `aes-256-gcm` and `scrypt`; no extra crypto
|
|
165
|
-
-
|
|
166
|
-
-
|
|
167
|
-
-
|
|
168
|
-
- Keychain operations
|
|
169
|
-
-
|
|
212
|
+
- Authenticated encryption uses Node built-in `aes-256-gcm` and asynchronous `scrypt`; no extra crypto dependency is added for the file backend.
|
|
213
|
+
- Envelope parsing rejects unknown shape, non-canonical/oversized base64, wrong salt/IV/tag size, unsupported algorithms/version, and excessive KDF work before scrypt. `N` must be a power of two from 16,384–262,144; `r≤32`, `p≤16`, `keyLength=32`, `N*r*p≤2,097,152`, and `128*N*r` must fit `maxScryptMemoryBytes`.
|
|
214
|
+
- Existing Unix vaults are checked before content read and must deny group/other access. Atomic writes create a random exclusive temp file at the requested restrictive mode, then rename; Windows skips Unix mode checks.
|
|
215
|
+
- Derived keys and package-owned plaintext buffers are zeroed after use. JavaScript passphrase strings and returned credentials remain host-owned.
|
|
216
|
+
- Keychain operations use `@napi-rs/keyring`'s abort-aware `AsyncEntry`, so native work runs outside the JavaScript event loop. A main-loop timer aborts and rejects at `timeoutMs`; native cancellation remains OS/backend-dependent and may briefly retain one libuv worker after rejection.
|
|
217
|
+
- Keychain payloads are bytes rather than password strings and are zeroed after parse/write. Unknown native errors are mapped to sanitized typed errors; no native message or secret value is echoed.
|
|
218
|
+
- Never log passphrases, derived keys, or decrypted credential payloads.
|
|
170
219
|
- Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
|
|
171
220
|
|
|
172
221
|
## Related APIs
|
|
@@ -107,7 +107,7 @@ console.log(error.message);
|
|
|
107
107
|
- Redaction only removes exact known secret values passed to the helper. It is not a general-purpose secret detector.
|
|
108
108
|
- Do not pass empty strings as secrets; they are ignored.
|
|
109
109
|
- `redactSecrets()` recursively walks arrays and object entries, so avoid using it on huge objects unless needed.
|
|
110
|
-
- Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via
|
|
110
|
+
- Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via an active-path `WeakSet`. Ancestor cycles render `"[Circular]"` at the back-reference instead of throwing; shared diamond references on separate branches stay structured (they are not collapsed to `"[Circular]"`). String object keys and `Map` keys are redacted like values, with deterministic `__2`/`__3`/… suffixes on collisions. `Date` and `RegExp` values are passed through unchanged; `ArrayBuffer` and typed arrays are passed through unchanged; `Map` is normalized to a plain object and `Set` to an array so the output stays JSON-compatible. `errorToErrorInfo()` tolerates a cyclic `error.cause` (rendered via `String()`). Symbols and non-enumerable properties remain outside JSON-shaped redaction.
|
|
111
111
|
- Use placeholders in tests and docs. Never commit real tokens.
|
|
112
112
|
- Live provider/worker tests are gated behind explicit environment variables and skipped by default: `PRISM_LIVE_PROVIDER_TESTS`, `PRISM_LIVE_COMPACTION_TESTS`, `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS`. Default `npm test` is network-free; do not add ungated network calls to default tests.
|
|
113
113
|
- Credentials are not eagerly resolved by the core runtime, serialized into provider requests/events/stores, or passed to loops/compaction.
|
|
@@ -128,7 +128,7 @@ A branch leaf (`leaf_entry_id`) is the current entry id for that branch. Rebuild
|
|
|
128
128
|
| --- | --- |
|
|
129
129
|
| `prism_session_entries` | `id` PK, `session_id` FK, `parent_id`, `run_id`, `timestamp`, `kind`, `schema_version`, `message` JSONB, `event` JSONB, `model` JSONB, `previous_model` JSONB, `label`, `summary`, `data` JSONB, `metadata` JSONB |
|
|
130
130
|
|
|
131
|
-
Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version`
|
|
131
|
+
Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version` is nullable in schema version 3 for existing entry compatibility; hosts writing new rows should persist the entry's current schema version. `parent_id` may be null for the root entry of a session.
|
|
132
132
|
|
|
133
133
|
`SessionAppendOptions.idempotencyKey` is not part of `SessionEntry`; store it in an adapter-owned side table when you need durable retry detection:
|
|
134
134
|
|
|
@@ -214,7 +214,8 @@ await runRunLedgerConformance(() => createLedger(testDatabase), { exerciseReopen
|
|
|
214
214
|
| Primitive | Purpose |
|
|
215
215
|
| --- | --- |
|
|
216
216
|
| `PersistenceSchemaModel` | Versioned table/column/index model covering sessions, entries, parent chain, idempotency side table, runs, events, tool calls, usage, tenant columns, and `prism_migrations` |
|
|
217
|
-
| `createPersistenceMigrationContract()` | Strictly increasing migration steps, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
|
|
217
|
+
| `createPersistenceMigrationContract()` | Strictly increasing checked-in migration steps with deterministic SHA-256 checksums, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
|
|
218
|
+
| `assertAppliedPersistenceMigrations()` / `assertPersistenceSchemaShape()` | Testable fail-closed history and normalized SQLite/PostgreSQL catalog checks for every required table, column/type/null/default, PK/unique/FK, and named index. |
|
|
218
219
|
| `getPersistencePaginationCursors()` | Indexed `(session_id, timestamp, id)`, `(run_id, sequence)`, `(run_id, recorded_at, id)` cursor shapes that avoid offset scans |
|
|
219
220
|
| `assertPersistenceQueryPaginationConforms()` | Generic cursor pagination fixture for `queryEntries` |
|
|
220
221
|
| `assertTenantScopedQueryIsolation()` | Tenant-filtered reads must not leak rows or primary-id collisions across tenants |
|
|
@@ -345,7 +346,11 @@ Retention jobs should not run inside the agent/session runtime. They are a host
|
|
|
345
346
|
|
|
346
347
|
## Migrations
|
|
347
348
|
|
|
348
|
-
Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library.
|
|
349
|
+
Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library. First-party SQLite/PostgreSQL adapters automatically verify their checked-in schema-v3 history and catalog at open, before runtime writes. Their catalog reads are bounded metadata queries/PRAGMAs, not table-data scans.
|
|
350
|
+
|
|
351
|
+
Each new adapter-owned migration row records the contract SHA-256 checksum. A complete known v0.0.5 history whose checksum values are all `NULL` is a one-time compatibility case: under the SQLite transaction or PostgreSQL advisory transaction lock, the adapter verifies the full current shape, backfills all checksums, and continues. Unknown/duplicate/out-of-order/name-version/checksum mismatch, mixed/partial legacy values, or any missing/renamed/wrong-type/null/default/key/index artifact fails closed. Restore or apply a reviewed host migration; never edit checksums to silence drift.
|
|
352
|
+
|
|
353
|
+
Recommended migration practices:
|
|
349
354
|
|
|
350
355
|
- Use a sequential or timestamped migration naming convention.
|
|
351
356
|
- Store applied migrations in `prism_migrations` with `name`, `version`, `applied_at`, `applied_by`, and `checksum`.
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
# Guardrails
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Guardrails are typed, fail-closed checks at input, completed provider output, tool input, and raw tool output boundaries. `session.run()` evaluates configured stages through one core runner; `dispatchToolCall()` uses same runner for direct, MCP-server, and workflow tool calls.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
Use guardrails to block unsafe prompts, model responses, tool arguments, or tool results before their next boundary. Use a redactor for known secrets. Do not treat guardrails as a sandbox, secret detector, permission policy, or validation replacement.
|
|
10
|
+
|
|
11
|
+
## Inputs / request
|
|
12
|
+
|
|
13
|
+
```ts
|
|
14
|
+
import type { Guardrail, Guardrails } from "@arnilo/prism";
|
|
15
|
+
|
|
16
|
+
const pii: Guardrail<"input"> = {
|
|
17
|
+
name: "pii",
|
|
18
|
+
stage: "input",
|
|
19
|
+
evaluate: ({ value }) => JSON.stringify(value).includes("SSN")
|
|
20
|
+
? { action: "tripwire", reason: "pii" }
|
|
21
|
+
: { action: "allow" },
|
|
22
|
+
};
|
|
23
|
+
|
|
24
|
+
const guardrails: Guardrails = { input: [pii], maxConcurrency: 1 };
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Set `AgentConfig.guardrails` for every session run or `RunOptions.guardrails` to append checks for one run. `DispatchToolCallOptions.guardrails`, workflow `RunWorkflowOptions.guardrails`, and MCP server `CreatePrismMcpServerOptions.guardrails` apply tool stages to direct calls. A stage has `Guardrail<"input" | "output" | "tool_input" | "tool_output">`, a name, optional revision, and `evaluate(context)` result.
|
|
28
|
+
|
|
29
|
+
Decisions are `allow`, `block`, `tripwire`, or `interrupt`. Evaluation defaults to declaration-order sequential. `maxConcurrency` may be 1–16; records are emitted in declaration order. Thrown or malformed decisions become a fail-closed tripwire. Decision reasons are capped at 4 KiB and metadata at 16 KiB after JSON normalization and optional redaction.
|
|
30
|
+
|
|
31
|
+
## Outputs / response / events
|
|
32
|
+
|
|
33
|
+
Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
|
|
34
|
+
|
|
35
|
+
Ordering is fixed:
|
|
36
|
+
|
|
37
|
+
1. input before session append, compaction, or provider work;
|
|
38
|
+
2. provider output is privately collected, then output checks run before any assistant message event or persistence;
|
|
39
|
+
3. tool input runs after tool-call middleware normalization and before lookup, permission, validation, execution policy, and side effect;
|
|
40
|
+
4. tool output runs after the side effect but before redaction, tool events, ledger rows, transcript append, or next turn.
|
|
41
|
+
|
|
42
|
+
With no output guardrails, provider streaming retains existing behavior. With output guardrails, message events are buffered until the completed provider turn is allowed.
|
|
43
|
+
|
|
44
|
+
## Request/response example
|
|
45
|
+
|
|
46
|
+
```json
|
|
47
|
+
{
|
|
48
|
+
"event": {
|
|
49
|
+
"type": "guardrail_decision",
|
|
50
|
+
"record": { "guardrail": "pii", "stage": "input", "action": "tripwire", "reason": "pii" }
|
|
51
|
+
}
|
|
52
|
+
}
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
## Implementation example
|
|
56
|
+
|
|
57
|
+
```ts
|
|
58
|
+
const agent = createAgent({ model, provider, guardrails: { input: [pii], output: [responseGuard] } });
|
|
59
|
+
await agent.createSession().run("Draft reply", { guardrails: { toolInput: [commandGuard] } });
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
## Extension and configuration notes
|
|
63
|
+
|
|
64
|
+
Guardrails are callbacks supplied by the host. Prism does not discover, load, retry, or persist callback code. `createSecureAgent()` keeps configured guardrails and only appends run-level checks; it never lets a run remove secure defaults. Custom loops receive guarded `LoopContext.generate()` and `LoopContext.dispatchToolCall()`; host code that directly calls a provider or `ToolDefinition.execute()` is outside the runtime boundary.
|
|
65
|
+
|
|
66
|
+
## Security and performance notes
|
|
67
|
+
|
|
68
|
+
Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work.
|
|
69
|
+
|
|
70
|
+
## Related APIs
|
|
71
|
+
|
|
72
|
+
- [Agent/session runtime](agent-session-runtime.md)
|
|
73
|
+
- [Tools](tools.md)
|
|
74
|
+
- [Agent events](agent-events.md)
|
|
75
|
+
- [Host security](host-security.md)
|