@arnilo/prism 0.0.4 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/CHANGELOG.md +46 -1
  2. package/README.md +34 -10
  3. package/dist/agent-loops.d.ts +1 -0
  4. package/dist/agent-loops.js +26 -16
  5. package/dist/agents.js +147 -21
  6. package/dist/cli-init.d.ts +41 -0
  7. package/dist/cli-init.js +390 -0
  8. package/dist/cli-runner.d.ts +7 -1
  9. package/dist/cli-runner.js +13 -1
  10. package/dist/content.d.ts +19 -0
  11. package/dist/content.js +197 -69
  12. package/dist/contracts.d.ts +96 -9
  13. package/dist/contracts.js +8 -0
  14. package/dist/feedback.d.ts +48 -0
  15. package/dist/feedback.js +230 -0
  16. package/dist/ids.d.ts +2 -0
  17. package/dist/ids.js +6 -0
  18. package/dist/index.d.ts +10 -4
  19. package/dist/index.js +6 -3
  20. package/dist/providers/media.d.ts +3 -1
  21. package/dist/providers/media.js +11 -1
  22. package/dist/session-stores.js +2 -3
  23. package/dist/testing/feedback.d.ts +6 -0
  24. package/dist/testing/feedback.js +37 -0
  25. package/dist/testing/persistence-schema.d.ts +48 -10
  26. package/dist/testing/persistence-schema.js +166 -22
  27. package/dist/testing/run-ledger-conformance.js +7 -1
  28. package/dist/thinking.d.ts +42 -0
  29. package/dist/thinking.js +92 -0
  30. package/dist/tools.js +2 -3
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +75 -0
  34. package/docs/agent-events.md +14 -21
  35. package/docs/agent-loops.md +12 -9
  36. package/docs/agent-session-runtime.md +14 -16
  37. package/docs/cli-rpc.md +35 -7
  38. package/docs/coding-agent-tools.md +35 -14
  39. package/docs/coding-security.md +7 -3
  40. package/docs/compaction-llm.md +17 -7
  41. package/docs/compaction-observational-memory.md +30 -4
  42. package/docs/context-and-skills.md +1 -0
  43. package/docs/credential-storage.md +58 -9
  44. package/docs/credentials-and-redaction.md +3 -3
  45. package/docs/database-persistence.md +17 -9
  46. package/docs/evaluations.md +122 -0
  47. package/docs/extensions.md +2 -2
  48. package/docs/host-security.md +26 -5
  49. package/docs/index.md +43 -28
  50. package/docs/mcp-tools.md +74 -13
  51. package/docs/migration.md +177 -3
  52. package/docs/multimodal-content.md +14 -6
  53. package/docs/node-filesystem-config.md +1 -0
  54. package/docs/node-jsonl-session-store.md +5 -4
  55. package/docs/observability.md +14 -6
  56. package/docs/performance.md +209 -0
  57. package/docs/postgres-persistence.md +8 -6
  58. package/docs/provider-caching.md +16 -4
  59. package/docs/provider-conformance.md +40 -1
  60. package/docs/provider-packages.md +62 -3
  61. package/docs/providers/ai-sdk.md +149 -0
  62. package/docs/providers/kimi.md +124 -61
  63. package/docs/providers/neuralwatt.md +19 -13
  64. package/docs/providers/openai.md +56 -13
  65. package/docs/providers/opencode-go.md +118 -30
  66. package/docs/providers/openrouter.md +105 -35
  67. package/docs/providers/zai.md +94 -45
  68. package/docs/public-contracts.md +6 -5
  69. package/docs/rag.md +113 -0
  70. package/docs/release-and-install.md +100 -79
  71. package/docs/review-coverage-2026-07-15.md +193 -0
  72. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  73. package/docs/runs-and-usage.md +42 -5
  74. package/docs/server.md +139 -0
  75. package/docs/settings-auth-trust-security.md +5 -5
  76. package/docs/sqlite-persistence.md +6 -5
  77. package/docs/structured-output.md +1 -1
  78. package/docs/supervisors.md +71 -0
  79. package/docs/thinking-and-reasoning.md +98 -0
  80. package/docs/tool-execution-primitives.md +3 -3
  81. package/docs/tools.md +15 -0
  82. package/docs/use-case-model-selection.md +109 -0
  83. package/docs/workflow-orchestration-primitives.md +20 -3
  84. package/docs/workflows.md +114 -33
  85. package/docs/working-and-semantic-memory.md +170 -0
  86. package/package.json +13 -3
  87. package/templates/init/README.md.tmpl +28 -0
  88. package/templates/init/env.example.tmpl +1 -0
  89. package/templates/init/gitignore.tmpl +11 -0
  90. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  91. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  92. package/templates/init/package.json.tmpl +22 -0
  93. package/templates/init/providers.json +76 -0
  94. package/templates/init/src/agent.ts.tmpl +10 -0
  95. package/templates/init/src/index.ts.tmpl +12 -0
  96. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  97. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -25,15 +25,16 @@ Key exports:
25
25
  | Field | Purpose |
26
26
  | --- | --- |
27
27
  | `provider` / `summaryProvider` | Explicit `AIProvider`, or factory receiving a resolved credential. |
28
- | `model` / `summaryModel` | Explicit summary `ModelConfig`. |
28
+ | `model` / `summaryModel` | Summary `ModelConfig`. `summaryModel` wins; `model` is the session/host fallback via `resolveUseCaseModel`. See [Use-case model selection](use-case-model-selection.md). |
29
29
  | `credential`, `credentialRequest` | Optional per-call credential resolution for provider factories. |
30
30
  | `providerOptions` | Generic `ProviderRequest.options`, including cache fields. |
31
31
  | `providerRequestPolicies` | Optional Prism provider request policies applied before the summary call. |
32
32
  | `customInstructions` | Additional summary focus appended to prompts. |
33
- | `thinkingLevel` | Passed as `ProviderRequest.options.extra.thinkingLevel`. |
34
- | `reserveTokens` | Output budget basis; defaults to `16384`. |
33
+ | `thinkingLevel` | Mapped into `ProviderRequest.options.compat` via `applyThinkingLevel` / `thinkingFamilyForModel` (not inert `extra.thinkingLevel`). See [Thinking and reasoning](thinking-and-reasoning.md). |
34
+ | `reserveTokens` | Output budget basis; defaults to `16384`, hard cap `131072`. |
35
35
  | `keepRecentTokens` | Approximate recent-token budget; defaults to `20000`. |
36
- | `maxSummaryTokens` / `maxOutputTokens` | Computes the summary output budget, writes it to `summaryModel/model.parameters.maxTokens`, and truncates oversized collected summaries. First-party providers serialize that generic field to their real request field (`max_output_tokens` for OpenAI Responses, `max_tokens` for OpenAI-compatible/Anthropic-style providers). |
36
+ | `maxSummaryTokens` / `maxOutputTokens` | Summary retention/request ceiling; default `16384`, hard cap `131072`. `maxSummaryTokens` wins over the compatibility alias. The finite value is written to `model.parameters.maxTokens`; first-party providers map it to their wire field. |
37
+ | `maxErrorBytes` | Retained provider/factory/policy error detail; default `1024`, hard cap `8192`, UTF-8-safe and known-secret redacted. |
37
38
  | `maxToolResultChars` | Tool-result JSON truncation limit; defaults to `2000`. |
38
39
  | `trackFileOperations`, `includeFileOperations` | Control file path extraction and final summary blocks. |
39
40
  | `secrets` | Exact strings to redact from serialized prompts and final summaries. |
@@ -64,7 +65,8 @@ const strategy = createLlmCompactionStrategy({
64
65
  model: { provider: "openai", model: "gpt-4.1-mini" },
65
66
  keepRecentTokens: 20_000,
66
67
  reserveTokens: 16_384,
67
- maxOutputTokens: 800,
68
+ maxSummaryTokens: 800,
69
+ maxErrorBytes: 1_024,
68
70
  providerOptions: { cacheRetention: "short" },
69
71
  customInstructions: "Focus on current files and failing tests.",
70
72
  });
@@ -102,10 +104,18 @@ const agent = createAgent({ model, provider, compaction: { strategy, thresholdEn
102
104
  Registration only contributes an inert strategy. The host must resolve and pass it to runtime config.
103
105
 
104
106
  ## Security and performance notes
105
- Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization. Output-budget calculation is O(1) and does not add an extra summarization call. The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. It does not discover credentials, read files, start background jobs, or add provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history or provider output.
107
+ Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization. Limit options must be positive safe integers at or below their hard caps and reject during strategy creation. Missing output options use a 16,384-token summary ceiling; reserve ratio/model metadata may narrow the provider request, never remove its finite `maxTokens`. A request policy that replaces `maxTokens` with NaN, Infinity, zero, an unsafe integer, or above-hard-cap input fails before provider generation.
108
+
109
+ Provider deltas are redacted while retained and stop at `maxSummaryTokens * 4` UTF-16 code units without splitting a surrogate pair. Provider iteration is closed/aborted on overflow. A derived finite event ceiling also stops endless empty/non-text deltas. Final history/turn/file composition receives the same cap. Provider error events, generator throws, provider-factory failures, and policy failures expose only bounded redacted detail; host abort remains authoritative.
110
+
111
+ The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. It does not discover credentials, read files, start background jobs, or add provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history or provider output.
106
112
 
107
113
  ## Related APIs
108
- - [Compaction and retry policies](compaction-and-retry.md): core compaction strategy surface.
114
+
115
+ - [Use-case model selection](use-case-model-selection.md): `summaryModel` vs session `model` fallback.
116
+ - [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → `compat`.
117
+ - [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary and core compaction strategy surface.
118
+ - [Observational memory compaction package](compaction-observational-memory.md): source-backed memory workers with the same use-case binding pattern.
109
119
  - [Agent/session runtime](agent-session-runtime.md): `AgentSession.compact()` and opt-in auto-compaction.
110
120
  - [Provider layer](provider-layer.md): mock providers and provider request contracts.
111
121
  - [Credentials and redaction](credentials-and-redaction.md): exact known-secret redaction behavior.
@@ -6,6 +6,8 @@
6
6
 
7
7
  Current status: ledger/projection/render/recall utilities, explicit worker runtime, fast compaction strategy, inert extension helper, recall tool, and status/view command factories are available.
8
8
 
9
+ This package is distinct from `@arnilo/prism-memory` working/semantic memory: observational memory compresses and recalls source-backed observations/reflections; semantic memory retrieves embeddings; working memory stores the current structured profile/state. Hosts may compose both.
10
+
9
11
  ## When to use it
10
12
 
11
13
  Use it when a host wants to opt in to long-session memory that records observations/reflections as session custom entries, renders prepared memory during compaction, and supports exact-id recall.
@@ -25,6 +27,20 @@ Memory records use `SessionEntry.kind: "custom"` with `entry.data.type` markers:
25
27
 
26
28
  Ids are known, source-backed 12-character lowercase hex strings matching `^[a-f0-9]{12}$`.
27
29
 
30
+ Worker limits are finite positive safe integers:
31
+
32
+ | Runtime option | Default | Hard cap | Scope |
33
+ | --- | ---: | ---: | --- |
34
+ | `maxWorkerTurns` | 16 | 64 | Provider turns per observer/reflector/dropper run; overrides settings `agentMaxTurns` |
35
+ | `maxWorkerToolCallsPerTurn` | 32 | 256 | Calls retained from one provider response |
36
+ | `maxWorkerToolCalls` | 128 | 1,024 | Calls across all turns in one worker run |
37
+ | `maxWorkerArgumentBytes` | 64 KiB | 1 MiB | Each raw and redacted JSON argument object |
38
+ | `maxWorkerResultBytes` | 64 KiB | 1 MiB | Full tool result and replayed value/error payload |
39
+ | `maxWorkerMessageBytes` | 1 MiB | 8 MiB | System/prompt plus assistant-call/tool-result transcript |
40
+ | `maxWorkerErrorBytes` | 1 KiB | 8 KiB | Provider/tool/runtime error text after exact known-secret redaction |
41
+
42
+ Direct `runObserver()` / `runReflector()` / `runDropper()` calls retain required `maxTurns` and accept the corresponding shorter worker fields (`maxToolCalls`, `maxResultBytes`, etc.). Named default/hard constants and `resolveMemoryWorkerLimits()` are exported.
43
+
28
44
  ## Outputs / response / events
29
45
 
30
46
  Key exports:
@@ -76,7 +92,12 @@ const memory = createObservationalMemoryRuntime({
76
92
  session,
77
93
  appendEntry: (entry) => store.append(entry),
78
94
  workerProvider,
79
- workerModel: { provider: "mock", model: "memory" },
95
+ sessionModel: agent.config.model, // fallback when workerModel unset
96
+ // workerModel: { provider: "mock", model: "memory" }, // optional override
97
+ maxWorkerTurns: 8,
98
+ maxWorkerToolCalls: 64,
99
+ maxWorkerResultBytes: 64 * 1024,
100
+ overrides: { thinkingLevel: "low" },
80
101
  });
81
102
  await memory.flush();
82
103
  await session.compact({ strategy: createObservationalMemoryCompactionStrategy({ keepRecentEntries: 8 }) });
@@ -90,9 +111,9 @@ await kernel.load([createObservationalMemoryExtension({ recallTool: { getEntries
90
111
 
91
112
  ## Extension and configuration notes
92
113
 
93
- Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`.
114
+ Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`. `agentMaxTurns` now rejects non-integer/non-finite/out-of-range input (hard 64) instead of flooring/falling back. Runtime `maxWorkerTurns` takes precedence.
94
115
 
95
- The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, `workerProvider`, and `workerModel`. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution.
116
+ The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, and `workerProvider`. Model selection uses [use-case model selection](use-case-model-selection.md): pass optional `workerModel` (or settings `workerModel`) to override, and `sessionModel: agent.config.model` so workers fall back to the session model when no worker model is configured. `requireExplicitModel: true` restores the historical `missing_model` skip when no explicit worker model is set. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution. Default credential requests use the **resolved** model's provider id.
96
117
 
97
118
  `createObservationalMemoryCompactionStrategy()` keeps recent message entries like the default compaction strategy, renders existing observations/reflections as the summary, and returns a standard Prism compaction entry. Its `data` includes `throughEntryId`, `keepEntryIds`, `strategy`, `trigger`, and `memory: { type: "om.folded", version: 1, fullFold, observations, reflections, droppedObservationIds }`. When active observations exceed `observationsPoolMaxTokens`, it performs a full fold into `data.memory`.
98
119
 
@@ -108,13 +129,18 @@ The runtime requires host-supplied `session`, an `appendEntry` callback bound to
108
129
  - Recall tool and commands only see current-branch entries supplied by the host callback.
109
130
  - Invalid or missing ids fail closed; invalid recall tool ids skip entry lookup.
110
131
  - Utilities and fast compaction are O(n) over supplied entries and use no provider, network, filesystem, timer, worker, credential, or settings access.
111
- - Workers serialize only supplied branch entries, enforce `agentMaxTurns`, and run one consolidation pipeline at a time per runtime. Worker transcripts replay assistant `tool_call` messages before matching role `tool` `tool_result` messages so provider requests stay valid for call/result-pairing providers.
132
+ - Workers serialize only supplied branch entries within `maxWorkerMessageBytes`, enforce finite turns/calls/arguments/results/messages/errors, and run one consolidation pipeline at a time per runtime. Source serialization and reflection/drop prompts fail before joining beyond the transcript cap.
133
+ - Every provider call must name a registered worker tool. Unknown calls, call overflow, oversized/deep/cyclic/non-JSON arguments/results, and transcript overflow fail deterministically; no excess call enters the assistant transcript or executes.
134
+ - Raw arguments are measured before tool execution. Full results are measured before redaction/replay; the bounded redacted value/error is then measured again because replacement text can grow. Replayed call arguments, tool values/errors, runtime `lastError`, and debug error data contain exact known-secret redaction. Host tools may already have caused side effects before returning an invalid oversized result; keep worker tools small/idempotent.
135
+ - Worker transcripts replay assistant `tool_call` messages before matching role `tool` `tool_result` messages so provider requests stay valid for call/result-pairing providers. Calls produced on the final allowed turn execute and persist, but no additional provider turn starts.
112
136
  - Compaction preserves raw history; Prism appends one standard compaction entry and rebuilds provider context from its summary plus kept recent messages.
113
137
  - Pass known secrets to render/recall/runtime/tool/command helpers to redact exact values from prompts, records, structured results, and text output.
114
138
  - Live tests are opt-in with `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS=1`.
115
139
 
116
140
  ## Related APIs
117
141
 
142
+ - [Use-case model selection](use-case-model-selection.md): session vs worker model binding and `resolveUseCaseModel`.
143
+ - [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → provider `compat`.
118
144
  - [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary.
119
145
  - [LLM compaction package](compaction-llm.md): existing optional compaction-package pattern.
120
146
  - [Session stores and branching](session-stores-and-branching.md): branch entries that observational memory reads and appends to.
@@ -179,6 +179,7 @@ Use `activateAllCapabilities: true` only as a temporary all-skills/all-tools com
179
179
  - [Agent/session runtime](agent-session-runtime.md): consumes host-selected context providers and skills from explicit agent config.
180
180
  - [Input and prompt assembly](input-and-prompt-assembly.md): default prompt builder and provider-input assembly helper.
181
181
  - [Instruction injection](instruction-injection.md): package injectors contribute `contextBlocks` that merge after host+skill provider blocks.
182
+ - [Retrieval-augmented generation](rag.md): optional retrieved citations contribute through the same explicit inert context seam.
182
183
  - [Public contracts](public-contracts.md): `ContextProvider`, `ContextResolutionContext`, `ContextBlock`, `Skill`, `SkillRegistry`, `PromptBuilder`, and `PromptBuildRequest`.
183
184
  - [Middleware hooks](middleware-hooks.md): `context` and `prompt_build` hooks.
184
185
  - [Contribution registries](contribution-registries.md): inert context provider and skill contributions.
@@ -44,8 +44,11 @@ import {
44
44
  | --- | --- | --- |
45
45
  | `path` | `string` | Vault file path. Parent directories are created as needed. |
46
46
  | `getPassphrase` | `() => string \| Promise<string>` | Host-owned passphrase retrieval. Never logged by the adapter. |
47
- | `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32`. Minimum `N=16384`. |
48
- | `fileMode` | `number` | Unix mode for newly written files. Defaults to `0o600`. |
47
+ | `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32`; limits are listed below. |
48
+ | `fileMode` | `number` | Unix mode for files. Defaults to `0o600`; group/other permissions are rejected. |
49
+ | `limits.maxFileBytes` | `number` | Encrypted envelope file: 4 MiB default, 16 MiB hard cap. |
50
+ | `limits.maxVaultBytes` | `number` | Decrypted vault/plaintext: 3 MiB default, 12 MiB hard cap. |
51
+ | `limits.maxScryptMemoryBytes` | `number` | `128*N*r` memory estimate: 256 MiB default and hard cap. |
49
52
 
50
53
  ### System keychain
51
54
 
@@ -53,7 +56,8 @@ import {
53
56
  | --- | --- | --- |
54
57
  | `service` | `string` | Keychain service name (application identifier). |
55
58
  | `namespace` | `string` | Optional prefix separating environments or tenants within one service. |
56
- | `timeoutMs` | `number` | Operation timeout. Defaults to `5000`. |
59
+ | `timeoutMs` | `number` | Operation timeout. Defaults to 5,000 ms; hard cap 60,000 ms. |
60
+ | `maxPayloadBytes` | `number` | Decrypted keychain payload: 3 MiB default, 12 MiB hard cap. |
57
61
 
58
62
  ## Outputs / response / events
59
63
 
@@ -71,6 +75,8 @@ Encrypted file stores also expose:
71
75
  - `reload()` — re-read and decrypt from disk
72
76
  - `flush()` — force rewrite of the encrypted envelope
73
77
 
78
+ `encryptBytes()` and `decryptBytes()` are Promise-based because they use asynchronous `node:crypto.scrypt`.
79
+
74
80
  Errors are explicit and fail closed:
75
81
 
76
82
  | Error | Code | When |
@@ -123,6 +129,7 @@ import {
123
129
  const store = await openEncryptedCredentialStore({
124
130
  path: "./credentials.vault",
125
131
  getPassphrase: () => process.env.MY_APP_CREDENTIAL_PASSPHRASE!,
132
+ limits: { maxFileBytes: 4 * 1024 * 1024, maxVaultBytes: 3 * 1024 * 1024 },
126
133
  });
127
134
 
128
135
  const resolver = createExplicitCredentialResolver([
@@ -151,6 +158,47 @@ await rotateEncryptedCredentialStorePassphrase({
151
158
  });
152
159
  ```
153
160
 
161
+ ### Desktop keychain and explicit overrides
162
+
163
+ ```ts
164
+ import {
165
+ createEnvCredentialResolver,
166
+ createExplicitCredentialResolver,
167
+ createMemoryCredentialStore,
168
+ } from "@arnilo/prism";
169
+ import {
170
+ createKeychainCredentialStore,
171
+ createStoredCredentialResolver,
172
+ } from "@arnilo/prism-credentials-node";
173
+ import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
174
+
175
+ const keychain = createKeychainCredentialStore({
176
+ service: "com.example.my-app",
177
+ namespace: "production",
178
+ });
179
+ await keychain.set({
180
+ name: "apiKey",
181
+ provider: "openai",
182
+ credential: { type: "api_key", value: userSuppliedKey },
183
+ });
184
+
185
+ const runtimeOverrides = createMemoryCredentialStore();
186
+ // Set only for this user/agent instance; it wins over stored and env values.
187
+ runtimeOverrides.set({
188
+ name: "apiKey",
189
+ provider: "openai",
190
+ credential: { type: "api_key", value: temporaryOverride },
191
+ });
192
+
193
+ const apiKey = createExplicitCredentialResolver([
194
+ { name: "runtime", resolver: runtimeOverrides },
195
+ { name: "keychain", resolver: createStoredCredentialResolver(keychain) },
196
+ { name: "env", resolver: createEnvCredentialResolver(process.env, { openai: "OPENAI_API_KEY" }) },
197
+ ]);
198
+
199
+ const providers = createOpenAIProviderPackage({ apiKey });
200
+ ```
201
+
154
202
  ## Extension and configuration notes
155
203
 
156
204
  - Passphrase retrieval, TLS, and OS permission prompts remain host-owned.
@@ -161,12 +209,13 @@ await rotateEncryptedCredentialStorePassphrase({
161
209
 
162
210
  ## Security and performance notes
163
211
 
164
- - Authenticated encryption uses Node built-in `aes-256-gcm` and `scrypt`; no extra crypto dependencies for the file backend.
165
- - Atomic writes use temp file + rename; partial writes cannot replace a valid vault.
166
- - Derived keys are zeroed after encrypt/decrypt operations where practical.
167
- - Default scrypt `N=32768` targets interactive CLI unlock; raise `N` for higher security at the cost of unlock latency.
168
- - Keychain operations honor `timeoutMs` and surface `CredentialStoreTimeoutError` instead of blocking indefinitely.
169
- - Never log passphrases, derived keys, or decrypted credential payloads. Error messages do not echo secret values.
212
+ - Authenticated encryption uses Node built-in `aes-256-gcm` and asynchronous `scrypt`; no extra crypto dependency is added for the file backend.
213
+ - Envelope parsing rejects unknown shape, non-canonical/oversized base64, wrong salt/IV/tag size, unsupported algorithms/version, and excessive KDF work before scrypt. `N` must be a power of two from 16,384–262,144; `r≤32`, `p≤16`, `keyLength=32`, `N*r*p≤2,097,152`, and `128*N*r` must fit `maxScryptMemoryBytes`.
214
+ - Existing Unix vaults are checked before content read and must deny group/other access. Atomic writes create a random exclusive temp file at the requested restrictive mode, then rename; Windows skips Unix mode checks.
215
+ - Derived keys and package-owned plaintext buffers are zeroed after use. JavaScript passphrase strings and returned credentials remain host-owned.
216
+ - Keychain operations use `@napi-rs/keyring`'s abort-aware `AsyncEntry`, so native work runs outside the JavaScript event loop. A main-loop timer aborts and rejects at `timeoutMs`; native cancellation remains OS/backend-dependent and may briefly retain one libuv worker after rejection.
217
+ - Keychain payloads are bytes rather than password strings and are zeroed after parse/write. Unknown native errors are mapped to sanitized typed errors; no native message or secret value is echoed.
218
+ - Never log passphrases, derived keys, or decrypted credential payloads.
170
219
  - Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
171
220
 
172
221
  ## Related APIs
@@ -95,7 +95,7 @@ console.log(error.message);
95
95
  ## Extension and configuration notes
96
96
 
97
97
  - Hosts and extension packages can implement `CredentialResolver` and pass it explicitly to code that needs credentials.
98
- - `AgentConfig.credentials` is host-owned metadata for compatibility; `createAgent()` / `session.run()` do not call `credentials.resolve()`. Provider adapters, compaction workers, or request policies should receive and resolve credentials at the provider edge.
98
+ - Credentials stay host-owned outside `AgentConfig`. `createAgent()` / `session.run()` do not call `credentials.resolve()`. Provider adapters, compaction workers, or request policies should receive and resolve credentials at the provider edge.
99
99
  - Use `createExplicitCredentialResolver()` when documenting a fixed order such as runtime override, stored credential, caller-provided env object, then fallback resolver.
100
100
  - Use `createEnvCredentialResolver()` only with an object supplied by the host; Prism does not read `process.env` for you.
101
101
  - Provider adapters should resolve credentials as late as possible, per request.
@@ -107,10 +107,10 @@ console.log(error.message);
107
107
  - Redaction only removes exact known secret values passed to the helper. It is not a general-purpose secret detector.
108
108
  - Do not pass empty strings as secrets; they are ignored.
109
109
  - `redactSecrets()` recursively walks arrays and object entries, so avoid using it on huge objects unless needed.
110
- - Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via a `WeakSet` visited-set. Self-referential or mutually referenced objects render `"[Circular]"` at the back-reference instead of throwing. `Date` and `RegExp` values are passed through unchanged; `ArrayBuffer` and typed arrays are passed through unchanged; `Map` is normalized to a plain object and `Set` to an array so the output stays JSON-compatible. `errorToErrorInfo()` tolerates a cyclic `error.cause` (rendered via `String()`).
110
+ - Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via an active-path `WeakSet`. Ancestor cycles render `"[Circular]"` at the back-reference instead of throwing; shared diamond references on separate branches stay structured (they are not collapsed to `"[Circular]"`). String object keys and `Map` keys are redacted like values, with deterministic `__2`/`__3`/… suffixes on collisions. `Date` and `RegExp` values are passed through unchanged; `ArrayBuffer` and typed arrays are passed through unchanged; `Map` is normalized to a plain object and `Set` to an array so the output stays JSON-compatible. `errorToErrorInfo()` tolerates a cyclic `error.cause` (rendered via `String()`). Symbols and non-enumerable properties remain outside JSON-shaped redaction.
111
111
  - Use placeholders in tests and docs. Never commit real tokens.
112
112
  - Live provider/worker tests are gated behind explicit environment variables and skipped by default: `PRISM_LIVE_PROVIDER_TESTS`, `PRISM_LIVE_COMPACTION_TESTS`, `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS`. Default `npm test` is network-free; do not add ungated network calls to default tests.
113
- - `AgentConfig.credentials` is not eagerly resolved, serialized into provider requests/events/stores, or passed to loops/compaction by the core runtime.
113
+ - Credentials are not eagerly resolved by the core runtime, serialized into provider requests/events/stores, or passed to loops/compaction.
114
114
  - `resolveCredentialValue()` and `createExplicitCredentialResolver()` do not cache values. Add host-side caching only if a real credential source needs it.
115
115
  - `refreshOAuthCredential()` only calls the supplied OAuth provider and optional store; it has no built-in persistence or retry loop.
116
116
  - OpenAI Codex device-code OAuth polls inside `createOpenAICodexOAuthProvider().login()` with bounded delays and abort support via `OAuthLoginCallbacks.signal`. Token-endpoint failures redact authorization codes, PKCE verifiers, device/user codes, and access/refresh tokens when those values are known.
@@ -67,7 +67,7 @@ Important shapes:
67
67
  | `RunRecord` | Stored run with `sessionId`, `branchId`, status (`queued` \| `running` \| `succeeded` \| `failed` \| `aborted`), `model`, `provider`, `idempotencyKey`, `abortReason`, and `error`. |
68
68
  | `AgentEventRecord` | Event ledger row with `event: AgentEvent` and a `redacted` flag. Hosts redact before storage. |
69
69
  | `ToolCallRecord` | Tool-call row with `arguments`, optional `result: ToolResult`, `reason`, `progress` snapshots, status, and a `redacted` flag. |
70
- | `UsageRecord` | Usage row wrapping `Usage` with session/run/entry linkage. |
70
+ | `UsageRecord` | Scoped provider-turn or aggregate run usage with session/run/entry and turn/attempt linkage. |
71
71
  | `AgentDefinitionRecord` | Versioned agent definition snapshot. Only stores `AgentDefinition` data; never provider credentials/resolvers/instances. |
72
72
  | `RetentionPolicy` | Policy with `maxAgeDays`, `maxEntriesPerSession`, `maxTotalBytes`, `archiveStore`, and `appliedKinds`. |
73
73
  | `MigrationRecord` | Applied migration with name, version, timestamp, checksum, and applied-by. |
@@ -128,7 +128,7 @@ A branch leaf (`leaf_entry_id`) is the current entry id for that branch. Rebuild
128
128
  | --- | --- |
129
129
  | `prism_session_entries` | `id` PK, `session_id` FK, `parent_id`, `run_id`, `timestamp`, `kind`, `schema_version`, `message` JSONB, `event` JSONB, `model` JSONB, `previous_model` JSONB, `label`, `summary`, `data` JSONB, `metadata` JSONB |
130
130
 
131
- Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version` defaults to `1`. `parent_id` may be null for the root entry of a session.
131
+ Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version` is nullable in schema version 3 for existing entry compatibility; hosts writing new rows should persist the entry's current schema version. `parent_id` may be null for the root entry of a session.
132
132
 
133
133
  `SessionAppendOptions.idempotencyKey` is not part of `SessionEntry`; store it in an adapter-owned side table when you need durable retry detection:
134
134
 
@@ -166,7 +166,7 @@ The `event` JSONB stores a redacted `AgentEvent`. The `sequence` column is an im
166
166
 
167
167
  | Table | Key columns |
168
168
  | --- | --- |
169
- | `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
169
+ | `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `scope`, `turn`, `attempt`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
170
170
 
171
171
  The `usage` JSONB stores the `Usage` shape: input/output/total/cache tokens, cost, and currency.
172
172
 
@@ -214,7 +214,8 @@ await runRunLedgerConformance(() => createLedger(testDatabase), { exerciseReopen
214
214
  | Primitive | Purpose |
215
215
  | --- | --- |
216
216
  | `PersistenceSchemaModel` | Versioned table/column/index model covering sessions, entries, parent chain, idempotency side table, runs, events, tool calls, usage, tenant columns, and `prism_migrations` |
217
- | `createPersistenceMigrationContract()` | Strictly increasing migration steps, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
217
+ | `createPersistenceMigrationContract()` | Strictly increasing checked-in migration steps with deterministic SHA-256 checksums, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
218
+ | `assertAppliedPersistenceMigrations()` / `assertPersistenceSchemaShape()` | Testable fail-closed history and normalized SQLite/PostgreSQL catalog checks for every required table, column/type/null/default, PK/unique/FK, and named index. |
218
219
  | `getPersistencePaginationCursors()` | Indexed `(session_id, timestamp, id)`, `(run_id, sequence)`, `(run_id, recorded_at, id)` cursor shapes that avoid offset scans |
219
220
  | `assertPersistenceQueryPaginationConforms()` | Generic cursor pagination fixture for `queryEntries` |
220
221
  | `assertTenantScopedQueryIsolation()` | Tenant-filtered reads must not leak rows or primary-id collisions across tenants |
@@ -274,7 +275,8 @@ Recommended indexes for the reference schema. Hosts should add DB-specific parti
274
275
  | `prism_tool_calls` | `(session_id, name, started_at)` | tool usage by name |
275
276
  | `prism_tool_calls` | `(run_id, started_at)` | run tool-call listing |
276
277
  | `prism_tool_calls` | `(tool_call_id)` | deduplication / replay |
277
- | `prism_usage` | `(session_id, recorded_at)` | usage aggregation |
278
+ | `prism_usage` | `(session_id, recorded_at)` | usage pagination |
279
+ | `prism_usage` | `(session_id, scope, recorded_at)` | scope-safe billing/aggregate queries |
278
280
  | `prism_usage` | `(run_id, recorded_at)` | run usage |
279
281
  | `prism_agent_definitions` | `(name, version)` | definition lookup |
280
282
  | `prism_retention_policies` | `(tenant_id, account_id, user_id)` | policy listing |
@@ -344,7 +346,11 @@ Retention jobs should not run inside the agent/session runtime. They are a host
344
346
 
345
347
  ## Migrations
346
348
 
347
- Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library. Recommended migration practices:
349
+ Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library. First-party SQLite/PostgreSQL adapters automatically verify their checked-in schema-v3 history and catalog at open, before runtime writes. Their catalog reads are bounded metadata queries/PRAGMAs, not table-data scans.
350
+
351
+ Each new adapter-owned migration row records the contract SHA-256 checksum. A complete known v0.0.5 history whose checksum values are all `NULL` is a one-time compatibility case: under the SQLite transaction or PostgreSQL advisory transaction lock, the adapter verifies the full current shape, backfills all checksums, and continues. Unknown/duplicate/out-of-order/name-version/checksum mismatch, mixed/partial legacy values, or any missing/renamed/wrong-type/null/default/key/index artifact fails closed. Restore or apply a reviewed host migration; never edit checksums to silence drift.
352
+
353
+ Recommended migration practices:
348
354
 
349
355
  - Use a sequential or timestamped migration naming convention.
350
356
  - Store applied migrations in `prism_migrations` with `name`, `version`, `applied_at`, `applied_by`, and `checksum`.
@@ -422,18 +428,20 @@ const dbStore: ProductionPersistenceStore = {
422
428
  - `ProductionPersistenceStore` is an optional extension point. The runtime does not require it.
423
429
  - `ProductionPersistenceStore.checkpoints?: CheckpointStore` exposes generic versioned save/load/bounded-list/delete with compare-and-swap and fencing tokens, without workflow vocabulary.
424
430
  - `ProductionPersistenceStore.leases?: LeaseStore` exposes atomic acquire/renew/release/get with opaque claim tokens, expiries, ownership scope, and monotonic fencing tokens.
431
+ - `ProductionPersistenceStore.feedback?: RunFeedbackStore` exposes immutable append, bounded owned query, and owned deletion. First-party adapters store migration-003 rows in `prism_run_feedback`, FK-link `run_id`, and index owner/run/trace creation cursors.
425
432
  - Hosts choose the database, schema, transaction, and indexing strategy. The contracts specify query and checkpoint capability shapes.
426
433
  - `SessionStore` (`append`/`list`/`get`/optional `readBranchPath`) can be implemented on top of `ProductionPersistenceStore` or kept separate.
427
434
  - Cursor values and idempotency keys are host-defined and opaque to Prism.
428
- - First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume and multi-process coordination; workflow code owns no SQL table.
435
+ - First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume, human suspension, multi-process coordination, Phase 11 schedule records/fire leases, shared state, and replay lineage; workflow code owns no SQL table. `suspended`/`denied`, schedules, state history, and replay lineage remain namespaces/categories plus bounded checkpoint JSON values, so Phases 8 and 11 need no database migration.
429
436
 
430
437
  ## Security and performance notes
431
438
 
432
439
  - **No credentials in storage.** The contracts never include `CredentialResolver`, `AIProvider`, `ProviderResolver`, provider API keys, or credential values.
433
440
  - **Redact before storage.** Runtime session entries are redacted before `SessionStore.append`; `AgentEventRecord.event` and `ToolCallRecord.result` may contain secrets, so hosts must redact them (for example with `redactAgentEvent()` and a `SecretRedactor`) before writing to durable storage and set `redacted: true`.
434
- - **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility.
441
+ - **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility. First-party feedback stores require tenant plus account/user and use exact run ownership on append, query, and delete.
442
+ - **Feedback retention/deletion.** Comments, tags, and metadata are redacted and bounded before insert. `RunFeedbackStore.delete()` provides explicit erasure; deleting a run cascades its feedback in first-party SQL schemas. Apply host retention policy to `created_at`.
435
443
  - **Pagination and branch reads.** Every query supports `cursor`/`limit`/`order` so hosts can avoid full-table or full-session scans. Loading an entire large session into memory to serve a provider context is an anti-pattern; implement `readBranchPath` and use branch-relevant filters / recursive ancestor queries.
436
- - **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, and entry kind. See the reference indexes above.
444
+ - **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, entry kind, and feedback owner/run/trace creation cursors. See the reference indexes above.
437
445
 
438
446
  ## Related APIs
439
447
 
@@ -0,0 +1,122 @@
1
+ # Evaluations
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, live post-run scoring, and bounded batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
6
+
7
+ ## When to use it
8
+
9
+ Use this package when a host needs offline quality checks or sampled live scoring without coupling scorers into core agent execution. Install it directly or through `@arnilo/prism-all`; installation does not attach scorers to runs.
10
+
11
+ ## Inputs / request
12
+
13
+ | API | Key inputs |
14
+ | --- | --- |
15
+ | `defineScorer` | `id`, `score({ result, item?, expected?, signal? })` |
16
+ | `defineDataset` | `id`, `version?`, immutable `items[]` with unique ids |
17
+ | `scoreRun` / `scoreRunLive` | `AgentRunResult`, scorers, optional `sampleRate`, store, ownership, redactor |
18
+ | `runExperiment` | `agent`, dataset, scorers, bounded `concurrency`, optional store/ownership |
19
+ | `createMemoryEvaluationStore` | optional seed records |
20
+ | `appendEvaluationFeedback` | `RunFeedbackStore`, `EvaluationStore`, feedback fields, and 1–64 known evaluation IDs |
21
+
22
+ ## Outputs / response / events
23
+
24
+ | API | Output |
25
+ | --- | --- |
26
+ | `scoreRun` | `EvaluationRecord[]` with `scored` / `skipped` / `failed` |
27
+ | `scoreRunLive` | same records; never mutates the agent result; host may ignore the promise |
28
+ | `runExperiment` | `ExperimentReport` with stable item order, evaluations, and aggregates |
29
+ | `EvaluationStore.query` | cursor-paginated, ownership-filtered page |
30
+ | `appendEvaluationFeedback` | immutable `RunFeedbackRecord` containing only evaluation/scorer IDs |
31
+
32
+ ## Request/response example
33
+
34
+ ```json
35
+ {
36
+ "scorerId": "contains-citation",
37
+ "status": "scored",
38
+ "score": 1,
39
+ "runId": "run_1",
40
+ "sessionId": "session_1",
41
+ "experimentId": "exp_1",
42
+ "sampled": true
43
+ }
44
+ ```
45
+
46
+ ## Implementation example
47
+
48
+ ```ts
49
+ import { createAgent, createMemoryRunFeedbackStore, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
50
+ import {
51
+ appendEvaluationFeedback,
52
+ createMemoryEvaluationStore,
53
+ defineDataset,
54
+ defineScorer,
55
+ runExperiment,
56
+ scoreRunLive,
57
+ } from "@arnilo/prism-evals";
58
+
59
+ const scorer = defineScorer({
60
+ id: "contains-citation",
61
+ score: ({ result }) => ({ score: result.text.includes("[") ? 1 : 0 }),
62
+ });
63
+
64
+ const dataset = defineDataset({
65
+ id: "citations",
66
+ version: "1",
67
+ items: [{ id: "1", input: "Summarize with a citation" }],
68
+ });
69
+
70
+ const agent = createAgent({
71
+ model: { provider: "mock", model: "demo" },
72
+ provider: createMockProvider([providerTextDelta("ok [1]"), providerDone()]),
73
+ });
74
+
75
+ const store = createMemoryEvaluationStore();
76
+ const report = await runExperiment({
77
+ agent,
78
+ dataset,
79
+ scorers: [scorer],
80
+ concurrency: 2,
81
+ store,
82
+ ownership: { tenantId: "t1", userId: "u1" },
83
+ });
84
+
85
+ const result = await agent.createSession().run("Follow up");
86
+ void scoreRunLive(result, { scorers: [scorer], store });
87
+ const evaluation = report.evaluations[0]!;
88
+ const feedbackStore = createMemoryRunFeedbackStore({
89
+ resolveRun: ({ runId }) => runId === evaluation.runId
90
+ ? { runId, sessionId: evaluation.sessionId!, tenantId: "t1", userId: "u1" }
91
+ : false,
92
+ });
93
+ const linked = await appendEvaluationFeedback({
94
+ feedbackStore,
95
+ evaluationStore: store,
96
+ evaluationIds: [evaluation.id],
97
+ feedback: { id: "fb_1", runId: evaluation.runId!, rating: 1, tenantId: "t1", userId: "u1" },
98
+ });
99
+ console.log(report.aggregate.meanScore, linked.evaluationIds);
100
+ ```
101
+
102
+ ## Extension and configuration notes
103
+
104
+ - Function scorers are the base primitive. No mandatory LLM judge, dashboard, or schema library is included.
105
+ - Evaluation-result persistence remains package-local (`EvaluationStore`) and in-memory by default. Linked feedback is separately durable through optional `ProductionPersistenceStore.feedback`; evaluation score/reason payloads are not copied there.
106
+ - `sampleRate` is explicit (`0`–`1`). Inject `random` for deterministic tests.
107
+ - Dataset snapshots are frozen; duplicate item ids fail closed.
108
+ - `appendEvaluationFeedback()` resolves every supplied ID from `EvaluationStore`, rejects missing IDs, verifies each evaluation has the same run, optional trace, and exact ownership as feedback, then copies only deduplicated `evaluationIds`/`scorerIds`. Evaluation scores, reasons, errors, and metadata are not duplicated.
109
+
110
+ ## Security and performance notes
111
+
112
+ - Scorers receive result/item data only. Credentials, tools, and workspace access are not provided unless the host deliberately closes over them.
113
+ - Records pass through `SecretRedactor` / `secrets` before store append.
114
+ - Queries filter by ownership scope. Feedback linkage additionally requires tenant plus account/user and the feedback store re-verifies the run.
115
+ - Experiment concurrency defaults to `1` and is capped at `32`. Scoring can reference run IDs without duplicating unbounded event payloads.
116
+
117
+ ## Related APIs
118
+
119
+ - [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
120
+ - [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
121
+ - [Observability](observability.md): trace/run metadata hosts may copy into `traceId`
122
+ - [Release and install](release-and-install.md): optional package install
@@ -101,7 +101,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
101
101
  ## Extension and configuration notes
102
102
 
103
103
  - Extension loading is explicit. Prism does not discover packages, read manifests, or load filesystem config in the kernel.
104
- - `AgentConfig.extensions` is host-owned metadata for compatibility; `createAgent()` and `session.run()` do not load it or call `Extension.setup()`. Load extensions with `createExtensionKernel().load(...)`, then pass selected contributions (`tools`, `context`, `skills`, middleware, etc.) into `createAgent()`.
104
+ - Extensions stay host-owned outside `AgentConfig`. `createAgent()` and `session.run()` do not load extension lists or call `Extension.setup()`. Load extensions with `createExtensionKernel().load(...)`, then pass selected contributions (`tools`, `context`, `skills`, middleware, etc.) into `createAgent()`.
105
105
  - Setup order is the order provided by the host.
106
106
  - The kernel writes only to explicit registries returned by `createContributionRegistries()` or provided by the host.
107
107
  - `api.registerTool()` contributes an inert `ToolDefinition` to `registries.tools`; it does not add the tool to an active tool registry, allow list, or dispatch loop.
@@ -117,7 +117,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
117
117
  ## Security and performance notes
118
118
 
119
119
  - No hidden global extension kernel, provider registry, credential resolver, settings provider, store, or resource loader is created.
120
- - `AgentConfig.extensions` does not auto-execute, so constructing or running an agent cannot unexpectedly run extension code.
120
+ - Extension packages do not auto-execute from `createAgent()`, so constructing or running an agent cannot unexpectedly run extension code.
121
121
  - Error events use `ErrorInfo` and redact only known secret values passed in `secrets`.
122
122
  - Do not put resolved credential values in extension events, registry metadata, docs, logs, prompts, or session stores.
123
123
  - Event and middleware dispatch are ordered and dependency-free. They use no timers, background workers, filesystem discovery, network calls, provider calls, or tool execution.
@@ -26,9 +26,12 @@ Start from explicit host inputs. Do not let runtime code discover security state
26
26
  | Tool allow-list | active tools for this agent/session/run | `createToolRegistry`, `filterTools()`, `dispatchToolCall()` |
27
27
  | Tool argument rules | host validator | `AgentConfig.validator`, `RunOptions.validate`, `ToolValidator` |
28
28
  | Coding execution policy | path/command approval adapter | `ExecutionPolicy`, `@arnilo/prism-coding-security` |
29
+ | Remote media policy | public/default pinned DNS or explicit trusted transport | `SsrfPolicy`, `resolveMediaContentBlock()` |
29
30
  | Durable history | host database adapter | `SessionStore`, `assertSessionStoreConforms()` |
30
31
  | Durable audit | host ledger adapter | `RunLedger`, `redactRunLedgerRecord()` |
31
32
  | Extensions | explicit package imports only | `createExtensionKernel`, `ExtensionAPI` |
33
+ | Remote agent/workflow API | host authentication + ownership mapping | `@arnilo/prism-server`, `createPrismHandler()` |
34
+ | MCP server exposure | host MCP auth + selected capability list | `createPrismMcpServer()`, `createPrismMcpWebHandler()` |
32
35
 
33
36
  ## Outputs / response / events
34
37
 
@@ -105,7 +108,7 @@ Wire those values where they matter: provider adapters receive the resolved cred
105
108
 
106
109
  ## Extension and configuration notes
107
110
 
108
- - Keep security state explicit. `AgentConfig.settings` and `AgentConfig.credentials` are host-owned metadata; `createAgent()` and `session.run()` do not automatically call `settings.get()` or `credentials.resolve()`.
111
+ - Keep security state explicit. Settings and credentials are host-owned outside `AgentConfig`; `createAgent()` and `session.run()` do not automatically call `settings.get()` or `credentials.resolve()`.
109
112
  - Resolve credentials at the provider/request edge, as late as possible. Do not put resolved credentials in configs, manifests, registries, prompts, messages, events, session entries, run ledgers, idempotency keys, cache keys, or logs.
110
113
  - Use `createExplicitCredentialResolver()` to document source order such as runtime override → stored credential → caller-supplied env object → fallback.
111
114
  - Use `createEnvCredentialResolver()` only with an object the host passes in. Prism does not read `process.env` for credentials.
@@ -121,9 +124,15 @@ Wire those values where they matter: provider adapters receive the resolved cred
121
124
  - Prism does not sandbox host tools, extensions, provider adapters, credential resolvers, or custom middleware. Use OS/container/process isolation when code is untrusted.
122
125
  - Redaction is exact known-secret replacement only. It is not arbitrary secret detection, entropy scanning, or DLP.
123
126
  - Known secrets must be passed into redactors before data is emitted or persisted. Redact again in host adapters if they transform records after Prism redaction.
124
- - Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects.
125
- - MCP tools from `@arnilo/prism-mcp` are untrusted remote servers. Configure stdio commands and HTTP URLs explicitly; bound output with `maxResultBytes`; register prefixed tools only after trust review. See [MCP client bridge](mcp-tools.md).
126
- - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects. Use `@arnilo/prism-coding-security` for path roots, command rules, and approval caching. Prism does not provide OS sandboxing unless the host supplies a sandbox adapter.
127
+ - Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects. Its untrusted-schema adapter rejects non-local refs, forbidden keys/cycles/non-finite values and bounds bytes/depth/properties/keywords/refs plus its LRU cache before Ajv compilation; do not raise caps above documented hard limits.
128
+ - Treat embeddings as untrusted numeric input. `@arnilo/prism-memory` rejects empty, non-number, NaN, and infinite vectors before in-memory similarity or pgvector parameters; custom `Embedder`/`VectorStore` implementations must retain the same boundary.
129
+ - Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
130
+ - MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands, requires per-call `authorize`, and retains core gates. Its web handler still needs host `resolveAuthInfo`, TLS, edge rate limiting, and exact host/origin policy. See [MCP client/server exposure](mcp-tools.md).
131
+ - `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
132
+ - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Limits are not containment: Prism provides no OS sandbox unless the host supplies one.
133
+ - `@arnilo/prism-credentials-node` rejects oversized/malformed envelopes and excessive scrypt work before KDF allocation, uses async scrypt, and requires restrictive existing/new Unix vault modes. Keep vault ownership and parent-directory access host-controlled; review before `chmod 600`, never auto-weaken a file policy. Keychain calls use abort-aware native async work with finite timeout/payload caps and sanitized errors. OS prompts, service availability, and whether a native backend promptly honors cancellation remain host/platform boundaries; no plaintext fallback is attempted.
134
+ - LLM compaction always sends finite summary `maxTokens`, retains bounded deltas/events, and bounds/redacts provider/factory/policy error detail. Observational-memory workers cap turns, calls, arguments, results, transcript, and surfaced errors; unknown tools fail before execution, while invalid results can only be rejected after a host tool returns and may therefore follow side effects. Pass all known provider/credential/tool secrets into compaction/runtime options; exact replacement is not secret discovery.
135
+ - Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
127
136
  - Permission checks happen before tool validation and before `tool.execute()`. Middleware cannot grant permission by renaming a tool.
128
137
  - Session stores and ledgers receive redacted values when a redactor is active, but durable storage remains host-owned. Enforce tenant/account/user ownership and retention in the database layer.
129
138
  - Provider-owned auth/content/session/cache/security headers win over caller headers in adapters that merge headers.
@@ -138,10 +147,22 @@ Wire those values where they matter: provider adapters receive the resolved cred
138
147
  - Secret scan: source, tests, docs, workflow files, package metadata, built tests, packed-install canary, and tarball deny-list checks found no private-key block or common live-token prefix. Runtime redaction fixtures cover requests, events, ledgers, stores, checkpoints, provider/OAuth errors, and credential ciphertext.
139
148
  - Threat suites pass for parameterized SQL/tenant isolation, HTTP URL/SSRF rejection, realpath/symlink containment, shell-metacharacter approval, schema prototype-pollution/remote-reference bounds, OAuth polling/abort/redaction, credential tamper/wrong-key/KDF floors, MCP result bounds/timeouts, and coding approval/path policy.
140
149
 
141
- PostgreSQL TLS/network policy, MCP endpoint allow-listing, provider base URLs, OS keychain availability, process sandboxing, workflow tenant identity, and ANSI/control-sequence sanitization in any host terminal renderer remain host boundaries. Prism 0.0.4 ships JSON-line RPC, not an interactive TUI; hosts must render untrusted model/tool text safely. Credential-gated PostgreSQL/provider/keychain tests are separate operator/CI gates, not silently replaced by mocks.
150
+ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy beyond package origin/DNS pinning, provider base URLs, OS keychain availability, process sandboxing, workflow tenant identity, and ANSI/control-sequence sanitization in any host terminal renderer remain host boundaries. Prism 0.0.4 ships JSON-line RPC, not an interactive TUI; hosts must render untrusted model/tool text safely. Credential-gated PostgreSQL/provider/keychain tests are separate operator/CI gates, not silently replaced by mocks.
151
+
152
+ ## Supervisor and A2A boundaries
153
+
154
+ - Register children explicitly. AND-compose parent/child/hook permissions; never let a delegation hook replace broader parent policy.
155
+ - Build each child's context/memory with supervisor-provided `resourceId`/`threadId`; resolve provider credentials inside that child factory.
156
+ - Keep depth, active children, input, turn/tool/token, timeout, and queue ceilings finite; propagate abort through nested calls.
157
+ - Expose A2A only behind per-request authentication/authorization, TLS, edge rate limits, and replay policy. Public card discovery grants no invoke access.
158
+ - Remote A2A endpoints/card URLs require exact HTTPS origin allow-lists and redirect rejection. Pin ES256 card keys/expiry; never auto-fetch untrusted `jku`.
159
+ - Treat cards, task status, errors, artifacts, and SSE frames as untrusted bounded input and redact before logs/hooks/events. Client streaming uses one fatal UTF-8 decoder, bounded incremental LF/CRLF/multiline SSE parsing, and rejects partial/post-terminal frames; malformed remote bytes never become replacement characters in JSON.
142
160
 
143
161
  ## Related APIs
144
162
 
163
+ - [Web-standard server handler](server.md): remote agent/workflow route, ownership, limits, abort, and deployment boundary.
164
+ - [Supervisor delegation](supervisors.md): local child permission/memory/budget boundary.
165
+ - [A2A interoperability](a2a.md): remote card/auth/origin/signature boundary.
145
166
  - [Settings, auth, trust, and security controls](settings-auth-trust-security.md): low-level helpers and boundary hardening table.
146
167
  - [Credentials and redaction](credentials-and-redaction.md): credential resolver order, caller-supplied env objects, OAuth refresh, exact redaction, and no persistent secret store.
147
168
  - [Tools](tools.md): active tool registry, allow/deny filters, permission order, validator order, blocked events, and no sandbox.