@arnilo/prism 0.0.4 → 0.0.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +46 -1
- package/README.md +34 -10
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +26 -16
- package/dist/agents.js +147 -21
- package/dist/cli-init.d.ts +41 -0
- package/dist/cli-init.js +390 -0
- package/dist/cli-runner.d.ts +7 -1
- package/dist/cli-runner.js +13 -1
- package/dist/content.d.ts +19 -0
- package/dist/content.js +197 -69
- package/dist/contracts.d.ts +96 -9
- package/dist/contracts.js +8 -0
- package/dist/feedback.d.ts +48 -0
- package/dist/feedback.js +230 -0
- package/dist/ids.d.ts +2 -0
- package/dist/ids.js +6 -0
- package/dist/index.d.ts +10 -4
- package/dist/index.js +6 -3
- package/dist/providers/media.d.ts +3 -1
- package/dist/providers/media.js +11 -1
- package/dist/session-stores.js +2 -3
- package/dist/testing/feedback.d.ts +6 -0
- package/dist/testing/feedback.js +37 -0
- package/dist/testing/persistence-schema.d.ts +48 -10
- package/dist/testing/persistence-schema.js +166 -22
- package/dist/testing/run-ledger-conformance.js +7 -1
- package/dist/thinking.d.ts +42 -0
- package/dist/thinking.js +92 -0
- package/dist/tools.js +2 -3
- package/dist/use-case-model.d.ts +63 -0
- package/dist/use-case-model.js +52 -0
- package/docs/a2a.md +75 -0
- package/docs/agent-events.md +14 -21
- package/docs/agent-loops.md +12 -9
- package/docs/agent-session-runtime.md +14 -16
- package/docs/cli-rpc.md +35 -7
- package/docs/coding-agent-tools.md +35 -14
- package/docs/coding-security.md +7 -3
- package/docs/compaction-llm.md +17 -7
- package/docs/compaction-observational-memory.md +30 -4
- package/docs/context-and-skills.md +1 -0
- package/docs/credential-storage.md +58 -9
- package/docs/credentials-and-redaction.md +3 -3
- package/docs/database-persistence.md +17 -9
- package/docs/evaluations.md +122 -0
- package/docs/extensions.md +2 -2
- package/docs/host-security.md +26 -5
- package/docs/index.md +43 -28
- package/docs/mcp-tools.md +74 -13
- package/docs/migration.md +177 -3
- package/docs/multimodal-content.md +14 -6
- package/docs/node-filesystem-config.md +1 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/observability.md +14 -6
- package/docs/performance.md +209 -0
- package/docs/postgres-persistence.md +8 -6
- package/docs/provider-caching.md +16 -4
- package/docs/provider-conformance.md +40 -1
- package/docs/provider-packages.md +62 -3
- package/docs/providers/ai-sdk.md +149 -0
- package/docs/providers/kimi.md +124 -61
- package/docs/providers/neuralwatt.md +19 -13
- package/docs/providers/openai.md +56 -13
- package/docs/providers/opencode-go.md +118 -30
- package/docs/providers/openrouter.md +105 -35
- package/docs/providers/zai.md +94 -45
- package/docs/public-contracts.md +6 -5
- package/docs/rag.md +113 -0
- package/docs/release-and-install.md +100 -79
- package/docs/review-coverage-2026-07-15.md +193 -0
- package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
- package/docs/runs-and-usage.md +42 -5
- package/docs/server.md +139 -0
- package/docs/settings-auth-trust-security.md +5 -5
- package/docs/sqlite-persistence.md +6 -5
- package/docs/structured-output.md +1 -1
- package/docs/supervisors.md +71 -0
- package/docs/thinking-and-reasoning.md +98 -0
- package/docs/tool-execution-primitives.md +3 -3
- package/docs/tools.md +15 -0
- package/docs/use-case-model-selection.md +109 -0
- package/docs/workflow-orchestration-primitives.md +20 -3
- package/docs/workflows.md +114 -33
- package/docs/working-and-semantic-memory.md +170 -0
- package/package.json +13 -3
- package/templates/init/README.md.tmpl +28 -0
- package/templates/init/env.example.tmpl +1 -0
- package/templates/init/gitignore.tmpl +11 -0
- package/templates/init/optional/evals-example.ts.tmpl +17 -0
- package/templates/init/optional/workflows-example.ts.tmpl +27 -0
- package/templates/init/package.json.tmpl +22 -0
- package/templates/init/providers.json +76 -0
- package/templates/init/src/agent.ts.tmpl +10 -0
- package/templates/init/src/index.ts.tmpl +12 -0
- package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
- package/templates/init/tsconfig.json.tmpl +15 -0
package/docs/compaction-llm.md
CHANGED
|
@@ -25,15 +25,16 @@ Key exports:
|
|
|
25
25
|
| Field | Purpose |
|
|
26
26
|
| --- | --- |
|
|
27
27
|
| `provider` / `summaryProvider` | Explicit `AIProvider`, or factory receiving a resolved credential. |
|
|
28
|
-
| `model` / `summaryModel` |
|
|
28
|
+
| `model` / `summaryModel` | Summary `ModelConfig`. `summaryModel` wins; `model` is the session/host fallback via `resolveUseCaseModel`. See [Use-case model selection](use-case-model-selection.md). |
|
|
29
29
|
| `credential`, `credentialRequest` | Optional per-call credential resolution for provider factories. |
|
|
30
30
|
| `providerOptions` | Generic `ProviderRequest.options`, including cache fields. |
|
|
31
31
|
| `providerRequestPolicies` | Optional Prism provider request policies applied before the summary call. |
|
|
32
32
|
| `customInstructions` | Additional summary focus appended to prompts. |
|
|
33
|
-
| `thinkingLevel` |
|
|
34
|
-
| `reserveTokens` | Output budget basis; defaults to `16384`. |
|
|
33
|
+
| `thinkingLevel` | Mapped into `ProviderRequest.options.compat` via `applyThinkingLevel` / `thinkingFamilyForModel` (not inert `extra.thinkingLevel`). See [Thinking and reasoning](thinking-and-reasoning.md). |
|
|
34
|
+
| `reserveTokens` | Output budget basis; defaults to `16384`, hard cap `131072`. |
|
|
35
35
|
| `keepRecentTokens` | Approximate recent-token budget; defaults to `20000`. |
|
|
36
|
-
| `maxSummaryTokens` / `maxOutputTokens` |
|
|
36
|
+
| `maxSummaryTokens` / `maxOutputTokens` | Summary retention/request ceiling; default `16384`, hard cap `131072`. `maxSummaryTokens` wins over the compatibility alias. The finite value is written to `model.parameters.maxTokens`; first-party providers map it to their wire field. |
|
|
37
|
+
| `maxErrorBytes` | Retained provider/factory/policy error detail; default `1024`, hard cap `8192`, UTF-8-safe and known-secret redacted. |
|
|
37
38
|
| `maxToolResultChars` | Tool-result JSON truncation limit; defaults to `2000`. |
|
|
38
39
|
| `trackFileOperations`, `includeFileOperations` | Control file path extraction and final summary blocks. |
|
|
39
40
|
| `secrets` | Exact strings to redact from serialized prompts and final summaries. |
|
|
@@ -64,7 +65,8 @@ const strategy = createLlmCompactionStrategy({
|
|
|
64
65
|
model: { provider: "openai", model: "gpt-4.1-mini" },
|
|
65
66
|
keepRecentTokens: 20_000,
|
|
66
67
|
reserveTokens: 16_384,
|
|
67
|
-
|
|
68
|
+
maxSummaryTokens: 800,
|
|
69
|
+
maxErrorBytes: 1_024,
|
|
68
70
|
providerOptions: { cacheRetention: "short" },
|
|
69
71
|
customInstructions: "Focus on current files and failing tests.",
|
|
70
72
|
});
|
|
@@ -102,10 +104,18 @@ const agent = createAgent({ model, provider, compaction: { strategy, thresholdEn
|
|
|
102
104
|
Registration only contributes an inert strategy. The host must resolve and pass it to runtime config.
|
|
103
105
|
|
|
104
106
|
## Security and performance notes
|
|
105
|
-
Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization.
|
|
107
|
+
Preparation is O(n) over branch entries and uses only arrays, strings, and JSON serialization. Limit options must be positive safe integers at or below their hard caps and reject during strategy creation. Missing output options use a 16,384-token summary ceiling; reserve ratio/model metadata may narrow the provider request, never remove its finite `maxTokens`. A request policy that replaces `maxTokens` with NaN, Infinity, zero, an unsafe integer, or above-hard-cap input fails before provider generation.
|
|
108
|
+
|
|
109
|
+
Provider deltas are redacted while retained and stop at `maxSummaryTokens * 4` UTF-16 code units without splitting a surrogate pair. Provider iteration is closed/aborted on overflow. A derived finite event ceiling also stops endless empty/non-text deltas. Final history/turn/file composition receives the same cap. Provider error events, generator throws, provider-factory failures, and policy failures expose only bounded redacted detail; host abort remains authoritative.
|
|
110
|
+
|
|
111
|
+
The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. It does not discover credentials, read files, start background jobs, or add provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history or provider output.
|
|
106
112
|
|
|
107
113
|
## Related APIs
|
|
108
|
-
|
|
114
|
+
|
|
115
|
+
- [Use-case model selection](use-case-model-selection.md): `summaryModel` vs session `model` fallback.
|
|
116
|
+
- [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → `compat`.
|
|
117
|
+
- [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary and core compaction strategy surface.
|
|
118
|
+
- [Observational memory compaction package](compaction-observational-memory.md): source-backed memory workers with the same use-case binding pattern.
|
|
109
119
|
- [Agent/session runtime](agent-session-runtime.md): `AgentSession.compact()` and opt-in auto-compaction.
|
|
110
120
|
- [Provider layer](provider-layer.md): mock providers and provider request contracts.
|
|
111
121
|
- [Credentials and redaction](credentials-and-redaction.md): exact known-secret redaction behavior.
|
|
@@ -6,6 +6,8 @@
|
|
|
6
6
|
|
|
7
7
|
Current status: ledger/projection/render/recall utilities, explicit worker runtime, fast compaction strategy, inert extension helper, recall tool, and status/view command factories are available.
|
|
8
8
|
|
|
9
|
+
This package is distinct from `@arnilo/prism-memory` working/semantic memory: observational memory compresses and recalls source-backed observations/reflections; semantic memory retrieves embeddings; working memory stores the current structured profile/state. Hosts may compose both.
|
|
10
|
+
|
|
9
11
|
## When to use it
|
|
10
12
|
|
|
11
13
|
Use it when a host wants to opt in to long-session memory that records observations/reflections as session custom entries, renders prepared memory during compaction, and supports exact-id recall.
|
|
@@ -25,6 +27,20 @@ Memory records use `SessionEntry.kind: "custom"` with `entry.data.type` markers:
|
|
|
25
27
|
|
|
26
28
|
Ids are known, source-backed 12-character lowercase hex strings matching `^[a-f0-9]{12}$`.
|
|
27
29
|
|
|
30
|
+
Worker limits are finite positive safe integers:
|
|
31
|
+
|
|
32
|
+
| Runtime option | Default | Hard cap | Scope |
|
|
33
|
+
| --- | ---: | ---: | --- |
|
|
34
|
+
| `maxWorkerTurns` | 16 | 64 | Provider turns per observer/reflector/dropper run; overrides settings `agentMaxTurns` |
|
|
35
|
+
| `maxWorkerToolCallsPerTurn` | 32 | 256 | Calls retained from one provider response |
|
|
36
|
+
| `maxWorkerToolCalls` | 128 | 1,024 | Calls across all turns in one worker run |
|
|
37
|
+
| `maxWorkerArgumentBytes` | 64 KiB | 1 MiB | Each raw and redacted JSON argument object |
|
|
38
|
+
| `maxWorkerResultBytes` | 64 KiB | 1 MiB | Full tool result and replayed value/error payload |
|
|
39
|
+
| `maxWorkerMessageBytes` | 1 MiB | 8 MiB | System/prompt plus assistant-call/tool-result transcript |
|
|
40
|
+
| `maxWorkerErrorBytes` | 1 KiB | 8 KiB | Provider/tool/runtime error text after exact known-secret redaction |
|
|
41
|
+
|
|
42
|
+
Direct `runObserver()` / `runReflector()` / `runDropper()` calls retain required `maxTurns` and accept the corresponding shorter worker fields (`maxToolCalls`, `maxResultBytes`, etc.). Named default/hard constants and `resolveMemoryWorkerLimits()` are exported.
|
|
43
|
+
|
|
28
44
|
## Outputs / response / events
|
|
29
45
|
|
|
30
46
|
Key exports:
|
|
@@ -76,7 +92,12 @@ const memory = createObservationalMemoryRuntime({
|
|
|
76
92
|
session,
|
|
77
93
|
appendEntry: (entry) => store.append(entry),
|
|
78
94
|
workerProvider,
|
|
79
|
-
|
|
95
|
+
sessionModel: agent.config.model, // fallback when workerModel unset
|
|
96
|
+
// workerModel: { provider: "mock", model: "memory" }, // optional override
|
|
97
|
+
maxWorkerTurns: 8,
|
|
98
|
+
maxWorkerToolCalls: 64,
|
|
99
|
+
maxWorkerResultBytes: 64 * 1024,
|
|
100
|
+
overrides: { thinkingLevel: "low" },
|
|
80
101
|
});
|
|
81
102
|
await memory.flush();
|
|
82
103
|
await session.compact({ strategy: createObservationalMemoryCompactionStrategy({ keepRecentEntries: 8 }) });
|
|
@@ -90,9 +111,9 @@ await kernel.load([createObservationalMemoryExtension({ recallTool: { getEntries
|
|
|
90
111
|
|
|
91
112
|
## Extension and configuration notes
|
|
92
113
|
|
|
93
|
-
Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`.
|
|
114
|
+
Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`. `agentMaxTurns` now rejects non-integer/non-finite/out-of-range input (hard 64) instead of flooring/falling back. Runtime `maxWorkerTurns` takes precedence.
|
|
94
115
|
|
|
95
|
-
The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, `workerProvider
|
|
116
|
+
The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, and `workerProvider`. Model selection uses [use-case model selection](use-case-model-selection.md): pass optional `workerModel` (or settings `workerModel`) to override, and `sessionModel: agent.config.model` so workers fall back to the session model when no worker model is configured. `requireExplicitModel: true` restores the historical `missing_model` skip when no explicit worker model is set. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution. Default credential requests use the **resolved** model's provider id.
|
|
96
117
|
|
|
97
118
|
`createObservationalMemoryCompactionStrategy()` keeps recent message entries like the default compaction strategy, renders existing observations/reflections as the summary, and returns a standard Prism compaction entry. Its `data` includes `throughEntryId`, `keepEntryIds`, `strategy`, `trigger`, and `memory: { type: "om.folded", version: 1, fullFold, observations, reflections, droppedObservationIds }`. When active observations exceed `observationsPoolMaxTokens`, it performs a full fold into `data.memory`.
|
|
98
119
|
|
|
@@ -108,13 +129,18 @@ The runtime requires host-supplied `session`, an `appendEntry` callback bound to
|
|
|
108
129
|
- Recall tool and commands only see current-branch entries supplied by the host callback.
|
|
109
130
|
- Invalid or missing ids fail closed; invalid recall tool ids skip entry lookup.
|
|
110
131
|
- Utilities and fast compaction are O(n) over supplied entries and use no provider, network, filesystem, timer, worker, credential, or settings access.
|
|
111
|
-
- Workers serialize only supplied branch entries
|
|
132
|
+
- Workers serialize only supplied branch entries within `maxWorkerMessageBytes`, enforce finite turns/calls/arguments/results/messages/errors, and run one consolidation pipeline at a time per runtime. Source serialization and reflection/drop prompts fail before joining beyond the transcript cap.
|
|
133
|
+
- Every provider call must name a registered worker tool. Unknown calls, call overflow, oversized/deep/cyclic/non-JSON arguments/results, and transcript overflow fail deterministically; no excess call enters the assistant transcript or executes.
|
|
134
|
+
- Raw arguments are measured before tool execution. Full results are measured before redaction/replay; the bounded redacted value/error is then measured again because replacement text can grow. Replayed call arguments, tool values/errors, runtime `lastError`, and debug error data contain exact known-secret redaction. Host tools may already have caused side effects before returning an invalid oversized result; keep worker tools small/idempotent.
|
|
135
|
+
- Worker transcripts replay assistant `tool_call` messages before matching role `tool` `tool_result` messages so provider requests stay valid for call/result-pairing providers. Calls produced on the final allowed turn execute and persist, but no additional provider turn starts.
|
|
112
136
|
- Compaction preserves raw history; Prism appends one standard compaction entry and rebuilds provider context from its summary plus kept recent messages.
|
|
113
137
|
- Pass known secrets to render/recall/runtime/tool/command helpers to redact exact values from prompts, records, structured results, and text output.
|
|
114
138
|
- Live tests are opt-in with `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS=1`.
|
|
115
139
|
|
|
116
140
|
## Related APIs
|
|
117
141
|
|
|
142
|
+
- [Use-case model selection](use-case-model-selection.md): session vs worker model binding and `resolveUseCaseModel`.
|
|
143
|
+
- [Thinking and reasoning](thinking-and-reasoning.md): `thinkingLevel` → provider `compat`.
|
|
118
144
|
- [Compaction and retry policies](compaction-and-retry.md): replaceable compaction strategy boundary.
|
|
119
145
|
- [LLM compaction package](compaction-llm.md): existing optional compaction-package pattern.
|
|
120
146
|
- [Session stores and branching](session-stores-and-branching.md): branch entries that observational memory reads and appends to.
|
|
@@ -179,6 +179,7 @@ Use `activateAllCapabilities: true` only as a temporary all-skills/all-tools com
|
|
|
179
179
|
- [Agent/session runtime](agent-session-runtime.md): consumes host-selected context providers and skills from explicit agent config.
|
|
180
180
|
- [Input and prompt assembly](input-and-prompt-assembly.md): default prompt builder and provider-input assembly helper.
|
|
181
181
|
- [Instruction injection](instruction-injection.md): package injectors contribute `contextBlocks` that merge after host+skill provider blocks.
|
|
182
|
+
- [Retrieval-augmented generation](rag.md): optional retrieved citations contribute through the same explicit inert context seam.
|
|
182
183
|
- [Public contracts](public-contracts.md): `ContextProvider`, `ContextResolutionContext`, `ContextBlock`, `Skill`, `SkillRegistry`, `PromptBuilder`, and `PromptBuildRequest`.
|
|
183
184
|
- [Middleware hooks](middleware-hooks.md): `context` and `prompt_build` hooks.
|
|
184
185
|
- [Contribution registries](contribution-registries.md): inert context provider and skill contributions.
|
|
@@ -44,8 +44,11 @@ import {
|
|
|
44
44
|
| --- | --- | --- |
|
|
45
45
|
| `path` | `string` | Vault file path. Parent directories are created as needed. |
|
|
46
46
|
| `getPassphrase` | `() => string \| Promise<string>` | Host-owned passphrase retrieval. Never logged by the adapter. |
|
|
47
|
-
| `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32
|
|
48
|
-
| `fileMode` | `number` | Unix mode for
|
|
47
|
+
| `scrypt` | `{ N?, r?, p?, keyLength? }` | Optional KDF tuning. Defaults: `N=32768`, `r=8`, `p=1`, `keyLength=32`; limits are listed below. |
|
|
48
|
+
| `fileMode` | `number` | Unix mode for files. Defaults to `0o600`; group/other permissions are rejected. |
|
|
49
|
+
| `limits.maxFileBytes` | `number` | Encrypted envelope file: 4 MiB default, 16 MiB hard cap. |
|
|
50
|
+
| `limits.maxVaultBytes` | `number` | Decrypted vault/plaintext: 3 MiB default, 12 MiB hard cap. |
|
|
51
|
+
| `limits.maxScryptMemoryBytes` | `number` | `128*N*r` memory estimate: 256 MiB default and hard cap. |
|
|
49
52
|
|
|
50
53
|
### System keychain
|
|
51
54
|
|
|
@@ -53,7 +56,8 @@ import {
|
|
|
53
56
|
| --- | --- | --- |
|
|
54
57
|
| `service` | `string` | Keychain service name (application identifier). |
|
|
55
58
|
| `namespace` | `string` | Optional prefix separating environments or tenants within one service. |
|
|
56
|
-
| `timeoutMs` | `number` | Operation timeout. Defaults to
|
|
59
|
+
| `timeoutMs` | `number` | Operation timeout. Defaults to 5,000 ms; hard cap 60,000 ms. |
|
|
60
|
+
| `maxPayloadBytes` | `number` | Decrypted keychain payload: 3 MiB default, 12 MiB hard cap. |
|
|
57
61
|
|
|
58
62
|
## Outputs / response / events
|
|
59
63
|
|
|
@@ -71,6 +75,8 @@ Encrypted file stores also expose:
|
|
|
71
75
|
- `reload()` — re-read and decrypt from disk
|
|
72
76
|
- `flush()` — force rewrite of the encrypted envelope
|
|
73
77
|
|
|
78
|
+
`encryptBytes()` and `decryptBytes()` are Promise-based because they use asynchronous `node:crypto.scrypt`.
|
|
79
|
+
|
|
74
80
|
Errors are explicit and fail closed:
|
|
75
81
|
|
|
76
82
|
| Error | Code | When |
|
|
@@ -123,6 +129,7 @@ import {
|
|
|
123
129
|
const store = await openEncryptedCredentialStore({
|
|
124
130
|
path: "./credentials.vault",
|
|
125
131
|
getPassphrase: () => process.env.MY_APP_CREDENTIAL_PASSPHRASE!,
|
|
132
|
+
limits: { maxFileBytes: 4 * 1024 * 1024, maxVaultBytes: 3 * 1024 * 1024 },
|
|
126
133
|
});
|
|
127
134
|
|
|
128
135
|
const resolver = createExplicitCredentialResolver([
|
|
@@ -151,6 +158,47 @@ await rotateEncryptedCredentialStorePassphrase({
|
|
|
151
158
|
});
|
|
152
159
|
```
|
|
153
160
|
|
|
161
|
+
### Desktop keychain and explicit overrides
|
|
162
|
+
|
|
163
|
+
```ts
|
|
164
|
+
import {
|
|
165
|
+
createEnvCredentialResolver,
|
|
166
|
+
createExplicitCredentialResolver,
|
|
167
|
+
createMemoryCredentialStore,
|
|
168
|
+
} from "@arnilo/prism";
|
|
169
|
+
import {
|
|
170
|
+
createKeychainCredentialStore,
|
|
171
|
+
createStoredCredentialResolver,
|
|
172
|
+
} from "@arnilo/prism-credentials-node";
|
|
173
|
+
import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
|
|
174
|
+
|
|
175
|
+
const keychain = createKeychainCredentialStore({
|
|
176
|
+
service: "com.example.my-app",
|
|
177
|
+
namespace: "production",
|
|
178
|
+
});
|
|
179
|
+
await keychain.set({
|
|
180
|
+
name: "apiKey",
|
|
181
|
+
provider: "openai",
|
|
182
|
+
credential: { type: "api_key", value: userSuppliedKey },
|
|
183
|
+
});
|
|
184
|
+
|
|
185
|
+
const runtimeOverrides = createMemoryCredentialStore();
|
|
186
|
+
// Set only for this user/agent instance; it wins over stored and env values.
|
|
187
|
+
runtimeOverrides.set({
|
|
188
|
+
name: "apiKey",
|
|
189
|
+
provider: "openai",
|
|
190
|
+
credential: { type: "api_key", value: temporaryOverride },
|
|
191
|
+
});
|
|
192
|
+
|
|
193
|
+
const apiKey = createExplicitCredentialResolver([
|
|
194
|
+
{ name: "runtime", resolver: runtimeOverrides },
|
|
195
|
+
{ name: "keychain", resolver: createStoredCredentialResolver(keychain) },
|
|
196
|
+
{ name: "env", resolver: createEnvCredentialResolver(process.env, { openai: "OPENAI_API_KEY" }) },
|
|
197
|
+
]);
|
|
198
|
+
|
|
199
|
+
const providers = createOpenAIProviderPackage({ apiKey });
|
|
200
|
+
```
|
|
201
|
+
|
|
154
202
|
## Extension and configuration notes
|
|
155
203
|
|
|
156
204
|
- Passphrase retrieval, TLS, and OS permission prompts remain host-owned.
|
|
@@ -161,12 +209,13 @@ await rotateEncryptedCredentialStorePassphrase({
|
|
|
161
209
|
|
|
162
210
|
## Security and performance notes
|
|
163
211
|
|
|
164
|
-
- Authenticated encryption uses Node built-in `aes-256-gcm` and `scrypt`; no extra crypto
|
|
165
|
-
-
|
|
166
|
-
-
|
|
167
|
-
-
|
|
168
|
-
- Keychain operations
|
|
169
|
-
-
|
|
212
|
+
- Authenticated encryption uses Node built-in `aes-256-gcm` and asynchronous `scrypt`; no extra crypto dependency is added for the file backend.
|
|
213
|
+
- Envelope parsing rejects unknown shape, non-canonical/oversized base64, wrong salt/IV/tag size, unsupported algorithms/version, and excessive KDF work before scrypt. `N` must be a power of two from 16,384–262,144; `r≤32`, `p≤16`, `keyLength=32`, `N*r*p≤2,097,152`, and `128*N*r` must fit `maxScryptMemoryBytes`.
|
|
214
|
+
- Existing Unix vaults are checked before content read and must deny group/other access. Atomic writes create a random exclusive temp file at the requested restrictive mode, then rename; Windows skips Unix mode checks.
|
|
215
|
+
- Derived keys and package-owned plaintext buffers are zeroed after use. JavaScript passphrase strings and returned credentials remain host-owned.
|
|
216
|
+
- Keychain operations use `@napi-rs/keyring`'s abort-aware `AsyncEntry`, so native work runs outside the JavaScript event loop. A main-loop timer aborts and rejects at `timeoutMs`; native cancellation remains OS/backend-dependent and may briefly retain one libuv worker after rejection.
|
|
217
|
+
- Keychain payloads are bytes rather than password strings and are zeroed after parse/write. Unknown native errors are mapped to sanitized typed errors; no native message or secret value is echoed.
|
|
218
|
+
- Never log passphrases, derived keys, or decrypted credential payloads.
|
|
170
219
|
- Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
|
|
171
220
|
|
|
172
221
|
## Related APIs
|
|
@@ -95,7 +95,7 @@ console.log(error.message);
|
|
|
95
95
|
## Extension and configuration notes
|
|
96
96
|
|
|
97
97
|
- Hosts and extension packages can implement `CredentialResolver` and pass it explicitly to code that needs credentials.
|
|
98
|
-
-
|
|
98
|
+
- Credentials stay host-owned outside `AgentConfig`. `createAgent()` / `session.run()` do not call `credentials.resolve()`. Provider adapters, compaction workers, or request policies should receive and resolve credentials at the provider edge.
|
|
99
99
|
- Use `createExplicitCredentialResolver()` when documenting a fixed order such as runtime override, stored credential, caller-provided env object, then fallback resolver.
|
|
100
100
|
- Use `createEnvCredentialResolver()` only with an object supplied by the host; Prism does not read `process.env` for you.
|
|
101
101
|
- Provider adapters should resolve credentials as late as possible, per request.
|
|
@@ -107,10 +107,10 @@ console.log(error.message);
|
|
|
107
107
|
- Redaction only removes exact known secret values passed to the helper. It is not a general-purpose secret detector.
|
|
108
108
|
- Do not pass empty strings as secrets; they are ignored.
|
|
109
109
|
- `redactSecrets()` recursively walks arrays and object entries, so avoid using it on huge objects unless needed.
|
|
110
|
-
- Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via
|
|
110
|
+
- Cycle and non-JSON value handling: `redactSecrets()` is cycle-safe via an active-path `WeakSet`. Ancestor cycles render `"[Circular]"` at the back-reference instead of throwing; shared diamond references on separate branches stay structured (they are not collapsed to `"[Circular]"`). String object keys and `Map` keys are redacted like values, with deterministic `__2`/`__3`/… suffixes on collisions. `Date` and `RegExp` values are passed through unchanged; `ArrayBuffer` and typed arrays are passed through unchanged; `Map` is normalized to a plain object and `Set` to an array so the output stays JSON-compatible. `errorToErrorInfo()` tolerates a cyclic `error.cause` (rendered via `String()`). Symbols and non-enumerable properties remain outside JSON-shaped redaction.
|
|
111
111
|
- Use placeholders in tests and docs. Never commit real tokens.
|
|
112
112
|
- Live provider/worker tests are gated behind explicit environment variables and skipped by default: `PRISM_LIVE_PROVIDER_TESTS`, `PRISM_LIVE_COMPACTION_TESTS`, `PRISM_LIVE_OBSERVATIONAL_MEMORY_TESTS`. Default `npm test` is network-free; do not add ungated network calls to default tests.
|
|
113
|
-
-
|
|
113
|
+
- Credentials are not eagerly resolved by the core runtime, serialized into provider requests/events/stores, or passed to loops/compaction.
|
|
114
114
|
- `resolveCredentialValue()` and `createExplicitCredentialResolver()` do not cache values. Add host-side caching only if a real credential source needs it.
|
|
115
115
|
- `refreshOAuthCredential()` only calls the supplied OAuth provider and optional store; it has no built-in persistence or retry loop.
|
|
116
116
|
- OpenAI Codex device-code OAuth polls inside `createOpenAICodexOAuthProvider().login()` with bounded delays and abort support via `OAuthLoginCallbacks.signal`. Token-endpoint failures redact authorization codes, PKCE verifiers, device/user codes, and access/refresh tokens when those values are known.
|
|
@@ -67,7 +67,7 @@ Important shapes:
|
|
|
67
67
|
| `RunRecord` | Stored run with `sessionId`, `branchId`, status (`queued` \| `running` \| `succeeded` \| `failed` \| `aborted`), `model`, `provider`, `idempotencyKey`, `abortReason`, and `error`. |
|
|
68
68
|
| `AgentEventRecord` | Event ledger row with `event: AgentEvent` and a `redacted` flag. Hosts redact before storage. |
|
|
69
69
|
| `ToolCallRecord` | Tool-call row with `arguments`, optional `result: ToolResult`, `reason`, `progress` snapshots, status, and a `redacted` flag. |
|
|
70
|
-
| `UsageRecord` |
|
|
70
|
+
| `UsageRecord` | Scoped provider-turn or aggregate run usage with session/run/entry and turn/attempt linkage. |
|
|
71
71
|
| `AgentDefinitionRecord` | Versioned agent definition snapshot. Only stores `AgentDefinition` data; never provider credentials/resolvers/instances. |
|
|
72
72
|
| `RetentionPolicy` | Policy with `maxAgeDays`, `maxEntriesPerSession`, `maxTotalBytes`, `archiveStore`, and `appliedKinds`. |
|
|
73
73
|
| `MigrationRecord` | Applied migration with name, version, timestamp, checksum, and applied-by. |
|
|
@@ -128,7 +128,7 @@ A branch leaf (`leaf_entry_id`) is the current entry id for that branch. Rebuild
|
|
|
128
128
|
| --- | --- |
|
|
129
129
|
| `prism_session_entries` | `id` PK, `session_id` FK, `parent_id`, `run_id`, `timestamp`, `kind`, `schema_version`, `message` JSONB, `event` JSONB, `model` JSONB, `previous_model` JSONB, `label`, `summary`, `data` JSONB, `metadata` JSONB |
|
|
130
130
|
|
|
131
|
-
Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version`
|
|
131
|
+
Maps directly to `SessionEntry`. `kind` is one of the `SessionEntryKind` values. `schema_version` is nullable in schema version 3 for existing entry compatibility; hosts writing new rows should persist the entry's current schema version. `parent_id` may be null for the root entry of a session.
|
|
132
132
|
|
|
133
133
|
`SessionAppendOptions.idempotencyKey` is not part of `SessionEntry`; store it in an adapter-owned side table when you need durable retry detection:
|
|
134
134
|
|
|
@@ -166,7 +166,7 @@ The `event` JSONB stores a redacted `AgentEvent`. The `sequence` column is an im
|
|
|
166
166
|
|
|
167
167
|
| Table | Key columns |
|
|
168
168
|
| --- | --- |
|
|
169
|
-
| `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
|
|
169
|
+
| `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `scope`, `turn`, `attempt`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
|
|
170
170
|
|
|
171
171
|
The `usage` JSONB stores the `Usage` shape: input/output/total/cache tokens, cost, and currency.
|
|
172
172
|
|
|
@@ -214,7 +214,8 @@ await runRunLedgerConformance(() => createLedger(testDatabase), { exerciseReopen
|
|
|
214
214
|
| Primitive | Purpose |
|
|
215
215
|
| --- | --- |
|
|
216
216
|
| `PersistenceSchemaModel` | Versioned table/column/index model covering sessions, entries, parent chain, idempotency side table, runs, events, tool calls, usage, tenant columns, and `prism_migrations` |
|
|
217
|
-
| `createPersistenceMigrationContract()` | Strictly increasing migration steps, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
|
|
217
|
+
| `createPersistenceMigrationContract()` | Strictly increasing checked-in migration steps with deterministic SHA-256 checksums, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
|
|
218
|
+
| `assertAppliedPersistenceMigrations()` / `assertPersistenceSchemaShape()` | Testable fail-closed history and normalized SQLite/PostgreSQL catalog checks for every required table, column/type/null/default, PK/unique/FK, and named index. |
|
|
218
219
|
| `getPersistencePaginationCursors()` | Indexed `(session_id, timestamp, id)`, `(run_id, sequence)`, `(run_id, recorded_at, id)` cursor shapes that avoid offset scans |
|
|
219
220
|
| `assertPersistenceQueryPaginationConforms()` | Generic cursor pagination fixture for `queryEntries` |
|
|
220
221
|
| `assertTenantScopedQueryIsolation()` | Tenant-filtered reads must not leak rows or primary-id collisions across tenants |
|
|
@@ -274,7 +275,8 @@ Recommended indexes for the reference schema. Hosts should add DB-specific parti
|
|
|
274
275
|
| `prism_tool_calls` | `(session_id, name, started_at)` | tool usage by name |
|
|
275
276
|
| `prism_tool_calls` | `(run_id, started_at)` | run tool-call listing |
|
|
276
277
|
| `prism_tool_calls` | `(tool_call_id)` | deduplication / replay |
|
|
277
|
-
| `prism_usage` | `(session_id, recorded_at)` | usage
|
|
278
|
+
| `prism_usage` | `(session_id, recorded_at)` | usage pagination |
|
|
279
|
+
| `prism_usage` | `(session_id, scope, recorded_at)` | scope-safe billing/aggregate queries |
|
|
278
280
|
| `prism_usage` | `(run_id, recorded_at)` | run usage |
|
|
279
281
|
| `prism_agent_definitions` | `(name, version)` | definition lookup |
|
|
280
282
|
| `prism_retention_policies` | `(tenant_id, account_id, user_id)` | policy listing |
|
|
@@ -344,7 +346,11 @@ Retention jobs should not run inside the agent/session runtime. They are a host
|
|
|
344
346
|
|
|
345
347
|
## Migrations
|
|
346
348
|
|
|
347
|
-
Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library.
|
|
349
|
+
Hosts own schema migrations. Prism publishes only the TypeScript contracts; no DDL is generated or executed by the core library. First-party SQLite/PostgreSQL adapters automatically verify their checked-in schema-v3 history and catalog at open, before runtime writes. Their catalog reads are bounded metadata queries/PRAGMAs, not table-data scans.
|
|
350
|
+
|
|
351
|
+
Each new adapter-owned migration row records the contract SHA-256 checksum. A complete known v0.0.5 history whose checksum values are all `NULL` is a one-time compatibility case: under the SQLite transaction or PostgreSQL advisory transaction lock, the adapter verifies the full current shape, backfills all checksums, and continues. Unknown/duplicate/out-of-order/name-version/checksum mismatch, mixed/partial legacy values, or any missing/renamed/wrong-type/null/default/key/index artifact fails closed. Restore or apply a reviewed host migration; never edit checksums to silence drift.
|
|
352
|
+
|
|
353
|
+
Recommended migration practices:
|
|
348
354
|
|
|
349
355
|
- Use a sequential or timestamped migration naming convention.
|
|
350
356
|
- Store applied migrations in `prism_migrations` with `name`, `version`, `applied_at`, `applied_by`, and `checksum`.
|
|
@@ -422,18 +428,20 @@ const dbStore: ProductionPersistenceStore = {
|
|
|
422
428
|
- `ProductionPersistenceStore` is an optional extension point. The runtime does not require it.
|
|
423
429
|
- `ProductionPersistenceStore.checkpoints?: CheckpointStore` exposes generic versioned save/load/bounded-list/delete with compare-and-swap and fencing tokens, without workflow vocabulary.
|
|
424
430
|
- `ProductionPersistenceStore.leases?: LeaseStore` exposes atomic acquire/renew/release/get with opaque claim tokens, expiries, ownership scope, and monotonic fencing tokens.
|
|
431
|
+
- `ProductionPersistenceStore.feedback?: RunFeedbackStore` exposes immutable append, bounded owned query, and owned deletion. First-party adapters store migration-003 rows in `prism_run_feedback`, FK-link `run_id`, and index owner/run/trace creation cursors.
|
|
425
432
|
- Hosts choose the database, schema, transaction, and indexing strategy. The contracts specify query and checkpoint capability shapes.
|
|
426
433
|
- `SessionStore` (`append`/`list`/`get`/optional `readBranchPath`) can be implemented on top of `ProductionPersistenceStore` or kept separate.
|
|
427
434
|
- Cursor values and idempotency keys are host-defined and opaque to Prism.
|
|
428
|
-
- First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume
|
|
435
|
+
- First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume, human suspension, multi-process coordination, Phase 11 schedule records/fire leases, shared state, and replay lineage; workflow code owns no SQL table. `suspended`/`denied`, schedules, state history, and replay lineage remain namespaces/categories plus bounded checkpoint JSON values, so Phases 8 and 11 need no database migration.
|
|
429
436
|
|
|
430
437
|
## Security and performance notes
|
|
431
438
|
|
|
432
439
|
- **No credentials in storage.** The contracts never include `CredentialResolver`, `AIProvider`, `ProviderResolver`, provider API keys, or credential values.
|
|
433
440
|
- **Redact before storage.** Runtime session entries are redacted before `SessionStore.append`; `AgentEventRecord.event` and `ToolCallRecord.result` may contain secrets, so hosts must redact them (for example with `redactAgentEvent()` and a `SecretRedactor`) before writing to durable storage and set `redacted: true`.
|
|
434
|
-
- **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility.
|
|
441
|
+
- **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility. First-party feedback stores require tenant plus account/user and use exact run ownership on append, query, and delete.
|
|
442
|
+
- **Feedback retention/deletion.** Comments, tags, and metadata are redacted and bounded before insert. `RunFeedbackStore.delete()` provides explicit erasure; deleting a run cascades its feedback in first-party SQL schemas. Apply host retention policy to `created_at`.
|
|
435
443
|
- **Pagination and branch reads.** Every query supports `cursor`/`limit`/`order` so hosts can avoid full-table or full-session scans. Loading an entire large session into memory to serve a provider context is an anti-pattern; implement `readBranchPath` and use branch-relevant filters / recursive ancestor queries.
|
|
436
|
-
- **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, and
|
|
444
|
+
- **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, entry kind, and feedback owner/run/trace creation cursors. See the reference indexes above.
|
|
437
445
|
|
|
438
446
|
## Related APIs
|
|
439
447
|
|
|
@@ -0,0 +1,122 @@
|
|
|
1
|
+
# Evaluations
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, live post-run scoring, and bounded batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
Use this package when a host needs offline quality checks or sampled live scoring without coupling scorers into core agent execution. Install it directly or through `@arnilo/prism-all`; installation does not attach scorers to runs.
|
|
10
|
+
|
|
11
|
+
## Inputs / request
|
|
12
|
+
|
|
13
|
+
| API | Key inputs |
|
|
14
|
+
| --- | --- |
|
|
15
|
+
| `defineScorer` | `id`, `score({ result, item?, expected?, signal? })` |
|
|
16
|
+
| `defineDataset` | `id`, `version?`, immutable `items[]` with unique ids |
|
|
17
|
+
| `scoreRun` / `scoreRunLive` | `AgentRunResult`, scorers, optional `sampleRate`, store, ownership, redactor |
|
|
18
|
+
| `runExperiment` | `agent`, dataset, scorers, bounded `concurrency`, optional store/ownership |
|
|
19
|
+
| `createMemoryEvaluationStore` | optional seed records |
|
|
20
|
+
| `appendEvaluationFeedback` | `RunFeedbackStore`, `EvaluationStore`, feedback fields, and 1–64 known evaluation IDs |
|
|
21
|
+
|
|
22
|
+
## Outputs / response / events
|
|
23
|
+
|
|
24
|
+
| API | Output |
|
|
25
|
+
| --- | --- |
|
|
26
|
+
| `scoreRun` | `EvaluationRecord[]` with `scored` / `skipped` / `failed` |
|
|
27
|
+
| `scoreRunLive` | same records; never mutates the agent result; host may ignore the promise |
|
|
28
|
+
| `runExperiment` | `ExperimentReport` with stable item order, evaluations, and aggregates |
|
|
29
|
+
| `EvaluationStore.query` | cursor-paginated, ownership-filtered page |
|
|
30
|
+
| `appendEvaluationFeedback` | immutable `RunFeedbackRecord` containing only evaluation/scorer IDs |
|
|
31
|
+
|
|
32
|
+
## Request/response example
|
|
33
|
+
|
|
34
|
+
```json
|
|
35
|
+
{
|
|
36
|
+
"scorerId": "contains-citation",
|
|
37
|
+
"status": "scored",
|
|
38
|
+
"score": 1,
|
|
39
|
+
"runId": "run_1",
|
|
40
|
+
"sessionId": "session_1",
|
|
41
|
+
"experimentId": "exp_1",
|
|
42
|
+
"sampled": true
|
|
43
|
+
}
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
## Implementation example
|
|
47
|
+
|
|
48
|
+
```ts
|
|
49
|
+
import { createAgent, createMemoryRunFeedbackStore, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
|
|
50
|
+
import {
|
|
51
|
+
appendEvaluationFeedback,
|
|
52
|
+
createMemoryEvaluationStore,
|
|
53
|
+
defineDataset,
|
|
54
|
+
defineScorer,
|
|
55
|
+
runExperiment,
|
|
56
|
+
scoreRunLive,
|
|
57
|
+
} from "@arnilo/prism-evals";
|
|
58
|
+
|
|
59
|
+
const scorer = defineScorer({
|
|
60
|
+
id: "contains-citation",
|
|
61
|
+
score: ({ result }) => ({ score: result.text.includes("[") ? 1 : 0 }),
|
|
62
|
+
});
|
|
63
|
+
|
|
64
|
+
const dataset = defineDataset({
|
|
65
|
+
id: "citations",
|
|
66
|
+
version: "1",
|
|
67
|
+
items: [{ id: "1", input: "Summarize with a citation" }],
|
|
68
|
+
});
|
|
69
|
+
|
|
70
|
+
const agent = createAgent({
|
|
71
|
+
model: { provider: "mock", model: "demo" },
|
|
72
|
+
provider: createMockProvider([providerTextDelta("ok [1]"), providerDone()]),
|
|
73
|
+
});
|
|
74
|
+
|
|
75
|
+
const store = createMemoryEvaluationStore();
|
|
76
|
+
const report = await runExperiment({
|
|
77
|
+
agent,
|
|
78
|
+
dataset,
|
|
79
|
+
scorers: [scorer],
|
|
80
|
+
concurrency: 2,
|
|
81
|
+
store,
|
|
82
|
+
ownership: { tenantId: "t1", userId: "u1" },
|
|
83
|
+
});
|
|
84
|
+
|
|
85
|
+
const result = await agent.createSession().run("Follow up");
|
|
86
|
+
void scoreRunLive(result, { scorers: [scorer], store });
|
|
87
|
+
const evaluation = report.evaluations[0]!;
|
|
88
|
+
const feedbackStore = createMemoryRunFeedbackStore({
|
|
89
|
+
resolveRun: ({ runId }) => runId === evaluation.runId
|
|
90
|
+
? { runId, sessionId: evaluation.sessionId!, tenantId: "t1", userId: "u1" }
|
|
91
|
+
: false,
|
|
92
|
+
});
|
|
93
|
+
const linked = await appendEvaluationFeedback({
|
|
94
|
+
feedbackStore,
|
|
95
|
+
evaluationStore: store,
|
|
96
|
+
evaluationIds: [evaluation.id],
|
|
97
|
+
feedback: { id: "fb_1", runId: evaluation.runId!, rating: 1, tenantId: "t1", userId: "u1" },
|
|
98
|
+
});
|
|
99
|
+
console.log(report.aggregate.meanScore, linked.evaluationIds);
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
## Extension and configuration notes
|
|
103
|
+
|
|
104
|
+
- Function scorers are the base primitive. No mandatory LLM judge, dashboard, or schema library is included.
|
|
105
|
+
- Evaluation-result persistence remains package-local (`EvaluationStore`) and in-memory by default. Linked feedback is separately durable through optional `ProductionPersistenceStore.feedback`; evaluation score/reason payloads are not copied there.
|
|
106
|
+
- `sampleRate` is explicit (`0`–`1`). Inject `random` for deterministic tests.
|
|
107
|
+
- Dataset snapshots are frozen; duplicate item ids fail closed.
|
|
108
|
+
- `appendEvaluationFeedback()` resolves every supplied ID from `EvaluationStore`, rejects missing IDs, verifies each evaluation has the same run, optional trace, and exact ownership as feedback, then copies only deduplicated `evaluationIds`/`scorerIds`. Evaluation scores, reasons, errors, and metadata are not duplicated.
|
|
109
|
+
|
|
110
|
+
## Security and performance notes
|
|
111
|
+
|
|
112
|
+
- Scorers receive result/item data only. Credentials, tools, and workspace access are not provided unless the host deliberately closes over them.
|
|
113
|
+
- Records pass through `SecretRedactor` / `secrets` before store append.
|
|
114
|
+
- Queries filter by ownership scope. Feedback linkage additionally requires tenant plus account/user and the feedback store re-verifies the run.
|
|
115
|
+
- Experiment concurrency defaults to `1` and is capped at `32`. Scoring can reference run IDs without duplicating unbounded event payloads.
|
|
116
|
+
|
|
117
|
+
## Related APIs
|
|
118
|
+
|
|
119
|
+
- [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
|
|
120
|
+
- [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
|
|
121
|
+
- [Observability](observability.md): trace/run metadata hosts may copy into `traceId`
|
|
122
|
+
- [Release and install](release-and-install.md): optional package install
|
package/docs/extensions.md
CHANGED
|
@@ -101,7 +101,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
|
|
|
101
101
|
## Extension and configuration notes
|
|
102
102
|
|
|
103
103
|
- Extension loading is explicit. Prism does not discover packages, read manifests, or load filesystem config in the kernel.
|
|
104
|
-
-
|
|
104
|
+
- Extensions stay host-owned outside `AgentConfig`. `createAgent()` and `session.run()` do not load extension lists or call `Extension.setup()`. Load extensions with `createExtensionKernel().load(...)`, then pass selected contributions (`tools`, `context`, `skills`, middleware, etc.) into `createAgent()`.
|
|
105
105
|
- Setup order is the order provided by the host.
|
|
106
106
|
- The kernel writes only to explicit registries returned by `createContributionRegistries()` or provided by the host.
|
|
107
107
|
- `api.registerTool()` contributes an inert `ToolDefinition` to `registries.tools`; it does not add the tool to an active tool registry, allow list, or dispatch loop.
|
|
@@ -117,7 +117,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
|
|
|
117
117
|
## Security and performance notes
|
|
118
118
|
|
|
119
119
|
- No hidden global extension kernel, provider registry, credential resolver, settings provider, store, or resource loader is created.
|
|
120
|
-
-
|
|
120
|
+
- Extension packages do not auto-execute from `createAgent()`, so constructing or running an agent cannot unexpectedly run extension code.
|
|
121
121
|
- Error events use `ErrorInfo` and redact only known secret values passed in `secrets`.
|
|
122
122
|
- Do not put resolved credential values in extension events, registry metadata, docs, logs, prompts, or session stores.
|
|
123
123
|
- Event and middleware dispatch are ordered and dependency-free. They use no timers, background workers, filesystem discovery, network calls, provider calls, or tool execution.
|
package/docs/host-security.md
CHANGED
|
@@ -26,9 +26,12 @@ Start from explicit host inputs. Do not let runtime code discover security state
|
|
|
26
26
|
| Tool allow-list | active tools for this agent/session/run | `createToolRegistry`, `filterTools()`, `dispatchToolCall()` |
|
|
27
27
|
| Tool argument rules | host validator | `AgentConfig.validator`, `RunOptions.validate`, `ToolValidator` |
|
|
28
28
|
| Coding execution policy | path/command approval adapter | `ExecutionPolicy`, `@arnilo/prism-coding-security` |
|
|
29
|
+
| Remote media policy | public/default pinned DNS or explicit trusted transport | `SsrfPolicy`, `resolveMediaContentBlock()` |
|
|
29
30
|
| Durable history | host database adapter | `SessionStore`, `assertSessionStoreConforms()` |
|
|
30
31
|
| Durable audit | host ledger adapter | `RunLedger`, `redactRunLedgerRecord()` |
|
|
31
32
|
| Extensions | explicit package imports only | `createExtensionKernel`, `ExtensionAPI` |
|
|
33
|
+
| Remote agent/workflow API | host authentication + ownership mapping | `@arnilo/prism-server`, `createPrismHandler()` |
|
|
34
|
+
| MCP server exposure | host MCP auth + selected capability list | `createPrismMcpServer()`, `createPrismMcpWebHandler()` |
|
|
32
35
|
|
|
33
36
|
## Outputs / response / events
|
|
34
37
|
|
|
@@ -105,7 +108,7 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
105
108
|
|
|
106
109
|
## Extension and configuration notes
|
|
107
110
|
|
|
108
|
-
- Keep security state explicit.
|
|
111
|
+
- Keep security state explicit. Settings and credentials are host-owned outside `AgentConfig`; `createAgent()` and `session.run()` do not automatically call `settings.get()` or `credentials.resolve()`.
|
|
109
112
|
- Resolve credentials at the provider/request edge, as late as possible. Do not put resolved credentials in configs, manifests, registries, prompts, messages, events, session entries, run ledgers, idempotency keys, cache keys, or logs.
|
|
110
113
|
- Use `createExplicitCredentialResolver()` to document source order such as runtime override → stored credential → caller-supplied env object → fallback.
|
|
111
114
|
- Use `createEnvCredentialResolver()` only with an object the host passes in. Prism does not read `process.env` for credentials.
|
|
@@ -121,9 +124,15 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
121
124
|
- Prism does not sandbox host tools, extensions, provider adapters, credential resolvers, or custom middleware. Use OS/container/process isolation when code is untrusted.
|
|
122
125
|
- Redaction is exact known-secret replacement only. It is not arbitrary secret detection, entropy scanning, or DLP.
|
|
123
126
|
- Known secrets must be passed into redactors before data is emitted or persisted. Redact again in host adapters if they transform records after Prism redaction.
|
|
124
|
-
- Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects.
|
|
125
|
-
-
|
|
126
|
-
-
|
|
127
|
+
- Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects. Its untrusted-schema adapter rejects non-local refs, forbidden keys/cycles/non-finite values and bounds bytes/depth/properties/keywords/refs plus its LRU cache before Ajv compilation; do not raise caps above documented hard limits.
|
|
128
|
+
- Treat embeddings as untrusted numeric input. `@arnilo/prism-memory` rejects empty, non-number, NaN, and infinite vectors before in-memory similarity or pgvector parameters; custom `Embedder`/`VectorStore` implementations must retain the same boundary.
|
|
129
|
+
- Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
|
|
130
|
+
- MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands, requires per-call `authorize`, and retains core gates. Its web handler still needs host `resolveAuthInfo`, TLS, edge rate limiting, and exact host/origin policy. See [MCP client/server exposure](mcp-tools.md).
|
|
131
|
+
- `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
|
|
132
|
+
- Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Limits are not containment: Prism provides no OS sandbox unless the host supplies one.
|
|
133
|
+
- `@arnilo/prism-credentials-node` rejects oversized/malformed envelopes and excessive scrypt work before KDF allocation, uses async scrypt, and requires restrictive existing/new Unix vault modes. Keep vault ownership and parent-directory access host-controlled; review before `chmod 600`, never auto-weaken a file policy. Keychain calls use abort-aware native async work with finite timeout/payload caps and sanitized errors. OS prompts, service availability, and whether a native backend promptly honors cancellation remain host/platform boundaries; no plaintext fallback is attempted.
|
|
134
|
+
- LLM compaction always sends finite summary `maxTokens`, retains bounded deltas/events, and bounds/redacts provider/factory/policy error detail. Observational-memory workers cap turns, calls, arguments, results, transcript, and surfaced errors; unknown tools fail before execution, while invalid results can only be rejected after a host tool returns and may therefore follow side effects. Pass all known provider/credential/tool secrets into compaction/runtime options; exact replacement is not secret discovery.
|
|
135
|
+
- Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
|
|
127
136
|
- Permission checks happen before tool validation and before `tool.execute()`. Middleware cannot grant permission by renaming a tool.
|
|
128
137
|
- Session stores and ledgers receive redacted values when a redactor is active, but durable storage remains host-owned. Enforce tenant/account/user ownership and retention in the database layer.
|
|
129
138
|
- Provider-owned auth/content/session/cache/security headers win over caller headers in adapters that merge headers.
|
|
@@ -138,10 +147,22 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
138
147
|
- Secret scan: source, tests, docs, workflow files, package metadata, built tests, packed-install canary, and tarball deny-list checks found no private-key block or common live-token prefix. Runtime redaction fixtures cover requests, events, ledgers, stores, checkpoints, provider/OAuth errors, and credential ciphertext.
|
|
139
148
|
- Threat suites pass for parameterized SQL/tenant isolation, HTTP URL/SSRF rejection, realpath/symlink containment, shell-metacharacter approval, schema prototype-pollution/remote-reference bounds, OAuth polling/abort/redaction, credential tamper/wrong-key/KDF floors, MCP result bounds/timeouts, and coding approval/path policy.
|
|
140
149
|
|
|
141
|
-
PostgreSQL TLS/network policy, MCP endpoint
|
|
150
|
+
PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy beyond package origin/DNS pinning, provider base URLs, OS keychain availability, process sandboxing, workflow tenant identity, and ANSI/control-sequence sanitization in any host terminal renderer remain host boundaries. Prism 0.0.4 ships JSON-line RPC, not an interactive TUI; hosts must render untrusted model/tool text safely. Credential-gated PostgreSQL/provider/keychain tests are separate operator/CI gates, not silently replaced by mocks.
|
|
151
|
+
|
|
152
|
+
## Supervisor and A2A boundaries
|
|
153
|
+
|
|
154
|
+
- Register children explicitly. AND-compose parent/child/hook permissions; never let a delegation hook replace broader parent policy.
|
|
155
|
+
- Build each child's context/memory with supervisor-provided `resourceId`/`threadId`; resolve provider credentials inside that child factory.
|
|
156
|
+
- Keep depth, active children, input, turn/tool/token, timeout, and queue ceilings finite; propagate abort through nested calls.
|
|
157
|
+
- Expose A2A only behind per-request authentication/authorization, TLS, edge rate limits, and replay policy. Public card discovery grants no invoke access.
|
|
158
|
+
- Remote A2A endpoints/card URLs require exact HTTPS origin allow-lists and redirect rejection. Pin ES256 card keys/expiry; never auto-fetch untrusted `jku`.
|
|
159
|
+
- Treat cards, task status, errors, artifacts, and SSE frames as untrusted bounded input and redact before logs/hooks/events. Client streaming uses one fatal UTF-8 decoder, bounded incremental LF/CRLF/multiline SSE parsing, and rejects partial/post-terminal frames; malformed remote bytes never become replacement characters in JSON.
|
|
142
160
|
|
|
143
161
|
## Related APIs
|
|
144
162
|
|
|
163
|
+
- [Web-standard server handler](server.md): remote agent/workflow route, ownership, limits, abort, and deployment boundary.
|
|
164
|
+
- [Supervisor delegation](supervisors.md): local child permission/memory/budget boundary.
|
|
165
|
+
- [A2A interoperability](a2a.md): remote card/auth/origin/signature boundary.
|
|
145
166
|
- [Settings, auth, trust, and security controls](settings-auth-trust-security.md): low-level helpers and boundary hardening table.
|
|
146
167
|
- [Credentials and redaction](credentials-and-redaction.md): credential resolver order, caller-supplied env objects, OAuth refresh, exact redaction, and no persistent secret store.
|
|
147
168
|
- [Tools](tools.md): active tool registry, allow/deny filters, permission order, validator order, blocked events, and no sandbox.
|