@arnilo/prism 0.0.3 → 0.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/CHANGELOG.md +40 -0
  2. package/README.md +62 -26
  3. package/dist/agent-loops.d.ts +8 -1
  4. package/dist/agent-loops.js +57 -11
  5. package/dist/agents.js +212 -32
  6. package/dist/checkpoints.d.ts +11 -0
  7. package/dist/checkpoints.js +144 -0
  8. package/dist/cli-init.d.ts +41 -0
  9. package/dist/cli-init.js +390 -0
  10. package/dist/cli-runner.d.ts +7 -1
  11. package/dist/cli-runner.js +13 -1
  12. package/dist/compaction.js +9 -1
  13. package/dist/content.d.ts +121 -0
  14. package/dist/content.js +538 -0
  15. package/dist/contracts.d.ts +236 -11
  16. package/dist/contracts.js +8 -0
  17. package/dist/event-multiplexer.d.ts +23 -0
  18. package/dist/event-multiplexer.js +136 -0
  19. package/dist/execution-policy.d.ts +28 -0
  20. package/dist/execution-policy.js +24 -0
  21. package/dist/feedback.d.ts +48 -0
  22. package/dist/feedback.js +230 -0
  23. package/dist/index.d.ts +20 -6
  24. package/dist/index.js +13 -5
  25. package/dist/input.js +11 -1
  26. package/dist/leases.d.ts +8 -0
  27. package/dist/leases.js +111 -0
  28. package/dist/node/agent-definitions.js +3 -5
  29. package/dist/node/config.d.ts +1 -0
  30. package/dist/node/config.js +5 -3
  31. package/dist/node/contribution-discovery.js +5 -8
  32. package/dist/node/session-store-jsonl.js +8 -5
  33. package/dist/node/settings.js +2 -2
  34. package/dist/node/trust.js +2 -4
  35. package/dist/observability.d.ts +3 -0
  36. package/dist/observability.js +18 -0
  37. package/dist/providers/media.d.ts +44 -0
  38. package/dist/providers/media.js +126 -0
  39. package/dist/providers/openai-compatible.js +18 -119
  40. package/dist/providers/openai-primitives.d.ts +9 -0
  41. package/dist/providers/openai-primitives.js +129 -0
  42. package/dist/providers/transport.d.ts +40 -0
  43. package/dist/providers/transport.js +221 -0
  44. package/dist/redaction.js +40 -13
  45. package/dist/resources.d.ts +5 -0
  46. package/dist/resources.js +4 -0
  47. package/dist/structured-output.d.ts +11 -0
  48. package/dist/structured-output.js +59 -0
  49. package/dist/testing/feedback.d.ts +6 -0
  50. package/dist/testing/feedback.js +37 -0
  51. package/dist/testing/persistence-schema.d.ts +102 -0
  52. package/dist/testing/persistence-schema.js +487 -0
  53. package/dist/testing/provider-conformance.js +10 -1
  54. package/dist/testing/run-ledger-conformance.d.ts +33 -0
  55. package/dist/testing/run-ledger-conformance.js +178 -0
  56. package/dist/testing/session-store-conformance.d.ts +16 -0
  57. package/dist/testing/session-store-conformance.js +73 -0
  58. package/dist/tools.d.ts +17 -0
  59. package/dist/tools.js +29 -2
  60. package/docs/a2a.md +73 -0
  61. package/docs/agent-events.md +17 -10
  62. package/docs/agent-loops.md +11 -5
  63. package/docs/agent-session-runtime.md +15 -16
  64. package/docs/cli-rpc.md +36 -5
  65. package/docs/coding-agent-tools.md +43 -9
  66. package/docs/coding-security.md +88 -0
  67. package/docs/compaction-observational-memory.md +2 -0
  68. package/docs/context-and-skills.md +1 -0
  69. package/docs/credential-storage.md +177 -0
  70. package/docs/credentials-and-redaction.md +4 -3
  71. package/docs/database-persistence.md +52 -7
  72. package/docs/evaluations.md +122 -0
  73. package/docs/extensions.md +2 -2
  74. package/docs/host-security.md +33 -2
  75. package/docs/index.md +46 -18
  76. package/docs/input-and-prompt-assembly.md +6 -5
  77. package/docs/mcp-tools.md +184 -0
  78. package/docs/middleware-hooks.md +2 -0
  79. package/docs/migration.md +51 -28
  80. package/docs/model-registry.md +5 -3
  81. package/docs/multimodal-content.md +156 -0
  82. package/docs/observability.md +171 -0
  83. package/docs/performance.md +249 -1
  84. package/docs/persistence-credentials-multimodality-primitives.md +303 -0
  85. package/docs/postgres-persistence.md +143 -0
  86. package/docs/provider-conformance.md +18 -0
  87. package/docs/provider-layer.md +1 -1
  88. package/docs/provider-packages.md +2 -0
  89. package/docs/provider-primitives.md +281 -0
  90. package/docs/providers/ai-sdk.md +113 -0
  91. package/docs/providers/kimi.md +1 -0
  92. package/docs/providers/neuralwatt.md +1 -0
  93. package/docs/providers/openai-compatible.md +2 -1
  94. package/docs/providers/openai.md +8 -1
  95. package/docs/providers/opencode-go.md +1 -0
  96. package/docs/providers/openrouter.md +1 -0
  97. package/docs/providers/zai.md +1 -0
  98. package/docs/public-contracts.md +13 -5
  99. package/docs/rag.md +113 -0
  100. package/docs/release-and-install.md +237 -30
  101. package/docs/resource-loading.md +14 -4
  102. package/docs/review-coverage-2026-07-14.md +260 -0
  103. package/docs/review-coverage-2026-07-15.md +193 -0
  104. package/docs/run-ledger-conformance.md +96 -0
  105. package/docs/runs-and-usage.md +43 -4
  106. package/docs/server.md +139 -0
  107. package/docs/session-store-conformance.md +16 -0
  108. package/docs/session-stores-and-branching.md +1 -0
  109. package/docs/settings-auth-trust-security.md +6 -5
  110. package/docs/sqlite-persistence.md +123 -0
  111. package/docs/structured-output.md +9 -0
  112. package/docs/supervisors.md +71 -0
  113. package/docs/tool-conformance.md +1 -0
  114. package/docs/tool-execution-primitives.md +374 -0
  115. package/docs/tools.md +39 -1
  116. package/docs/workflow-orchestration-primitives.md +581 -0
  117. package/docs/workflow-tui-primitives.md +5 -0
  118. package/docs/workflows.md +293 -0
  119. package/docs/working-and-semantic-memory.md +169 -0
  120. package/package.json +43 -5
  121. package/templates/init/README.md.tmpl +28 -0
  122. package/templates/init/env.example.tmpl +1 -0
  123. package/templates/init/gitignore.tmpl +11 -0
  124. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  125. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  126. package/templates/init/package.json.tmpl +22 -0
  127. package/templates/init/providers.json +76 -0
  128. package/templates/init/src/agent.ts.tmpl +10 -0
  129. package/templates/init/src/index.ts.tmpl +12 -0
  130. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  131. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -4,7 +4,9 @@
4
4
 
5
5
  The production persistence contracts describe database-neutral types for durable, multi-tenant storage of Prism sessions, branch handles, session entries, runs, agent-event ledger rows, tool-call rows, usage rows, agent-definition versions, retention policies, and migration records. They also define cursor-paginated query shapes so hosts can implement SQL, NoSQL, or object-store adapters without changing Prism runtime internals.
6
6
 
7
- Prism itself does not ship a production database adapter. The built-in `SessionStore` contract (`append` / `list` / optional `get`) remains the runtime seam; `ProductionPersistenceStore` is the optional adapter-facing contract for hosts that need paginated reads, tenant isolation, audit tables, and retention.
7
+ Prism itself does not ship a production database adapter. The built-in `SessionStore` contract (`append` / `list` / optional `get`) remains the runtime seam; `ProductionPersistenceStore` is the optional adapter-facing contract for hosts that need paginated reads, tenant isolation, audit tables, retention, and optional generic `CheckpointStore` / `LeaseStore` capabilities.
8
+
9
+ Plan 056 Task 1 adds dialect-neutral shared primitives under `@arnilo/prism/testing/persistence-schema`, `@arnilo/prism/testing/session-store-conformance`, and `@arnilo/prism/testing/run-ledger-conformance`. Task 2 ships `@arnilo/prism-session-store-sqlite` (see [SQLite persistence](sqlite-persistence.md)); Task 3 ships `@arnilo/prism-session-store-postgres` (see [PostgreSQL persistence](postgres-persistence.md)). Both implement dialect-local SQL against the shared model; Prism core still ships no ORM, driver, or migration runner.
8
10
 
9
11
  ## When to use it
10
12
 
@@ -65,7 +67,7 @@ Important shapes:
65
67
  | `RunRecord` | Stored run with `sessionId`, `branchId`, status (`queued` \| `running` \| `succeeded` \| `failed` \| `aborted`), `model`, `provider`, `idempotencyKey`, `abortReason`, and `error`. |
66
68
  | `AgentEventRecord` | Event ledger row with `event: AgentEvent` and a `redacted` flag. Hosts redact before storage. |
67
69
  | `ToolCallRecord` | Tool-call row with `arguments`, optional `result: ToolResult`, `reason`, `progress` snapshots, status, and a `redacted` flag. |
68
- | `UsageRecord` | Usage row wrapping `Usage` with session/run/entry linkage. |
70
+ | `UsageRecord` | Scoped provider-turn or aggregate run usage with session/run/entry and turn/attempt linkage. |
69
71
  | `AgentDefinitionRecord` | Versioned agent definition snapshot. Only stores `AgentDefinition` data; never provider credentials/resolvers/instances. |
70
72
  | `RetentionPolicy` | Policy with `maxAgeDays`, `maxEntriesPerSession`, `maxTotalBytes`, `archiveStore`, and `appliedKinds`. |
71
73
  | `MigrationRecord` | Applied migration with name, version, timestamp, checksum, and applied-by. |
@@ -164,7 +166,7 @@ The `event` JSONB stores a redacted `AgentEvent`. The `sequence` column is an im
164
166
 
165
167
  | Table | Key columns |
166
168
  | --- | --- |
167
- | `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
169
+ | `prism_usage` | `id` PK, `session_id` FK, `run_id`, `entry_id`, `scope`, `turn`, `attempt`, `usage` JSONB, `recorded_at`, `tenant_id`, `account_id`, `user_id`, `metadata` JSONB |
168
170
 
169
171
  The `usage` JSONB stores the `Usage` shape: input/output/total/cache tokens, cost, and currency.
170
172
 
@@ -184,6 +186,42 @@ The `usage` JSONB stores the `Usage` shape: input/output/total/cache tokens, cos
184
186
 
185
187
  Prism does not run migrations; hosts own migration tooling and use this table to record applied changes.
186
188
 
189
+ ## Shared schema model and migration contract
190
+
191
+ Adapter packages import the shared model instead of copying table names piecemeal:
192
+
193
+ ```ts
194
+ import {
195
+ createPersistenceSchemaModel,
196
+ createPersistenceMigrationContract,
197
+ assertPersistenceSchemaModel,
198
+ assertAdapterSchemaMatchesModel,
199
+ assertPersistenceQueryPaginationConforms,
200
+ assertTenantScopedQueryIsolation,
201
+ getPersistencePaginationCursors,
202
+ PARAMETERIZED_QUERY_GUIDANCE,
203
+ } from "@arnilo/prism/testing/persistence-schema";
204
+ import { assertSessionStoreConforms, runSessionStoreConformance } from "@arnilo/prism/testing/session-store-conformance";
205
+ import { assertRunLedgerConforms, runRunLedgerConformance } from "@arnilo/prism/testing/run-ledger-conformance";
206
+
207
+ const model = createPersistenceSchemaModel();
208
+ assertPersistenceSchemaModel(model);
209
+
210
+ await runSessionStoreConformance(() => createStore(testDatabase), { exerciseReopen: true });
211
+ await runRunLedgerConformance(() => createLedger(testDatabase), { exerciseReopen: true });
212
+ ```
213
+
214
+ | Primitive | Purpose |
215
+ | --- | --- |
216
+ | `PersistenceSchemaModel` | Versioned table/column/index model covering sessions, entries, parent chain, idempotency side table, runs, events, tool calls, usage, tenant columns, and `prism_migrations` |
217
+ | `createPersistenceMigrationContract()` | Strictly increasing migration steps, `prism_migrations` recording, advisory-lock guidance, and least-privilege migration/runtime role guidance |
218
+ | `getPersistencePaginationCursors()` | Indexed `(session_id, timestamp, id)`, `(run_id, sequence)`, `(run_id, recorded_at, id)` cursor shapes that avoid offset scans |
219
+ | `assertPersistenceQueryPaginationConforms()` | Generic cursor pagination fixture for `queryEntries` |
220
+ | `assertTenantScopedQueryIsolation()` | Tenant-filtered reads must not leak rows or primary-id collisions across tenants |
221
+ | `PARAMETERIZED_QUERY_GUIDANCE` | Values are always bound parameters; only validated identifiers may be quoted |
222
+
223
+ Dialect-local SQL remains in optional adapter packages. The shared model is the contract both adapters must satisfy before release.
224
+
187
225
  ## Adapter readiness checklist
188
226
 
189
227
  Before using a host database adapter in production:
@@ -236,7 +274,8 @@ Recommended indexes for the reference schema. Hosts should add DB-specific parti
236
274
  | `prism_tool_calls` | `(session_id, name, started_at)` | tool usage by name |
237
275
  | `prism_tool_calls` | `(run_id, started_at)` | run tool-call listing |
238
276
  | `prism_tool_calls` | `(tool_call_id)` | deduplication / replay |
239
- | `prism_usage` | `(session_id, recorded_at)` | usage aggregation |
277
+ | `prism_usage` | `(session_id, recorded_at)` | usage pagination |
278
+ | `prism_usage` | `(session_id, scope, recorded_at)` | scope-safe billing/aggregate queries |
240
279
  | `prism_usage` | `(run_id, recorded_at)` | run usage |
241
280
  | `prism_agent_definitions` | `(name, version)` | definition lookup |
242
281
  | `prism_retention_policies` | `(tenant_id, account_id, user_id)` | policy listing |
@@ -382,17 +421,22 @@ const dbStore: ProductionPersistenceStore = {
382
421
  ## Extension and configuration notes
383
422
 
384
423
  - `ProductionPersistenceStore` is an optional extension point. The runtime does not require it.
385
- - Hosts choose the database, schema, transaction, and indexing strategy. The contract only specifies query shapes.
424
+ - `ProductionPersistenceStore.checkpoints?: CheckpointStore` exposes generic versioned save/load/bounded-list/delete with compare-and-swap and fencing tokens, without workflow vocabulary.
425
+ - `ProductionPersistenceStore.leases?: LeaseStore` exposes atomic acquire/renew/release/get with opaque claim tokens, expiries, ownership scope, and monotonic fencing tokens.
426
+ - `ProductionPersistenceStore.feedback?: RunFeedbackStore` exposes immutable append, bounded owned query, and owned deletion. First-party adapters store migration-003 rows in `prism_run_feedback`, FK-link `run_id`, and index owner/run/trace creation cursors.
427
+ - Hosts choose the database, schema, transaction, and indexing strategy. The contracts specify query and checkpoint capability shapes.
386
428
  - `SessionStore` (`append`/`list`/`get`/optional `readBranchPath`) can be implemented on top of `ProductionPersistenceStore` or kept separate.
387
429
  - Cursor values and idempotency keys are host-defined and opaque to Prism.
430
+ - First-party SQLite/PostgreSQL adapters expose `persistence.checkpoints` and `persistence.leases`, backed by package-owned `prism_checkpoints` / `prism_leases` tables. `@arnilo/prism-workflows` consumes them for durable resume, human suspension, multi-process coordination, Phase 11 schedule records/fire leases, shared state, and replay lineage; workflow code owns no SQL table. `suspended`/`denied`, schedules, state history, and replay lineage remain namespaces/categories plus bounded checkpoint JSON values, so Phases 8 and 11 need no database migration.
388
431
 
389
432
  ## Security and performance notes
390
433
 
391
434
  - **No credentials in storage.** The contracts never include `CredentialResolver`, `AIProvider`, `ProviderResolver`, provider API keys, or credential values.
392
435
  - **Redact before storage.** Runtime session entries are redacted before `SessionStore.append`; `AgentEventRecord.event` and `ToolCallRecord.result` may contain secrets, so hosts must redact them (for example with `redactAgentEvent()` and a `SecretRedactor`) before writing to durable storage and set `redacted: true`.
393
- - **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility.
436
+ - **Tenant isolation.** `OwnershipScope` fields are available on records and queries, but enforcement is the host's responsibility. First-party feedback stores require tenant plus account/user and use exact run ownership on append, query, and delete.
437
+ - **Feedback retention/deletion.** Comments, tags, and metadata are redacted and bounded before insert. `RunFeedbackStore.delete()` provides explicit erasure; deleting a run cascades its feedback in first-party SQL schemas. Apply host retention policy to `created_at`.
394
438
  - **Pagination and branch reads.** Every query supports `cursor`/`limit`/`order` so hosts can avoid full-table or full-session scans. Loading an entire large session into memory to serve a provider context is an anti-pattern; implement `readBranchPath` and use branch-relevant filters / recursive ancestor queries.
395
- - **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, and entry kind. See the reference indexes above.
439
+ - **Indexes.** Production schemas should index `sessionId`, `runId`, `parentId`, `leafId`, timestamps, tenant/account/user, event type, entry kind, and feedback owner/run/trace creation cursors. See the reference indexes above.
396
440
 
397
441
  ## Related APIs
398
442
 
@@ -405,3 +449,4 @@ const dbStore: ProductionPersistenceStore = {
405
449
  - [Agent events](agent-events.md): `AgentEvent` variants and redaction.
406
450
  - [Tools](tools.md): `ToolResult`, `ToolCallContent`, and tool execution events.
407
451
  - [Public contracts](public-contracts.md): full public contract inventory.
452
+ - [Workflows](workflows.md): package-local durable checkpoint adapters on shared SQLite/Postgres handles.
@@ -0,0 +1,122 @@
1
+ # Evaluations
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, live post-run scoring, and bounded batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
6
+
7
+ ## When to use it
8
+
9
+ Use this package when a host needs offline quality checks or sampled live scoring without coupling scorers into core agent execution. Install it directly or through `@arnilo/prism-all`; installation does not attach scorers to runs.
10
+
11
+ ## Inputs / request
12
+
13
+ | API | Key inputs |
14
+ | --- | --- |
15
+ | `defineScorer` | `id`, `score({ result, item?, expected?, signal? })` |
16
+ | `defineDataset` | `id`, `version?`, immutable `items[]` with unique ids |
17
+ | `scoreRun` / `scoreRunLive` | `AgentRunResult`, scorers, optional `sampleRate`, store, ownership, redactor |
18
+ | `runExperiment` | `agent`, dataset, scorers, bounded `concurrency`, optional store/ownership |
19
+ | `createMemoryEvaluationStore` | optional seed records |
20
+ | `appendEvaluationFeedback` | `RunFeedbackStore`, `EvaluationStore`, feedback fields, and 1–64 known evaluation IDs |
21
+
22
+ ## Outputs / response / events
23
+
24
+ | API | Output |
25
+ | --- | --- |
26
+ | `scoreRun` | `EvaluationRecord[]` with `scored` / `skipped` / `failed` |
27
+ | `scoreRunLive` | same records; never mutates the agent result; host may ignore the promise |
28
+ | `runExperiment` | `ExperimentReport` with stable item order, evaluations, and aggregates |
29
+ | `EvaluationStore.query` | cursor-paginated, ownership-filtered page |
30
+ | `appendEvaluationFeedback` | immutable `RunFeedbackRecord` containing only evaluation/scorer IDs |
31
+
32
+ ## Request/response example
33
+
34
+ ```json
35
+ {
36
+ "scorerId": "contains-citation",
37
+ "status": "scored",
38
+ "score": 1,
39
+ "runId": "run_1",
40
+ "sessionId": "session_1",
41
+ "experimentId": "exp_1",
42
+ "sampled": true
43
+ }
44
+ ```
45
+
46
+ ## Implementation example
47
+
48
+ ```ts
49
+ import { createAgent, createMemoryRunFeedbackStore, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
50
+ import {
51
+ appendEvaluationFeedback,
52
+ createMemoryEvaluationStore,
53
+ defineDataset,
54
+ defineScorer,
55
+ runExperiment,
56
+ scoreRunLive,
57
+ } from "@arnilo/prism-evals";
58
+
59
+ const scorer = defineScorer({
60
+ id: "contains-citation",
61
+ score: ({ result }) => ({ score: result.text.includes("[") ? 1 : 0 }),
62
+ });
63
+
64
+ const dataset = defineDataset({
65
+ id: "citations",
66
+ version: "1",
67
+ items: [{ id: "1", input: "Summarize with a citation" }],
68
+ });
69
+
70
+ const agent = createAgent({
71
+ model: { provider: "mock", model: "demo" },
72
+ provider: createMockProvider([providerTextDelta("ok [1]"), providerDone()]),
73
+ });
74
+
75
+ const store = createMemoryEvaluationStore();
76
+ const report = await runExperiment({
77
+ agent,
78
+ dataset,
79
+ scorers: [scorer],
80
+ concurrency: 2,
81
+ store,
82
+ ownership: { tenantId: "t1", userId: "u1" },
83
+ });
84
+
85
+ const result = await agent.createSession().run("Follow up");
86
+ void scoreRunLive(result, { scorers: [scorer], store });
87
+ const evaluation = report.evaluations[0]!;
88
+ const feedbackStore = createMemoryRunFeedbackStore({
89
+ resolveRun: ({ runId }) => runId === evaluation.runId
90
+ ? { runId, sessionId: evaluation.sessionId!, tenantId: "t1", userId: "u1" }
91
+ : false,
92
+ });
93
+ const linked = await appendEvaluationFeedback({
94
+ feedbackStore,
95
+ evaluationStore: store,
96
+ evaluationIds: [evaluation.id],
97
+ feedback: { id: "fb_1", runId: evaluation.runId!, rating: 1, tenantId: "t1", userId: "u1" },
98
+ });
99
+ console.log(report.aggregate.meanScore, linked.evaluationIds);
100
+ ```
101
+
102
+ ## Extension and configuration notes
103
+
104
+ - Function scorers are the base primitive. No mandatory LLM judge, dashboard, or schema library is included.
105
+ - Evaluation-result persistence remains package-local (`EvaluationStore`) and in-memory by default. Linked feedback is separately durable through optional `ProductionPersistenceStore.feedback`; evaluation score/reason payloads are not copied there.
106
+ - `sampleRate` is explicit (`0`–`1`). Inject `random` for deterministic tests.
107
+ - Dataset snapshots are frozen; duplicate item ids fail closed.
108
+ - `appendEvaluationFeedback()` resolves every supplied ID from `EvaluationStore`, rejects missing IDs, verifies each evaluation has the same run, optional trace, and exact ownership as feedback, then copies only deduplicated `evaluationIds`/`scorerIds`. Evaluation scores, reasons, errors, and metadata are not duplicated.
109
+
110
+ ## Security and performance notes
111
+
112
+ - Scorers receive result/item data only. Credentials, tools, and workspace access are not provided unless the host deliberately closes over them.
113
+ - Records pass through `SecretRedactor` / `secrets` before store append.
114
+ - Queries filter by ownership scope. Feedback linkage additionally requires tenant plus account/user and the feedback store re-verifies the run.
115
+ - Experiment concurrency defaults to `1` and is capped at `32`. Scoring can reference run IDs without duplicating unbounded event payloads.
116
+
117
+ ## Related APIs
118
+
119
+ - [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
120
+ - [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
121
+ - [Observability](observability.md): trace/run metadata hosts may copy into `traceId`
122
+ - [Release and install](release-and-install.md): optional package install
@@ -101,7 +101,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
101
101
  ## Extension and configuration notes
102
102
 
103
103
  - Extension loading is explicit. Prism does not discover packages, read manifests, or load filesystem config in the kernel.
104
- - `AgentConfig.extensions` is host-owned metadata for compatibility; `createAgent()` and `session.run()` do not load it or call `Extension.setup()`. Load extensions with `createExtensionKernel().load(...)`, then pass selected contributions (`tools`, `context`, `skills`, middleware, etc.) into `createAgent()`.
104
+ - Extensions stay host-owned outside `AgentConfig`. `createAgent()` and `session.run()` do not load extension lists or call `Extension.setup()`. Load extensions with `createExtensionKernel().load(...)`, then pass selected contributions (`tools`, `context`, `skills`, middleware, etc.) into `createAgent()`.
105
105
  - Setup order is the order provided by the host.
106
106
  - The kernel writes only to explicit registries returned by `createContributionRegistries()` or provided by the host.
107
107
  - `api.registerTool()` contributes an inert `ToolDefinition` to `registries.tools`; it does not add the tool to an active tool registry, allow list, or dispatch loop.
@@ -117,7 +117,7 @@ await kernel.middleware.run("provider_request", { metadata: {} });
117
117
  ## Security and performance notes
118
118
 
119
119
  - No hidden global extension kernel, provider registry, credential resolver, settings provider, store, or resource loader is created.
120
- - `AgentConfig.extensions` does not auto-execute, so constructing or running an agent cannot unexpectedly run extension code.
120
+ - Extension packages do not auto-execute from `createAgent()`, so constructing or running an agent cannot unexpectedly run extension code.
121
121
  - Error events use `ErrorInfo` and redact only known secret values passed in `secrets`.
122
122
  - Do not put resolved credential values in extension events, registry metadata, docs, logs, prompts, or session stores.
123
123
  - Event and middleware dispatch are ordered and dependency-free. They use no timers, background workers, filesystem discovery, network calls, provider calls, or tool execution.
@@ -25,9 +25,13 @@ Start from explicit host inputs. Do not let runtime code discover security state
25
25
  | Permission decisions | allow/deny rules or approval UI result | `createStaticPermissionPolicy`, `assertPermission()` |
26
26
  | Tool allow-list | active tools for this agent/session/run | `createToolRegistry`, `filterTools()`, `dispatchToolCall()` |
27
27
  | Tool argument rules | host validator | `AgentConfig.validator`, `RunOptions.validate`, `ToolValidator` |
28
+ | Coding execution policy | path/command approval adapter | `ExecutionPolicy`, `@arnilo/prism-coding-security` |
29
+ | Remote media policy | public/default pinned DNS or explicit trusted transport | `SsrfPolicy`, `resolveMediaContentBlock()` |
28
30
  | Durable history | host database adapter | `SessionStore`, `assertSessionStoreConforms()` |
29
31
  | Durable audit | host ledger adapter | `RunLedger`, `redactRunLedgerRecord()` |
30
32
  | Extensions | explicit package imports only | `createExtensionKernel`, `ExtensionAPI` |
33
+ | Remote agent/workflow API | host authentication + ownership mapping | `@arnilo/prism-server`, `createPrismHandler()` |
34
+ | MCP server exposure | host MCP auth + selected capability list | `createPrismMcpServer()`, `createPrismMcpWebHandler()` |
31
35
 
32
36
  ## Outputs / response / events
33
37
 
@@ -104,7 +108,7 @@ Wire those values where they matter: provider adapters receive the resolved cred
104
108
 
105
109
  ## Extension and configuration notes
106
110
 
107
- - Keep security state explicit. `AgentConfig.settings` and `AgentConfig.credentials` are host-owned metadata; `createAgent()` and `session.run()` do not automatically call `settings.get()` or `credentials.resolve()`.
111
+ - Keep security state explicit. Settings and credentials are host-owned outside `AgentConfig`; `createAgent()` and `session.run()` do not automatically call `settings.get()` or `credentials.resolve()`.
108
112
  - Resolve credentials at the provider/request edge, as late as possible. Do not put resolved credentials in configs, manifests, registries, prompts, messages, events, session entries, run ledgers, idempotency keys, cache keys, or logs.
109
113
  - Use `createExplicitCredentialResolver()` to document source order such as runtime override → stored credential → caller-supplied env object → fallback.
110
114
  - Use `createEnvCredentialResolver()` only with an object the host passes in. Prism does not read `process.env` for credentials.
@@ -120,14 +124,41 @@ Wire those values where they matter: provider adapters receive the resolved cred
120
124
  - Prism does not sandbox host tools, extensions, provider adapters, credential resolvers, or custom middleware. Use OS/container/process isolation when code is untrusted.
121
125
  - Redaction is exact known-secret replacement only. It is not arbitrary secret detection, entropy scanning, or DLP.
122
126
  - Known secrets must be passed into redactors before data is emitted or persisted. Redact again in host adapters if they transform records after Prism redaction.
123
- - Tool `parameters` metadata is not validation. Add a `ToolValidator` or validate inside the tool before side effects.
127
+ - Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects.
128
+ - MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Configure stdio commands and HTTP URLs explicitly; bound output with `maxResultBytes`; register prefixed tools only after trust review. MCP server direction exposes only passed tools/commands, requires per-call `authorize`, and retains `PermissionPolicy`/`ToolValidator` gates for tools. Its web handler needs host `resolveAuthInfo`, TLS, rate limiting, and exact host/origin policy. See [MCP client/server exposure](mcp-tools.md).
129
+ - `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive non-empty ownership from validated host identity, never request JSON. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
130
+ - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Caching defaults to none; missing run/session identity never falls back to a global cache. Prism does not provide OS sandboxing unless the host supplies a sandbox adapter.
131
+ - Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
124
132
  - Permission checks happen before tool validation and before `tool.execute()`. Middleware cannot grant permission by renaming a tool.
125
133
  - Session stores and ledgers receive redacted values when a redactor is active, but durable storage remains host-owned. Enforce tenant/account/user ownership and retention in the database layer.
126
134
  - Provider-owned auth/content/session/cache/security headers win over caller headers in adapters that merge headers.
127
135
  - Security checks are bounded explicit calls on the active path. Prism adds no hidden global middleware, background workers, watchers, network calls, or filesystem scans.
128
136
 
137
+ ### 0.0.4 release security audit (2026-07-14)
138
+
139
+ - `npm audit --audit-level=high`: 0 vulnerabilities at every severity.
140
+ - Lockfile: 162 registry dependency records, all with `resolved` provenance URL and integrity hash; `npm ls --all` reports a clean graph.
141
+ - License inventory: 160 locked third-party packages; all declare permissive MIT, ISC, BSD, Apache-2.0, or compatible dual licenses. No GPL, AGPL, SSPL, or missing lockfile license metadata.
142
+ - Install scripts: only `better-sqlite3@12.11.1` runs an install script (`prebuild-install || node-gyp rebuild --release`), required by the explicitly installed SQLite adapter. Core and other optional packages add no install hook.
143
+ - Secret scan: source, tests, docs, workflow files, package metadata, built tests, packed-install canary, and tarball deny-list checks found no private-key block or common live-token prefix. Runtime redaction fixtures cover requests, events, ledgers, stores, checkpoints, provider/OAuth errors, and credential ciphertext.
144
+ - Threat suites pass for parameterized SQL/tenant isolation, HTTP URL/SSRF rejection, realpath/symlink containment, shell-metacharacter approval, schema prototype-pollution/remote-reference bounds, OAuth polling/abort/redaction, credential tamper/wrong-key/KDF floors, MCP result bounds/timeouts, and coding approval/path policy.
145
+
146
+ PostgreSQL TLS/network policy, MCP endpoint allow-listing, provider base URLs, OS keychain availability, process sandboxing, workflow tenant identity, and ANSI/control-sequence sanitization in any host terminal renderer remain host boundaries. Prism 0.0.4 ships JSON-line RPC, not an interactive TUI; hosts must render untrusted model/tool text safely. Credential-gated PostgreSQL/provider/keychain tests are separate operator/CI gates, not silently replaced by mocks.
147
+
148
+ ## Supervisor and A2A boundaries
149
+
150
+ - Register children explicitly. AND-compose parent/child/hook permissions; never let a delegation hook replace broader parent policy.
151
+ - Build each child's context/memory with supervisor-provided `resourceId`/`threadId`; resolve provider credentials inside that child factory.
152
+ - Keep depth, active children, input, turn/tool/token, timeout, and queue ceilings finite; propagate abort through nested calls.
153
+ - Expose A2A only behind per-request authentication/authorization, TLS, edge rate limits, and replay policy. Public card discovery grants no invoke access.
154
+ - Remote A2A endpoints/card URLs require exact HTTPS origin allow-lists and redirect rejection. Pin ES256 card keys/expiry; never auto-fetch untrusted `jku`.
155
+ - Treat cards, task status, errors, artifacts, and SSE frames as untrusted bounded input and redact before logs/hooks/events.
156
+
129
157
  ## Related APIs
130
158
 
159
+ - [Web-standard server handler](server.md): remote agent/workflow route, ownership, limits, abort, and deployment boundary.
160
+ - [Supervisor delegation](supervisors.md): local child permission/memory/budget boundary.
161
+ - [A2A interoperability](a2a.md): remote card/auth/origin/signature boundary.
131
162
  - [Settings, auth, trust, and security controls](settings-auth-trust-security.md): low-level helpers and boundary hardening table.
132
163
  - [Credentials and redaction](credentials-and-redaction.md): credential resolver order, caller-supplied env objects, OAuth refresh, exact redaction, and no persistent secret store.
133
164
  - [Tools](tools.md): active tool registry, allow/deny filters, permission order, validator order, blocked events, and no sandbox.
package/docs/index.md CHANGED
@@ -3,49 +3,63 @@
3
3
  Prism is a TypeScript/Node.js agent harness. Host apps and extension packages own providers, tools, resources, credentials, storage, UI, and business behavior. Prism supplies contracts, registries, streaming events, and replaceable runtime primitives.
4
4
 
5
5
  ## Public contracts
6
- - [Public contracts](public-contracts.md): type shapes for messages, content, agents, sessions, providers, tools, context, skills, extensions, stores, resources, settings, credentials, and events.
6
+ - [Public contracts](public-contracts.md): type shapes for messages, agents, tools, stores, generic `CheckpointStore`, atomic `LeaseStore`, bounded `EventMultiplexer`, resources, credentials, and events.
7
7
 
8
8
  ## Agent/session runtime
9
- - [Agent/session runtime](agent-session-runtime.md): create agents and sessions, run prompts, subscribe to normalized events, and see which `AgentConfig` fields are runtime-consumed vs host-owned metadata. Covers tool-call loop transcript shape and prior-reasoning preservation across turns.
9
+ - [Agent/session runtime](agent-session-runtime.md): create agents and sessions, get direct `AgentRunResult` values from `run`/`prompt`, use integrated `stream()`, and subscribe to normalized events. Covers tool-call loop transcript shape and prior-reasoning preservation across turns.
10
10
  - [Agent definitions](agent-definitions.md): resolve declarative `AgentDefinition` values via `resolveAgentDefinition`, and turn app-config `<configRoot>/agents/<name>/AGENT.md` bundles into runnable agents via `discoverAgentBundles` / `resolveAgentBundle` (explicit tool/skill activation by name, fail-closed omitted capabilities, migration-only `activateAllCapabilities`, strict duplicate scope checks, configurable prompt layers, no auto-discovery).
11
11
  - [Agent loops](agent-loops.md): replaceable per-run control loops — `singleShotLoop` default and `generate-validate-revise` with host-supplied `validator`/`parser`/`repairer` callbacks.
12
- - [Agent events](agent-events.md): the `AgentEvent` stream — agent/turn/message (including live `tool_call_delta` fragments), tool execution, queue/subscriber overflow, compaction/retry, artifact validation/refinement, and error variants, redacted via `redactAgentEvent`.
13
- - [Runs and usage ledger](runs-and-usage.md): `RunLedger` adapter for durable run, event, tool-call, usage persistence, cache diagnostics, ownership/idempotency, and redaction guidance.
12
+ - [Agent events](agent-events.md): the `AgentEvent` stream — agent/turn/message (including live `tool_call_delta` fragments), provider turn timing, tool execution, queue/subscriber overflow, compaction/retry, artifact validation/refinement, and error variants, redacted via `redactAgentEvent`.
13
+ - [Observability](observability.md): metadata-only provider/tool and run-feedback/evaluation projection, terminal span cleanup, low-cardinality metrics, and optional `@arnilo/prism-observability-opentelemetry` adapter.
14
+ - [Evaluations](evaluations.md): optional deterministic scorers/datasets/experiments plus ID-only linkage from evaluation records to immutable owned run feedback.
15
+ - [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence plus bounded immutable run/trace feedback, evaluation links, ownership, redaction, query, and deletion semantics.
14
16
  - [Performance limits](performance.md): bounded live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
15
- - [Structured output](structured-output.md): the `Artifact*` seam (parser/validator/repairer, host-defined `T`) — the only typed-output path from a loop, with a Synapta-style schema→`ArtifactValidation` mapping example and an end-to-end third-party integration walkthrough.
17
+ - [Structured output](structured-output.md): the `Artifact*` seam plus provider-native `StructuredOutputOptions` / `structuredOutputMode` for capable models.
16
18
 
17
19
  ## Compaction/session memory
18
20
  - [Compaction and retry policies](compaction-and-retry.md): summarize branch history and retry transient provider failures with host-replaceable policies.
19
21
  - [LLM compaction package](compaction-llm.md): optional provider-backed compaction strategy package with max-output budgets mapped through `model.parameters.maxTokens` to provider wire fields.
20
22
  - [Observational memory compaction package](compaction-observational-memory.md): optional source-backed memory, owned runtime append callback, provider-valid worker transcripts, fast compaction, recall tool, and status/view command package.
23
+ - [Working and semantic memory](working-and-semantic-memory.md): optional `@arnilo/prism-memory` working-memory store, semantic recall, Embedder/VectorStore contracts, in-memory adapters, and PostgreSQL/pgvector path.
21
24
  - [Session stores](session-stores.md): `SessionStore` contract, `SessionAppendOptions`, `SessionAppendConflictError`, branch handles, `readBranchPath`, and dev-vs-production branch reads — start here for session persistence.
22
25
  - [Session stores and branching](session-stores-and-branching.md): detailed branch semantics and helper reference (kept for compatibility; links back to the canonical atomic append / branch-handle sections).
23
- - [Database persistence](database-persistence.md): production persistence contracts, conditional append transaction pattern, idempotency indexes, `readBranchPath`, reference relational schema, retention, migrations, and NoSQL mapping.
24
- - [Migration guide](migration.md): the two cross-cutting app migrations in one place — in-memory/JSONL → database-backed `ProductionPersistenceStore` persistence (+ `RunLedger`) and permissive capability defaults → Phase 38 explicit `tools`/`skills` activation, with before/after shapes and links to the detailed pages.
26
+ - [Database persistence](database-persistence.md): production persistence contracts, shared schema/migration primitives (`@arnilo/prism/testing/persistence-schema`), conditional append transaction pattern, idempotency indexes, `readBranchPath`, reference relational schema, retention, migrations, and NoSQL mapping.
27
+ - [SQLite persistence](sqlite-persistence.md): optional adapter with session/run storage, generic checkpoints/leases, and migration-003 owned run feedback over `better-sqlite3`.
28
+ - [PostgreSQL persistence](postgres-persistence.md): optional pooled session/run/checkpoint/lease/feedback adapter over `pg`, with advisory-lock migrations and opt-in live conformance.
29
+ - [Migration guide](migration.md): 0.0.3 compatibility and 0.0.5 adoption — first-party/custom database persistence plus explicit fail-closed tool/skill activation.
25
30
  - [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety.
31
+ - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
26
32
 
27
33
  ## Provider and model connection
34
+ - [Provider primitives](provider-primitives.md): shared bounded transport and OpenAI serialization helpers — migrated across first-party providers; native structured-output and observability contracts.
28
35
  - [Provider layer](provider-layer.md): register and resolve host-owned providers/models, choose replace-or-error duplicate policy, create provider events, stream/reconstruct tool-call deltas, use generic provider request options, and test with the mock provider; deprecated provider-level timeout/retry hints point to runtime abort/retry.
29
36
  - [Model registry](model-registry.md): register and resolve `ModelConfig` records with capabilities, limits, cost, cache support metadata, compat data, and duplicate policy.
30
37
  - [Provider caching](provider-caching.md): use `PromptCacheHints`, `PromptCacheBreakpoint`, `ModelCacheCapabilities`, cache-aware stable-prefix guidance, and shared cache diagnostics helpers; includes a per-provider explicit/implicit cache matrix for OpenAI, OpenRouter, OpenCode Go, Z.AI, Kimi, and NeuralWatt; cache hints are best-effort and cache keys are never secrets.
31
38
  - [Provider request policies](provider-request-policies.md): chain `ProviderRequestPolicy` hooks, use `createSessionCachePolicy`, and merge legacy/structured cache options safely.
32
39
  - [Provider packages](provider-packages.md): define explicit provider packages, model metadata, auth descriptors, request/cache policies, and provider-owned header precedence without package discovery or provider-specific core behavior; includes a first-party cache behavior summary.
33
40
  - Phase 12 package workspaces: [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md) with implicit vLLM prefix caching, reasoning controls (`reasoning_effort`/`thinking_token_budget`/`enable_thinking`/`preserve_thinking`/`clear_thinking`), reasoning preservation, OpenAI-style tool-call loop, quota, telemetry, and retry classification helpers.
41
+ - Optional AI SDK adapter: [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md) maps host-owned `LanguageModelV4` models onto Prism `AIProvider` streams (specification v4; available directly or through provider/all umbrellas).
34
42
  - [OpenAI-compatible provider](providers/openai-compatible.md): optional provider subpath using native or injected `fetch` for Chat Completions streaming.
35
43
 
36
44
  ## Input, prompt, and context assembly
37
45
  - [SDK customization guide](customization.md): map provider resolution, middleware, context, builders, injectors, loops, compaction, retry, stores, and skills to explicit host-wired APIs.
38
- - [Input and prompt assembly](input-and-prompt-assembly.md): render tiny prompt templates and turn common host input, history, attachments, explicit resources, summaries, and tool results into messages with replaceable builders, provider-input assembly, legacy default order, and opt-in cache-aware ordering.
46
+ - [Input and prompt assembly](input-and-prompt-assembly.md): render tiny prompt templates and turn common host input, history, attachments, explicit resources, summaries, and tool results into messages with replaceable builders, provider-input assembly, legacy default order, and opt-in cache-aware ordering. Audio/file/document `ContentBlock` types and capability checks are documented there.
47
+ - [Multimodal content](multimodal-content.md): complete-request media resolution and aggregate bounds, DNS-classified/address-pinned URLs, SSRF/MIME policy, and `ModelCapabilities.input` tags.
39
48
  - [System prompts](system-prompts.md): compose explicit user/package/app/run system prompt layers, auto-load the standard `AGENTS.md` (workspace) / `SYSTEM.md` prompt files via the Node `loadSystemPromptFiles` loader (trust-gated for `AGENTS.md`), and append `SYSTEM.md` → per-agent `AGENT.md` body → repo `AGENTS.md` layers from a discovered agent bundle via `resolveAgentBundle`.
40
49
  - [Instruction injection](instruction-injection.md): register package injectors that layer redacted instructions/context blocks without granting tools, permissions, or resource escapes.
41
50
  - [Context and skills](context-and-skills.md): resolve ordered context providers and keep context/skill selection host-owned; omitted declarative skills stay inactive by default, `toolNames` fail closed before provider turns, and strict skill registries prevent silent shadowing.
51
+ - [Retrieval-augmented generation](rag.md): optional bounded text/Markdown chunking, Phase 7 vector indexing/retrieval, stable citations, and explicit inert context injection.
42
52
 
43
53
  ## Tools
44
54
  - [Tools](tools.md): register host-owned active tools with replace-or-error duplicate policy, apply exact allow/deny filtering, and dispatch tool calls.
45
- - [Coding agent tools](coding-agent-tools.md): optional first-party package `@arnilo/prism-coding-agent` providing `shell`, `read`, `write`, and `edit` tools (ported from pi) as `ToolDefinition`s a host registers; pluggable operation backends, per-path mutation serialization, and read-only/coding aggregators. Host shell/filesystem access — gate with permission/trust policies.
55
+ - [Tool execution primitives](tool-execution-primitives.md): JSON Schema validation, exclusive-aware bounded parallel dispatch, MCP bridge mapping, coding execution policy, and image-read bounds.
56
+ - [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
57
+ - [MCP client bridge and server exposure](mcp-tools.md): optional `@arnilo/prism-mcp` package mapping remote tools into Prism and explicitly selected Prism tools/commands onto SDK `McpServer`.
58
+ - [Coding agent tools](coding-agent-tools.md): optional first-party package `@arnilo/prism-coding-agent` providing `shell`, `read`, `write`, and `edit` tools (ported from pi) as `ToolDefinition`s a host registers; pluggable operation backends, per-path mutation serialization, optional `ExecutionPolicy`, bounded image reads (`maxImageBytes`, `transformImage`), and read-only/coding aggregators. Host shell/filesystem access — gate with permission/trust policies and `@arnilo/prism-coding-security` approval.
59
+ - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, and abort-aware streaming sandbox adapters for coding tools.
46
60
 
47
61
  ## Extensions/plugins
48
- - [Contribution discovery (workspace)](contribution-discovery.md): opt-in, realpath-contained directory scanner turning `SKILL.md`/`manifest.json` into inert `DiscoveredContribution` envelopes the host registers — no `import()`, no auto-activate, no provider scanning. (Per-agent `AGENT.md` bundles live under an app-controlled `configRoot`; see [Agent definitions](agent-definitions.md).)
62
+ - [Contribution discovery (workspace)](contribution-discovery.md): opt-in, realpath-contained directory scanner turning `SKILL.md`/`manifest.json` into inert `DiscoveredContribution` envelopes the host registers — no `import()`, no auto-activate, no provider scanning. Per-agent bundles remain app-controlled and are documented under Agent/session runtime.
49
63
  - [Contribution registries](contribution-registries.md): explicit host-owned registries for extension/package contributions without hidden globals, with `duplicate: "error"` strict mode for provider/model/tool/skill shadowing prevention.
50
64
  - [Extension kernel and event bus](extensions.md): load host-provided extensions in order, register contributions, emit lifecycle events, and isolate extension errors.
51
65
  - [Extension authoring guide](extension-authoring.md): publish third-party extension packages that register inert contributions and show host-owned activation, trust, permissions, redaction, and no-sandbox boundaries.
@@ -54,25 +68,39 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
54
68
  ## Configuration/manifests
55
69
  - [Configuration and manifests](configuration-and-manifests.md): merge in-memory JSON config layers and validate data-only package manifests with prototype-pollution key rejection.
56
70
  - [Node filesystem config loader](node-filesystem-config.md): explicitly read caller-named JSON config files in Node hosts.
57
- - [Resource loading](resource-loading.md): decode text, JSON, and manifest resources through caller-provided loaders.
71
+ - [Resource loading](resource-loading.md): decode text, JSON, binary, and manifest resources through caller-provided loaders with bounded byte limits.
72
+
73
+ ## Server/API
74
+ - [Web-standard server handler](server.md): optional framework-free authorized direct/SSE agent and durable workflow run/status/cancel/resume routes with explicit bounds and zero default exposure.
75
+
76
+ ## Multi-agent and interoperability
77
+ - [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, and finite budgets.
78
+ - [A2A interoperability](a2a.md): optional A2A 1.0 cards, ES256 signatures, authorized JSON-RPC/SSE handler, and exact-origin remote client.
58
79
 
59
80
  ## CLI/RPC
60
- - [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including branch-handle results, fixed `forkSession`, and `checkout`.
81
+ - [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
82
+ - [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — durable human suspend/resume, schedules/background execution, nested workflows, bounded validated state, immutable-lineage replay, multi-process coordination, events, and optional RPC/Web bindings. Interactive TUI (C-012) deferred.
83
+ - [Workflow orchestration primitives](workflow-orchestration-primitives.md): architecture inventory — workflow adapters consume core `CheckpointStore`, `LeaseStore`, and bounded `EventMultiplexer`; run control and optional RPC commands stay package-local.
84
+ - [Workflow/TUI scope](workflow-tui-primitives.md): records why 0.0.5 ships workflow APIs/RPC control but no interactive terminal UI.
61
85
 
62
86
  ## Security and credentials
63
- - [Host security guide](host-security.md): fail-closed checklist for credentials, settings, redaction, trust roots, permission policies, persistence, extension loading, and tool validation.
64
- - [Security/auth/trust](settings-auth-trust-security.md): settings providers, credential helpers, trust/permission policies, redaction controls, host-owned `AgentConfig.settings`/`credentials`, and security-boundary hardening summary.
65
- - [Credentials and redaction](credentials-and-redaction.md): compose explicit credential resolver order, use caller-supplied env objects/OAuth refresh helpers, avoid eager `AgentConfig.credentials` resolution, and redact known secret values.
87
+ - [Host security guide](host-security.md): fail-closed checklist for credentials, settings, redaction, trust roots, remote media, permission/approval policies, persistence, extension loading, and tool validation.
88
+ - [Security/auth/trust](settings-auth-trust-security.md): settings providers, credential helpers, trust/permission policies, redaction controls, host-owned settings/credentials wiring outside `AgentConfig`, and security-boundary hardening summary.
89
+ - [Credentials and redaction](credentials-and-redaction.md): compose explicit credential resolver order, use caller-supplied env objects/OAuth refresh helpers, resolve credentials only at the provider edge, and redact known secret values.
90
+ - [Credential storage](credential-storage.md): optional `@arnilo/prism-credentials-node` encrypted-file and system-keychain adapters for durable host-owned credentials.
66
91
 
67
92
  ## Testing and examples
68
- - [Provider layer](provider-layer.md): use `createMockProvider()` and provider event helpers for deterministic tests without timers, credentials, or network.
93
+ - Provider test doubles: `createMockProvider()` and provider event helpers are documented on the canonical Provider layer page above.
69
94
  - [Provider conformance](provider-conformance.md): run network-free provider adapter assertions (stream order, abort, tool-call reconstruction, cache usage, content coverage, protected header ownership, secret leak) from `@arnilo/prism/testing/provider-conformance`.
70
95
  - [Session store conformance](session-store-conformance.md): assert any `SessionStore` adapter satisfies append/idempotency/conflict/branch invariants from `@arnilo/prism/testing/session-store-conformance`.
96
+ - [Run ledger conformance](run-ledger-conformance.md): assert durable run/event/tool/usage writes and reopen survival. Run-feedback stores use `@arnilo/prism/testing/feedback` for append/query/delete/ownership linkage conformance.
71
97
  - [Compaction conformance](compaction-conformance.md): assert any `CompactionStrategy` returns a non-empty redacted summary and observes abort from `@arnilo/prism/testing/compaction-conformance`.
72
98
  - [Tool conformance](tool-conformance.md): assert the tool-dispatch blocked-reason matrix (unknown/denied/invalid/permission/validator) and success path from `@arnilo/prism/testing/tool-conformance`.
73
99
  - [Extension conformance](extension-conformance.md): assert an `Extension` setup runs, contributions stay inert, and setup errors are redacted or rethrown from `@arnilo/prism/testing/extension-conformance`.
74
- - `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, cache-aware prompt assembly, NeuralWatt agent run, stores/branching, compaction, observational-memory recall, structured-output/artifact-loop, CLI, RPC).
100
+ - `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, cache-aware prompt assembly, NeuralWatt agent run, stores/branching, compaction, observational-memory recall, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
75
101
 
76
102
  ## Release and install
77
- - [Release and install](release-and-install.md): package layout, install specifiers, required `@arnilo/prism` peer, tarball contents and exclusions, the map-retention knob, the release workflow, and the offline test budget.
103
+ - [Release and install](release-and-install.md): 30-package graph and profiles, install/tarball rules, deterministic resumable provenance publication, and offline test budget.
104
+ - [Review coverage (2026-07-15)](review-coverage-2026-07-15.md): frozen 0.0.5 finding/feature ownership, existing-primitive inventory, package decisions, threat boundaries, exclusions, and measured Phase 0 baseline.
105
+ - [Review coverage (2026-07-14)](review-coverage-2026-07-14.md): traceability matrix linking review findings and bug-report fixes to plan tasks, tests, and documentation for release 0.0.4.
78
106
 
@@ -60,7 +60,7 @@ Useful exported types:
60
60
  - `DefaultInputBuilder`: the default `InputBuilder` with typed default context.
61
61
  - `InputAssemblyLayout`: `"legacy" | "cache_aware"`; legacy is default.
62
62
  - `DefaultInputBuildContext`: optional input layout, instructions, history, summaries, attachments, resource loader/URIs, tool results, middleware, ids, metadata, and abort signal.
63
- - `InputAttachment`: already-loaded text/content or an explicit URI loaded through a caller-provided `ResourceLoader`.
63
+ - `InputAttachment`: already-loaded text/content blocks (including `audio`, `file`, and `document`) or an explicit URI loaded through a caller-provided `ResourceLoader`.
64
64
  - `PromptInstruction`: labeled system instruction text.
65
65
  - `DefaultPromptBuilder`: the default `PromptBuilder`.
66
66
  - `AssembleProviderInputOptions`: model, input, optional builders, context providers, selected skills, active tools, generic provider options, metadata, and signal.
@@ -82,10 +82,10 @@ The builder returns `readonly Message[]`.
82
82
  The default prompt builder still prepends context, selected skills, and tool declarations before those input messages. Cache-aware ordering gives cache-capable providers a stable prefix only while those stable inputs stay byte-stable; changing tools, context, resources, summaries, history, or attachments changes the prefix too.
83
83
  - History is prepended before current input.
84
84
  - Instructions and summaries are system messages; compacted branch summaries from `rebuildSessionContext()` use the same path.
85
- - Text attachments and explicit text resources are user messages.
85
+ - Text attachments and explicit text resources are user messages; inline `audio`/`file`/`document` blocks pass through unchanged on attachments with `content`.
86
86
  - Tool results are tool messages containing `tool_result` content; the agent/session runtime uses this to feed dispatched tool results into the next provider turn, placing the assistant `tool_call` and the matching role `tool` `tool_result` before any final assistant content. Cache-aware layout keeps tool results before the current user suffix so it does not split tool transcripts.
87
87
  - Middleware runs only when `middleware` is supplied in the context.
88
- - `assembleProviderInput()` returns a `ProviderRequest` with the caller's model/tools/provider options/metadata/signal and composed messages/context.
88
+ - `assembleProviderInput()` returns a `ProviderRequest` with the caller's model/tools/provider options/metadata/signal and composed messages/context. It also calls `assertMessagesSupportModelCapabilities()` so unsupported `audio`/`file`/`document`/`image` blocks fail with `UnsupportedModalityError` when the model declares `capabilities.input`.
89
89
  - `renderPromptTemplate()` replaces top-level `{{name}}` variables with caller-supplied JSON-compatible values. Strings are inserted directly; numbers, booleans, `null`, arrays, and objects are stringified deterministically with sorted object keys. Missing variables throw by default or stay unchanged with `{ missing: "preserve" }`.
90
90
 
91
91
  ## Request/response example
@@ -165,7 +165,7 @@ const request = await assembleProviderInput({
165
165
  - The builder is linear in supplied messages, attachments, resources, and tool results. Layout selection is one flattening branch over already-built groups.
166
166
  - Template expansion is dependency-free string replacement over `{{name}}` variables. It does not evaluate expressions, filters, loops, partials, JavaScript, globals, or prototype properties.
167
167
  - It performs no provider calls, tool execution, credential resolution, package discovery, filesystem scan, network access, timers, or watchers.
168
- - URI attachments/resources load only through the caller-provided `ResourceLoader`.
168
+ - URI attachments/resources load only through the caller-provided `ResourceLoader`. Binary media uses `resolveMediaContentBlock()` / `loadBinaryResource()` with bounded bytes, SSRF checks for URLs, and MIME magic validation — see [Multimodal content](multimodal-content.md).
169
169
  - Do not place secrets in templates, variables, instructions, messages, attachments, tool results, metadata, middleware payloads, or docs examples.
170
170
  - Active tools are passed through from the host; prompt middleware cannot grant additional provider tools.
171
171
  - Skill selection is handled by the host/skill registry path; this builder only includes selected skills passed by the caller.
@@ -175,7 +175,8 @@ const request = await assembleProviderInput({
175
175
  - [SDK customization guide](customization.md): high-level map of replaceable provider resolution, middleware, context, builder, injector, loop, compaction, retry, store, and skill seams.
176
176
  - [Public contracts](public-contracts.md): `Message`, `ContentBlock`, `InputBuilder`, `InputBuildContext`, `ToolResult`, and `ResourceLoader` shapes.
177
177
  - [Context and skills](context-and-skills.md): ordered context resolution feeding prompt composition.
178
- - [Resource loading](resource-loading.md): `loadTextResource()` behavior used for explicit URI resources.
178
+ - [Multimodal content](multimodal-content.md): `audio`/`file`/`document` blocks, bounded media resolution, and capability checks.
179
+ - [Resource loading](resource-loading.md): `loadTextResource()` and `loadBinaryResource()` behavior used for explicit URI resources.
179
180
  - [Middleware hooks](middleware-hooks.md): ordered middleware registry and `input_assembly`, `context`, and `prompt_build` hooks.
180
181
  - [System prompts](system-prompts.md): compose layered package/app/user/run prompts before input assembly.
181
182
  - [Contribution registries](contribution-registries.md): inert input, prompt, context, and skill contributions.