@arnilo/prism 0.0.4 → 0.0.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +18 -0
- package/README.md +34 -10
- package/dist/agents.js +146 -19
- package/dist/cli-init.d.ts +41 -0
- package/dist/cli-init.js +390 -0
- package/dist/cli-runner.d.ts +7 -1
- package/dist/cli-runner.js +13 -1
- package/dist/content.d.ts +19 -0
- package/dist/content.js +197 -69
- package/dist/contracts.d.ts +94 -9
- package/dist/contracts.js +8 -0
- package/dist/feedback.d.ts +48 -0
- package/dist/feedback.js +230 -0
- package/dist/index.d.ts +6 -4
- package/dist/index.js +4 -3
- package/dist/providers/media.d.ts +3 -1
- package/dist/providers/media.js +11 -1
- package/dist/testing/feedback.d.ts +6 -0
- package/dist/testing/feedback.js +37 -0
- package/dist/testing/persistence-schema.d.ts +3 -3
- package/dist/testing/persistence-schema.js +32 -2
- package/dist/testing/run-ledger-conformance.js +7 -1
- package/docs/a2a.md +73 -0
- package/docs/agent-events.md +4 -6
- package/docs/agent-loops.md +1 -1
- package/docs/agent-session-runtime.md +14 -16
- package/docs/cli-rpc.md +35 -7
- package/docs/coding-agent-tools.md +2 -2
- package/docs/coding-security.md +7 -3
- package/docs/compaction-observational-memory.md +2 -0
- package/docs/context-and-skills.md +1 -0
- package/docs/credentials-and-redaction.md +2 -2
- package/docs/database-persistence.md +9 -6
- package/docs/evaluations.md +122 -0
- package/docs/extensions.md +2 -2
- package/docs/host-security.md +20 -3
- package/docs/index.md +29 -17
- package/docs/mcp-tools.md +49 -4
- package/docs/migration.md +33 -3
- package/docs/multimodal-content.md +14 -6
- package/docs/observability.md +14 -6
- package/docs/performance.md +209 -0
- package/docs/postgres-persistence.md +6 -4
- package/docs/provider-conformance.md +1 -0
- package/docs/provider-packages.md +2 -0
- package/docs/providers/ai-sdk.md +113 -0
- package/docs/public-contracts.md +6 -5
- package/docs/rag.md +113 -0
- package/docs/release-and-install.md +100 -77
- package/docs/review-coverage-2026-07-15.md +193 -0
- package/docs/runs-and-usage.md +41 -4
- package/docs/server.md +139 -0
- package/docs/settings-auth-trust-security.md +5 -5
- package/docs/sqlite-persistence.md +4 -3
- package/docs/supervisors.md +71 -0
- package/docs/workflow-orchestration-primitives.md +19 -3
- package/docs/workflows.md +97 -23
- package/docs/working-and-semantic-memory.md +169 -0
- package/package.json +12 -2
- package/templates/init/README.md.tmpl +28 -0
- package/templates/init/env.example.tmpl +1 -0
- package/templates/init/gitignore.tmpl +11 -0
- package/templates/init/optional/evals-example.ts.tmpl +17 -0
- package/templates/init/optional/workflows-example.ts.tmpl +27 -0
- package/templates/init/package.json.tmpl +22 -0
- package/templates/init/providers.json +76 -0
- package/templates/init/src/agent.ts.tmpl +10 -0
- package/templates/init/src/index.ts.tmpl +12 -0
- package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
- package/templates/init/tsconfig.json.tmpl +15 -0
package/docs/performance.md
CHANGED
|
@@ -156,6 +156,215 @@ Provider SSE remained at the frozen 380 MiB/s / +1.7 MiB heap snapshot. Media an
|
|
|
156
156
|
|
|
157
157
|
The ledger percentage overhead is intentionally not a threshold: its no-ledger baseline is below 1 ms, making the percentage unstable while absolute added latency remains about 1 ms. JSONL's append path is intentionally O(n²) across repeated appends because it rereads for corruption/conflict checks; move production or high-volume workloads to SQLite/PostgreSQL rather than weakening validation.
|
|
158
158
|
|
|
159
|
+
### 0.0.5 Phase 0 baseline (2026-07-15)
|
|
160
|
+
|
|
161
|
+
Scope froze at commit `f5128a816ae204c52f3e2f089de71c99bd5de6d4`. Measurement host: Node v24.18.0, npm 11.16.0, Linux 7.1.3 x86_64, AMD Ryzen 9 PRO 7940HS (16 logical CPUs). Supported package runtime remains Node >=20. These are dated local comparison points, not portable CI wall-clock assertions.
|
|
162
|
+
|
|
163
|
+
| Surface | Workload | Result |
|
|
164
|
+
| --- | --- | --- |
|
|
165
|
+
| Network-free tests | `npm test` | 25.750 s; 1,475 tests, 1,450 pass, 25 explicit live skips, 0 fail |
|
|
166
|
+
| Release readiness | `npm run sdk:ready` | 54.341 s; typecheck, tests, examples, builds, and 24 dry-run packs pass |
|
|
167
|
+
| Provider/agent stream | One mock run with 5,000 one-character text deltas and a concurrently drained 8,192-event subscriber | 3.78 ms median |
|
|
168
|
+
| Tool dispatch | Six independent 20 ms tools, concurrency 1 | 121.05 ms median |
|
|
169
|
+
| Tool dispatch | Same calls, concurrency 2 | 60.65 ms median (2.00x speedup) |
|
|
170
|
+
| Workflow runner | Existing bounded 1,000-node chain, configured concurrency 8 | 9.66 ms median |
|
|
171
|
+
| Package artifacts | All 24 dry-run tarballs | 542,993 packed bytes; 2,084,900 unpacked bytes aggregate |
|
|
172
|
+
| Root artifact | `@arnilo/prism@0.0.4` dry-run tarball | 346.0 kB packed; 1.3 MB unpacked; 196 files |
|
|
173
|
+
| Installed workspace | Current root `node_modules` | 72 MiB |
|
|
174
|
+
|
|
175
|
+
Synthetic stream/tool/workflow values are medians of seven measured runs after one warm-up and contain no network, database, or exporter I/O. The temporary benchmark reused public `AgentSession`, `dispatchToolCallsInOrder`, and `@arnilo/prism-workflows` APIs; it was not added to CI because this phase records a baseline rather than creating hardware-sensitive tests.
|
|
176
|
+
|
|
177
|
+
Repository size at the same commit, counted from `src/` and `packages/` while excluding `dist/`:
|
|
178
|
+
|
|
179
|
+
| Area | Files | Lines |
|
|
180
|
+
| --- | ---: | ---: |
|
|
181
|
+
| Production TypeScript | 189 | 26,828 |
|
|
182
|
+
| Test TypeScript | 144 | 23,535 |
|
|
183
|
+
| Documentation Markdown | 70 | 12,662 |
|
|
184
|
+
| Numbered plans | 58 | 24,270 |
|
|
185
|
+
| TypeScript examples | 39 | 3,134 |
|
|
186
|
+
|
|
187
|
+
Prism has no project generator before Phase 5, so a generated-Prism-project install/build size is **not applicable** at this baseline. The closest current install figure is the 72 MiB development workspace; it is not a scaffold target. The comparison Mastra default scaffold measured during the review used 439 MB `node_modules`, 300 MB build output, and 427 installed packages. Phase 5 must establish a real generated Prism project baseline and keep unselected storage, telemetry, eval, memory, server, and workflow dependencies absent.
|
|
188
|
+
|
|
189
|
+
See [Review coverage — 2026-07-15](review-coverage-2026-07-15.md) for scope, primitive, package, and threat-boundary ownership.
|
|
190
|
+
|
|
191
|
+
### 0.0.5 Phase 2 verification (2026-07-15)
|
|
192
|
+
|
|
193
|
+
Same Phase 0 host and seven-run warm benchmark. Runtime correctness changes stayed inside frozen ceilings:
|
|
194
|
+
|
|
195
|
+
| Surface | Result |
|
|
196
|
+
| --- | --- |
|
|
197
|
+
| Network-free tests | 27.992 s; 1,485 tests, 1,460 pass, 25 explicit live skips, 0 fail |
|
|
198
|
+
| `npm run sdk:ready` | 55.598 s; typecheck, examples, tests, builds, and all 24 dry-run packs pass |
|
|
199
|
+
| Provider/agent stream, 5,000 deltas | 3.54 ms median (Phase 0: 3.78 ms) |
|
|
200
|
+
| Six 20 ms tools, concurrency 1 / 2 | 121.22 ms / 60.63 ms (2.00x speedup retained) |
|
|
201
|
+
| Workflow 1,000-node chain | 10.31 ms median (well below 1 s ceiling) |
|
|
202
|
+
| Root dry-run tarball | 361.2 kB packed, 1.3 MB unpacked, 197 files |
|
|
203
|
+
|
|
204
|
+
Usage aggregation performs one constant-size accumulator update per terminal provider turn. Telemetry retains only active span metadata and removes every terminal/detached entry. Complete media resolution is sequential, rejects item count and inline estimates before I/O, and retains at most the request budget plus one per-item-bounded candidate before failing an aggregate overflow. Sandbox output still streams into the existing bounded `OutputAccumulator`; no adapter-side response buffer was added.
|
|
205
|
+
|
|
206
|
+
### 0.0.5 Phase 4 verification (2026-07-15)
|
|
207
|
+
|
|
208
|
+
Optional `@arnilo/prism-evals` adds package-local scoring without changing core run latency. Validation stayed within the frozen release gate:
|
|
209
|
+
|
|
210
|
+
| Surface | Result |
|
|
211
|
+
| --- | --- |
|
|
212
|
+
| Network-free tests | 1,503 tests, 1,478 pass, 25 explicit live skips, 0 fail |
|
|
213
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 25 dry-run packs pass |
|
|
214
|
+
| Evals dry-run tarball | 35.4 kB unpacked package payload |
|
|
215
|
+
| Profile bundles | unchanged; evals remains opt-in until size/use review |
|
|
216
|
+
|
|
217
|
+
Experiment concurrency is capped at 32 workers and defaults to 1. Scorers operate on `AgentRunResult` references plus dataset item metadata rather than duplicating event ledgers.
|
|
218
|
+
|
|
219
|
+
### 0.0.5 Phase 5 verification (2026-07-15)
|
|
220
|
+
|
|
221
|
+
`prism init` lands as a stdlib-only CLI subcommand with checked-in templates under `templates/init/`.
|
|
222
|
+
|
|
223
|
+
| Surface | Result |
|
|
224
|
+
| --- | --- |
|
|
225
|
+
| Default generated sources | 8 files / ~3.3 KB |
|
|
226
|
+
| Default clean consumer install (`@arnilo/prism` + TypeScript tooling) | ~27.5 MB `node_modules` |
|
|
227
|
+
| Mastra comparator | 439 MB install / 300 MB build / 427 packages |
|
|
228
|
+
| Default dependencies | `@arnilo/prism` only; no storage, telemetry, eval, memory, server, or workflow packages unless `--with-*` / provider flags select them |
|
|
229
|
+
| Offline proof | packed core tarball → `npm install` → `npm run typecheck` → `npm test` (mock provider) |
|
|
230
|
+
|
|
231
|
+
### 0.0.5 Phase 6 verification (2026-07-15)
|
|
232
|
+
|
|
233
|
+
Optional `@arnilo/prism-provider-ai-sdk` adapts AI SDK `LanguageModelV4` streams to Prism without adding an AI SDK dependency to core.
|
|
234
|
+
|
|
235
|
+
| Surface | Result |
|
|
236
|
+
| --- | --- |
|
|
237
|
+
| Supported specification | `@ai-sdk/provider@^4` (`LanguageModelV4`) |
|
|
238
|
+
| Adapter behavior | incremental stream translation; unsupported content fails before `doStream`; abort owned by Prism `request.signal` |
|
|
239
|
+
| Network-free tests | 1,522 tests, 1,497 pass, 25 explicit live skips, 0 fail |
|
|
240
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 26 dry-run packs pass |
|
|
241
|
+
| AI SDK adapter dry-run tarball | 6.5 kB packed / 22.5 kB unpacked / 16 files |
|
|
242
|
+
| Profile bundles | unchanged; AI SDK adapter remains opt-in until size/use review |
|
|
243
|
+
| Publishable graph | 26 packages |
|
|
244
|
+
|
|
245
|
+
### 0.0.5 Phase 7 verification (2026-07-15)
|
|
246
|
+
|
|
247
|
+
Optional `@arnilo/prism-memory` adds working memory and semantic recall without changing core session stores.
|
|
248
|
+
|
|
249
|
+
| Surface | Result |
|
|
250
|
+
| --- | --- |
|
|
251
|
+
| Contracts | package-owned `Embedder`, `VectorStore`, `WorkingMemoryStore`, `createMemory` |
|
|
252
|
+
| Adapters | in-memory reference + PostgreSQL/pgvector production path |
|
|
253
|
+
| Injection | existing `ContextProvider` seam; opt-in working-memory processor |
|
|
254
|
+
| Profile bundles | unchanged; memory remains opt-in until size/use review |
|
|
255
|
+
| Publishable graph | 27 packages |
|
|
256
|
+
| Network-free tests | 1,538 tests, 1,513 pass, 25 explicit live skips, 0 fail |
|
|
257
|
+
| `npm run sdk:ready` | pass |
|
|
258
|
+
| Memory dry-run tarball | 17.9 kB packed / 76.6 kB unpacked / 32 files |
|
|
259
|
+
|
|
260
|
+
### 0.0.5 Phase 8 verification (2026-07-15)
|
|
261
|
+
|
|
262
|
+
Durable human suspension extends existing workflow checkpoint JSON/CAS; no worker polling loop, package, dependency, or database migration was added.
|
|
263
|
+
|
|
264
|
+
| Surface | Result |
|
|
265
|
+
| --- | --- |
|
|
266
|
+
| Focused workflow suite | 43 tests pass, 0 fail |
|
|
267
|
+
| Network-free tests | 1,547 tests, 1,522 pass, 25 explicit live skips, 0 fail |
|
|
268
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 27 dry-run packs pass |
|
|
269
|
+
| Workflow dry-run tarball | 25.7 kB packed / 121.6 kB unpacked / 34 files |
|
|
270
|
+
| Coordinator behavior | `suspended` absent from queued/running poll; zero worker/lease retained |
|
|
271
|
+
| Storage | existing bounded checkpoint JSON/category; no SQLite/PostgreSQL migration |
|
|
272
|
+
|
|
273
|
+
### 0.0.5 Phase 9 verification (2026-07-16)
|
|
274
|
+
|
|
275
|
+
Optional `@arnilo/prism-rag` reuses Phase 7 vector contracts and adds no core path, parser dependency, network loader, or profile activation.
|
|
276
|
+
|
|
277
|
+
| Surface | Result |
|
|
278
|
+
| --- | --- |
|
|
279
|
+
| Focused RAG suite | 9 tests pass, 0 fail |
|
|
280
|
+
| Network-free tests | 1,561 tests, 1,536 pass, 25 explicit live skips, 0 fail |
|
|
281
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 28 dry-run packs pass |
|
|
282
|
+
| RAG dry-run tarball | 9.0 kB packed / 34.6 kB unpacked / 22 files |
|
|
283
|
+
| Index bounds | chunk/document/count/metadata caps; embed batches default 32, hard 128 |
|
|
284
|
+
| Retrieval bounds | top-K default 5/hard 32; candidates default 20/hard 128; result 64/512 KiB; context 2,000/8,000 estimated tokens |
|
|
285
|
+
| Profile bundles | unchanged; RAG and memory remain explicit opt-ins |
|
|
286
|
+
|
|
287
|
+
### 0.0.5 Phase 10 verification (2026-07-16)
|
|
288
|
+
|
|
289
|
+
Optional `@arnilo/prism-server` and MCP server-direction APIs compose existing agent/workflow/tool/SDK primitives; no core path, framework/listener, auth provider, database, or profile activation was added.
|
|
290
|
+
|
|
291
|
+
| Surface | Result |
|
|
292
|
+
| --- | --- |
|
|
293
|
+
| Focused server suites | 6 Web handler tests + 4 MCP server tests pass; existing 12 MCP client tests remain green |
|
|
294
|
+
| Network-free tests | 1,576 tests, 1,551 pass, 25 explicit live skips, 0 fail |
|
|
295
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 29 dry-run packs pass |
|
|
296
|
+
| Server dry-run tarball | 8.4 kB packed / 34.4 kB unpacked / 12 files |
|
|
297
|
+
| MCP dry-run tarball | 11.6 kB packed / 45.0 kB unpacked / 20 files |
|
|
298
|
+
| Web handler bounds | request 64 KiB, result 1 MiB, event 64 KiB, stream 10 MiB/10k events, queue 128, concurrency 16, timeout 120 s by default; all have hard caps |
|
|
299
|
+
| MCP server bounds | call result 1 MiB, calls 16, timeout 60 s; HTTP request 1 MiB, response 2 MiB, requests 32 by default; all have hard caps |
|
|
300
|
+
| Profile bundles | unchanged; server remains explicit opt-in |
|
|
301
|
+
|
|
302
|
+
### 0.0.5 Phase 11 verification (2026-07-16)
|
|
303
|
+
|
|
304
|
+
Workflow schedules, background runs, composition, state, and replay reuse the existing workflow package plus generic checkpoint/lease stores. No package, runtime dependency, SQL migration, listener, cron parser, or auto-started worker was added.
|
|
305
|
+
|
|
306
|
+
| Surface | Result |
|
|
307
|
+
| --- | --- |
|
|
308
|
+
| Focused workflow/server suites | 54 workflow tests + 8 Web handler tests pass, 0 fail |
|
|
309
|
+
| Network-free tests | 1,589 tests, 1,564 pass, 25 explicit live skips, 0 fail |
|
|
310
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 29 dry-run packs pass |
|
|
311
|
+
| Workflow dry-run tarball | 34.7 kB packed / 171.5 kB unpacked / 38 files |
|
|
312
|
+
| Server dry-run tarball | 9.9 kB packed / 45.2 kB unpacked / 12 files |
|
|
313
|
+
| Synthetic schedule bound | 100 in-memory creates: 1.28 ms; scan 100 / claim+enqueue 16 due fires: 6.64 ms |
|
|
314
|
+
| Synthetic composition/replay | depth-8 nested run: 2.73 ms; 100-node source: 35.42 ms; replay 50 nodes: 21.17 ms |
|
|
315
|
+
| State/replay ceilings | state 64/512 KiB; history 32/128; nested depth 8/32; replay depth 8/32 default/hard |
|
|
316
|
+
| Schedule ceilings | page 100/500; claims 16/256; input 256 KiB/1 MiB; 1s idle timer; 30s fire lease defaults |
|
|
317
|
+
|
|
318
|
+
Synthetic timings are one local Node v24.18.0 run over memory adapters with no network/database I/O; finite limits and behavior tests, not wall-clock numbers, are CI gates.
|
|
319
|
+
|
|
320
|
+
### 0.0.5 Phase 12 verification (2026-07-16)
|
|
321
|
+
|
|
322
|
+
Run feedback adds no package or runtime dependency. Memory/SQLite/PostgreSQL implementations share bounded append/query/delete semantics; OTel projection accepts only fixed scalar metadata.
|
|
323
|
+
|
|
324
|
+
| Surface | Result |
|
|
325
|
+
| --- | --- |
|
|
326
|
+
| Focused feedback/eval/SQLite/OTel tests | 35 tests pass, 0 fail; PostgreSQL DDL suite passes and live feedback conformance is env-gated |
|
|
327
|
+
| Synthetic memory feedback | 1,000 bounded appends: 3.82 ms; 100 filtered 100-row queries over 1,000 records: 11.53 ms |
|
|
328
|
+
| Core dry-run tarball | 398.1 kB packed / 1.4 MB unpacked / 219 files |
|
|
329
|
+
| Evals dry-run tarball | 9.8 kB packed / 38.4 kB unpacked / 26 files |
|
|
330
|
+
| OTel dry-run tarball | 6.3 kB packed / 26.5 kB unpacked / 8 files |
|
|
331
|
+
| SQLite/PostgreSQL tarballs | 17.7/18.1 kB packed; 89.8/89.8 kB unpacked |
|
|
332
|
+
| Feedback limits | comment 4/16 KiB; tags 16/64; links 16/64; metadata 16/64 KiB; pages 100/500 default/hard |
|
|
333
|
+
|
|
334
|
+
Metrics came from one local Node v24.18.0 memory-adapter run. SQL correctness/indexing/migration behavior and hard bounds are gates; local timings are not release thresholds.
|
|
335
|
+
|
|
336
|
+
### 0.0.5 Phase 13 verification (2026-07-16)
|
|
337
|
+
|
|
338
|
+
Supervisor/A2A stays in one optional zero-runtime-dependency package; core and profile bundles gained no import, listener, worker, protocol SDK, or network activation.
|
|
339
|
+
|
|
340
|
+
| Surface | Result |
|
|
341
|
+
| --- | --- |
|
|
342
|
+
| Focused supervisor/A2A suite | 11 tests pass, 0 fail; local delegation, policy/budget/abort/redaction, card signatures, server/client/stream bounds |
|
|
343
|
+
| Synthetic local delegation | 100 sequential mock child results: 11.83 ms |
|
|
344
|
+
| Synthetic in-process A2A | 100 card discovery + JSON-RPC mock round trips: 34.17 ms |
|
|
345
|
+
| Supervisor dry-run tarball | 15.3 kB packed / 69.4 kB unpacked / 22 files |
|
|
346
|
+
| Local hard ceilings | depth 16; active 32; message 1 MiB; steps 64; tools 256; tokens 1m; timeout 30m; event queue 4096 |
|
|
347
|
+
| A2A hard ceilings | request/card/event 1 MiB; response 8 MiB; stream 64 MiB/100k events; concurrency 256; timeout 30m |
|
|
348
|
+
|
|
349
|
+
Timings are one local Node v24.18.0 run over mock agents and an in-process fetch adapter. Bounds, protocol validation, signature/auth/origin checks, and offline behavior tests are release gates; timings are not thresholds.
|
|
350
|
+
|
|
351
|
+
### 0.0.5 Phase 14 release-candidate verification (2026-07-16)
|
|
352
|
+
|
|
353
|
+
| Surface | Result |
|
|
354
|
+
| --- | --- |
|
|
355
|
+
| Default network-free test | 32.247 s, below 60 s budget |
|
|
356
|
+
| Full SDK readiness | 70.560 s; build/typecheck/examples/tests/30 pack dry-runs |
|
|
357
|
+
| Test matrix | 1,618 total; 1,593 pass; 25 explicit live skips; 0 fail |
|
|
358
|
+
| Node compatibility | Node 20.20.2 imports 44 built root/package export targets; Node 24.18.0 runs full matrix |
|
|
359
|
+
| PostgreSQL/pgvector | 29 live checks pass in fresh `pgvector/pgvector:pg16` container |
|
|
360
|
+
| Packed artifact set | 30 tarballs / 699 files; post-bundle snapshot ~690.6 kB packed / 2.64 MB unpacked |
|
|
361
|
+
| Core artifact | post-bundle snapshot ~403.7 kB packed / 1.46 MB unpacked / 221 files |
|
|
362
|
+
| Generated default project | under 50 KiB source and under 50 MiB installed; packed-core typecheck/test pass |
|
|
363
|
+
| Fresh packed journey | 30 packages install/import and Phase 1-13 optional composition pass in ~8.0 s |
|
|
364
|
+
| Registry/publish preview | 30/30 versions available; 30/30 dependency-ordered provenance dry-runs pass |
|
|
365
|
+
|
|
366
|
+
No performance ceiling was raised. Core grew from Phase 0's 346.0 kB packed baseline to ~403.7 kB after documented APIs/templates, while the full package set remains ~690.6 kB packed. Follow-up review includes all six Phase 4-13 capability packages through `prism-all` and AI SDK interoperability through `prism-providers`; focused base/code/SDK profiles remain unchanged and no capability auto-activates. Manifest tarballs remain tiny: providers 1.4 kB and all 1.6 kB packed.
|
|
367
|
+
|
|
159
368
|
## Related APIs
|
|
160
369
|
|
|
161
370
|
- [Agent events](agent-events.md): `SubscribeOptions` and `event_subscriber_overflow` event details.
|
|
@@ -39,6 +39,7 @@ import { createPostgresPersistence } from "@arnilo/prism-session-store-postgres"
|
|
|
39
39
|
| `connectionString` | `string` | Create an adapter-owned bounded pool when `pool` is omitted. |
|
|
40
40
|
| `schema` | `string` | PostgreSQL schema for Prism tables. Defaults to `"prism"`. Validated and double-quoted. |
|
|
41
41
|
| `poolMax` | `number` | Maximum pool size for adapter-owned pools. Defaults to `10`. |
|
|
42
|
+
| `feedbackRedactor` | `SecretRedactor` | Optional redaction for feedback comment/tags/metadata before insert. |
|
|
42
43
|
| `poolConfig` | `PoolConfig` | Additional `pg` options (TLS, idle timeout, application name, etc.). |
|
|
43
44
|
|
|
44
45
|
Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backup/retention enforcement.
|
|
@@ -54,7 +55,7 @@ Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backu
|
|
|
54
55
|
| `SessionStore.readBranchPath` | Recursive ancestor query from `leafId` (or latest leaf) in root→leaf order. |
|
|
55
56
|
| `RunLedger.append*` | Inserts run/event/tool/usage rows; events receive monotonic per-run `sequence` values. |
|
|
56
57
|
| `ProductionPersistenceStore.query*` | Parameterized cursor pagination on indexed columns with tenant/account/user filters. |
|
|
57
|
-
| `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, and
|
|
58
|
+
| `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, bounded pagination, and workflow suspended/denied/schedule/state/replay values without a schema migration. |
|
|
58
59
|
| `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
|
|
59
60
|
| `close()` | Ends the pool when the adapter created it from `connectionString`. |
|
|
60
61
|
|
|
@@ -115,7 +116,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
|
|
|
115
116
|
- The package is optional and workspace-local; `@arnilo/prism` core has no PostgreSQL dependency.
|
|
116
117
|
- Schema names must match `^[a-zA-Z_][a-zA-Z0-9_]*$`; the adapter quotes them and never interpolates user values into identifier positions.
|
|
117
118
|
- `SessionAppendOptions` idempotency rows are durable in `prism_session_append_idempotency` and survive reopen.
|
|
118
|
-
- Schema version **
|
|
119
|
+
- Schema version **3** applies `001_init`, additive `002_usage_scope`, and `003_run_feedback`. Migration 003 adds immutable `prism_run_feedback` rows with run FK/cascade deletion and owner/run/trace cursor indexes. `persistence.feedback` validates exact run ownership, bounds/redacts through optional `feedbackRedactor`, queries bounded pages, and deletes only exact-owned IDs. SQLite shares the same model with dialect-local DDL.
|
|
119
120
|
- Pass an existing `pg` `Pool` when your host already manages pooling, TLS, and credential rotation.
|
|
120
121
|
|
|
121
122
|
## Security and performance notes
|
|
@@ -126,7 +127,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
|
|
|
126
127
|
- **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
|
|
127
128
|
- **Bounded pool.** Adapter-owned pools default to `max: 10`. Hosts with heavy concurrency should supply their own pool sizing.
|
|
128
129
|
- **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid sequential scans.
|
|
129
|
-
- **Migration locking.** `pg_advisory_xact_lock` prevents concurrent
|
|
130
|
+
- **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once.
|
|
130
131
|
- **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns participate in query filters; hosts must still scope writes correctly.
|
|
131
132
|
- **Benchmark target.** Indexed append + paginated branch read on a warm pool should stay under **50 ms p95** for local/CI-sized datasets (≤100k entries per session); measure with your pool size and hardware before production sizing.
|
|
132
133
|
|
|
@@ -137,5 +138,6 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
|
|
|
137
138
|
- [Session store conformance](session-store-conformance.md): `assertSessionStoreConforms` / `runSessionStoreConformance`.
|
|
138
139
|
- [Run ledger conformance](run-ledger-conformance.md): `assertRunLedgerConforms` / `runRunLedgerConformance`.
|
|
139
140
|
- [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): package matrix and threat model.
|
|
140
|
-
- [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` for durable
|
|
141
|
+
- [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` and `createWorkflowSchedules()` for durable background execution and schedules.
|
|
142
|
+
- [Working and semantic memory](working-and-semantic-memory.md): optional `@arnilo/prism-memory` PostgreSQL/pgvector working + semantic stores (separate from session/run persistence).
|
|
141
143
|
- [Migration guide](migration.md): moving from JSONL/in-memory to database-backed persistence.
|
|
@@ -149,5 +149,6 @@ The helpers are a testing subpath only. Provider packages can use them with thei
|
|
|
149
149
|
|
|
150
150
|
- [Provider layer](provider-layer.md): `AIProvider`, provider events, and mock provider.
|
|
151
151
|
- [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters.
|
|
152
|
+
- [AI SDK provider adapter](providers/ai-sdk.md): optional `LanguageModelV4` bridge tested with a fake AI SDK model.
|
|
152
153
|
- [OpenAI-compatible provider](providers/openai-compatible.md): optional provider adapter tested with mocked streams.
|
|
153
154
|
- [Public contracts](public-contracts.md): provider request/event/usage contracts.
|
|
@@ -63,6 +63,8 @@ First-party providers map generic `ModelConfig.parameters.maxTokens` to real out
|
|
|
63
63
|
|
|
64
64
|
Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md). Each package starts with a side-effect-free `create*ProviderPackage()` export, README, TypeScript build, network-free default tests, and real opt-in live smoke tests.
|
|
65
65
|
|
|
66
|
+
Phase 6 also adds optional [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md), which adapts a host-owned AI SDK `LanguageModelV4` to Prism's `AIProvider`. It joins `@arnilo/prism-providers` as the seventh adapter while remaining independent from the six HTTP implementations.
|
|
67
|
+
|
|
66
68
|
Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
|
|
67
69
|
|
|
68
70
|
These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
|
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
# AI SDK provider adapter
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-provider-ai-sdk` adapts a host-supplied AI SDK `LanguageModelV4` into a Prism `AIProvider`. It maps Prism messages, tools, and structured-output options into `doStream` call options, then translates stream parts into Prism provider events incrementally.
|
|
6
|
+
|
|
7
|
+
Supported specification: `@ai-sdk/provider` **v4** (`specificationVersion: "v4"`). Core `@arnilo/prism` does not depend on the AI SDK.
|
|
8
|
+
|
|
9
|
+
## When to use it
|
|
10
|
+
|
|
11
|
+
Use this package when a host already creates AI SDK language models and wants them inside Prism agent/session loops without adding another first-party HTTP provider.
|
|
12
|
+
|
|
13
|
+
Do not use it as a credential store, model catalog, or high-level `streamText`/`generateText` replacement. Hosts keep owning credentials inside the supplied model.
|
|
14
|
+
|
|
15
|
+
## Inputs / request
|
|
16
|
+
|
|
17
|
+
```ts
|
|
18
|
+
import { createAiSdkProvider } from "@arnilo/prism-provider-ai-sdk";
|
|
19
|
+
|
|
20
|
+
createAiSdkProvider(options: {
|
|
21
|
+
model: LanguageModelV4;
|
|
22
|
+
id?: string;
|
|
23
|
+
}): AIProvider
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
| Field | Type | Purpose |
|
|
27
|
+
| --- | --- | --- |
|
|
28
|
+
| `model` | `LanguageModelV4` | Host-owned AI SDK language model. |
|
|
29
|
+
| `id` | `string` | Prism provider id. Defaults to `ai-sdk:<model.provider>` or `ai-sdk`. |
|
|
30
|
+
|
|
31
|
+
Mapped request surfaces:
|
|
32
|
+
|
|
33
|
+
| Prism | AI SDK |
|
|
34
|
+
| --- | --- |
|
|
35
|
+
| `messages` | `LanguageModelV4Prompt` |
|
|
36
|
+
| `tools` | `LanguageModelV4FunctionTool[]` with JSON Schema `inputSchema` |
|
|
37
|
+
| `options.structuredOutput` | `responseFormat: { type: "json", name, schema }` |
|
|
38
|
+
| `model.parameters` | `maxOutputTokens`, `temperature`, `topP`, `topK`, penalties, `seed`, `stopSequences` |
|
|
39
|
+
| `request.signal` | `abortSignal` (always wins over adapter options) |
|
|
40
|
+
| `options.headers` | extension headers only; model owns auth |
|
|
41
|
+
|
|
42
|
+
Unsupported content fails before `doStream` (for example unresolved `resourceUri`, audio/file/document without declared capability, `tool_call_delta` in history, non-text system content).
|
|
43
|
+
|
|
44
|
+
## Outputs / response / events
|
|
45
|
+
|
|
46
|
+
| AI SDK stream part | Prism event |
|
|
47
|
+
| --- | --- |
|
|
48
|
+
| `text-delta` | `content_delta` text |
|
|
49
|
+
| `reasoning-delta` | `content_delta` thinking |
|
|
50
|
+
| `tool-input-start` / `tool-input-delta` | `tool_call_delta` |
|
|
51
|
+
| `tool-call` (client-executed) | `tool_call` |
|
|
52
|
+
| `finish` usage | `usage` then `done` |
|
|
53
|
+
| `error` / thrown / abort | redacted `error` |
|
|
54
|
+
|
|
55
|
+
Provider-executed tool calls, files/sources/custom parts, warnings, and raw chunks are ignored rather than silently converted into unsupported Prism content.
|
|
56
|
+
|
|
57
|
+
## Request/response example
|
|
58
|
+
|
|
59
|
+
```json
|
|
60
|
+
{
|
|
61
|
+
"prompt": [{ "role": "user", "content": [{ "type": "text", "text": "hello" }] }],
|
|
62
|
+
"tools": [{ "type": "function", "name": "echo", "inputSchema": { "type": "object" } }],
|
|
63
|
+
"responseFormat": {
|
|
64
|
+
"type": "json",
|
|
65
|
+
"name": "Answer",
|
|
66
|
+
"schema": { "type": "object", "properties": { "ok": { "type": "boolean" } } }
|
|
67
|
+
}
|
|
68
|
+
}
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
## Implementation example
|
|
72
|
+
|
|
73
|
+
```ts
|
|
74
|
+
import { createAgent } from "@arnilo/prism";
|
|
75
|
+
import { createAiSdkProvider } from "@arnilo/prism-provider-ai-sdk";
|
|
76
|
+
|
|
77
|
+
const provider = createAiSdkProvider({ model: hostCreatedLanguageModelV4 });
|
|
78
|
+
|
|
79
|
+
const agent = createAgent({
|
|
80
|
+
provider,
|
|
81
|
+
model: {
|
|
82
|
+
provider: provider.id,
|
|
83
|
+
model: hostCreatedLanguageModelV4.modelId,
|
|
84
|
+
capabilities: { tools: true, streaming: true, structuredOutput: true, input: ["text"] },
|
|
85
|
+
},
|
|
86
|
+
});
|
|
87
|
+
|
|
88
|
+
const result = await agent.createSession().run("Summarize this");
|
|
89
|
+
console.log(result.text);
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
## Extension and configuration notes
|
|
93
|
+
|
|
94
|
+
- Peer dependency: `@ai-sdk/provider@^4.0.0`. Upgrade policy tracks one specification major at a time.
|
|
95
|
+
- First-party HTTP providers remain independent; this adapter is available directly, through `@arnilo/prism-providers`, or through `@arnilo/prism-all`. Installation does not select a model or invoke AI SDK.
|
|
96
|
+
- `options.compat` / `options.extra` pass through as AI SDK `providerOptions.prism`.
|
|
97
|
+
- Export helpers `toAiSdkCallOptions`, `toAiSdkPrompt`, and `mapAiSdkStream` for tests and custom hosts.
|
|
98
|
+
|
|
99
|
+
## Security and performance notes
|
|
100
|
+
|
|
101
|
+
- Host credentials stay inside the supplied AI SDK model. The adapter never reads env keys or credential stores.
|
|
102
|
+
- Abort and resource limits come from Prism `request.signal`; adapter options cannot replace that bound.
|
|
103
|
+
- Stream parts are translated incrementally with no full-response buffering and no duplicate model call.
|
|
104
|
+
- Unsupported content fails closed before model invocation. Errors use Prism `providerError` redaction.
|
|
105
|
+
- Provider metadata/warnings are not emitted as prompt or tool content.
|
|
106
|
+
|
|
107
|
+
## Related APIs
|
|
108
|
+
|
|
109
|
+
- [Provider packages](../provider-packages.md)
|
|
110
|
+
- [Provider conformance](../provider-conformance.md)
|
|
111
|
+
- [Provider layer](../provider-layer.md)
|
|
112
|
+
- [Structured output](../structured-output.md)
|
|
113
|
+
- [Agent session runtime](../agent-session-runtime.md)
|
package/docs/public-contracts.md
CHANGED
|
@@ -15,7 +15,7 @@ Current contract groups:
|
|
|
15
15
|
- Extensions/middleware: `ExtensionLifecycleEventName`, `ExtensionEvent`, `Extension`, `ExtensionAPI`, `MiddlewareHookName`, `Middleware`, `MiddlewareNext`, `MiddlewareRegistry`
|
|
16
16
|
- Configuration/manifests: `ConfigProvider`, `ConfigLayer`, `ConfigLoadContext`, `PrismManifest`, `ManifestContributionDeclaration`, `ManifestResourceDeclaration`, `ManifestContributionKind`
|
|
17
17
|
- Stores/resources/settings/credentials/compaction/retry/cache helpers: `SessionEntry`, `SessionStore`, `StoreFactory`, `Resource`, `ResourceLoader`, `ResourceLoadContext`, `SettingsProvider`, `CredentialRequest`, `Credential`, `CredentialResolver`, `CompactionStrategy`, `CompactionContext`, `CompactionResult`, `CompactionOptions`, `CompactionMiddlewarePayload`, `CompactionEntryData`, `DefaultCompactionStrategyOptions`, `RetryPolicy`, `RetryContext`, `RetryDecision`, `RetryOptions`, `RetryMiddlewarePayload`, `DefaultRetryPolicyOptions`, `CacheUsageReport`, `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, `cacheUsageReport`
|
|
18
|
-
- Production persistence (adapter-facing): `ProductionPersistenceStore`, `CheckpointStore`, `CheckpointKey`, `CheckpointSaveInput`, `CheckpointRecord`, `CheckpointQuery`, `LeaseStore`, `LeaseKey`, `LeaseAcquireInput`, `LeaseClaimInput`, `LeaseRecord`, `PersistencePage`, `PersistenceQuery`, `OwnershipScope`, `SessionRecord`, `SessionQuery`, `BranchRecord`, `BranchQuery`, `SessionEntryQuery`, `RunRecord`, `RunQuery`, `AgentEventRecord`, `AgentEventQuery`, `ToolCallRecord`, `ToolCallQuery`, `UsageRecord`, `UsageQuery`, `AgentDefinitionRecord`, `AgentDefinitionQuery`, `RetentionPolicy`, `RetentionPolicyQuery`, `MigrationRecord`, `MigrationQuery`
|
|
18
|
+
- Production persistence (adapter-facing): `ProductionPersistenceStore`, `CheckpointStore`, `CheckpointKey`, `CheckpointSaveInput`, `CheckpointRecord`, `CheckpointQuery`, `LeaseStore`, `LeaseKey`, `LeaseAcquireInput`, `LeaseClaimInput`, `LeaseRecord`, `PersistencePage`, `PersistenceQuery`, `OwnershipScope`, `SessionRecord`, `SessionQuery`, `BranchRecord`, `BranchQuery`, `SessionEntryQuery`, `RunRecord`, `RunQuery`, `RunFeedbackRecord`, `RunFeedbackStore`, `RunFeedbackQuery`, `AgentEventRecord`, `AgentEventQuery`, `ToolCallRecord`, `ToolCallQuery`, `UsageRecord`, `UsageQuery`, `AgentDefinitionRecord`, `AgentDefinitionQuery`, `RetentionPolicy`, `RetentionPolicyQuery`, `MigrationRecord`, `MigrationQuery`
|
|
19
19
|
|
|
20
20
|
## When to use it
|
|
21
21
|
|
|
@@ -142,9 +142,10 @@ Important request shapes:
|
|
|
142
142
|
| `SystemPromptContribution` | Explicit caller-selected prompt layer with source, mode, text, and metadata. |
|
|
143
143
|
| `ConfigLayer` | Named JSON config layer consumed by `mergeConfigLayers()`. |
|
|
144
144
|
| `PrismManifest` | Data-only package manifest with config defaults, contribution declarations, and resource declarations. |
|
|
145
|
-
| `ProductionPersistenceStore` | Adapter-facing interface for durable, paginated, multi-tenant storage plus optional `checkpoints?: CheckpointStore` and `
|
|
145
|
+
| `ProductionPersistenceStore` | Adapter-facing interface for durable, paginated, multi-tenant storage plus optional `checkpoints?: CheckpointStore`, `leases?: LeaseStore`, and `feedback?: RunFeedbackStore`. No SQL/ORM/host file storage/network dependency. |
|
|
146
146
|
| `CheckpointStore` | Generic versioned checkpoint capability: save/load/bounded-list/delete by namespace and key, with ownership, exact-version CAS, and lease fencing. `createMemoryCheckpointStore()` is the reference implementation. |
|
|
147
147
|
| `LeaseStore` | Atomic acquire/renew/release/get by namespace and key, with opaque claim tokens, expiry, ownership scope, and monotonically increasing takeover fences. `createMemoryLeaseStore()` is the reference implementation. |
|
|
148
|
+
| `RunFeedbackStore` | Immutable append, bounded owned query, and owned deletion for ratings/comments/tags linked to existing run/trace/evaluation IDs. `createMemoryRunFeedbackStore()` is the reference implementation. |
|
|
148
149
|
| `EventMultiplexer<T>` | Generic bounded fan-in from async sources. `createEventMultiplexer()` owns queue limits, overflow policy, abort, source teardown, and close behavior. |
|
|
149
150
|
| `PersistencePage<T>` | Cursor-paginated result page: `items`, optional `nextCursor`, optional `total`. |
|
|
150
151
|
| `PersistenceQuery` | Common pagination controls: `cursor?`, `limit?`, `order?: "asc" \| "desc"`. |
|
|
@@ -155,7 +156,7 @@ Important request shapes:
|
|
|
155
156
|
| `RunRecord` / `RunQuery` | Stored run and filters: session, branch, status, timestamps, ownership. |
|
|
156
157
|
| `AgentEventRecord` / `AgentEventQuery` | Event ledger row with `redacted` flag and filters by type, session, run, entry, timestamp, ownership. |
|
|
157
158
|
| `ToolCallRecord` / `ToolCallQuery` | Tool-call row with `redacted` flag and filters by name, status, session, run, entry, timestamps, ownership. |
|
|
158
|
-
| `UsageRecord` / `UsageQuery` | Usage row and filters: session, run, entry, recorded-at range, ownership. |
|
|
159
|
+
| `UsageRecord` / `UsageQuery` | Usage row and filters: `provider_turn`/`run_total` scope, turn/attempt, session, run, entry, recorded-at range, ownership. |
|
|
159
160
|
| `CacheUsageReport` | Numeric cache diagnostics from normalized `Usage`: read/write tokens, hit rate, estimated savings, and optional currency. |
|
|
160
161
|
| `AgentDefinitionRecord` / `AgentDefinitionQuery` | Versioned agent-definition snapshot and filters. Does not store credentials or provider instances. |
|
|
161
162
|
| `RetentionPolicy` / `RetentionPolicyQuery` | Retention policy and filters: age, entry count, byte limits, archive store, applied kinds. |
|
|
@@ -412,8 +413,8 @@ void credentials;
|
|
|
412
413
|
- Contracts are host-owned and package-friendly. External packages can implement `AIProvider`, `ToolDefinition`, `CommandDefinition`, `AgentDefinition`, `InputBuilder`, `PromptBuilder`, `Middleware`, `ContextProvider`, `Skill`, `Extension`, config providers, data-only manifests, compaction strategies, store factories, resource loaders, settings providers, and credential resolvers.
|
|
413
414
|
- `ExtensionAPI` is implemented by the extension kernel. It exposes explicit registries, ordered middleware registration, ordered event subscription/emission, and registration methods for Phase 2 contribution categories.
|
|
414
415
|
- `AgentConfig.provider` can hold a direct provider instance for simple host wiring. Hosts that need config-driven selection should use `ModelConfig.provider` with explicit `createProviderRegistry()` / `createModelRegistry()` objects; Prism does not create a hidden global provider registry.
|
|
415
|
-
- `SettingsProvider` and `CredentialResolver` are explicit dependencies. Prism must not hide global settings or credentials behind these contracts, and `CredentialResolver` should be passed only to the edge that needs a credential.
|
|
416
|
-
-
|
|
416
|
+
- `SettingsProvider` and `CredentialResolver` are explicit dependencies. Prism must not hide global settings or credentials behind these contracts, and `CredentialResolver` should be passed only to the edge that needs a credential. These seams are host-owned outside `AgentConfig`; the session runtime does not call `settings.get()` or `credentials.resolve()`.
|
|
417
|
+
- Extension loading is host-owned outside `AgentConfig`; the session runtime does not load extensions or call `Extension.setup()`. Use `createExtensionKernel().load(...)` before creating an agent, then pass selected contributions into `AgentConfig`.
|
|
417
418
|
- `PrismManifest` is data-only. It can describe contribution modules/resources and config defaults, but parsing it does not import modules, execute package code, or mutate registries.
|
|
418
419
|
- Resource helper functions decode resources from a caller-provided `ResourceLoader`; Prism does not include host file storage, network, package, or URI router loaders.
|
|
419
420
|
- `createDefaultInputBuilder()` is a small default implementation of `InputBuilder`. It is replaceable and only loads explicit URI resources through a caller-provided `ResourceLoader`.
|
package/docs/rag.md
ADDED
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
# Retrieval-augmented generation (RAG)
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-rag` is an optional package for deterministic plain-text/Markdown chunking, bounded embedding/vector indexing, filtered semantic retrieval, stable citations, and explicit `ContextProvider` injection. It reuses `Embedder` and `VectorStore` from `@arnilo/prism-memory`; Prism core input assembly is unchanged.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
Use it when a host already owns trusted document text and needs small retrieval primitives without a document framework. Do not use it for PDF/HTML/LaTeX parsing, semantic chunking, metadata extraction agents, reranker pipelines, GraphRAG, crawling, URL fetching, or filesystem discovery.
|
|
10
|
+
|
|
11
|
+
## Inputs / request
|
|
12
|
+
|
|
13
|
+
Chunking:
|
|
14
|
+
|
|
15
|
+
| API/field | Meaning |
|
|
16
|
+
| --- | --- |
|
|
17
|
+
| `chunkText(text, options)` | Character-bounded plain-text chunks |
|
|
18
|
+
| `chunkMarkdown(markdown, options)` | Same engine, preferring heading/paragraph boundaries |
|
|
19
|
+
| `sourceId` | Required stable, non-secret source identifier |
|
|
20
|
+
| `size` / `overlap` | Character ceiling and repeated context |
|
|
21
|
+
| `metadata` | JSON metadata copied to every chunk |
|
|
22
|
+
|
|
23
|
+
Index/retrieve:
|
|
24
|
+
|
|
25
|
+
| Field | Required | Meaning |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| `embedder` / `store` | yes | Phase 7 `Embedder` and `VectorStore` |
|
|
28
|
+
| `scope` | yes | `{ tenantId, resourceId, corpusId }`; corpus maps to vector thread isolation |
|
|
29
|
+
| `chunks` | indexing | `RagChunk[]` from package chunkers or compatible host parser |
|
|
30
|
+
| `topK` / `queryCandidates` | retrieval | Returned result count and bounded pre-filter candidates |
|
|
31
|
+
| `filter` | no | Shallow JSON metadata equality filter |
|
|
32
|
+
| `redactor` / `secrets` | no | Redact before embedding, persistence, and injection |
|
|
33
|
+
| `signal` | no | Abort embedding, vector operations, and batch progression |
|
|
34
|
+
|
|
35
|
+
## Outputs / response / events
|
|
36
|
+
|
|
37
|
+
- `chunkText()` / `chunkMarkdown()` return frozen `RagChunk[]` with `sourceId`, zero-based index, offsets, and stable IDs such as `guide#0001`.
|
|
38
|
+
- `indexChunks()` returns `{ indexed, sourceIds }` after bounded batch upserts.
|
|
39
|
+
- `retrieveContext()` returns `{ query, text, hits, citations, truncated }`. Rendered text uses `[citation-id] text` blocks.
|
|
40
|
+
- `createRagContextProvider()` returns one ordinary context provider. Empty queries/results contribute no block.
|
|
41
|
+
- No events, tools, permissions, provider calls, loaders, or network requests are added.
|
|
42
|
+
|
|
43
|
+
Default hard ceilings include 1,000/16,384 chunk characters, 100/4,096 overlap, 1,048,576/8,388,608 document characters, 2,048/8,192 chunks, 32/128 embed batch, top-K 5/32, candidates 20/128, result 64/512 KiB, and context 2,000/8,000 estimated tokens.
|
|
44
|
+
|
|
45
|
+
## Request/response example
|
|
46
|
+
|
|
47
|
+
```json
|
|
48
|
+
{
|
|
49
|
+
"scope": { "tenantId": "t1", "resourceId": "docs", "corpusId": "handbook" },
|
|
50
|
+
"query": "How do approvals work?",
|
|
51
|
+
"topK": 1,
|
|
52
|
+
"result": {
|
|
53
|
+
"text": "[security-guide#0001] Recheck policy before side effects.",
|
|
54
|
+
"citations": [{ "id": "security-guide#0001", "sourceId": "security-guide" }]
|
|
55
|
+
}
|
|
56
|
+
}
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## Implementation example
|
|
60
|
+
|
|
61
|
+
```ts
|
|
62
|
+
import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
|
|
63
|
+
import { createHashEmbedder, createMemoryVectorStore } from "@arnilo/prism-memory";
|
|
64
|
+
import { chunkMarkdown, createRagContextProvider, indexChunks, retrieveContext } from "@arnilo/prism-rag";
|
|
65
|
+
|
|
66
|
+
const embedder = createHashEmbedder(); // deterministic demo/test helper, not production semantic quality
|
|
67
|
+
const store = createMemoryVectorStore();
|
|
68
|
+
const scope = { tenantId: "t1", resourceId: "docs", corpusId: "handbook" };
|
|
69
|
+
const chunks = chunkMarkdown("# Approval\n\nRecheck current policy before side effects.", {
|
|
70
|
+
sourceId: "security-guide",
|
|
71
|
+
metadata: { category: "security" },
|
|
72
|
+
});
|
|
73
|
+
await indexChunks({ chunks, embedder, store, scope });
|
|
74
|
+
|
|
75
|
+
const found = await retrieveContext("approval policy", {
|
|
76
|
+
embedder,
|
|
77
|
+
store,
|
|
78
|
+
scope,
|
|
79
|
+
topK: 4,
|
|
80
|
+
filter: { category: "security" },
|
|
81
|
+
});
|
|
82
|
+
|
|
83
|
+
const agent = createAgent({
|
|
84
|
+
model: { provider: "mock", model: "demo" },
|
|
85
|
+
provider: createMockProvider([providerTextDelta("Policy checked."), providerDone()]),
|
|
86
|
+
context: [createRagContextProvider({ embedder, store, scope })],
|
|
87
|
+
});
|
|
88
|
+
console.log(found.text, await agent.createSession().run("How do approvals work?"));
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
## Extension and configuration notes
|
|
92
|
+
|
|
93
|
+
- Supply any Phase 7-conforming embedder/vector store, including the in-memory reference or PostgreSQL/pgvector adapter.
|
|
94
|
+
- Metadata filtering is package-local after a bounded candidate query so existing vector contracts/adapters remain unchanged. Increase `queryCandidates` only when selective filters measurably need it.
|
|
95
|
+
- `createRagContextProvider()` derives its query from latest user text by default; pass a fixed string or callback for host-controlled query generation.
|
|
96
|
+
- Load source text separately with a host-owned `ResourceLoader`. The RAG package intentionally accepts text, not URLs or filesystem paths.
|
|
97
|
+
- Package is available directly or through `@arnilo/prism-all`; installation does not create an embedder, vector store, or context provider.
|
|
98
|
+
|
|
99
|
+
## Security and performance notes
|
|
100
|
+
|
|
101
|
+
- Every index/query includes exact tenant/resource/corpus scope; returned records are rechecked and malformed/foreign records fail closed.
|
|
102
|
+
- Source IDs become citation/storage IDs and must be stable non-secret identifiers. Text and user metadata can be redacted before external embedding and persistence.
|
|
103
|
+
- Retrieved documents are untrusted inert context. Prompt-injection text cannot activate tools, skills, credentials, permissions, or extensions.
|
|
104
|
+
- Remote sources must pass existing resource/media trust, SSRF, MIME, and byte policies before their decoded text reaches this package.
|
|
105
|
+
- Indexing is bounded per batch and checks abort between embed/upsert operations. A failure can leave completed batches persisted; retry is idempotent for the same stable source/chunk IDs.
|
|
106
|
+
- Filtering scans at most `queryCandidates` hits; rendering stops at top-K, UTF-8 result bytes, or estimated context-token ceiling.
|
|
107
|
+
|
|
108
|
+
## Related APIs
|
|
109
|
+
|
|
110
|
+
- [Working and semantic memory](working-and-semantic-memory.md): shared `Embedder`/`VectorStore` contracts and adapters.
|
|
111
|
+
- [Context and skills](context-and-skills.md): explicit `ContextProvider` injection and inert context semantics.
|
|
112
|
+
- [Resource loading](resource-loading.md): host-owned trusted source loading.
|
|
113
|
+
- [Multimodal content](multimodal-content.md): remote media SSRF/MIME/byte policies before text extraction.
|