@arnilo/prism 0.0.4 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/CHANGELOG.md +46 -1
  2. package/README.md +34 -10
  3. package/dist/agent-loops.d.ts +1 -0
  4. package/dist/agent-loops.js +26 -16
  5. package/dist/agents.js +147 -21
  6. package/dist/cli-init.d.ts +41 -0
  7. package/dist/cli-init.js +390 -0
  8. package/dist/cli-runner.d.ts +7 -1
  9. package/dist/cli-runner.js +13 -1
  10. package/dist/content.d.ts +19 -0
  11. package/dist/content.js +197 -69
  12. package/dist/contracts.d.ts +96 -9
  13. package/dist/contracts.js +8 -0
  14. package/dist/feedback.d.ts +48 -0
  15. package/dist/feedback.js +230 -0
  16. package/dist/ids.d.ts +2 -0
  17. package/dist/ids.js +6 -0
  18. package/dist/index.d.ts +10 -4
  19. package/dist/index.js +6 -3
  20. package/dist/providers/media.d.ts +3 -1
  21. package/dist/providers/media.js +11 -1
  22. package/dist/session-stores.js +2 -3
  23. package/dist/testing/feedback.d.ts +6 -0
  24. package/dist/testing/feedback.js +37 -0
  25. package/dist/testing/persistence-schema.d.ts +48 -10
  26. package/dist/testing/persistence-schema.js +166 -22
  27. package/dist/testing/run-ledger-conformance.js +7 -1
  28. package/dist/thinking.d.ts +42 -0
  29. package/dist/thinking.js +92 -0
  30. package/dist/tools.js +2 -3
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +75 -0
  34. package/docs/agent-events.md +14 -21
  35. package/docs/agent-loops.md +12 -9
  36. package/docs/agent-session-runtime.md +14 -16
  37. package/docs/cli-rpc.md +35 -7
  38. package/docs/coding-agent-tools.md +35 -14
  39. package/docs/coding-security.md +7 -3
  40. package/docs/compaction-llm.md +17 -7
  41. package/docs/compaction-observational-memory.md +30 -4
  42. package/docs/context-and-skills.md +1 -0
  43. package/docs/credential-storage.md +58 -9
  44. package/docs/credentials-and-redaction.md +3 -3
  45. package/docs/database-persistence.md +17 -9
  46. package/docs/evaluations.md +122 -0
  47. package/docs/extensions.md +2 -2
  48. package/docs/host-security.md +26 -5
  49. package/docs/index.md +43 -28
  50. package/docs/mcp-tools.md +74 -13
  51. package/docs/migration.md +177 -3
  52. package/docs/multimodal-content.md +14 -6
  53. package/docs/node-filesystem-config.md +1 -0
  54. package/docs/node-jsonl-session-store.md +5 -4
  55. package/docs/observability.md +14 -6
  56. package/docs/performance.md +209 -0
  57. package/docs/postgres-persistence.md +8 -6
  58. package/docs/provider-caching.md +16 -4
  59. package/docs/provider-conformance.md +40 -1
  60. package/docs/provider-packages.md +62 -3
  61. package/docs/providers/ai-sdk.md +149 -0
  62. package/docs/providers/kimi.md +124 -61
  63. package/docs/providers/neuralwatt.md +19 -13
  64. package/docs/providers/openai.md +56 -13
  65. package/docs/providers/opencode-go.md +118 -30
  66. package/docs/providers/openrouter.md +105 -35
  67. package/docs/providers/zai.md +94 -45
  68. package/docs/public-contracts.md +6 -5
  69. package/docs/rag.md +113 -0
  70. package/docs/release-and-install.md +100 -79
  71. package/docs/review-coverage-2026-07-15.md +193 -0
  72. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  73. package/docs/runs-and-usage.md +42 -5
  74. package/docs/server.md +139 -0
  75. package/docs/settings-auth-trust-security.md +5 -5
  76. package/docs/sqlite-persistence.md +6 -5
  77. package/docs/structured-output.md +1 -1
  78. package/docs/supervisors.md +71 -0
  79. package/docs/thinking-and-reasoning.md +98 -0
  80. package/docs/tool-execution-primitives.md +3 -3
  81. package/docs/tools.md +15 -0
  82. package/docs/use-case-model-selection.md +109 -0
  83. package/docs/workflow-orchestration-primitives.md +20 -3
  84. package/docs/workflows.md +114 -33
  85. package/docs/working-and-semantic-memory.md +170 -0
  86. package/package.json +13 -3
  87. package/templates/init/README.md.tmpl +28 -0
  88. package/templates/init/env.example.tmpl +1 -0
  89. package/templates/init/gitignore.tmpl +11 -0
  90. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  91. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  92. package/templates/init/package.json.tmpl +22 -0
  93. package/templates/init/providers.json +76 -0
  94. package/templates/init/src/agent.ts.tmpl +10 -0
  95. package/templates/init/src/index.ts.tmpl +12 -0
  96. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  97. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -67,6 +67,7 @@ console.log(config);
67
67
 
68
68
  - This loader is an explicit Node subpath. Importing `@arnilo/prism` does not read files or compute config layers.
69
69
  - Hosts choose which paths to read and which missing files are optional.
70
+ - Optional missing files are detected with typed Node `error.code === "ENOENT"` via `isNodeErrorCode()` — not by matching `"ENOENT"` in `error.message`.
70
71
  - The loader returns `ConfigLayer[]`; use `mergeConfigLayers()` from the root package to combine layers.
71
72
  - It does not discover packages, scan directories, watch files, import extension modules, load manifests, or start agent/session runtime behavior.
72
73
 
@@ -32,12 +32,12 @@ import { createJsonlSessionStore } from "@arnilo/prism/node/session-store-jsonl"
32
32
 
33
33
  `createJsonlSessionStore()` returns a `SessionStore`:
34
34
 
35
- - `append(entry, options?)` appends one JSON line, rejects duplicate entry ids, honors `expectedParentId` existence checks, and deduplicates exact idempotency retries within this store instance.
35
+ - `append(entry, options?)` appends one JSON line, rejects duplicate entry ids, honors `expectedParentId` existence checks, and deduplicates exact idempotency retries within this store instance. Append **fails closed** when the file already contains any corrupt or shape-invalid line (`Invalid JSONL at line N: …`) so writers cannot extend a damaged log.
36
36
  - `list(sessionId)` reads the file and returns valid entries for that session id. Corrupt or shape-invalid lines are skipped; they do not poison the whole file.
37
37
  - `get(id)` reads the file and returns the matching valid entry, if any.
38
38
  - `readJsonlSessionEntries(path)` returns `{ entries: SessionEntry[]; errors: SessionEntryParseError[] }` so hosts/tests can inspect per-line parse errors.
39
39
 
40
- Missing files read as empty stores. Invalid JSON, missing required fields, unsupported `schemaVersion`, unknown `kind`, or wrong per-kind shapes (`message`, `summary`, `model_change`, `custom`, `compaction`, `label`, `event`, `metadata`, or non-string `parentId`) are quarantined per line with line number and reason; the raw line is included in `SessionEntryParseError.raw`. Unknown entry kinds and future schema versions fail closed: the line is skipped and never returned by `list()` or `get()`.
40
+ Missing files read as empty stores (typed Node `ENOENT`). Invalid JSON, missing required fields, unsupported `schemaVersion`, unknown `kind`, or wrong per-kind shapes (`message`, `summary`, `model_change`, `custom`, `compaction`, `label`, `event`, `metadata`, or non-string `parentId`) are quarantined per line with line number and reason; the raw line is included in `SessionEntryParseError.raw`. Unknown entry kinds and future schema versions fail closed for reads: the line is skipped and never returned by `list()` or `get()`. For writes, any parse error blocks `append()` until the host repairs or replaces the file.
41
41
 
42
42
  ## Request/response example
43
43
 
@@ -65,15 +65,16 @@ Use `createMemorySessionStore()` for tests or throwaway sessions; use the JSONL
65
65
  - This adapter is an explicit Node subpath. Importing `@arnilo/prism` does not touch the filesystem.
66
66
  - Hosts choose the file path. Prism does not discover, watch, rotate, compact, or migrate files.
67
67
  - The adapter stores only `SessionEntry` data passed to `append()`.
68
- - `SessionAppendOptions` idempotency tracking is in memory for the store instance. It is a development guard, not a durable cross-process coordination mechanism.
68
+ - `SessionAppendOptions` idempotency tracking is in memory for the store instance. It resets on process restart and is a development guard, not a durable cross-process coordination mechanism.
69
69
 
70
70
  ## Security and performance notes
71
71
 
72
72
  - Reads and writes use only the caller-provided path.
73
73
  - Errors include path/reason or line number, not file contents.
74
74
  - Do not put secrets in messages, metadata, summaries, labels, or custom entries.
75
- - Reads are linear in file size. Appends are serialized per store instance.
75
+ - Reads are linear in file size. Appends also re-read and re-parse the whole file for duplicate/parent/corruption checks before writing one line, and are serialized per store instance.
76
76
  - There is no cross-process lock or durable idempotency table; two processes writing the same file can race. Add a database or external lock if multiple processes write the same file.
77
+ - Treat this adapter as development/single-process storage. Production multi-writer hosts should use an indexed database `SessionStore` adapter.
77
78
 
78
79
  ## Related APIs
79
80
 
@@ -11,6 +11,7 @@ APIs:
11
11
  - `ProviderTurnMetadata`, `ToolExecutionMetadata` on `AgentEvent`
12
12
  - `createProviderTurnMetadata()`, `readProviderHttpStatus()` in `@arnilo/prism`
13
13
  - `createOpenTelemetryInstrumentation()`, `wrapOpenTelemetryApi()`, `createInMemoryTelemetry()` in `@arnilo/prism-observability-opentelemetry`
14
+ - `handleRunFeedback()` / `handleEvaluation()` for explicit safe post-run projection
14
15
 
15
16
  ## When to use it
16
17
 
@@ -58,7 +59,7 @@ const detach = telemetry.attachSession(session);
58
59
  // or: for await (const event of session.subscribe()) telemetry.handleAgentEvent(event);
59
60
  ```
60
61
 
61
- Set `enabled: false` or omit `tracer`/`meter` for a no-op adapter.
62
+ Set `enabled: false` or omit `tracer`/`meter` for a no-op adapter. Feedback handlers accept only `runId`, rating/score, booleans, bounded counts, and fixed status — never comment, tag values, scorer/evaluation IDs, or arbitrary metadata.
62
63
 
63
64
  ## Outputs / response / events
64
65
 
@@ -78,10 +79,13 @@ OpenTelemetry mapping (when enabled):
78
79
 
79
80
  | Agent event | Span | Metric labels |
80
81
  | --- | --- | --- |
81
- | `agent_started` / `agent_finished` | `prism.agent.run` | — |
82
+ | `agent_started` / `agent_finished` / run `error` | `prism.agent.run` | `prism.run.tokens` on successful aggregate usage |
82
83
  | `provider_turn_*` | `prism.provider.turn` | `provider_id`, `outcome` on duration histogram |
83
84
  | `tool_execution_*` (terminal) | `prism.tool.execute` when started | `status` on duration histogram |
84
- | `provider_turn_finished` / `agent_finished` usage | span attributes | `prism.provider.tokens` counter (`kind`: input/output/cache_*) |
85
+ | `provider_turn_finished` usage | span attributes | `prism.provider.tokens` (`kind`: input/output/cache_*`) |
86
+ | `agent_finished` aggregate usage | — | `prism.run.tokens` (`kind`: input/output) |
87
+ | `handleRunFeedback` | active-run `prism.run.feedback` event or ended-run span | `prism.run.feedback` (`rating`, `linked_evaluation`) |
88
+ | `handleEvaluation` | active-run `prism.run.evaluation` event or ended-run span | `prism.run.evaluation` (`status`) |
85
89
 
86
90
  High-cardinality identifiers (`sessionId`, `runId`, `requestId`, `toolCallId`) are **span attributes only**, never metric labels.
87
91
 
@@ -130,8 +134,10 @@ const session = createAgent({
130
134
  }).createSession();
131
135
 
132
136
  const detach = telemetry.attachSession(session);
133
- await session.run("hello");
137
+ const result = await session.run("hello");
134
138
  detach();
139
+ telemetry.handleRunFeedback({ runId: result.runId, rating: 1, hasComment: true, tagCount: 1, scorerCount: 1, evaluationCount: 1 });
140
+ telemetry.handleEvaluation({ runId: result.runId, status: "scored", score: 0.9, hasReason: true });
135
141
 
136
142
  console.log(memory.spans.map((span) => span.name));
137
143
  ```
@@ -142,18 +148,20 @@ console.log(memory.spans.map((span) => span.name));
142
148
  - `retry_scheduled` still signals backoff; each retry attempt emits its own `provider_turn_*` pair with `metadata.attempt`.
143
149
  - NeuralWatt `neuralwatt:telemetry` provider events remain package-local; hosts may forward numeric cost/energy into custom metrics.
144
150
  - `@arnilo/prism-observability-opentelemetry` is optional and included through `@arnilo/prism-sdk` and `@arnilo/prism-all`; instrumentation remains disabled until a host configures it.
145
- - Exporter failures are isolated: instrumentation catches tracer/meter errors and invokes `onExporterError` without affecting the run.
151
+ - Exporter failures are isolated: instrumentation catches tracer/meter errors and invokes `onExporterError` without affecting the run, feedback persistence, or evaluation scoring.
152
+ - Run `error` events close every outstanding span attributable to that run. Detaching a session closes any remaining session spans; repeated terminal events are idempotent and cannot end a span twice.
146
153
  - Disabled instrumentation performs no per-delta span work (`enabled: false` or missing tracer/meter).
147
154
 
148
155
  ## Security and performance notes
149
156
 
150
157
  - Default events are metadata-only — no prompts, streamed deltas, tool arguments, or credentials.
151
158
  - Opt-in content in other event types (`message_delta`, tool `result`) is still subject to `redactAgentEvent`.
152
- - Metric labels stay low-cardinality (`provider_id`, `outcome`, `status`, token `kind`); never use `sessionId`/`runId` as labels.
159
+ - Metric labels stay low-cardinality (`provider_id`, `outcome`, `status`, token `kind`, feedback rating bucket/link presence); never use `sessionId`/`runId`, comments, tag values, scorer/evaluation IDs, or arbitrary metadata as labels. Provider-turn and run-total tokens use distinct instruments, so one counter cannot double count both scopes.
153
160
  - Target overhead when enabled is under 5% excluding exporter I/O; disabled hooks allocate no spans.
154
161
  - Provider transport limits and redaction order are documented in [Provider primitives](provider-primitives.md).
155
162
 
156
163
  ## Related APIs
164
+ - [Evaluations](evaluations.md): optional scorers can link scores to run/session/trace IDs from agent events.
157
165
 
158
166
  - [Agent events](agent-events.md): full `AgentEvent` union and subscriber semantics.
159
167
  - [Runs and usage ledger](runs-and-usage.md): durable `AgentEventRecord` persistence.
@@ -156,6 +156,215 @@ Provider SSE remained at the frozen 380 MiB/s / +1.7 MiB heap snapshot. Media an
156
156
 
157
157
  The ledger percentage overhead is intentionally not a threshold: its no-ledger baseline is below 1 ms, making the percentage unstable while absolute added latency remains about 1 ms. JSONL's append path is intentionally O(n²) across repeated appends because it rereads for corruption/conflict checks; move production or high-volume workloads to SQLite/PostgreSQL rather than weakening validation.
158
158
 
159
+ ### 0.0.5 Phase 0 baseline (2026-07-15)
160
+
161
+ Scope froze at commit `f5128a816ae204c52f3e2f089de71c99bd5de6d4`. Measurement host: Node v24.18.0, npm 11.16.0, Linux 7.1.3 x86_64, AMD Ryzen 9 PRO 7940HS (16 logical CPUs). Supported package runtime remains Node >=20. These are dated local comparison points, not portable CI wall-clock assertions.
162
+
163
+ | Surface | Workload | Result |
164
+ | --- | --- | --- |
165
+ | Network-free tests | `npm test` | 25.750 s; 1,475 tests, 1,450 pass, 25 explicit live skips, 0 fail |
166
+ | Release readiness | `npm run sdk:ready` | 54.341 s; typecheck, tests, examples, builds, and 24 dry-run packs pass |
167
+ | Provider/agent stream | One mock run with 5,000 one-character text deltas and a concurrently drained 8,192-event subscriber | 3.78 ms median |
168
+ | Tool dispatch | Six independent 20 ms tools, concurrency 1 | 121.05 ms median |
169
+ | Tool dispatch | Same calls, concurrency 2 | 60.65 ms median (2.00x speedup) |
170
+ | Workflow runner | Existing bounded 1,000-node chain, configured concurrency 8 | 9.66 ms median |
171
+ | Package artifacts | All 24 dry-run tarballs | 542,993 packed bytes; 2,084,900 unpacked bytes aggregate |
172
+ | Root artifact | `@arnilo/prism@0.0.4` dry-run tarball | 346.0 kB packed; 1.3 MB unpacked; 196 files |
173
+ | Installed workspace | Current root `node_modules` | 72 MiB |
174
+
175
+ Synthetic stream/tool/workflow values are medians of seven measured runs after one warm-up and contain no network, database, or exporter I/O. The temporary benchmark reused public `AgentSession`, `dispatchToolCallsInOrder`, and `@arnilo/prism-workflows` APIs; it was not added to CI because this phase records a baseline rather than creating hardware-sensitive tests.
176
+
177
+ Repository size at the same commit, counted from `src/` and `packages/` while excluding `dist/`:
178
+
179
+ | Area | Files | Lines |
180
+ | --- | ---: | ---: |
181
+ | Production TypeScript | 189 | 26,828 |
182
+ | Test TypeScript | 144 | 23,535 |
183
+ | Documentation Markdown | 70 | 12,662 |
184
+ | Numbered plans | 58 | 24,270 |
185
+ | TypeScript examples | 39 | 3,134 |
186
+
187
+ Prism has no project generator before Phase 5, so a generated-Prism-project install/build size is **not applicable** at this baseline. The closest current install figure is the 72 MiB development workspace; it is not a scaffold target. The comparison Mastra default scaffold measured during the review used 439 MB `node_modules`, 300 MB build output, and 427 installed packages. Phase 5 must establish a real generated Prism project baseline and keep unselected storage, telemetry, eval, memory, server, and workflow dependencies absent.
188
+
189
+ See [Review coverage — 2026-07-15](review-coverage-2026-07-15.md) for scope, primitive, package, and threat-boundary ownership.
190
+
191
+ ### 0.0.5 Phase 2 verification (2026-07-15)
192
+
193
+ Same Phase 0 host and seven-run warm benchmark. Runtime correctness changes stayed inside frozen ceilings:
194
+
195
+ | Surface | Result |
196
+ | --- | --- |
197
+ | Network-free tests | 27.992 s; 1,485 tests, 1,460 pass, 25 explicit live skips, 0 fail |
198
+ | `npm run sdk:ready` | 55.598 s; typecheck, examples, tests, builds, and all 24 dry-run packs pass |
199
+ | Provider/agent stream, 5,000 deltas | 3.54 ms median (Phase 0: 3.78 ms) |
200
+ | Six 20 ms tools, concurrency 1 / 2 | 121.22 ms / 60.63 ms (2.00x speedup retained) |
201
+ | Workflow 1,000-node chain | 10.31 ms median (well below 1 s ceiling) |
202
+ | Root dry-run tarball | 361.2 kB packed, 1.3 MB unpacked, 197 files |
203
+
204
+ Usage aggregation performs one constant-size accumulator update per terminal provider turn. Telemetry retains only active span metadata and removes every terminal/detached entry. Complete media resolution is sequential, rejects item count and inline estimates before I/O, and retains at most the request budget plus one per-item-bounded candidate before failing an aggregate overflow. Sandbox output still streams into the existing bounded `OutputAccumulator`; no adapter-side response buffer was added.
205
+
206
+ ### 0.0.5 Phase 4 verification (2026-07-15)
207
+
208
+ Optional `@arnilo/prism-evals` adds package-local scoring without changing core run latency. Validation stayed within the frozen release gate:
209
+
210
+ | Surface | Result |
211
+ | --- | --- |
212
+ | Network-free tests | 1,503 tests, 1,478 pass, 25 explicit live skips, 0 fail |
213
+ | `npm run sdk:ready` | typecheck, examples, tests, builds, and all 25 dry-run packs pass |
214
+ | Evals dry-run tarball | 35.4 kB unpacked package payload |
215
+ | Profile bundles | unchanged; evals remains opt-in until size/use review |
216
+
217
+ Experiment concurrency is capped at 32 workers and defaults to 1. Scorers operate on `AgentRunResult` references plus dataset item metadata rather than duplicating event ledgers.
218
+
219
+ ### 0.0.5 Phase 5 verification (2026-07-15)
220
+
221
+ `prism init` lands as a stdlib-only CLI subcommand with checked-in templates under `templates/init/`.
222
+
223
+ | Surface | Result |
224
+ | --- | --- |
225
+ | Default generated sources | 8 files / ~3.3 KB |
226
+ | Default clean consumer install (`@arnilo/prism` + TypeScript tooling) | ~27.5 MB `node_modules` |
227
+ | Mastra comparator | 439 MB install / 300 MB build / 427 packages |
228
+ | Default dependencies | `@arnilo/prism` only; no storage, telemetry, eval, memory, server, or workflow packages unless `--with-*` / provider flags select them |
229
+ | Offline proof | packed core tarball → `npm install` → `npm run typecheck` → `npm test` (mock provider) |
230
+
231
+ ### 0.0.5 Phase 6 verification (2026-07-15)
232
+
233
+ Optional `@arnilo/prism-provider-ai-sdk` adapts AI SDK `LanguageModelV4` streams to Prism without adding an AI SDK dependency to core.
234
+
235
+ | Surface | Result |
236
+ | --- | --- |
237
+ | Supported specification | `@ai-sdk/provider@^4` (`LanguageModelV4`) |
238
+ | Adapter behavior | incremental stream translation; unsupported content fails before `doStream`; abort owned by Prism `request.signal` |
239
+ | Network-free tests | 1,522 tests, 1,497 pass, 25 explicit live skips, 0 fail |
240
+ | `npm run sdk:ready` | typecheck, examples, tests, builds, and all 26 dry-run packs pass |
241
+ | AI SDK adapter dry-run tarball | 6.5 kB packed / 22.5 kB unpacked / 16 files |
242
+ | Profile bundles | unchanged; AI SDK adapter remains opt-in until size/use review |
243
+ | Publishable graph | 26 packages |
244
+
245
+ ### 0.0.5 Phase 7 verification (2026-07-15)
246
+
247
+ Optional `@arnilo/prism-memory` adds working memory and semantic recall without changing core session stores.
248
+
249
+ | Surface | Result |
250
+ | --- | --- |
251
+ | Contracts | package-owned `Embedder`, `VectorStore`, `WorkingMemoryStore`, `createMemory` |
252
+ | Adapters | in-memory reference + PostgreSQL/pgvector production path |
253
+ | Injection | existing `ContextProvider` seam; opt-in working-memory processor |
254
+ | Profile bundles | unchanged; memory remains opt-in until size/use review |
255
+ | Publishable graph | 27 packages |
256
+ | Network-free tests | 1,538 tests, 1,513 pass, 25 explicit live skips, 0 fail |
257
+ | `npm run sdk:ready` | pass |
258
+ | Memory dry-run tarball | 17.9 kB packed / 76.6 kB unpacked / 32 files |
259
+
260
+ ### 0.0.5 Phase 8 verification (2026-07-15)
261
+
262
+ Durable human suspension extends existing workflow checkpoint JSON/CAS; no worker polling loop, package, dependency, or database migration was added.
263
+
264
+ | Surface | Result |
265
+ | --- | --- |
266
+ | Focused workflow suite | 43 tests pass, 0 fail |
267
+ | Network-free tests | 1,547 tests, 1,522 pass, 25 explicit live skips, 0 fail |
268
+ | `npm run sdk:ready` | typecheck, examples, tests, builds, and all 27 dry-run packs pass |
269
+ | Workflow dry-run tarball | 25.7 kB packed / 121.6 kB unpacked / 34 files |
270
+ | Coordinator behavior | `suspended` absent from queued/running poll; zero worker/lease retained |
271
+ | Storage | existing bounded checkpoint JSON/category; no SQLite/PostgreSQL migration |
272
+
273
+ ### 0.0.5 Phase 9 verification (2026-07-16)
274
+
275
+ Optional `@arnilo/prism-rag` reuses Phase 7 vector contracts and adds no core path, parser dependency, network loader, or profile activation.
276
+
277
+ | Surface | Result |
278
+ | --- | --- |
279
+ | Focused RAG suite | 9 tests pass, 0 fail |
280
+ | Network-free tests | 1,561 tests, 1,536 pass, 25 explicit live skips, 0 fail |
281
+ | `npm run sdk:ready` | typecheck, examples, tests, builds, and all 28 dry-run packs pass |
282
+ | RAG dry-run tarball | 9.0 kB packed / 34.6 kB unpacked / 22 files |
283
+ | Index bounds | chunk/document/count/metadata caps; embed batches default 32, hard 128 |
284
+ | Retrieval bounds | top-K default 5/hard 32; candidates default 20/hard 128; result 64/512 KiB; context 2,000/8,000 estimated tokens |
285
+ | Profile bundles | unchanged; RAG and memory remain explicit opt-ins |
286
+
287
+ ### 0.0.5 Phase 10 verification (2026-07-16)
288
+
289
+ Optional `@arnilo/prism-server` and MCP server-direction APIs compose existing agent/workflow/tool/SDK primitives; no core path, framework/listener, auth provider, database, or profile activation was added.
290
+
291
+ | Surface | Result |
292
+ | --- | --- |
293
+ | Focused server suites | 6 Web handler tests + 4 MCP server tests pass; existing 12 MCP client tests remain green |
294
+ | Network-free tests | 1,576 tests, 1,551 pass, 25 explicit live skips, 0 fail |
295
+ | `npm run sdk:ready` | typecheck, examples, tests, builds, and all 29 dry-run packs pass |
296
+ | Server dry-run tarball | 8.4 kB packed / 34.4 kB unpacked / 12 files |
297
+ | MCP dry-run tarball | 11.6 kB packed / 45.0 kB unpacked / 20 files |
298
+ | Web handler bounds | request 64 KiB, result 1 MiB, event 64 KiB, stream 10 MiB/10k events, queue 128, concurrency 16, timeout 120 s by default; all have hard caps |
299
+ | MCP server bounds | call result 1 MiB, calls 16, timeout 60 s; HTTP request 1 MiB, response 2 MiB, requests 32 by default; all have hard caps |
300
+ | Profile bundles | unchanged; server remains explicit opt-in |
301
+
302
+ ### 0.0.5 Phase 11 verification (2026-07-16)
303
+
304
+ Workflow schedules, background runs, composition, state, and replay reuse the existing workflow package plus generic checkpoint/lease stores. No package, runtime dependency, SQL migration, listener, cron parser, or auto-started worker was added.
305
+
306
+ | Surface | Result |
307
+ | --- | --- |
308
+ | Focused workflow/server suites | 54 workflow tests + 8 Web handler tests pass, 0 fail |
309
+ | Network-free tests | 1,589 tests, 1,564 pass, 25 explicit live skips, 0 fail |
310
+ | `npm run sdk:ready` | typecheck, examples, tests, builds, and all 29 dry-run packs pass |
311
+ | Workflow dry-run tarball | 34.7 kB packed / 171.5 kB unpacked / 38 files |
312
+ | Server dry-run tarball | 9.9 kB packed / 45.2 kB unpacked / 12 files |
313
+ | Synthetic schedule bound | 100 in-memory creates: 1.28 ms; scan 100 / claim+enqueue 16 due fires: 6.64 ms |
314
+ | Synthetic composition/replay | depth-8 nested run: 2.73 ms; 100-node source: 35.42 ms; replay 50 nodes: 21.17 ms |
315
+ | State/replay ceilings | state 64/512 KiB; history 32/128; nested depth 8/32; replay depth 8/32 default/hard |
316
+ | Schedule ceilings | page 100/500; claims 16/256; input 256 KiB/1 MiB; 1s idle timer; 30s fire lease defaults |
317
+
318
+ Synthetic timings are one local Node v24.18.0 run over memory adapters with no network/database I/O; finite limits and behavior tests, not wall-clock numbers, are CI gates.
319
+
320
+ ### 0.0.5 Phase 12 verification (2026-07-16)
321
+
322
+ Run feedback adds no package or runtime dependency. Memory/SQLite/PostgreSQL implementations share bounded append/query/delete semantics; OTel projection accepts only fixed scalar metadata.
323
+
324
+ | Surface | Result |
325
+ | --- | --- |
326
+ | Focused feedback/eval/SQLite/OTel tests | 35 tests pass, 0 fail; PostgreSQL DDL suite passes and live feedback conformance is env-gated |
327
+ | Synthetic memory feedback | 1,000 bounded appends: 3.82 ms; 100 filtered 100-row queries over 1,000 records: 11.53 ms |
328
+ | Core dry-run tarball | 398.1 kB packed / 1.4 MB unpacked / 219 files |
329
+ | Evals dry-run tarball | 9.8 kB packed / 38.4 kB unpacked / 26 files |
330
+ | OTel dry-run tarball | 6.3 kB packed / 26.5 kB unpacked / 8 files |
331
+ | SQLite/PostgreSQL tarballs | 17.7/18.1 kB packed; 89.8/89.8 kB unpacked |
332
+ | Feedback limits | comment 4/16 KiB; tags 16/64; links 16/64; metadata 16/64 KiB; pages 100/500 default/hard |
333
+
334
+ Metrics came from one local Node v24.18.0 memory-adapter run. SQL correctness/indexing/migration behavior and hard bounds are gates; local timings are not release thresholds.
335
+
336
+ ### 0.0.5 Phase 13 verification (2026-07-16)
337
+
338
+ Supervisor/A2A stays in one optional zero-runtime-dependency package; core and profile bundles gained no import, listener, worker, protocol SDK, or network activation.
339
+
340
+ | Surface | Result |
341
+ | --- | --- |
342
+ | Focused supervisor/A2A suite | 11 tests pass, 0 fail; local delegation, policy/budget/abort/redaction, card signatures, server/client/stream bounds |
343
+ | Synthetic local delegation | 100 sequential mock child results: 11.83 ms |
344
+ | Synthetic in-process A2A | 100 card discovery + JSON-RPC mock round trips: 34.17 ms |
345
+ | Supervisor dry-run tarball | 15.3 kB packed / 69.4 kB unpacked / 22 files |
346
+ | Local hard ceilings | depth 16; active 32; message 1 MiB; steps 64; tools 256; tokens 1m; timeout 30m; event queue 4096 |
347
+ | A2A hard ceilings | request/card/event 1 MiB; response 8 MiB; stream 64 MiB/100k events; concurrency 256; timeout 30m |
348
+
349
+ Timings are one local Node v24.18.0 run over mock agents and an in-process fetch adapter. Bounds, protocol validation, signature/auth/origin checks, and offline behavior tests are release gates; timings are not thresholds.
350
+
351
+ ### 0.0.5 Phase 14 release-candidate verification (2026-07-16)
352
+
353
+ | Surface | Result |
354
+ | --- | --- |
355
+ | Default network-free test | 32.247 s, below 60 s budget |
356
+ | Full SDK readiness | 70.560 s; build/typecheck/examples/tests/30 pack dry-runs |
357
+ | Test matrix | 1,618 total; 1,593 pass; 25 explicit live skips; 0 fail |
358
+ | Node compatibility | Node 20.20.2 imports 44 built root/package export targets; Node 24.18.0 runs full matrix |
359
+ | PostgreSQL/pgvector | 29 live checks pass in fresh `pgvector/pgvector:pg16` container |
360
+ | Packed artifact set | 30 tarballs / 699 files; post-bundle snapshot ~690.6 kB packed / 2.64 MB unpacked |
361
+ | Core artifact | post-bundle snapshot ~403.7 kB packed / 1.46 MB unpacked / 221 files |
362
+ | Generated default project | under 50 KiB source and under 50 MiB installed; packed-core typecheck/test pass |
363
+ | Fresh packed journey | 30 packages install/import and Phase 1-13 optional composition pass in ~8.0 s |
364
+ | Registry/publish preview | 30/30 versions available; 30/30 dependency-ordered provenance dry-runs pass |
365
+
366
+ No performance ceiling was raised. Core grew from Phase 0's 346.0 kB packed baseline to ~403.7 kB after documented APIs/templates, while the full package set remains ~690.6 kB packed. Follow-up review includes all six Phase 4-13 capability packages through `prism-all` and AI SDK interoperability through `prism-providers`; focused base/code/SDK profiles remain unchanged and no capability auto-activates. Manifest tarballs remain tiny: providers 1.4 kB and all 1.6 kB packed.
367
+
159
368
  ## Related APIs
160
369
 
161
370
  - [Agent events](agent-events.md): `SubscribeOptions` and `event_subscriber_overflow` event details.
@@ -24,7 +24,7 @@ Use this package when you need server-backed persistence with pooled connections
24
24
  - managed cloud databases (RDS, Cloud SQL, Neon, Supabase, etc.)
25
25
  - CI integration tests against a real PostgreSQL service
26
26
 
27
- Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests.
27
+ Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests. This adapter stores sessions/runs, not semantic vectors; use the separate [`@arnilo/prism-memory` pgvector path](working-and-semantic-memory.md), which rejects non-finite vectors before SQL, when vector recall is needed.
28
28
 
29
29
  ## Inputs / request
30
30
 
@@ -39,6 +39,7 @@ import { createPostgresPersistence } from "@arnilo/prism-session-store-postgres"
39
39
  | `connectionString` | `string` | Create an adapter-owned bounded pool when `pool` is omitted. |
40
40
  | `schema` | `string` | PostgreSQL schema for Prism tables. Defaults to `"prism"`. Validated and double-quoted. |
41
41
  | `poolMax` | `number` | Maximum pool size for adapter-owned pools. Defaults to `10`. |
42
+ | `feedbackRedactor` | `SecretRedactor` | Optional redaction for feedback comment/tags/metadata before insert. |
42
43
  | `poolConfig` | `PoolConfig` | Additional `pg` options (TLS, idle timeout, application name, etc.). |
43
44
 
44
45
  Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backup/retention enforcement.
@@ -54,11 +55,11 @@ Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backu
54
55
  | `SessionStore.readBranchPath` | Recursive ancestor query from `leafId` (or latest leaf) in root→leaf order. |
55
56
  | `RunLedger.append*` | Inserts run/event/tool/usage rows; events receive monotonic per-run `sequence` values. |
56
57
  | `ProductionPersistenceStore.query*` | Parameterized cursor pagination on indexed columns with tenant/account/user filters. |
57
- | `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, and bounded pagination. |
58
+ | `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, bounded pagination, and workflow suspended/denied/schedule/state/replay values without a schema migration. |
58
59
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
59
60
  | `close()` | Ends the pool when the adapter created it from `connectionString`. |
60
61
 
61
- Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks.
62
+ Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks. While holding that lock, startup verifies ordered contract name/version/SHA-256 rows and full schema-v3 `information_schema`/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
62
63
 
63
64
  ## Request/response example
64
65
 
@@ -115,7 +116,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
115
116
  - The package is optional and workspace-local; `@arnilo/prism` core has no PostgreSQL dependency.
116
117
  - Schema names must match `^[a-zA-Z_][a-zA-Z0-9_]*$`; the adapter quotes them and never interpolates user values into identifier positions.
117
118
  - `SessionAppendOptions` idempotency rows are durable in `prism_session_append_idempotency` and survive reopen.
118
- - Schema version **1** (`001_init`) matches `@arnilo/prism/testing/persistence-schema` — SQLite adapters share the same model with dialect-local DDL.
119
+ - Schema version **3** applies `001_init`, additive `002_usage_scope`, and `003_run_feedback`. Migration 003 adds immutable `prism_run_feedback` rows with run FK/cascade deletion and owner/run/trace cursor indexes. `persistence.feedback` validates exact run ownership, bounds/redacts through optional `feedbackRedactor`, queries bounded pages, and deletes only exact-owned IDs. SQLite shares the same model with dialect-local DDL.
119
120
  - Pass an existing `pg` `Pool` when your host already manages pooling, TLS, and credential rotation.
120
121
 
121
122
  ## Security and performance notes
@@ -126,7 +127,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
126
127
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
127
128
  - **Bounded pool.** Adapter-owned pools default to `max: 10`. Hosts with heavy concurrency should supply their own pool sizing.
128
129
  - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid sequential scans.
129
- - **Migration locking.** `pg_advisory_xact_lock` prevents concurrent `001_init` races when multiple processes open the adapter at once.
130
+ - **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once. Startup catalog reads are bounded metadata queries, not application-row scans.
130
131
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns participate in query filters; hosts must still scope writes correctly.
131
132
  - **Benchmark target.** Indexed append + paginated branch read on a warm pool should stay under **50 ms p95** for local/CI-sized datasets (≤100k entries per session); measure with your pool size and hardware before production sizing.
132
133
 
@@ -137,5 +138,6 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
137
138
  - [Session store conformance](session-store-conformance.md): `assertSessionStoreConforms` / `runSessionStoreConformance`.
138
139
  - [Run ledger conformance](run-ledger-conformance.md): `assertRunLedgerConforms` / `runRunLedgerConformance`.
139
140
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): package matrix and threat model.
140
- - [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` for durable multi-process workflow execution.
141
+ - [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` and `createWorkflowSchedules()` for durable background execution and schedules.
142
+ - [Working and semantic memory](working-and-semantic-memory.md): optional `@arnilo/prism-memory` PostgreSQL/pgvector working + semantic stores (separate from session/run persistence).
141
143
  - [Migration guide](migration.md): moving from JSONL/in-memory to database-backed persistence.
@@ -21,6 +21,7 @@ Use this page when a host or provider package needs to:
21
21
  - Carry a stable cache key across turns without putting provider-specific fields in core.
22
22
  - Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
23
23
  - Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
24
+ - Understand when **caller-gated model discovery** may fill `ModelConfig.cache` / `ModelConfig.cost` from a live `/models` response (see [Discovery and live cache/cost metadata](#discovery-and-live-cache-cost-metadata)).
24
25
 
25
26
  Do not use cache keys for credentials, bearer tokens, API keys, OAuth tokens, user secrets, or raw private prompts.
26
27
 
@@ -145,21 +146,23 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
145
146
  | Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
146
147
  | --- | --- | --- | --- | --- |
147
148
  | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
148
- | `@arnilo/prism-provider-openrouter` | `cache_control` | Applies `cache_control` markers only to caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. | Breakpoint-stable prefixes can be reused by upstream providers. | Best-effort only; no marker is added to every block. |
149
+ | `@arnilo/prism-provider-openrouter` | `cache_control` | Top-level automatic `cache_control` when enabled without breakpoints; otherwise markers only on caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable / automatic prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
149
150
  | `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
150
151
  | `@arnilo/prism-provider-zai` | `implicit` | No explicit cache payload; GLM context caching is automatic. | Resend unchanged prior history for implicit context-cache reuse. | Best-effort only; cache options do not force hits. |
151
152
  | `@arnilo/prism-provider-kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
152
153
  | `@arnilo/prism-provider-neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
154
+ | `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
153
155
 
154
156
  Detailed first-party provider notes:
155
157
 
156
- - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
158
+ - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. GPT-5.6+ official docs use `prompt_cache_options` / breakpoints instead of retention — `listOpenAIModels` sets `longRetention: false` for those ids. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
157
159
  - OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
158
- - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars; applies Anthropic-style `cache_control` markers only to caller-selected `cache.breakpoints` (last content block of each selected message), not every block; `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
159
- - OpenCode Go (`@arnilo/prism-provider-opencode-go`): `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route sends none. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`.
160
+ - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing; with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
161
+ - OpenCode Go (`@arnilo/prism-provider-opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
160
162
  - Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
161
163
  - NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
162
164
  - Kimi (`@arnilo/prism-provider-kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
165
+ - AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
163
166
 
164
167
  ### NeuralWatt cache-aware limiter
165
168
 
@@ -185,6 +188,15 @@ sessions differently from one-shot chat:
185
188
  See [NeuralWatt provider](providers/neuralwatt.md) for the package-level cache,
186
189
  usage, and retry details.
187
190
 
191
+ ## Discovery and live cache/cost metadata
192
+
193
+ Caller-gated `list*Models()` helpers (see [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery)) may map official list-models fields onto `ModelConfig`:
194
+
195
+ - `cache` — when the provider documents cache kind / long-retention / breakpoint support in model metadata (otherwise keep the package's known default, e.g. NeuralWatt/Z.AI `implicit`, OpenAI `openai_key`).
196
+ - `cost` — when the provider documents per-token or per-million rates, including cache-read rates such as NeuralWatt `cached_input_per_million`.
197
+
198
+ Static featured catalogs remain offline bootstrap and must **not** invent pricing or cache capabilities the official docs do not state. Discovery is never invoked by `create*ProviderPackage()`; hosts that want live `cost`/`cache` pass the returned models into package `models:` (or register them themselves).
199
+
188
200
  ## Security and performance notes
189
201
 
190
202
  - Cache hints are best-effort and do not guarantee cache hits.
@@ -133,6 +133,44 @@ await assertProviderStreamConforms({
133
133
  });
134
134
  ```
135
135
 
136
+ ## Model discovery checklist
137
+
138
+ Every first-party package that ships (or plans) a `list*Models()` helper must keep setup network-free. Add these assertions in the package suite (pattern from NeuralWatt):
139
+
140
+ 1. **`*_provider_setup_does_not_call_model_discovery`** — inject a counting `fetch` into `create*ProviderPackage({ fetch })`, run `setup`, assert `calls === 0`.
141
+ 2. **`list_*_models_maps_fixture_…`** — fixture response maps to `ModelConfig` (`id` → `model`, documented capabilities/limits/cost/cache); no credentials in returned objects.
142
+ 3. **`list_*_models_forwards_auth_abort_baseurl`** (as applicable) — Authorization owned by helper when key present; auth omitted when optional and unset; `signal` / `baseUrl` forwarded.
143
+ 4. **`list_*_models_redacts_token_in_errors`** — non-OK bodies use `readBoundedResponseText` + `redactSecrets`; secret canaries absent from thrown messages.
144
+ 5. **Malformed payload rejects** — missing `data` array (or provider-equivalent) throws a clear discovery error.
145
+
146
+ OpenRouter stays app-registration-first: an optional list helper must still not run during setup. AI SDK has no discovery export. Packages without a public list API document curated official-doc refresh instead of inventing a fake helper.
147
+
148
+ Canonical contract: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery).
149
+
150
+ ## Thinking / reasoning checklist
151
+
152
+ Every first-party package that exposes thinking or reasoning controls should cover:
153
+
154
+ 1. **Model default** — `ModelConfig.compat` (or documented capability) sets the official wire field when no per-turn override is present.
155
+ 2. **Per-turn override wins** — `ProviderRequestOptions.compat` via `mergeProviderRequestOptions` / `applyThinkingLevel` overrides the model default.
156
+ 3. **Shared family mapping** — effort levels from `ThinkingLevel` land in the package's recommended family (`openai_reasoning` / `reasoning_effort` / `thinking_type` / `noop`) per [Thinking and reasoning](thinking-and-reasoning.md).
157
+ 4. **Non-reasoning / noop** — applying a level with `noop` (or omitting compat) must not invent unsupported body fields.
158
+ 5. **No inert `extra.thinkingLevel`** — package code must not rely on `options.extra.thinkingLevel` for wire mapping.
159
+
160
+ Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md).
161
+
162
+ ## AI SDK adapter checklist
163
+
164
+ `@arnilo/prism-provider-ai-sdk` is a host-owned `LanguageModelV4` bridge. It does not participate in the discovery or thinking/reasoning checklists above. Cover instead:
165
+
166
+ 1. **No catalog / no setup fetch** — package exports no `list*Models()`; `createAiSdkProvider` wraps a host model only.
167
+ 2. **Specification gate** — rejects non-v4 models (`specificationVersion !== "v4"` or missing `doStream`).
168
+ 3. **Cache usage mapping** — `finish.usage.inputTokens.cacheRead`/`cacheWrite` map to `Usage.cacheReadTokens`/`cacheWriteTokens`; adapter does not emit cache request fields.
169
+ 4. **Reasoning stream mapping** — `reasoning-delta` → thinking deltas; assistant `thinking` blocks replay as AI SDK `reasoning` prompt parts.
170
+ 5. **Host-owned controls** — `options.compat` / `options.extra` forward as `providerOptions.prism`; reasoning effort stays on the host model.
171
+
172
+ Canonical contract: [AI SDK provider adapter](providers/ai-sdk.md).
173
+
136
174
  ## Extension and configuration notes
137
175
 
138
176
  The helpers are a testing subpath only. Provider packages can use them with their own mocked fetch/transport or `createMockProvider()`. Live provider tests should stay opt-in and env-gated outside Prism's default test suite.
@@ -148,6 +186,7 @@ The helpers are a testing subpath only. Provider packages can use them with thei
148
186
  ## Related APIs
149
187
 
150
188
  - [Provider layer](provider-layer.md): `AIProvider`, provider events, and mock provider.
151
- - [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters.
189
+ - [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters; includes the caller-gated discovery contract and setup zero-fetch rule.
190
+ - [AI SDK provider adapter](providers/ai-sdk.md): optional `LanguageModelV4` bridge tested with a fake AI SDK model.
152
191
  - [OpenAI-compatible provider](providers/openai-compatible.md): optional provider adapter tested with mocked streams.
153
192
  - [Public contracts](public-contracts.md): provider request/event/usage contracts.
@@ -63,9 +63,11 @@ First-party providers map generic `ModelConfig.parameters.maxTokens` to real out
63
63
 
64
64
  Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md). Each package starts with a side-effect-free `create*ProviderPackage()` export, README, TypeScript build, network-free default tests, and real opt-in live smoke tests.
65
65
 
66
+ Phase 6 also adds optional [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md), which adapts a host-owned AI SDK `LanguageModelV4` to Prism's `AIProvider`. It joins `@arnilo/prism-providers` as the seventh adapter while remaining independent from the six HTTP implementations.
67
+
66
68
  Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
67
69
 
68
- These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
70
+ These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only, with optional `models`/`codexModels` overrides and an opt-in `listOpenAIModels()` helper for official `GET /models` discovery. `@arnilo/prism-provider-opencode-go` now registers docs-verified OpenCode Go open coding models with dual OpenAI/Anthropic routes (`compat.route`), official default base `https://opencode.ai/zen/go/v1`, `reasoning_content`/thinking preserve, and an opt-in `listOpenCodeGoModels()` helper for official `GET /zen/go/v1/models`. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/`reasoning`/cache passthrough, assistant `reasoning` replay, optional top-level automatic `cache_control`, and an opt-in `listOpenRouterModels()` helper for official `GET /api/v1/models` (setup still never fetches). `@arnilo/prism-provider-zai` now registers featured GLM-5.x/4.x metadata with official `thinking`/`reasoning_effort`/`tool_stream`/`clear_thinking` mapping, Preserved Thinking `reasoning_content` replay, implicit context caching, and an opt-in `listZaiModels()` helper for OpenAI-compatible `GET /models`. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default, optional callable Moonshot Open Platform Chat Completions when `includeMoonshotModels` is requested, official Coding/Open Platform featured ids, thinking/`reasoning_effort` compat mapping, and an opt-in `listKimiModels()` helper for Moonshot `GET /v1/models`. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
69
71
 
70
72
  ### First-party cache behavior
71
73
 
@@ -73,14 +75,71 @@ Every first-party provider package hardens prompt-cache behavior so it cannot em
73
75
 
74
76
  - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
75
77
  - **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
76
- - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; `cache_control` markers applied only to caller-selected `cache.breakpoints` (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
77
- - **OpenCode Go**: `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; the Anthropic route applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`), the OpenAI route sends none. Per-route usage mapping.
78
+ - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
79
+ - **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
78
80
  - **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
79
81
  - **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
80
82
  - **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
81
83
 
82
84
  See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
83
85
 
86
+ ## Caller-gated model discovery
87
+
88
+ First-party packages keep `create*ProviderPackage()` network-free. Latest models come from **caller-gated** `list*Models()` helpers that hosts invoke explicitly and then pass back via `models:` (or register themselves). Plan 015's "no setup catalog fetch" rule still holds; Plan 067 adds on-demand discovery without hidden latency.
89
+
90
+ ### Contract
91
+
92
+ ```ts
93
+ export async function listExampleModels(options: {
94
+ apiKey?: CredentialValueSource;
95
+ fetch?: typeof fetch;
96
+ baseUrl?: string;
97
+ signal?: AbortSignal;
98
+ headers?: Readonly<Record<string, string>>;
99
+ }): Promise<ModelConfig[]> {
100
+ // GET {baseUrl}/models — never called from create*ProviderPackage()
101
+ }
102
+ ```
103
+
104
+ | Rule | Requirement |
105
+ | --- | --- |
106
+ | Setup | `create*ProviderPackage().setup` performs **zero** fetches / discovery calls |
107
+ | Shape | Package-local `list*Models(options) → Promise<ModelConfig[]>` + optional `map*Model(entry)` |
108
+ | Injectables | `fetch`, `baseUrl`, `signal`, optional `apiKey` / `headers` |
109
+ | Transport | Error bodies via `@arnilo/prism/providers/transport` `readBoundedResponseText`; credentials via `resolveCredentialValue` + `redactSecrets` |
110
+ | Return | `ModelConfig[]` only — never embed API keys, tokens, or auth headers in returned metadata |
111
+ | Static catalog | Featured aliases / offline bootstrap only; may omit live pricing until discovery fills `cost` / `cache` |
112
+ | Core | Prefer package-local helpers. Do **not** add a core model-discovery registry. Extract a shared HTTP/list helper only when ≥2 packages share identical parsing |
113
+
114
+ Template: [`listNeuralWattModels`](providers/neuralwatt.md) in `@arnilo/prism-provider-neuralwatt`.
115
+
116
+ ### Per-package policy
117
+
118
+ | Package | Discovery helper | Setup catalog | Notes |
119
+ | --- | --- | --- | --- |
120
+ | OpenAI | **`listOpenAIModels` (exists)** | Featured Responses/Codex aliases; factory accepts `models?` / `codexModels?` | Official `GET /v1/models`; Codex not listed by api.openai.com |
121
+ | Kimi | **`listKimiModels`** (Moonshot `GET /v1/models`) | Featured Coding ids + optional callable Moonshot | Official Moonshot/Kimi list-models; Coding curated |
122
+ | Z.AI | **`listZaiModels`** (OpenAI-compatible `GET /models`) + curated featured refresh | Featured GLM-5.2…4.5 aliases | No first-class docs.z.ai list page; discovery is best-effort; featured set from Chat Completions enum / overview |
123
+ | OpenRouter | **`listOpenRouterModels`** (official `GET /api/v1/models`) | **App-controlled** `models:` only — no bundled mega-catalog | Helper feeds host registration; setup still does not fetch |
124
+ | OpenCode Go | **`listOpenCodeGoModels`** (official `GET /zen/go/v1/models`) | Featured dual-route official Go aliases | Official Go docs endpoint table + sparse list API |
125
+ | NeuralWatt | **`listNeuralWattModels` (exists)** | Featured aliases without guessed pricing | Auth optional for public models |
126
+ | AI SDK | None | Host-owned `LanguageModelV4` | No Prism-side catalog by design |
127
+
128
+ Host pattern:
129
+
130
+ ```ts
131
+ const models = await listNeuralWattModels({ apiKey, fetch });
132
+ await kernel.load([createNeuralWattProviderPackage({ apiKey, models })]);
133
+ ```
134
+
135
+ Discovery may populate `ModelConfig.cache` and `ModelConfig.cost` from live metadata when the provider documents those fields; see [Provider caching](provider-caching.md#discovery-and-live-cache-cost-metadata). Package authors: include the [setup zero-fetch checklist](provider-conformance.md#model-discovery-checklist) in every first-party suite that ships or plans a `list*Models` helper.
136
+
137
+ ## Per-turn thinking / reasoning
138
+
139
+ Hosts set effort with portable helpers from `@arnilo/prism` (`applyThinkingLevel`, `thinkingCompatFor`) that write official fields into `ProviderRequestOptions.compat`. Model defaults stay on `ModelConfig.compat`; per-turn patches win via `mergeProviderRequestOptions`. Providers keep reading `options.compat` / `model.compat` — do not invent a parallel options tree or put effort only in `extra`.
140
+
141
+ Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md). Package-local knobs (NeuralWatt budgets, Z.AI `tool_stream`, Kimi keep/all) remain on `compat` beside the shared families.
142
+
84
143
  ## Third-party provider packaging
85
144
 
86
145
  A third party ships their own providers the same way Prism ships first-party