@arnilo/prism 0.0.5 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/CHANGELOG.md +28 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +26 -16
  4. package/dist/agents.js +2 -3
  5. package/dist/contracts.d.ts +2 -0
  6. package/dist/ids.d.ts +2 -0
  7. package/dist/ids.js +6 -0
  8. package/dist/index.d.ts +5 -1
  9. package/dist/index.js +3 -1
  10. package/dist/session-stores.js +2 -3
  11. package/dist/testing/persistence-schema.d.ts +45 -7
  12. package/dist/testing/persistence-schema.js +138 -24
  13. package/dist/thinking.d.ts +42 -0
  14. package/dist/thinking.js +92 -0
  15. package/dist/tools.js +2 -3
  16. package/dist/use-case-model.d.ts +63 -0
  17. package/dist/use-case-model.js +52 -0
  18. package/docs/a2a.md +4 -2
  19. package/docs/agent-events.md +10 -15
  20. package/docs/agent-loops.md +11 -8
  21. package/docs/coding-agent-tools.md +33 -12
  22. package/docs/coding-security.md +2 -2
  23. package/docs/compaction-llm.md +17 -7
  24. package/docs/compaction-observational-memory.md +28 -4
  25. package/docs/credential-storage.md +58 -9
  26. package/docs/credentials-and-redaction.md +1 -1
  27. package/docs/database-persistence.md +8 -3
  28. package/docs/host-security.md +10 -6
  29. package/docs/index.md +23 -20
  30. package/docs/mcp-tools.md +26 -10
  31. package/docs/migration.md +146 -2
  32. package/docs/node-filesystem-config.md +1 -0
  33. package/docs/node-jsonl-session-store.md +5 -4
  34. package/docs/postgres-persistence.md +3 -3
  35. package/docs/provider-caching.md +16 -4
  36. package/docs/provider-conformance.md +39 -1
  37. package/docs/provider-packages.md +60 -3
  38. package/docs/providers/ai-sdk.md +36 -0
  39. package/docs/providers/kimi.md +124 -61
  40. package/docs/providers/neuralwatt.md +19 -13
  41. package/docs/providers/openai.md +56 -13
  42. package/docs/providers/opencode-go.md +118 -30
  43. package/docs/providers/openrouter.md +105 -35
  44. package/docs/providers/zai.md +94 -45
  45. package/docs/release-and-install.md +47 -49
  46. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  47. package/docs/runs-and-usage.md +1 -1
  48. package/docs/sqlite-persistence.md +2 -2
  49. package/docs/structured-output.md +1 -1
  50. package/docs/thinking-and-reasoning.md +98 -0
  51. package/docs/tool-execution-primitives.md +3 -3
  52. package/docs/tools.md +15 -0
  53. package/docs/use-case-model-selection.md +109 -0
  54. package/docs/workflow-orchestration-primitives.md +1 -0
  55. package/docs/workflows.md +17 -10
  56. package/docs/working-and-semantic-memory.md +1 -0
  57. package/package.json +2 -2
package/docs/migration.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- Prism 0.0.5 preserves documented 0.0.3 agent construction except for two intentional Phase 3 public-API cleanups:
5
+ Prism 0.0.6 preserves documented 0.0.3 agent construction except for two intentional Phase 3 public-API cleanups:
6
6
 
7
7
  1. **`session.run()` / `session.prompt()` return `AgentRunResult`** and `session.stream()` starts one owned run after subscribing. Callers that ignored the previous `Promise<void>` keep working; failed/aborted runs reject with `AgentRunError` (`.result` attached).
8
8
  2. **`AgentConfig.extensions` / `settings` / `credentials` are removed.** Wire extensions through `createExtensionKernel()`, read settings in the host, and pass credential resolvers to the provider edge.
@@ -23,6 +23,150 @@ Phase 10 adds optional `@arnilo/prism-server` and extends `@arnilo/prism-mcp` wi
23
23
 
24
24
  Phase 11 compatibly extends `@arnilo/prism-workflows` with `workflowNode`, shared state fields/context updates, replay lineage, explicit background enqueue, and ownership-scoped schedules. Existing workflow definitions and direct runs remain valid; `WorkflowRunResult` now always includes `state`. State schemas require a host `validateState` callback. Schedules reuse generic checkpoint/lease stores, so SQLite/PostgreSQL need no migration and no scheduler starts automatically.
25
25
 
26
+ Prism 0.0.6 intentionally hardens workflow identity and resource limits:
27
+
28
+ - Every `defineWorkflow()` input requires a non-empty host-authored `revision`. Revision and nested workflow revisions enter `definitionHash`; bump revision whenever function/tool behavior changes. Existing checkpoints with a different hash fail resume/replay/cancel before mutation.
29
+ - `cancelWorkflowRun()` now requires `workflow` as well as IDs/checkpoints. Cancellation compares exact tenant/account/user ownership; tenant-only or missing ownership no longer matches a run stored with account/user identity.
30
+ - Active runs are keyed by workflow ID, run ID, and exact ownership. Duplicate exact registration throws `ERR_PRISM_WORKFLOW_ALREADY_ACTIVE`; same IDs under distinct exact owners remain isolated.
31
+ - All `WorkflowLimits`, runtime `concurrency`, node retries/timeouts, and checkpoint byte options reject non-finite, unsafe, zero/negative, or above-hard-cap values instead of accepting/clamping them.
32
+
33
+ ```ts
34
+ // Before
35
+ const workflow = defineWorkflow({ id: "publish", nodes });
36
+ await cancelWorkflowRun({ workflowId: workflow.id, runId, checkpoints, ownership });
37
+
38
+ // 0.0.6
39
+ const workflow = defineWorkflow({ id: "publish", revision: "2026-07-19.1", nodes });
40
+ await cancelWorkflowRun({ workflowId: workflow.id, runId, workflow, checkpoints, ownership });
41
+ ```
42
+
43
+ Checkpoint schema remains version 1; no table migration is required. Pre-0.0.6 checkpoint hashes do not include revision and therefore fail against 0.0.6 definitions. Complete them before upgrade, or perform an explicit host-owned checkpoint rewrite only after verifying the exact old/new definition; do not guess a revision to bypass evidence checks.
44
+
45
+ Prism 0.0.6 also makes coding-agent I/O finite:
46
+
47
+ - `shell` now defaults to 600 seconds and 64 MiB combined output; request/config timeout cannot exceed 3,600 seconds. Timeout, abort, overflow, and spill failure kill signal-aware operations and remove unpublished spills.
48
+ - `read` streams one text page with a 64 MiB scan ceiling instead of calling full-file `readFile()`. Custom `ReadOperations` must implement `readText(path, ReadTextOptions)` and `statFile()`; text results must stay within requested caps.
49
+ - `write` rejects UTF-8 input over `maxInputBytes` before policy/filesystem mutation.
50
+ - `edit` requires custom `EditOperations.statFile()`, caps the target, aggregate old/new input, and replacement count, and passes caps/signals into operation methods.
51
+
52
+ ```ts
53
+ const tools = createCodingTools(root, {
54
+ shell: { timeout: 600, maxTotalOutputBytes: 64 * 1024 * 1024 },
55
+ read: { maxScanBytes: 64 * 1024 * 1024 },
56
+ write: { maxInputBytes: 8 * 1024 * 1024 },
57
+ edit: { maxFileBytes: 8 * 1024 * 1024, maxInputBytes: 2 * 1024 * 1024, maxEdits: 100 },
58
+ });
59
+ ```
60
+
61
+ Custom shell/sandbox adapters must honor the composed `signal` and finite `timeout`; Prism cannot kill an opaque remote operation that ignores its host contract. Successful truncated local output remains at `metadata.fullOutputPath` for the host to consume and delete.
62
+
63
+ Prism 0.0.6 also bounds JSON Schema, vectors, and generated IDs:
64
+
65
+ - `@arnilo/prism-tool-validator-json-schema` now rejects invalid instance/schema/cache limits during construction, then rejects schemas over default 256 KiB, depth 64, 10,000 properties/keywords, or 128 refs before Ajv compilation. Only `#` fragment refs remain valid; the compiled cache is a finite 256-entry LRU. Configure an explicit lower cap where tools accept third-party schemas.
66
+ - `@arnilo/prism-memory` now fails before scoring/storage for empty, non-number, NaN, or infinite embeddings and for dimension mismatches in configured PostgreSQL/pgvector stores. Fix the host embedder/data rather than filtering invalid values after a query.
67
+ - Generated core/workflow/evaluation IDs are cryptographic UUIDs. No API shape changes, but tests or parsers that assumed timestamp/base36 IDs must treat IDs as opaque strings.
68
+
69
+ Prism 0.0.6 hardens `@arnilo/prism-credentials-node`:
70
+
71
+ - `encryptBytes()` and `decryptBytes()` now return Promises because scrypt runs asynchronously instead of blocking the JavaScript event loop.
72
+ - Encrypted files default to 4 MiB and decrypted vaults to 3 MiB (hard 16 MiB/12 MiB). Strict envelope parsing rejects unknown properties, non-canonical base64, invalid salt/IV/tag lengths, unsupported algorithms/version, and excessive KDF work before scrypt.
73
+ - scrypt requires power-of-two `N` from 16,384–262,144, `r≤32`, `p≤16`, exact 32-byte keys, `N*r*p≤2,097,152`, and `128*N*r≤256 MiB`.
74
+ - Existing Unix vault files with group/other permissions now fail on open/rotate before content read. Fix deliberately with `chmod 600 <vault>` after confirming ownership; Prism does not silently chmod an existing file.
75
+ - Keychain calls use abort-aware native async operations, a 5-second default/60-second hard timeout, and a 3 MiB default/12 MiB hard payload bound. Unknown native messages are no longer rethrown.
76
+
77
+ ```ts
78
+ // Before
79
+ const envelope = encryptBytes(plaintext, passphrase);
80
+ const bytes = decryptBytes(envelope, passphrase);
81
+
82
+ // 0.0.6
83
+ const envelope = await encryptBytes(plaintext, passphrase);
84
+ const bytes = await decryptBytes(envelope, passphrase);
85
+
86
+ const store = await openEncryptedCredentialStore({
87
+ path: "./credentials.vault",
88
+ getPassphrase,
89
+ limits: { maxFileBytes: 4 * 1024 * 1024, maxVaultBytes: 3 * 1024 * 1024 },
90
+ });
91
+ ```
92
+
93
+ Version-1 AES-GCM envelopes written with documented 0.0.5 defaults remain compatible when canonical and within limits. Oversized, permissive-mode, malformed, or previously out-of-policy custom KDF files require explicit host review; no automatic rewrite bypass is provided.
94
+
95
+ Prism 0.0.6 makes MCP client discovery/results and Streamable HTTP fail closed:
96
+
97
+ - Every `streamable-http` config now requires `allowedOrigins` with exact HTTPS origins. URLs with credentials/fragments, redirects, public plaintext HTTP, private/mixed DNS, and origin changes fail. Every SDK POST/GET/DELETE/reconnect pins one validated address and defaults to a 16 MiB response cap (64 MiB hard).
98
+ - Local development plaintext requires `allowLoopbackHttp: true`; both hostname and every DNS answer must remain loopback. This does not enable arbitrary private-network endpoints.
99
+ - Discovery defaults to 20 pages, 500 tools, 4 KiB cursors, 256-byte names, 16 KiB descriptions, 256 KiB schema/tool, and 4 MiB aggregate schemas. Repeated cursors and failed refreshes reject without replacing the previous tools.
100
+ - `content`, `structuredContent`, and legacy SDK `toolResult` now share `maxResultBytes` plus JSON depth/property limits. `structuredContent` remains `ToolResult.value` but is no longer duplicated under metadata.
101
+ - `listAllMcpTools(client, signal?, limits?)` accepts an optional third finite-limits object. Bridge options expose the same discovery/result fields. Invalid, non-finite, unsafe, zero/negative, or above-hard-cap values reject at setup.
102
+
103
+ ```ts
104
+ // Before: HTTP accepted without package-enforced origin/DNS policy.
105
+ transport: { type: "streamable-http", url: "http://mcp.example.test/mcp" }
106
+
107
+ // 0.0.6: exact HTTPS origin and finite discovery/result configuration.
108
+ const bridge = await connectMcpTools({
109
+ serverId: "docs",
110
+ transport: {
111
+ type: "streamable-http",
112
+ url: "https://mcp.example.test/mcp",
113
+ allowedOrigins: ["https://mcp.example.test"],
114
+ },
115
+ maxListPages: 20,
116
+ maxTools: 500,
117
+ maxToolSchemaBytes: 256 * 1024,
118
+ maxResultBytes: 2 * 1024 * 1024,
119
+ });
120
+ ```
121
+
122
+ Stdio remains an explicit host-selected executable and does not gain network policy. MCP bridge calls should still pass through core dispatch with a host `SecretRedactor`, `PermissionPolicy`, and `ToolValidator`; package limits do not establish server trust or sandbox subprocesses.
123
+
124
+ Prism 0.0.6 makes first-party persistence startup fail closed on migration/schema drift:
125
+
126
+ - `@arnilo/prism-session-store-sqlite` and `@arnilo/prism-session-store-postgres` now write deterministic SHA-256 checksums for every new `prism_migrations` row and validate exact ordered name/version/checksum history before applying DDL or exposing runtime writes.
127
+ - Open also checks full schema version 3 metadata: required tables, columns/types/nullability/defaults, primary/unique/foreign keys, and named index definitions. SQLite uses bounded PRAGMAs/catalog reads; PostgreSQL uses bounded `information_schema`/system-catalog reads while its existing per-schema advisory transaction lock is held. Neither scans application rows.
128
+ - Existing complete 0.0.5 histories with all `checksum` values `NULL` are accepted exactly once: Prism verifies full current shape, backfills every checksum inside the migration transaction, and then opens. Unknown, duplicate, out-of-order, name/version/checksum-mismatched, mixed/partial legacy rows or shape drift now reject before runtime writes.
129
+
130
+ ```ts
131
+ // No call-site API change. Open either verifies/backfills safely or fails.
132
+ const sqlite = createSqlitePersistence({ filename: "./prism.db" });
133
+ const postgres = await createPostgresPersistence({ pool, schema: "prism" });
134
+ ```
135
+
136
+ Before upgrade, back up the database and complete any in-flight migration. On a drift error, restore a known schema or apply a reviewed DDL repair that matches version 3, then reopen. Do not update `prism_migrations.checksum` manually: that bypasses evidence rather than repairing the schema.
137
+
138
+ Prism 0.0.6 makes compaction workers and A2A stream decoding finite:
139
+
140
+ - LLM compaction now defaults `maxSummaryTokens` to 16,384 (131,072 hard), `reserveTokens` to 16,384 (131,072 hard), and `maxErrorBytes` to 1 KiB (8 KiB hard). `maxOutputTokens` remains an alias. Invalid values reject when the strategy is created. Every post-policy provider request must retain finite `model.parameters.maxTokens`; streamed text and even empty/non-text event counts terminate at derived finite bounds.
141
+ - Final summaries are capped at four UTF-16 code units per configured token without splitting a surrogate pair. Tiny caps may omit the human truncation marker to honor the actual ceiling. Provider error/factory/policy text is exact-known-secret redacted and UTF-8 bounded.
142
+ - Observational-memory runtime adds flat `maxWorkerTurns`, `maxWorkerToolCallsPerTurn`, `maxWorkerToolCalls`, `maxWorkerArgumentBytes`, `maxWorkerResultBytes`, `maxWorkerMessageBytes`, and `maxWorkerErrorBytes` options. Defaults are 16 turns, 32/128 calls, 64 KiB arguments/results, 1 MiB messages, and 1 KiB errors; hard caps are 64, 256/1,024, 1 MiB, 1 MiB, 8 MiB, and 8 KiB.
143
+ - Settings `agentMaxTurns` now rejects fractions, non-finite values, zero/negative values, and values above 64 instead of flooring or falling back. Runtime `maxWorkerTurns` overrides it. Direct worker calls retain required `maxTurns` and use the shorter corresponding option names.
144
+ - Unknown/excess worker calls and oversized/deep/cyclic/non-JSON arguments/results now reject. Replayed arguments/results and runtime status/debug errors are bounded/redacted; pass all known secrets explicitly.
145
+ - A2A public limit defaults/options do not change. Client streaming now correctly preserves split UTF-8, accepts LF/CRLF/mixed separators and multiline `data:`, and rejects malformed UTF-8, unterminated frames, missing terminal state, or events after completion.
146
+
147
+ ```ts
148
+ const strategy = createLlmCompactionStrategy({
149
+ provider: summaryProvider,
150
+ model: summaryModel,
151
+ maxSummaryTokens: 4_096,
152
+ maxErrorBytes: 1_024,
153
+ });
154
+
155
+ const memory = createObservationalMemoryRuntime({
156
+ session,
157
+ appendEntry,
158
+ workerProvider,
159
+ sessionModel,
160
+ maxWorkerTurns: 8,
161
+ maxWorkerToolCalls: 64,
162
+ maxWorkerResultBytes: 64 * 1024,
163
+ });
164
+ ```
165
+
166
+ No background worker, provider call, or network connection activates at import/setup. Host-provided observational-memory tools remain trusted code: Prism can reject an oversized result after return but cannot undo tool side effects.
167
+
168
+ Prism 0.0.6 also adds opt-in bounded artifact-loop tools. Set `loop: { strategy: "generate-validate-revise", toolCalls: "bounded", validator }` with `maxToolRounds`; calls dispatch sequentially through normal permission, validation, redaction, ledger, and lifecycle paths. Tool-call turns do not consume artifact revisions or parse/validate an artifact. The shared round cap emits terminal `artifact_failed` metadata `{ reason: "tool_round_limit" }`; omitted or `"disabled"` preserves prior inert-call behavior.
169
+
26
170
  This page also covers two optional adoption paths:
27
171
 
28
172
  1. **In-memory / JSONL → database-backed persistence** — replace the single-process development `SessionStore` with `@arnilo/prism-session-store-sqlite`, `@arnilo/prism-session-store-postgres`, or a host implementation, and optionally attach its durable `RunLedger`.
@@ -36,7 +180,7 @@ Read this page when:
36
180
 
37
181
  - you are taking an app from the `createMemorySessionStore()` / `createJsonlSessionStore()` path to a multi-process, multi-tenant, or durable database backend;
38
182
  - you are hardening an agent that previously relied on "every scoped tool/skill is active" and need to name capabilities explicitly;
39
- - you are adopting 0.0.5 persistence, checkpoints/leases, workflows, structured output, multimodality, or explicit tool safety for the first time.
183
+ - you are adopting 0.0.6 persistence, checkpoints/leases, workflows, structured output, multimodality, or explicit tool safety for the first time.
40
184
 
41
185
  If you are new to Prism, start at [Session stores](session-stores.md) and [Agent/session runtime](agent-session-runtime.md) instead.
42
186
 
@@ -67,6 +67,7 @@ console.log(config);
67
67
 
68
68
  - This loader is an explicit Node subpath. Importing `@arnilo/prism` does not read files or compute config layers.
69
69
  - Hosts choose which paths to read and which missing files are optional.
70
+ - Optional missing files are detected with typed Node `error.code === "ENOENT"` via `isNodeErrorCode()` — not by matching `"ENOENT"` in `error.message`.
70
71
  - The loader returns `ConfigLayer[]`; use `mergeConfigLayers()` from the root package to combine layers.
71
72
  - It does not discover packages, scan directories, watch files, import extension modules, load manifests, or start agent/session runtime behavior.
72
73
 
@@ -32,12 +32,12 @@ import { createJsonlSessionStore } from "@arnilo/prism/node/session-store-jsonl"
32
32
 
33
33
  `createJsonlSessionStore()` returns a `SessionStore`:
34
34
 
35
- - `append(entry, options?)` appends one JSON line, rejects duplicate entry ids, honors `expectedParentId` existence checks, and deduplicates exact idempotency retries within this store instance.
35
+ - `append(entry, options?)` appends one JSON line, rejects duplicate entry ids, honors `expectedParentId` existence checks, and deduplicates exact idempotency retries within this store instance. Append **fails closed** when the file already contains any corrupt or shape-invalid line (`Invalid JSONL at line N: …`) so writers cannot extend a damaged log.
36
36
  - `list(sessionId)` reads the file and returns valid entries for that session id. Corrupt or shape-invalid lines are skipped; they do not poison the whole file.
37
37
  - `get(id)` reads the file and returns the matching valid entry, if any.
38
38
  - `readJsonlSessionEntries(path)` returns `{ entries: SessionEntry[]; errors: SessionEntryParseError[] }` so hosts/tests can inspect per-line parse errors.
39
39
 
40
- Missing files read as empty stores. Invalid JSON, missing required fields, unsupported `schemaVersion`, unknown `kind`, or wrong per-kind shapes (`message`, `summary`, `model_change`, `custom`, `compaction`, `label`, `event`, `metadata`, or non-string `parentId`) are quarantined per line with line number and reason; the raw line is included in `SessionEntryParseError.raw`. Unknown entry kinds and future schema versions fail closed: the line is skipped and never returned by `list()` or `get()`.
40
+ Missing files read as empty stores (typed Node `ENOENT`). Invalid JSON, missing required fields, unsupported `schemaVersion`, unknown `kind`, or wrong per-kind shapes (`message`, `summary`, `model_change`, `custom`, `compaction`, `label`, `event`, `metadata`, or non-string `parentId`) are quarantined per line with line number and reason; the raw line is included in `SessionEntryParseError.raw`. Unknown entry kinds and future schema versions fail closed for reads: the line is skipped and never returned by `list()` or `get()`. For writes, any parse error blocks `append()` until the host repairs or replaces the file.
41
41
 
42
42
  ## Request/response example
43
43
 
@@ -65,15 +65,16 @@ Use `createMemorySessionStore()` for tests or throwaway sessions; use the JSONL
65
65
  - This adapter is an explicit Node subpath. Importing `@arnilo/prism` does not touch the filesystem.
66
66
  - Hosts choose the file path. Prism does not discover, watch, rotate, compact, or migrate files.
67
67
  - The adapter stores only `SessionEntry` data passed to `append()`.
68
- - `SessionAppendOptions` idempotency tracking is in memory for the store instance. It is a development guard, not a durable cross-process coordination mechanism.
68
+ - `SessionAppendOptions` idempotency tracking is in memory for the store instance. It resets on process restart and is a development guard, not a durable cross-process coordination mechanism.
69
69
 
70
70
  ## Security and performance notes
71
71
 
72
72
  - Reads and writes use only the caller-provided path.
73
73
  - Errors include path/reason or line number, not file contents.
74
74
  - Do not put secrets in messages, metadata, summaries, labels, or custom entries.
75
- - Reads are linear in file size. Appends are serialized per store instance.
75
+ - Reads are linear in file size. Appends also re-read and re-parse the whole file for duplicate/parent/corruption checks before writing one line, and are serialized per store instance.
76
76
  - There is no cross-process lock or durable idempotency table; two processes writing the same file can race. Add a database or external lock if multiple processes write the same file.
77
+ - Treat this adapter as development/single-process storage. Production multi-writer hosts should use an indexed database `SessionStore` adapter.
77
78
 
78
79
  ## Related APIs
79
80
 
@@ -24,7 +24,7 @@ Use this package when you need server-backed persistence with pooled connections
24
24
  - managed cloud databases (RDS, Cloud SQL, Neon, Supabase, etc.)
25
25
  - CI integration tests against a real PostgreSQL service
26
26
 
27
- Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests.
27
+ Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests. This adapter stores sessions/runs, not semantic vectors; use the separate [`@arnilo/prism-memory` pgvector path](working-and-semantic-memory.md), which rejects non-finite vectors before SQL, when vector recall is needed.
28
28
 
29
29
  ## Inputs / request
30
30
 
@@ -59,7 +59,7 @@ Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backu
59
59
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
60
60
  | `close()` | Ends the pool when the adapter created it from `connectionString`. |
61
61
 
62
- Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks.
62
+ Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks. While holding that lock, startup verifies ordered contract name/version/SHA-256 rows and full schema-v3 `information_schema`/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
63
63
 
64
64
  ## Request/response example
65
65
 
@@ -127,7 +127,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
127
127
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
128
128
  - **Bounded pool.** Adapter-owned pools default to `max: 10`. Hosts with heavy concurrency should supply their own pool sizing.
129
129
  - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid sequential scans.
130
- - **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once.
130
+ - **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once. Startup catalog reads are bounded metadata queries, not application-row scans.
131
131
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns participate in query filters; hosts must still scope writes correctly.
132
132
  - **Benchmark target.** Indexed append + paginated branch read on a warm pool should stay under **50 ms p95** for local/CI-sized datasets (≤100k entries per session); measure with your pool size and hardware before production sizing.
133
133
 
@@ -21,6 +21,7 @@ Use this page when a host or provider package needs to:
21
21
  - Carry a stable cache key across turns without putting provider-specific fields in core.
22
22
  - Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
23
23
  - Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
24
+ - Understand when **caller-gated model discovery** may fill `ModelConfig.cache` / `ModelConfig.cost` from a live `/models` response (see [Discovery and live cache/cost metadata](#discovery-and-live-cache-cost-metadata)).
24
25
 
25
26
  Do not use cache keys for credentials, bearer tokens, API keys, OAuth tokens, user secrets, or raw private prompts.
26
27
 
@@ -145,21 +146,23 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
145
146
  | Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
146
147
  | --- | --- | --- | --- | --- |
147
148
  | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
148
- | `@arnilo/prism-provider-openrouter` | `cache_control` | Applies `cache_control` markers only to caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. | Breakpoint-stable prefixes can be reused by upstream providers. | Best-effort only; no marker is added to every block. |
149
+ | `@arnilo/prism-provider-openrouter` | `cache_control` | Top-level automatic `cache_control` when enabled without breakpoints; otherwise markers only on caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable / automatic prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
149
150
  | `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
150
151
  | `@arnilo/prism-provider-zai` | `implicit` | No explicit cache payload; GLM context caching is automatic. | Resend unchanged prior history for implicit context-cache reuse. | Best-effort only; cache options do not force hits. |
151
152
  | `@arnilo/prism-provider-kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
152
153
  | `@arnilo/prism-provider-neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
154
+ | `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
153
155
 
154
156
  Detailed first-party provider notes:
155
157
 
156
- - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
158
+ - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. GPT-5.6+ official docs use `prompt_cache_options` / breakpoints instead of retention — `listOpenAIModels` sets `longRetention: false` for those ids. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
157
159
  - OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
158
- - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars; applies Anthropic-style `cache_control` markers only to caller-selected `cache.breakpoints` (last content block of each selected message), not every block; `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
159
- - OpenCode Go (`@arnilo/prism-provider-opencode-go`): `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route sends none. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`.
160
+ - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing; with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
161
+ - OpenCode Go (`@arnilo/prism-provider-opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
160
162
  - Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
161
163
  - NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
162
164
  - Kimi (`@arnilo/prism-provider-kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
165
+ - AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
163
166
 
164
167
  ### NeuralWatt cache-aware limiter
165
168
 
@@ -185,6 +188,15 @@ sessions differently from one-shot chat:
185
188
  See [NeuralWatt provider](providers/neuralwatt.md) for the package-level cache,
186
189
  usage, and retry details.
187
190
 
191
+ ## Discovery and live cache/cost metadata
192
+
193
+ Caller-gated `list*Models()` helpers (see [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery)) may map official list-models fields onto `ModelConfig`:
194
+
195
+ - `cache` — when the provider documents cache kind / long-retention / breakpoint support in model metadata (otherwise keep the package's known default, e.g. NeuralWatt/Z.AI `implicit`, OpenAI `openai_key`).
196
+ - `cost` — when the provider documents per-token or per-million rates, including cache-read rates such as NeuralWatt `cached_input_per_million`.
197
+
198
+ Static featured catalogs remain offline bootstrap and must **not** invent pricing or cache capabilities the official docs do not state. Discovery is never invoked by `create*ProviderPackage()`; hosts that want live `cost`/`cache` pass the returned models into package `models:` (or register them themselves).
199
+
188
200
  ## Security and performance notes
189
201
 
190
202
  - Cache hints are best-effort and do not guarantee cache hits.
@@ -133,6 +133,44 @@ await assertProviderStreamConforms({
133
133
  });
134
134
  ```
135
135
 
136
+ ## Model discovery checklist
137
+
138
+ Every first-party package that ships (or plans) a `list*Models()` helper must keep setup network-free. Add these assertions in the package suite (pattern from NeuralWatt):
139
+
140
+ 1. **`*_provider_setup_does_not_call_model_discovery`** — inject a counting `fetch` into `create*ProviderPackage({ fetch })`, run `setup`, assert `calls === 0`.
141
+ 2. **`list_*_models_maps_fixture_…`** — fixture response maps to `ModelConfig` (`id` → `model`, documented capabilities/limits/cost/cache); no credentials in returned objects.
142
+ 3. **`list_*_models_forwards_auth_abort_baseurl`** (as applicable) — Authorization owned by helper when key present; auth omitted when optional and unset; `signal` / `baseUrl` forwarded.
143
+ 4. **`list_*_models_redacts_token_in_errors`** — non-OK bodies use `readBoundedResponseText` + `redactSecrets`; secret canaries absent from thrown messages.
144
+ 5. **Malformed payload rejects** — missing `data` array (or provider-equivalent) throws a clear discovery error.
145
+
146
+ OpenRouter stays app-registration-first: an optional list helper must still not run during setup. AI SDK has no discovery export. Packages without a public list API document curated official-doc refresh instead of inventing a fake helper.
147
+
148
+ Canonical contract: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery).
149
+
150
+ ## Thinking / reasoning checklist
151
+
152
+ Every first-party package that exposes thinking or reasoning controls should cover:
153
+
154
+ 1. **Model default** — `ModelConfig.compat` (or documented capability) sets the official wire field when no per-turn override is present.
155
+ 2. **Per-turn override wins** — `ProviderRequestOptions.compat` via `mergeProviderRequestOptions` / `applyThinkingLevel` overrides the model default.
156
+ 3. **Shared family mapping** — effort levels from `ThinkingLevel` land in the package's recommended family (`openai_reasoning` / `reasoning_effort` / `thinking_type` / `noop`) per [Thinking and reasoning](thinking-and-reasoning.md).
157
+ 4. **Non-reasoning / noop** — applying a level with `noop` (or omitting compat) must not invent unsupported body fields.
158
+ 5. **No inert `extra.thinkingLevel`** — package code must not rely on `options.extra.thinkingLevel` for wire mapping.
159
+
160
+ Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md).
161
+
162
+ ## AI SDK adapter checklist
163
+
164
+ `@arnilo/prism-provider-ai-sdk` is a host-owned `LanguageModelV4` bridge. It does not participate in the discovery or thinking/reasoning checklists above. Cover instead:
165
+
166
+ 1. **No catalog / no setup fetch** — package exports no `list*Models()`; `createAiSdkProvider` wraps a host model only.
167
+ 2. **Specification gate** — rejects non-v4 models (`specificationVersion !== "v4"` or missing `doStream`).
168
+ 3. **Cache usage mapping** — `finish.usage.inputTokens.cacheRead`/`cacheWrite` map to `Usage.cacheReadTokens`/`cacheWriteTokens`; adapter does not emit cache request fields.
169
+ 4. **Reasoning stream mapping** — `reasoning-delta` → thinking deltas; assistant `thinking` blocks replay as AI SDK `reasoning` prompt parts.
170
+ 5. **Host-owned controls** — `options.compat` / `options.extra` forward as `providerOptions.prism`; reasoning effort stays on the host model.
171
+
172
+ Canonical contract: [AI SDK provider adapter](providers/ai-sdk.md).
173
+
136
174
  ## Extension and configuration notes
137
175
 
138
176
  The helpers are a testing subpath only. Provider packages can use them with their own mocked fetch/transport or `createMockProvider()`. Live provider tests should stay opt-in and env-gated outside Prism's default test suite.
@@ -148,7 +186,7 @@ The helpers are a testing subpath only. Provider packages can use them with thei
148
186
  ## Related APIs
149
187
 
150
188
  - [Provider layer](provider-layer.md): `AIProvider`, provider events, and mock provider.
151
- - [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters.
189
+ - [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters; includes the caller-gated discovery contract and setup zero-fetch rule.
152
190
  - [AI SDK provider adapter](providers/ai-sdk.md): optional `LanguageModelV4` bridge tested with a fake AI SDK model.
153
191
  - [OpenAI-compatible provider](providers/openai-compatible.md): optional provider adapter tested with mocked streams.
154
192
  - [Public contracts](public-contracts.md): provider request/event/usage contracts.
@@ -67,7 +67,7 @@ Phase 6 also adds optional [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md
67
67
 
68
68
  Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
69
69
 
70
- These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
70
+ These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only, with optional `models`/`codexModels` overrides and an opt-in `listOpenAIModels()` helper for official `GET /models` discovery. `@arnilo/prism-provider-opencode-go` now registers docs-verified OpenCode Go open coding models with dual OpenAI/Anthropic routes (`compat.route`), official default base `https://opencode.ai/zen/go/v1`, `reasoning_content`/thinking preserve, and an opt-in `listOpenCodeGoModels()` helper for official `GET /zen/go/v1/models`. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/`reasoning`/cache passthrough, assistant `reasoning` replay, optional top-level automatic `cache_control`, and an opt-in `listOpenRouterModels()` helper for official `GET /api/v1/models` (setup still never fetches). `@arnilo/prism-provider-zai` now registers featured GLM-5.x/4.x metadata with official `thinking`/`reasoning_effort`/`tool_stream`/`clear_thinking` mapping, Preserved Thinking `reasoning_content` replay, implicit context caching, and an opt-in `listZaiModels()` helper for OpenAI-compatible `GET /models`. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default, optional callable Moonshot Open Platform Chat Completions when `includeMoonshotModels` is requested, official Coding/Open Platform featured ids, thinking/`reasoning_effort` compat mapping, and an opt-in `listKimiModels()` helper for Moonshot `GET /v1/models`. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
71
71
 
72
72
  ### First-party cache behavior
73
73
 
@@ -75,14 +75,71 @@ Every first-party provider package hardens prompt-cache behavior so it cannot em
75
75
 
76
76
  - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
77
77
  - **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
78
- - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; `cache_control` markers applied only to caller-selected `cache.breakpoints` (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
79
- - **OpenCode Go**: `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; the Anthropic route applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`), the OpenAI route sends none. Per-route usage mapping.
78
+ - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
79
+ - **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
80
80
  - **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
81
81
  - **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
82
82
  - **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
83
83
 
84
84
  See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
85
85
 
86
+ ## Caller-gated model discovery
87
+
88
+ First-party packages keep `create*ProviderPackage()` network-free. Latest models come from **caller-gated** `list*Models()` helpers that hosts invoke explicitly and then pass back via `models:` (or register themselves). Plan 015's "no setup catalog fetch" rule still holds; Plan 067 adds on-demand discovery without hidden latency.
89
+
90
+ ### Contract
91
+
92
+ ```ts
93
+ export async function listExampleModels(options: {
94
+ apiKey?: CredentialValueSource;
95
+ fetch?: typeof fetch;
96
+ baseUrl?: string;
97
+ signal?: AbortSignal;
98
+ headers?: Readonly<Record<string, string>>;
99
+ }): Promise<ModelConfig[]> {
100
+ // GET {baseUrl}/models — never called from create*ProviderPackage()
101
+ }
102
+ ```
103
+
104
+ | Rule | Requirement |
105
+ | --- | --- |
106
+ | Setup | `create*ProviderPackage().setup` performs **zero** fetches / discovery calls |
107
+ | Shape | Package-local `list*Models(options) → Promise<ModelConfig[]>` + optional `map*Model(entry)` |
108
+ | Injectables | `fetch`, `baseUrl`, `signal`, optional `apiKey` / `headers` |
109
+ | Transport | Error bodies via `@arnilo/prism/providers/transport` `readBoundedResponseText`; credentials via `resolveCredentialValue` + `redactSecrets` |
110
+ | Return | `ModelConfig[]` only — never embed API keys, tokens, or auth headers in returned metadata |
111
+ | Static catalog | Featured aliases / offline bootstrap only; may omit live pricing until discovery fills `cost` / `cache` |
112
+ | Core | Prefer package-local helpers. Do **not** add a core model-discovery registry. Extract a shared HTTP/list helper only when ≥2 packages share identical parsing |
113
+
114
+ Template: [`listNeuralWattModels`](providers/neuralwatt.md) in `@arnilo/prism-provider-neuralwatt`.
115
+
116
+ ### Per-package policy
117
+
118
+ | Package | Discovery helper | Setup catalog | Notes |
119
+ | --- | --- | --- | --- |
120
+ | OpenAI | **`listOpenAIModels` (exists)** | Featured Responses/Codex aliases; factory accepts `models?` / `codexModels?` | Official `GET /v1/models`; Codex not listed by api.openai.com |
121
+ | Kimi | **`listKimiModels`** (Moonshot `GET /v1/models`) | Featured Coding ids + optional callable Moonshot | Official Moonshot/Kimi list-models; Coding curated |
122
+ | Z.AI | **`listZaiModels`** (OpenAI-compatible `GET /models`) + curated featured refresh | Featured GLM-5.2…4.5 aliases | No first-class docs.z.ai list page; discovery is best-effort; featured set from Chat Completions enum / overview |
123
+ | OpenRouter | **`listOpenRouterModels`** (official `GET /api/v1/models`) | **App-controlled** `models:` only — no bundled mega-catalog | Helper feeds host registration; setup still does not fetch |
124
+ | OpenCode Go | **`listOpenCodeGoModels`** (official `GET /zen/go/v1/models`) | Featured dual-route official Go aliases | Official Go docs endpoint table + sparse list API |
125
+ | NeuralWatt | **`listNeuralWattModels` (exists)** | Featured aliases without guessed pricing | Auth optional for public models |
126
+ | AI SDK | None | Host-owned `LanguageModelV4` | No Prism-side catalog by design |
127
+
128
+ Host pattern:
129
+
130
+ ```ts
131
+ const models = await listNeuralWattModels({ apiKey, fetch });
132
+ await kernel.load([createNeuralWattProviderPackage({ apiKey, models })]);
133
+ ```
134
+
135
+ Discovery may populate `ModelConfig.cache` and `ModelConfig.cost` from live metadata when the provider documents those fields; see [Provider caching](provider-caching.md#discovery-and-live-cache-cost-metadata). Package authors: include the [setup zero-fetch checklist](provider-conformance.md#model-discovery-checklist) in every first-party suite that ships or plans a `list*Models` helper.
136
+
137
+ ## Per-turn thinking / reasoning
138
+
139
+ Hosts set effort with portable helpers from `@arnilo/prism` (`applyThinkingLevel`, `thinkingCompatFor`) that write official fields into `ProviderRequestOptions.compat`. Model defaults stay on `ModelConfig.compat`; per-turn patches win via `mergeProviderRequestOptions`. Providers keep reading `options.compat` / `model.compat` — do not invent a parallel options tree or put effort only in `extra`.
140
+
141
+ Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md). Package-local knobs (NeuralWatt budgets, Z.AI `tool_stream`, Kimi keep/all) remain on `compat` beside the shared families.
142
+
86
143
  ## Third-party provider packaging
87
144
 
88
145
  A third party ships their own providers the same way Prism ships first-party
@@ -52,6 +52,8 @@ Unsupported content fails before `doStream` (for example unresolved `resourceUri
52
52
  | `finish` usage | `usage` then `done` |
53
53
  | `error` / thrown / abort | redacted `error` |
54
54
 
55
+ `finish.usage.inputTokens.cacheRead` / `cacheWrite` map to Prism `Usage.cacheReadTokens` / `cacheWriteTokens`. The adapter does not invent cache request fields; prompt caching is owned by the host `LanguageModelV4` and its upstream provider.
56
+
55
57
  Provider-executed tool calls, files/sources/custom parts, warnings, and raw chunks are ignored rather than silently converted into unsupported Prism content.
56
58
 
57
59
  ## Request/response example
@@ -89,6 +91,40 @@ const result = await agent.createSession().run("Summarize this");
89
91
  console.log(result.text);
90
92
  ```
91
93
 
94
+ ## Model catalog and discovery
95
+
96
+ There is **no Prism-side model catalog** and **no `list*Models()` export** by design. Hosts supply a ready-made `LanguageModelV4` instance (typically from `@ai-sdk/openai`, `@ai-sdk/anthropic`, AI Gateway, or a custom provider) and register a matching `ModelConfig` for capabilities/limits.
97
+
98
+ Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials.
99
+
100
+ ## Prompt caching
101
+
102
+ The adapter is **host-owned for request caching**. It does not emit `cache_control`, `prompt_cache_key`, `cacheKey`, or `cacheRetention` on AI SDK call options. Hosts configure caching on the underlying AI SDK model/provider (for example via AI SDK `providerOptions` on the model factory or per-call options forwarded through `options.compat` / `options.extra` → `providerOptions.prism`).
103
+
104
+ When the host model reports cache accounting on the `finish` stream part, Prism maps official AI SDK v4 usage fields:
105
+
106
+ | AI SDK `LanguageModelV4Usage` | Prism `Usage` |
107
+ | --- | --- |
108
+ | `inputTokens.cacheRead` | `cacheReadTokens` |
109
+ | `inputTokens.cacheWrite` | `cacheWriteTokens` |
110
+ | `inputTokens.total` | `inputTokens` |
111
+ | `outputTokens.total` | `outputTokens` |
112
+
113
+ See [Provider caching](../provider-caching.md) for the cross-provider matrix.
114
+
115
+ ## Thinking and reasoning
116
+
117
+ Reasoning effort, budgets, and provider-specific thinking controls are **host-model-owned**. Prism does not map `ThinkingLevel` into AI SDK call options (`thinkingFamilyForModel` → `noop`). Hosts configure reasoning on the AI SDK model (for example OpenAI `reasoning.effort` via AI SDK `providerOptions`) and may pass per-turn overrides through `ProviderRequestOptions.compat` / `extra`, which the adapter forwards as `providerOptions.prism`.
118
+
119
+ Stream mapping:
120
+
121
+ | Direction | Mapping |
122
+ | --- | --- |
123
+ | AI SDK `reasoning-delta` → Prism | `content_delta` thinking |
124
+ | Prism `thinking` blocks → AI SDK prompt | `{ type: "reasoning", text }` on assistant messages |
125
+
126
+ Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); [Language Model Specification V4](https://github.com/vercel/ai/tree/main/packages/provider/src/language-model/v4); `@ai-sdk/provider` `LanguageModelV4Usage` (`inputTokens.cacheRead` / `cacheWrite`).
127
+
92
128
  ## Extension and configuration notes
93
129
 
94
130
  - Peer dependency: `@ai-sdk/provider@^4.0.0`. Upgrade policy tracks one specification major at a time.