@arnilo/prism 0.0.13 → 0.0.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.md +23 -2
  2. package/README.md +9 -2
  3. package/dist/agent-loops.d.ts +4 -0
  4. package/dist/agent-loops.js +16 -3
  5. package/dist/artifacts.d.ts +78 -0
  6. package/dist/artifacts.js +24 -0
  7. package/dist/contracts.d.ts +86 -0
  8. package/dist/contracts.js +8 -0
  9. package/dist/conversations.d.ts +50 -0
  10. package/dist/conversations.js +97 -0
  11. package/dist/credentials.d.ts +14 -0
  12. package/dist/credentials.js +9 -0
  13. package/dist/devices.d.ts +94 -0
  14. package/dist/devices.js +138 -0
  15. package/dist/index.d.ts +12 -6
  16. package/dist/index.js +7 -4
  17. package/dist/provider-events.d.ts +1 -0
  18. package/dist/provider-events.js +3 -0
  19. package/dist/providers/openai-primitives.js +5 -2
  20. package/docs/ag-ui.md +5 -0
  21. package/docs/browser-automation.md +3 -0
  22. package/docs/conversations.md +135 -0
  23. package/docs/credential-storage.md +28 -1
  24. package/docs/credentials-and-redaction.md +2 -0
  25. package/docs/database-persistence.md +5 -1
  26. package/docs/device-adapters.md +97 -0
  27. package/docs/host-security.md +7 -2
  28. package/docs/index.md +24 -19
  29. package/docs/migration.md +50 -1
  30. package/docs/multimodal-content.md +8 -5
  31. package/docs/performance.md +36 -0
  32. package/docs/policy-and-audit.md +1 -0
  33. package/docs/provider-caching.md +12 -0
  34. package/docs/provider-conformance.md +29 -5
  35. package/docs/provider-packages.md +26 -2
  36. package/docs/providers/ai-sdk.md +23 -7
  37. package/docs/providers/alibaba.md +179 -0
  38. package/docs/providers/ollama.md +166 -0
  39. package/docs/providers/openai.md +22 -3
  40. package/docs/rag.md +41 -12
  41. package/docs/release-and-install.md +135 -17
  42. package/docs/resource-loading.md +3 -0
  43. package/docs/review-coverage-2026-07-25-phase-9.md +256 -0
  44. package/docs/review-coverage-2026-07-26-phase-10.md +132 -0
  45. package/docs/server.md +4 -0
  46. package/docs/work-artifacts-and-review.md +100 -0
  47. package/docs/work-connectors.md +5 -1
  48. package/docs/work-tools.md +3 -0
  49. package/docs/workflows.md +4 -0
  50. package/docs/working-and-semantic-memory.md +40 -7
  51. package/package.json +1 -1
  52. package/templates/init/providers.json +22 -0
@@ -21,7 +21,28 @@ Use these helpers in provider package tests to check event order, terminal event
21
21
 
22
22
  Do not use them as a live integration runner, provider simulator, retry framework, credential loader, or test framework replacement.
23
23
 
24
- For real network smoke tests, each first-party provider package ships an env-gated `src/__tests__/live.test.ts` that exercises the live API when `PRISM_LIVE_PROVIDER_TESTS=1` and a provider-specific API key are set. These live tests reuse the same conformance helpers (`assertProviderStreamConforms`, `assertAbortIsObserved`, `assertNoSecretLeak`) against the real provider, so offline and live assertions stay consistent. The default `npm test` never sets these gates and stays network-free; see [Release and install](release-and-install.md) for the full env-var list.
24
+ Offline conformance is mandatory for every package; credentialed probes are not uniform. Packages with a checked-in `live.test.ts` use `PRISM_LIVE_PROVIDER_TESTS=1` plus their provider key. Realtime, hosted tools, AI SDK host models, Alibaba/Ollama account or daemon paths, and enterprise workload identities need host-owned protected probes instead of a generic fixture. The default `npm test` never sets these gates and stays network-free; see [Release and install](release-and-install.md#015-protected-live-canary-matrix) for the exact environment/key boundary.
25
+
26
+ ## Phase 10 provider conformance matrix
27
+
28
+ | Package | Required offline evidence | Restricted live evidence |
29
+ | --- | --- | --- |
30
+ | OpenAI | Responses serialization/stream ordering, provider-hosted authority, continuation cap/cursor, Realtime fake WebSocket caps | Standard API-key smoke; separate protected hosted-tool/Realtime entitlement probe |
31
+ | AI SDK | Exact 4.0.3/V4 gate; every mapped stream part; authority, cache usage, redaction, unsupported mapping | Host-created V4 model only; no Prism credential fixture |
32
+ | Anthropic | Messages serialization, cache/thinking/tools, header/redaction/abort assertions | Protected `ANTHROPIC_API_KEY` smoke |
33
+ | Google | `generateContent` serialization, complete tool calls, media/abort/redaction assertions | Protected `GOOGLE_API_KEY` or `GEMINI_API_KEY` smoke |
34
+ | Kimi | Coding/Moonshot route fixtures, thinking/tool reconstruction, headers/redaction | Protected `KIMI_API_KEY` smoke |
35
+ | Z.AI | GLM thinking/tool-stream fixtures, implicit-cache usage, headers/redaction | Protected `ZAI_API_KEY` smoke |
36
+ | OpenRouter | routing/reasoning/cache-control fixture, stream/tool reconstruction, headers/redaction | Protected `OPENROUTER_API_KEY` smoke |
37
+ | OpenCode Go | OpenAI/Anthropic route fixture, completion proof, PDF/media boundary, headers/redaction | Protected `OPENCODE_API_KEY` smoke |
38
+ | Alibaba | DashScope presets, Qwen thinking, image rejection/mapping, cache/usage fixture | Protected account/region host probe; no generic key fixture |
39
+ | Ollama | cloud/local preset, reasoning/image mapping, implicit-cache fixture | Protected cloud or host-local authenticated daemon probe; no daemon starts in tests |
40
+ | NeuralWatt | stream/retry/quota/telemetry fixtures, implicit-cache usage, headers/redaction | Protected `NEURALWATT_API_KEY` smoke |
41
+ | Azure | endpoint preservation, Entra/resource-key header and OpenAI-compatible stream fixture | Protected host workload-identity probe |
42
+ | Bedrock | SigV4/region/PrivateLink and OpenAI-compatible stream fixture | Protected host IAM/IRSA probe |
43
+ | Vertex | location/endpoint preservation, ADC header and OpenAI-compatible stream fixture | Protected host ADC/WIF probe |
44
+
45
+ All rows must retain bounded request/response fixtures, abort propagation, provider-owned-header precedence, and fake-secret leak assertions where the package surfaces those values. A successful fake transport proves Prism mapping, not account entitlement or vendor availability.
25
46
 
26
47
  ## Inputs / request
27
48
 
@@ -50,6 +71,7 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
50
71
  - `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
51
72
  - `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.
52
73
  - `assertNoSecretLeak()` stringifies all collected events and fails if any known secret string is present.
74
+ - Provider-hosted calls must surface as `tool_call` with `authority: "provider-hosted"`; loops record them but never dispatch them or append a host `tool_result`. Bounded continuations must emit an opaque cursor event and end in `done` or redacted `error`, never silent truncation.
53
75
 
54
76
  ## Request/response example
55
77
 
@@ -164,10 +186,12 @@ Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md).
164
186
  `@arnilo/prism-provider-ai-sdk` is a host-owned `LanguageModelV4` bridge. It does not participate in the discovery or thinking/reasoning checklists above. Cover instead:
165
187
 
166
188
  1. **No catalog / no setup fetch** — package exports no `list*Models()`; `createAiSdkProvider` wraps a host model only.
167
- 2. **Specification gate** — rejects non-v4 models (`specificationVersion !== "v4"` or missing `doStream`).
168
- 3. **Cache usage mapping** — `finish.usage.inputTokens.cacheRead`/`cacheWrite` map to `Usage.cacheReadTokens`/`cacheWriteTokens`; adapter does not emit cache request fields.
169
- 4. **Reasoning stream mapping** — `reasoning-delta` → thinking deltas; assistant `thinking` blocks replay as AI SDK `reasoning` prompt parts.
170
- 5. **Host-owned controls** — `options.compat` / `options.extra` forward as `providerOptions.prism`; reasoning effort stays on the host model.
189
+ 2. **Version + specification gate** — exact `@ai-sdk/provider` matrix version is verified at setup; rejects version skew, non-v4 models (`specificationVersion !== "v4"`), or missing `doStream`.
190
+ 3. **Mapping table** — fixture covers every supported matrix row: response metadata id, text/reasoning/tool deltas, client/provider-hosted tool authority, structured output, cache usage, finish/error/abort; unmappable stream parts and `structuredOutput.strict` fail typed instead of dropping.
191
+ 4. **Cache usage mapping** — `finish.usage.inputTokens.cacheRead`/`cacheWrite` map to `Usage.cacheReadTokens`/`cacheWriteTokens`; adapter does not emit cache request fields.
192
+ 5. **Reasoning stream mapping** — `reasoning-delta` → thinking deltas; assistant `thinking` blocks replay as AI SDK `reasoning` prompt parts.
193
+ 6. **Redaction** — direct adapter errors use its supplied `SecretRedactor`; agent runs use their active redactor; opaque provider metadata is never emitted.
194
+ 7. **Host-owned controls** — `options.compat` / `options.extra` forward as `providerOptions.prism`; reasoning effort stays on the host model.
171
195
 
172
196
  Canonical contract: [AI SDK provider adapter](providers/ai-sdk.md).
173
197
 
@@ -80,7 +80,28 @@ Phase 6 also adds optional [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md
80
80
 
81
81
  Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
82
82
 
83
- These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only, with optional `models`/`codexModels` overrides and an opt-in `listOpenAIModels()` helper for official `GET /models` discovery. `@arnilo/prism-provider-opencode-go` now registers docs-verified OpenCode Go open coding models with dual OpenAI/Anthropic routes (`compat.route`), official default base `https://opencode.ai/zen/go/v1`, `reasoning_content`/thinking preserve, and an opt-in `listOpenCodeGoModels()` helper for official `GET /zen/go/v1/models`. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/`reasoning`/cache passthrough, assistant `reasoning` replay, optional top-level automatic `cache_control`, and an opt-in `listOpenRouterModels()` helper for official `GET /api/v1/models` (setup still never fetches). `@arnilo/prism-provider-zai` now registers featured GLM-5.x/4.x metadata with official `thinking`/`reasoning_effort`/`tool_stream`/`clear_thinking` mapping, Preserved Thinking `reasoning_content` replay, implicit context caching, and an opt-in `listZaiModels()` helper for OpenAI-compatible `GET /models`. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default, optional callable Moonshot Open Platform Chat Completions when `includeMoonshotModels` is requested, official Coding/Open Platform featured ids, thinking/`reasoning_effort` compat mapping, and an opt-in `listKimiModels()` helper for Moonshot `GET /v1/models`. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation. `@arnilo/prism-provider-anthropic` registers native Anthropic Messages (`createAnthropicProviderPackage` / `listAnthropicModels`). `@arnilo/prism-provider-google` registers native Gemini `generateContent` streaming (`createGoogleProviderPackage` / `listGoogleModels`; Vertex identity deferred). Both follow the same zero-setup-network / host-owned credential / provider-owned-header rules; see [`docs/providers/anthropic.md`](providers/anthropic.md) and [`docs/providers/google.md`](providers/google.md). `@arnilo/prism-provider-anthropic` registers native Anthropic Messages (`createAnthropicProviderPackage` / `listAnthropicModels`). `@arnilo/prism-provider-google` registers native Gemini `generateContent` streaming (`createGoogleProviderPackage` / `listGoogleModels`; Vertex identity deferred). Both follow the same zero-setup-network / host-owned credential / provider-owned-header rules; see [`docs/providers/anthropic.md`](providers/anthropic.md) and [`docs/providers/google.md`](providers/google.md).
83
+ These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only, with optional `models`/`codexModels` overrides and an opt-in `listOpenAIModels()` helper for official `GET /models` discovery. `@arnilo/prism-provider-opencode-go` now registers docs-verified OpenCode Go open coding models with dual OpenAI/Anthropic routes (`compat.route`), official default base `https://opencode.ai/zen/go/v1`, `reasoning_content`/thinking preserve, and an opt-in `listOpenCodeGoModels()` helper for official `GET /zen/go/v1/models`. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/`reasoning`/cache passthrough, assistant `reasoning` replay, optional top-level automatic `cache_control`, and an opt-in `listOpenRouterModels()` helper for official `GET /api/v1/models` (setup still never fetches). `@arnilo/prism-provider-zai` now registers featured GLM-5.x/4.x metadata with official `thinking`/`reasoning_effort`/`tool_stream`/`clear_thinking` mapping, Preserved Thinking `reasoning_content` replay, implicit context caching, and an opt-in `listZaiModels()` helper for OpenAI-compatible `GET /models`. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default, optional callable Moonshot Open Platform Chat Completions when `includeMoonshotModels` is requested, official Coding/Open Platform featured ids, thinking/`reasoning_effort` compat mapping, and an opt-in `listKimiModels()` helper for Moonshot `GET /v1/models`. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation. `@arnilo/prism-provider-anthropic` registers native Anthropic Messages (`createAnthropicProviderPackage` / `listAnthropicModels`). `@arnilo/prism-provider-google` registers native Gemini `generateContent` streaming (`createGoogleProviderPackage` / `listGoogleModels`; Vertex identity stays in the separate package). Both follow the same zero-setup-network / host-owned credential / provider-owned-header rules; see [`docs/providers/anthropic.md`](providers/anthropic.md) and [`docs/providers/google.md`](providers/google.md).
84
+
85
+ ### Phase 10 compatibility matrix
86
+
87
+ Every package remains explicit, setup-zero-fetch, and late-credential-bound. `ModelConfig.capabilities.input` is authoritative: a listed wire mapping is usable only when the selected model declares that tag; unsupported blocks reject before provider I/O. “Protected” means an operator/release-environment probe, never default CI; its exact key/command boundary is in [Release and install](release-and-install.md#015-protected-live-canary-matrix).
88
+
89
+ | Package | Protocol / model source | Content mapping | Stream, tools, and reasoning | Cache / canary |
90
+ | --- | --- | --- | --- | --- |
91
+ | OpenAI | Responses; featured or caller-gated `listOpenAIModels` | text, image, audio, file, document | Host and provider-hosted tools; 8-hop continuation; Realtime seam; Responses reasoning | `openai_key`; checked-in standard smoke + protected hosted/Realtime probe |
92
+ | AI SDK | Host `LanguageModelV4`; no Prism catalog | declared text/image/audio/file/document prompt parts (role-limited) | v4 mapping; provider-executed tool authority; host-owned reasoning | host-owned; exact 4.0.3 matrix; protected host integration |
93
+ | Anthropic | Messages; caller-gated list | text, image, PDF document/file | tool deltas, thinking | `cache_control`; protected API-key smoke |
94
+ | Google | Gemini `generateContent`; caller-gated list | text, image, audio, document/file | complete tool calls, thinking | no Prism cache marker; protected API-key smoke |
95
+ | Kimi | Coding Messages or opt-in Moonshot; caller-gated list | text, image, PDF document/file by route/model | tool deltas, route-native thinking replay | implicit / optional Anthropic markers; protected API-key smoke |
96
+ | Z.AI | OpenAI-compatible; caller-gated list | text, image | tool deltas, `reasoning_content` | implicit; protected API-key smoke |
97
+ | OpenRouter | OpenAI-compatible; host catalog + caller-gated list | text, image | tool deltas, reasoning replay/routing metadata | `cache_control`; protected API-key smoke |
98
+ | OpenCode Go | OpenAI or Anthropic route; caller-gated list | text/image OpenAI route; PDF document/file Anthropic route | tool deltas, route-native thinking | route-specific; protected API-key smoke |
99
+ | Alibaba | DashScope OpenAI-compatible; caller-gated list | text, image | tool deltas, Qwen thinking | implicit / optional markers; protected host probe |
100
+ | Ollama | Cloud/local OpenAI-compatible; caller-gated list | text, image | tool deltas, reasoning effort | implicit only; protected host/daemon probe |
101
+ | NeuralWatt | OpenAI-compatible; caller-gated list | text, image | tool deltas, reasoning and telemetry | implicit; protected API-key smoke |
102
+ | Azure | Azure/Foundry OpenAI-compatible; host models | selected endpoint/model capability | normalized OpenAI-compatible tools | no Prism cache mapping; protected host workload-identity probe |
103
+ | Bedrock | Bedrock OpenAI-compatible; host models | selected endpoint/model capability | normalized OpenAI-compatible tools | no Prism cache mapping; protected host IAM/IRSA probe |
104
+ | Vertex | Vertex OpenAPI-compatible; host models | selected endpoint/model capability | normalized OpenAI-compatible tools | no Prism cache mapping; protected host ADC/WIF probe |
84
105
 
85
106
  ### First-party cache behavior
86
107
 
@@ -93,6 +114,8 @@ Every first-party provider package hardens prompt-cache behavior so it cannot em
93
114
  - **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
94
115
  - **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
95
116
  - **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
117
+ - **Alibaba Cloud** (implicit by default, optional `cache_control`): DashScope implicit prefix caching is automatic; hosts opt in via `ModelConfig.cache.kind: cache_control`, then `cache_control: {"type":"ephemeral"}` markers apply only to selected breakpoints, capped at 4. `prompt_tokens_details.cached_tokens`/`cache_creation_input_tokens` map to cache usage. Caller-gated `listAlibabaModels`.
118
+ - **Ollama** (`kind: implicit`): Ollama KV/prefix caching is automatic with no request knob; sends no explicit cache payload. Ollama reports no cached-token count, so `Usage.cacheReadTokens` stays `undefined`. Caller-gated `listOllamaModels`.
96
119
 
97
120
  See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
98
121
 
@@ -160,7 +183,8 @@ provider packages: an `Extension` whose `setup(api)` calls
160
183
  `api.registerProvider(provider)` for each provider it owns. First-party
161
184
  provider packages (`@arnilo/prism-provider-openai`, `@arnilo/prism-provider-openrouter`,
162
185
  `@arnilo/prism-provider-kimi`, `@arnilo/prism-provider-zai`,
163
- `@arnilo/prism-provider-opencode-go`) are **opt-in and individually installable**;
186
+ `@arnilo/prism-provider-opencode-go`, `@arnilo/prism-provider-alibaba`,
187
+ `@arnilo/prism-provider-ollama`) are **opt-in and individually installable**;
164
188
  `@arnilo/prism` core runs without any first-party provider package (mock-only).
165
189
 
166
190
  A host mixes first-party packages and third-party providers in one resolver.
@@ -4,7 +4,15 @@
4
4
 
5
5
  `@arnilo/prism-provider-ai-sdk` adapts a host-supplied AI SDK `LanguageModelV4` into a Prism `AIProvider`. It maps Prism messages, tools, and structured-output options into `doStream` call options, then translates stream parts into Prism provider events incrementally.
6
6
 
7
- Supported specification: `@ai-sdk/provider` **v4** (`specificationVersion: "v4"`). Core `@arnilo/prism` does not depend on the AI SDK.
7
+ Core `@arnilo/prism` does not depend on the AI SDK.
8
+
9
+ ### Supported-version matrix
10
+
11
+ | `@ai-sdk/provider` | `LanguageModel` ABI | Status |
12
+ | --- | --- | --- |
13
+ | `4.0.3` | `LanguageModelV4`, `specificationVersion: "v4"` | Supported and offline-tested |
14
+
15
+ The peer dependency is intentionally exact. `createAiSdkProvider()` reads its resolved `@ai-sdk/provider/package.json` version during setup and throws typed `AiSdkProviderError { code: "unsupported_version" }` for an unlisted version; it does not infer compatibility from a matching `"v4"` string.
8
16
 
9
17
  ## When to use it
10
18
 
@@ -19,6 +27,7 @@ import { createAiSdkProvider } from "@arnilo/prism-provider-ai-sdk";
19
27
 
20
28
  createAiSdkProvider(options: {
21
29
  model: LanguageModelV4;
30
+ redactor?: SecretRedactor;
22
31
  id?: string;
23
32
  }): AIProvider
24
33
  ```
@@ -26,6 +35,7 @@ createAiSdkProvider(options: {
26
35
  | Field | Type | Purpose |
27
36
  | --- | --- | --- |
28
37
  | `model` | `LanguageModelV4` | Host-owned AI SDK language model. |
38
+ | `redactor` | `SecretRedactor` | Optional direct-provider error redactor; agent runs use their active redactor. |
29
39
  | `id` | `string` | Prism provider id. Defaults to `ai-sdk:<model.provider>` or `ai-sdk`. |
30
40
 
31
41
  Mapped request surfaces:
@@ -34,7 +44,7 @@ Mapped request surfaces:
34
44
  | --- | --- |
35
45
  | `messages` | `LanguageModelV4Prompt` |
36
46
  | `tools` | `LanguageModelV4FunctionTool[]` with JSON Schema `inputSchema` |
37
- | `options.structuredOutput` | `responseFormat: { type: "json", name, schema }` |
47
+ | `options.structuredOutput` | `responseFormat: { type: "json", name, schema }`; `strict` fails explicitly (V4 has no equivalent) |
38
48
  | `model.parameters` | `maxOutputTokens`, `temperature`, `topP`, `topK`, penalties, `seed`, `stopSequences` |
39
49
  | `request.signal` | `abortSignal` (always wins over adapter options) |
40
50
  | `options.headers` | extension headers only; model owns auth |
@@ -45,16 +55,22 @@ Unsupported content fails before `doStream` (for example unresolved `resourceUri
45
55
 
46
56
  | AI SDK stream part | Prism event |
47
57
  | --- | --- |
58
+ | `response-metadata.id` | `message_start.messageId` |
48
59
  | `text-delta` | `content_delta` text |
49
60
  | `reasoning-delta` | `content_delta` thinking |
50
61
  | `tool-input-start` / `tool-input-delta` | `tool_call_delta` |
51
- | `tool-call` (client-executed) | `tool_call` |
62
+ | `tool-call` | `tool_call`; `providerExecuted` becomes `authority: "provider-hosted"` |
52
63
  | `finish` usage | `usage` then `done` |
53
64
  | `error` / thrown / abort | redacted `error` |
65
+ | `stream-start`, boundaries, raw diagnostics | intentionally not emitted: no normalized safe payload |
66
+ | provider-executed `tool-result` | remains provider-side; no host result or dispatch |
67
+ | file / reasoning-file / source / custom / approval request | typed `unsupported_mapping` error |
68
+
69
+ `response-metadata.modelId`/timestamp and opaque `providerMetadata` have no normalized Prism counterpart and are not emitted, preventing provider-private metadata from entering prompt, event, or telemetry paths.
54
70
 
55
71
  `finish.usage.inputTokens.cacheRead` / `cacheWrite` map to Prism `Usage.cacheReadTokens` / `cacheWriteTokens`. The adapter does not invent cache request fields; prompt caching is owned by the host `LanguageModelV4` and its upstream provider.
56
72
 
57
- Provider-executed tool calls, files/sources/custom parts, warnings, and raw chunks are ignored rather than silently converted into unsupported Prism content.
73
+ No AI SDK stream part is silently coerced into Prism content: the table above maps safe normalized semantics, deliberately withholds provider-private diagnostics/results, and fails unsupported output types explicitly.
58
74
 
59
75
  ## Request/response example
60
76
 
@@ -127,7 +143,7 @@ Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/provi
127
143
 
128
144
  ## Extension and configuration notes
129
145
 
130
- - Peer dependency: `@ai-sdk/provider@^4.0.0`. Upgrade policy tracks one specification major at a time.
146
+ - Peer dependency: `@ai-sdk/provider@4.0.3`. Upgrade policy adds a matrix row and offline conformance fixture before accepting any new version.
131
147
  - First-party HTTP providers remain independent; this adapter is available directly, through `@arnilo/prism-providers`, or through `@arnilo/prism-all`. Installation does not select a model or invoke AI SDK.
132
148
  - `options.compat` / `options.extra` pass through as AI SDK `providerOptions.prism`.
133
149
  - Export helpers `toAiSdkCallOptions`, `toAiSdkPrompt`, and `mapAiSdkStream` for tests and custom hosts.
@@ -137,8 +153,8 @@ Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/provi
137
153
  - Host credentials stay inside the supplied AI SDK model. The adapter never reads env keys or credential stores.
138
154
  - Abort and resource limits come from Prism `request.signal`; adapter options cannot replace that bound.
139
155
  - Stream parts are translated incrementally with no full-response buffering and no duplicate model call.
140
- - Unsupported content fails closed before model invocation. Errors use Prism `providerError` redaction.
141
- - Provider metadata/warnings are not emitted as prompt or tool content.
156
+ - Unsupported content and stream parts fail closed before/at mapping; `structuredOutput.strict` is rejected because V4 cannot carry it.
157
+ - Pass `redactor` for direct use; agent runs apply their active redactor. Provider metadata/warnings never become prompt, tool, event, or telemetry content.
142
158
 
143
159
  ## Related APIs
144
160
 
@@ -0,0 +1,179 @@
1
+ # Alibaba Cloud provider package
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-alibaba` is a side-effect-free adapter for Alibaba Cloud
6
+ Model Studio / DashScope (including the Coding Plan) over the **OpenAI-compatible**
7
+ `POST {base}/chat/completions` endpoint.
8
+
9
+ - **Dynamic model discovery** — `listAlibabaModels()` calls the OpenAI-compatible
10
+ `GET {base}/models`. No model catalog is hard-coded in the package: available
11
+ models vary by region, workspace, and billing plan, so discovery is the source of
12
+ truth. Package setup never fetches.
13
+ - **Context cache** — DashScope implicit prefix caching is automatic. Explicit
14
+ caching is opt-in via Anthropic-style `cache_control: {"type":"ephemeral"}`
15
+ markers (at most 4 per request). Cache hits are accounted from
16
+ `usage.prompt_tokens_details.cached_tokens` (read) and
17
+ `cache_creation_input_tokens` (write).
18
+ - **Qwen thinking** — `enable_thinking` passthrough toggles reasoning on Qwen models.
19
+
20
+ The API key is region/plan-scoped: it must match the base URL's billing plan
21
+ (pay-as-you-go regional, workspace-dedicated, or Coding Plan).
22
+
23
+ ## When to use it
24
+
25
+ Use it when a host app wants Alibaba Cloud Qwen models (Model Studio / DashScope or
26
+ the Coding Plan) through Prism's `AgentSession` runtime with OpenAI-compatible
27
+ serialization, dynamic model discovery, and explicit/implicit cache accounting.
28
+
29
+ Do not use it for automatic credential discovery, setup-time catalog fetches, or
30
+ real-network tests (live tests stay opt-in).
31
+
32
+ ## Inputs / request
33
+
34
+ ```ts
35
+ import {
36
+ createAlibabaProviderPackage,
37
+ createAlibabaProvider,
38
+ listAlibabaModels,
39
+ defineAlibabaModel,
40
+ alibabaBaseUrl,
41
+ } from "@arnilo/prism-provider-alibaba";
42
+
43
+ createAlibabaProviderPackage(options: AlibabaProviderPackageOptions): ProviderPackage
44
+ createAlibabaProvider(options?: AlibabaProviderOptions): AIProvider
45
+ listAlibabaModels(options?: ListAlibabaModelsOptions): Promise<ModelConfig[]>
46
+ defineAlibabaModel(config: AlibabaModelConfig): ModelConfig
47
+ alibabaBaseUrl(options?: { baseUrl?: string; preset?: AlibabaBasePreset }): string
48
+ ```
49
+
50
+ | Field | Type | Purpose |
51
+ | --- | --- | --- |
52
+ | `apiKey` | `CredentialValueSource` | DashScope API key (`DASHSCOPE_API_KEY`), region/plan-scoped. |
53
+ | `baseUrl` | `string` | Explicit OpenAI-compatible base URL (wins over `preset`). |
54
+ | `preset` | `AlibabaBasePreset` | `"singapore"` (default) / `"beijing"` / `"us"` / `"coding-plan"`. |
55
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
56
+ | `id` | `string` | Provider id (default `alibaba`). |
57
+ | `models` | `readonly ModelConfig[]` | Host-supplied models (from `listAlibabaModels`) to register. |
58
+
59
+ Base URLs resolved by preset:
60
+
61
+ | Preset | Base URL |
62
+ | --- | --- |
63
+ | `singapore` | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` |
64
+ | `beijing` | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
65
+ | `us` | `https://dashscope-us.aliyuncs.com/compatible-mode/v1` |
66
+ | `coding-plan` | `https://coding-intl.dashscope.aliyuncs.com/v1` |
67
+
68
+ Workspace-dedicated endpoints
69
+ (`https://{workspaceId}.{region}.maas.aliyuncs.com/compatible-mode/v1`) are supplied
70
+ verbatim via `baseUrl`.
71
+
72
+ ## Outputs / response / events
73
+
74
+ | Surface | Behavior |
75
+ | --- | --- |
76
+ | Stream | Prism text deltas, `delta.reasoning_content` → thinking deltas, tool-call delta/final, `usage`, `done`, redacted `error`. |
77
+ | Usage | `prompt_tokens`/`completion_tokens`/`total_tokens`; `prompt_tokens_details.cached_tokens` → `cacheReadTokens`, `cache_creation_input_tokens` → `cacheWriteTokens`. |
78
+ | Discovery | `listAlibabaModels()` maps `GET {base}/models` entries → `ModelConfig` (reasoning/vision inferred from id). |
79
+ | Auth methods | `api_key` for `alibaba`. |
80
+
81
+ The stream parser emits `done` only on completion evidence (`[DONE]` plus a terminal
82
+ `finish_reason` with no dangling tool calls). Truncated streams terminate with an
83
+ `error` event instead. Unsupported block placements or unclaimed images fail before
84
+ fetch.
85
+
86
+ ## Request/response example
87
+
88
+ ```bash
89
+ curl 'https://dashscope-intl.aliyuncs.com/compatible-mode/v1/chat/completions' \
90
+ -H "Authorization: Bearer $DASHSCOPE_API_KEY" \
91
+ -H 'Content-Type: application/json' \
92
+ -d '{
93
+ "model": "qwen-plus",
94
+ "messages": [{ "role": "user", "content": "Hello" }],
95
+ "stream": true,
96
+ "stream_options": { "include_usage": true }
97
+ }'
98
+ ```
99
+
100
+ Usage in the final streamed chunk:
101
+
102
+ ```json
103
+ {
104
+ "usage": {
105
+ "prompt_tokens": 100,
106
+ "completion_tokens": 5,
107
+ "total_tokens": 105,
108
+ "prompt_tokens_details": { "cached_tokens": 80, "cache_creation_input_tokens": 10 }
109
+ }
110
+ }
111
+ ```
112
+
113
+ ## Implementation example
114
+
115
+ ```ts
116
+ import { createExtensionKernel } from "@arnilo/prism";
117
+ import {
118
+ createAlibabaProviderPackage,
119
+ listAlibabaModels,
120
+ } from "@arnilo/prism-provider-alibaba";
121
+
122
+ const kernel = createExtensionKernel();
123
+
124
+ // Caller-gated discovery — never runs during setup.
125
+ const models = await listAlibabaModels({ apiKey: process.env.DASHSCOPE_API_KEY });
126
+
127
+ await kernel.load([
128
+ createAlibabaProviderPackage({
129
+ apiKey: process.env.DASHSCOPE_API_KEY,
130
+ preset: "singapore", // or "coding-plan" with a Coding Plan key
131
+ models,
132
+ }),
133
+ ]);
134
+ ```
135
+
136
+ ## Extension and configuration notes
137
+
138
+ - Hosts choose base URL/preset, provider id, model list, credential source, and
139
+ `fetch` impl. Nothing is hard-coded; register discovered models via `models:`.
140
+ - Qwen thinking: `compat.enable_thinking` (request wins over model default) maps to
141
+ the top-level `enable_thinking` wire field; omitted unless explicitly boolean.
142
+ - Provider-owned compat keys (`route`, `enable_thinking`, `alibaba`) are stripped
143
+ before the opaque `compat` spread so they never leak into wire bodies.
144
+
145
+ ### Cache behavior
146
+
147
+ - **Implicit** prefix caching is automatic upstream and sends no markers.
148
+ - **Explicit** caching is opt-in: when `ModelConfig.cache.kind === "cache_control"`
149
+ (or `cache.mode === "on"`) and the caller supplies
150
+ `ProviderRequestOptions.cache.breakpoints`, `cache_control: {"type":"ephemeral"}`
151
+ markers land on the last content block of each selected message, capped at
152
+ `ALIBABA_MAX_CACHE_BREAKPOINTS` (4). Each cached prefix needs ≥1024 tokens and
153
+ lives ~5 minutes upstream.
154
+ - Usage accounting: `cached_tokens` → `Usage.cacheReadTokens`,
155
+ `cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
156
+
157
+ ## Security and performance notes
158
+
159
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
160
+ helpers (`readSseData`, `readBoundedResponseText`).
161
+ - No network calls during import, setup, build, or default tests.
162
+ - No automatic environment, file, keychain, or shell credential lookup.
163
+ - The API key is resolved per request via `resolveCredentialValue` and sent only as
164
+ `Authorization: Bearer`; keys are redacted from all thrown errors (including
165
+ discovery failures). No local filesystem paths enter request payloads.
166
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
167
+ provider-owned headers (`content-type`, `authorization`) are applied last and
168
+ cannot be overridden.
169
+ - Model discovery is caller-gated and never invoked in the provider hot path.
170
+ - Live tests stay opt-in; default tests are network-free.
171
+
172
+ ## Related APIs
173
+
174
+ - [Provider packages](../provider-packages.md): `defineProviderPackage`,
175
+ caller-gated discovery, OpenAI-compatible routes.
176
+ - [Provider caching](../provider-caching.md): explicit/implicit matrix.
177
+ - [Credentials and redaction](../credentials-and-redaction.md):
178
+ `resolveCredentialValue`, `redactSecrets`.
179
+ - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -0,0 +1,166 @@
1
+ # Ollama Cloud provider package
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-ollama` is a side-effect-free adapter for Ollama — both
6
+ **Ollama Cloud** (`https://ollama.com`) and a **local** `ollama serve`
7
+ (`http://localhost:11434`) — over the OpenAI-compatible
8
+ `POST {base}/chat/completions` endpoint.
9
+
10
+ - **Dynamic model discovery** — `listOllamaModels()` calls the OpenAI-compatible
11
+ `GET {base}/models`. No model catalog is hard-coded: available models vary by cloud
12
+ account or local pull, so discovery is the source of truth. Package setup never
13
+ fetches. (The native `GET {base}/api/tags` endpoint is an alternate catalog source;
14
+ Prism uses the OpenAI-compatible route for a uniform shape.)
15
+ - **Implicit cache only** — Ollama reuses its KV/prompt cache automatically. There is
16
+ no request knob and no cached-token count in usage, so `Usage.cacheReadTokens` is
17
+ intentionally left undefined (documented ceiling below).
18
+ - **Reasoning** — `reasoning_effort` passthrough (e.g. gpt-oss models).
19
+
20
+ Cloud auth is an ollama.com API key sent as `Authorization: Bearer`; local instances
21
+ are typically unauthenticated (omit the key).
22
+
23
+ ## When to use it
24
+
25
+ Use it when a host app wants Ollama Cloud or local Ollama models through Prism's
26
+ `AgentSession` runtime with OpenAI-compatible serialization and dynamic model
27
+ discovery.
28
+
29
+ Do not use it for automatic credential discovery, setup-time catalog fetches, explicit
30
+ cache control (Ollama has none), or real-network tests (live tests stay opt-in).
31
+
32
+ ## Inputs / request
33
+
34
+ ```ts
35
+ import {
36
+ createOllamaProviderPackage,
37
+ createOllamaProvider,
38
+ listOllamaModels,
39
+ defineOllamaModel,
40
+ ollamaBaseUrl,
41
+ } from "@arnilo/prism-provider-ollama";
42
+
43
+ createOllamaProviderPackage(options: OllamaProviderPackageOptions): ProviderPackage
44
+ createOllamaProvider(options?: OllamaProviderOptions): AIProvider
45
+ listOllamaModels(options?: ListOllamaModelsOptions): Promise<ModelConfig[]>
46
+ defineOllamaModel(config: OllamaModelConfig): ModelConfig
47
+ ollamaBaseUrl(options?: { baseUrl?: string; preset?: OllamaBasePreset }): string
48
+ ```
49
+
50
+ | Field | Type | Purpose |
51
+ | --- | --- | --- |
52
+ | `apiKey` | `CredentialValueSource` | Ollama Cloud API key; omit for unauthenticated local. |
53
+ | `baseUrl` | `string` | Explicit OpenAI-compatible base URL (wins over `preset`). |
54
+ | `preset` | `OllamaBasePreset` | `"cloud"` (default) / `"local"`. |
55
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
56
+ | `id` | `string` | Provider id (default `ollama`). |
57
+ | `models` | `readonly ModelConfig[]` | Host-supplied models (from `listOllamaModels`) to register. |
58
+
59
+ Base URLs resolved by preset (each includes the `/v1` segment):
60
+
61
+ | Preset | Base URL |
62
+ | --- | --- |
63
+ | `cloud` | `https://ollama.com/v1` |
64
+ | `local` | `http://localhost:11434/v1` |
65
+
66
+ ## Outputs / response / events
67
+
68
+ | Surface | Behavior |
69
+ | --- | --- |
70
+ | Stream | Prism text deltas, `delta.reasoning_content` → thinking deltas, tool-call delta/final, `usage`, `done`, redacted `error`. |
71
+ | Usage | `prompt_tokens` → `inputTokens`, `completion_tokens` → `outputTokens` (native `prompt_eval_count`/`eval_count` are the equivalent). `cacheReadTokens` stays undefined. |
72
+ | Discovery | `listOllamaModels()` maps `GET {base}/models` entries → `ModelConfig` (reasoning/vision inferred from id). |
73
+ | Auth methods | `api_key` for `ollama`. |
74
+
75
+ The stream parser emits `done` only on completion evidence (`[DONE]` plus a terminal
76
+ `finish_reason` with no dangling tool calls). Truncated streams terminate with an
77
+ `error` event instead. Unsupported block placements or unclaimed images fail before
78
+ fetch.
79
+
80
+ ## Request/response example
81
+
82
+ ```bash
83
+ curl 'https://ollama.com/v1/chat/completions' \
84
+ -H "Authorization: Bearer $OLLAMA_API_KEY" \
85
+ -H 'Content-Type: application/json' \
86
+ -d '{
87
+ "model": "gpt-oss:20b",
88
+ "messages": [{ "role": "user", "content": "Hello" }],
89
+ "stream": true,
90
+ "stream_options": { "include_usage": true }
91
+ }'
92
+
93
+ curl 'https://ollama.com/v1/models' -H "Authorization: Bearer $OLLAMA_API_KEY"
94
+ ```
95
+
96
+ Usage in the final streamed chunk:
97
+
98
+ ```json
99
+ { "usage": { "prompt_tokens": 100, "completion_tokens": 5, "total_tokens": 105 } }
100
+ ```
101
+
102
+ ## Implementation example
103
+
104
+ ```ts
105
+ import { createExtensionKernel } from "@arnilo/prism";
106
+ import {
107
+ createOllamaProviderPackage,
108
+ listOllamaModels,
109
+ } from "@arnilo/prism-provider-ollama";
110
+
111
+ const kernel = createExtensionKernel();
112
+
113
+ // Caller-gated discovery — never runs during setup.
114
+ const models = await listOllamaModels({ apiKey: process.env.OLLAMA_API_KEY });
115
+
116
+ await kernel.load([
117
+ createOllamaProviderPackage({
118
+ apiKey: process.env.OLLAMA_API_KEY, // omit for local
119
+ preset: "cloud", // or "local"
120
+ models,
121
+ }),
122
+ ]);
123
+ ```
124
+
125
+ ## Extension and configuration notes
126
+
127
+ - Hosts choose base URL/preset, provider id, model list, credential source, and
128
+ `fetch` impl. Nothing is hard-coded; register discovered models via `models:`.
129
+ - Reasoning: `compat.reasoning_effort` (request wins over model default) maps to the
130
+ top-level `reasoning_effort` wire field; omitted unless explicitly a string.
131
+ - Provider-owned compat keys (`route`, `reasoning_effort`, `ollama`) are stripped
132
+ before the opaque `compat` spread so they never leak into wire bodies.
133
+
134
+ ### Cache behavior
135
+
136
+ - **Implicit only.** Ollama reuses its KV/prompt cache automatically; there is no
137
+ request knob and no wire marker. Prism never emits `cache_control` for Ollama.
138
+ - **Documented ceiling:** Ollama exposes no cached-token count, so
139
+ `Usage.cacheReadTokens` is intentionally left `undefined` (not `0`). If a future
140
+ Ollama release reports cached tokens, map them in `mapOllamaModel`/usage handling.
141
+
142
+ ## Security and performance notes
143
+
144
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
145
+ helpers (`readSseData`, `readBoundedResponseText`).
146
+ - No network calls during import, setup, build, or default tests.
147
+ - No automatic environment, file, keychain, or shell credential lookup.
148
+ - The cloud API key is resolved per request via `resolveCredentialValue` and sent only
149
+ as `Authorization: Bearer`; keys are redacted from all thrown errors (including
150
+ discovery failures). Local presets send no auth header when no key is configured.
151
+ No local filesystem paths enter request payloads.
152
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
153
+ provider-owned headers (`content-type`, `authorization`) are applied last and
154
+ cannot be overridden.
155
+ - Model discovery is caller-gated and never invoked in the provider hot path.
156
+ - Live tests stay opt-in; default tests are network-free.
157
+
158
+ ## Related APIs
159
+
160
+ - [Provider packages](../provider-packages.md): `defineProviderPackage`,
161
+ caller-gated discovery, OpenAI-compatible routes.
162
+ - [Provider caching](../provider-caching.md): explicit/implicit matrix (Ollama =
163
+ implicit only).
164
+ - [Credentials and redaction](../credentials-and-redaction.md):
165
+ `resolveCredentialValue`, `redactSecrets`.
166
+ - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -50,11 +50,13 @@ uses official Responses `reasoning: { effort, summary? }` via
50
50
 
51
51
  | Surface | Behavior |
52
52
  | --- | --- |
53
- | Provider stream | Prism text, thinking (downgraded to text), `tool_call` deltas/finals, `usage`, `done`, redacted `error` events. |
54
- | Block preservation | User/system text → `input_text`; assistant text → `output_text`; assistant `tool_call` → top-level `function_call` with `call_id`; `tool_result` → top-level `function_call_output`; images/files/audio when declared on the model. Bare thinking without an encrypted Responses reasoning item is omitted on replay. |
53
+ | Provider stream | Prism text, thinking (downgraded to text), host `tool_call` deltas/finals, provider-hosted `tool_call` events (`authority: "provider-hosted"`), `continuation_required`, `usage`, `done`, and redacted `error` events. |
54
+ | Continuation | An incomplete Responses stream self-resumes at most eight HTTP hops using opaque `previous_response_id`; a cursor is at most 4 KiB, is never replayed, and is observable as `continuation_required`. |
55
+ | Realtime | `createOpenAIRealtimeSession()` exposes server-session creation, audio in/out, transcript deltas, provider-hosted calls, interrupt, and idempotent close through the neutral `RealtimeSession` seam. |
56
+ | Block preservation | User/system text → `input_text`; assistant text → `output_text`; assistant host `tool_call` → top-level `function_call` with `call_id`; provider-hosted calls are not replayed; `tool_result` → top-level `function_call_output`; images/files/audio when declared on the model. Bare thinking without an encrypted Responses reasoning item is omitted on replay. |
55
57
  | Auth methods | `api_key` for `openai`; host-invoked subscription `oauth` for `openai-codex`. This is Prism's only first-party subscription OAuth flow in 0.0.12. |
56
58
 
57
- Unsupported block placements or unclaimed images fail before `fetch`.
59
+ Unsupported block placements or unclaimed images fail before `fetch`. Provider-hosted calls are telemetry only: Prism never dispatches them as host tools or sends a `tool_result`.
58
60
 
59
61
  ## Request/response example
60
62
 
@@ -76,6 +78,21 @@ Responses request body (Codex subscription shape, abbreviated):
76
78
  }
77
79
  ```
78
80
 
81
+ Realtime session (OpenAI session creation, abbreviated):
82
+
83
+ ```ts
84
+ import { createOpenAIRealtimeSession } from "@arnilo/prism-provider-openai";
85
+
86
+ const session = createOpenAIRealtimeSession({
87
+ model: { provider: "openai", model: "gpt-realtime-2.1" },
88
+ ownerId: "hashed-host-user-id",
89
+ apiKey,
90
+ });
91
+ for await (const event of session.events()) {
92
+ if (event.type === "audio_delta") play(event.audio);
93
+ }
94
+ ```
95
+
79
96
  OAuth authorize URL (PKCE, `S256`):
80
97
 
81
98
  ```
@@ -125,6 +142,7 @@ const challenge = computeS256Challenge(verifier);
125
142
  `expires_in`, honors RFC 8628 `authorization_pending` / `slow_down`, and stops
126
143
  on terminal errors or expiry. Pass `signal` on `OAuthLoginCallbacks` to abort
127
144
  polling promptly.
145
+ - `createOpenAIRealtimeSession` requires a stable host `ownerId`; it sends that value as OpenAI's safety identifier and binds the stream to the server `session.created` id. An injected `webSocket(url, { headers })` factory supports Node 22 hosts whose global WebSocket does not expose header options.
128
146
 
129
147
  ### Cache behavior
130
148
 
@@ -189,6 +207,7 @@ Official: [Reasoning models](https://developers.openai.com/api/docs/guides/reaso
189
207
  access/refresh tokens echoed in token-endpoint failures.
190
208
  - The PKCE verifier is exchanged at the token endpoint, never sent on the authorize
191
209
  URL.
210
+ - Realtime uses the documented WebSocket `Authorization` header, never a credential query parameter. API keys are redacted from transcript/error events; audio/transcript input is untrusted, realtime queues are bounded, and disconnect, abort, malformed session identity, or audio/byte/wall-time cap breach closes the session.
192
211
  - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
193
212
  provider-specific env names; default `npm test` is network-free.
194
213