@arnilo/prism 0.0.12 → 0.0.14

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/CHANGELOG.md +23 -0
  2. package/README.md +8 -2
  3. package/dist/agents.js +21 -2
  4. package/dist/artifacts.d.ts +78 -0
  5. package/dist/artifacts.js +24 -0
  6. package/dist/contracts.d.ts +35 -1
  7. package/dist/contracts.js +8 -0
  8. package/dist/conversations.d.ts +50 -0
  9. package/dist/conversations.js +97 -0
  10. package/dist/credentials.d.ts +14 -0
  11. package/dist/credentials.js +9 -0
  12. package/dist/devices.d.ts +94 -0
  13. package/dist/devices.js +138 -0
  14. package/dist/extensions.d.ts +11 -0
  15. package/dist/extensions.js +15 -0
  16. package/dist/identity.d.ts +92 -0
  17. package/dist/identity.js +257 -0
  18. package/dist/index.d.ts +15 -5
  19. package/dist/index.js +8 -3
  20. package/dist/persistence-lifecycle.d.ts +103 -0
  21. package/dist/persistence-lifecycle.js +204 -0
  22. package/dist/providers/openai-compatible.d.ts +5 -1
  23. package/dist/providers/openai-compatible.js +15 -6
  24. package/dist/providers/openai-primitives.js +5 -2
  25. package/dist/secure-agent.js +7 -1
  26. package/dist/testing/persistence-schema.d.ts +2 -2
  27. package/dist/testing/persistence-schema.js +35 -2
  28. package/dist/tools.d.ts +2 -0
  29. package/dist/tools.js +6 -0
  30. package/docs/a2a.md +2 -0
  31. package/docs/ag-ui.md +5 -0
  32. package/docs/agent-identity.md +111 -0
  33. package/docs/browser-automation.md +3 -0
  34. package/docs/conversations.md +135 -0
  35. package/docs/credential-storage.md +31 -1
  36. package/docs/credentials-and-redaction.md +2 -0
  37. package/docs/database-persistence.md +22 -7
  38. package/docs/device-adapters.md +97 -0
  39. package/docs/extensions.md +1 -0
  40. package/docs/guardrails.md +3 -0
  41. package/docs/host-security.md +9 -3
  42. package/docs/index.md +26 -13
  43. package/docs/mcp-tools.md +2 -0
  44. package/docs/migration.md +48 -0
  45. package/docs/model-routing.md +102 -0
  46. package/docs/observability.md +2 -0
  47. package/docs/performance.md +21 -0
  48. package/docs/policy-and-audit.md +128 -0
  49. package/docs/postgres-persistence.md +1 -1
  50. package/docs/provider-caching.md +4 -0
  51. package/docs/provider-packages.md +12 -2
  52. package/docs/provider-request-policies.md +2 -0
  53. package/docs/providers/alibaba.md +179 -0
  54. package/docs/providers/azure.md +74 -0
  55. package/docs/providers/bedrock.md +72 -0
  56. package/docs/providers/google.md +1 -0
  57. package/docs/providers/ollama.md +166 -0
  58. package/docs/providers/openai-compatible.md +3 -1
  59. package/docs/providers/openrouter.md +2 -0
  60. package/docs/providers/vertex.md +71 -0
  61. package/docs/public-contracts.md +3 -1
  62. package/docs/release-and-install.md +149 -7
  63. package/docs/review-coverage-2026-07-23-phase-8.md +245 -0
  64. package/docs/review-coverage-2026-07-25-phase-9.md +256 -0
  65. package/docs/runs-and-usage.md +2 -0
  66. package/docs/server.md +37 -4
  67. package/docs/sqlite-persistence.md +1 -1
  68. package/docs/supervisors.md +2 -0
  69. package/docs/work-artifacts-and-review.md +100 -0
  70. package/docs/work-connectors.md +32 -0
  71. package/docs/work-tools.md +117 -0
  72. package/docs/workflows.md +4 -0
  73. package/docs/working-and-semantic-memory.md +20 -5
  74. package/package.json +4 -1
  75. package/templates/init/providers.json +22 -0
@@ -0,0 +1,179 @@
1
+ # Alibaba Cloud provider package
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-alibaba` is a side-effect-free adapter for Alibaba Cloud
6
+ Model Studio / DashScope (including the Coding Plan) over the **OpenAI-compatible**
7
+ `POST {base}/chat/completions` endpoint.
8
+
9
+ - **Dynamic model discovery** — `listAlibabaModels()` calls the OpenAI-compatible
10
+ `GET {base}/models`. No model catalog is hard-coded in the package: available
11
+ models vary by region, workspace, and billing plan, so discovery is the source of
12
+ truth. Package setup never fetches.
13
+ - **Context cache** — DashScope implicit prefix caching is automatic. Explicit
14
+ caching is opt-in via Anthropic-style `cache_control: {"type":"ephemeral"}`
15
+ markers (at most 4 per request). Cache hits are accounted from
16
+ `usage.prompt_tokens_details.cached_tokens` (read) and
17
+ `cache_creation_input_tokens` (write).
18
+ - **Qwen thinking** — `enable_thinking` passthrough toggles reasoning on Qwen models.
19
+
20
+ The API key is region/plan-scoped: it must match the base URL's billing plan
21
+ (pay-as-you-go regional, workspace-dedicated, or Coding Plan).
22
+
23
+ ## When to use it
24
+
25
+ Use it when a host app wants Alibaba Cloud Qwen models (Model Studio / DashScope or
26
+ the Coding Plan) through Prism's `AgentSession` runtime with OpenAI-compatible
27
+ serialization, dynamic model discovery, and explicit/implicit cache accounting.
28
+
29
+ Do not use it for automatic credential discovery, setup-time catalog fetches, or
30
+ real-network tests (live tests stay opt-in).
31
+
32
+ ## Inputs / request
33
+
34
+ ```ts
35
+ import {
36
+ createAlibabaProviderPackage,
37
+ createAlibabaProvider,
38
+ listAlibabaModels,
39
+ defineAlibabaModel,
40
+ alibabaBaseUrl,
41
+ } from "@arnilo/prism-provider-alibaba";
42
+
43
+ createAlibabaProviderPackage(options: AlibabaProviderPackageOptions): ProviderPackage
44
+ createAlibabaProvider(options?: AlibabaProviderOptions): AIProvider
45
+ listAlibabaModels(options?: ListAlibabaModelsOptions): Promise<ModelConfig[]>
46
+ defineAlibabaModel(config: AlibabaModelConfig): ModelConfig
47
+ alibabaBaseUrl(options?: { baseUrl?: string; preset?: AlibabaBasePreset }): string
48
+ ```
49
+
50
+ | Field | Type | Purpose |
51
+ | --- | --- | --- |
52
+ | `apiKey` | `CredentialValueSource` | DashScope API key (`DASHSCOPE_API_KEY`), region/plan-scoped. |
53
+ | `baseUrl` | `string` | Explicit OpenAI-compatible base URL (wins over `preset`). |
54
+ | `preset` | `AlibabaBasePreset` | `"singapore"` (default) / `"beijing"` / `"us"` / `"coding-plan"`. |
55
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
56
+ | `id` | `string` | Provider id (default `alibaba`). |
57
+ | `models` | `readonly ModelConfig[]` | Host-supplied models (from `listAlibabaModels`) to register. |
58
+
59
+ Base URLs resolved by preset:
60
+
61
+ | Preset | Base URL |
62
+ | --- | --- |
63
+ | `singapore` | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` |
64
+ | `beijing` | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
65
+ | `us` | `https://dashscope-us.aliyuncs.com/compatible-mode/v1` |
66
+ | `coding-plan` | `https://coding-intl.dashscope.aliyuncs.com/v1` |
67
+
68
+ Workspace-dedicated endpoints
69
+ (`https://{workspaceId}.{region}.maas.aliyuncs.com/compatible-mode/v1`) are supplied
70
+ verbatim via `baseUrl`.
71
+
72
+ ## Outputs / response / events
73
+
74
+ | Surface | Behavior |
75
+ | --- | --- |
76
+ | Stream | Prism text deltas, `delta.reasoning_content` → thinking deltas, tool-call delta/final, `usage`, `done`, redacted `error`. |
77
+ | Usage | `prompt_tokens`/`completion_tokens`/`total_tokens`; `prompt_tokens_details.cached_tokens` → `cacheReadTokens`, `cache_creation_input_tokens` → `cacheWriteTokens`. |
78
+ | Discovery | `listAlibabaModels()` maps `GET {base}/models` entries → `ModelConfig` (reasoning/vision inferred from id). |
79
+ | Auth methods | `api_key` for `alibaba`. |
80
+
81
+ The stream parser emits `done` only on completion evidence (`[DONE]` plus a terminal
82
+ `finish_reason` with no dangling tool calls). Truncated streams terminate with an
83
+ `error` event instead. Unsupported block placements or unclaimed images fail before
84
+ fetch.
85
+
86
+ ## Request/response example
87
+
88
+ ```bash
89
+ curl 'https://dashscope-intl.aliyuncs.com/compatible-mode/v1/chat/completions' \
90
+ -H "Authorization: Bearer $DASHSCOPE_API_KEY" \
91
+ -H 'Content-Type: application/json' \
92
+ -d '{
93
+ "model": "qwen-plus",
94
+ "messages": [{ "role": "user", "content": "Hello" }],
95
+ "stream": true,
96
+ "stream_options": { "include_usage": true }
97
+ }'
98
+ ```
99
+
100
+ Usage in the final streamed chunk:
101
+
102
+ ```json
103
+ {
104
+ "usage": {
105
+ "prompt_tokens": 100,
106
+ "completion_tokens": 5,
107
+ "total_tokens": 105,
108
+ "prompt_tokens_details": { "cached_tokens": 80, "cache_creation_input_tokens": 10 }
109
+ }
110
+ }
111
+ ```
112
+
113
+ ## Implementation example
114
+
115
+ ```ts
116
+ import { createExtensionKernel } from "@arnilo/prism";
117
+ import {
118
+ createAlibabaProviderPackage,
119
+ listAlibabaModels,
120
+ } from "@arnilo/prism-provider-alibaba";
121
+
122
+ const kernel = createExtensionKernel();
123
+
124
+ // Caller-gated discovery — never runs during setup.
125
+ const models = await listAlibabaModels({ apiKey: process.env.DASHSCOPE_API_KEY });
126
+
127
+ await kernel.load([
128
+ createAlibabaProviderPackage({
129
+ apiKey: process.env.DASHSCOPE_API_KEY,
130
+ preset: "singapore", // or "coding-plan" with a Coding Plan key
131
+ models,
132
+ }),
133
+ ]);
134
+ ```
135
+
136
+ ## Extension and configuration notes
137
+
138
+ - Hosts choose base URL/preset, provider id, model list, credential source, and
139
+ `fetch` impl. Nothing is hard-coded; register discovered models via `models:`.
140
+ - Qwen thinking: `compat.enable_thinking` (request wins over model default) maps to
141
+ the top-level `enable_thinking` wire field; omitted unless explicitly boolean.
142
+ - Provider-owned compat keys (`route`, `enable_thinking`, `alibaba`) are stripped
143
+ before the opaque `compat` spread so they never leak into wire bodies.
144
+
145
+ ### Cache behavior
146
+
147
+ - **Implicit** prefix caching is automatic upstream and sends no markers.
148
+ - **Explicit** caching is opt-in: when `ModelConfig.cache.kind === "cache_control"`
149
+ (or `cache.mode === "on"`) and the caller supplies
150
+ `ProviderRequestOptions.cache.breakpoints`, `cache_control: {"type":"ephemeral"}`
151
+ markers land on the last content block of each selected message, capped at
152
+ `ALIBABA_MAX_CACHE_BREAKPOINTS` (4). Each cached prefix needs ≥1024 tokens and
153
+ lives ~5 minutes upstream.
154
+ - Usage accounting: `cached_tokens` → `Usage.cacheReadTokens`,
155
+ `cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
156
+
157
+ ## Security and performance notes
158
+
159
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
160
+ helpers (`readSseData`, `readBoundedResponseText`).
161
+ - No network calls during import, setup, build, or default tests.
162
+ - No automatic environment, file, keychain, or shell credential lookup.
163
+ - The API key is resolved per request via `resolveCredentialValue` and sent only as
164
+ `Authorization: Bearer`; keys are redacted from all thrown errors (including
165
+ discovery failures). No local filesystem paths enter request payloads.
166
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
167
+ provider-owned headers (`content-type`, `authorization`) are applied last and
168
+ cannot be overridden.
169
+ - Model discovery is caller-gated and never invoked in the provider hot path.
170
+ - Live tests stay opt-in; default tests are network-free.
171
+
172
+ ## Related APIs
173
+
174
+ - [Provider packages](../provider-packages.md): `defineProviderPackage`,
175
+ caller-gated discovery, OpenAI-compatible routes.
176
+ - [Provider caching](../provider-caching.md): explicit/implicit matrix.
177
+ - [Credentials and redaction](../credentials-and-redaction.md):
178
+ `resolveCredentialValue`, `redactSecrets`.
179
+ - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -0,0 +1,74 @@
1
+ # Azure OpenAI / Foundry
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-azure` registers an Azure OpenAI / Foundry Chat Completions provider that uses host-supplied Entra workload identity (Bearer) or Azure resource keys (`api-key`). Deployment URLs keep the configured endpoint host (custom subdomain, private endpoint, or VNet FQDN).
6
+
7
+ ## When to use it
8
+
9
+ Use it for enterprise Azure OpenAI / Foundry deployments with Managed Identity or host token providers. Do not fold this into consumer OpenAI packages. Do not embed static keys in fixtures.
10
+
11
+ ## Inputs / request
12
+
13
+ ```ts
14
+ import { createAzureOpenAIProviderPackage } from "@arnilo/prism-provider-azure";
15
+
16
+ createAzureOpenAIProviderPackage({
17
+ endpoint: "https://my-resource.openai.azure.com",
18
+ deployment: "gpt-4o",
19
+ apiVersion: "2024-10-21",
20
+ credential: hostEntraToken, // late-bound
21
+ authStyle: "bearer",
22
+ models: [{ provider: "azure", model: "gpt-4o" }],
23
+ });
24
+ ```
25
+
26
+ | Field | Meaning |
27
+ | --- | --- |
28
+ | `endpoint` | Absolute https resource URL; host preserved |
29
+ | `deployment` | Deployment name (defaults to `model.model`) |
30
+ | `apiVersion` | Query `api-version` (default `2024-10-21`) |
31
+ | `credential` | `CredentialValueSource` — Entra token or resource key |
32
+ | `authStyle` | `bearer` (default) or `api-key` |
33
+
34
+ ## Outputs / response / events
35
+
36
+ Reuses `@arnilo/prism/providers/openai-compatible` streaming events (text, tool deltas, usage, done, redacted errors). Missing credentials fail closed before `fetch`.
37
+
38
+ ## Request/response example
39
+
40
+ ```http
41
+ POST https://my-resource.privatelink.openai.azure.com/openai/deployments/gpt-4o/chat/completions?api-version=2024-10-21
42
+ Authorization: Bearer <entra-token>
43
+ ```
44
+
45
+ ## Implementation example
46
+
47
+ ```ts
48
+ const provider = createAzureOpenAIProvider({
49
+ endpoint: process.env.AZURE_OPENAI_ENDPOINT!,
50
+ deployment: "gpt-4o",
51
+ credential: () => entra.getToken("https://cognitiveservices.azure.com/.default").then((t) => t.token),
52
+ });
53
+ ```
54
+
55
+ Opt-in live canaries: inject real `fetch` + host credential behind host CI secrets — no secrets in repo.
56
+
57
+ ## Extension and configuration notes
58
+
59
+ Register via `createExtensionKernel().load([createAzureOpenAIProviderPackage(...)])`. Pair with `@arnilo/prism-model-router` for residency allow-lists on Azure regions/endpoints.
60
+
61
+ ## Security and performance notes
62
+
63
+ - No credential prefetch at import; resolve per request.
64
+ - Endpoint host is never rewritten to public DNS.
65
+ - Errors redact credential values via shared transport helpers.
66
+ - No Azure SDK dependency.
67
+
68
+ ## Related APIs
69
+
70
+ - [OpenAI-compatible provider](openai-compatible.md)
71
+ - [Provider packages](../provider-packages.md)
72
+ - [Model routing](../model-routing.md)
73
+ - [Credential storage](../credential-storage.md)
74
+ - Package README: [`@arnilo/prism-provider-azure`](../../packages/provider-azure/README.md)
@@ -0,0 +1,72 @@
1
+ # Amazon Bedrock
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-bedrock` registers an Amazon Bedrock Runtime OpenAI-compatible Chat Completions provider. Hosts supply IAM/IRSA/assumed-role credentials; the package signs requests with SigV4 (no AWS SDK). Region and optional PrivateLink endpoint URLs are preserved.
6
+
7
+ ## When to use it
8
+
9
+ Use it for enterprise Bedrock access under workload identity. Do not embed long-lived keys in fixtures. Use model-router residency policy to deny disallowed regions.
10
+
11
+ ## Inputs / request
12
+
13
+ ```ts
14
+ import { createBedrockProviderPackage } from "@arnilo/prism-provider-bedrock";
15
+
16
+ createBedrockProviderPackage({
17
+ region: "eu-west-1",
18
+ // endpoint: "https://vpce-….bedrock-runtime.eu-west-1.vpce.amazonaws.com",
19
+ credential: () => hostAwsCredentials(),
20
+ models: [{ provider: "bedrock", model: "anthropic.claude-3-haiku-20240307-v1:0" }],
21
+ });
22
+ ```
23
+
24
+ | Field | Meaning |
25
+ | --- | --- |
26
+ | `region` | AWS region for signing + default endpoint |
27
+ | `endpoint` | Optional https PrivateLink / VPC interface base URL |
28
+ | `credential` | `{ accessKeyId, secretAccessKey, sessionToken? }` or async callback |
29
+ | `signRequest` | Optional host SigV4 override |
30
+
31
+ Default public base: `https://bedrock-runtime.{region}.amazonaws.com` → `/openai/v1/chat/completions`.
32
+
33
+ ## Outputs / response / events
34
+
35
+ OpenAI-compatible SSE mapped to Prism provider events. Missing credentials fail closed before network I/O.
36
+
37
+ ## Request/response example
38
+
39
+ ```http
40
+ POST https://bedrock-runtime.eu-west-1.amazonaws.com/openai/v1/chat/completions
41
+ Authorization: AWS4-HMAC-SHA256 Credential=…/eu-west-1/bedrock/aws4_request, …
42
+ X-Amz-Security-Token: …
43
+ ```
44
+
45
+ ## Implementation example
46
+
47
+ ```ts
48
+ const provider = createBedrockProvider({
49
+ region: "us-east-1",
50
+ credential: async () => fromNodeProviderChain()(),
51
+ });
52
+ ```
53
+
54
+ Live canaries stay opt-in behind host credentials; default tests are network-free.
55
+
56
+ ## Extension and configuration notes
57
+
58
+ Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hosts needing Converse-only models should supply a custom provider or AI SDK bridge.
59
+
60
+ ## Security and performance notes
61
+
62
+ - No AWS SDK; package-local SigV4 only for `bedrock` service.
63
+ - Private endpoint hosts are not rewritten to public DNS.
64
+ - Credential secrets are redacted from provider errors.
65
+ - No credential prefetch at import.
66
+
67
+ ## Related APIs
68
+
69
+ - [OpenAI-compatible provider](openai-compatible.md)
70
+ - [Model routing](../model-routing.md)
71
+ - [Provider packages](../provider-packages.md)
72
+ - Package README: [`@arnilo/prism-provider-bedrock`](../../packages/provider-bedrock/README.md)
@@ -82,6 +82,7 @@ api.registerProviderPackage(createGoogleProviderPackage({ apiKey: hostKey, model
82
82
 
83
83
  ## Related APIs
84
84
 
85
+ - [Google Vertex AI](vertex.md): enterprise ADC/workload-identity package (separate from this consumer API-key package).
85
86
  - [Provider packages](../provider-packages.md): package setup + discovery contract.
86
87
  - [Thinking and reasoning](../thinking-and-reasoning.md): portable thinking helpers.
87
88
  - [Provider conformance](../provider-conformance.md): network-free assertions.
@@ -0,0 +1,166 @@
1
+ # Ollama Cloud provider package
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-ollama` is a side-effect-free adapter for Ollama — both
6
+ **Ollama Cloud** (`https://ollama.com`) and a **local** `ollama serve`
7
+ (`http://localhost:11434`) — over the OpenAI-compatible
8
+ `POST {base}/chat/completions` endpoint.
9
+
10
+ - **Dynamic model discovery** — `listOllamaModels()` calls the OpenAI-compatible
11
+ `GET {base}/models`. No model catalog is hard-coded: available models vary by cloud
12
+ account or local pull, so discovery is the source of truth. Package setup never
13
+ fetches. (The native `GET {base}/api/tags` endpoint is an alternate catalog source;
14
+ Prism uses the OpenAI-compatible route for a uniform shape.)
15
+ - **Implicit cache only** — Ollama reuses its KV/prompt cache automatically. There is
16
+ no request knob and no cached-token count in usage, so `Usage.cacheReadTokens` is
17
+ intentionally left undefined (documented ceiling below).
18
+ - **Reasoning** — `reasoning_effort` passthrough (e.g. gpt-oss models).
19
+
20
+ Cloud auth is an ollama.com API key sent as `Authorization: Bearer`; local instances
21
+ are typically unauthenticated (omit the key).
22
+
23
+ ## When to use it
24
+
25
+ Use it when a host app wants Ollama Cloud or local Ollama models through Prism's
26
+ `AgentSession` runtime with OpenAI-compatible serialization and dynamic model
27
+ discovery.
28
+
29
+ Do not use it for automatic credential discovery, setup-time catalog fetches, explicit
30
+ cache control (Ollama has none), or real-network tests (live tests stay opt-in).
31
+
32
+ ## Inputs / request
33
+
34
+ ```ts
35
+ import {
36
+ createOllamaProviderPackage,
37
+ createOllamaProvider,
38
+ listOllamaModels,
39
+ defineOllamaModel,
40
+ ollamaBaseUrl,
41
+ } from "@arnilo/prism-provider-ollama";
42
+
43
+ createOllamaProviderPackage(options: OllamaProviderPackageOptions): ProviderPackage
44
+ createOllamaProvider(options?: OllamaProviderOptions): AIProvider
45
+ listOllamaModels(options?: ListOllamaModelsOptions): Promise<ModelConfig[]>
46
+ defineOllamaModel(config: OllamaModelConfig): ModelConfig
47
+ ollamaBaseUrl(options?: { baseUrl?: string; preset?: OllamaBasePreset }): string
48
+ ```
49
+
50
+ | Field | Type | Purpose |
51
+ | --- | --- | --- |
52
+ | `apiKey` | `CredentialValueSource` | Ollama Cloud API key; omit for unauthenticated local. |
53
+ | `baseUrl` | `string` | Explicit OpenAI-compatible base URL (wins over `preset`). |
54
+ | `preset` | `OllamaBasePreset` | `"cloud"` (default) / `"local"`. |
55
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
56
+ | `id` | `string` | Provider id (default `ollama`). |
57
+ | `models` | `readonly ModelConfig[]` | Host-supplied models (from `listOllamaModels`) to register. |
58
+
59
+ Base URLs resolved by preset (each includes the `/v1` segment):
60
+
61
+ | Preset | Base URL |
62
+ | --- | --- |
63
+ | `cloud` | `https://ollama.com/v1` |
64
+ | `local` | `http://localhost:11434/v1` |
65
+
66
+ ## Outputs / response / events
67
+
68
+ | Surface | Behavior |
69
+ | --- | --- |
70
+ | Stream | Prism text deltas, `delta.reasoning_content` → thinking deltas, tool-call delta/final, `usage`, `done`, redacted `error`. |
71
+ | Usage | `prompt_tokens` → `inputTokens`, `completion_tokens` → `outputTokens` (native `prompt_eval_count`/`eval_count` are the equivalent). `cacheReadTokens` stays undefined. |
72
+ | Discovery | `listOllamaModels()` maps `GET {base}/models` entries → `ModelConfig` (reasoning/vision inferred from id). |
73
+ | Auth methods | `api_key` for `ollama`. |
74
+
75
+ The stream parser emits `done` only on completion evidence (`[DONE]` plus a terminal
76
+ `finish_reason` with no dangling tool calls). Truncated streams terminate with an
77
+ `error` event instead. Unsupported block placements or unclaimed images fail before
78
+ fetch.
79
+
80
+ ## Request/response example
81
+
82
+ ```bash
83
+ curl 'https://ollama.com/v1/chat/completions' \
84
+ -H "Authorization: Bearer $OLLAMA_API_KEY" \
85
+ -H 'Content-Type: application/json' \
86
+ -d '{
87
+ "model": "gpt-oss:20b",
88
+ "messages": [{ "role": "user", "content": "Hello" }],
89
+ "stream": true,
90
+ "stream_options": { "include_usage": true }
91
+ }'
92
+
93
+ curl 'https://ollama.com/v1/models' -H "Authorization: Bearer $OLLAMA_API_KEY"
94
+ ```
95
+
96
+ Usage in the final streamed chunk:
97
+
98
+ ```json
99
+ { "usage": { "prompt_tokens": 100, "completion_tokens": 5, "total_tokens": 105 } }
100
+ ```
101
+
102
+ ## Implementation example
103
+
104
+ ```ts
105
+ import { createExtensionKernel } from "@arnilo/prism";
106
+ import {
107
+ createOllamaProviderPackage,
108
+ listOllamaModels,
109
+ } from "@arnilo/prism-provider-ollama";
110
+
111
+ const kernel = createExtensionKernel();
112
+
113
+ // Caller-gated discovery — never runs during setup.
114
+ const models = await listOllamaModels({ apiKey: process.env.OLLAMA_API_KEY });
115
+
116
+ await kernel.load([
117
+ createOllamaProviderPackage({
118
+ apiKey: process.env.OLLAMA_API_KEY, // omit for local
119
+ preset: "cloud", // or "local"
120
+ models,
121
+ }),
122
+ ]);
123
+ ```
124
+
125
+ ## Extension and configuration notes
126
+
127
+ - Hosts choose base URL/preset, provider id, model list, credential source, and
128
+ `fetch` impl. Nothing is hard-coded; register discovered models via `models:`.
129
+ - Reasoning: `compat.reasoning_effort` (request wins over model default) maps to the
130
+ top-level `reasoning_effort` wire field; omitted unless explicitly a string.
131
+ - Provider-owned compat keys (`route`, `reasoning_effort`, `ollama`) are stripped
132
+ before the opaque `compat` spread so they never leak into wire bodies.
133
+
134
+ ### Cache behavior
135
+
136
+ - **Implicit only.** Ollama reuses its KV/prompt cache automatically; there is no
137
+ request knob and no wire marker. Prism never emits `cache_control` for Ollama.
138
+ - **Documented ceiling:** Ollama exposes no cached-token count, so
139
+ `Usage.cacheReadTokens` is intentionally left `undefined` (not `0`). If a future
140
+ Ollama release reports cached tokens, map them in `mapOllamaModel`/usage handling.
141
+
142
+ ## Security and performance notes
143
+
144
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
145
+ helpers (`readSseData`, `readBoundedResponseText`).
146
+ - No network calls during import, setup, build, or default tests.
147
+ - No automatic environment, file, keychain, or shell credential lookup.
148
+ - The cloud API key is resolved per request via `resolveCredentialValue` and sent only
149
+ as `Authorization: Bearer`; keys are redacted from all thrown errors (including
150
+ discovery failures). Local presets send no auth header when no key is configured.
151
+ No local filesystem paths enter request payloads.
152
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
153
+ provider-owned headers (`content-type`, `authorization`) are applied last and
154
+ cannot be overridden.
155
+ - Model discovery is caller-gated and never invoked in the provider hot path.
156
+ - Live tests stay opt-in; default tests are network-free.
157
+
158
+ ## Related APIs
159
+
160
+ - [Provider packages](../provider-packages.md): `defineProviderPackage`,
161
+ caller-gated discovery, OpenAI-compatible routes.
162
+ - [Provider caching](../provider-caching.md): explicit/implicit matrix (Ollama =
163
+ implicit only).
164
+ - [Credentials and redaction](../credentials-and-redaction.md):
165
+ `resolveCredentialValue`, `redactSecrets`.
166
+ - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -27,9 +27,11 @@ Options:
27
27
  | Field | Type | Purpose |
28
28
  | --- | --- | --- |
29
29
  | `id` | `string` | Optional provider id. Defaults to `openai-compatible`. |
30
- | `baseUrl` | `string` | Base API URL; `/chat/completions` is appended. |
30
+ | `baseUrl` | `string` | Base API URL; `/chat/completions` is appended unless `chatCompletionsUrl` is set. |
31
31
  | `apiKey` | `CredentialValueSource` | Optional direct/callback/resolver credential source. |
32
32
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests or custom hosts. |
33
+ | `chatCompletionsUrl` | `string \| ((request) => string)` | Optional full chat-completions URL override (Azure deployment paths). |
34
+ | `authStyle` | `"bearer" \| "api-key" \| "none"` | Auth header style. Default `bearer`. |
33
35
 
34
36
  Provider requests use the standard `ProviderRequest` shape: `model`, `messages`, optional `tools`, `metadata`, and `signal`.
35
37
 
@@ -188,9 +188,11 @@ cache-read pricing exists), and seeds `compat.reasoning.effort` from
188
188
  hidden app identity.
189
189
  - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
190
190
  provider-specific env names; default tests are network-free.
191
+ - Enterprise hosts that must gate `compat.openRouterRouting` should wrap selection with `@arnilo/prism-model-router` (`allowOpenRouterRouting`); the OpenRouter adapter itself still passthroughs routing when present on the request.
191
192
 
192
193
  ## Related APIs
193
194
 
195
+ - [Model routing](../model-routing.md): optional allow-list/residency/budget/circuit facade and OpenRouter routing gate.
194
196
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
195
197
  `ModelConfig`/`compat`, cache policy, caller-gated discovery.
196
198
  - [Thinking and reasoning](../thinking-and-reasoning.md): `applyThinkingLevel`
@@ -0,0 +1,71 @@
1
+ # Google Vertex AI
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-vertex` registers a Vertex AI OpenAPI-compatible Chat Completions provider authenticated with host ADC / workload identity tokens. It is intentionally separate from `@arnilo/prism-provider-google` (consumer Gemini API keys).
6
+
7
+ ## When to use it
8
+
9
+ Use it for GCP enterprise Vertex deployments with Application Default Credentials or workload identity federation. Do not use the consumer Google package for Vertex auth semantics.
10
+
11
+ ## Inputs / request
12
+
13
+ ```ts
14
+ import { createVertexProviderPackage } from "@arnilo/prism-provider-vertex";
15
+
16
+ createVertexProviderPackage({
17
+ projectId: "my-gcp-project",
18
+ location: "europe-west1",
19
+ credential: () => hostAdcAccessToken(),
20
+ models: [{ provider: "vertex", model: "google/gemini-2.0-flash-001" }],
21
+ });
22
+ ```
23
+
24
+ | Field | Meaning |
25
+ | --- | --- |
26
+ | `projectId` | GCP project |
27
+ | `location` | Vertex location / region |
28
+ | `endpoint` | Optional full https OpenAPI base (private/custom); otherwise location-scoped default |
29
+ | `credential` | Bearer access token source (ADC / WIF) |
30
+
31
+ Default base: `https://{location}-aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/endpoints/openapi`.
32
+
33
+ ## Outputs / response / events
34
+
35
+ OpenAI-compatible SSE → Prism provider events. Missing ADC token fails closed before `fetch`.
36
+
37
+ ## Request/response example
38
+
39
+ ```http
40
+ POST https://europe-west1-aiplatform.googleapis.com/v1/projects/my-gcp-project/locations/europe-west1/endpoints/openapi/chat/completions
41
+ Authorization: Bearer <adc-token>
42
+ ```
43
+
44
+ ## Implementation example
45
+
46
+ ```ts
47
+ const provider = createVertexProvider({
48
+ projectId: "my-gcp-project",
49
+ location: "us-central1",
50
+ credential: async () => (await GoogleAuth.getAccessToken()),
51
+ });
52
+ ```
53
+
54
+ ## Extension and configuration notes
55
+
56
+ `@arnilo/prism-provider-google` remains API-key Gemini (`generativelanguage.googleapis.com`) and must not register Vertex OAuth/ADC. Load this package explicitly for Vertex.
57
+
58
+ ## Security and performance notes
59
+
60
+ - No Google Cloud SDK dependency in the package.
61
+ - Custom/private endpoint hosts are preserved.
62
+ - Tokens redacted from errors; no import-time credential prefetch.
63
+ - Pair with model-router residency allow-lists on `location`.
64
+
65
+ ## Related APIs
66
+
67
+ - [Google Gemini (consumer)](google.md)
68
+ - [OpenAI-compatible provider](openai-compatible.md)
69
+ - [Provider packages](../provider-packages.md)
70
+ - [Model routing](../model-routing.md)
71
+ - Package README: [`@arnilo/prism-provider-vertex`](../../packages/provider-vertex/README.md)
@@ -15,7 +15,8 @@ Current contract groups:
15
15
  - Extensions/middleware: `ExtensionLifecycleEventName`, `ExtensionEvent`, `Extension`, `ExtensionAPI`, `MiddlewareHookName`, `Middleware`, `MiddlewareNext`, `MiddlewareRegistry`
16
16
  - Configuration/manifests: `ConfigProvider`, `ConfigLayer`, `ConfigLoadContext`, `PrismManifest`, `ManifestContributionDeclaration`, `ManifestResourceDeclaration`, `ManifestContributionKind`
17
17
  - Stores/resources/settings/credentials/compaction/retry/cache helpers: `SessionEntry`, `SessionStore`, `StoreFactory`, `Resource`, `ResourceLoader`, `ResourceLoadContext`, `SettingsProvider`, `CredentialRequest`, `Credential`, `CredentialResolver`, `CompactionStrategy`, `CompactionContext`, `CompactionResult`, `CompactionOptions`, `CompactionMiddlewarePayload`, `CompactionEntryData`, `DefaultCompactionStrategyOptions`, `RetryPolicy`, `RetryContext`, `RetryDecision`, `RetryOptions`, `RetryMiddlewarePayload`, `DefaultRetryPolicyOptions`, `CacheUsageReport`, `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, `cacheUsageReport`
18
- - Production persistence (adapter-facing): `ProductionPersistenceStore`, `CheckpointStore`, `CheckpointKey`, `CheckpointSaveInput`, `CheckpointRecord`, `CheckpointQuery`, `LeaseStore`, `LeaseKey`, `LeaseAcquireInput`, `LeaseClaimInput`, `LeaseRecord`, `PersistencePage`, `PersistenceQuery`, `OwnershipScope`, `SessionRecord`, `SessionQuery`, `BranchRecord`, `BranchQuery`, `SessionEntryQuery`, `RunRecord`, `RunQuery`, `RunFeedbackRecord`, `RunFeedbackStore`, `RunFeedbackQuery`, `AgentEventRecord`, `AgentEventQuery`, `ToolCallRecord`, `ToolCallQuery`, `UsageRecord`, `UsageQuery`, `AgentDefinitionRecord`, `AgentDefinitionQuery`, `RetentionPolicy`, `RetentionPolicyQuery`, `MigrationRecord`, `MigrationQuery`
18
+ - Production persistence (adapter-facing): `ProductionPersistenceStore` (optional `lifecycle`), `CheckpointStore`, `CheckpointKey`, `CheckpointSaveInput`, `CheckpointRecord`, `CheckpointQuery`, `LeaseStore`, `LeaseKey`, `LeaseAcquireInput`, `LeaseClaimInput`, `LeaseRecord`, `PersistencePage`, `PersistenceQuery`, `OwnershipScope`, `SessionRecord`, `SessionQuery`, `BranchRecord`, `BranchQuery`, `SessionEntryQuery`, `RunRecord`, `RunQuery`, `RunFeedbackRecord`, `RunFeedbackStore`, `RunFeedbackQuery`, `AgentEventRecord`, `AgentEventQuery`, `ToolCallRecord`, `ToolCallQuery`, `UsageRecord`, `UsageQuery`, `AgentDefinitionRecord`, `AgentDefinitionQuery`, `RetentionPolicy`, `RetentionPolicyQuery`, `MigrationRecord`, `MigrationQuery`, `PersistenceLifecycleStore`, `LegalHoldRecord`, `TenantQuota`
19
+ - Identity: `Principal`, `AgentIdentity`, `IdentityVerifier`, `assertIdentityActive`, `narrowIdentity`, `ownershipFromIdentity`, `assertIdentityMatchesOwnership`, `assertIdentityPropagation`, `identityTelemetryAttributes`, `resolveRunIdentity`, `IdentityError`, identity limit constants
19
20
 
20
21
  ## When to use it
21
22
 
@@ -165,6 +166,7 @@ Important request shapes:
165
166
  | `CacheUsageReport` | Numeric cache diagnostics from normalized `Usage`: read/write tokens, hit rate, estimated savings, and optional currency. |
166
167
  | `AgentDefinitionRecord` / `AgentDefinitionQuery` | Versioned agent-definition snapshot and filters. Does not store credentials or provider instances. |
167
168
  | `RetentionPolicy` / `RetentionPolicyQuery` | Retention policy and filters: age, entry count, byte limits, archive store, applied kinds. |
169
+ | `PersistenceLifecycleStore` / `LegalHoldRecord` / `TenantQuota` | Optional hold/retention apply/export/quota capability on `ProductionPersistenceStore.lifecycle`. |
168
170
  | `MigrationRecord` / `MigrationQuery` | Applied migration record and filters. |
169
171
 
170
172
  ## Outputs / response / events