@arnilo/prism 0.0.1 → 0.0.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +19 -2
- package/README.md +17 -7
- package/dist/agent-definitions.d.ts +12 -0
- package/dist/agent-definitions.js +131 -0
- package/dist/agent-loops.d.ts +14 -0
- package/dist/agent-loops.js +161 -0
- package/dist/agents.js +263 -76
- package/dist/cache-helpers.d.ts +28 -0
- package/dist/cache-helpers.js +73 -0
- package/dist/cli-runner.d.ts +38 -2
- package/dist/cli-runner.js +167 -5
- package/dist/compaction.js +2 -0
- package/dist/config.js +47 -12
- package/dist/contracts.d.ts +581 -6
- package/dist/contracts.js +41 -1
- package/dist/contribution-parsing.d.ts +19 -0
- package/dist/contribution-parsing.js +124 -0
- package/dist/contributions.d.ts +13 -3
- package/dist/contributions.js +96 -20
- package/dist/extensions.js +3 -0
- package/dist/index.d.ts +19 -9
- package/dist/index.js +10 -4
- package/dist/input.d.ts +7 -1
- package/dist/input.js +52 -11
- package/dist/instruction-injection.d.ts +28 -0
- package/dist/instruction-injection.js +55 -0
- package/dist/manifests.d.ts +1 -1
- package/dist/manifests.js +3 -3
- package/dist/models.d.ts +4 -1
- package/dist/models.js +5 -2
- package/dist/node/agent-definitions.d.ts +98 -0
- package/dist/node/agent-definitions.js +389 -0
- package/dist/node/contribution-discovery.d.ts +17 -0
- package/dist/node/contribution-discovery.js +163 -0
- package/dist/node/instruction-injectors.d.ts +32 -0
- package/dist/node/instruction-injectors.js +72 -0
- package/dist/node/session-store-jsonl.d.ts +1 -1
- package/dist/node/session-store-jsonl.js +42 -4
- package/dist/node/system-project-prompts.d.ts +30 -0
- package/dist/node/system-project-prompts.js +53 -0
- package/dist/provider-events.d.ts +3 -1
- package/dist/provider-events.js +34 -0
- package/dist/provider-request-policy.js +15 -1
- package/dist/providers/openai-compatible.js +1 -1
- package/dist/providers.d.ts +6 -2
- package/dist/providers.js +15 -1
- package/dist/redaction.d.ts +2 -1
- package/dist/redaction.js +3 -0
- package/dist/registry-options.d.ts +5 -0
- package/dist/registry-options.js +5 -0
- package/dist/rpc.d.ts +6 -2
- package/dist/rpc.js +71 -13
- package/dist/session-stores.d.ts +3 -1
- package/dist/session-stores.js +67 -6
- package/dist/skills.d.ts +4 -1
- package/dist/skills.js +3 -1
- package/dist/system-prompts.js +6 -2
- package/dist/testing/compaction-conformance.d.ts +17 -0
- package/dist/testing/compaction-conformance.js +61 -0
- package/dist/testing/extension-conformance.d.ts +26 -0
- package/dist/testing/extension-conformance.js +55 -0
- package/dist/testing/provider-conformance.d.ts +7 -0
- package/dist/testing/provider-conformance.js +18 -31
- package/dist/testing/session-store-conformance.d.ts +20 -0
- package/dist/testing/session-store-conformance.js +92 -0
- package/dist/testing/tool-conformance.d.ts +39 -0
- package/dist/testing/tool-conformance.js +79 -0
- package/dist/tools.d.ts +7 -2
- package/dist/tools.js +50 -13
- package/docs/agent-definitions.md +251 -0
- package/docs/agent-events.md +199 -0
- package/docs/agent-loops.md +217 -0
- package/docs/agent-session-runtime.md +20 -8
- package/docs/cli-rpc.md +39 -4
- package/docs/coding-agent-tools.md +208 -0
- package/docs/compaction-and-retry.md +2 -2
- package/docs/compaction-conformance.md +76 -0
- package/docs/compaction-llm.md +6 -3
- package/docs/compaction-observational-memory.md +4 -4
- package/docs/configuration-and-manifests.md +6 -1
- package/docs/context-and-skills.md +79 -6
- package/docs/contribution-discovery.md +149 -0
- package/docs/contribution-registries.md +9 -6
- package/docs/credentials-and-redaction.md +2 -0
- package/docs/customization.md +191 -0
- package/docs/database-persistence.md +407 -0
- package/docs/extension-authoring.md +193 -0
- package/docs/extension-conformance.md +80 -0
- package/docs/extensions.md +6 -0
- package/docs/host-security.md +141 -0
- package/docs/index.md +41 -19
- package/docs/input-and-prompt-assembly.md +19 -3
- package/docs/instruction-injection.md +183 -0
- package/docs/migration.md +201 -0
- package/docs/model-registry.md +122 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/performance.md +127 -0
- package/docs/provider-caching.md +206 -0
- package/docs/provider-conformance.md +32 -5
- package/docs/provider-layer.md +51 -11
- package/docs/provider-packages.md +65 -5
- package/docs/provider-request-policies.md +113 -0
- package/docs/providers/kimi.md +22 -0
- package/docs/providers/neuralwatt.md +388 -0
- package/docs/providers/openai-compatible.md +1 -0
- package/docs/providers/openai.md +21 -0
- package/docs/providers/opencode-go.md +31 -3
- package/docs/providers/openrouter.md +29 -0
- package/docs/providers/zai.md +17 -0
- package/docs/public-contracts.md +87 -12
- package/docs/release-and-install.md +79 -27
- package/docs/runs-and-usage.md +236 -0
- package/docs/session-store-conformance.md +78 -0
- package/docs/session-stores-and-branching.md +10 -6
- package/docs/session-stores.md +126 -0
- package/docs/settings-auth-trust-security.md +18 -4
- package/docs/structured-output.md +247 -0
- package/docs/system-prompts.md +104 -2
- package/docs/tool-conformance.md +87 -0
- package/docs/tools.md +65 -8
- package/package.json +36 -2
|
@@ -0,0 +1,206 @@
|
|
|
1
|
+
# Provider caching
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Provider caching documents Prism's cache intent surface:
|
|
6
|
+
|
|
7
|
+
- `ProviderRequestOptions.cache?: PromptCacheHints` for structured, provider-agnostic cache hints.
|
|
8
|
+
- Legacy aliases `cacheKey` and `cacheRetention`, still supported for backwards compatibility.
|
|
9
|
+
- `PromptCacheBreakpoint` locations for reusable prompt regions.
|
|
10
|
+
- `ModelCacheCapabilities` for model/provider cache support metadata.
|
|
11
|
+
- Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
|
|
12
|
+
|
|
13
|
+
Cache hints are best-effort. They describe intent; providers decide whether their native API can use them. Prism does not guarantee cache hits.
|
|
14
|
+
|
|
15
|
+
## When to use it
|
|
16
|
+
|
|
17
|
+
Use this page when a host or provider package needs to:
|
|
18
|
+
|
|
19
|
+
- Mark stable system prompts, tools, context, or messages as cacheable.
|
|
20
|
+
- Opt into cache-aware default input ordering so stable attachments/resources, summaries, and prior history form a reusable prefix before the current user turn.
|
|
21
|
+
- Carry a stable cache key across turns without putting provider-specific fields in core.
|
|
22
|
+
- Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
|
|
23
|
+
- Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
|
|
24
|
+
|
|
25
|
+
Do not use cache keys for credentials, bearer tokens, API keys, OAuth tokens, user secrets, or raw private prompts.
|
|
26
|
+
|
|
27
|
+
## Inputs / request
|
|
28
|
+
|
|
29
|
+
```ts
|
|
30
|
+
import type {
|
|
31
|
+
ModelCacheCapabilities,
|
|
32
|
+
PromptCacheBreakpoint,
|
|
33
|
+
PromptCacheHints,
|
|
34
|
+
ProviderRequestOptions,
|
|
35
|
+
} from "@arnilo/prism";
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
| Type / field | Purpose |
|
|
39
|
+
| --- | --- |
|
|
40
|
+
| `PromptCacheHints.mode?: "auto" | "on" | "off"` | Host intent. Providers may ignore unsupported modes. |
|
|
41
|
+
| `PromptCacheHints.key?: string` | Stable, untrusted cache key. Sanitize before sending to provider APIs. |
|
|
42
|
+
| `PromptCacheHints.retention?: "none" | "short" | "long"` | Desired retention. `mapCacheRetention()` downgrades unsupported long retention. |
|
|
43
|
+
| `PromptCacheHints.breakpoints?: readonly PromptCacheBreakpoint[]` | Stable prompt locations to mark for cache-control style providers. |
|
|
44
|
+
| `PromptCacheBreakpoint.location` | `system_prompt`, `tools`, `stable_context`, `last_stable_message`, `last_user_message`, or `message_id`. |
|
|
45
|
+
| `PromptCacheBreakpoint.messageId?` | Required when `location: "message_id"`. |
|
|
46
|
+
| `PromptCacheBreakpoint.ttl?` | Generic `short` / `long` hint. Provider packages map to native TTL shape. |
|
|
47
|
+
| `ModelConfig.cache?: ModelCacheCapabilities` | Static model/provider cache support metadata. |
|
|
48
|
+
|
|
49
|
+
`ModelCacheCapabilities.kind` values are generic: `implicit`, `openai_key`, `cache_control`, `provider_specific`, or `none`. Core never branches on provider names; provider packages read the metadata and map it to native requests.
|
|
50
|
+
|
|
51
|
+
Legacy alias note: `cacheKey` maps to `cache.key`, and `cacheRetention` maps to `cache.retention`. When both are present, structured `cache.key` / `cache.retention` is the authoritative cache intent for providers that read structured hints; legacy fields remain for older adapters.
|
|
52
|
+
|
|
53
|
+
## Outputs / response / events
|
|
54
|
+
|
|
55
|
+
Cache helpers return plain data:
|
|
56
|
+
|
|
57
|
+
| Helper | Output |
|
|
58
|
+
| --- | --- |
|
|
59
|
+
| `sanitizeCacheKey(value, maxLength)` | Safe key string or `undefined`. |
|
|
60
|
+
| `mapCacheRetention(retention, model)` | `"short"`, `"long"`, or `undefined`. |
|
|
61
|
+
| `applyCacheControl(messages, breakpoints, options)` | New message array with `cache_control: { type: "ephemeral" }` on selected message anchors. |
|
|
62
|
+
| `cacheHitRate(usage)` | Cached input ratio or `undefined`. |
|
|
63
|
+
| `cacheSavings(usage, model)` | Estimated read-token savings or `undefined` without pricing. |
|
|
64
|
+
| `cacheUsageReport(usage, model?)` | Normalized read/write tokens, hit rate, estimated savings, and currency when available; `undefined` when no usage is supplied. |
|
|
65
|
+
|
|
66
|
+
Provider events do not change. Cache accounting stays in normalized `Usage.cacheReadTokens` and `Usage.cacheWriteTokens`.
|
|
67
|
+
|
|
68
|
+
For stable-prefix payloads, set `inputLayout: "cache_aware"` on the default input builder, `assembleProviderInput()`, `AgentConfig`, or `RunOptions`. The default prompt builder already places context, selected skills, and tool declarations before input messages; cache-aware input ordering then places attachments/resources, summaries, prior history, and pending tool results before the current user suffix. The prefix is byte-stable only when those stable inputs are unchanged; Prism still does not guarantee provider cache hits.
|
|
69
|
+
|
|
70
|
+
## Request/response example
|
|
71
|
+
|
|
72
|
+
```json
|
|
73
|
+
{
|
|
74
|
+
"providerRequest.options": {
|
|
75
|
+
"sessionId": "sess_123",
|
|
76
|
+
"cache": {
|
|
77
|
+
"mode": "on",
|
|
78
|
+
"key": "sess_123",
|
|
79
|
+
"retention": "long",
|
|
80
|
+
"breakpoints": [
|
|
81
|
+
{ "location": "system_prompt" },
|
|
82
|
+
{ "location": "last_user_message" }
|
|
83
|
+
]
|
|
84
|
+
}
|
|
85
|
+
},
|
|
86
|
+
"model.cache": {
|
|
87
|
+
"kind": "cache_control",
|
|
88
|
+
"maxBreakpoints": 4,
|
|
89
|
+
"minCacheableTokens": 1024,
|
|
90
|
+
"longRetention": true
|
|
91
|
+
}
|
|
92
|
+
}
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
## Implementation example
|
|
96
|
+
|
|
97
|
+
```ts
|
|
98
|
+
import {
|
|
99
|
+
applyCacheControl,
|
|
100
|
+
cacheHitRate,
|
|
101
|
+
cacheUsageReport,
|
|
102
|
+
mapCacheRetention,
|
|
103
|
+
sanitizeCacheKey,
|
|
104
|
+
type ModelConfig,
|
|
105
|
+
type PromptCacheHints,
|
|
106
|
+
} from "@arnilo/prism";
|
|
107
|
+
|
|
108
|
+
const model: ModelConfig = {
|
|
109
|
+
provider: "demo",
|
|
110
|
+
model: "demo-large",
|
|
111
|
+
cache: { kind: "cache_control", maxBreakpoints: 4, longRetention: true },
|
|
112
|
+
};
|
|
113
|
+
|
|
114
|
+
const hints: PromptCacheHints = {
|
|
115
|
+
mode: "on",
|
|
116
|
+
key: "workspace:agent#1",
|
|
117
|
+
retention: "long",
|
|
118
|
+
breakpoints: [{ location: "system_prompt" }, { location: "last_user_message" }],
|
|
119
|
+
};
|
|
120
|
+
|
|
121
|
+
const key = sanitizeCacheKey(hints.key, model.cache?.maxKeyLength ?? 128);
|
|
122
|
+
const retention = mapCacheRetention(hints.retention, model);
|
|
123
|
+
const stamped = applyCacheControl(messages, hints.breakpoints ?? [], { maxBreakpoints: model.cache?.maxBreakpoints });
|
|
124
|
+
const hitRate = cacheHitRate({ inputTokens: 1000, cacheReadTokens: 800 });
|
|
125
|
+
const report = cacheUsageReport({ inputTokens: 1000, cacheReadTokens: 800 }, model);
|
|
126
|
+
// { cacheReadTokens: 800, cacheWriteTokens: 0, hitRate: 0.8, ... }
|
|
127
|
+
|
|
128
|
+
await session.run("Explain this", { inputLayout: "cache_aware" });
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
## Extension and configuration notes
|
|
132
|
+
|
|
133
|
+
Provider request policies can set `ProviderRequestOptions.cache` or the legacy `cacheKey` / `cacheRetention` aliases. Provider packages decide how to map hints to native payloads:
|
|
134
|
+
|
|
135
|
+
| `ModelCacheCapabilities.kind` | Typical mapping |
|
|
136
|
+
| --- | --- |
|
|
137
|
+
| `implicit` | No request mutation; provider caches automatically. |
|
|
138
|
+
| `openai_key` | Send sanitized cache key and mapped retention where supported. |
|
|
139
|
+
| `cache_control` | Use `applyCacheControl()` on provider-native message anchors. |
|
|
140
|
+
| `provider_specific` | Provider package uses `compat`/native options intentionally. |
|
|
141
|
+
| `none` | Do not send cache fields. |
|
|
142
|
+
|
|
143
|
+
### Per-provider cache behavior
|
|
144
|
+
|
|
145
|
+
| Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
|
|
146
|
+
| --- | --- | --- | --- | --- |
|
|
147
|
+
| `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
|
|
148
|
+
| `@arnilo/prism-provider-openrouter` | `cache_control` | Applies `cache_control` markers only to caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. | Breakpoint-stable prefixes can be reused by upstream providers. | Best-effort only; no marker is added to every block. |
|
|
149
|
+
| `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
|
|
150
|
+
| `@arnilo/prism-provider-zai` | `implicit` | No explicit cache payload; GLM context caching is automatic. | Resend unchanged prior history for implicit context-cache reuse. | Best-effort only; cache options do not force hits. |
|
|
151
|
+
| `@arnilo/prism-provider-kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
|
|
152
|
+
| `@arnilo/prism-provider-neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
|
|
153
|
+
|
|
154
|
+
Detailed first-party provider notes:
|
|
155
|
+
|
|
156
|
+
- OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
157
|
+
- OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
|
|
158
|
+
- OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars; applies Anthropic-style `cache_control` markers only to caller-selected `cache.breakpoints` (last content block of each selected message), not every block; `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
159
|
+
- OpenCode Go (`@arnilo/prism-provider-opencode-go`): `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route sends none. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`.
|
|
160
|
+
- Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
161
|
+
- NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
|
|
162
|
+
- Kimi (`@arnilo/prism-provider-kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
163
|
+
|
|
164
|
+
### NeuralWatt cache-aware limiter
|
|
165
|
+
|
|
166
|
+
NeuralWatt (`@arnilo/prism-provider-neuralwatt`) runs a cache-aware backend rate
|
|
167
|
+
limiter on top of its implicit vLLM prefix cache. This shapes long-running agent
|
|
168
|
+
sessions differently from one-shot chat:
|
|
169
|
+
|
|
170
|
+
- **Uncached TPM counts cold prefill only.** The tokens-per-minute budget charges the
|
|
171
|
+
prefix that is not already cached. A request whose prefix is fully cached consumes
|
|
172
|
+
far less TPM than a cold request of the same total prompt length.
|
|
173
|
+
- **Warm-prefix requests can avoid some `503` fleet-capacity blocks.** Near fleet
|
|
174
|
+
capacity, requests that reuse a cached prefix are more likely to be admitted than
|
|
175
|
+
fully cold requests. Prefix reuse is both an availability and a latency lever.
|
|
176
|
+
- **Full prior history is required for multi-turn cache reuse.** The prefix cache is
|
|
177
|
+
keyed by request content, so each follow-up turn must resend the entire prior
|
|
178
|
+
transcript (system prompt + all prior turns) unchanged, with only the new turn
|
|
179
|
+
appended. Use `inputLayout: "cache_aware"` so Prism keeps the stable prefix first.
|
|
180
|
+
- Cache behavior is best-effort and **does not guarantee cache hits**. Admission and
|
|
181
|
+
eviction are server-side decisions and vary with fleet load. `cacheRetention:
|
|
182
|
+
"none"` disables Prism cache-control hints only; it does not disable the implicit
|
|
183
|
+
backend prefix cache.
|
|
184
|
+
|
|
185
|
+
See [NeuralWatt provider](providers/neuralwatt.md) for the package-level cache,
|
|
186
|
+
usage, and retry details.
|
|
187
|
+
|
|
188
|
+
## Security and performance notes
|
|
189
|
+
|
|
190
|
+
- Cache hints are best-effort and do not guarantee cache hits.
|
|
191
|
+
- Cache keys are untrusted input; sanitize and truncate with `sanitizeCacheKey()` before provider I/O.
|
|
192
|
+
- Cache keys must never be credentials or secrets.
|
|
193
|
+
- Provider-owned auth/session/security headers always win over caller headers.
|
|
194
|
+
- Helpers are pure, network-free, and O(messages) at most. `cacheUsageReport()` is O(1).
|
|
195
|
+
- Cache-aware input ordering does not change resource loading: URI attachments/resources still load only through the caller-provided `ResourceLoader`.
|
|
196
|
+
- Cache usage reports contain only usage counts and optional pricing/currency; they do not include prompt text, cache keys, headers, credentials, or provider payloads.
|
|
197
|
+
- `applyCacheControl()` returns new message objects for stamped anchors and does not mutate input messages.
|
|
198
|
+
|
|
199
|
+
## Related APIs
|
|
200
|
+
|
|
201
|
+
- [Input and prompt assembly](input-and-prompt-assembly.md): opt-in cache-aware ordering for stable provider payload prefixes.
|
|
202
|
+
- [Provider request policies](provider-request-policies.md): set cache hints before provider calls.
|
|
203
|
+
- [Model registry](model-registry.md): register `ModelConfig.cache` capability metadata.
|
|
204
|
+
- [Provider layer](provider-layer.md): provider/model registries and provider events.
|
|
205
|
+
- [Provider packages](provider-packages.md): package-owned mapping to provider-native cache APIs.
|
|
206
|
+
- [Public contracts](public-contracts.md): public type list for cache contracts and helpers.
|
|
@@ -12,14 +12,17 @@ Exported from `@arnilo/prism/testing/provider-conformance`:
|
|
|
12
12
|
- `assertToolCallDeltasReconstruct(events, expected)`
|
|
13
13
|
- `assertUsageAccounting(events, expected)`
|
|
14
14
|
- `assertSerializedRequestCoversContent(request, body, options?)`
|
|
15
|
+
- `assertProviderOwnedHeadersWin(captured, options)`
|
|
15
16
|
- `assertNoSecretLeak(events, secrets)`
|
|
16
17
|
|
|
17
18
|
## When to use it
|
|
18
19
|
|
|
19
|
-
Use these helpers in provider package tests to check event order, terminal events, abort propagation
|
|
20
|
+
Use these helpers in provider package tests to check event order, terminal events, abort propagation via `ProviderRequest.signal`, streamed tool-call deltas, usage/cache accounting, request body content preservation, protected header ownership, and secret redaction. Do not treat deprecated `ProviderRequestOptions.timeoutMs`/`maxRetries`/`maxRetryDelayMs` as conformance requirements; first-party providers use runtime abort signals and `AgentConfig.retry`/`RunOptions.retry` instead.
|
|
20
21
|
|
|
21
22
|
Do not use them as a live integration runner, provider simulator, retry framework, credential loader, or test framework replacement.
|
|
22
23
|
|
|
24
|
+
For real network smoke tests, each first-party provider package ships an env-gated `src/__tests__/live.test.ts` that exercises the live API when `PRISM_LIVE_PROVIDER_TESTS=1` and a provider-specific API key are set. These live tests reuse the same conformance helpers (`assertProviderStreamConforms`, `assertAbortIsObserved`, `assertNoSecretLeak`) against the real provider, so offline and live assertions stay consistent. The default `npm test` never sets these gates and stays network-free; see [Release and install](release-and-install.md) for the full env-var list.
|
|
25
|
+
|
|
23
26
|
## Inputs / request
|
|
24
27
|
|
|
25
28
|
```ts
|
|
@@ -41,10 +44,11 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
|
|
|
41
44
|
|
|
42
45
|
- `collectProviderEvents()` returns provider events in stream order.
|
|
43
46
|
- `assertProviderStreamConforms()` returns collected events after verifying the stream ends with `done` or `error`, terminal events are last, and optional text/usage expectations match.
|
|
44
|
-
- `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject.
|
|
45
|
-
- `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments.
|
|
46
|
-
- `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`.
|
|
47
|
-
- `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped.
|
|
47
|
+
- `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject. This is the supported timeout primitive; use a host abort controller or `RunOptions.signal` rather than deprecated provider-level `timeoutMs`.
|
|
48
|
+
- `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. The runtime uses the same reconstruction behavior before tool execution when a provider streams deltas.
|
|
49
|
+
- `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`. This is the provider-neutral check for normalized cache read/write token extraction; every first-party provider package exercises it against server-specific fields (`cached_tokens`, `cache_read_input_tokens`, etc.).
|
|
50
|
+
- `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
|
|
51
|
+
- `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.
|
|
48
52
|
- `assertNoSecretLeak()` stringifies all collected events and fails if any known secret string is present.
|
|
49
53
|
|
|
50
54
|
## Request/response example
|
|
@@ -56,6 +60,18 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
|
|
|
56
60
|
}
|
|
57
61
|
```
|
|
58
62
|
|
|
63
|
+
Tool-call delta reconstruction example:
|
|
64
|
+
|
|
65
|
+
```ts
|
|
66
|
+
import { assertToolCallDeltasReconstruct } from "@arnilo/prism/testing/provider-conformance";
|
|
67
|
+
|
|
68
|
+
assertToolCallDeltasReconstruct([
|
|
69
|
+
{ type: "tool_call_delta", index: 0, id: "call_1", name: "lookup", argumentsText: "{\"q\":" },
|
|
70
|
+
{ type: "tool_call_delta", index: 0, argumentsText: "\"prism\"}" },
|
|
71
|
+
{ type: "done" },
|
|
72
|
+
], [{ index: 0, id: "call_1", name: "lookup", arguments: { q: "prism" } }]);
|
|
73
|
+
```
|
|
74
|
+
|
|
59
75
|
Content-preservation example:
|
|
60
76
|
|
|
61
77
|
```ts
|
|
@@ -77,6 +93,17 @@ const body = JSON.parse(String(fetchInit.body));
|
|
|
77
93
|
assertSerializedRequestCoversContent(request, body, { unsupported: ["image"] });
|
|
78
94
|
```
|
|
79
95
|
|
|
96
|
+
Protected header-ownership example:
|
|
97
|
+
|
|
98
|
+
```ts
|
|
99
|
+
import { assertProviderOwnedHeadersWin } from "@arnilo/prism/testing/provider-conformance";
|
|
100
|
+
|
|
101
|
+
assertProviderOwnedHeadersWin(capturedHeaders, {
|
|
102
|
+
owned: { authorization: "Bearer provider-key", "content-type": "application/json" },
|
|
103
|
+
caller: { authorization: "Bearer attacker", "content-type": "text/plain", "x-caller": "kept" },
|
|
104
|
+
});
|
|
105
|
+
```
|
|
106
|
+
|
|
80
107
|
## Implementation example
|
|
81
108
|
|
|
82
109
|
```ts
|
package/docs/provider-layer.md
CHANGED
|
@@ -5,9 +5,10 @@
|
|
|
5
5
|
The provider layer contains the small runtime pieces Prism already ships for host-owned model access:
|
|
6
6
|
|
|
7
7
|
- `createProviderRegistry()` / `ProviderRegistry`: register and resolve `AIProvider` instances by id.
|
|
8
|
-
- `
|
|
9
|
-
- `
|
|
10
|
-
-
|
|
8
|
+
- `createProviderResolver()` / `ProviderResolver`: build a resolver that maps a `ModelConfig` to an `AIProvider` (or `undefined`), from a `ProviderRegistry` or a plain `AIProvider[]`.
|
|
9
|
+
- `createModelRegistry()` / `ModelRegistry`: register and resolve `ModelConfig` values by provider/model key. See [Model registry](model-registry.md).
|
|
10
|
+
- `ModelConfig` metadata fields for display names, capabilities, limits, cost/cache pricing, cache support metadata, opaque provider compat data, and host metadata. See [Provider caching](provider-caching.md).
|
|
11
|
+
- Provider event helpers: create normalized `ProviderEvent` values for text, thinking, streamed tool-call deltas, final tool calls, usage, done, and errors, including optional cache read/write usage fields.
|
|
11
12
|
- `toolCallContent()`: create a `ToolCallContent` block.
|
|
12
13
|
- `createMockProvider()` / `MockProviderOptions`: create a deterministic scripted `AIProvider` for tests and examples.
|
|
13
14
|
- `@arnilo/prism/testing/provider-conformance`: optional network-free assertion helpers for provider adapter tests.
|
|
@@ -23,21 +24,21 @@ Use this layer when a host app, extension package, or test needs to:
|
|
|
23
24
|
- Emit provider events without hand-writing event objects.
|
|
24
25
|
- Test agent/provider flows without timers, credentials, SDKs, or network calls.
|
|
25
26
|
|
|
26
|
-
Do not use this layer for credential storage, settings loading, tool dispatch, agent loops, package discovery, cache stores, or provider SDK configuration. Those stay host-owned or belong to provider packages.
|
|
27
|
+
Do not use this layer for credential storage, settings loading, tool dispatch, agent loops, package discovery, cache stores, or provider SDK configuration. Those stay host-owned or belong to provider packages. Request option hooks are covered in [Provider request policies](provider-request-policies.md).
|
|
27
28
|
|
|
28
29
|
## Inputs / request
|
|
29
30
|
|
|
30
31
|
### Provider registry
|
|
31
32
|
|
|
32
33
|
```ts
|
|
33
|
-
createProviderRegistry(providers?: readonly AIProvider[]): ProviderRegistry
|
|
34
|
+
createProviderRegistry(providers?: readonly AIProvider[], options?: { duplicate?: "replace" | "error" }): ProviderRegistry
|
|
34
35
|
```
|
|
35
36
|
|
|
36
37
|
`ProviderRegistry` methods:
|
|
37
38
|
|
|
38
39
|
| Method | Input | Result |
|
|
39
40
|
| --- | --- | --- |
|
|
40
|
-
| `register(provider)` | `AIProvider` | Stores provider by `provider.id`. |
|
|
41
|
+
| `register(provider)` | `AIProvider` | Stores/replaces provider by `provider.id`; throws `Duplicate provider: <id>` when `duplicate: "error"`. |
|
|
41
42
|
| `get(id)` | provider id string | Returns provider or `undefined`. |
|
|
42
43
|
| `resolve(model)` | provider id string or `{ provider: string }` | Returns provider or throws `Unknown provider: <id>`. |
|
|
43
44
|
| `list()` | none | Returns registered providers in insertion order. |
|
|
@@ -45,14 +46,14 @@ createProviderRegistry(providers?: readonly AIProvider[]): ProviderRegistry
|
|
|
45
46
|
### Model registry
|
|
46
47
|
|
|
47
48
|
```ts
|
|
48
|
-
createModelRegistry(models?: readonly ModelConfig[]): ModelRegistry
|
|
49
|
+
createModelRegistry(models?: readonly ModelConfig[], options?: { duplicate?: "replace" | "error" }): ModelRegistry
|
|
49
50
|
```
|
|
50
51
|
|
|
51
52
|
`ModelRegistry` methods:
|
|
52
53
|
|
|
53
54
|
| Method | Input | Result |
|
|
54
55
|
| --- | --- | --- |
|
|
55
|
-
| `register(model)` | `ModelConfig` | Stores model by provider/model key, preserving inert metadata
|
|
56
|
+
| `register(model)` | `ModelConfig` | Stores/replaces model by provider/model key, preserving inert metadata; throws `Duplicate model: <provider>/<model>` when `duplicate: "error"`. |
|
|
56
57
|
| `get(provider, model)` | provider id and model id | Returns model config or `undefined`. |
|
|
57
58
|
| `resolve(provider, model)` | provider id and model id | Returns model config or throws `Unknown model: <provider>/<model>`. |
|
|
58
59
|
| `list()` | none | Returns registered model configs in insertion order. |
|
|
@@ -71,6 +72,8 @@ providerError(error: unknown, secrets?: readonly (string | undefined)[]): Provid
|
|
|
71
72
|
toolCallContent(id: string, name: string, args?: JsonObject): ToolCallContent
|
|
72
73
|
```
|
|
73
74
|
|
|
75
|
+
`tool_call_delta` fragments use the same `{ index, id?, name?, argumentsText? }` shape as live `message_delta` content. Runtime reconstructs final `ToolCallContent` before tool execution; conformance tests use the same reconstruction rules.
|
|
76
|
+
|
|
74
77
|
### Mock provider
|
|
75
78
|
|
|
76
79
|
```ts
|
|
@@ -84,13 +87,49 @@ createMockProvider(events?: readonly ProviderEvent[], options?: MockProviderOpti
|
|
|
84
87
|
| `id` | `string` | Optional provider id. Defaults to `mock`. |
|
|
85
88
|
| `onRequest` | `(request: ProviderRequest) => void` | Optional request observer for tests. |
|
|
86
89
|
|
|
90
|
+
### Provider resolver
|
|
91
|
+
|
|
92
|
+
```ts
|
|
93
|
+
export type ProviderResolver = (model: ModelConfig) => AIProvider | undefined;
|
|
94
|
+
|
|
95
|
+
createProviderResolver(source: ProviderRegistry | readonly AIProvider[]): ProviderResolver
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
A `ProviderResolver` maps a `ModelConfig` to an `AIProvider` for the current
|
|
99
|
+
run. `createProviderResolver()` builds one from a `ProviderRegistry` (reuses
|
|
100
|
+
`ProviderRegistry.get`) or a plain `AIProvider[]` (builds an id-keyed lookup
|
|
101
|
+
once at construction; last duplicate id wins). A custom function is the
|
|
102
|
+
zero-helper path for hosts with their own provider map (lazy construction,
|
|
103
|
+
per-request routing).
|
|
104
|
+
|
|
105
|
+
The resolver returns `undefined` on a miss; the agent runtime fails closed with
|
|
106
|
+
`Unknown provider: ${model.provider}` before any provider turn (see
|
|
107
|
+
[Agent/session runtime](agent-session-runtime.md)).
|
|
108
|
+
|
|
109
|
+
Wire `providerSource` on `AgentConfig`, override per run with
|
|
110
|
+
`RunOptions.providerSource` (RunOptions wins). When `AgentConfig.provider` is
|
|
111
|
+
set it takes first precedence and the resolver is bypassed. The resolver is
|
|
112
|
+
called once per run with `options.model ?? config.model`; per-turn
|
|
113
|
+
re-resolution is unnecessary.
|
|
114
|
+
|
|
115
|
+
```ts
|
|
116
|
+
import { createAgent, createProviderResolver, createProviderRegistry } from "@arnilo/prism";
|
|
117
|
+
|
|
118
|
+
const own = createMyProvider();
|
|
119
|
+
const providerSource = createProviderResolver(createProviderRegistry([own]));
|
|
120
|
+
// or mix first-party + own in one list:
|
|
121
|
+
// const providerSource = createProviderResolver([firstPartyProvider, own]);
|
|
122
|
+
|
|
123
|
+
const agent = createAgent({ model: { provider: own.id, model: "demo" }, providerSource });
|
|
124
|
+
```
|
|
125
|
+
|
|
87
126
|
## Outputs / response / events
|
|
88
127
|
|
|
89
128
|
- Registry `resolve()` returns the matching provider/model or throws before any provider `generate()` call.
|
|
90
129
|
- Provider event helpers return plain `ProviderEvent` objects.
|
|
91
130
|
- `providerError()` converts unknown errors to redacted `ErrorInfo` through `errorToErrorInfo()` and preserves safe string/number `code` fields for retry classification.
|
|
92
131
|
- `createMockProvider()` returns an `AIProvider` whose `generate()` yields the scripted events in order and checks `request.signal?.aborted` before each event.
|
|
93
|
-
- The agent/session runtime passes its per-run abort signal as `ProviderRequest.signal`.
|
|
132
|
+
- The agent/session runtime passes its per-run abort signal as `ProviderRequest.signal`. `ProviderRequestOptions.timeoutMs`, `maxRetries`, and `maxRetryDelayMs` are deprecated inert hints in first-party providers; use `RunOptions.signal`/host abort controllers for timeouts and `AgentConfig.retry`/`RunOptions.retry` for retry.
|
|
94
133
|
|
|
95
134
|
## Request/response example
|
|
96
135
|
|
|
@@ -147,6 +186,7 @@ for await (const event of resolvedProvider.generate({
|
|
|
147
186
|
## Extension and configuration notes
|
|
148
187
|
|
|
149
188
|
- Registries are explicit objects returned by factories. Prism does not create a hidden global provider/model registry.
|
|
189
|
+
- Default duplicate policy is `"replace"` for compatibility. Hosts that load third-party provider/model contributions can pass `duplicate: "error"` to reject silent shadowing.
|
|
150
190
|
- Extension packages can contribute `AIProvider` and `ModelConfig` values by registering them with host-owned registries.
|
|
151
191
|
- Model resolution and provider resolution are separate on purpose: hosts can validate a model exists before selecting a provider.
|
|
152
192
|
- Credential resolvers stay outside these registries; pass credentials directly to the provider adapter or runtime edge that needs them.
|
|
@@ -154,7 +194,7 @@ for await (const event of resolvedProvider.generate({
|
|
|
154
194
|
|
|
155
195
|
## Security and performance notes
|
|
156
196
|
|
|
157
|
-
- Provider/model registries are `Map`-backed and perform O(1) lookup.
|
|
197
|
+
- Provider/model registries are `Map`-backed and perform O(1) lookup. Strict duplicate mode adds one O(1) `Map.has()` check during registration only.
|
|
158
198
|
- Registries store providers and model metadata only. Do not store API keys, credential resolvers, headers, tokens, or secret-bearing settings in them.
|
|
159
199
|
- Unknown provider/model resolution fails before provider execution or network I/O.
|
|
160
200
|
- `createMockProvider()` uses scripted events only: no timers, credentials, SDKs, or network.
|
|
@@ -164,7 +204,7 @@ for await (const event of resolvedProvider.generate({
|
|
|
164
204
|
|
|
165
205
|
## Related APIs
|
|
166
206
|
|
|
167
|
-
- [Agent/session runtime](agent-session-runtime.md): passes abort signals to providers, maps provider errors to session `error` events, and can retry configured transient provider-turn failures before output.
|
|
207
|
+
- [Agent/session runtime](agent-session-runtime.md): passes abort signals to providers, maps provider errors to session `error` events, and can retry configured transient provider-turn failures before output; this is the supported replacement for deprecated provider-level retry options.
|
|
168
208
|
- [Provider packages](provider-packages.md): explicit package primitive for registering providers, models, auth descriptors, request/cache policies, and prompt contributions.
|
|
169
209
|
- [Public contracts](public-contracts.md): `AIProvider`, `ProviderRequest`, `ProviderEvent`, `ModelConfig`, `Usage`, and content/tool-call contracts.
|
|
170
210
|
- [Credentials and redaction](credentials-and-redaction.md): credential and redaction helpers used by provider adapters.
|
|
@@ -53,13 +53,70 @@ api.registerProviderRequestPolicy(createSessionCachePolicy({ retention: "short"
|
|
|
53
53
|
api.registerSystemPromptContribution({ id: "demo-prompt", source: "package", mode: "append", text: "Use demo provider rules." });
|
|
54
54
|
```
|
|
55
55
|
|
|
56
|
-
Hosts decide which credential resolvers, env objects, OAuth stores, request policies, and prompt contributions become active. Request policies can set generic `ProviderRequest.options` such as `sessionId`, `cacheRetention`, `headers`,
|
|
56
|
+
Hosts decide which credential resolvers, env objects, OAuth stores, request policies, and prompt contributions become active. Request policies can set generic `ProviderRequest.options` such as `sessionId`, `cacheRetention`, `headers`, `compat`, and opaque `extra`; provider adapters decide how to map those options to provider payloads. Caller headers are extension headers only: provider adapters must apply provider-owned headers (auth, content type, session/cache/security, attribution) after caller headers so requests cannot override credentials or provider policy.
|
|
57
|
+
|
|
58
|
+
Deprecated provider request options: `timeoutMs`, `maxRetries`, and `maxRetryDelayMs` are inert in first-party providers. Use `RunOptions.signal`/host abort controllers for timeouts and `AgentConfig.retry`/`RunOptions.retry` for retry. Provider packages should not add provider-specific retry loops unless the vendor protocol requires it and runtime retry cannot cover the failure mode.
|
|
59
|
+
|
|
60
|
+
First-party providers map generic `ModelConfig.parameters.maxTokens` to real output-token request fields instead of sending `maxTokens` on the wire: OpenAI Responses uses `max_output_tokens`; OpenRouter, OpenCode Go OpenAI-compatible, OpenCode Go Anthropic-style, Z.AI, Kimi, and NeuralWatt use `max_tokens`. Other `model.parameters` values pass through unchanged unless the provider docs say otherwise.
|
|
57
61
|
|
|
58
62
|
## First-party provider package skeletons
|
|
59
63
|
|
|
60
|
-
Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md),
|
|
64
|
+
Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md). Each package starts with a side-effect-free `create*ProviderPackage()` export, README, TypeScript build, network-free default tests, and real opt-in live smoke tests.
|
|
65
|
+
|
|
66
|
+
Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
|
|
67
|
+
|
|
68
|
+
These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
|
|
69
|
+
|
|
70
|
+
### First-party cache behavior
|
|
71
|
+
|
|
72
|
+
Every first-party provider package hardens prompt-cache behavior so it cannot emit invalid cache retention values or over-broad cache-control markers, and so provider-owned `authorization`/session/security headers cannot be overridden by caller `ProviderRequest.options.headers`. Cache behavior is provider-specific and best-effort: OpenAI/OpenRouter use explicit hints, NeuralWatt/Z.AI use implicit caching, and OpenCode Go/Kimi are route/model-dependent. See [Provider caching](provider-caching.md#per-provider-cache-behavior) for the canonical explicit/implicit matrix.
|
|
73
|
+
|
|
74
|
+
- **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
75
|
+
- **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
76
|
+
- **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; `cache_control` markers applied only to caller-selected `cache.breakpoints` (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
|
|
77
|
+
- **OpenCode Go**: `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; the Anthropic route applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`), the OpenAI route sends none. Per-route usage mapping.
|
|
78
|
+
- **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
|
|
79
|
+
- **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
|
|
80
|
+
- **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
|
|
81
|
+
|
|
82
|
+
See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
|
|
83
|
+
|
|
84
|
+
## Third-party provider packaging
|
|
85
|
+
|
|
86
|
+
A third party ships their own providers the same way Prism ships first-party
|
|
87
|
+
provider packages: an `Extension` whose `setup(api)` calls
|
|
88
|
+
`api.registerProvider(provider)` for each provider it owns. First-party
|
|
89
|
+
provider packages (`@arnilo/prism-provider-openai`, `@arnilo/prism-provider-openrouter`,
|
|
90
|
+
`@arnilo/prism-provider-kimi`, `@arnilo/prism-provider-zai`,
|
|
91
|
+
`@arnilo/prism-provider-opencode-go`) are **opt-in and individually installable**;
|
|
92
|
+
`@arnilo/prism` core runs without any first-party provider package (mock-only).
|
|
61
93
|
|
|
62
|
-
|
|
94
|
+
A host mixes first-party packages and third-party providers in one resolver.
|
|
95
|
+
The host owns the resolver — declaring a provider does not activate it:
|
|
96
|
+
|
|
97
|
+
```ts
|
|
98
|
+
import { createExtensionKernel, createProviderResolver, createAgent } from "@arnilo/prism";
|
|
99
|
+
import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
|
|
100
|
+
|
|
101
|
+
// First-party package, inert until loaded.
|
|
102
|
+
const kernel = createExtensionKernel();
|
|
103
|
+
await kernel.load([createOpenAIProviderPackage({ apiKey: () => process.env.OPENAI_API_KEY })]);
|
|
104
|
+
|
|
105
|
+
// Third-party own provider (bring your own adapter). Combined with first-party
|
|
106
|
+
// providers in one resolver passed to the agent as `providerSource`.
|
|
107
|
+
const own = createMyProvider(/* credentials */);
|
|
108
|
+
const providerSource = createProviderResolver([...kernel.registries.providers.list(), own]);
|
|
109
|
+
|
|
110
|
+
const agent = createAgent({ model: { provider: own.id, model: "demo" }, providerSource });
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
The resolver is the selection mechanism: `model.provider` selects which
|
|
114
|
+
provider runs per turn. Hosts can build the resolver from a `ProviderRegistry`,
|
|
115
|
+
a plain `AIProvider[]`, or implement `ProviderResolver` directly as a one-line
|
|
116
|
+
function over their own map (lazy construction, per-request routing). Declaring
|
|
117
|
+
a provider grants no permissions and forces no activation; the host always has
|
|
118
|
+
final say. See [Provider layer § Provider resolver](provider-layer.md#provider-resolver)
|
|
119
|
+
for the resolver contract.
|
|
63
120
|
|
|
64
121
|
## Outputs / response / events
|
|
65
122
|
|
|
@@ -67,7 +124,7 @@ These workspaces still follow the same rule as external packages: no provider SD
|
|
|
67
124
|
|
|
68
125
|
## Request/response example
|
|
69
126
|
|
|
70
|
-
Provider package manifest contribution and the generic request options a provider request policy can set:
|
|
127
|
+
Provider package manifest contribution and the generic request options a provider request policy can set (see [Provider request policies](provider-request-policies.md) and [Provider caching](provider-caching.md)):
|
|
71
128
|
|
|
72
129
|
```json
|
|
73
130
|
{
|
|
@@ -83,7 +140,8 @@ Provider package manifest contribution and the generic request options a provide
|
|
|
83
140
|
"cacheKey": "demo",
|
|
84
141
|
"cacheRetention": "short",
|
|
85
142
|
"headers": { "x-demo": "1" }
|
|
86
|
-
}
|
|
143
|
+
},
|
|
144
|
+
"runOptions.retry": { "maxAttempts": 3, "maxDelayMs": 1000 }
|
|
87
145
|
}
|
|
88
146
|
```
|
|
89
147
|
|
|
@@ -124,6 +182,7 @@ await kernel.load([pkg]);
|
|
|
124
182
|
(`provider_request`) that sets generic `ProviderRequest.options`
|
|
125
183
|
(`sessionId`, `cacheKey`, `cacheRetention`, `headers`, opaque `extra`) before
|
|
126
184
|
`AIProvider.generate()`; provider adapters map those options to provider payloads.
|
|
185
|
+
- `ModelConfig.cache` is the generic cache capability metadata documented in [Model registry](model-registry.md); `ModelConfig.compat` remains provider-owned inert JSON for behavior that has no generic field yet.
|
|
127
186
|
- `ModelConfig.compat` is provider-owned inert JSON: cache policy overrides,
|
|
128
187
|
reasoning/thinking formats, and provider-specific usage mapping live there
|
|
129
188
|
rather than in core, so Prism never branches on provider names.
|
|
@@ -139,6 +198,7 @@ await kernel.load([pkg]);
|
|
|
139
198
|
- Registration is in-memory only and does no filesystem, network, env, OAuth refresh, or command access.
|
|
140
199
|
- Provider-specific behavior belongs in provider packages, not Prism core.
|
|
141
200
|
- Adapter serializers should preserve Prism content blocks (text, thinking, tool_call, tool_result, and image when the model declares image input) in provider-native request shape, or fail explicitly when a block is unsupported.
|
|
201
|
+
- Adapter header merging must put caller-supplied `ProviderRequest.options.headers` first and provider-owned headers last. Caller headers may add non-owned headers, but cannot replace resolved credentials, content type, session/cache/security headers, or provider attribution headers.
|
|
142
202
|
|
|
143
203
|
## Manifest declarations
|
|
144
204
|
|