@arnilo/prism 0.0.1 → 0.0.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +19 -2
- package/README.md +17 -7
- package/dist/agent-definitions.d.ts +12 -0
- package/dist/agent-definitions.js +131 -0
- package/dist/agent-loops.d.ts +14 -0
- package/dist/agent-loops.js +161 -0
- package/dist/agents.js +263 -76
- package/dist/cache-helpers.d.ts +28 -0
- package/dist/cache-helpers.js +73 -0
- package/dist/cli-runner.d.ts +38 -2
- package/dist/cli-runner.js +167 -5
- package/dist/compaction.js +2 -0
- package/dist/config.js +47 -12
- package/dist/contracts.d.ts +581 -6
- package/dist/contracts.js +41 -1
- package/dist/contribution-parsing.d.ts +19 -0
- package/dist/contribution-parsing.js +124 -0
- package/dist/contributions.d.ts +13 -3
- package/dist/contributions.js +96 -20
- package/dist/extensions.js +3 -0
- package/dist/index.d.ts +19 -9
- package/dist/index.js +10 -4
- package/dist/input.d.ts +7 -1
- package/dist/input.js +52 -11
- package/dist/instruction-injection.d.ts +28 -0
- package/dist/instruction-injection.js +55 -0
- package/dist/manifests.d.ts +1 -1
- package/dist/manifests.js +3 -3
- package/dist/models.d.ts +4 -1
- package/dist/models.js +5 -2
- package/dist/node/agent-definitions.d.ts +98 -0
- package/dist/node/agent-definitions.js +389 -0
- package/dist/node/contribution-discovery.d.ts +17 -0
- package/dist/node/contribution-discovery.js +163 -0
- package/dist/node/instruction-injectors.d.ts +32 -0
- package/dist/node/instruction-injectors.js +72 -0
- package/dist/node/session-store-jsonl.d.ts +1 -1
- package/dist/node/session-store-jsonl.js +42 -4
- package/dist/node/system-project-prompts.d.ts +30 -0
- package/dist/node/system-project-prompts.js +53 -0
- package/dist/provider-events.d.ts +3 -1
- package/dist/provider-events.js +34 -0
- package/dist/provider-request-policy.js +15 -1
- package/dist/providers/openai-compatible.js +1 -1
- package/dist/providers.d.ts +6 -2
- package/dist/providers.js +15 -1
- package/dist/redaction.d.ts +2 -1
- package/dist/redaction.js +3 -0
- package/dist/registry-options.d.ts +5 -0
- package/dist/registry-options.js +5 -0
- package/dist/rpc.d.ts +6 -2
- package/dist/rpc.js +71 -13
- package/dist/session-stores.d.ts +3 -1
- package/dist/session-stores.js +67 -6
- package/dist/skills.d.ts +4 -1
- package/dist/skills.js +3 -1
- package/dist/system-prompts.js +6 -2
- package/dist/testing/compaction-conformance.d.ts +17 -0
- package/dist/testing/compaction-conformance.js +61 -0
- package/dist/testing/extension-conformance.d.ts +26 -0
- package/dist/testing/extension-conformance.js +55 -0
- package/dist/testing/provider-conformance.d.ts +7 -0
- package/dist/testing/provider-conformance.js +18 -31
- package/dist/testing/session-store-conformance.d.ts +20 -0
- package/dist/testing/session-store-conformance.js +92 -0
- package/dist/testing/tool-conformance.d.ts +39 -0
- package/dist/testing/tool-conformance.js +79 -0
- package/dist/tools.d.ts +7 -2
- package/dist/tools.js +50 -13
- package/docs/agent-definitions.md +251 -0
- package/docs/agent-events.md +199 -0
- package/docs/agent-loops.md +217 -0
- package/docs/agent-session-runtime.md +20 -8
- package/docs/cli-rpc.md +39 -4
- package/docs/coding-agent-tools.md +208 -0
- package/docs/compaction-and-retry.md +2 -2
- package/docs/compaction-conformance.md +76 -0
- package/docs/compaction-llm.md +6 -3
- package/docs/compaction-observational-memory.md +4 -4
- package/docs/configuration-and-manifests.md +6 -1
- package/docs/context-and-skills.md +79 -6
- package/docs/contribution-discovery.md +149 -0
- package/docs/contribution-registries.md +9 -6
- package/docs/credentials-and-redaction.md +2 -0
- package/docs/customization.md +191 -0
- package/docs/database-persistence.md +407 -0
- package/docs/extension-authoring.md +193 -0
- package/docs/extension-conformance.md +80 -0
- package/docs/extensions.md +6 -0
- package/docs/host-security.md +141 -0
- package/docs/index.md +41 -19
- package/docs/input-and-prompt-assembly.md +19 -3
- package/docs/instruction-injection.md +183 -0
- package/docs/migration.md +201 -0
- package/docs/model-registry.md +122 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/performance.md +127 -0
- package/docs/provider-caching.md +206 -0
- package/docs/provider-conformance.md +32 -5
- package/docs/provider-layer.md +51 -11
- package/docs/provider-packages.md +65 -5
- package/docs/provider-request-policies.md +113 -0
- package/docs/providers/kimi.md +22 -0
- package/docs/providers/neuralwatt.md +388 -0
- package/docs/providers/openai-compatible.md +1 -0
- package/docs/providers/openai.md +21 -0
- package/docs/providers/opencode-go.md +31 -3
- package/docs/providers/openrouter.md +29 -0
- package/docs/providers/zai.md +17 -0
- package/docs/public-contracts.md +87 -12
- package/docs/release-and-install.md +79 -27
- package/docs/runs-and-usage.md +236 -0
- package/docs/session-store-conformance.md +78 -0
- package/docs/session-stores-and-branching.md +10 -6
- package/docs/session-stores.md +126 -0
- package/docs/settings-auth-trust-security.md +18 -4
- package/docs/structured-output.md +247 -0
- package/docs/system-prompts.md +104 -2
- package/docs/tool-conformance.md +87 -0
- package/docs/tools.md +65 -8
- package/package.json +36 -2
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
# Provider request policies
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Provider request policies are small host/package hooks that can adjust `ProviderRequest.options` before `AIProvider.generate()` runs.
|
|
6
|
+
|
|
7
|
+
Public helpers:
|
|
8
|
+
|
|
9
|
+
- `createProviderRequestPolicyChain(policies)` runs policies in order.
|
|
10
|
+
- `createSessionCachePolicy(options)` sets legacy `cacheKey` / `cacheRetention` aliases from `sessionId`.
|
|
11
|
+
- `mergeProviderRequestOptions(base, patch)` merges request options, including structured `cache` hints.
|
|
12
|
+
|
|
13
|
+
## When to use it
|
|
14
|
+
|
|
15
|
+
Use provider request policies when an app or provider package needs to set generic per-request options such as cache hints, caller-owned headers, `compat`, or `extra` without changing every provider call site.
|
|
16
|
+
|
|
17
|
+
Do not use request policies to resolve credentials, read env vars, perform OAuth refresh, fetch model lists, or override provider-owned auth/session/security headers.
|
|
18
|
+
|
|
19
|
+
## Inputs / request
|
|
20
|
+
|
|
21
|
+
```ts
|
|
22
|
+
import type { ProviderRequestPolicy, ProviderRequestPolicyContext, ProviderRequestOptions } from "@arnilo/prism";
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
| API | Input | Purpose |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| `ProviderRequestPolicy.apply(context)` | `{ sessionId?, request, options? }` | Returns a patched request or options. |
|
|
28
|
+
| `createProviderRequestPolicyChain(policies)` | ordered policies | Applies patches in order. |
|
|
29
|
+
| `createSessionCachePolicy({ retention?, cacheKey? })` | optional cache defaults | Sets legacy aliases. |
|
|
30
|
+
| `mergeProviderRequestOptions(base, patch)` | two option bags | Shallow merges scalars and structurally merges `cache`. |
|
|
31
|
+
|
|
32
|
+
`mergeProviderRequestOptions()` behavior:
|
|
33
|
+
|
|
34
|
+
- Patch scalar fields win.
|
|
35
|
+
- `headers`, `compat`, and `extra` shallow-merge.
|
|
36
|
+
- `cache` shallow-merges; patch `mode`, `key`, and `retention` win.
|
|
37
|
+
- `cache.breakpoints` concatenate in base-then-patch order.
|
|
38
|
+
- Legacy-only `cacheKey` / `cacheRetention` merges remain unchanged and do not add a `cache` property.
|
|
39
|
+
|
|
40
|
+
## Outputs / response / events
|
|
41
|
+
|
|
42
|
+
A policy chain returns either a full `ProviderRequest` or `{ request, options }` style result, normalized by the chain before the next policy runs. The final request is what the agent/session runtime passes to the provider.
|
|
43
|
+
|
|
44
|
+
No agent events are emitted by the policy chain itself.
|
|
45
|
+
|
|
46
|
+
## Request/response example
|
|
47
|
+
|
|
48
|
+
```json
|
|
49
|
+
{
|
|
50
|
+
"before": { "options": { "cacheRetention": "short" } },
|
|
51
|
+
"patch": { "options": { "cache": { "key": "stable", "retention": "long" } } },
|
|
52
|
+
"after": {
|
|
53
|
+
"options": {
|
|
54
|
+
"cacheRetention": "short",
|
|
55
|
+
"cache": { "key": "stable", "retention": "long" }
|
|
56
|
+
}
|
|
57
|
+
}
|
|
58
|
+
}
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
## Implementation example
|
|
62
|
+
|
|
63
|
+
```ts
|
|
64
|
+
import {
|
|
65
|
+
createProviderRequestPolicyChain,
|
|
66
|
+
createSessionCachePolicy,
|
|
67
|
+
mergeProviderRequestOptions,
|
|
68
|
+
type ProviderRequestPolicy,
|
|
69
|
+
} from "@arnilo/prism";
|
|
70
|
+
|
|
71
|
+
const structuredCache: ProviderRequestPolicy = {
|
|
72
|
+
name: "demo.structured-cache",
|
|
73
|
+
apply({ request }) {
|
|
74
|
+
return {
|
|
75
|
+
...request,
|
|
76
|
+
options: mergeProviderRequestOptions(request.options, {
|
|
77
|
+
cache: {
|
|
78
|
+
mode: "on",
|
|
79
|
+
key: request.options?.sessionId,
|
|
80
|
+
retention: "long",
|
|
81
|
+
breakpoints: [{ location: "system_prompt" }],
|
|
82
|
+
},
|
|
83
|
+
}),
|
|
84
|
+
};
|
|
85
|
+
},
|
|
86
|
+
};
|
|
87
|
+
|
|
88
|
+
const chain = createProviderRequestPolicyChain([
|
|
89
|
+
createSessionCachePolicy({ retention: "short" }),
|
|
90
|
+
structuredCache,
|
|
91
|
+
]);
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
## Extension and configuration notes
|
|
95
|
+
|
|
96
|
+
Provider packages can register request policies during `defineProviderPackage().setup(api)`. Hosts decide which packages/policies load and in which order. Prism has no hidden provider request policy registry and no automatic provider-specific cache behavior in core.
|
|
97
|
+
|
|
98
|
+
Policy output should stay generic: use `ProviderRequestOptions.cache`, `headers`, `compat`, and `extra` instead of provider-name branches in core.
|
|
99
|
+
|
|
100
|
+
## Security and performance notes
|
|
101
|
+
|
|
102
|
+
- Request policies must not store or log credentials.
|
|
103
|
+
- Caller headers are advisory; provider adapters must apply provider-owned auth/session/security headers last.
|
|
104
|
+
- Cache keys must never be credentials.
|
|
105
|
+
- Policy chains are O(number of policies) plus option merge cost.
|
|
106
|
+
- Policies should be pure and synchronous unless the host explicitly accepts async work.
|
|
107
|
+
|
|
108
|
+
## Related APIs
|
|
109
|
+
|
|
110
|
+
- [Provider caching](provider-caching.md): structured cache hints and helpers.
|
|
111
|
+
- [Provider packages](provider-packages.md): registering policies from extension packages.
|
|
112
|
+
- [Provider layer](provider-layer.md): provider request flow and `AIProvider.generate()`.
|
|
113
|
+
- [Public contracts](public-contracts.md): `ProviderRequestPolicy`, `ProviderRequestOptions`, and cache types.
|
package/docs/providers/kimi.md
CHANGED
|
@@ -90,12 +90,34 @@ await kernel.load([
|
|
|
90
90
|
`includeMoonshotModels: true`; it is not core behavior.
|
|
91
91
|
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
92
92
|
|
|
93
|
+
### Cache behavior
|
|
94
|
+
|
|
95
|
+
- Default catalog models (e.g. `kimi-k2.7-code` on the Anthropic-compatible
|
|
96
|
+
`/messages` route) use **implicit caching** and send no explicit `cache_control`
|
|
97
|
+
fields. `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention` have no
|
|
98
|
+
effect on the request body unless the model opts in.
|
|
99
|
+
- Hosts may opt a model into Anthropic-style `cache_control` by declaring
|
|
100
|
+
`ModelConfig.cache.kind: "cache_control"` on the Anthropic route. When opted in,
|
|
101
|
+
`cache_control: { type: "ephemeral" }` markers are applied only to the
|
|
102
|
+
caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
|
|
103
|
+
shared `applyCacheControl()` helper) on the last content block of each selected
|
|
104
|
+
message — not to every block. `cacheRetention: "long"` adds `ttl: "1h"` when the
|
|
105
|
+
model allows long retention (`ModelConfig.cache.longRetention !== false`).
|
|
106
|
+
- The Moonshot Open Platform route (`compat.route: "openai"`) never receives
|
|
107
|
+
Anthropic `cache_control` fields.
|
|
108
|
+
- Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
|
|
109
|
+
`Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
|
|
110
|
+
`Usage.cacheWriteTokens`.
|
|
111
|
+
|
|
93
112
|
## Security and performance notes
|
|
94
113
|
|
|
95
114
|
- No network calls during import, setup, build, or default tests.
|
|
96
115
|
- No automatic environment, file, keychain, or shell credential lookup.
|
|
97
116
|
- Kimi credentials are resolved per request from caller-supplied values or resolvers
|
|
98
117
|
and redacted from errors.
|
|
118
|
+
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
|
|
119
|
+
but provider-owned headers (`content-type`, `user-agent`, `authorization`)
|
|
120
|
+
are applied last and cannot be overridden by caller headers.
|
|
99
121
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
|
|
100
122
|
provider-specific env names; default tests are network-free.
|
|
101
123
|
|
|
@@ -0,0 +1,388 @@
|
|
|
1
|
+
# NeuralWatt provider package
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-provider-neuralwatt` provides explicit, side-effect-free setup for the
|
|
6
|
+
NeuralWatt OpenAI-compatible Chat Completions provider using Prism's OpenAI-compatible
|
|
7
|
+
route with NeuralWatt-specific reasoning/template escape hatches, SSE comment tolerance,
|
|
8
|
+
and implicit prefix caching.
|
|
9
|
+
|
|
10
|
+
The package registers a provider, default model metadata for the featured NeuralWatt
|
|
11
|
+
aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
|
|
12
|
+
`kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
|
|
13
|
+
`qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
|
|
14
|
+
method through `createExtensionKernel().load([...])`.
|
|
15
|
+
|
|
16
|
+
## When to use it
|
|
17
|
+
|
|
18
|
+
Use it when a host app wants to run the NeuralWatt endpoint (`https://api.neuralwatt.com/v1`)
|
|
19
|
+
through Prism's `AgentSession` runtime with NeuralWatt-specific `reasoning_effort`,
|
|
20
|
+
`thinking_token_budget`, and `chat_template_kwargs` handling.
|
|
21
|
+
|
|
22
|
+
Do not use it for automatic credential discovery, catalog fetches, or real-network tests.
|
|
23
|
+
|
|
24
|
+
## Inputs / request
|
|
25
|
+
|
|
26
|
+
```ts
|
|
27
|
+
import {
|
|
28
|
+
classifyNeuralWattError,
|
|
29
|
+
createNeuralWattProviderPackage,
|
|
30
|
+
defineNeuralWattModel,
|
|
31
|
+
getNeuralWattQuota,
|
|
32
|
+
listNeuralWattModels,
|
|
33
|
+
mapNeuralWattTelemetry,
|
|
34
|
+
neuralWattEventsWithTelemetry,
|
|
35
|
+
neuralWattModels,
|
|
36
|
+
parseNeuralWattComment,
|
|
37
|
+
} from "@arnilo/prism-provider-neuralwatt";
|
|
38
|
+
|
|
39
|
+
createNeuralWattProviderPackage(options: NeuralWattProviderPackageOptions): ProviderPackage
|
|
40
|
+
defineNeuralWattModel(config: NeuralWattModelConfig): ModelConfig
|
|
41
|
+
listNeuralWattModels(options?: ListNeuralWattModelsOptions): Promise<ModelConfig[]>
|
|
42
|
+
getNeuralWattQuota(options: GetNeuralWattQuotaOptions): Promise<NeuralWattQuota>
|
|
43
|
+
classifyNeuralWattError(input: NeuralWattErrorInput): NeuralWattRetryDecision
|
|
44
|
+
mapNeuralWattTelemetry(body: unknown): { energy?: NeuralWattEnergyTelemetry; cost?: NeuralWattCostTelemetry }
|
|
45
|
+
parseNeuralWattComment(text: string): NeuralWattTelemetryEvent | undefined
|
|
46
|
+
neuralWattEventsWithTelemetry(body: ReadableStream<Uint8Array>): AsyncIterable<NeuralWattEvent>
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
| Field | Type | Purpose |
|
|
50
|
+
| --- | --- | --- |
|
|
51
|
+
| `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
|
|
52
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
53
|
+
| `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
|
|
54
|
+
| `id` | `string` | Overrides the provider id (default `neuralwatt`). |
|
|
55
|
+
| `models` | `readonly ModelConfig[]` | Overrides `neuralWattModels` defaults. |
|
|
56
|
+
|
|
57
|
+
`listNeuralWattModels()` options:
|
|
58
|
+
|
|
59
|
+
| Field | Type | Purpose |
|
|
60
|
+
| --- | --- | --- |
|
|
61
|
+
| `apiKey` | `CredentialValueSource` | Optional API-key source. Unauthenticated calls return the public catalog; authenticated calls may include private models. |
|
|
62
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
63
|
+
| `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
|
|
64
|
+
| `signal` | `AbortSignal` | Cancels the single discovery request. |
|
|
65
|
+
| `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
|
|
66
|
+
|
|
67
|
+
`getNeuralWattQuota()` options:
|
|
68
|
+
|
|
69
|
+
| Field | Type | Purpose |
|
|
70
|
+
| --- | --- | --- |
|
|
71
|
+
| `apiKey` | `CredentialValueSource` | **Required.** NeuralWatt returns 401 for unauthenticated quota calls. |
|
|
72
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
73
|
+
| `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
|
|
74
|
+
| `signal` | `AbortSignal` | Cancels the single quota request. |
|
|
75
|
+
| `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
|
|
76
|
+
|
|
77
|
+
The endpoint is rate-limited to **1 request per second per customer** (429 with
|
|
78
|
+
`Retry-After: 1`). The helper performs no polling or caching and is never called
|
|
79
|
+
from `generate()` or package setup; the caller owns throttling.
|
|
80
|
+
|
|
81
|
+
NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
|
|
82
|
+
/ `extra` escape hatches: `compat.reasoning_effort` (`"low" | "medium" | "high"`),
|
|
83
|
+
`compat.thinking_token_budget`, `compat.chat_template_kwargs` (including `enable_thinking`),
|
|
84
|
+
`compat.preserve_thinking`, `compat.clear_thinking`, and `compat.tool_choice`.
|
|
85
|
+
`preserve_thinking: true` keeps prior assistant reasoning in request history so
|
|
86
|
+
multi-turn reasoning continues with the earlier chain of thought; `clear_thinking:
|
|
87
|
+
true` drops it for the next turn, resetting the chain. `options.extra` spreads after
|
|
88
|
+
`compat` so per-call values and overrides win.
|
|
89
|
+
|
|
90
|
+
## Outputs / response / events
|
|
91
|
+
|
|
92
|
+
| Surface | Behavior |
|
|
93
|
+
| --- | --- |
|
|
94
|
+
| Provider stream | Prism text, thinking (`delta.reasoning_content` → `providerThinkingDelta`), tool-call delta/final, `usage`, `done`, redacted `error` with HTTP-status `code` for retry classification. |
|
|
95
|
+
| Block preservation | Text, thinking, assistant `tool_call` → `tool_calls`, `tool_result` → role `tool` messages, images when `capabilities.input` includes `"image"`. |
|
|
96
|
+
| Model catalog | Featured aliases declare provider id, display name, context limit, text/image input support, tools, reasoning/fast variants, streaming, implicit cache, and NeuralWatt JSON-mode compat metadata where documented. |
|
|
97
|
+
| Pricing | Static aliases do not guess rates. Exact per-alias input/output/cache-read prices are advertised by NeuralWatt's `/v1/models` response and mapped by `listNeuralWattModels()` when present. |
|
|
98
|
+
| SSE comments | `: energy` / `: cost` comment lines are parsed by `neuralWattEventsWithTelemetry()` into `neuralwatt:telemetry` events; the standard `neuralWattEvents()` stream (used by `generate()`) tolerates them without spurious events. |
|
|
99
|
+
| `[DONE]` | Terminates the stream; final `providerDone(usage)` always emitted on a clean stream. |
|
|
100
|
+
| Malformed data | Yields `providerError` rather than crashing the generator (more robust than the Z.AI parser). |
|
|
101
|
+
| Auth method | `api_key` for the configured provider id, credential name `apiKey`. |
|
|
102
|
+
|
|
103
|
+
Unsupported block placements or unclaimed images fail before fetch.
|
|
104
|
+
|
|
105
|
+
## Request/response example
|
|
106
|
+
|
|
107
|
+
Example request body (OpenAI-compatible Chat Completions shape):
|
|
108
|
+
|
|
109
|
+
```json
|
|
110
|
+
{
|
|
111
|
+
"model": "glm-5.2",
|
|
112
|
+
"messages": [{ "role": "user", "content": "Hello" }],
|
|
113
|
+
"stream": true,
|
|
114
|
+
"stream_options": { "include_usage": true },
|
|
115
|
+
"reasoning_effort": "medium",
|
|
116
|
+
"thinking_token_budget": 8192
|
|
117
|
+
}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
## Implementation example
|
|
121
|
+
|
|
122
|
+
```ts
|
|
123
|
+
import { createExtensionKernel } from "@arnilo/prism";
|
|
124
|
+
import { createNeuralWattProviderPackage } from "@arnilo/prism-provider-neuralwatt";
|
|
125
|
+
|
|
126
|
+
const kernel = createExtensionKernel();
|
|
127
|
+
await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake-neuralwatt-key" })]);
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
Override the provider id and models:
|
|
131
|
+
|
|
132
|
+
```ts
|
|
133
|
+
import { createNeuralWattProviderPackage, defineNeuralWattModel, neuralWattModels } from "@arnilo/prism-provider-neuralwatt";
|
|
134
|
+
|
|
135
|
+
await kernel.load([
|
|
136
|
+
createNeuralWattProviderPackage({ id: "neuralwatt", apiKey: "fake", models: neuralWattModels }),
|
|
137
|
+
]);
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
Explicit catalog discovery:
|
|
141
|
+
|
|
142
|
+
```ts
|
|
143
|
+
import { listNeuralWattModels } from "@arnilo/prism-provider-neuralwatt";
|
|
144
|
+
|
|
145
|
+
const models = await listNeuralWattModels({ apiKey: "fake-neuralwatt-key", fetch });
|
|
146
|
+
await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake", models })]);
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
Discovery performs exactly one `GET /v1/models` call when invoked. Provider package
|
|
150
|
+
setup and `generate()` never call model discovery implicitly.
|
|
151
|
+
|
|
152
|
+
Account quota:
|
|
153
|
+
|
|
154
|
+
```ts
|
|
155
|
+
import { getNeuralWattQuota } from "@arnilo/prism-provider-neuralwatt";
|
|
156
|
+
|
|
157
|
+
const quota = await getNeuralWattQuota({ apiKey: "fake-neuralwatt-key", fetch });
|
|
158
|
+
console.log(quota.usage?.current_month?.energy_kwh, quota.balance?.balance_usd);
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
Returns typed `NeuralWattQuota` (`balance`, `usage.lifetime`/`usage.current_month`,
|
|
162
|
+
`limits`, `subscription`, `key`). All fields optional; minimal structural
|
|
163
|
+
validation. The caller owns throttling/caching — the helper makes one explicit
|
|
164
|
+
`GET /v1/quota` call and is never invoked from `generate()` or setup.
|
|
165
|
+
|
|
166
|
+
## Extension and configuration notes
|
|
167
|
+
|
|
168
|
+
- Hosts choose base URL, provider id, model list, credential source, and `fetch`
|
|
169
|
+
impl.
|
|
170
|
+
- `defineNeuralWattModel` lets apps set NeuralWatt-specific `compat`
|
|
171
|
+
(`reasoning_effort`, `thinking_token_budget`, `chat_template_kwargs`,
|
|
172
|
+
`preserve_thinking`, `clear_thinking`, `tool_choice`).
|
|
173
|
+
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
174
|
+
- Curated aliases are static and network-free. They include documented context windows
|
|
175
|
+
and capabilities only; `ModelConfig.cost` is left unset until exact per-alias pricing
|
|
176
|
+
is read from NeuralWatt's `/v1/models` catalog.
|
|
177
|
+
- `listNeuralWattModels()` maps `/v1/models` entries to `ModelConfig`: id/display
|
|
178
|
+
name, capabilities, limits, implicit cache metadata, `ModelCost` pricing, and
|
|
179
|
+
provider-owned NeuralWatt metadata in `compat.neuralwatt`.
|
|
180
|
+
- `getNeuralWattQuota()` calls `GET /v1/quota` once with a required API key and
|
|
181
|
+
returns typed account quota (`balance`, `usage`, `limits`, `subscription`, `key`).
|
|
182
|
+
It is opt-in, never called from `generate()` or setup, and the caller owns
|
|
183
|
+
throttling (NeuralWatt limits the endpoint to 1 rps per customer).
|
|
184
|
+
|
|
185
|
+
### Model catalog and pricing
|
|
186
|
+
|
|
187
|
+
`neuralWattModels` includes NeuralWatt's featured aliases:
|
|
188
|
+
|
|
189
|
+
| Alias | Context | Notable metadata |
|
|
190
|
+
| --- | ---: | --- |
|
|
191
|
+
| `glm-5.2` | 1024K | Tools, reasoning |
|
|
192
|
+
| `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
|
|
193
|
+
| `glm-5.2-short` | 195K | Tools, reasoning |
|
|
194
|
+
| `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
|
|
195
|
+
| `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
|
|
196
|
+
| `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
|
|
197
|
+
| `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
|
|
198
|
+
| `qwen3.5-397b` | 256K | Tools, reasoning, JSON mode |
|
|
199
|
+
| `qwen3.5-397b-fast` | 256K | Tools, JSON mode, fast/no reasoning |
|
|
200
|
+
| `qwen3.6-35b` | 128K | Tools, reasoning, vision, JSON mode |
|
|
201
|
+
| `qwen3.6-35b-fast` | 128K | Tools, vision, JSON mode, fast/no reasoning |
|
|
202
|
+
|
|
203
|
+
NeuralWatt exposes exact pricing from `GET /v1/models` as per-million-token
|
|
204
|
+
`input_per_million`, `output_per_million`, `cached_input_per_million`,
|
|
205
|
+
`cached_output_per_million`, `currency`, and `pricing_tbd`. Cache reads for
|
|
206
|
+
NeuralWatt-hosted models are advertised by the API and default to 25% of the input
|
|
207
|
+
rate; there is no separate cache-write price (`cached_output_per_million` is `null`).
|
|
208
|
+
The static catalog does not copy or infer prices that are not published as fixed
|
|
209
|
+
alias values in these docs.
|
|
210
|
+
|
|
211
|
+
### Cache behavior
|
|
212
|
+
|
|
213
|
+
- NeuralWatt models use **implicit prefix caching**: the server caches prompt prefixes
|
|
214
|
+
automatically based on request content, with no explicit request-side cache payload.
|
|
215
|
+
Catalog models declare `cache: { kind: "implicit" }`. For the cross-provider
|
|
216
|
+
explicit/implicit cache matrix, see [Provider caching](../provider-caching.md).
|
|
217
|
+
- The provider sends no `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention`
|
|
218
|
+
fields regardless of `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention`
|
|
219
|
+
settings — those options have no effect on the NeuralWatt request body.
|
|
220
|
+
`cacheRetention: "none"` disables Prism cache-control hints only; it does **not**
|
|
221
|
+
disable the implicit backend prefix cache. Hosts relying on cache hits should keep
|
|
222
|
+
their stable prompt prefix byte-stable and stable inputs unchanged.
|
|
223
|
+
- Usage accounting is read-only: `prompt_tokens_details.cached_tokens` maps to
|
|
224
|
+
`Usage.cacheReadTokens`. NeuralWatt does not report a cache-write token today, so
|
|
225
|
+
`Usage.cacheWriteTokens` is never fabricated (stays `undefined`).
|
|
226
|
+
|
|
227
|
+
#### Cache-aware limiter behavior
|
|
228
|
+
|
|
229
|
+
NeuralWatt's backend rate limiter is cache-aware, which affects when long-running
|
|
230
|
+
agent sessions are throttled versus served:
|
|
231
|
+
|
|
232
|
+
- **Uncached TPM counts cold prefill only.** Tokens-per-minute accounting charges the
|
|
233
|
+
cold prefill of a request — the prefix that is not already in the vLLM prefix cache.
|
|
234
|
+
A request whose prefix is fully cached consumes far less of the TPM budget than a
|
|
235
|
+
cold request of the same total prompt length.
|
|
236
|
+
- **Warm-prefix requests can avoid some `503` fleet-capacity blocks.** When the fleet
|
|
237
|
+
is near capacity, requests that can reuse a cached prefix are more likely to be
|
|
238
|
+
admitted than fully cold requests. Prefix reuse is therefore both a latency and an
|
|
239
|
+
availability lever, not just a cost lever.
|
|
240
|
+
- **Full prior history is required for multi-turn cache reuse.** The prefix cache is
|
|
241
|
+
keyed by request content, so each follow-up turn must resend the entire prior
|
|
242
|
+
transcript (system prompt + all prior turns) unchanged, with only the new turn
|
|
243
|
+
appended. Prism's `inputLayout: "cache_aware"` ordering keeps the stable prefix
|
|
244
|
+
first; see [Provider caching](../provider-caching.md).
|
|
245
|
+
- Cache behavior is best-effort and **does not guarantee cache hits**. Prefix cache
|
|
246
|
+
admission and eviction are server-side decisions and can vary with fleet load.
|
|
247
|
+
`cacheRetention: "none"` disables Prism cache-control hints only; it does not
|
|
248
|
+
disable the implicit backend prefix cache.
|
|
249
|
+
|
|
250
|
+
### Reasoning preservation across turns
|
|
251
|
+
|
|
252
|
+
NeuralWatt reasoning-capable models (Kimi-, GLM-, Qwen-style aliases with
|
|
253
|
+
`capabilities.reasoning: true`) accept prior assistant reasoning in request history
|
|
254
|
+
so multi-turn sessions continue the earlier chain of thought:
|
|
255
|
+
|
|
256
|
+
- Prior `thinking` content blocks on an assistant message are serialized under a
|
|
257
|
+
`reasoning_content` field on that message (matching the streaming
|
|
258
|
+
`delta.reasoning_content` field). They are **not** flattened into text `content`, so
|
|
259
|
+
the model sees reasoning and answer as distinct.
|
|
260
|
+
- Preservation is gated on `model.capabilities.reasoning === true` **or**
|
|
261
|
+
`compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
|
|
262
|
+
field and prior `thinking` blocks are dropped — they never leak into text content for
|
|
263
|
+
providers/models that do not support reasoning.
|
|
264
|
+
- `compat.clear_thinking: true` drops prior reasoning for the next turn even on
|
|
265
|
+
reasoning-capable models, resetting the chain of thought. `clear_thinking` takes
|
|
266
|
+
precedence over `preserve_thinking`.
|
|
267
|
+
- The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
|
|
268
|
+
reasoning.
|
|
269
|
+
|
|
270
|
+
### Tool calls and the tool-call loop
|
|
271
|
+
|
|
272
|
+
NeuralWatt exposes OpenAI-compatible function calling. The provider carries tools and
|
|
273
|
+
prior tool turns through a multi-turn loop:
|
|
274
|
+
|
|
275
|
+
- **Request serialization.** `ProviderRequest.tools` (`ToolDefinition[]`) is serialized
|
|
276
|
+
to OpenAI `tools: [{ type: "function", function: { name, description, parameters } }]`.
|
|
277
|
+
Missing `parameters` default to `{ type: "object" }`. `compat.tool_choice` passes
|
|
278
|
+
through as `tool_choice` (string or `{ type: "function", function: { name } }`).
|
|
279
|
+
- **Streaming reconstruction.** `delta.tool_calls` fragments (keyed by `index`) are
|
|
280
|
+
accumulated and re-emitted as `tool_call_delta` events for UI consumers, then
|
|
281
|
+
reconstructed into a final `tool_call` event per call with parsed JSON arguments.
|
|
282
|
+
Parallel calls are tracked by index.
|
|
283
|
+
- **Next-turn ordering.** On the following turn the assistant `tool_call` block is
|
|
284
|
+
serialized to `role: "assistant"` with a `tool_calls` array (arguments stringified to
|
|
285
|
+
JSON), immediately followed by a `role: "tool"` message carrying `tool_call_id` and
|
|
286
|
+
the stringified `tool_result` — matching the OpenAI requirement that a tool result
|
|
287
|
+
follows the call that produced it. `tool_result` blocks must appear in `role: "tool"
|
|
288
|
+
messages; `tool_call` blocks must be the only content on their assistant message.
|
|
289
|
+
|
|
290
|
+
### Energy and cost telemetry
|
|
291
|
+
|
|
292
|
+
NeuralWatt streams energy and cost data as SSE comment lines (`: energy {...}`
|
|
293
|
+
and `: cost {...}`) before `data: [DONE]`, and as top-level `energy`/`cost` JSON
|
|
294
|
+
fields on non-streaming responses. Standard SSE clients ignore comments, so
|
|
295
|
+
these values are invisible unless the raw stream is parsed.
|
|
296
|
+
|
|
297
|
+
Prism's core `ProviderEvent` union has no generic telemetry event, so NeuralWatt
|
|
298
|
+
exposes telemetry through package-specific helpers:
|
|
299
|
+
|
|
300
|
+
- `neuralWattEventsWithTelemetry(body)` yields the standard provider events plus
|
|
301
|
+
`neuralwatt:telemetry` events (`{ type: "neuralwatt:telemetry", energy?, cost? }`)
|
|
302
|
+
in stream order. Use it when a host wants to observe telemetry alongside text,
|
|
303
|
+
tool, usage, and done events.
|
|
304
|
+
- `parseNeuralWattComment(text)` parses a single `: energy`/`: cost` comment line
|
|
305
|
+
into a `NeuralWattTelemetryEvent` (`undefined` for unknown/malformed comments).
|
|
306
|
+
- `parseNeuralWattEnergy(payload)` / `parseNeuralWattCost(payload)` parse the JSON
|
|
307
|
+
payload of a single comment into typed `NeuralWattEnergyTelemetry` /
|
|
308
|
+
`NeuralWattCostTelemetry`.
|
|
309
|
+
- `mapNeuralWattTelemetry(body)` maps a non-streaming response body's top-level
|
|
310
|
+
`energy`/`cost` fields into the same typed telemetry.
|
|
311
|
+
|
|
312
|
+
`generate()` stays streaming-only and uses `neuralWattEvents()`, so telemetry is
|
|
313
|
+
opt-in via `neuralWattEventsWithTelemetry()`. Telemetry contains usage/cost
|
|
314
|
+
numbers only — never prompts, API keys, or headers. All documented fields are
|
|
315
|
+
optional and tolerated when absent; malformed comments yield no telemetry event
|
|
316
|
+
and never crash the stream.
|
|
317
|
+
|
|
318
|
+
```ts
|
|
319
|
+
import { neuralWattEventsWithTelemetry } from "@arnilo/prism-provider-neuralwatt";
|
|
320
|
+
|
|
321
|
+
for await (const event of neuralWattEventsWithTelemetry(response.body)) {
|
|
322
|
+
if (event.type === "neuralwatt:telemetry") {
|
|
323
|
+
console.log(event.energy?.energy_kwh, event.cost?.request_cost_usd);
|
|
324
|
+
}
|
|
325
|
+
}
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
### Retry classification
|
|
329
|
+
|
|
330
|
+
NeuralWatt error responses are classified by `classifyNeuralWattError()` so the
|
|
331
|
+
Prism runtime retry policy can decide retryability without provider-specific
|
|
332
|
+
core branches:
|
|
333
|
+
|
|
334
|
+
| Status | Retryable | Notes |
|
|
335
|
+
| --- | --- | --- |
|
|
336
|
+
| `400` `401` `402` `403` `404` | no | Client/payment/auth errors fail closed. |
|
|
337
|
+
| `429` | yes | Reads `Retry-After` header and `error.retry_after`; preserves `error.retry_strategy` (`type`, `suggested_initial_delay_s`, `max_delay_s`, `backoff`, `jitter`). |
|
|
338
|
+
| `500` `502` `503` | yes | Transient server/fleet-capacity errors; `503` `Retry-After` honored when present. |
|
|
339
|
+
|
|
340
|
+
The provider emits `providerError` with `ErrorInfo.code` set to the numeric HTTP
|
|
341
|
+
status. Prism's default retry policy (`createDefaultRetryPolicy()`) treats `429`/
|
|
342
|
+
`500`/`502`/`503` as transient and `400`/`401`/`402`/`403`/`404` as non-transient,
|
|
343
|
+
so NeuralWatt errors retry correctly out of the box. `classifyNeuralWattError()`
|
|
344
|
+
and `neuralWattHttpError()` are exported for hosts/tests that want structured
|
|
345
|
+
retry metadata (`retryAfterMs`, `errorCode`, `strategy`). The host retry policy
|
|
346
|
+
owns the exact delay; `retryAfterMs` is surfaced but not enforced by the
|
|
347
|
+
provider. Classification is O(1) over status/headers/body and makes no extra
|
|
348
|
+
provider calls.
|
|
349
|
+
|
|
350
|
+
```ts
|
|
351
|
+
import { classifyNeuralWattError } from "@arnilo/prism-provider-neuralwatt";
|
|
352
|
+
|
|
353
|
+
const decision = classifyNeuralWattError({ status: 429, headers: { "retry-after": "1" }, body: { error: { code: "concurrent_budget_exceeded", retry_after: 1 } } });
|
|
354
|
+
// { retryable: true, code: 429, retryAfterMs: 1000, errorCode: "concurrent_budget_exceeded", strategy: undefined }
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
## Security and performance notes
|
|
358
|
+
|
|
359
|
+
- No network calls during import, setup, build, default tests, or generation beyond
|
|
360
|
+
the explicit Chat Completions request. `listNeuralWattModels()` is opt-in and
|
|
361
|
+
makes one `GET /v1/models` call per invocation; `getNeuralWattQuota()` is opt-in
|
|
362
|
+
and makes one `GET /v1/quota` call per invocation (endpoint limited to 1 rps per
|
|
363
|
+
customer; caller owns throttling).
|
|
364
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
365
|
+
- API keys are resolved per request/helper call from caller-supplied values or resolvers
|
|
366
|
+
and redacted from errors via `redactSecrets`. `listNeuralWattModels()` and
|
|
367
|
+
`getNeuralWattQuota()` apply provider-owned `authorization` after caller headers so
|
|
368
|
+
callers cannot override it. Quota values never enter provider events unless the caller
|
|
369
|
+
emits them.
|
|
370
|
+
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
|
|
371
|
+
but provider-owned headers (`content-type`, `authorization`) are applied last
|
|
372
|
+
and cannot be overridden by caller headers.
|
|
373
|
+
- Live tests stay opt-in behind `NEURALWATT_API_KEY` (plus `PRISM_LIVE_PROVIDER_TESTS=1`);
|
|
374
|
+
default tests are network-free.
|
|
375
|
+
|
|
376
|
+
## Related APIs
|
|
377
|
+
|
|
378
|
+
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
379
|
+
`ModelConfig`/`compat`, thinking formats.
|
|
380
|
+
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
381
|
+
`resolveCredentialValue`, `redactSecrets`.
|
|
382
|
+
- [OpenAI-compatible provider](openai-compatible.md): underlying Chat Completions
|
|
383
|
+
adapter.
|
|
384
|
+
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
385
|
+
- [Provider caching](../provider-caching.md): implicit cache behavior and
|
|
386
|
+
`cacheUsageReport`.
|
|
387
|
+
- [NeuralWatt agent example](../../examples/neuralwatt-agent-run.ts): runnable mocked
|
|
388
|
+
agent turn with tools, reasoning controls, streamed cache tokens, and energy/cost telemetry.
|
|
@@ -108,6 +108,7 @@ const provider = createOpenAICompatibleProvider({
|
|
|
108
108
|
- The adapter resolves `apiKey` per request through `resolveCredentialValue()`.
|
|
109
109
|
- This adapter currently targets Chat Completions streaming only.
|
|
110
110
|
- The serializer preserves text, thinking (downgraded to text), assistant `tool_call` blocks as `tool_calls`, `tool_result` blocks as role `tool` messages, and image blocks when the model declares `capabilities.input` includes `"image"`. Unsupported block placements or unclaimed images fail before fetch.
|
|
111
|
+
- Cache behavior is intentionally minimal: this Chat Completions adapter sends no `prompt_cache_key`, `prompt_cache_retention`, or `cache_control` fields. Endpoints that cache implicitly do so automatically; hosts needing OpenAI `prompt_cache_key`/`prompt_cache_retention` should use the [`@arnilo/prism-provider-openai`](openai.md) Responses package. The adapter still normalizes cache usage from `prompt_tokens_details.cached_tokens` (and `prompt_cache_hit_tokens`) into `Usage.cacheReadTokens`.
|
|
111
112
|
|
|
112
113
|
## Security and performance notes
|
|
113
114
|
|
package/docs/providers/openai.md
CHANGED
|
@@ -109,6 +109,27 @@ const challenge = computeS256Challenge(verifier);
|
|
|
109
109
|
- OAuth browser/device-code flows run only when the caller explicitly invokes the
|
|
110
110
|
OAuth provider.
|
|
111
111
|
|
|
112
|
+
### Cache behavior
|
|
113
|
+
|
|
114
|
+
- `prompt_cache_key` is derived from `ProviderRequestOptions.cacheKey` (falling
|
|
115
|
+
back to `sessionId`) and sanitized + clamped to 64 characters via the shared
|
|
116
|
+
`sanitizeCacheKey()` helper. Cache keys are session/customer identifiers only;
|
|
117
|
+
never credentials or raw prompts.
|
|
118
|
+
- `prompt_cache_retention` accepts only `"24h"` on the OpenAI Responses API
|
|
119
|
+
(extended caching). Prism `cacheRetention: "short"` and `"none"` omit the field
|
|
120
|
+
so default automatic/implicit caching applies and no invalid literal is sent.
|
|
121
|
+
`cacheRetention: "long"` maps to `prompt_cache_retention: "24h"` only when the
|
|
122
|
+
model declares `ModelConfig.cache.longRetention === true`; models without that
|
|
123
|
+
metadata omit the field. The catalog `gpt-5.1` model declares
|
|
124
|
+
`cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
|
|
125
|
+
- Cache accounting is preserved in normalized `Usage`: OpenAI
|
|
126
|
+
`input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
|
|
127
|
+
Responses does not report a cache-write token field.
|
|
128
|
+
- Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
|
|
129
|
+
are applied after caller `ProviderRequestOptions.headers` so caller config
|
|
130
|
+
cannot replace credentials, content type, or the session request id; non-owned
|
|
131
|
+
caller headers are kept.
|
|
132
|
+
|
|
112
133
|
## Security and performance notes
|
|
113
134
|
|
|
114
135
|
- No network calls during import, setup, build, or default tests.
|
|
@@ -34,8 +34,9 @@ createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): Prov
|
|
|
34
34
|
| `baseUrl` | `string` | Overrides the OpenCode Go base URL. |
|
|
35
35
|
| `models` | `readonly ModelConfig[]` | Overrides `openCodeGoModels` defaults. |
|
|
36
36
|
|
|
37
|
-
`ProviderRequest.options.sessionId` maps to the
|
|
38
|
-
`
|
|
37
|
+
`ProviderRequest.options.cacheKey` (falling back to `sessionId`) maps to the
|
|
38
|
+
`x-opencode-session` header; the Anthropic-compatible route accepts
|
|
39
|
+
`cache_control` breakpoints; `cacheRetention` maps to cache retention.
|
|
39
40
|
|
|
40
41
|
## Outputs / response / events
|
|
41
42
|
|
|
@@ -52,7 +53,7 @@ Example headers added before fetch:
|
|
|
52
53
|
```json
|
|
53
54
|
{
|
|
54
55
|
"Authorization": "Bearer <resolved-key>",
|
|
55
|
-
"x-opencode-session": "<ProviderRequest.options.sessionId>"
|
|
56
|
+
"x-opencode-session": "<ProviderRequest.options.cacheKey ?? sessionId>"
|
|
56
57
|
}
|
|
57
58
|
```
|
|
58
59
|
|
|
@@ -83,12 +84,39 @@ await kernel.load([
|
|
|
83
84
|
routes preserve `tool_use`/`tool_result` blocks.
|
|
84
85
|
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
85
86
|
|
|
87
|
+
### Cache and session behavior
|
|
88
|
+
|
|
89
|
+
- `x-opencode-session` is derived from `ProviderRequestOptions.cacheKey` (falling
|
|
90
|
+
back to `sessionId`) and sanitized + clamped to 128 characters via the shared
|
|
91
|
+
`sanitizeCacheKey()` helper. Session ids route/stick requests and identify
|
|
92
|
+
conversations; never credentials or raw prompts.
|
|
93
|
+
- The Anthropic-compatible route (`compat.route: "anthropic"`) applies
|
|
94
|
+
Anthropic-style `cache_control: { type: "ephemeral" }` markers only to the
|
|
95
|
+
caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
|
|
96
|
+
shared `applyCacheControl()` helper) on the last content block of each selected
|
|
97
|
+
message — not to every block. Caching is enabled unless disabled
|
|
98
|
+
(`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
|
|
99
|
+
`ModelConfig.cache.kind: "cache_control"` (or `cache.mode: "on"`).
|
|
100
|
+
- `cacheRetention: "long"` emits `cache_control: { type: "ephemeral", ttl: "1h" }`
|
|
101
|
+
markers when the model allows long retention
|
|
102
|
+
(`ModelConfig.cache.longRetention !== false`); otherwise the default ephemeral
|
|
103
|
+
window applies.
|
|
104
|
+
- The OpenAI-compatible chat route (`compat.route: "openai"`, the default) sends
|
|
105
|
+
no Anthropic `cache_control` fields; it relies on OpenAI-style implicit caching.
|
|
106
|
+
- Usage accounting is preserved per route: the OpenAI route maps
|
|
107
|
+
`prompt_tokens_details.cached_tokens`/`cache_write_tokens` to
|
|
108
|
+
`Usage.cacheReadTokens`/`cacheWriteTokens`; the Anthropic route maps
|
|
109
|
+
`cache_read_input_tokens`/`cache_creation_input_tokens`.
|
|
110
|
+
|
|
86
111
|
## Security and performance notes
|
|
87
112
|
|
|
88
113
|
- No network calls during import, setup, build, or default tests.
|
|
89
114
|
- No automatic environment, file, keychain, or shell credential lookup.
|
|
90
115
|
- API keys are resolved per request from caller-supplied values or resolvers and
|
|
91
116
|
redacted from errors.
|
|
117
|
+
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
|
|
118
|
+
provider-owned headers (`content-type`, `x-opencode-session`, `authorization`)
|
|
119
|
+
are applied last and cannot be overridden by caller headers.
|
|
92
120
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
|
|
93
121
|
provider-specific env names; default tests are network-free.
|
|
94
122
|
|