@arnilo/prism 0.0.1 → 0.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (121) hide show
  1. package/CHANGELOG.md +19 -2
  2. package/README.md +17 -7
  3. package/dist/agent-definitions.d.ts +12 -0
  4. package/dist/agent-definitions.js +131 -0
  5. package/dist/agent-loops.d.ts +14 -0
  6. package/dist/agent-loops.js +161 -0
  7. package/dist/agents.js +263 -76
  8. package/dist/cache-helpers.d.ts +28 -0
  9. package/dist/cache-helpers.js +73 -0
  10. package/dist/cli-runner.d.ts +38 -2
  11. package/dist/cli-runner.js +167 -5
  12. package/dist/compaction.js +2 -0
  13. package/dist/config.js +47 -12
  14. package/dist/contracts.d.ts +581 -6
  15. package/dist/contracts.js +41 -1
  16. package/dist/contribution-parsing.d.ts +19 -0
  17. package/dist/contribution-parsing.js +124 -0
  18. package/dist/contributions.d.ts +13 -3
  19. package/dist/contributions.js +96 -20
  20. package/dist/extensions.js +3 -0
  21. package/dist/index.d.ts +19 -9
  22. package/dist/index.js +10 -4
  23. package/dist/input.d.ts +7 -1
  24. package/dist/input.js +52 -11
  25. package/dist/instruction-injection.d.ts +28 -0
  26. package/dist/instruction-injection.js +55 -0
  27. package/dist/manifests.d.ts +1 -1
  28. package/dist/manifests.js +3 -3
  29. package/dist/models.d.ts +4 -1
  30. package/dist/models.js +5 -2
  31. package/dist/node/agent-definitions.d.ts +98 -0
  32. package/dist/node/agent-definitions.js +389 -0
  33. package/dist/node/contribution-discovery.d.ts +17 -0
  34. package/dist/node/contribution-discovery.js +163 -0
  35. package/dist/node/instruction-injectors.d.ts +32 -0
  36. package/dist/node/instruction-injectors.js +72 -0
  37. package/dist/node/session-store-jsonl.d.ts +1 -1
  38. package/dist/node/session-store-jsonl.js +42 -4
  39. package/dist/node/system-project-prompts.d.ts +30 -0
  40. package/dist/node/system-project-prompts.js +53 -0
  41. package/dist/provider-events.d.ts +3 -1
  42. package/dist/provider-events.js +34 -0
  43. package/dist/provider-request-policy.js +15 -1
  44. package/dist/providers/openai-compatible.js +1 -1
  45. package/dist/providers.d.ts +6 -2
  46. package/dist/providers.js +15 -1
  47. package/dist/redaction.d.ts +2 -1
  48. package/dist/redaction.js +3 -0
  49. package/dist/registry-options.d.ts +5 -0
  50. package/dist/registry-options.js +5 -0
  51. package/dist/rpc.d.ts +6 -2
  52. package/dist/rpc.js +71 -13
  53. package/dist/session-stores.d.ts +3 -1
  54. package/dist/session-stores.js +67 -6
  55. package/dist/skills.d.ts +4 -1
  56. package/dist/skills.js +3 -1
  57. package/dist/system-prompts.js +6 -2
  58. package/dist/testing/compaction-conformance.d.ts +17 -0
  59. package/dist/testing/compaction-conformance.js +61 -0
  60. package/dist/testing/extension-conformance.d.ts +26 -0
  61. package/dist/testing/extension-conformance.js +55 -0
  62. package/dist/testing/provider-conformance.d.ts +7 -0
  63. package/dist/testing/provider-conformance.js +18 -31
  64. package/dist/testing/session-store-conformance.d.ts +20 -0
  65. package/dist/testing/session-store-conformance.js +92 -0
  66. package/dist/testing/tool-conformance.d.ts +39 -0
  67. package/dist/testing/tool-conformance.js +79 -0
  68. package/dist/tools.d.ts +7 -2
  69. package/dist/tools.js +50 -13
  70. package/docs/agent-definitions.md +251 -0
  71. package/docs/agent-events.md +199 -0
  72. package/docs/agent-loops.md +217 -0
  73. package/docs/agent-session-runtime.md +20 -8
  74. package/docs/cli-rpc.md +39 -4
  75. package/docs/coding-agent-tools.md +208 -0
  76. package/docs/compaction-and-retry.md +2 -2
  77. package/docs/compaction-conformance.md +76 -0
  78. package/docs/compaction-llm.md +6 -3
  79. package/docs/compaction-observational-memory.md +4 -4
  80. package/docs/configuration-and-manifests.md +6 -1
  81. package/docs/context-and-skills.md +79 -6
  82. package/docs/contribution-discovery.md +149 -0
  83. package/docs/contribution-registries.md +9 -6
  84. package/docs/credentials-and-redaction.md +2 -0
  85. package/docs/customization.md +191 -0
  86. package/docs/database-persistence.md +407 -0
  87. package/docs/extension-authoring.md +193 -0
  88. package/docs/extension-conformance.md +80 -0
  89. package/docs/extensions.md +6 -0
  90. package/docs/host-security.md +141 -0
  91. package/docs/index.md +41 -19
  92. package/docs/input-and-prompt-assembly.md +19 -3
  93. package/docs/instruction-injection.md +183 -0
  94. package/docs/migration.md +201 -0
  95. package/docs/model-registry.md +122 -0
  96. package/docs/node-jsonl-session-store.md +5 -4
  97. package/docs/performance.md +127 -0
  98. package/docs/provider-caching.md +206 -0
  99. package/docs/provider-conformance.md +32 -5
  100. package/docs/provider-layer.md +51 -11
  101. package/docs/provider-packages.md +65 -5
  102. package/docs/provider-request-policies.md +113 -0
  103. package/docs/providers/kimi.md +22 -0
  104. package/docs/providers/neuralwatt.md +388 -0
  105. package/docs/providers/openai-compatible.md +1 -0
  106. package/docs/providers/openai.md +21 -0
  107. package/docs/providers/opencode-go.md +31 -3
  108. package/docs/providers/openrouter.md +29 -0
  109. package/docs/providers/zai.md +17 -0
  110. package/docs/public-contracts.md +87 -12
  111. package/docs/release-and-install.md +79 -27
  112. package/docs/runs-and-usage.md +236 -0
  113. package/docs/session-store-conformance.md +78 -0
  114. package/docs/session-stores-and-branching.md +10 -6
  115. package/docs/session-stores.md +126 -0
  116. package/docs/settings-auth-trust-security.md +18 -4
  117. package/docs/structured-output.md +247 -0
  118. package/docs/system-prompts.md +104 -2
  119. package/docs/tool-conformance.md +87 -0
  120. package/docs/tools.md +65 -8
  121. package/package.json +36 -2
@@ -0,0 +1,113 @@
1
+ # Provider request policies
2
+
3
+ ## What it does
4
+
5
+ Provider request policies are small host/package hooks that can adjust `ProviderRequest.options` before `AIProvider.generate()` runs.
6
+
7
+ Public helpers:
8
+
9
+ - `createProviderRequestPolicyChain(policies)` runs policies in order.
10
+ - `createSessionCachePolicy(options)` sets legacy `cacheKey` / `cacheRetention` aliases from `sessionId`.
11
+ - `mergeProviderRequestOptions(base, patch)` merges request options, including structured `cache` hints.
12
+
13
+ ## When to use it
14
+
15
+ Use provider request policies when an app or provider package needs to set generic per-request options such as cache hints, caller-owned headers, `compat`, or `extra` without changing every provider call site.
16
+
17
+ Do not use request policies to resolve credentials, read env vars, perform OAuth refresh, fetch model lists, or override provider-owned auth/session/security headers.
18
+
19
+ ## Inputs / request
20
+
21
+ ```ts
22
+ import type { ProviderRequestPolicy, ProviderRequestPolicyContext, ProviderRequestOptions } from "@arnilo/prism";
23
+ ```
24
+
25
+ | API | Input | Purpose |
26
+ | --- | --- | --- |
27
+ | `ProviderRequestPolicy.apply(context)` | `{ sessionId?, request, options? }` | Returns a patched request or options. |
28
+ | `createProviderRequestPolicyChain(policies)` | ordered policies | Applies patches in order. |
29
+ | `createSessionCachePolicy({ retention?, cacheKey? })` | optional cache defaults | Sets legacy aliases. |
30
+ | `mergeProviderRequestOptions(base, patch)` | two option bags | Shallow merges scalars and structurally merges `cache`. |
31
+
32
+ `mergeProviderRequestOptions()` behavior:
33
+
34
+ - Patch scalar fields win.
35
+ - `headers`, `compat`, and `extra` shallow-merge.
36
+ - `cache` shallow-merges; patch `mode`, `key`, and `retention` win.
37
+ - `cache.breakpoints` concatenate in base-then-patch order.
38
+ - Legacy-only `cacheKey` / `cacheRetention` merges remain unchanged and do not add a `cache` property.
39
+
40
+ ## Outputs / response / events
41
+
42
+ A policy chain returns either a full `ProviderRequest` or `{ request, options }` style result, normalized by the chain before the next policy runs. The final request is what the agent/session runtime passes to the provider.
43
+
44
+ No agent events are emitted by the policy chain itself.
45
+
46
+ ## Request/response example
47
+
48
+ ```json
49
+ {
50
+ "before": { "options": { "cacheRetention": "short" } },
51
+ "patch": { "options": { "cache": { "key": "stable", "retention": "long" } } },
52
+ "after": {
53
+ "options": {
54
+ "cacheRetention": "short",
55
+ "cache": { "key": "stable", "retention": "long" }
56
+ }
57
+ }
58
+ }
59
+ ```
60
+
61
+ ## Implementation example
62
+
63
+ ```ts
64
+ import {
65
+ createProviderRequestPolicyChain,
66
+ createSessionCachePolicy,
67
+ mergeProviderRequestOptions,
68
+ type ProviderRequestPolicy,
69
+ } from "@arnilo/prism";
70
+
71
+ const structuredCache: ProviderRequestPolicy = {
72
+ name: "demo.structured-cache",
73
+ apply({ request }) {
74
+ return {
75
+ ...request,
76
+ options: mergeProviderRequestOptions(request.options, {
77
+ cache: {
78
+ mode: "on",
79
+ key: request.options?.sessionId,
80
+ retention: "long",
81
+ breakpoints: [{ location: "system_prompt" }],
82
+ },
83
+ }),
84
+ };
85
+ },
86
+ };
87
+
88
+ const chain = createProviderRequestPolicyChain([
89
+ createSessionCachePolicy({ retention: "short" }),
90
+ structuredCache,
91
+ ]);
92
+ ```
93
+
94
+ ## Extension and configuration notes
95
+
96
+ Provider packages can register request policies during `defineProviderPackage().setup(api)`. Hosts decide which packages/policies load and in which order. Prism has no hidden provider request policy registry and no automatic provider-specific cache behavior in core.
97
+
98
+ Policy output should stay generic: use `ProviderRequestOptions.cache`, `headers`, `compat`, and `extra` instead of provider-name branches in core.
99
+
100
+ ## Security and performance notes
101
+
102
+ - Request policies must not store or log credentials.
103
+ - Caller headers are advisory; provider adapters must apply provider-owned auth/session/security headers last.
104
+ - Cache keys must never be credentials.
105
+ - Policy chains are O(number of policies) plus option merge cost.
106
+ - Policies should be pure and synchronous unless the host explicitly accepts async work.
107
+
108
+ ## Related APIs
109
+
110
+ - [Provider caching](provider-caching.md): structured cache hints and helpers.
111
+ - [Provider packages](provider-packages.md): registering policies from extension packages.
112
+ - [Provider layer](provider-layer.md): provider request flow and `AIProvider.generate()`.
113
+ - [Public contracts](public-contracts.md): `ProviderRequestPolicy`, `ProviderRequestOptions`, and cache types.
@@ -90,12 +90,34 @@ await kernel.load([
90
90
  `includeMoonshotModels: true`; it is not core behavior.
91
91
  - Package contributes models via the extension `api` and an `api_key` auth method.
92
92
 
93
+ ### Cache behavior
94
+
95
+ - Default catalog models (e.g. `kimi-k2.7-code` on the Anthropic-compatible
96
+ `/messages` route) use **implicit caching** and send no explicit `cache_control`
97
+ fields. `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention` have no
98
+ effect on the request body unless the model opts in.
99
+ - Hosts may opt a model into Anthropic-style `cache_control` by declaring
100
+ `ModelConfig.cache.kind: "cache_control"` on the Anthropic route. When opted in,
101
+ `cache_control: { type: "ephemeral" }` markers are applied only to the
102
+ caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
103
+ shared `applyCacheControl()` helper) on the last content block of each selected
104
+ message — not to every block. `cacheRetention: "long"` adds `ttl: "1h"` when the
105
+ model allows long retention (`ModelConfig.cache.longRetention !== false`).
106
+ - The Moonshot Open Platform route (`compat.route: "openai"`) never receives
107
+ Anthropic `cache_control` fields.
108
+ - Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
109
+ `Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
110
+ `Usage.cacheWriteTokens`.
111
+
93
112
  ## Security and performance notes
94
113
 
95
114
  - No network calls during import, setup, build, or default tests.
96
115
  - No automatic environment, file, keychain, or shell credential lookup.
97
116
  - Kimi credentials are resolved per request from caller-supplied values or resolvers
98
117
  and redacted from errors.
118
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
119
+ but provider-owned headers (`content-type`, `user-agent`, `authorization`)
120
+ are applied last and cannot be overridden by caller headers.
99
121
  - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
100
122
  provider-specific env names; default tests are network-free.
101
123
 
@@ -0,0 +1,388 @@
1
+ # NeuralWatt provider package
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-neuralwatt` provides explicit, side-effect-free setup for the
6
+ NeuralWatt OpenAI-compatible Chat Completions provider using Prism's OpenAI-compatible
7
+ route with NeuralWatt-specific reasoning/template escape hatches, SSE comment tolerance,
8
+ and implicit prefix caching.
9
+
10
+ The package registers a provider, default model metadata for the featured NeuralWatt
11
+ aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
12
+ `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
13
+ `qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
14
+ method through `createExtensionKernel().load([...])`.
15
+
16
+ ## When to use it
17
+
18
+ Use it when a host app wants to run the NeuralWatt endpoint (`https://api.neuralwatt.com/v1`)
19
+ through Prism's `AgentSession` runtime with NeuralWatt-specific `reasoning_effort`,
20
+ `thinking_token_budget`, and `chat_template_kwargs` handling.
21
+
22
+ Do not use it for automatic credential discovery, catalog fetches, or real-network tests.
23
+
24
+ ## Inputs / request
25
+
26
+ ```ts
27
+ import {
28
+ classifyNeuralWattError,
29
+ createNeuralWattProviderPackage,
30
+ defineNeuralWattModel,
31
+ getNeuralWattQuota,
32
+ listNeuralWattModels,
33
+ mapNeuralWattTelemetry,
34
+ neuralWattEventsWithTelemetry,
35
+ neuralWattModels,
36
+ parseNeuralWattComment,
37
+ } from "@arnilo/prism-provider-neuralwatt";
38
+
39
+ createNeuralWattProviderPackage(options: NeuralWattProviderPackageOptions): ProviderPackage
40
+ defineNeuralWattModel(config: NeuralWattModelConfig): ModelConfig
41
+ listNeuralWattModels(options?: ListNeuralWattModelsOptions): Promise<ModelConfig[]>
42
+ getNeuralWattQuota(options: GetNeuralWattQuotaOptions): Promise<NeuralWattQuota>
43
+ classifyNeuralWattError(input: NeuralWattErrorInput): NeuralWattRetryDecision
44
+ mapNeuralWattTelemetry(body: unknown): { energy?: NeuralWattEnergyTelemetry; cost?: NeuralWattCostTelemetry }
45
+ parseNeuralWattComment(text: string): NeuralWattTelemetryEvent | undefined
46
+ neuralWattEventsWithTelemetry(body: ReadableStream<Uint8Array>): AsyncIterable<NeuralWattEvent>
47
+ ```
48
+
49
+ | Field | Type | Purpose |
50
+ | --- | --- | --- |
51
+ | `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
52
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
53
+ | `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
54
+ | `id` | `string` | Overrides the provider id (default `neuralwatt`). |
55
+ | `models` | `readonly ModelConfig[]` | Overrides `neuralWattModels` defaults. |
56
+
57
+ `listNeuralWattModels()` options:
58
+
59
+ | Field | Type | Purpose |
60
+ | --- | --- | --- |
61
+ | `apiKey` | `CredentialValueSource` | Optional API-key source. Unauthenticated calls return the public catalog; authenticated calls may include private models. |
62
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
63
+ | `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
64
+ | `signal` | `AbortSignal` | Cancels the single discovery request. |
65
+ | `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
66
+
67
+ `getNeuralWattQuota()` options:
68
+
69
+ | Field | Type | Purpose |
70
+ | --- | --- | --- |
71
+ | `apiKey` | `CredentialValueSource` | **Required.** NeuralWatt returns 401 for unauthenticated quota calls. |
72
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
73
+ | `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
74
+ | `signal` | `AbortSignal` | Cancels the single quota request. |
75
+ | `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
76
+
77
+ The endpoint is rate-limited to **1 request per second per customer** (429 with
78
+ `Retry-After: 1`). The helper performs no polling or caching and is never called
79
+ from `generate()` or package setup; the caller owns throttling.
80
+
81
+ NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
82
+ / `extra` escape hatches: `compat.reasoning_effort` (`"low" | "medium" | "high"`),
83
+ `compat.thinking_token_budget`, `compat.chat_template_kwargs` (including `enable_thinking`),
84
+ `compat.preserve_thinking`, `compat.clear_thinking`, and `compat.tool_choice`.
85
+ `preserve_thinking: true` keeps prior assistant reasoning in request history so
86
+ multi-turn reasoning continues with the earlier chain of thought; `clear_thinking:
87
+ true` drops it for the next turn, resetting the chain. `options.extra` spreads after
88
+ `compat` so per-call values and overrides win.
89
+
90
+ ## Outputs / response / events
91
+
92
+ | Surface | Behavior |
93
+ | --- | --- |
94
+ | Provider stream | Prism text, thinking (`delta.reasoning_content` → `providerThinkingDelta`), tool-call delta/final, `usage`, `done`, redacted `error` with HTTP-status `code` for retry classification. |
95
+ | Block preservation | Text, thinking, assistant `tool_call` → `tool_calls`, `tool_result` → role `tool` messages, images when `capabilities.input` includes `"image"`. |
96
+ | Model catalog | Featured aliases declare provider id, display name, context limit, text/image input support, tools, reasoning/fast variants, streaming, implicit cache, and NeuralWatt JSON-mode compat metadata where documented. |
97
+ | Pricing | Static aliases do not guess rates. Exact per-alias input/output/cache-read prices are advertised by NeuralWatt's `/v1/models` response and mapped by `listNeuralWattModels()` when present. |
98
+ | SSE comments | `: energy` / `: cost` comment lines are parsed by `neuralWattEventsWithTelemetry()` into `neuralwatt:telemetry` events; the standard `neuralWattEvents()` stream (used by `generate()`) tolerates them without spurious events. |
99
+ | `[DONE]` | Terminates the stream; final `providerDone(usage)` always emitted on a clean stream. |
100
+ | Malformed data | Yields `providerError` rather than crashing the generator (more robust than the Z.AI parser). |
101
+ | Auth method | `api_key` for the configured provider id, credential name `apiKey`. |
102
+
103
+ Unsupported block placements or unclaimed images fail before fetch.
104
+
105
+ ## Request/response example
106
+
107
+ Example request body (OpenAI-compatible Chat Completions shape):
108
+
109
+ ```json
110
+ {
111
+ "model": "glm-5.2",
112
+ "messages": [{ "role": "user", "content": "Hello" }],
113
+ "stream": true,
114
+ "stream_options": { "include_usage": true },
115
+ "reasoning_effort": "medium",
116
+ "thinking_token_budget": 8192
117
+ }
118
+ ```
119
+
120
+ ## Implementation example
121
+
122
+ ```ts
123
+ import { createExtensionKernel } from "@arnilo/prism";
124
+ import { createNeuralWattProviderPackage } from "@arnilo/prism-provider-neuralwatt";
125
+
126
+ const kernel = createExtensionKernel();
127
+ await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake-neuralwatt-key" })]);
128
+ ```
129
+
130
+ Override the provider id and models:
131
+
132
+ ```ts
133
+ import { createNeuralWattProviderPackage, defineNeuralWattModel, neuralWattModels } from "@arnilo/prism-provider-neuralwatt";
134
+
135
+ await kernel.load([
136
+ createNeuralWattProviderPackage({ id: "neuralwatt", apiKey: "fake", models: neuralWattModels }),
137
+ ]);
138
+ ```
139
+
140
+ Explicit catalog discovery:
141
+
142
+ ```ts
143
+ import { listNeuralWattModels } from "@arnilo/prism-provider-neuralwatt";
144
+
145
+ const models = await listNeuralWattModels({ apiKey: "fake-neuralwatt-key", fetch });
146
+ await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake", models })]);
147
+ ```
148
+
149
+ Discovery performs exactly one `GET /v1/models` call when invoked. Provider package
150
+ setup and `generate()` never call model discovery implicitly.
151
+
152
+ Account quota:
153
+
154
+ ```ts
155
+ import { getNeuralWattQuota } from "@arnilo/prism-provider-neuralwatt";
156
+
157
+ const quota = await getNeuralWattQuota({ apiKey: "fake-neuralwatt-key", fetch });
158
+ console.log(quota.usage?.current_month?.energy_kwh, quota.balance?.balance_usd);
159
+ ```
160
+
161
+ Returns typed `NeuralWattQuota` (`balance`, `usage.lifetime`/`usage.current_month`,
162
+ `limits`, `subscription`, `key`). All fields optional; minimal structural
163
+ validation. The caller owns throttling/caching — the helper makes one explicit
164
+ `GET /v1/quota` call and is never invoked from `generate()` or setup.
165
+
166
+ ## Extension and configuration notes
167
+
168
+ - Hosts choose base URL, provider id, model list, credential source, and `fetch`
169
+ impl.
170
+ - `defineNeuralWattModel` lets apps set NeuralWatt-specific `compat`
171
+ (`reasoning_effort`, `thinking_token_budget`, `chat_template_kwargs`,
172
+ `preserve_thinking`, `clear_thinking`, `tool_choice`).
173
+ - Package contributes models via the extension `api` and an `api_key` auth method.
174
+ - Curated aliases are static and network-free. They include documented context windows
175
+ and capabilities only; `ModelConfig.cost` is left unset until exact per-alias pricing
176
+ is read from NeuralWatt's `/v1/models` catalog.
177
+ - `listNeuralWattModels()` maps `/v1/models` entries to `ModelConfig`: id/display
178
+ name, capabilities, limits, implicit cache metadata, `ModelCost` pricing, and
179
+ provider-owned NeuralWatt metadata in `compat.neuralwatt`.
180
+ - `getNeuralWattQuota()` calls `GET /v1/quota` once with a required API key and
181
+ returns typed account quota (`balance`, `usage`, `limits`, `subscription`, `key`).
182
+ It is opt-in, never called from `generate()` or setup, and the caller owns
183
+ throttling (NeuralWatt limits the endpoint to 1 rps per customer).
184
+
185
+ ### Model catalog and pricing
186
+
187
+ `neuralWattModels` includes NeuralWatt's featured aliases:
188
+
189
+ | Alias | Context | Notable metadata |
190
+ | --- | ---: | --- |
191
+ | `glm-5.2` | 1024K | Tools, reasoning |
192
+ | `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
193
+ | `glm-5.2-short` | 195K | Tools, reasoning |
194
+ | `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
195
+ | `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
196
+ | `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
197
+ | `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
198
+ | `qwen3.5-397b` | 256K | Tools, reasoning, JSON mode |
199
+ | `qwen3.5-397b-fast` | 256K | Tools, JSON mode, fast/no reasoning |
200
+ | `qwen3.6-35b` | 128K | Tools, reasoning, vision, JSON mode |
201
+ | `qwen3.6-35b-fast` | 128K | Tools, vision, JSON mode, fast/no reasoning |
202
+
203
+ NeuralWatt exposes exact pricing from `GET /v1/models` as per-million-token
204
+ `input_per_million`, `output_per_million`, `cached_input_per_million`,
205
+ `cached_output_per_million`, `currency`, and `pricing_tbd`. Cache reads for
206
+ NeuralWatt-hosted models are advertised by the API and default to 25% of the input
207
+ rate; there is no separate cache-write price (`cached_output_per_million` is `null`).
208
+ The static catalog does not copy or infer prices that are not published as fixed
209
+ alias values in these docs.
210
+
211
+ ### Cache behavior
212
+
213
+ - NeuralWatt models use **implicit prefix caching**: the server caches prompt prefixes
214
+ automatically based on request content, with no explicit request-side cache payload.
215
+ Catalog models declare `cache: { kind: "implicit" }`. For the cross-provider
216
+ explicit/implicit cache matrix, see [Provider caching](../provider-caching.md).
217
+ - The provider sends no `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention`
218
+ fields regardless of `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention`
219
+ settings — those options have no effect on the NeuralWatt request body.
220
+ `cacheRetention: "none"` disables Prism cache-control hints only; it does **not**
221
+ disable the implicit backend prefix cache. Hosts relying on cache hits should keep
222
+ their stable prompt prefix byte-stable and stable inputs unchanged.
223
+ - Usage accounting is read-only: `prompt_tokens_details.cached_tokens` maps to
224
+ `Usage.cacheReadTokens`. NeuralWatt does not report a cache-write token today, so
225
+ `Usage.cacheWriteTokens` is never fabricated (stays `undefined`).
226
+
227
+ #### Cache-aware limiter behavior
228
+
229
+ NeuralWatt's backend rate limiter is cache-aware, which affects when long-running
230
+ agent sessions are throttled versus served:
231
+
232
+ - **Uncached TPM counts cold prefill only.** Tokens-per-minute accounting charges the
233
+ cold prefill of a request — the prefix that is not already in the vLLM prefix cache.
234
+ A request whose prefix is fully cached consumes far less of the TPM budget than a
235
+ cold request of the same total prompt length.
236
+ - **Warm-prefix requests can avoid some `503` fleet-capacity blocks.** When the fleet
237
+ is near capacity, requests that can reuse a cached prefix are more likely to be
238
+ admitted than fully cold requests. Prefix reuse is therefore both a latency and an
239
+ availability lever, not just a cost lever.
240
+ - **Full prior history is required for multi-turn cache reuse.** The prefix cache is
241
+ keyed by request content, so each follow-up turn must resend the entire prior
242
+ transcript (system prompt + all prior turns) unchanged, with only the new turn
243
+ appended. Prism's `inputLayout: "cache_aware"` ordering keeps the stable prefix
244
+ first; see [Provider caching](../provider-caching.md).
245
+ - Cache behavior is best-effort and **does not guarantee cache hits**. Prefix cache
246
+ admission and eviction are server-side decisions and can vary with fleet load.
247
+ `cacheRetention: "none"` disables Prism cache-control hints only; it does not
248
+ disable the implicit backend prefix cache.
249
+
250
+ ### Reasoning preservation across turns
251
+
252
+ NeuralWatt reasoning-capable models (Kimi-, GLM-, Qwen-style aliases with
253
+ `capabilities.reasoning: true`) accept prior assistant reasoning in request history
254
+ so multi-turn sessions continue the earlier chain of thought:
255
+
256
+ - Prior `thinking` content blocks on an assistant message are serialized under a
257
+ `reasoning_content` field on that message (matching the streaming
258
+ `delta.reasoning_content` field). They are **not** flattened into text `content`, so
259
+ the model sees reasoning and answer as distinct.
260
+ - Preservation is gated on `model.capabilities.reasoning === true` **or**
261
+ `compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
262
+ field and prior `thinking` blocks are dropped — they never leak into text content for
263
+ providers/models that do not support reasoning.
264
+ - `compat.clear_thinking: true` drops prior reasoning for the next turn even on
265
+ reasoning-capable models, resetting the chain of thought. `clear_thinking` takes
266
+ precedence over `preserve_thinking`.
267
+ - The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
268
+ reasoning.
269
+
270
+ ### Tool calls and the tool-call loop
271
+
272
+ NeuralWatt exposes OpenAI-compatible function calling. The provider carries tools and
273
+ prior tool turns through a multi-turn loop:
274
+
275
+ - **Request serialization.** `ProviderRequest.tools` (`ToolDefinition[]`) is serialized
276
+ to OpenAI `tools: [{ type: "function", function: { name, description, parameters } }]`.
277
+ Missing `parameters` default to `{ type: "object" }`. `compat.tool_choice` passes
278
+ through as `tool_choice` (string or `{ type: "function", function: { name } }`).
279
+ - **Streaming reconstruction.** `delta.tool_calls` fragments (keyed by `index`) are
280
+ accumulated and re-emitted as `tool_call_delta` events for UI consumers, then
281
+ reconstructed into a final `tool_call` event per call with parsed JSON arguments.
282
+ Parallel calls are tracked by index.
283
+ - **Next-turn ordering.** On the following turn the assistant `tool_call` block is
284
+ serialized to `role: "assistant"` with a `tool_calls` array (arguments stringified to
285
+ JSON), immediately followed by a `role: "tool"` message carrying `tool_call_id` and
286
+ the stringified `tool_result` — matching the OpenAI requirement that a tool result
287
+ follows the call that produced it. `tool_result` blocks must appear in `role: "tool"
288
+ messages; `tool_call` blocks must be the only content on their assistant message.
289
+
290
+ ### Energy and cost telemetry
291
+
292
+ NeuralWatt streams energy and cost data as SSE comment lines (`: energy {...}`
293
+ and `: cost {...}`) before `data: [DONE]`, and as top-level `energy`/`cost` JSON
294
+ fields on non-streaming responses. Standard SSE clients ignore comments, so
295
+ these values are invisible unless the raw stream is parsed.
296
+
297
+ Prism's core `ProviderEvent` union has no generic telemetry event, so NeuralWatt
298
+ exposes telemetry through package-specific helpers:
299
+
300
+ - `neuralWattEventsWithTelemetry(body)` yields the standard provider events plus
301
+ `neuralwatt:telemetry` events (`{ type: "neuralwatt:telemetry", energy?, cost? }`)
302
+ in stream order. Use it when a host wants to observe telemetry alongside text,
303
+ tool, usage, and done events.
304
+ - `parseNeuralWattComment(text)` parses a single `: energy`/`: cost` comment line
305
+ into a `NeuralWattTelemetryEvent` (`undefined` for unknown/malformed comments).
306
+ - `parseNeuralWattEnergy(payload)` / `parseNeuralWattCost(payload)` parse the JSON
307
+ payload of a single comment into typed `NeuralWattEnergyTelemetry` /
308
+ `NeuralWattCostTelemetry`.
309
+ - `mapNeuralWattTelemetry(body)` maps a non-streaming response body's top-level
310
+ `energy`/`cost` fields into the same typed telemetry.
311
+
312
+ `generate()` stays streaming-only and uses `neuralWattEvents()`, so telemetry is
313
+ opt-in via `neuralWattEventsWithTelemetry()`. Telemetry contains usage/cost
314
+ numbers only — never prompts, API keys, or headers. All documented fields are
315
+ optional and tolerated when absent; malformed comments yield no telemetry event
316
+ and never crash the stream.
317
+
318
+ ```ts
319
+ import { neuralWattEventsWithTelemetry } from "@arnilo/prism-provider-neuralwatt";
320
+
321
+ for await (const event of neuralWattEventsWithTelemetry(response.body)) {
322
+ if (event.type === "neuralwatt:telemetry") {
323
+ console.log(event.energy?.energy_kwh, event.cost?.request_cost_usd);
324
+ }
325
+ }
326
+ ```
327
+
328
+ ### Retry classification
329
+
330
+ NeuralWatt error responses are classified by `classifyNeuralWattError()` so the
331
+ Prism runtime retry policy can decide retryability without provider-specific
332
+ core branches:
333
+
334
+ | Status | Retryable | Notes |
335
+ | --- | --- | --- |
336
+ | `400` `401` `402` `403` `404` | no | Client/payment/auth errors fail closed. |
337
+ | `429` | yes | Reads `Retry-After` header and `error.retry_after`; preserves `error.retry_strategy` (`type`, `suggested_initial_delay_s`, `max_delay_s`, `backoff`, `jitter`). |
338
+ | `500` `502` `503` | yes | Transient server/fleet-capacity errors; `503` `Retry-After` honored when present. |
339
+
340
+ The provider emits `providerError` with `ErrorInfo.code` set to the numeric HTTP
341
+ status. Prism's default retry policy (`createDefaultRetryPolicy()`) treats `429`/
342
+ `500`/`502`/`503` as transient and `400`/`401`/`402`/`403`/`404` as non-transient,
343
+ so NeuralWatt errors retry correctly out of the box. `classifyNeuralWattError()`
344
+ and `neuralWattHttpError()` are exported for hosts/tests that want structured
345
+ retry metadata (`retryAfterMs`, `errorCode`, `strategy`). The host retry policy
346
+ owns the exact delay; `retryAfterMs` is surfaced but not enforced by the
347
+ provider. Classification is O(1) over status/headers/body and makes no extra
348
+ provider calls.
349
+
350
+ ```ts
351
+ import { classifyNeuralWattError } from "@arnilo/prism-provider-neuralwatt";
352
+
353
+ const decision = classifyNeuralWattError({ status: 429, headers: { "retry-after": "1" }, body: { error: { code: "concurrent_budget_exceeded", retry_after: 1 } } });
354
+ // { retryable: true, code: 429, retryAfterMs: 1000, errorCode: "concurrent_budget_exceeded", strategy: undefined }
355
+ ```
356
+
357
+ ## Security and performance notes
358
+
359
+ - No network calls during import, setup, build, default tests, or generation beyond
360
+ the explicit Chat Completions request. `listNeuralWattModels()` is opt-in and
361
+ makes one `GET /v1/models` call per invocation; `getNeuralWattQuota()` is opt-in
362
+ and makes one `GET /v1/quota` call per invocation (endpoint limited to 1 rps per
363
+ customer; caller owns throttling).
364
+ - No automatic environment, file, keychain, or shell credential lookup.
365
+ - API keys are resolved per request/helper call from caller-supplied values or resolvers
366
+ and redacted from errors via `redactSecrets`. `listNeuralWattModels()` and
367
+ `getNeuralWattQuota()` apply provider-owned `authorization` after caller headers so
368
+ callers cannot override it. Quota values never enter provider events unless the caller
369
+ emits them.
370
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
371
+ but provider-owned headers (`content-type`, `authorization`) are applied last
372
+ and cannot be overridden by caller headers.
373
+ - Live tests stay opt-in behind `NEURALWATT_API_KEY` (plus `PRISM_LIVE_PROVIDER_TESTS=1`);
374
+ default tests are network-free.
375
+
376
+ ## Related APIs
377
+
378
+ - [Provider packages](../provider-packages.md): `defineProviderPackage`,
379
+ `ModelConfig`/`compat`, thinking formats.
380
+ - [Credentials and redaction](../credentials-and-redaction.md):
381
+ `resolveCredentialValue`, `redactSecrets`.
382
+ - [OpenAI-compatible provider](openai-compatible.md): underlying Chat Completions
383
+ adapter.
384
+ - [Provider conformance](../provider-conformance.md): network-free adapter tests.
385
+ - [Provider caching](../provider-caching.md): implicit cache behavior and
386
+ `cacheUsageReport`.
387
+ - [NeuralWatt agent example](../../examples/neuralwatt-agent-run.ts): runnable mocked
388
+ agent turn with tools, reasoning controls, streamed cache tokens, and energy/cost telemetry.
@@ -108,6 +108,7 @@ const provider = createOpenAICompatibleProvider({
108
108
  - The adapter resolves `apiKey` per request through `resolveCredentialValue()`.
109
109
  - This adapter currently targets Chat Completions streaming only.
110
110
  - The serializer preserves text, thinking (downgraded to text), assistant `tool_call` blocks as `tool_calls`, `tool_result` blocks as role `tool` messages, and image blocks when the model declares `capabilities.input` includes `"image"`. Unsupported block placements or unclaimed images fail before fetch.
111
+ - Cache behavior is intentionally minimal: this Chat Completions adapter sends no `prompt_cache_key`, `prompt_cache_retention`, or `cache_control` fields. Endpoints that cache implicitly do so automatically; hosts needing OpenAI `prompt_cache_key`/`prompt_cache_retention` should use the [`@arnilo/prism-provider-openai`](openai.md) Responses package. The adapter still normalizes cache usage from `prompt_tokens_details.cached_tokens` (and `prompt_cache_hit_tokens`) into `Usage.cacheReadTokens`.
111
112
 
112
113
  ## Security and performance notes
113
114
 
@@ -109,6 +109,27 @@ const challenge = computeS256Challenge(verifier);
109
109
  - OAuth browser/device-code flows run only when the caller explicitly invokes the
110
110
  OAuth provider.
111
111
 
112
+ ### Cache behavior
113
+
114
+ - `prompt_cache_key` is derived from `ProviderRequestOptions.cacheKey` (falling
115
+ back to `sessionId`) and sanitized + clamped to 64 characters via the shared
116
+ `sanitizeCacheKey()` helper. Cache keys are session/customer identifiers only;
117
+ never credentials or raw prompts.
118
+ - `prompt_cache_retention` accepts only `"24h"` on the OpenAI Responses API
119
+ (extended caching). Prism `cacheRetention: "short"` and `"none"` omit the field
120
+ so default automatic/implicit caching applies and no invalid literal is sent.
121
+ `cacheRetention: "long"` maps to `prompt_cache_retention: "24h"` only when the
122
+ model declares `ModelConfig.cache.longRetention === true`; models without that
123
+ metadata omit the field. The catalog `gpt-5.1` model declares
124
+ `cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
125
+ - Cache accounting is preserved in normalized `Usage`: OpenAI
126
+ `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
127
+ Responses does not report a cache-write token field.
128
+ - Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
129
+ are applied after caller `ProviderRequestOptions.headers` so caller config
130
+ cannot replace credentials, content type, or the session request id; non-owned
131
+ caller headers are kept.
132
+
112
133
  ## Security and performance notes
113
134
 
114
135
  - No network calls during import, setup, build, or default tests.
@@ -34,8 +34,9 @@ createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): Prov
34
34
  | `baseUrl` | `string` | Overrides the OpenCode Go base URL. |
35
35
  | `models` | `readonly ModelConfig[]` | Overrides `openCodeGoModels` defaults. |
36
36
 
37
- `ProviderRequest.options.sessionId` maps to the `x-opencode-session` header;
38
- `cacheKey`/`cacheRetention` map to OpenCode cache retention.
37
+ `ProviderRequest.options.cacheKey` (falling back to `sessionId`) maps to the
38
+ `x-opencode-session` header; the Anthropic-compatible route accepts
39
+ `cache_control` breakpoints; `cacheRetention` maps to cache retention.
39
40
 
40
41
  ## Outputs / response / events
41
42
 
@@ -52,7 +53,7 @@ Example headers added before fetch:
52
53
  ```json
53
54
  {
54
55
  "Authorization": "Bearer <resolved-key>",
55
- "x-opencode-session": "<ProviderRequest.options.sessionId>"
56
+ "x-opencode-session": "<ProviderRequest.options.cacheKey ?? sessionId>"
56
57
  }
57
58
  ```
58
59
 
@@ -83,12 +84,39 @@ await kernel.load([
83
84
  routes preserve `tool_use`/`tool_result` blocks.
84
85
  - Package contributes models via the extension `api` and an `api_key` auth method.
85
86
 
87
+ ### Cache and session behavior
88
+
89
+ - `x-opencode-session` is derived from `ProviderRequestOptions.cacheKey` (falling
90
+ back to `sessionId`) and sanitized + clamped to 128 characters via the shared
91
+ `sanitizeCacheKey()` helper. Session ids route/stick requests and identify
92
+ conversations; never credentials or raw prompts.
93
+ - The Anthropic-compatible route (`compat.route: "anthropic"`) applies
94
+ Anthropic-style `cache_control: { type: "ephemeral" }` markers only to the
95
+ caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
96
+ shared `applyCacheControl()` helper) on the last content block of each selected
97
+ message — not to every block. Caching is enabled unless disabled
98
+ (`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
99
+ `ModelConfig.cache.kind: "cache_control"` (or `cache.mode: "on"`).
100
+ - `cacheRetention: "long"` emits `cache_control: { type: "ephemeral", ttl: "1h" }`
101
+ markers when the model allows long retention
102
+ (`ModelConfig.cache.longRetention !== false`); otherwise the default ephemeral
103
+ window applies.
104
+ - The OpenAI-compatible chat route (`compat.route: "openai"`, the default) sends
105
+ no Anthropic `cache_control` fields; it relies on OpenAI-style implicit caching.
106
+ - Usage accounting is preserved per route: the OpenAI route maps
107
+ `prompt_tokens_details.cached_tokens`/`cache_write_tokens` to
108
+ `Usage.cacheReadTokens`/`cacheWriteTokens`; the Anthropic route maps
109
+ `cache_read_input_tokens`/`cache_creation_input_tokens`.
110
+
86
111
  ## Security and performance notes
87
112
 
88
113
  - No network calls during import, setup, build, or default tests.
89
114
  - No automatic environment, file, keychain, or shell credential lookup.
90
115
  - API keys are resolved per request from caller-supplied values or resolvers and
91
116
  redacted from errors.
117
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
118
+ provider-owned headers (`content-type`, `x-opencode-session`, `authorization`)
119
+ are applied last and cannot be overridden by caller headers.
92
120
  - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
93
121
  provider-specific env names; default tests are network-free.
94
122