@arnilo/prism 0.0.1 → 0.0.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. package/CHANGELOG.md +4 -2
  2. package/README.md +17 -7
  3. package/dist/agent-definitions.d.ts +12 -0
  4. package/dist/agent-definitions.js +131 -0
  5. package/dist/agent-loops.d.ts +14 -0
  6. package/dist/agent-loops.js +161 -0
  7. package/dist/agents.js +263 -76
  8. package/dist/cache-helpers.d.ts +28 -0
  9. package/dist/cache-helpers.js +73 -0
  10. package/dist/cli-runner.d.ts +38 -2
  11. package/dist/cli-runner.js +167 -5
  12. package/dist/compaction.js +2 -0
  13. package/dist/config.js +47 -12
  14. package/dist/contracts.d.ts +581 -6
  15. package/dist/contracts.js +41 -1
  16. package/dist/contribution-parsing.d.ts +19 -0
  17. package/dist/contribution-parsing.js +124 -0
  18. package/dist/contributions.d.ts +13 -3
  19. package/dist/contributions.js +96 -20
  20. package/dist/extensions.js +3 -0
  21. package/dist/index.d.ts +19 -9
  22. package/dist/index.js +10 -4
  23. package/dist/input.d.ts +7 -1
  24. package/dist/input.js +52 -11
  25. package/dist/instruction-injection.d.ts +28 -0
  26. package/dist/instruction-injection.js +55 -0
  27. package/dist/manifests.d.ts +1 -1
  28. package/dist/manifests.js +3 -3
  29. package/dist/models.d.ts +4 -1
  30. package/dist/models.js +5 -2
  31. package/dist/node/agent-definitions.d.ts +98 -0
  32. package/dist/node/agent-definitions.js +389 -0
  33. package/dist/node/contribution-discovery.d.ts +17 -0
  34. package/dist/node/contribution-discovery.js +163 -0
  35. package/dist/node/instruction-injectors.d.ts +32 -0
  36. package/dist/node/instruction-injectors.js +72 -0
  37. package/dist/node/session-store-jsonl.d.ts +1 -1
  38. package/dist/node/session-store-jsonl.js +42 -4
  39. package/dist/node/system-project-prompts.d.ts +30 -0
  40. package/dist/node/system-project-prompts.js +53 -0
  41. package/dist/provider-events.d.ts +3 -1
  42. package/dist/provider-events.js +34 -0
  43. package/dist/provider-request-policy.js +15 -1
  44. package/dist/providers/openai-compatible.js +1 -1
  45. package/dist/providers.d.ts +6 -2
  46. package/dist/providers.js +15 -1
  47. package/dist/redaction.d.ts +2 -1
  48. package/dist/redaction.js +3 -0
  49. package/dist/registry-options.d.ts +5 -0
  50. package/dist/registry-options.js +5 -0
  51. package/dist/rpc.d.ts +6 -2
  52. package/dist/rpc.js +71 -13
  53. package/dist/session-stores.d.ts +3 -1
  54. package/dist/session-stores.js +67 -6
  55. package/dist/skills.d.ts +4 -1
  56. package/dist/skills.js +3 -1
  57. package/dist/system-prompts.js +6 -2
  58. package/dist/testing/compaction-conformance.d.ts +17 -0
  59. package/dist/testing/compaction-conformance.js +61 -0
  60. package/dist/testing/extension-conformance.d.ts +26 -0
  61. package/dist/testing/extension-conformance.js +55 -0
  62. package/dist/testing/provider-conformance.d.ts +7 -0
  63. package/dist/testing/provider-conformance.js +18 -31
  64. package/dist/testing/session-store-conformance.d.ts +20 -0
  65. package/dist/testing/session-store-conformance.js +92 -0
  66. package/dist/testing/tool-conformance.d.ts +39 -0
  67. package/dist/testing/tool-conformance.js +79 -0
  68. package/dist/tools.d.ts +7 -2
  69. package/dist/tools.js +50 -13
  70. package/docs/agent-definitions.md +251 -0
  71. package/docs/agent-events.md +199 -0
  72. package/docs/agent-loops.md +217 -0
  73. package/docs/agent-session-runtime.md +20 -8
  74. package/docs/cli-rpc.md +39 -4
  75. package/docs/compaction-and-retry.md +2 -2
  76. package/docs/compaction-conformance.md +76 -0
  77. package/docs/compaction-llm.md +6 -3
  78. package/docs/compaction-observational-memory.md +4 -4
  79. package/docs/configuration-and-manifests.md +6 -1
  80. package/docs/context-and-skills.md +79 -6
  81. package/docs/contribution-discovery.md +149 -0
  82. package/docs/contribution-registries.md +9 -6
  83. package/docs/credentials-and-redaction.md +2 -0
  84. package/docs/customization.md +191 -0
  85. package/docs/database-persistence.md +407 -0
  86. package/docs/extension-authoring.md +193 -0
  87. package/docs/extension-conformance.md +80 -0
  88. package/docs/extensions.md +6 -0
  89. package/docs/host-security.md +141 -0
  90. package/docs/index.md +40 -19
  91. package/docs/input-and-prompt-assembly.md +19 -3
  92. package/docs/instruction-injection.md +183 -0
  93. package/docs/migration.md +201 -0
  94. package/docs/model-registry.md +122 -0
  95. package/docs/node-jsonl-session-store.md +5 -4
  96. package/docs/performance.md +127 -0
  97. package/docs/provider-caching.md +206 -0
  98. package/docs/provider-conformance.md +32 -5
  99. package/docs/provider-layer.md +51 -11
  100. package/docs/provider-packages.md +65 -5
  101. package/docs/provider-request-policies.md +113 -0
  102. package/docs/providers/kimi.md +22 -0
  103. package/docs/providers/neuralwatt.md +388 -0
  104. package/docs/providers/openai-compatible.md +1 -0
  105. package/docs/providers/openai.md +21 -0
  106. package/docs/providers/opencode-go.md +31 -3
  107. package/docs/providers/openrouter.md +29 -0
  108. package/docs/providers/zai.md +17 -0
  109. package/docs/public-contracts.md +87 -12
  110. package/docs/release-and-install.md +76 -26
  111. package/docs/runs-and-usage.md +236 -0
  112. package/docs/session-store-conformance.md +78 -0
  113. package/docs/session-stores-and-branching.md +10 -6
  114. package/docs/session-stores.md +126 -0
  115. package/docs/settings-auth-trust-security.md +18 -4
  116. package/docs/structured-output.md +247 -0
  117. package/docs/system-prompts.md +104 -2
  118. package/docs/tool-conformance.md +87 -0
  119. package/docs/tools.md +64 -8
  120. package/package.json +35 -2
@@ -53,13 +53,70 @@ api.registerProviderRequestPolicy(createSessionCachePolicy({ retention: "short"
53
53
  api.registerSystemPromptContribution({ id: "demo-prompt", source: "package", mode: "append", text: "Use demo provider rules." });
54
54
  ```
55
55
 
56
- Hosts decide which credential resolvers, env objects, OAuth stores, request policies, and prompt contributions become active. Request policies can set generic `ProviderRequest.options` such as `sessionId`, `cacheRetention`, `headers`, retry/timeouts, and opaque `extra`; provider adapters decide how to map those options to provider payloads.
56
+ Hosts decide which credential resolvers, env objects, OAuth stores, request policies, and prompt contributions become active. Request policies can set generic `ProviderRequest.options` such as `sessionId`, `cacheRetention`, `headers`, `compat`, and opaque `extra`; provider adapters decide how to map those options to provider payloads. Caller headers are extension headers only: provider adapters must apply provider-owned headers (auth, content type, session/cache/security, attribution) after caller headers so requests cannot override credentials or provider policy.
57
+
58
+ Deprecated provider request options: `timeoutMs`, `maxRetries`, and `maxRetryDelayMs` are inert in first-party providers. Use `RunOptions.signal`/host abort controllers for timeouts and `AgentConfig.retry`/`RunOptions.retry` for retry. Provider packages should not add provider-specific retry loops unless the vendor protocol requires it and runtime retry cannot cover the failure mode.
59
+
60
+ First-party providers map generic `ModelConfig.parameters.maxTokens` to real output-token request fields instead of sending `maxTokens` on the wire: OpenAI Responses uses `max_output_tokens`; OpenRouter, OpenCode Go OpenAI-compatible, OpenCode Go Anthropic-style, Z.AI, Kimi, and NeuralWatt use `max_tokens`. Other `model.parameters` values pass through unchanged unless the provider docs say otherwise.
57
61
 
58
62
  ## First-party provider package skeletons
59
63
 
60
- Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md), and [`@arnilo/prism-provider-kimi`](providers/kimi.md). Each package starts with a side-effect-free `create*ProviderPackage()` export, README, TypeScript build, network-free default tests, and an env-gated live-test placeholder.
64
+ Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md). Each package starts with a side-effect-free `create*ProviderPackage()` export, README, TypeScript build, network-free default tests, and real opt-in live smoke tests.
65
+
66
+ Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
67
+
68
+ These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
69
+
70
+ ### First-party cache behavior
71
+
72
+ Every first-party provider package hardens prompt-cache behavior so it cannot emit invalid cache retention values or over-broad cache-control markers, and so provider-owned `authorization`/session/security headers cannot be overridden by caller `ProviderRequest.options.headers`. Cache behavior is provider-specific and best-effort: OpenAI/OpenRouter use explicit hints, NeuralWatt/Z.AI use implicit caching, and OpenCode Go/Kimi are route/model-dependent. See [Provider caching](provider-caching.md#per-provider-cache-behavior) for the canonical explicit/implicit matrix.
73
+
74
+ - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
75
+ - **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
76
+ - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; `cache_control` markers applied only to caller-selected `cache.breakpoints` (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
77
+ - **OpenCode Go**: `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; the Anthropic route applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`), the OpenAI route sends none. Per-route usage mapping.
78
+ - **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
79
+ - **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
80
+ - **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
81
+
82
+ See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
83
+
84
+ ## Third-party provider packaging
85
+
86
+ A third party ships their own providers the same way Prism ships first-party
87
+ provider packages: an `Extension` whose `setup(api)` calls
88
+ `api.registerProvider(provider)` for each provider it owns. First-party
89
+ provider packages (`@arnilo/prism-provider-openai`, `@arnilo/prism-provider-openrouter`,
90
+ `@arnilo/prism-provider-kimi`, `@arnilo/prism-provider-zai`,
91
+ `@arnilo/prism-provider-opencode-go`) are **opt-in and individually installable**;
92
+ `@arnilo/prism` core runs without any first-party provider package (mock-only).
61
93
 
62
- These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested.
94
+ A host mixes first-party packages and third-party providers in one resolver.
95
+ The host owns the resolver — declaring a provider does not activate it:
96
+
97
+ ```ts
98
+ import { createExtensionKernel, createProviderResolver, createAgent } from "@arnilo/prism";
99
+ import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
100
+
101
+ // First-party package, inert until loaded.
102
+ const kernel = createExtensionKernel();
103
+ await kernel.load([createOpenAIProviderPackage({ apiKey: () => process.env.OPENAI_API_KEY })]);
104
+
105
+ // Third-party own provider (bring your own adapter). Combined with first-party
106
+ // providers in one resolver passed to the agent as `providerSource`.
107
+ const own = createMyProvider(/* credentials */);
108
+ const providerSource = createProviderResolver([...kernel.registries.providers.list(), own]);
109
+
110
+ const agent = createAgent({ model: { provider: own.id, model: "demo" }, providerSource });
111
+ ```
112
+
113
+ The resolver is the selection mechanism: `model.provider` selects which
114
+ provider runs per turn. Hosts can build the resolver from a `ProviderRegistry`,
115
+ a plain `AIProvider[]`, or implement `ProviderResolver` directly as a one-line
116
+ function over their own map (lazy construction, per-request routing). Declaring
117
+ a provider grants no permissions and forces no activation; the host always has
118
+ final say. See [Provider layer § Provider resolver](provider-layer.md#provider-resolver)
119
+ for the resolver contract.
63
120
 
64
121
  ## Outputs / response / events
65
122
 
@@ -67,7 +124,7 @@ These workspaces still follow the same rule as external packages: no provider SD
67
124
 
68
125
  ## Request/response example
69
126
 
70
- Provider package manifest contribution and the generic request options a provider request policy can set:
127
+ Provider package manifest contribution and the generic request options a provider request policy can set (see [Provider request policies](provider-request-policies.md) and [Provider caching](provider-caching.md)):
71
128
 
72
129
  ```json
73
130
  {
@@ -83,7 +140,8 @@ Provider package manifest contribution and the generic request options a provide
83
140
  "cacheKey": "demo",
84
141
  "cacheRetention": "short",
85
142
  "headers": { "x-demo": "1" }
86
- }
143
+ },
144
+ "runOptions.retry": { "maxAttempts": 3, "maxDelayMs": 1000 }
87
145
  }
88
146
  ```
89
147
 
@@ -124,6 +182,7 @@ await kernel.load([pkg]);
124
182
  (`provider_request`) that sets generic `ProviderRequest.options`
125
183
  (`sessionId`, `cacheKey`, `cacheRetention`, `headers`, opaque `extra`) before
126
184
  `AIProvider.generate()`; provider adapters map those options to provider payloads.
185
+ - `ModelConfig.cache` is the generic cache capability metadata documented in [Model registry](model-registry.md); `ModelConfig.compat` remains provider-owned inert JSON for behavior that has no generic field yet.
127
186
  - `ModelConfig.compat` is provider-owned inert JSON: cache policy overrides,
128
187
  reasoning/thinking formats, and provider-specific usage mapping live there
129
188
  rather than in core, so Prism never branches on provider names.
@@ -139,6 +198,7 @@ await kernel.load([pkg]);
139
198
  - Registration is in-memory only and does no filesystem, network, env, OAuth refresh, or command access.
140
199
  - Provider-specific behavior belongs in provider packages, not Prism core.
141
200
  - Adapter serializers should preserve Prism content blocks (text, thinking, tool_call, tool_result, and image when the model declares image input) in provider-native request shape, or fail explicitly when a block is unsupported.
201
+ - Adapter header merging must put caller-supplied `ProviderRequest.options.headers` first and provider-owned headers last. Caller headers may add non-owned headers, but cannot replace resolved credentials, content type, session/cache/security headers, or provider attribution headers.
142
202
 
143
203
  ## Manifest declarations
144
204
 
@@ -0,0 +1,113 @@
1
+ # Provider request policies
2
+
3
+ ## What it does
4
+
5
+ Provider request policies are small host/package hooks that can adjust `ProviderRequest.options` before `AIProvider.generate()` runs.
6
+
7
+ Public helpers:
8
+
9
+ - `createProviderRequestPolicyChain(policies)` runs policies in order.
10
+ - `createSessionCachePolicy(options)` sets legacy `cacheKey` / `cacheRetention` aliases from `sessionId`.
11
+ - `mergeProviderRequestOptions(base, patch)` merges request options, including structured `cache` hints.
12
+
13
+ ## When to use it
14
+
15
+ Use provider request policies when an app or provider package needs to set generic per-request options such as cache hints, caller-owned headers, `compat`, or `extra` without changing every provider call site.
16
+
17
+ Do not use request policies to resolve credentials, read env vars, perform OAuth refresh, fetch model lists, or override provider-owned auth/session/security headers.
18
+
19
+ ## Inputs / request
20
+
21
+ ```ts
22
+ import type { ProviderRequestPolicy, ProviderRequestPolicyContext, ProviderRequestOptions } from "@arnilo/prism";
23
+ ```
24
+
25
+ | API | Input | Purpose |
26
+ | --- | --- | --- |
27
+ | `ProviderRequestPolicy.apply(context)` | `{ sessionId?, request, options? }` | Returns a patched request or options. |
28
+ | `createProviderRequestPolicyChain(policies)` | ordered policies | Applies patches in order. |
29
+ | `createSessionCachePolicy({ retention?, cacheKey? })` | optional cache defaults | Sets legacy aliases. |
30
+ | `mergeProviderRequestOptions(base, patch)` | two option bags | Shallow merges scalars and structurally merges `cache`. |
31
+
32
+ `mergeProviderRequestOptions()` behavior:
33
+
34
+ - Patch scalar fields win.
35
+ - `headers`, `compat`, and `extra` shallow-merge.
36
+ - `cache` shallow-merges; patch `mode`, `key`, and `retention` win.
37
+ - `cache.breakpoints` concatenate in base-then-patch order.
38
+ - Legacy-only `cacheKey` / `cacheRetention` merges remain unchanged and do not add a `cache` property.
39
+
40
+ ## Outputs / response / events
41
+
42
+ A policy chain returns either a full `ProviderRequest` or `{ request, options }` style result, normalized by the chain before the next policy runs. The final request is what the agent/session runtime passes to the provider.
43
+
44
+ No agent events are emitted by the policy chain itself.
45
+
46
+ ## Request/response example
47
+
48
+ ```json
49
+ {
50
+ "before": { "options": { "cacheRetention": "short" } },
51
+ "patch": { "options": { "cache": { "key": "stable", "retention": "long" } } },
52
+ "after": {
53
+ "options": {
54
+ "cacheRetention": "short",
55
+ "cache": { "key": "stable", "retention": "long" }
56
+ }
57
+ }
58
+ }
59
+ ```
60
+
61
+ ## Implementation example
62
+
63
+ ```ts
64
+ import {
65
+ createProviderRequestPolicyChain,
66
+ createSessionCachePolicy,
67
+ mergeProviderRequestOptions,
68
+ type ProviderRequestPolicy,
69
+ } from "@arnilo/prism";
70
+
71
+ const structuredCache: ProviderRequestPolicy = {
72
+ name: "demo.structured-cache",
73
+ apply({ request }) {
74
+ return {
75
+ ...request,
76
+ options: mergeProviderRequestOptions(request.options, {
77
+ cache: {
78
+ mode: "on",
79
+ key: request.options?.sessionId,
80
+ retention: "long",
81
+ breakpoints: [{ location: "system_prompt" }],
82
+ },
83
+ }),
84
+ };
85
+ },
86
+ };
87
+
88
+ const chain = createProviderRequestPolicyChain([
89
+ createSessionCachePolicy({ retention: "short" }),
90
+ structuredCache,
91
+ ]);
92
+ ```
93
+
94
+ ## Extension and configuration notes
95
+
96
+ Provider packages can register request policies during `defineProviderPackage().setup(api)`. Hosts decide which packages/policies load and in which order. Prism has no hidden provider request policy registry and no automatic provider-specific cache behavior in core.
97
+
98
+ Policy output should stay generic: use `ProviderRequestOptions.cache`, `headers`, `compat`, and `extra` instead of provider-name branches in core.
99
+
100
+ ## Security and performance notes
101
+
102
+ - Request policies must not store or log credentials.
103
+ - Caller headers are advisory; provider adapters must apply provider-owned auth/session/security headers last.
104
+ - Cache keys must never be credentials.
105
+ - Policy chains are O(number of policies) plus option merge cost.
106
+ - Policies should be pure and synchronous unless the host explicitly accepts async work.
107
+
108
+ ## Related APIs
109
+
110
+ - [Provider caching](provider-caching.md): structured cache hints and helpers.
111
+ - [Provider packages](provider-packages.md): registering policies from extension packages.
112
+ - [Provider layer](provider-layer.md): provider request flow and `AIProvider.generate()`.
113
+ - [Public contracts](public-contracts.md): `ProviderRequestPolicy`, `ProviderRequestOptions`, and cache types.
@@ -90,12 +90,34 @@ await kernel.load([
90
90
  `includeMoonshotModels: true`; it is not core behavior.
91
91
  - Package contributes models via the extension `api` and an `api_key` auth method.
92
92
 
93
+ ### Cache behavior
94
+
95
+ - Default catalog models (e.g. `kimi-k2.7-code` on the Anthropic-compatible
96
+ `/messages` route) use **implicit caching** and send no explicit `cache_control`
97
+ fields. `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention` have no
98
+ effect on the request body unless the model opts in.
99
+ - Hosts may opt a model into Anthropic-style `cache_control` by declaring
100
+ `ModelConfig.cache.kind: "cache_control"` on the Anthropic route. When opted in,
101
+ `cache_control: { type: "ephemeral" }` markers are applied only to the
102
+ caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
103
+ shared `applyCacheControl()` helper) on the last content block of each selected
104
+ message — not to every block. `cacheRetention: "long"` adds `ttl: "1h"` when the
105
+ model allows long retention (`ModelConfig.cache.longRetention !== false`).
106
+ - The Moonshot Open Platform route (`compat.route: "openai"`) never receives
107
+ Anthropic `cache_control` fields.
108
+ - Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
109
+ `Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
110
+ `Usage.cacheWriteTokens`.
111
+
93
112
  ## Security and performance notes
94
113
 
95
114
  - No network calls during import, setup, build, or default tests.
96
115
  - No automatic environment, file, keychain, or shell credential lookup.
97
116
  - Kimi credentials are resolved per request from caller-supplied values or resolvers
98
117
  and redacted from errors.
118
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
119
+ but provider-owned headers (`content-type`, `user-agent`, `authorization`)
120
+ are applied last and cannot be overridden by caller headers.
99
121
  - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
100
122
  provider-specific env names; default tests are network-free.
101
123
 
@@ -0,0 +1,388 @@
1
+ # NeuralWatt provider package
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-neuralwatt` provides explicit, side-effect-free setup for the
6
+ NeuralWatt OpenAI-compatible Chat Completions provider using Prism's OpenAI-compatible
7
+ route with NeuralWatt-specific reasoning/template escape hatches, SSE comment tolerance,
8
+ and implicit prefix caching.
9
+
10
+ The package registers a provider, default model metadata for the featured NeuralWatt
11
+ aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
12
+ `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
13
+ `qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
14
+ method through `createExtensionKernel().load([...])`.
15
+
16
+ ## When to use it
17
+
18
+ Use it when a host app wants to run the NeuralWatt endpoint (`https://api.neuralwatt.com/v1`)
19
+ through Prism's `AgentSession` runtime with NeuralWatt-specific `reasoning_effort`,
20
+ `thinking_token_budget`, and `chat_template_kwargs` handling.
21
+
22
+ Do not use it for automatic credential discovery, catalog fetches, or real-network tests.
23
+
24
+ ## Inputs / request
25
+
26
+ ```ts
27
+ import {
28
+ classifyNeuralWattError,
29
+ createNeuralWattProviderPackage,
30
+ defineNeuralWattModel,
31
+ getNeuralWattQuota,
32
+ listNeuralWattModels,
33
+ mapNeuralWattTelemetry,
34
+ neuralWattEventsWithTelemetry,
35
+ neuralWattModels,
36
+ parseNeuralWattComment,
37
+ } from "@arnilo/prism-provider-neuralwatt";
38
+
39
+ createNeuralWattProviderPackage(options: NeuralWattProviderPackageOptions): ProviderPackage
40
+ defineNeuralWattModel(config: NeuralWattModelConfig): ModelConfig
41
+ listNeuralWattModels(options?: ListNeuralWattModelsOptions): Promise<ModelConfig[]>
42
+ getNeuralWattQuota(options: GetNeuralWattQuotaOptions): Promise<NeuralWattQuota>
43
+ classifyNeuralWattError(input: NeuralWattErrorInput): NeuralWattRetryDecision
44
+ mapNeuralWattTelemetry(body: unknown): { energy?: NeuralWattEnergyTelemetry; cost?: NeuralWattCostTelemetry }
45
+ parseNeuralWattComment(text: string): NeuralWattTelemetryEvent | undefined
46
+ neuralWattEventsWithTelemetry(body: ReadableStream<Uint8Array>): AsyncIterable<NeuralWattEvent>
47
+ ```
48
+
49
+ | Field | Type | Purpose |
50
+ | --- | --- | --- |
51
+ | `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
52
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
53
+ | `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
54
+ | `id` | `string` | Overrides the provider id (default `neuralwatt`). |
55
+ | `models` | `readonly ModelConfig[]` | Overrides `neuralWattModels` defaults. |
56
+
57
+ `listNeuralWattModels()` options:
58
+
59
+ | Field | Type | Purpose |
60
+ | --- | --- | --- |
61
+ | `apiKey` | `CredentialValueSource` | Optional API-key source. Unauthenticated calls return the public catalog; authenticated calls may include private models. |
62
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
63
+ | `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
64
+ | `signal` | `AbortSignal` | Cancels the single discovery request. |
65
+ | `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
66
+
67
+ `getNeuralWattQuota()` options:
68
+
69
+ | Field | Type | Purpose |
70
+ | --- | --- | --- |
71
+ | `apiKey` | `CredentialValueSource` | **Required.** NeuralWatt returns 401 for unauthenticated quota calls. |
72
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
73
+ | `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
74
+ | `signal` | `AbortSignal` | Cancels the single quota request. |
75
+ | `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
76
+
77
+ The endpoint is rate-limited to **1 request per second per customer** (429 with
78
+ `Retry-After: 1`). The helper performs no polling or caching and is never called
79
+ from `generate()` or package setup; the caller owns throttling.
80
+
81
+ NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
82
+ / `extra` escape hatches: `compat.reasoning_effort` (`"low" | "medium" | "high"`),
83
+ `compat.thinking_token_budget`, `compat.chat_template_kwargs` (including `enable_thinking`),
84
+ `compat.preserve_thinking`, `compat.clear_thinking`, and `compat.tool_choice`.
85
+ `preserve_thinking: true` keeps prior assistant reasoning in request history so
86
+ multi-turn reasoning continues with the earlier chain of thought; `clear_thinking:
87
+ true` drops it for the next turn, resetting the chain. `options.extra` spreads after
88
+ `compat` so per-call values and overrides win.
89
+
90
+ ## Outputs / response / events
91
+
92
+ | Surface | Behavior |
93
+ | --- | --- |
94
+ | Provider stream | Prism text, thinking (`delta.reasoning_content` → `providerThinkingDelta`), tool-call delta/final, `usage`, `done`, redacted `error` with HTTP-status `code` for retry classification. |
95
+ | Block preservation | Text, thinking, assistant `tool_call` → `tool_calls`, `tool_result` → role `tool` messages, images when `capabilities.input` includes `"image"`. |
96
+ | Model catalog | Featured aliases declare provider id, display name, context limit, text/image input support, tools, reasoning/fast variants, streaming, implicit cache, and NeuralWatt JSON-mode compat metadata where documented. |
97
+ | Pricing | Static aliases do not guess rates. Exact per-alias input/output/cache-read prices are advertised by NeuralWatt's `/v1/models` response and mapped by `listNeuralWattModels()` when present. |
98
+ | SSE comments | `: energy` / `: cost` comment lines are parsed by `neuralWattEventsWithTelemetry()` into `neuralwatt:telemetry` events; the standard `neuralWattEvents()` stream (used by `generate()`) tolerates them without spurious events. |
99
+ | `[DONE]` | Terminates the stream; final `providerDone(usage)` always emitted on a clean stream. |
100
+ | Malformed data | Yields `providerError` rather than crashing the generator (more robust than the Z.AI parser). |
101
+ | Auth method | `api_key` for the configured provider id, credential name `apiKey`. |
102
+
103
+ Unsupported block placements or unclaimed images fail before fetch.
104
+
105
+ ## Request/response example
106
+
107
+ Example request body (OpenAI-compatible Chat Completions shape):
108
+
109
+ ```json
110
+ {
111
+ "model": "glm-5.2",
112
+ "messages": [{ "role": "user", "content": "Hello" }],
113
+ "stream": true,
114
+ "stream_options": { "include_usage": true },
115
+ "reasoning_effort": "medium",
116
+ "thinking_token_budget": 8192
117
+ }
118
+ ```
119
+
120
+ ## Implementation example
121
+
122
+ ```ts
123
+ import { createExtensionKernel } from "@arnilo/prism";
124
+ import { createNeuralWattProviderPackage } from "@arnilo/prism-provider-neuralwatt";
125
+
126
+ const kernel = createExtensionKernel();
127
+ await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake-neuralwatt-key" })]);
128
+ ```
129
+
130
+ Override the provider id and models:
131
+
132
+ ```ts
133
+ import { createNeuralWattProviderPackage, defineNeuralWattModel, neuralWattModels } from "@arnilo/prism-provider-neuralwatt";
134
+
135
+ await kernel.load([
136
+ createNeuralWattProviderPackage({ id: "neuralwatt", apiKey: "fake", models: neuralWattModels }),
137
+ ]);
138
+ ```
139
+
140
+ Explicit catalog discovery:
141
+
142
+ ```ts
143
+ import { listNeuralWattModels } from "@arnilo/prism-provider-neuralwatt";
144
+
145
+ const models = await listNeuralWattModels({ apiKey: "fake-neuralwatt-key", fetch });
146
+ await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake", models })]);
147
+ ```
148
+
149
+ Discovery performs exactly one `GET /v1/models` call when invoked. Provider package
150
+ setup and `generate()` never call model discovery implicitly.
151
+
152
+ Account quota:
153
+
154
+ ```ts
155
+ import { getNeuralWattQuota } from "@arnilo/prism-provider-neuralwatt";
156
+
157
+ const quota = await getNeuralWattQuota({ apiKey: "fake-neuralwatt-key", fetch });
158
+ console.log(quota.usage?.current_month?.energy_kwh, quota.balance?.balance_usd);
159
+ ```
160
+
161
+ Returns typed `NeuralWattQuota` (`balance`, `usage.lifetime`/`usage.current_month`,
162
+ `limits`, `subscription`, `key`). All fields optional; minimal structural
163
+ validation. The caller owns throttling/caching — the helper makes one explicit
164
+ `GET /v1/quota` call and is never invoked from `generate()` or setup.
165
+
166
+ ## Extension and configuration notes
167
+
168
+ - Hosts choose base URL, provider id, model list, credential source, and `fetch`
169
+ impl.
170
+ - `defineNeuralWattModel` lets apps set NeuralWatt-specific `compat`
171
+ (`reasoning_effort`, `thinking_token_budget`, `chat_template_kwargs`,
172
+ `preserve_thinking`, `clear_thinking`, `tool_choice`).
173
+ - Package contributes models via the extension `api` and an `api_key` auth method.
174
+ - Curated aliases are static and network-free. They include documented context windows
175
+ and capabilities only; `ModelConfig.cost` is left unset until exact per-alias pricing
176
+ is read from NeuralWatt's `/v1/models` catalog.
177
+ - `listNeuralWattModels()` maps `/v1/models` entries to `ModelConfig`: id/display
178
+ name, capabilities, limits, implicit cache metadata, `ModelCost` pricing, and
179
+ provider-owned NeuralWatt metadata in `compat.neuralwatt`.
180
+ - `getNeuralWattQuota()` calls `GET /v1/quota` once with a required API key and
181
+ returns typed account quota (`balance`, `usage`, `limits`, `subscription`, `key`).
182
+ It is opt-in, never called from `generate()` or setup, and the caller owns
183
+ throttling (NeuralWatt limits the endpoint to 1 rps per customer).
184
+
185
+ ### Model catalog and pricing
186
+
187
+ `neuralWattModels` includes NeuralWatt's featured aliases:
188
+
189
+ | Alias | Context | Notable metadata |
190
+ | --- | ---: | --- |
191
+ | `glm-5.2` | 1024K | Tools, reasoning |
192
+ | `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
193
+ | `glm-5.2-short` | 195K | Tools, reasoning |
194
+ | `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
195
+ | `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
196
+ | `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
197
+ | `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
198
+ | `qwen3.5-397b` | 256K | Tools, reasoning, JSON mode |
199
+ | `qwen3.5-397b-fast` | 256K | Tools, JSON mode, fast/no reasoning |
200
+ | `qwen3.6-35b` | 128K | Tools, reasoning, vision, JSON mode |
201
+ | `qwen3.6-35b-fast` | 128K | Tools, vision, JSON mode, fast/no reasoning |
202
+
203
+ NeuralWatt exposes exact pricing from `GET /v1/models` as per-million-token
204
+ `input_per_million`, `output_per_million`, `cached_input_per_million`,
205
+ `cached_output_per_million`, `currency`, and `pricing_tbd`. Cache reads for
206
+ NeuralWatt-hosted models are advertised by the API and default to 25% of the input
207
+ rate; there is no separate cache-write price (`cached_output_per_million` is `null`).
208
+ The static catalog does not copy or infer prices that are not published as fixed
209
+ alias values in these docs.
210
+
211
+ ### Cache behavior
212
+
213
+ - NeuralWatt models use **implicit prefix caching**: the server caches prompt prefixes
214
+ automatically based on request content, with no explicit request-side cache payload.
215
+ Catalog models declare `cache: { kind: "implicit" }`. For the cross-provider
216
+ explicit/implicit cache matrix, see [Provider caching](../provider-caching.md).
217
+ - The provider sends no `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention`
218
+ fields regardless of `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention`
219
+ settings — those options have no effect on the NeuralWatt request body.
220
+ `cacheRetention: "none"` disables Prism cache-control hints only; it does **not**
221
+ disable the implicit backend prefix cache. Hosts relying on cache hits should keep
222
+ their stable prompt prefix byte-stable and stable inputs unchanged.
223
+ - Usage accounting is read-only: `prompt_tokens_details.cached_tokens` maps to
224
+ `Usage.cacheReadTokens`. NeuralWatt does not report a cache-write token today, so
225
+ `Usage.cacheWriteTokens` is never fabricated (stays `undefined`).
226
+
227
+ #### Cache-aware limiter behavior
228
+
229
+ NeuralWatt's backend rate limiter is cache-aware, which affects when long-running
230
+ agent sessions are throttled versus served:
231
+
232
+ - **Uncached TPM counts cold prefill only.** Tokens-per-minute accounting charges the
233
+ cold prefill of a request — the prefix that is not already in the vLLM prefix cache.
234
+ A request whose prefix is fully cached consumes far less of the TPM budget than a
235
+ cold request of the same total prompt length.
236
+ - **Warm-prefix requests can avoid some `503` fleet-capacity blocks.** When the fleet
237
+ is near capacity, requests that can reuse a cached prefix are more likely to be
238
+ admitted than fully cold requests. Prefix reuse is therefore both a latency and an
239
+ availability lever, not just a cost lever.
240
+ - **Full prior history is required for multi-turn cache reuse.** The prefix cache is
241
+ keyed by request content, so each follow-up turn must resend the entire prior
242
+ transcript (system prompt + all prior turns) unchanged, with only the new turn
243
+ appended. Prism's `inputLayout: "cache_aware"` ordering keeps the stable prefix
244
+ first; see [Provider caching](../provider-caching.md).
245
+ - Cache behavior is best-effort and **does not guarantee cache hits**. Prefix cache
246
+ admission and eviction are server-side decisions and can vary with fleet load.
247
+ `cacheRetention: "none"` disables Prism cache-control hints only; it does not
248
+ disable the implicit backend prefix cache.
249
+
250
+ ### Reasoning preservation across turns
251
+
252
+ NeuralWatt reasoning-capable models (Kimi-, GLM-, Qwen-style aliases with
253
+ `capabilities.reasoning: true`) accept prior assistant reasoning in request history
254
+ so multi-turn sessions continue the earlier chain of thought:
255
+
256
+ - Prior `thinking` content blocks on an assistant message are serialized under a
257
+ `reasoning_content` field on that message (matching the streaming
258
+ `delta.reasoning_content` field). They are **not** flattened into text `content`, so
259
+ the model sees reasoning and answer as distinct.
260
+ - Preservation is gated on `model.capabilities.reasoning === true` **or**
261
+ `compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
262
+ field and prior `thinking` blocks are dropped — they never leak into text content for
263
+ providers/models that do not support reasoning.
264
+ - `compat.clear_thinking: true` drops prior reasoning for the next turn even on
265
+ reasoning-capable models, resetting the chain of thought. `clear_thinking` takes
266
+ precedence over `preserve_thinking`.
267
+ - The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
268
+ reasoning.
269
+
270
+ ### Tool calls and the tool-call loop
271
+
272
+ NeuralWatt exposes OpenAI-compatible function calling. The provider carries tools and
273
+ prior tool turns through a multi-turn loop:
274
+
275
+ - **Request serialization.** `ProviderRequest.tools` (`ToolDefinition[]`) is serialized
276
+ to OpenAI `tools: [{ type: "function", function: { name, description, parameters } }]`.
277
+ Missing `parameters` default to `{ type: "object" }`. `compat.tool_choice` passes
278
+ through as `tool_choice` (string or `{ type: "function", function: { name } }`).
279
+ - **Streaming reconstruction.** `delta.tool_calls` fragments (keyed by `index`) are
280
+ accumulated and re-emitted as `tool_call_delta` events for UI consumers, then
281
+ reconstructed into a final `tool_call` event per call with parsed JSON arguments.
282
+ Parallel calls are tracked by index.
283
+ - **Next-turn ordering.** On the following turn the assistant `tool_call` block is
284
+ serialized to `role: "assistant"` with a `tool_calls` array (arguments stringified to
285
+ JSON), immediately followed by a `role: "tool"` message carrying `tool_call_id` and
286
+ the stringified `tool_result` — matching the OpenAI requirement that a tool result
287
+ follows the call that produced it. `tool_result` blocks must appear in `role: "tool"
288
+ messages; `tool_call` blocks must be the only content on their assistant message.
289
+
290
+ ### Energy and cost telemetry
291
+
292
+ NeuralWatt streams energy and cost data as SSE comment lines (`: energy {...}`
293
+ and `: cost {...}`) before `data: [DONE]`, and as top-level `energy`/`cost` JSON
294
+ fields on non-streaming responses. Standard SSE clients ignore comments, so
295
+ these values are invisible unless the raw stream is parsed.
296
+
297
+ Prism's core `ProviderEvent` union has no generic telemetry event, so NeuralWatt
298
+ exposes telemetry through package-specific helpers:
299
+
300
+ - `neuralWattEventsWithTelemetry(body)` yields the standard provider events plus
301
+ `neuralwatt:telemetry` events (`{ type: "neuralwatt:telemetry", energy?, cost? }`)
302
+ in stream order. Use it when a host wants to observe telemetry alongside text,
303
+ tool, usage, and done events.
304
+ - `parseNeuralWattComment(text)` parses a single `: energy`/`: cost` comment line
305
+ into a `NeuralWattTelemetryEvent` (`undefined` for unknown/malformed comments).
306
+ - `parseNeuralWattEnergy(payload)` / `parseNeuralWattCost(payload)` parse the JSON
307
+ payload of a single comment into typed `NeuralWattEnergyTelemetry` /
308
+ `NeuralWattCostTelemetry`.
309
+ - `mapNeuralWattTelemetry(body)` maps a non-streaming response body's top-level
310
+ `energy`/`cost` fields into the same typed telemetry.
311
+
312
+ `generate()` stays streaming-only and uses `neuralWattEvents()`, so telemetry is
313
+ opt-in via `neuralWattEventsWithTelemetry()`. Telemetry contains usage/cost
314
+ numbers only — never prompts, API keys, or headers. All documented fields are
315
+ optional and tolerated when absent; malformed comments yield no telemetry event
316
+ and never crash the stream.
317
+
318
+ ```ts
319
+ import { neuralWattEventsWithTelemetry } from "@arnilo/prism-provider-neuralwatt";
320
+
321
+ for await (const event of neuralWattEventsWithTelemetry(response.body)) {
322
+ if (event.type === "neuralwatt:telemetry") {
323
+ console.log(event.energy?.energy_kwh, event.cost?.request_cost_usd);
324
+ }
325
+ }
326
+ ```
327
+
328
+ ### Retry classification
329
+
330
+ NeuralWatt error responses are classified by `classifyNeuralWattError()` so the
331
+ Prism runtime retry policy can decide retryability without provider-specific
332
+ core branches:
333
+
334
+ | Status | Retryable | Notes |
335
+ | --- | --- | --- |
336
+ | `400` `401` `402` `403` `404` | no | Client/payment/auth errors fail closed. |
337
+ | `429` | yes | Reads `Retry-After` header and `error.retry_after`; preserves `error.retry_strategy` (`type`, `suggested_initial_delay_s`, `max_delay_s`, `backoff`, `jitter`). |
338
+ | `500` `502` `503` | yes | Transient server/fleet-capacity errors; `503` `Retry-After` honored when present. |
339
+
340
+ The provider emits `providerError` with `ErrorInfo.code` set to the numeric HTTP
341
+ status. Prism's default retry policy (`createDefaultRetryPolicy()`) treats `429`/
342
+ `500`/`502`/`503` as transient and `400`/`401`/`402`/`403`/`404` as non-transient,
343
+ so NeuralWatt errors retry correctly out of the box. `classifyNeuralWattError()`
344
+ and `neuralWattHttpError()` are exported for hosts/tests that want structured
345
+ retry metadata (`retryAfterMs`, `errorCode`, `strategy`). The host retry policy
346
+ owns the exact delay; `retryAfterMs` is surfaced but not enforced by the
347
+ provider. Classification is O(1) over status/headers/body and makes no extra
348
+ provider calls.
349
+
350
+ ```ts
351
+ import { classifyNeuralWattError } from "@arnilo/prism-provider-neuralwatt";
352
+
353
+ const decision = classifyNeuralWattError({ status: 429, headers: { "retry-after": "1" }, body: { error: { code: "concurrent_budget_exceeded", retry_after: 1 } } });
354
+ // { retryable: true, code: 429, retryAfterMs: 1000, errorCode: "concurrent_budget_exceeded", strategy: undefined }
355
+ ```
356
+
357
+ ## Security and performance notes
358
+
359
+ - No network calls during import, setup, build, default tests, or generation beyond
360
+ the explicit Chat Completions request. `listNeuralWattModels()` is opt-in and
361
+ makes one `GET /v1/models` call per invocation; `getNeuralWattQuota()` is opt-in
362
+ and makes one `GET /v1/quota` call per invocation (endpoint limited to 1 rps per
363
+ customer; caller owns throttling).
364
+ - No automatic environment, file, keychain, or shell credential lookup.
365
+ - API keys are resolved per request/helper call from caller-supplied values or resolvers
366
+ and redacted from errors via `redactSecrets`. `listNeuralWattModels()` and
367
+ `getNeuralWattQuota()` apply provider-owned `authorization` after caller headers so
368
+ callers cannot override it. Quota values never enter provider events unless the caller
369
+ emits them.
370
+ - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
371
+ but provider-owned headers (`content-type`, `authorization`) are applied last
372
+ and cannot be overridden by caller headers.
373
+ - Live tests stay opt-in behind `NEURALWATT_API_KEY` (plus `PRISM_LIVE_PROVIDER_TESTS=1`);
374
+ default tests are network-free.
375
+
376
+ ## Related APIs
377
+
378
+ - [Provider packages](../provider-packages.md): `defineProviderPackage`,
379
+ `ModelConfig`/`compat`, thinking formats.
380
+ - [Credentials and redaction](../credentials-and-redaction.md):
381
+ `resolveCredentialValue`, `redactSecrets`.
382
+ - [OpenAI-compatible provider](openai-compatible.md): underlying Chat Completions
383
+ adapter.
384
+ - [Provider conformance](../provider-conformance.md): network-free adapter tests.
385
+ - [Provider caching](../provider-caching.md): implicit cache behavior and
386
+ `cacheUsageReport`.
387
+ - [NeuralWatt agent example](../../examples/neuralwatt-agent-run.ts): runnable mocked
388
+ agent turn with tools, reasoning controls, streamed cache tokens, and energy/cost telemetry.