@arnilo/prism 0.0.1 → 0.0.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +4 -2
- package/README.md +17 -7
- package/dist/agent-definitions.d.ts +12 -0
- package/dist/agent-definitions.js +131 -0
- package/dist/agent-loops.d.ts +14 -0
- package/dist/agent-loops.js +161 -0
- package/dist/agents.js +263 -76
- package/dist/cache-helpers.d.ts +28 -0
- package/dist/cache-helpers.js +73 -0
- package/dist/cli-runner.d.ts +38 -2
- package/dist/cli-runner.js +167 -5
- package/dist/compaction.js +2 -0
- package/dist/config.js +47 -12
- package/dist/contracts.d.ts +581 -6
- package/dist/contracts.js +41 -1
- package/dist/contribution-parsing.d.ts +19 -0
- package/dist/contribution-parsing.js +124 -0
- package/dist/contributions.d.ts +13 -3
- package/dist/contributions.js +96 -20
- package/dist/extensions.js +3 -0
- package/dist/index.d.ts +19 -9
- package/dist/index.js +10 -4
- package/dist/input.d.ts +7 -1
- package/dist/input.js +52 -11
- package/dist/instruction-injection.d.ts +28 -0
- package/dist/instruction-injection.js +55 -0
- package/dist/manifests.d.ts +1 -1
- package/dist/manifests.js +3 -3
- package/dist/models.d.ts +4 -1
- package/dist/models.js +5 -2
- package/dist/node/agent-definitions.d.ts +98 -0
- package/dist/node/agent-definitions.js +389 -0
- package/dist/node/contribution-discovery.d.ts +17 -0
- package/dist/node/contribution-discovery.js +163 -0
- package/dist/node/instruction-injectors.d.ts +32 -0
- package/dist/node/instruction-injectors.js +72 -0
- package/dist/node/session-store-jsonl.d.ts +1 -1
- package/dist/node/session-store-jsonl.js +42 -4
- package/dist/node/system-project-prompts.d.ts +30 -0
- package/dist/node/system-project-prompts.js +53 -0
- package/dist/provider-events.d.ts +3 -1
- package/dist/provider-events.js +34 -0
- package/dist/provider-request-policy.js +15 -1
- package/dist/providers/openai-compatible.js +1 -1
- package/dist/providers.d.ts +6 -2
- package/dist/providers.js +15 -1
- package/dist/redaction.d.ts +2 -1
- package/dist/redaction.js +3 -0
- package/dist/registry-options.d.ts +5 -0
- package/dist/registry-options.js +5 -0
- package/dist/rpc.d.ts +6 -2
- package/dist/rpc.js +71 -13
- package/dist/session-stores.d.ts +3 -1
- package/dist/session-stores.js +67 -6
- package/dist/skills.d.ts +4 -1
- package/dist/skills.js +3 -1
- package/dist/system-prompts.js +6 -2
- package/dist/testing/compaction-conformance.d.ts +17 -0
- package/dist/testing/compaction-conformance.js +61 -0
- package/dist/testing/extension-conformance.d.ts +26 -0
- package/dist/testing/extension-conformance.js +55 -0
- package/dist/testing/provider-conformance.d.ts +7 -0
- package/dist/testing/provider-conformance.js +18 -31
- package/dist/testing/session-store-conformance.d.ts +20 -0
- package/dist/testing/session-store-conformance.js +92 -0
- package/dist/testing/tool-conformance.d.ts +39 -0
- package/dist/testing/tool-conformance.js +79 -0
- package/dist/tools.d.ts +7 -2
- package/dist/tools.js +50 -13
- package/docs/agent-definitions.md +251 -0
- package/docs/agent-events.md +199 -0
- package/docs/agent-loops.md +217 -0
- package/docs/agent-session-runtime.md +20 -8
- package/docs/cli-rpc.md +39 -4
- package/docs/compaction-and-retry.md +2 -2
- package/docs/compaction-conformance.md +76 -0
- package/docs/compaction-llm.md +6 -3
- package/docs/compaction-observational-memory.md +4 -4
- package/docs/configuration-and-manifests.md +6 -1
- package/docs/context-and-skills.md +79 -6
- package/docs/contribution-discovery.md +149 -0
- package/docs/contribution-registries.md +9 -6
- package/docs/credentials-and-redaction.md +2 -0
- package/docs/customization.md +191 -0
- package/docs/database-persistence.md +407 -0
- package/docs/extension-authoring.md +193 -0
- package/docs/extension-conformance.md +80 -0
- package/docs/extensions.md +6 -0
- package/docs/host-security.md +141 -0
- package/docs/index.md +40 -19
- package/docs/input-and-prompt-assembly.md +19 -3
- package/docs/instruction-injection.md +183 -0
- package/docs/migration.md +201 -0
- package/docs/model-registry.md +122 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/performance.md +127 -0
- package/docs/provider-caching.md +206 -0
- package/docs/provider-conformance.md +32 -5
- package/docs/provider-layer.md +51 -11
- package/docs/provider-packages.md +65 -5
- package/docs/provider-request-policies.md +113 -0
- package/docs/providers/kimi.md +22 -0
- package/docs/providers/neuralwatt.md +388 -0
- package/docs/providers/openai-compatible.md +1 -0
- package/docs/providers/openai.md +21 -0
- package/docs/providers/opencode-go.md +31 -3
- package/docs/providers/openrouter.md +29 -0
- package/docs/providers/zai.md +17 -0
- package/docs/public-contracts.md +87 -12
- package/docs/release-and-install.md +76 -26
- package/docs/runs-and-usage.md +236 -0
- package/docs/session-store-conformance.md +78 -0
- package/docs/session-stores-and-branching.md +10 -6
- package/docs/session-stores.md +126 -0
- package/docs/settings-auth-trust-security.md +18 -4
- package/docs/structured-output.md +247 -0
- package/docs/system-prompts.md +104 -2
- package/docs/tool-conformance.md +87 -0
- package/docs/tools.md +64 -8
- package/package.json +35 -2
|
@@ -53,13 +53,70 @@ api.registerProviderRequestPolicy(createSessionCachePolicy({ retention: "short"
|
|
|
53
53
|
api.registerSystemPromptContribution({ id: "demo-prompt", source: "package", mode: "append", text: "Use demo provider rules." });
|
|
54
54
|
```
|
|
55
55
|
|
|
56
|
-
Hosts decide which credential resolvers, env objects, OAuth stores, request policies, and prompt contributions become active. Request policies can set generic `ProviderRequest.options` such as `sessionId`, `cacheRetention`, `headers`,
|
|
56
|
+
Hosts decide which credential resolvers, env objects, OAuth stores, request policies, and prompt contributions become active. Request policies can set generic `ProviderRequest.options` such as `sessionId`, `cacheRetention`, `headers`, `compat`, and opaque `extra`; provider adapters decide how to map those options to provider payloads. Caller headers are extension headers only: provider adapters must apply provider-owned headers (auth, content type, session/cache/security, attribution) after caller headers so requests cannot override credentials or provider policy.
|
|
57
|
+
|
|
58
|
+
Deprecated provider request options: `timeoutMs`, `maxRetries`, and `maxRetryDelayMs` are inert in first-party providers. Use `RunOptions.signal`/host abort controllers for timeouts and `AgentConfig.retry`/`RunOptions.retry` for retry. Provider packages should not add provider-specific retry loops unless the vendor protocol requires it and runtime retry cannot cover the failure mode.
|
|
59
|
+
|
|
60
|
+
First-party providers map generic `ModelConfig.parameters.maxTokens` to real output-token request fields instead of sending `maxTokens` on the wire: OpenAI Responses uses `max_output_tokens`; OpenRouter, OpenCode Go OpenAI-compatible, OpenCode Go Anthropic-style, Z.AI, Kimi, and NeuralWatt use `max_tokens`. Other `model.parameters` values pass through unchanged unless the provider docs say otherwise.
|
|
57
61
|
|
|
58
62
|
## First-party provider package skeletons
|
|
59
63
|
|
|
60
|
-
Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md),
|
|
64
|
+
Phase 12 adds explicit npm workspaces for [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md), [`@arnilo/prism-provider-zai`](providers/zai.md), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md). Each package starts with a side-effect-free `create*ProviderPackage()` export, README, TypeScript build, network-free default tests, and real opt-in live smoke tests.
|
|
65
|
+
|
|
66
|
+
Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
|
|
67
|
+
|
|
68
|
+
These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
|
|
69
|
+
|
|
70
|
+
### First-party cache behavior
|
|
71
|
+
|
|
72
|
+
Every first-party provider package hardens prompt-cache behavior so it cannot emit invalid cache retention values or over-broad cache-control markers, and so provider-owned `authorization`/session/security headers cannot be overridden by caller `ProviderRequest.options.headers`. Cache behavior is provider-specific and best-effort: OpenAI/OpenRouter use explicit hints, NeuralWatt/Z.AI use implicit caching, and OpenCode Go/Kimi are route/model-dependent. See [Provider caching](provider-caching.md#per-provider-cache-behavior) for the canonical explicit/implicit matrix.
|
|
73
|
+
|
|
74
|
+
- **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
75
|
+
- **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
76
|
+
- **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; `cache_control` markers applied only to caller-selected `cache.breakpoints` (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
|
|
77
|
+
- **OpenCode Go**: `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; the Anthropic route applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`), the OpenAI route sends none. Per-route usage mapping.
|
|
78
|
+
- **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
|
|
79
|
+
- **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
|
|
80
|
+
- **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
|
|
81
|
+
|
|
82
|
+
See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
|
|
83
|
+
|
|
84
|
+
## Third-party provider packaging
|
|
85
|
+
|
|
86
|
+
A third party ships their own providers the same way Prism ships first-party
|
|
87
|
+
provider packages: an `Extension` whose `setup(api)` calls
|
|
88
|
+
`api.registerProvider(provider)` for each provider it owns. First-party
|
|
89
|
+
provider packages (`@arnilo/prism-provider-openai`, `@arnilo/prism-provider-openrouter`,
|
|
90
|
+
`@arnilo/prism-provider-kimi`, `@arnilo/prism-provider-zai`,
|
|
91
|
+
`@arnilo/prism-provider-opencode-go`) are **opt-in and individually installable**;
|
|
92
|
+
`@arnilo/prism` core runs without any first-party provider package (mock-only).
|
|
61
93
|
|
|
62
|
-
|
|
94
|
+
A host mixes first-party packages and third-party providers in one resolver.
|
|
95
|
+
The host owns the resolver — declaring a provider does not activate it:
|
|
96
|
+
|
|
97
|
+
```ts
|
|
98
|
+
import { createExtensionKernel, createProviderResolver, createAgent } from "@arnilo/prism";
|
|
99
|
+
import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
|
|
100
|
+
|
|
101
|
+
// First-party package, inert until loaded.
|
|
102
|
+
const kernel = createExtensionKernel();
|
|
103
|
+
await kernel.load([createOpenAIProviderPackage({ apiKey: () => process.env.OPENAI_API_KEY })]);
|
|
104
|
+
|
|
105
|
+
// Third-party own provider (bring your own adapter). Combined with first-party
|
|
106
|
+
// providers in one resolver passed to the agent as `providerSource`.
|
|
107
|
+
const own = createMyProvider(/* credentials */);
|
|
108
|
+
const providerSource = createProviderResolver([...kernel.registries.providers.list(), own]);
|
|
109
|
+
|
|
110
|
+
const agent = createAgent({ model: { provider: own.id, model: "demo" }, providerSource });
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
The resolver is the selection mechanism: `model.provider` selects which
|
|
114
|
+
provider runs per turn. Hosts can build the resolver from a `ProviderRegistry`,
|
|
115
|
+
a plain `AIProvider[]`, or implement `ProviderResolver` directly as a one-line
|
|
116
|
+
function over their own map (lazy construction, per-request routing). Declaring
|
|
117
|
+
a provider grants no permissions and forces no activation; the host always has
|
|
118
|
+
final say. See [Provider layer § Provider resolver](provider-layer.md#provider-resolver)
|
|
119
|
+
for the resolver contract.
|
|
63
120
|
|
|
64
121
|
## Outputs / response / events
|
|
65
122
|
|
|
@@ -67,7 +124,7 @@ These workspaces still follow the same rule as external packages: no provider SD
|
|
|
67
124
|
|
|
68
125
|
## Request/response example
|
|
69
126
|
|
|
70
|
-
Provider package manifest contribution and the generic request options a provider request policy can set:
|
|
127
|
+
Provider package manifest contribution and the generic request options a provider request policy can set (see [Provider request policies](provider-request-policies.md) and [Provider caching](provider-caching.md)):
|
|
71
128
|
|
|
72
129
|
```json
|
|
73
130
|
{
|
|
@@ -83,7 +140,8 @@ Provider package manifest contribution and the generic request options a provide
|
|
|
83
140
|
"cacheKey": "demo",
|
|
84
141
|
"cacheRetention": "short",
|
|
85
142
|
"headers": { "x-demo": "1" }
|
|
86
|
-
}
|
|
143
|
+
},
|
|
144
|
+
"runOptions.retry": { "maxAttempts": 3, "maxDelayMs": 1000 }
|
|
87
145
|
}
|
|
88
146
|
```
|
|
89
147
|
|
|
@@ -124,6 +182,7 @@ await kernel.load([pkg]);
|
|
|
124
182
|
(`provider_request`) that sets generic `ProviderRequest.options`
|
|
125
183
|
(`sessionId`, `cacheKey`, `cacheRetention`, `headers`, opaque `extra`) before
|
|
126
184
|
`AIProvider.generate()`; provider adapters map those options to provider payloads.
|
|
185
|
+
- `ModelConfig.cache` is the generic cache capability metadata documented in [Model registry](model-registry.md); `ModelConfig.compat` remains provider-owned inert JSON for behavior that has no generic field yet.
|
|
127
186
|
- `ModelConfig.compat` is provider-owned inert JSON: cache policy overrides,
|
|
128
187
|
reasoning/thinking formats, and provider-specific usage mapping live there
|
|
129
188
|
rather than in core, so Prism never branches on provider names.
|
|
@@ -139,6 +198,7 @@ await kernel.load([pkg]);
|
|
|
139
198
|
- Registration is in-memory only and does no filesystem, network, env, OAuth refresh, or command access.
|
|
140
199
|
- Provider-specific behavior belongs in provider packages, not Prism core.
|
|
141
200
|
- Adapter serializers should preserve Prism content blocks (text, thinking, tool_call, tool_result, and image when the model declares image input) in provider-native request shape, or fail explicitly when a block is unsupported.
|
|
201
|
+
- Adapter header merging must put caller-supplied `ProviderRequest.options.headers` first and provider-owned headers last. Caller headers may add non-owned headers, but cannot replace resolved credentials, content type, session/cache/security headers, or provider attribution headers.
|
|
142
202
|
|
|
143
203
|
## Manifest declarations
|
|
144
204
|
|
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
# Provider request policies
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Provider request policies are small host/package hooks that can adjust `ProviderRequest.options` before `AIProvider.generate()` runs.
|
|
6
|
+
|
|
7
|
+
Public helpers:
|
|
8
|
+
|
|
9
|
+
- `createProviderRequestPolicyChain(policies)` runs policies in order.
|
|
10
|
+
- `createSessionCachePolicy(options)` sets legacy `cacheKey` / `cacheRetention` aliases from `sessionId`.
|
|
11
|
+
- `mergeProviderRequestOptions(base, patch)` merges request options, including structured `cache` hints.
|
|
12
|
+
|
|
13
|
+
## When to use it
|
|
14
|
+
|
|
15
|
+
Use provider request policies when an app or provider package needs to set generic per-request options such as cache hints, caller-owned headers, `compat`, or `extra` without changing every provider call site.
|
|
16
|
+
|
|
17
|
+
Do not use request policies to resolve credentials, read env vars, perform OAuth refresh, fetch model lists, or override provider-owned auth/session/security headers.
|
|
18
|
+
|
|
19
|
+
## Inputs / request
|
|
20
|
+
|
|
21
|
+
```ts
|
|
22
|
+
import type { ProviderRequestPolicy, ProviderRequestPolicyContext, ProviderRequestOptions } from "@arnilo/prism";
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
| API | Input | Purpose |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| `ProviderRequestPolicy.apply(context)` | `{ sessionId?, request, options? }` | Returns a patched request or options. |
|
|
28
|
+
| `createProviderRequestPolicyChain(policies)` | ordered policies | Applies patches in order. |
|
|
29
|
+
| `createSessionCachePolicy({ retention?, cacheKey? })` | optional cache defaults | Sets legacy aliases. |
|
|
30
|
+
| `mergeProviderRequestOptions(base, patch)` | two option bags | Shallow merges scalars and structurally merges `cache`. |
|
|
31
|
+
|
|
32
|
+
`mergeProviderRequestOptions()` behavior:
|
|
33
|
+
|
|
34
|
+
- Patch scalar fields win.
|
|
35
|
+
- `headers`, `compat`, and `extra` shallow-merge.
|
|
36
|
+
- `cache` shallow-merges; patch `mode`, `key`, and `retention` win.
|
|
37
|
+
- `cache.breakpoints` concatenate in base-then-patch order.
|
|
38
|
+
- Legacy-only `cacheKey` / `cacheRetention` merges remain unchanged and do not add a `cache` property.
|
|
39
|
+
|
|
40
|
+
## Outputs / response / events
|
|
41
|
+
|
|
42
|
+
A policy chain returns either a full `ProviderRequest` or `{ request, options }` style result, normalized by the chain before the next policy runs. The final request is what the agent/session runtime passes to the provider.
|
|
43
|
+
|
|
44
|
+
No agent events are emitted by the policy chain itself.
|
|
45
|
+
|
|
46
|
+
## Request/response example
|
|
47
|
+
|
|
48
|
+
```json
|
|
49
|
+
{
|
|
50
|
+
"before": { "options": { "cacheRetention": "short" } },
|
|
51
|
+
"patch": { "options": { "cache": { "key": "stable", "retention": "long" } } },
|
|
52
|
+
"after": {
|
|
53
|
+
"options": {
|
|
54
|
+
"cacheRetention": "short",
|
|
55
|
+
"cache": { "key": "stable", "retention": "long" }
|
|
56
|
+
}
|
|
57
|
+
}
|
|
58
|
+
}
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
## Implementation example
|
|
62
|
+
|
|
63
|
+
```ts
|
|
64
|
+
import {
|
|
65
|
+
createProviderRequestPolicyChain,
|
|
66
|
+
createSessionCachePolicy,
|
|
67
|
+
mergeProviderRequestOptions,
|
|
68
|
+
type ProviderRequestPolicy,
|
|
69
|
+
} from "@arnilo/prism";
|
|
70
|
+
|
|
71
|
+
const structuredCache: ProviderRequestPolicy = {
|
|
72
|
+
name: "demo.structured-cache",
|
|
73
|
+
apply({ request }) {
|
|
74
|
+
return {
|
|
75
|
+
...request,
|
|
76
|
+
options: mergeProviderRequestOptions(request.options, {
|
|
77
|
+
cache: {
|
|
78
|
+
mode: "on",
|
|
79
|
+
key: request.options?.sessionId,
|
|
80
|
+
retention: "long",
|
|
81
|
+
breakpoints: [{ location: "system_prompt" }],
|
|
82
|
+
},
|
|
83
|
+
}),
|
|
84
|
+
};
|
|
85
|
+
},
|
|
86
|
+
};
|
|
87
|
+
|
|
88
|
+
const chain = createProviderRequestPolicyChain([
|
|
89
|
+
createSessionCachePolicy({ retention: "short" }),
|
|
90
|
+
structuredCache,
|
|
91
|
+
]);
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
## Extension and configuration notes
|
|
95
|
+
|
|
96
|
+
Provider packages can register request policies during `defineProviderPackage().setup(api)`. Hosts decide which packages/policies load and in which order. Prism has no hidden provider request policy registry and no automatic provider-specific cache behavior in core.
|
|
97
|
+
|
|
98
|
+
Policy output should stay generic: use `ProviderRequestOptions.cache`, `headers`, `compat`, and `extra` instead of provider-name branches in core.
|
|
99
|
+
|
|
100
|
+
## Security and performance notes
|
|
101
|
+
|
|
102
|
+
- Request policies must not store or log credentials.
|
|
103
|
+
- Caller headers are advisory; provider adapters must apply provider-owned auth/session/security headers last.
|
|
104
|
+
- Cache keys must never be credentials.
|
|
105
|
+
- Policy chains are O(number of policies) plus option merge cost.
|
|
106
|
+
- Policies should be pure and synchronous unless the host explicitly accepts async work.
|
|
107
|
+
|
|
108
|
+
## Related APIs
|
|
109
|
+
|
|
110
|
+
- [Provider caching](provider-caching.md): structured cache hints and helpers.
|
|
111
|
+
- [Provider packages](provider-packages.md): registering policies from extension packages.
|
|
112
|
+
- [Provider layer](provider-layer.md): provider request flow and `AIProvider.generate()`.
|
|
113
|
+
- [Public contracts](public-contracts.md): `ProviderRequestPolicy`, `ProviderRequestOptions`, and cache types.
|
package/docs/providers/kimi.md
CHANGED
|
@@ -90,12 +90,34 @@ await kernel.load([
|
|
|
90
90
|
`includeMoonshotModels: true`; it is not core behavior.
|
|
91
91
|
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
92
92
|
|
|
93
|
+
### Cache behavior
|
|
94
|
+
|
|
95
|
+
- Default catalog models (e.g. `kimi-k2.7-code` on the Anthropic-compatible
|
|
96
|
+
`/messages` route) use **implicit caching** and send no explicit `cache_control`
|
|
97
|
+
fields. `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention` have no
|
|
98
|
+
effect on the request body unless the model opts in.
|
|
99
|
+
- Hosts may opt a model into Anthropic-style `cache_control` by declaring
|
|
100
|
+
`ModelConfig.cache.kind: "cache_control"` on the Anthropic route. When opted in,
|
|
101
|
+
`cache_control: { type: "ephemeral" }` markers are applied only to the
|
|
102
|
+
caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
|
|
103
|
+
shared `applyCacheControl()` helper) on the last content block of each selected
|
|
104
|
+
message — not to every block. `cacheRetention: "long"` adds `ttl: "1h"` when the
|
|
105
|
+
model allows long retention (`ModelConfig.cache.longRetention !== false`).
|
|
106
|
+
- The Moonshot Open Platform route (`compat.route: "openai"`) never receives
|
|
107
|
+
Anthropic `cache_control` fields.
|
|
108
|
+
- Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
|
|
109
|
+
`Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
|
|
110
|
+
`Usage.cacheWriteTokens`.
|
|
111
|
+
|
|
93
112
|
## Security and performance notes
|
|
94
113
|
|
|
95
114
|
- No network calls during import, setup, build, or default tests.
|
|
96
115
|
- No automatic environment, file, keychain, or shell credential lookup.
|
|
97
116
|
- Kimi credentials are resolved per request from caller-supplied values or resolvers
|
|
98
117
|
and redacted from errors.
|
|
118
|
+
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
|
|
119
|
+
but provider-owned headers (`content-type`, `user-agent`, `authorization`)
|
|
120
|
+
are applied last and cannot be overridden by caller headers.
|
|
99
121
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
|
|
100
122
|
provider-specific env names; default tests are network-free.
|
|
101
123
|
|
|
@@ -0,0 +1,388 @@
|
|
|
1
|
+
# NeuralWatt provider package
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-provider-neuralwatt` provides explicit, side-effect-free setup for the
|
|
6
|
+
NeuralWatt OpenAI-compatible Chat Completions provider using Prism's OpenAI-compatible
|
|
7
|
+
route with NeuralWatt-specific reasoning/template escape hatches, SSE comment tolerance,
|
|
8
|
+
and implicit prefix caching.
|
|
9
|
+
|
|
10
|
+
The package registers a provider, default model metadata for the featured NeuralWatt
|
|
11
|
+
aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
|
|
12
|
+
`kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
|
|
13
|
+
`qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
|
|
14
|
+
method through `createExtensionKernel().load([...])`.
|
|
15
|
+
|
|
16
|
+
## When to use it
|
|
17
|
+
|
|
18
|
+
Use it when a host app wants to run the NeuralWatt endpoint (`https://api.neuralwatt.com/v1`)
|
|
19
|
+
through Prism's `AgentSession` runtime with NeuralWatt-specific `reasoning_effort`,
|
|
20
|
+
`thinking_token_budget`, and `chat_template_kwargs` handling.
|
|
21
|
+
|
|
22
|
+
Do not use it for automatic credential discovery, catalog fetches, or real-network tests.
|
|
23
|
+
|
|
24
|
+
## Inputs / request
|
|
25
|
+
|
|
26
|
+
```ts
|
|
27
|
+
import {
|
|
28
|
+
classifyNeuralWattError,
|
|
29
|
+
createNeuralWattProviderPackage,
|
|
30
|
+
defineNeuralWattModel,
|
|
31
|
+
getNeuralWattQuota,
|
|
32
|
+
listNeuralWattModels,
|
|
33
|
+
mapNeuralWattTelemetry,
|
|
34
|
+
neuralWattEventsWithTelemetry,
|
|
35
|
+
neuralWattModels,
|
|
36
|
+
parseNeuralWattComment,
|
|
37
|
+
} from "@arnilo/prism-provider-neuralwatt";
|
|
38
|
+
|
|
39
|
+
createNeuralWattProviderPackage(options: NeuralWattProviderPackageOptions): ProviderPackage
|
|
40
|
+
defineNeuralWattModel(config: NeuralWattModelConfig): ModelConfig
|
|
41
|
+
listNeuralWattModels(options?: ListNeuralWattModelsOptions): Promise<ModelConfig[]>
|
|
42
|
+
getNeuralWattQuota(options: GetNeuralWattQuotaOptions): Promise<NeuralWattQuota>
|
|
43
|
+
classifyNeuralWattError(input: NeuralWattErrorInput): NeuralWattRetryDecision
|
|
44
|
+
mapNeuralWattTelemetry(body: unknown): { energy?: NeuralWattEnergyTelemetry; cost?: NeuralWattCostTelemetry }
|
|
45
|
+
parseNeuralWattComment(text: string): NeuralWattTelemetryEvent | undefined
|
|
46
|
+
neuralWattEventsWithTelemetry(body: ReadableStream<Uint8Array>): AsyncIterable<NeuralWattEvent>
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
| Field | Type | Purpose |
|
|
50
|
+
| --- | --- | --- |
|
|
51
|
+
| `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
|
|
52
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
53
|
+
| `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
|
|
54
|
+
| `id` | `string` | Overrides the provider id (default `neuralwatt`). |
|
|
55
|
+
| `models` | `readonly ModelConfig[]` | Overrides `neuralWattModels` defaults. |
|
|
56
|
+
|
|
57
|
+
`listNeuralWattModels()` options:
|
|
58
|
+
|
|
59
|
+
| Field | Type | Purpose |
|
|
60
|
+
| --- | --- | --- |
|
|
61
|
+
| `apiKey` | `CredentialValueSource` | Optional API-key source. Unauthenticated calls return the public catalog; authenticated calls may include private models. |
|
|
62
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
63
|
+
| `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
|
|
64
|
+
| `signal` | `AbortSignal` | Cancels the single discovery request. |
|
|
65
|
+
| `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
|
|
66
|
+
|
|
67
|
+
`getNeuralWattQuota()` options:
|
|
68
|
+
|
|
69
|
+
| Field | Type | Purpose |
|
|
70
|
+
| --- | --- | --- |
|
|
71
|
+
| `apiKey` | `CredentialValueSource` | **Required.** NeuralWatt returns 401 for unauthenticated quota calls. |
|
|
72
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
73
|
+
| `baseUrl` | `string` | Overrides `https://api.neuralwatt.com/v1`. |
|
|
74
|
+
| `signal` | `AbortSignal` | Cancels the single quota request. |
|
|
75
|
+
| `headers` | `Record<string, string>` | Optional non-owned headers. `authorization` is provider-owned and applied last. |
|
|
76
|
+
|
|
77
|
+
The endpoint is rate-limited to **1 request per second per customer** (429 with
|
|
78
|
+
`Retry-After: 1`). The helper performs no polling or caching and is never called
|
|
79
|
+
from `generate()` or package setup; the caller owns throttling.
|
|
80
|
+
|
|
81
|
+
NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
|
|
82
|
+
/ `extra` escape hatches: `compat.reasoning_effort` (`"low" | "medium" | "high"`),
|
|
83
|
+
`compat.thinking_token_budget`, `compat.chat_template_kwargs` (including `enable_thinking`),
|
|
84
|
+
`compat.preserve_thinking`, `compat.clear_thinking`, and `compat.tool_choice`.
|
|
85
|
+
`preserve_thinking: true` keeps prior assistant reasoning in request history so
|
|
86
|
+
multi-turn reasoning continues with the earlier chain of thought; `clear_thinking:
|
|
87
|
+
true` drops it for the next turn, resetting the chain. `options.extra` spreads after
|
|
88
|
+
`compat` so per-call values and overrides win.
|
|
89
|
+
|
|
90
|
+
## Outputs / response / events
|
|
91
|
+
|
|
92
|
+
| Surface | Behavior |
|
|
93
|
+
| --- | --- |
|
|
94
|
+
| Provider stream | Prism text, thinking (`delta.reasoning_content` → `providerThinkingDelta`), tool-call delta/final, `usage`, `done`, redacted `error` with HTTP-status `code` for retry classification. |
|
|
95
|
+
| Block preservation | Text, thinking, assistant `tool_call` → `tool_calls`, `tool_result` → role `tool` messages, images when `capabilities.input` includes `"image"`. |
|
|
96
|
+
| Model catalog | Featured aliases declare provider id, display name, context limit, text/image input support, tools, reasoning/fast variants, streaming, implicit cache, and NeuralWatt JSON-mode compat metadata where documented. |
|
|
97
|
+
| Pricing | Static aliases do not guess rates. Exact per-alias input/output/cache-read prices are advertised by NeuralWatt's `/v1/models` response and mapped by `listNeuralWattModels()` when present. |
|
|
98
|
+
| SSE comments | `: energy` / `: cost` comment lines are parsed by `neuralWattEventsWithTelemetry()` into `neuralwatt:telemetry` events; the standard `neuralWattEvents()` stream (used by `generate()`) tolerates them without spurious events. |
|
|
99
|
+
| `[DONE]` | Terminates the stream; final `providerDone(usage)` always emitted on a clean stream. |
|
|
100
|
+
| Malformed data | Yields `providerError` rather than crashing the generator (more robust than the Z.AI parser). |
|
|
101
|
+
| Auth method | `api_key` for the configured provider id, credential name `apiKey`. |
|
|
102
|
+
|
|
103
|
+
Unsupported block placements or unclaimed images fail before fetch.
|
|
104
|
+
|
|
105
|
+
## Request/response example
|
|
106
|
+
|
|
107
|
+
Example request body (OpenAI-compatible Chat Completions shape):
|
|
108
|
+
|
|
109
|
+
```json
|
|
110
|
+
{
|
|
111
|
+
"model": "glm-5.2",
|
|
112
|
+
"messages": [{ "role": "user", "content": "Hello" }],
|
|
113
|
+
"stream": true,
|
|
114
|
+
"stream_options": { "include_usage": true },
|
|
115
|
+
"reasoning_effort": "medium",
|
|
116
|
+
"thinking_token_budget": 8192
|
|
117
|
+
}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
## Implementation example
|
|
121
|
+
|
|
122
|
+
```ts
|
|
123
|
+
import { createExtensionKernel } from "@arnilo/prism";
|
|
124
|
+
import { createNeuralWattProviderPackage } from "@arnilo/prism-provider-neuralwatt";
|
|
125
|
+
|
|
126
|
+
const kernel = createExtensionKernel();
|
|
127
|
+
await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake-neuralwatt-key" })]);
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
Override the provider id and models:
|
|
131
|
+
|
|
132
|
+
```ts
|
|
133
|
+
import { createNeuralWattProviderPackage, defineNeuralWattModel, neuralWattModels } from "@arnilo/prism-provider-neuralwatt";
|
|
134
|
+
|
|
135
|
+
await kernel.load([
|
|
136
|
+
createNeuralWattProviderPackage({ id: "neuralwatt", apiKey: "fake", models: neuralWattModels }),
|
|
137
|
+
]);
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
Explicit catalog discovery:
|
|
141
|
+
|
|
142
|
+
```ts
|
|
143
|
+
import { listNeuralWattModels } from "@arnilo/prism-provider-neuralwatt";
|
|
144
|
+
|
|
145
|
+
const models = await listNeuralWattModels({ apiKey: "fake-neuralwatt-key", fetch });
|
|
146
|
+
await kernel.load([createNeuralWattProviderPackage({ apiKey: "fake", models })]);
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
Discovery performs exactly one `GET /v1/models` call when invoked. Provider package
|
|
150
|
+
setup and `generate()` never call model discovery implicitly.
|
|
151
|
+
|
|
152
|
+
Account quota:
|
|
153
|
+
|
|
154
|
+
```ts
|
|
155
|
+
import { getNeuralWattQuota } from "@arnilo/prism-provider-neuralwatt";
|
|
156
|
+
|
|
157
|
+
const quota = await getNeuralWattQuota({ apiKey: "fake-neuralwatt-key", fetch });
|
|
158
|
+
console.log(quota.usage?.current_month?.energy_kwh, quota.balance?.balance_usd);
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
Returns typed `NeuralWattQuota` (`balance`, `usage.lifetime`/`usage.current_month`,
|
|
162
|
+
`limits`, `subscription`, `key`). All fields optional; minimal structural
|
|
163
|
+
validation. The caller owns throttling/caching — the helper makes one explicit
|
|
164
|
+
`GET /v1/quota` call and is never invoked from `generate()` or setup.
|
|
165
|
+
|
|
166
|
+
## Extension and configuration notes
|
|
167
|
+
|
|
168
|
+
- Hosts choose base URL, provider id, model list, credential source, and `fetch`
|
|
169
|
+
impl.
|
|
170
|
+
- `defineNeuralWattModel` lets apps set NeuralWatt-specific `compat`
|
|
171
|
+
(`reasoning_effort`, `thinking_token_budget`, `chat_template_kwargs`,
|
|
172
|
+
`preserve_thinking`, `clear_thinking`, `tool_choice`).
|
|
173
|
+
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
174
|
+
- Curated aliases are static and network-free. They include documented context windows
|
|
175
|
+
and capabilities only; `ModelConfig.cost` is left unset until exact per-alias pricing
|
|
176
|
+
is read from NeuralWatt's `/v1/models` catalog.
|
|
177
|
+
- `listNeuralWattModels()` maps `/v1/models` entries to `ModelConfig`: id/display
|
|
178
|
+
name, capabilities, limits, implicit cache metadata, `ModelCost` pricing, and
|
|
179
|
+
provider-owned NeuralWatt metadata in `compat.neuralwatt`.
|
|
180
|
+
- `getNeuralWattQuota()` calls `GET /v1/quota` once with a required API key and
|
|
181
|
+
returns typed account quota (`balance`, `usage`, `limits`, `subscription`, `key`).
|
|
182
|
+
It is opt-in, never called from `generate()` or setup, and the caller owns
|
|
183
|
+
throttling (NeuralWatt limits the endpoint to 1 rps per customer).
|
|
184
|
+
|
|
185
|
+
### Model catalog and pricing
|
|
186
|
+
|
|
187
|
+
`neuralWattModels` includes NeuralWatt's featured aliases:
|
|
188
|
+
|
|
189
|
+
| Alias | Context | Notable metadata |
|
|
190
|
+
| --- | ---: | --- |
|
|
191
|
+
| `glm-5.2` | 1024K | Tools, reasoning |
|
|
192
|
+
| `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
|
|
193
|
+
| `glm-5.2-short` | 195K | Tools, reasoning |
|
|
194
|
+
| `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
|
|
195
|
+
| `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
|
|
196
|
+
| `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
|
|
197
|
+
| `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
|
|
198
|
+
| `qwen3.5-397b` | 256K | Tools, reasoning, JSON mode |
|
|
199
|
+
| `qwen3.5-397b-fast` | 256K | Tools, JSON mode, fast/no reasoning |
|
|
200
|
+
| `qwen3.6-35b` | 128K | Tools, reasoning, vision, JSON mode |
|
|
201
|
+
| `qwen3.6-35b-fast` | 128K | Tools, vision, JSON mode, fast/no reasoning |
|
|
202
|
+
|
|
203
|
+
NeuralWatt exposes exact pricing from `GET /v1/models` as per-million-token
|
|
204
|
+
`input_per_million`, `output_per_million`, `cached_input_per_million`,
|
|
205
|
+
`cached_output_per_million`, `currency`, and `pricing_tbd`. Cache reads for
|
|
206
|
+
NeuralWatt-hosted models are advertised by the API and default to 25% of the input
|
|
207
|
+
rate; there is no separate cache-write price (`cached_output_per_million` is `null`).
|
|
208
|
+
The static catalog does not copy or infer prices that are not published as fixed
|
|
209
|
+
alias values in these docs.
|
|
210
|
+
|
|
211
|
+
### Cache behavior
|
|
212
|
+
|
|
213
|
+
- NeuralWatt models use **implicit prefix caching**: the server caches prompt prefixes
|
|
214
|
+
automatically based on request content, with no explicit request-side cache payload.
|
|
215
|
+
Catalog models declare `cache: { kind: "implicit" }`. For the cross-provider
|
|
216
|
+
explicit/implicit cache matrix, see [Provider caching](../provider-caching.md).
|
|
217
|
+
- The provider sends no `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention`
|
|
218
|
+
fields regardless of `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention`
|
|
219
|
+
settings — those options have no effect on the NeuralWatt request body.
|
|
220
|
+
`cacheRetention: "none"` disables Prism cache-control hints only; it does **not**
|
|
221
|
+
disable the implicit backend prefix cache. Hosts relying on cache hits should keep
|
|
222
|
+
their stable prompt prefix byte-stable and stable inputs unchanged.
|
|
223
|
+
- Usage accounting is read-only: `prompt_tokens_details.cached_tokens` maps to
|
|
224
|
+
`Usage.cacheReadTokens`. NeuralWatt does not report a cache-write token today, so
|
|
225
|
+
`Usage.cacheWriteTokens` is never fabricated (stays `undefined`).
|
|
226
|
+
|
|
227
|
+
#### Cache-aware limiter behavior
|
|
228
|
+
|
|
229
|
+
NeuralWatt's backend rate limiter is cache-aware, which affects when long-running
|
|
230
|
+
agent sessions are throttled versus served:
|
|
231
|
+
|
|
232
|
+
- **Uncached TPM counts cold prefill only.** Tokens-per-minute accounting charges the
|
|
233
|
+
cold prefill of a request — the prefix that is not already in the vLLM prefix cache.
|
|
234
|
+
A request whose prefix is fully cached consumes far less of the TPM budget than a
|
|
235
|
+
cold request of the same total prompt length.
|
|
236
|
+
- **Warm-prefix requests can avoid some `503` fleet-capacity blocks.** When the fleet
|
|
237
|
+
is near capacity, requests that can reuse a cached prefix are more likely to be
|
|
238
|
+
admitted than fully cold requests. Prefix reuse is therefore both a latency and an
|
|
239
|
+
availability lever, not just a cost lever.
|
|
240
|
+
- **Full prior history is required for multi-turn cache reuse.** The prefix cache is
|
|
241
|
+
keyed by request content, so each follow-up turn must resend the entire prior
|
|
242
|
+
transcript (system prompt + all prior turns) unchanged, with only the new turn
|
|
243
|
+
appended. Prism's `inputLayout: "cache_aware"` ordering keeps the stable prefix
|
|
244
|
+
first; see [Provider caching](../provider-caching.md).
|
|
245
|
+
- Cache behavior is best-effort and **does not guarantee cache hits**. Prefix cache
|
|
246
|
+
admission and eviction are server-side decisions and can vary with fleet load.
|
|
247
|
+
`cacheRetention: "none"` disables Prism cache-control hints only; it does not
|
|
248
|
+
disable the implicit backend prefix cache.
|
|
249
|
+
|
|
250
|
+
### Reasoning preservation across turns
|
|
251
|
+
|
|
252
|
+
NeuralWatt reasoning-capable models (Kimi-, GLM-, Qwen-style aliases with
|
|
253
|
+
`capabilities.reasoning: true`) accept prior assistant reasoning in request history
|
|
254
|
+
so multi-turn sessions continue the earlier chain of thought:
|
|
255
|
+
|
|
256
|
+
- Prior `thinking` content blocks on an assistant message are serialized under a
|
|
257
|
+
`reasoning_content` field on that message (matching the streaming
|
|
258
|
+
`delta.reasoning_content` field). They are **not** flattened into text `content`, so
|
|
259
|
+
the model sees reasoning and answer as distinct.
|
|
260
|
+
- Preservation is gated on `model.capabilities.reasoning === true` **or**
|
|
261
|
+
`compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
|
|
262
|
+
field and prior `thinking` blocks are dropped — they never leak into text content for
|
|
263
|
+
providers/models that do not support reasoning.
|
|
264
|
+
- `compat.clear_thinking: true` drops prior reasoning for the next turn even on
|
|
265
|
+
reasoning-capable models, resetting the chain of thought. `clear_thinking` takes
|
|
266
|
+
precedence over `preserve_thinking`.
|
|
267
|
+
- The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
|
|
268
|
+
reasoning.
|
|
269
|
+
|
|
270
|
+
### Tool calls and the tool-call loop
|
|
271
|
+
|
|
272
|
+
NeuralWatt exposes OpenAI-compatible function calling. The provider carries tools and
|
|
273
|
+
prior tool turns through a multi-turn loop:
|
|
274
|
+
|
|
275
|
+
- **Request serialization.** `ProviderRequest.tools` (`ToolDefinition[]`) is serialized
|
|
276
|
+
to OpenAI `tools: [{ type: "function", function: { name, description, parameters } }]`.
|
|
277
|
+
Missing `parameters` default to `{ type: "object" }`. `compat.tool_choice` passes
|
|
278
|
+
through as `tool_choice` (string or `{ type: "function", function: { name } }`).
|
|
279
|
+
- **Streaming reconstruction.** `delta.tool_calls` fragments (keyed by `index`) are
|
|
280
|
+
accumulated and re-emitted as `tool_call_delta` events for UI consumers, then
|
|
281
|
+
reconstructed into a final `tool_call` event per call with parsed JSON arguments.
|
|
282
|
+
Parallel calls are tracked by index.
|
|
283
|
+
- **Next-turn ordering.** On the following turn the assistant `tool_call` block is
|
|
284
|
+
serialized to `role: "assistant"` with a `tool_calls` array (arguments stringified to
|
|
285
|
+
JSON), immediately followed by a `role: "tool"` message carrying `tool_call_id` and
|
|
286
|
+
the stringified `tool_result` — matching the OpenAI requirement that a tool result
|
|
287
|
+
follows the call that produced it. `tool_result` blocks must appear in `role: "tool"
|
|
288
|
+
messages; `tool_call` blocks must be the only content on their assistant message.
|
|
289
|
+
|
|
290
|
+
### Energy and cost telemetry
|
|
291
|
+
|
|
292
|
+
NeuralWatt streams energy and cost data as SSE comment lines (`: energy {...}`
|
|
293
|
+
and `: cost {...}`) before `data: [DONE]`, and as top-level `energy`/`cost` JSON
|
|
294
|
+
fields on non-streaming responses. Standard SSE clients ignore comments, so
|
|
295
|
+
these values are invisible unless the raw stream is parsed.
|
|
296
|
+
|
|
297
|
+
Prism's core `ProviderEvent` union has no generic telemetry event, so NeuralWatt
|
|
298
|
+
exposes telemetry through package-specific helpers:
|
|
299
|
+
|
|
300
|
+
- `neuralWattEventsWithTelemetry(body)` yields the standard provider events plus
|
|
301
|
+
`neuralwatt:telemetry` events (`{ type: "neuralwatt:telemetry", energy?, cost? }`)
|
|
302
|
+
in stream order. Use it when a host wants to observe telemetry alongside text,
|
|
303
|
+
tool, usage, and done events.
|
|
304
|
+
- `parseNeuralWattComment(text)` parses a single `: energy`/`: cost` comment line
|
|
305
|
+
into a `NeuralWattTelemetryEvent` (`undefined` for unknown/malformed comments).
|
|
306
|
+
- `parseNeuralWattEnergy(payload)` / `parseNeuralWattCost(payload)` parse the JSON
|
|
307
|
+
payload of a single comment into typed `NeuralWattEnergyTelemetry` /
|
|
308
|
+
`NeuralWattCostTelemetry`.
|
|
309
|
+
- `mapNeuralWattTelemetry(body)` maps a non-streaming response body's top-level
|
|
310
|
+
`energy`/`cost` fields into the same typed telemetry.
|
|
311
|
+
|
|
312
|
+
`generate()` stays streaming-only and uses `neuralWattEvents()`, so telemetry is
|
|
313
|
+
opt-in via `neuralWattEventsWithTelemetry()`. Telemetry contains usage/cost
|
|
314
|
+
numbers only — never prompts, API keys, or headers. All documented fields are
|
|
315
|
+
optional and tolerated when absent; malformed comments yield no telemetry event
|
|
316
|
+
and never crash the stream.
|
|
317
|
+
|
|
318
|
+
```ts
|
|
319
|
+
import { neuralWattEventsWithTelemetry } from "@arnilo/prism-provider-neuralwatt";
|
|
320
|
+
|
|
321
|
+
for await (const event of neuralWattEventsWithTelemetry(response.body)) {
|
|
322
|
+
if (event.type === "neuralwatt:telemetry") {
|
|
323
|
+
console.log(event.energy?.energy_kwh, event.cost?.request_cost_usd);
|
|
324
|
+
}
|
|
325
|
+
}
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
### Retry classification
|
|
329
|
+
|
|
330
|
+
NeuralWatt error responses are classified by `classifyNeuralWattError()` so the
|
|
331
|
+
Prism runtime retry policy can decide retryability without provider-specific
|
|
332
|
+
core branches:
|
|
333
|
+
|
|
334
|
+
| Status | Retryable | Notes |
|
|
335
|
+
| --- | --- | --- |
|
|
336
|
+
| `400` `401` `402` `403` `404` | no | Client/payment/auth errors fail closed. |
|
|
337
|
+
| `429` | yes | Reads `Retry-After` header and `error.retry_after`; preserves `error.retry_strategy` (`type`, `suggested_initial_delay_s`, `max_delay_s`, `backoff`, `jitter`). |
|
|
338
|
+
| `500` `502` `503` | yes | Transient server/fleet-capacity errors; `503` `Retry-After` honored when present. |
|
|
339
|
+
|
|
340
|
+
The provider emits `providerError` with `ErrorInfo.code` set to the numeric HTTP
|
|
341
|
+
status. Prism's default retry policy (`createDefaultRetryPolicy()`) treats `429`/
|
|
342
|
+
`500`/`502`/`503` as transient and `400`/`401`/`402`/`403`/`404` as non-transient,
|
|
343
|
+
so NeuralWatt errors retry correctly out of the box. `classifyNeuralWattError()`
|
|
344
|
+
and `neuralWattHttpError()` are exported for hosts/tests that want structured
|
|
345
|
+
retry metadata (`retryAfterMs`, `errorCode`, `strategy`). The host retry policy
|
|
346
|
+
owns the exact delay; `retryAfterMs` is surfaced but not enforced by the
|
|
347
|
+
provider. Classification is O(1) over status/headers/body and makes no extra
|
|
348
|
+
provider calls.
|
|
349
|
+
|
|
350
|
+
```ts
|
|
351
|
+
import { classifyNeuralWattError } from "@arnilo/prism-provider-neuralwatt";
|
|
352
|
+
|
|
353
|
+
const decision = classifyNeuralWattError({ status: 429, headers: { "retry-after": "1" }, body: { error: { code: "concurrent_budget_exceeded", retry_after: 1 } } });
|
|
354
|
+
// { retryable: true, code: 429, retryAfterMs: 1000, errorCode: "concurrent_budget_exceeded", strategy: undefined }
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
## Security and performance notes
|
|
358
|
+
|
|
359
|
+
- No network calls during import, setup, build, default tests, or generation beyond
|
|
360
|
+
the explicit Chat Completions request. `listNeuralWattModels()` is opt-in and
|
|
361
|
+
makes one `GET /v1/models` call per invocation; `getNeuralWattQuota()` is opt-in
|
|
362
|
+
and makes one `GET /v1/quota` call per invocation (endpoint limited to 1 rps per
|
|
363
|
+
customer; caller owns throttling).
|
|
364
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
365
|
+
- API keys are resolved per request/helper call from caller-supplied values or resolvers
|
|
366
|
+
and redacted from errors via `redactSecrets`. `listNeuralWattModels()` and
|
|
367
|
+
`getNeuralWattQuota()` apply provider-owned `authorization` after caller headers so
|
|
368
|
+
callers cannot override it. Quota values never enter provider events unless the caller
|
|
369
|
+
emits them.
|
|
370
|
+
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
|
|
371
|
+
but provider-owned headers (`content-type`, `authorization`) are applied last
|
|
372
|
+
and cannot be overridden by caller headers.
|
|
373
|
+
- Live tests stay opt-in behind `NEURALWATT_API_KEY` (plus `PRISM_LIVE_PROVIDER_TESTS=1`);
|
|
374
|
+
default tests are network-free.
|
|
375
|
+
|
|
376
|
+
## Related APIs
|
|
377
|
+
|
|
378
|
+
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
379
|
+
`ModelConfig`/`compat`, thinking formats.
|
|
380
|
+
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
381
|
+
`resolveCredentialValue`, `redactSecrets`.
|
|
382
|
+
- [OpenAI-compatible provider](openai-compatible.md): underlying Chat Completions
|
|
383
|
+
adapter.
|
|
384
|
+
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
385
|
+
- [Provider caching](../provider-caching.md): implicit cache behavior and
|
|
386
|
+
`cacheUsageReport`.
|
|
387
|
+
- [NeuralWatt agent example](../../examples/neuralwatt-agent-run.ts): runnable mocked
|
|
388
|
+
agent turn with tools, reasoning controls, streamed cache tokens, and energy/cost telemetry.
|