@arnilo/prism 0.0.5 → 0.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -1
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +27 -16
- package/dist/agent-run-lifecycle.d.ts +28 -0
- package/dist/agent-run-lifecycle.js +33 -0
- package/dist/agent-run-state.d.ts +53 -0
- package/dist/agent-run-state.js +127 -0
- package/dist/agents.d.ts +3 -1
- package/dist/agents.js +337 -46
- package/dist/contracts.d.ts +205 -3
- package/dist/contracts.js +4 -0
- package/dist/guardrails.d.ts +25 -0
- package/dist/guardrails.js +133 -0
- package/dist/ids.d.ts +2 -0
- package/dist/ids.js +6 -0
- package/dist/index.d.ts +17 -3
- package/dist/index.js +10 -3
- package/dist/input.js +2 -0
- package/dist/resources.js +2 -1
- package/dist/run-limits.d.ts +34 -0
- package/dist/run-limits.js +163 -0
- package/dist/secure-agent.d.ts +3 -0
- package/dist/secure-agent.js +63 -0
- package/dist/session-stores.js +2 -3
- package/dist/testing/persistence-schema.d.ts +45 -7
- package/dist/testing/persistence-schema.js +138 -24
- package/dist/thinking.d.ts +42 -0
- package/dist/thinking.js +92 -0
- package/dist/tools.d.ts +10 -2
- package/dist/tools.js +56 -7
- package/dist/use-case-model.d.ts +63 -0
- package/dist/use-case-model.js +52 -0
- package/docs/a2a.md +4 -2
- package/docs/agent-events.md +23 -16
- package/docs/agent-loops.md +19 -8
- package/docs/agent-session-runtime.md +33 -1
- package/docs/coding-agent-tools.md +33 -12
- package/docs/coding-security.md +2 -2
- package/docs/compaction-llm.md +17 -7
- package/docs/compaction-observational-memory.md +28 -4
- package/docs/credential-storage.md +58 -9
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/database-persistence.md +8 -3
- package/docs/guardrails.md +75 -0
- package/docs/host-security.md +16 -8
- package/docs/index.md +26 -22
- package/docs/mcp-tools.md +32 -12
- package/docs/migration.md +164 -2
- package/docs/node-filesystem-config.md +1 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/postgres-persistence.md +3 -3
- package/docs/provider-caching.md +16 -4
- package/docs/provider-conformance.md +39 -1
- package/docs/provider-packages.md +60 -3
- package/docs/providers/ai-sdk.md +36 -0
- package/docs/providers/kimi.md +124 -61
- package/docs/providers/neuralwatt.md +19 -13
- package/docs/providers/openai.md +56 -13
- package/docs/providers/opencode-go.md +118 -30
- package/docs/providers/openrouter.md +105 -35
- package/docs/providers/zai.md +94 -45
- package/docs/release-and-install.md +47 -49
- package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
- package/docs/runs-and-usage.md +30 -3
- package/docs/server.md +5 -2
- package/docs/sqlite-persistence.md +2 -2
- package/docs/structured-output.md +1 -1
- package/docs/thinking-and-reasoning.md +98 -0
- package/docs/tool-execution-primitives.md +3 -3
- package/docs/tools.md +21 -1
- package/docs/use-case-model-selection.md +109 -0
- package/docs/workflow-orchestration-primitives.md +1 -0
- package/docs/workflows.md +18 -10
- package/docs/working-and-semantic-memory.md +1 -0
- package/package.json +2 -2
|
@@ -9,7 +9,7 @@ and implicit prefix caching.
|
|
|
9
9
|
|
|
10
10
|
The package registers a provider, default model metadata for the featured NeuralWatt
|
|
11
11
|
aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
|
|
12
|
-
`kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
|
|
12
|
+
`gemma-4-31b`, `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
|
|
13
13
|
`qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
|
|
14
14
|
method through `createExtensionKernel().load([...])`.
|
|
15
15
|
|
|
@@ -79,13 +79,15 @@ The endpoint is rate-limited to **1 request per second per customer** (429 with
|
|
|
79
79
|
from `generate()` or package setup; the caller owns throttling.
|
|
80
80
|
|
|
81
81
|
NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
|
|
82
|
-
/ `extra` escape hatches: `compat.reasoning_effort` (
|
|
83
|
-
|
|
84
|
-
`compat.
|
|
85
|
-
`
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
`
|
|
82
|
+
/ `extra` escape hatches: `compat.reasoning_effort` (OpenAI-style scale; GLM-5.2 defaults
|
|
83
|
+
unset to `max` per official docs), `compat.thinking_token_budget`,
|
|
84
|
+
`compat.chat_template_kwargs` (including `enable_thinking`, `preserve_thinking`, and
|
|
85
|
+
`clear_thinking` per official gateway docs), and `compat.tool_choice`.
|
|
86
|
+
`preserve_thinking` / `clear_thinking` compat flags are routed into `chat_template_kwargs`
|
|
87
|
+
on the wire (not top-level body fields). Prism also uses these flags for client-side
|
|
88
|
+
message serialization: `preserve_thinking: true` keeps prior assistant reasoning in
|
|
89
|
+
request history; `clear_thinking: true` drops it for the next turn. `options.extra` spreads
|
|
90
|
+
after resolved fields so per-call values and overrides win.
|
|
89
91
|
|
|
90
92
|
## Outputs / response / events
|
|
91
93
|
|
|
@@ -112,7 +114,7 @@ Example request body (OpenAI-compatible Chat Completions shape):
|
|
|
112
114
|
"messages": [{ "role": "user", "content": "Hello" }],
|
|
113
115
|
"stream": true,
|
|
114
116
|
"stream_options": { "include_usage": true },
|
|
115
|
-
"reasoning_effort": "
|
|
117
|
+
"reasoning_effort": "max",
|
|
116
118
|
"thinking_token_budget": 8192
|
|
117
119
|
}
|
|
118
120
|
```
|
|
@@ -192,6 +194,7 @@ validation. The caller owns throttling/caching — the helper makes one explicit
|
|
|
192
194
|
| `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
|
|
193
195
|
| `glm-5.2-short` | 195K | Tools, reasoning |
|
|
194
196
|
| `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
|
|
197
|
+
| `gemma-4-31b` | 256K | Tools, vision, JSON mode |
|
|
195
198
|
| `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
|
|
196
199
|
| `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
|
|
197
200
|
| `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
|
|
@@ -255,15 +258,18 @@ so multi-turn sessions continue the earlier chain of thought:
|
|
|
255
258
|
|
|
256
259
|
- Prior `thinking` content blocks on an assistant message are serialized under a
|
|
257
260
|
`reasoning_content` field on that message (matching the streaming
|
|
258
|
-
`delta.reasoning_content` field
|
|
261
|
+
`delta.reasoning_content` field; the gateway also accepts `reasoning` as an alias).
|
|
262
|
+
They are **not** flattened into text `content`, so
|
|
259
263
|
the model sees reasoning and answer as distinct.
|
|
260
264
|
- Preservation is gated on `model.capabilities.reasoning === true` **or**
|
|
261
265
|
`compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
|
|
262
266
|
field and prior `thinking` blocks are dropped — they never leak into text content for
|
|
263
267
|
providers/models that do not support reasoning.
|
|
264
|
-
- `compat.
|
|
265
|
-
|
|
266
|
-
|
|
268
|
+
- `compat.preserve_thinking` / `compat.clear_thinking` map into `chat_template_kwargs`
|
|
269
|
+
on the request body per official NeuralWatt docs (Kimi K2.6 `preserve_thinking`, GLM
|
|
270
|
+
`clear_thinking: false` for full-history). Prism also uses `clear_thinking: true` to
|
|
271
|
+
drop prior reasoning client-side even on reasoning-capable models; `clear_thinking`
|
|
272
|
+
takes precedence over `preserve_thinking`.
|
|
267
273
|
- The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
|
|
268
274
|
reasoning.
|
|
269
275
|
|
package/docs/providers/openai.md
CHANGED
|
@@ -36,16 +36,22 @@ createOpenAIProviderPackage(options: OpenAIProviderPackageOptions): ProviderPack
|
|
|
36
36
|
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
37
37
|
| `baseUrl` | `string` | Overrides `https://api.openai.com/v1`. |
|
|
38
38
|
| `codexBaseUrl` | `string` | Overrides `https://chatgpt.com/backend-api/codex`. |
|
|
39
|
+
| `models` | `readonly ModelConfig[]` | Optional override for registered OpenAI Responses models (defaults to featured `openAIModels`). |
|
|
40
|
+
| `codexModels` | `readonly ModelConfig[]` | Optional override for registered Codex models (defaults to featured `openAICodexModels`). |
|
|
39
41
|
|
|
40
42
|
`ProviderRequest.options.sessionId`, `cacheKey`, `cacheRetention`, `headers`,
|
|
41
|
-
`compat`, and `extra` map to request headers/payload fields.
|
|
43
|
+
`compat`, and `extra` map to request headers/payload fields. Per-turn reasoning
|
|
44
|
+
uses official Responses `reasoning: { effort, summary? }` via
|
|
45
|
+
`ModelConfig.compat.reasoning` defaults merged with
|
|
46
|
+
`ProviderRequestOptions.compat.reasoning` (request wins). Prefer
|
|
47
|
+
`applyThinkingLevel(..., "openai_reasoning")` from `@arnilo/prism`.
|
|
42
48
|
|
|
43
49
|
## Outputs / response / events
|
|
44
50
|
|
|
45
51
|
| Surface | Behavior |
|
|
46
52
|
| --- | --- |
|
|
47
53
|
| Provider stream | Prism text, thinking (downgraded to text), `tool_call` deltas/finals, `usage`, `done`, redacted `error` events. |
|
|
48
|
-
| Block preservation |
|
|
54
|
+
| Block preservation | User/system text → `input_text`; assistant text → `output_text`; assistant `tool_call` → top-level `function_call` with `call_id`; `tool_result` → top-level `function_call_output`; images/files/audio when declared on the model. Bare thinking without an encrypted Responses reasoning item is omitted on replay. |
|
|
49
55
|
| Auth methods | `api_key` for `openai`; `oauth` for `openai-codex`. |
|
|
50
56
|
|
|
51
57
|
Unsupported block placements or unclaimed images fail before `fetch`.
|
|
@@ -56,10 +62,17 @@ Responses request body (Codex subscription shape, abbreviated):
|
|
|
56
62
|
|
|
57
63
|
```json
|
|
58
64
|
{
|
|
59
|
-
"model": "gpt-5
|
|
60
|
-
"
|
|
61
|
-
|
|
62
|
-
|
|
65
|
+
"model": "gpt-5.1",
|
|
66
|
+
"input": [
|
|
67
|
+
{ "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] },
|
|
68
|
+
{ "role": "assistant", "content": [{ "type": "output_text", "text": "Calling lookup" }] },
|
|
69
|
+
{ "type": "function_call", "call_id": "call_1", "name": "lookup", "arguments": "{\"q\":\"x\"}" },
|
|
70
|
+
{ "type": "function_call_output", "call_id": "call_1", "output": "{\"ok\":true}" }
|
|
71
|
+
],
|
|
72
|
+
"reasoning": { "effort": "high" },
|
|
73
|
+
"prompt_cache_key": "session-1",
|
|
74
|
+
"stream": true,
|
|
75
|
+
"store": false
|
|
63
76
|
}
|
|
64
77
|
```
|
|
65
78
|
|
|
@@ -73,13 +86,13 @@ https://auth.openai.com/authorize?response_type=code&client_id=...&code_challeng
|
|
|
73
86
|
|
|
74
87
|
```ts
|
|
75
88
|
import { createExtensionKernel, createEnvCredentialResolver } from "@arnilo/prism";
|
|
76
|
-
import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
|
|
89
|
+
import { createOpenAIProviderPackage, listOpenAIModels } from "@arnilo/prism-provider-openai";
|
|
77
90
|
|
|
91
|
+
const apiKey = createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" });
|
|
92
|
+
const models = await listOpenAIModels({ apiKey }); // caller-gated; never runs during setup
|
|
78
93
|
const kernel = createExtensionKernel();
|
|
79
94
|
await kernel.load([
|
|
80
|
-
createOpenAIProviderPackage({
|
|
81
|
-
apiKey: createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" }),
|
|
82
|
-
}),
|
|
95
|
+
createOpenAIProviderPackage({ apiKey, models }),
|
|
83
96
|
]);
|
|
84
97
|
```
|
|
85
98
|
|
|
@@ -115,25 +128,55 @@ const challenge = computeS256Challenge(verifier);
|
|
|
115
128
|
|
|
116
129
|
### Cache behavior
|
|
117
130
|
|
|
131
|
+
Official: [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching).
|
|
132
|
+
|
|
118
133
|
- `prompt_cache_key` is derived from `ProviderRequestOptions.cacheKey` (falling
|
|
119
134
|
back to `sessionId`) and sanitized + clamped to 64 characters via the shared
|
|
120
135
|
`sanitizeCacheKey()` helper. Cache keys are session/customer identifiers only;
|
|
121
136
|
never credentials or raw prompts.
|
|
122
|
-
- `prompt_cache_retention` accepts only `"24h"` on
|
|
137
|
+
- `prompt_cache_retention` accepts only `"24h"` on pre-GPT-5.6 Responses models
|
|
123
138
|
(extended caching). Prism `cacheRetention: "short"` and `"none"` omit the field
|
|
124
139
|
so default automatic/implicit caching applies and no invalid literal is sent.
|
|
125
140
|
`cacheRetention: "long"` maps to `prompt_cache_retention: "24h"` only when the
|
|
126
141
|
model declares `ModelConfig.cache.longRetention === true`; models without that
|
|
127
|
-
metadata omit the field.
|
|
142
|
+
metadata omit the field. Featured `gpt-5.1` declares
|
|
128
143
|
`cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
|
|
144
|
+
- GPT-5.6+ official docs prefer `prompt_cache_options` / explicit breakpoints;
|
|
145
|
+
`listOpenAIModels` sets `longRetention: false` for those ids so Prism does not
|
|
146
|
+
emit deprecated `prompt_cache_retention` for them. Breakpoint helpers are not
|
|
147
|
+
shipped in this package yet — hosts may pass `prompt_cache_options` through
|
|
148
|
+
`compat` / `extra` when needed.
|
|
129
149
|
- Cache accounting is preserved in normalized `Usage`: OpenAI
|
|
130
150
|
`input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
|
|
131
|
-
Responses does not report a cache-write token field.
|
|
151
|
+
Responses does not report a cache-write token field on older models.
|
|
132
152
|
- Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
|
|
133
153
|
are applied after caller `ProviderRequestOptions.headers` so caller config
|
|
134
154
|
cannot replace credentials, content type, or the session request id; non-owned
|
|
135
155
|
caller headers are kept.
|
|
136
156
|
|
|
157
|
+
### Model discovery
|
|
158
|
+
|
|
159
|
+
- `listOpenAIModels({ apiKey, fetch, baseUrl, signal, headers })` calls official
|
|
160
|
+
[`GET /models`](https://developers.openai.com/api/reference/resources/models/methods/list)
|
|
161
|
+
and maps sparse `{ id, created, owned_by }` entries to `ModelConfig` with
|
|
162
|
+
`cache.kind: "openai_key"` and heuristic `longRetention` / `capabilities.reasoning`.
|
|
163
|
+
- `createOpenAIProviderPackage` never calls discovery; pass results via `models:`.
|
|
164
|
+
- Codex subscription models are **not** listed by `api.openai.com` — keep using
|
|
165
|
+
featured `openAICodexModels` or `codexModels:` override.
|
|
166
|
+
- Static `openAIModels` / `openAICodexModels` are offline bootstrap / featured aliases only.
|
|
167
|
+
|
|
168
|
+
### Reasoning
|
|
169
|
+
|
|
170
|
+
Official: [Reasoning models](https://developers.openai.com/api/docs/guides/reasoning).
|
|
171
|
+
|
|
172
|
+
- Body field is top-level `reasoning: { effort, summary?, mode?, context? }`.
|
|
173
|
+
- Model defaults: `ModelConfig.compat.reasoning`; per-turn override:
|
|
174
|
+
`ProviderRequestOptions.compat.reasoning` (shallow-merged; request wins).
|
|
175
|
+
- Portable helper: `applyThinkingLevel(options, level, "openai_reasoning")`.
|
|
176
|
+
- Streaming tool args follow official
|
|
177
|
+
`response.output_item.added` + `response.function_call_arguments.delta`
|
|
178
|
+
(string `delta`), not Chat Completions object deltas.
|
|
179
|
+
|
|
137
180
|
## Security and performance notes
|
|
138
181
|
|
|
139
182
|
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
|
|
@@ -2,27 +2,41 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-provider-opencode-go` provides explicit, side-effect-free setup for
|
|
6
|
-
OpenCode Go
|
|
7
|
-
|
|
8
|
-
cache/session headers.
|
|
5
|
+
`@arnilo/prism-provider-opencode-go` provides explicit, side-effect-free setup for
|
|
6
|
+
[OpenCode Go](https://opencode.ai/docs/go/) — a low-cost subscription gateway for
|
|
7
|
+
open coding models. The package dual-routes by `ModelConfig.compat.route`:
|
|
9
8
|
|
|
10
|
-
|
|
11
|
-
|
|
9
|
+
| Route | Endpoint | Official model families |
|
|
10
|
+
| --- | --- | --- |
|
|
11
|
+
| `"openai"` (default) | `POST {baseUrl}/chat/completions` | Grok, GLM, Kimi, MiMo, DeepSeek |
|
|
12
|
+
| `"anthropic"` | `POST {baseUrl}/messages` | MiniMax, Qwen |
|
|
13
|
+
|
|
14
|
+
Default base URL is the official Go API root:
|
|
15
|
+
|
|
16
|
+
```txt
|
|
17
|
+
https://opencode.ai/zen/go/v1
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
Session stickiness uses the sanitized `x-opencode-session` header. Anthropic-route
|
|
21
|
+
models may emit selected `cache_control` breakpoints; OpenAI-route models use
|
|
22
|
+
implicit caching and never receive Anthropic cache fields.
|
|
12
23
|
|
|
13
24
|
## When to use it
|
|
14
25
|
|
|
15
|
-
Use it when a host app wants
|
|
16
|
-
|
|
17
|
-
|
|
26
|
+
Use it when a host app wants OpenCode Go models through Prism's `AgentSession`
|
|
27
|
+
runtime with dual-route serialization, per-request session headers, and optional
|
|
28
|
+
caller-gated model discovery.
|
|
18
29
|
|
|
19
|
-
Do not use it for automatic credential discovery, catalog fetches, or
|
|
30
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches, or
|
|
20
31
|
real-network tests.
|
|
21
32
|
|
|
22
33
|
## Inputs / request
|
|
23
34
|
|
|
24
35
|
```ts
|
|
25
|
-
import {
|
|
36
|
+
import {
|
|
37
|
+
createOpenCodeGoProviderPackage,
|
|
38
|
+
listOpenCodeGoModels,
|
|
39
|
+
} from "@arnilo/prism-provider-opencode-go";
|
|
26
40
|
|
|
27
41
|
createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): ProviderPackage
|
|
28
42
|
```
|
|
@@ -31,57 +45,126 @@ createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): Prov
|
|
|
31
45
|
| --- | --- | --- |
|
|
32
46
|
| `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
|
|
33
47
|
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
34
|
-
| `baseUrl` | `string` | Overrides
|
|
35
|
-
| `models` | `readonly ModelConfig[]` | Overrides `openCodeGoModels` defaults. |
|
|
48
|
+
| `baseUrl` | `string` | Overrides official `https://opencode.ai/zen/go/v1`. |
|
|
49
|
+
| `models` | `readonly ModelConfig[]` | Overrides featured `openCodeGoModels` defaults. |
|
|
36
50
|
|
|
37
51
|
`ProviderRequest.options.cacheKey` (falling back to `sessionId`) maps to the
|
|
38
|
-
`x-opencode-session` header
|
|
39
|
-
`
|
|
52
|
+
`x-opencode-session` header. Anthropic-route `cache_control` breakpoints and
|
|
53
|
+
`cacheRetention` map as documented below.
|
|
40
54
|
|
|
41
55
|
## Outputs / response / events
|
|
42
56
|
|
|
43
57
|
| Surface | Behavior |
|
|
44
58
|
| --- | --- |
|
|
45
59
|
| Provider stream | Prism text, thinking, tool-call delta/final, `usage`, `done`, redacted `error`. |
|
|
46
|
-
|
|
|
60
|
+
| OpenAI thinking | `delta.reasoning_content` → thinking deltas; replay via `reasoning_content` when `preserveThinking`. |
|
|
61
|
+
| Anthropic thinking | `thinking_delta` → thinking deltas; replay via Anthropic thinking blocks when `preserveThinking`. |
|
|
62
|
+
| Session/cache | `x-opencode-session` + route-specific cache markers. |
|
|
47
63
|
| Auth method | `api_key` for `opencode-go`, credential name `apiKey`. |
|
|
48
64
|
|
|
49
65
|
## Request/response example
|
|
50
66
|
|
|
51
|
-
Example headers added before fetch:
|
|
52
|
-
|
|
53
67
|
```json
|
|
54
68
|
{
|
|
55
69
|
"Authorization": "Bearer <resolved-key>",
|
|
70
|
+
"content-type": "application/json",
|
|
56
71
|
"x-opencode-session": "<ProviderRequest.options.cacheKey ?? sessionId>"
|
|
57
72
|
}
|
|
58
73
|
```
|
|
59
74
|
|
|
75
|
+
OpenAI-route body (thinking passthrough + preserved reasoning):
|
|
76
|
+
|
|
77
|
+
```json
|
|
78
|
+
{
|
|
79
|
+
"model": "kimi-k3",
|
|
80
|
+
"stream": true,
|
|
81
|
+
"stream_options": { "include_usage": true },
|
|
82
|
+
"reasoning_effort": "high",
|
|
83
|
+
"messages": [
|
|
84
|
+
{
|
|
85
|
+
"role": "assistant",
|
|
86
|
+
"content": "calling",
|
|
87
|
+
"tool_calls": [{ "id": "call_1", "type": "function", "function": { "name": "lookup", "arguments": "{\"q\":\"x\"}" } }],
|
|
88
|
+
"reasoning_content": "plan the lookup"
|
|
89
|
+
}
|
|
90
|
+
]
|
|
91
|
+
}
|
|
92
|
+
```
|
|
93
|
+
|
|
60
94
|
## Implementation example
|
|
61
95
|
|
|
62
96
|
```ts
|
|
63
97
|
import { createExtensionKernel } from "@arnilo/prism";
|
|
64
|
-
import {
|
|
98
|
+
import {
|
|
99
|
+
createOpenCodeGoProviderPackage,
|
|
100
|
+
listOpenCodeGoModels,
|
|
101
|
+
openCodeGoModels,
|
|
102
|
+
} from "@arnilo/prism-provider-opencode-go";
|
|
65
103
|
|
|
66
104
|
const kernel = createExtensionKernel();
|
|
67
|
-
await kernel.load([createOpenCodeGoProviderPackage({ apiKey:
|
|
105
|
+
await kernel.load([createOpenCodeGoProviderPackage({ apiKey: process.env.OPENCODE_API_KEY })]);
|
|
68
106
|
```
|
|
69
107
|
|
|
70
|
-
|
|
108
|
+
Caller-gated live catalog (never runs during package setup):
|
|
71
109
|
|
|
72
110
|
```ts
|
|
73
|
-
|
|
111
|
+
const models = await listOpenCodeGoModels({ apiKey: process.env.OPENCODE_API_KEY });
|
|
112
|
+
await kernel.load([createOpenCodeGoProviderPackage({ apiKey: process.env.OPENCODE_API_KEY, models })]);
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
Offline bootstrap with featured docs-verified aliases:
|
|
74
116
|
|
|
117
|
+
```ts
|
|
75
118
|
await kernel.load([
|
|
76
119
|
createOpenCodeGoProviderPackage({ apiKey: "fake", models: openCodeGoModels }),
|
|
77
120
|
]);
|
|
78
121
|
```
|
|
79
122
|
|
|
123
|
+
## Featured models and routes
|
|
124
|
+
|
|
125
|
+
Featured `openCodeGoModels` mirrors the official Go docs list (open coding models
|
|
126
|
+
only — **not** Zen GPT/Claude ids). Route selection follows the official endpoint
|
|
127
|
+
table; Pi secondary metadata is used only for context/output limits when docs omit them.
|
|
128
|
+
|
|
129
|
+
| Model ID | Route | Cache kind |
|
|
130
|
+
| --- | --- | --- |
|
|
131
|
+
| `grok-4.5`, `glm-5.2`, `glm-5.1`, `kimi-k3`, `kimi-k2.7-code`, `kimi-k2.6`, `mimo-v2.5`, `mimo-v2.5-pro`, `deepseek-v4-pro`, `deepseek-v4-flash` | `openai` | `implicit` |
|
|
132
|
+
| `minimax-m3`, `minimax-m2.7`, `minimax-m2.5`, `qwen3.7-max`, `qwen3.7-plus`, `qwen3.6-plus` | `anthropic` | `cache_control` |
|
|
133
|
+
|
|
134
|
+
## Model discovery
|
|
135
|
+
|
|
136
|
+
Official list endpoint (sparse OpenAI-compatible shape):
|
|
137
|
+
|
|
138
|
+
```txt
|
|
139
|
+
GET https://opencode.ai/zen/go/v1/models
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
`listOpenCodeGoModels({ apiKey?, fetch?, baseUrl?, signal?, headers? })` maps each
|
|
143
|
+
`{ id, owned_by }` entry to `ModelConfig` with route/cache heuristics from the docs
|
|
144
|
+
endpoint table. Featured metadata (pricing/limits/thinking defaults) is applied when
|
|
145
|
+
the id matches `openCodeGoModels`. Discovery is **caller-gated** — setup performs
|
|
146
|
+
zero fetches.
|
|
147
|
+
|
|
148
|
+
## Thinking / reasoning
|
|
149
|
+
|
|
150
|
+
OpenCode Go does not document gateway-owned thinking fields; Prism forwards
|
|
151
|
+
upstream-compatible compat and preserves prior reasoning for tool-call continuity:
|
|
152
|
+
|
|
153
|
+
| Surface | Behavior |
|
|
154
|
+
| --- | --- |
|
|
155
|
+
| OpenAI route stream | `reasoning_content` → thinking deltas |
|
|
156
|
+
| OpenAI route replay | thinking blocks → top-level `reasoning_content` when `preserveThinking` (default for reasoning models); never folded into text |
|
|
157
|
+
| OpenAI route body | optional `thinking` / `reasoning_effort` / `reasoning` from model + per-turn `options.compat` (request wins) |
|
|
158
|
+
| Anthropic route stream | `thinking_delta` → thinking deltas |
|
|
159
|
+
| Anthropic route replay | thinking blocks with optional `signature` when `preserveThinking` |
|
|
160
|
+
|
|
161
|
+
Owned compat keys (`route`, `thinking`, `reasoning`, `reasoning_effort`,
|
|
162
|
+
`preserveThinking`) are stripped before opaque compat spread so resolved values win.
|
|
163
|
+
|
|
80
164
|
## Extension and configuration notes
|
|
81
165
|
|
|
82
166
|
- Hosts choose base URL, model list, credential source, and `fetch` impl.
|
|
83
|
-
-
|
|
84
|
-
routes preserve `tool_use`/`tool_result` blocks.
|
|
167
|
+
- Route selection is explicit via `compat.route` (`"anthropic"` or default `"openai"`).
|
|
85
168
|
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
86
169
|
|
|
87
170
|
### Cache and session behavior
|
|
@@ -114,19 +197,24 @@ await kernel.load([
|
|
|
114
197
|
- No network calls during import, setup, build, or default tests.
|
|
115
198
|
- No automatic environment, file, keychain, or shell credential lookup.
|
|
116
199
|
- API keys are resolved per request from caller-supplied values or resolvers and
|
|
117
|
-
redacted from errors.
|
|
200
|
+
redacted from errors (including discovery failures).
|
|
118
201
|
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
|
|
119
202
|
provider-owned headers (`content-type`, `x-opencode-session`, `authorization`)
|
|
120
203
|
are applied last and cannot be overridden by caller headers.
|
|
121
|
-
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus
|
|
122
|
-
|
|
204
|
+
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus `OPENCODE_API_KEY`;
|
|
205
|
+
default tests are network-free.
|
|
206
|
+
|
|
207
|
+
## Official evidence
|
|
208
|
+
|
|
209
|
+
- [OpenCode Go](https://opencode.ai/docs/go/) — model list, dual endpoints, pricing/usage, `GET /zen/go/v1/models`
|
|
210
|
+
- Pi secondary (ids/limits only): `packages/ai/src/providers/opencode-go.ts`, `opencode-go.models.ts`
|
|
123
211
|
|
|
124
212
|
## Related APIs
|
|
125
213
|
|
|
126
214
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
127
|
-
`ModelConfig`, request/cache policies.
|
|
215
|
+
`ModelConfig`, discovery contract, request/cache policies.
|
|
216
|
+
- [Thinking and reasoning](../thinking-and-reasoning.md): per-turn `ThinkingLevel` → compat families.
|
|
128
217
|
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
129
218
|
`resolveCredentialValue`, `redactSecrets`.
|
|
130
|
-
- [
|
|
131
|
-
adapter.
|
|
219
|
+
- [Provider caching](../provider-caching.md): route-specific OpenCode Go cache matrix.
|
|
132
220
|
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
@@ -3,12 +3,15 @@
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
5
|
`@arnilo/prism-provider-openrouter` provides explicit, side-effect-free setup for the
|
|
6
|
-
OpenRouter API-key provider with app-controlled model
|
|
7
|
-
|
|
6
|
+
OpenRouter API-key provider with **app-controlled** model registration, routing
|
|
7
|
+
passthrough, official `reasoning` controls, and Anthropic-style `cache_control`
|
|
8
|
+
(plus sticky `session_id` routing).
|
|
8
9
|
|
|
9
10
|
The package registers a provider, caller-supplied model metadata, and an
|
|
10
|
-
`api_key` auth method through `createExtensionKernel().load([...])`.
|
|
11
|
-
|
|
11
|
+
`api_key` auth method through `createExtensionKernel().load([...])`. There is
|
|
12
|
+
**no bundled mega-catalog**. Optional `listOpenRouterModels()` lets hosts fetch
|
|
13
|
+
the live official catalog and pass a filtered subset via `models:` — setup
|
|
14
|
+
itself never fetches.
|
|
12
15
|
|
|
13
16
|
## When to use it
|
|
14
17
|
|
|
@@ -16,16 +19,22 @@ Use it when a host app wants OpenRouter routing passthrough, reasoning controls,
|
|
|
16
19
|
and per-model cache policy through Prism's `AgentSession` runtime, and needs to
|
|
17
20
|
override cache behavior per model rather than accept a single hard-coded policy.
|
|
18
21
|
|
|
19
|
-
Do not use it for catalog
|
|
20
|
-
real-network tests.
|
|
22
|
+
Do not use it for automatic catalog fetch during setup, automatic credential
|
|
23
|
+
discovery, or real-network tests in CI defaults.
|
|
21
24
|
|
|
22
25
|
## Inputs / request
|
|
23
26
|
|
|
24
27
|
```ts
|
|
25
|
-
import {
|
|
28
|
+
import {
|
|
29
|
+
createOpenRouterProviderPackage,
|
|
30
|
+
defineOpenRouterModel,
|
|
31
|
+
listOpenRouterModels,
|
|
32
|
+
} from "@arnilo/prism-provider-openrouter";
|
|
26
33
|
|
|
27
34
|
createOpenRouterProviderPackage(options: OpenRouterProviderPackageOptions): ProviderPackage
|
|
28
|
-
defineOpenRouterModel(config: OpenRouterModelConfig):
|
|
35
|
+
defineOpenRouterModel(config: OpenRouterModelConfig): ModelConfig
|
|
36
|
+
listOpenRouterModels(options?: ListOpenRouterModelsOptions): Promise<ModelConfig[]>
|
|
37
|
+
mapOpenRouterModel(entry: OpenRouterModelEntry): ModelConfig
|
|
29
38
|
```
|
|
30
39
|
|
|
31
40
|
| Field | Type | Purpose |
|
|
@@ -37,25 +46,32 @@ defineOpenRouterModel(config: OpenRouterModelConfig): OpenRouterModelConfig
|
|
|
37
46
|
| `appTitle` | `string` | App title for the `X-Title` attribution header. |
|
|
38
47
|
| `models` | `readonly ModelConfig[]` | App-supplied model catalog (no default fetch). |
|
|
39
48
|
|
|
40
|
-
`OpenRouterModelConfig.compat.openRouterRouting` controls routing order
|
|
41
|
-
`data_collection
|
|
49
|
+
`OpenRouterModelConfig.compat.openRouterRouting` controls routing order /
|
|
50
|
+
`data_collection`. `compat.reasoning` carries the official OpenRouter
|
|
51
|
+
`reasoning` object (`effort`, `max_tokens`, `exclude`, …). Per-turn
|
|
52
|
+
`providerOptions.compat.reasoning` merges over model defaults (request wins
|
|
53
|
+
key-by-key). `compat.preserveThinking` replays assistant thinking as body
|
|
54
|
+
`reasoning` for tool-call continuity.
|
|
42
55
|
|
|
43
56
|
## Outputs / response / events
|
|
44
57
|
|
|
45
58
|
| Surface | Behavior |
|
|
46
59
|
| --- | --- |
|
|
47
|
-
| Provider stream | Prism text, thinking, tool-call delta/final, `usage` (with cache read/write mapped), `done`, redacted `error`. |
|
|
60
|
+
| Provider stream | Prism text, thinking (`delta.reasoning` / `reasoning_content`), tool-call delta/final, `usage` (with cache read/write mapped), `done`, redacted `error`. |
|
|
48
61
|
| Attribution | `HTTP-Referer`/`X-Title` headers sent only when `appUrl`/`appTitle` are supplied. |
|
|
49
62
|
| Auth method | `api_key` for `openrouter`, credential name `apiKey`. |
|
|
50
63
|
|
|
51
64
|
## Request/response example
|
|
52
65
|
|
|
53
|
-
Per-model routing override:
|
|
66
|
+
Per-model routing + reasoning override:
|
|
54
67
|
|
|
55
68
|
```json
|
|
56
69
|
{
|
|
57
70
|
"model": "anthropic/claude-sonnet-4",
|
|
58
|
-
"
|
|
71
|
+
"provider": { "order": ["anthropic"], "data_collection": "deny" },
|
|
72
|
+
"reasoning": { "effort": "high" },
|
|
73
|
+
"session_id": "session-with-spaces",
|
|
74
|
+
"cache_control": { "type": "ephemeral" }
|
|
59
75
|
}
|
|
60
76
|
```
|
|
61
77
|
|
|
@@ -63,58 +79,107 @@ Per-model routing override:
|
|
|
63
79
|
|
|
64
80
|
```ts
|
|
65
81
|
import { createExtensionKernel } from "@arnilo/prism";
|
|
66
|
-
import {
|
|
82
|
+
import {
|
|
83
|
+
createOpenRouterProviderPackage,
|
|
84
|
+
defineOpenRouterModel,
|
|
85
|
+
listOpenRouterModels,
|
|
86
|
+
} from "@arnilo/prism-provider-openrouter";
|
|
67
87
|
|
|
88
|
+
// App-controlled registration (default — no fetch):
|
|
68
89
|
const sonnet = defineOpenRouterModel({
|
|
69
90
|
model: "anthropic/claude-sonnet-4",
|
|
70
|
-
compat: {
|
|
91
|
+
compat: {
|
|
92
|
+
openRouterRouting: { order: ["anthropic"], data_collection: "deny" },
|
|
93
|
+
openRouterCache: true,
|
|
94
|
+
reasoning: { effort: "medium" },
|
|
95
|
+
},
|
|
71
96
|
});
|
|
72
97
|
|
|
98
|
+
// Optional live discovery — caller-gated, never run by setup:
|
|
99
|
+
const live = await listOpenRouterModels({ apiKey: process.env.OPENROUTER_API_KEY });
|
|
100
|
+
const filtered = live.filter((m) => m.model.startsWith("anthropic/"));
|
|
101
|
+
|
|
73
102
|
const kernel = createExtensionKernel();
|
|
74
103
|
await kernel.load([
|
|
75
|
-
createOpenRouterProviderPackage({
|
|
104
|
+
createOpenRouterProviderPackage({
|
|
105
|
+
apiKey: process.env.OPENROUTER_API_KEY,
|
|
106
|
+
models: filtered.length ? filtered : [sonnet],
|
|
107
|
+
}),
|
|
76
108
|
]);
|
|
77
109
|
```
|
|
78
110
|
|
|
79
111
|
## Extension and configuration notes
|
|
80
112
|
|
|
81
113
|
- Apps supply the model catalog via `models`; no catalog is fetched during setup.
|
|
114
|
+
- `listOpenRouterModels()` is the official `GET https://openrouter.ai/api/v1/models`
|
|
115
|
+
helper (auth optional for the public catalog). Map pricing/context/modalities/
|
|
116
|
+
reasoning metadata into `ModelConfig`; hosts still decide what to register.
|
|
82
117
|
- `defineOpenRouterModel` lets apps override cache policy and routing per model.
|
|
83
118
|
- Hosts choose base URL, attribution, credential source, and `fetch` impl.
|
|
84
119
|
- Package contributes models and an `api_key` auth method.
|
|
85
120
|
|
|
121
|
+
### Reasoning
|
|
122
|
+
|
|
123
|
+
- Body field is the official OpenRouter `reasoning` object
|
|
124
|
+
(`effort`: `max`/`xhigh`/`high`/`medium`/`low`/`minimal`/`none`, plus
|
|
125
|
+
`max_tokens`, `exclude`, `enabled`, `context`, `mode` as documented).
|
|
126
|
+
- Model `compat.reasoning` defaults merge with per-turn `options.compat.reasoning`
|
|
127
|
+
(request keys win). Task 4 `applyThinkingLevel(..., "openai_reasoning")` writes
|
|
128
|
+
`{ reasoning: { effort } }` into that path.
|
|
129
|
+
- Owned compat keys (`reasoning`, `openRouterRouting`, `openRouterCache`,
|
|
130
|
+
`preserveThinking`) are stripped from opaque compat spreads so resolved values
|
|
131
|
+
cannot be overwritten accidentally.
|
|
132
|
+
- When `preserveThinking` is enabled (default for reasoning-capable models),
|
|
133
|
+
assistant `thinking` blocks replay as top-level `reasoning` — not folded into
|
|
134
|
+
text — matching OpenRouter's tool-call continuity guidance.
|
|
135
|
+
|
|
86
136
|
### Cache and session behavior
|
|
87
137
|
|
|
88
138
|
- `session_id` (request body) and the `X-Session-Id` header are derived from
|
|
89
139
|
`ProviderRequestOptions.cacheKey` (falling back to `sessionId`) and sanitized
|
|
90
140
|
+ clamped to 256 characters via the shared `sanitizeCacheKey()` helper.
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
`
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
141
|
+
OpenRouter uses this for provider sticky routing to maximize cache hits.
|
|
142
|
+
- **Automatic caching** (no breakpoints): when caching is enabled for an
|
|
143
|
+
explicit `cache_control` model (or `compat.openRouterCache` /
|
|
144
|
+
`cache.mode: "on"`), Prism emits a top-level
|
|
145
|
+
`cache_control: { type: "ephemeral" }` per OpenRouter's Anthropic automatic
|
|
146
|
+
caching docs. Note: top-level `cache_control` can exclude some backends
|
|
147
|
+
(e.g. Bedrock/Vertex) from routing.
|
|
148
|
+
- **Explicit breakpoints**: Anthropic-style markers are applied only to the
|
|
149
|
+
Prism `PromptCacheBreakpoint` locations the caller selects via
|
|
150
|
+
`ProviderRequestOptions.cache.breakpoints` (last content block of each
|
|
151
|
+
selected message). When breakpoints are present, top-level automatic
|
|
152
|
+
`cache_control` is omitted.
|
|
100
153
|
- Caching is enabled unless disabled (`cacheRetention: "none"` /
|
|
101
154
|
`cache.mode: "off"`) and the model opts in via `ModelConfig.cache.kind`
|
|
102
155
|
(`"cache_control"`) or the legacy `compat.openRouterCache: true` flag.
|
|
103
156
|
- `cacheRetention: "long"` (or `cache.retention: "long"`) emits
|
|
104
|
-
`
|
|
105
|
-
long retention (`ModelConfig.cache.longRetention !== false`)
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
157
|
+
`ttl: "1h"` on markers / top-level automatic control when the model allows
|
|
158
|
+
long retention (`ModelConfig.cache.longRetention !== false`).
|
|
159
|
+
- Usage accounting: OpenRouter `prompt_tokens_details.cached_tokens` →
|
|
160
|
+
`Usage.cacheReadTokens`; `prompt_tokens_details.cache_write_tokens` →
|
|
161
|
+
`Usage.cacheWriteTokens`.
|
|
162
|
+
|
|
163
|
+
### Model discovery
|
|
164
|
+
|
|
165
|
+
```ts
|
|
166
|
+
const models = await listOpenRouterModels({ apiKey, fetch, signal, baseUrl });
|
|
167
|
+
createOpenRouterProviderPackage({ apiKey, models: models.filter(...) });
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
`mapOpenRouterModel` converts official per-token USD pricing to
|
|
171
|
+
`ModelCost` with `unit: "per_million_tokens"`, infers `cache.kind` (`cache_control`
|
|
172
|
+
for Anthropic/Qwen/Gemini families with cache pricing; otherwise `implicit` when
|
|
173
|
+
cache-read pricing exists), and seeds `compat.reasoning.effort` from
|
|
174
|
+
`reasoning.default_effort` when present.
|
|
110
175
|
|
|
111
176
|
## Security and performance notes
|
|
112
177
|
|
|
113
178
|
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
|
|
114
|
-
- No catalog fetch during setup;
|
|
115
|
-
|
|
179
|
+
- No catalog fetch during setup; discovery is caller-gated and bounded.
|
|
180
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
116
181
|
- API keys are resolved per request from caller-supplied values or resolvers and
|
|
117
|
-
redacted from errors.
|
|
182
|
+
redacted from errors (including discovery failures).
|
|
118
183
|
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
|
|
119
184
|
but OpenRouter-owned headers are applied last: `Authorization`,
|
|
120
185
|
`Content-Type`, `X-Session-Id`, `HTTP-Referer`, and `X-Title` cannot be
|
|
@@ -127,9 +192,14 @@ await kernel.load([
|
|
|
127
192
|
## Related APIs
|
|
128
193
|
|
|
129
194
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
130
|
-
`ModelConfig`/`compat`, cache policy,
|
|
195
|
+
`ModelConfig`/`compat`, cache policy, caller-gated discovery.
|
|
196
|
+
- [Thinking and reasoning](../thinking-and-reasoning.md): `applyThinkingLevel`
|
|
197
|
+
/ `openai_reasoning` family for OpenRouter.
|
|
131
198
|
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
132
199
|
`resolveCredentialValue`, `redactSecrets`.
|
|
133
200
|
- [Provider layer](../provider-layer.md): `ProviderRequest.options` and usage
|
|
134
201
|
mapping.
|
|
135
202
|
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
203
|
+
- Official: [Models API](https://openrouter.ai/docs/api/api-reference/models/get-models),
|
|
204
|
+
[Prompt caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching),
|
|
205
|
+
[Reasoning tokens](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens).
|