@arnilo/prism 0.0.5 → 0.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/CHANGELOG.md +39 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +27 -16
  4. package/dist/agent-run-lifecycle.d.ts +28 -0
  5. package/dist/agent-run-lifecycle.js +33 -0
  6. package/dist/agent-run-state.d.ts +53 -0
  7. package/dist/agent-run-state.js +127 -0
  8. package/dist/agents.d.ts +3 -1
  9. package/dist/agents.js +337 -46
  10. package/dist/contracts.d.ts +205 -3
  11. package/dist/contracts.js +4 -0
  12. package/dist/guardrails.d.ts +25 -0
  13. package/dist/guardrails.js +133 -0
  14. package/dist/ids.d.ts +2 -0
  15. package/dist/ids.js +6 -0
  16. package/dist/index.d.ts +17 -3
  17. package/dist/index.js +10 -3
  18. package/dist/input.js +2 -0
  19. package/dist/resources.js +2 -1
  20. package/dist/run-limits.d.ts +34 -0
  21. package/dist/run-limits.js +163 -0
  22. package/dist/secure-agent.d.ts +3 -0
  23. package/dist/secure-agent.js +63 -0
  24. package/dist/session-stores.js +2 -3
  25. package/dist/testing/persistence-schema.d.ts +45 -7
  26. package/dist/testing/persistence-schema.js +138 -24
  27. package/dist/thinking.d.ts +42 -0
  28. package/dist/thinking.js +92 -0
  29. package/dist/tools.d.ts +10 -2
  30. package/dist/tools.js +56 -7
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +4 -2
  34. package/docs/agent-events.md +23 -16
  35. package/docs/agent-loops.md +19 -8
  36. package/docs/agent-session-runtime.md +33 -1
  37. package/docs/coding-agent-tools.md +33 -12
  38. package/docs/coding-security.md +2 -2
  39. package/docs/compaction-llm.md +17 -7
  40. package/docs/compaction-observational-memory.md +28 -4
  41. package/docs/credential-storage.md +58 -9
  42. package/docs/credentials-and-redaction.md +1 -1
  43. package/docs/database-persistence.md +8 -3
  44. package/docs/guardrails.md +75 -0
  45. package/docs/host-security.md +16 -8
  46. package/docs/index.md +26 -22
  47. package/docs/mcp-tools.md +32 -12
  48. package/docs/migration.md +164 -2
  49. package/docs/node-filesystem-config.md +1 -0
  50. package/docs/node-jsonl-session-store.md +5 -4
  51. package/docs/postgres-persistence.md +3 -3
  52. package/docs/provider-caching.md +16 -4
  53. package/docs/provider-conformance.md +39 -1
  54. package/docs/provider-packages.md +60 -3
  55. package/docs/providers/ai-sdk.md +36 -0
  56. package/docs/providers/kimi.md +124 -61
  57. package/docs/providers/neuralwatt.md +19 -13
  58. package/docs/providers/openai.md +56 -13
  59. package/docs/providers/opencode-go.md +118 -30
  60. package/docs/providers/openrouter.md +105 -35
  61. package/docs/providers/zai.md +94 -45
  62. package/docs/release-and-install.md +47 -49
  63. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  64. package/docs/runs-and-usage.md +30 -3
  65. package/docs/server.md +5 -2
  66. package/docs/sqlite-persistence.md +2 -2
  67. package/docs/structured-output.md +1 -1
  68. package/docs/thinking-and-reasoning.md +98 -0
  69. package/docs/tool-execution-primitives.md +3 -3
  70. package/docs/tools.md +21 -1
  71. package/docs/use-case-model-selection.md +109 -0
  72. package/docs/workflow-orchestration-primitives.md +1 -0
  73. package/docs/workflows.md +18 -10
  74. package/docs/working-and-semantic-memory.md +1 -0
  75. package/package.json +2 -2
@@ -9,7 +9,7 @@ and implicit prefix caching.
9
9
 
10
10
  The package registers a provider, default model metadata for the featured NeuralWatt
11
11
  aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
12
- `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
12
+ `gemma-4-31b`, `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
13
13
  `qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
14
14
  method through `createExtensionKernel().load([...])`.
15
15
 
@@ -79,13 +79,15 @@ The endpoint is rate-limited to **1 request per second per customer** (429 with
79
79
  from `generate()` or package setup; the caller owns throttling.
80
80
 
81
81
  NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
82
- / `extra` escape hatches: `compat.reasoning_effort` (`"low" | "medium" | "high"`),
83
- `compat.thinking_token_budget`, `compat.chat_template_kwargs` (including `enable_thinking`),
84
- `compat.preserve_thinking`, `compat.clear_thinking`, and `compat.tool_choice`.
85
- `preserve_thinking: true` keeps prior assistant reasoning in request history so
86
- multi-turn reasoning continues with the earlier chain of thought; `clear_thinking:
87
- true` drops it for the next turn, resetting the chain. `options.extra` spreads after
88
- `compat` so per-call values and overrides win.
82
+ / `extra` escape hatches: `compat.reasoning_effort` (OpenAI-style scale; GLM-5.2 defaults
83
+ unset to `max` per official docs), `compat.thinking_token_budget`,
84
+ `compat.chat_template_kwargs` (including `enable_thinking`, `preserve_thinking`, and
85
+ `clear_thinking` per official gateway docs), and `compat.tool_choice`.
86
+ `preserve_thinking` / `clear_thinking` compat flags are routed into `chat_template_kwargs`
87
+ on the wire (not top-level body fields). Prism also uses these flags for client-side
88
+ message serialization: `preserve_thinking: true` keeps prior assistant reasoning in
89
+ request history; `clear_thinking: true` drops it for the next turn. `options.extra` spreads
90
+ after resolved fields so per-call values and overrides win.
89
91
 
90
92
  ## Outputs / response / events
91
93
 
@@ -112,7 +114,7 @@ Example request body (OpenAI-compatible Chat Completions shape):
112
114
  "messages": [{ "role": "user", "content": "Hello" }],
113
115
  "stream": true,
114
116
  "stream_options": { "include_usage": true },
115
- "reasoning_effort": "medium",
117
+ "reasoning_effort": "max",
116
118
  "thinking_token_budget": 8192
117
119
  }
118
120
  ```
@@ -192,6 +194,7 @@ validation. The caller owns throttling/caching — the helper makes one explicit
192
194
  | `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
193
195
  | `glm-5.2-short` | 195K | Tools, reasoning |
194
196
  | `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
197
+ | `gemma-4-31b` | 256K | Tools, vision, JSON mode |
195
198
  | `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
196
199
  | `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
197
200
  | `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
@@ -255,15 +258,18 @@ so multi-turn sessions continue the earlier chain of thought:
255
258
 
256
259
  - Prior `thinking` content blocks on an assistant message are serialized under a
257
260
  `reasoning_content` field on that message (matching the streaming
258
- `delta.reasoning_content` field). They are **not** flattened into text `content`, so
261
+ `delta.reasoning_content` field; the gateway also accepts `reasoning` as an alias).
262
+ They are **not** flattened into text `content`, so
259
263
  the model sees reasoning and answer as distinct.
260
264
  - Preservation is gated on `model.capabilities.reasoning === true` **or**
261
265
  `compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
262
266
  field and prior `thinking` blocks are dropped — they never leak into text content for
263
267
  providers/models that do not support reasoning.
264
- - `compat.clear_thinking: true` drops prior reasoning for the next turn even on
265
- reasoning-capable models, resetting the chain of thought. `clear_thinking` takes
266
- precedence over `preserve_thinking`.
268
+ - `compat.preserve_thinking` / `compat.clear_thinking` map into `chat_template_kwargs`
269
+ on the request body per official NeuralWatt docs (Kimi K2.6 `preserve_thinking`, GLM
270
+ `clear_thinking: false` for full-history). Prism also uses `clear_thinking: true` to
271
+ drop prior reasoning client-side even on reasoning-capable models; `clear_thinking`
272
+ takes precedence over `preserve_thinking`.
267
273
  - The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
268
274
  reasoning.
269
275
 
@@ -36,16 +36,22 @@ createOpenAIProviderPackage(options: OpenAIProviderPackageOptions): ProviderPack
36
36
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
37
37
  | `baseUrl` | `string` | Overrides `https://api.openai.com/v1`. |
38
38
  | `codexBaseUrl` | `string` | Overrides `https://chatgpt.com/backend-api/codex`. |
39
+ | `models` | `readonly ModelConfig[]` | Optional override for registered OpenAI Responses models (defaults to featured `openAIModels`). |
40
+ | `codexModels` | `readonly ModelConfig[]` | Optional override for registered Codex models (defaults to featured `openAICodexModels`). |
39
41
 
40
42
  `ProviderRequest.options.sessionId`, `cacheKey`, `cacheRetention`, `headers`,
41
- `compat`, and `extra` map to request headers/payload fields.
43
+ `compat`, and `extra` map to request headers/payload fields. Per-turn reasoning
44
+ uses official Responses `reasoning: { effort, summary? }` via
45
+ `ModelConfig.compat.reasoning` defaults merged with
46
+ `ProviderRequestOptions.compat.reasoning` (request wins). Prefer
47
+ `applyThinkingLevel(..., "openai_reasoning")` from `@arnilo/prism`.
42
48
 
43
49
  ## Outputs / response / events
44
50
 
45
51
  | Surface | Behavior |
46
52
  | --- | --- |
47
53
  | Provider stream | Prism text, thinking (downgraded to text), `tool_call` deltas/finals, `usage`, `done`, redacted `error` events. |
48
- | Block preservation | Text, thinking (downgraded), assistant `tool_call` → `function_call` input items, `tool_result` → `function_call_output` input items, images when `capabilities.input` includes `"image"`. |
54
+ | Block preservation | User/system text → `input_text`; assistant text → `output_text`; assistant `tool_call` → top-level `function_call` with `call_id`; `tool_result` → top-level `function_call_output`; images/files/audio when declared on the model. Bare thinking without an encrypted Responses reasoning item is omitted on replay. |
49
55
  | Auth methods | `api_key` for `openai`; `oauth` for `openai-codex`. |
50
56
 
51
57
  Unsupported block placements or unclaimed images fail before `fetch`.
@@ -56,10 +62,17 @@ Responses request body (Codex subscription shape, abbreviated):
56
62
 
57
63
  ```json
58
64
  {
59
- "model": "gpt-5-codex",
60
- "instructions": "You are a coding agent.",
61
- "input": [{ "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] }],
62
- "stream": true
65
+ "model": "gpt-5.1",
66
+ "input": [
67
+ { "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] },
68
+ { "role": "assistant", "content": [{ "type": "output_text", "text": "Calling lookup" }] },
69
+ { "type": "function_call", "call_id": "call_1", "name": "lookup", "arguments": "{\"q\":\"x\"}" },
70
+ { "type": "function_call_output", "call_id": "call_1", "output": "{\"ok\":true}" }
71
+ ],
72
+ "reasoning": { "effort": "high" },
73
+ "prompt_cache_key": "session-1",
74
+ "stream": true,
75
+ "store": false
63
76
  }
64
77
  ```
65
78
 
@@ -73,13 +86,13 @@ https://auth.openai.com/authorize?response_type=code&client_id=...&code_challeng
73
86
 
74
87
  ```ts
75
88
  import { createExtensionKernel, createEnvCredentialResolver } from "@arnilo/prism";
76
- import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
89
+ import { createOpenAIProviderPackage, listOpenAIModels } from "@arnilo/prism-provider-openai";
77
90
 
91
+ const apiKey = createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" });
92
+ const models = await listOpenAIModels({ apiKey }); // caller-gated; never runs during setup
78
93
  const kernel = createExtensionKernel();
79
94
  await kernel.load([
80
- createOpenAIProviderPackage({
81
- apiKey: createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" }),
82
- }),
95
+ createOpenAIProviderPackage({ apiKey, models }),
83
96
  ]);
84
97
  ```
85
98
 
@@ -115,25 +128,55 @@ const challenge = computeS256Challenge(verifier);
115
128
 
116
129
  ### Cache behavior
117
130
 
131
+ Official: [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching).
132
+
118
133
  - `prompt_cache_key` is derived from `ProviderRequestOptions.cacheKey` (falling
119
134
  back to `sessionId`) and sanitized + clamped to 64 characters via the shared
120
135
  `sanitizeCacheKey()` helper. Cache keys are session/customer identifiers only;
121
136
  never credentials or raw prompts.
122
- - `prompt_cache_retention` accepts only `"24h"` on the OpenAI Responses API
137
+ - `prompt_cache_retention` accepts only `"24h"` on pre-GPT-5.6 Responses models
123
138
  (extended caching). Prism `cacheRetention: "short"` and `"none"` omit the field
124
139
  so default automatic/implicit caching applies and no invalid literal is sent.
125
140
  `cacheRetention: "long"` maps to `prompt_cache_retention: "24h"` only when the
126
141
  model declares `ModelConfig.cache.longRetention === true`; models without that
127
- metadata omit the field. The catalog `gpt-5.1` model declares
142
+ metadata omit the field. Featured `gpt-5.1` declares
128
143
  `cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
144
+ - GPT-5.6+ official docs prefer `prompt_cache_options` / explicit breakpoints;
145
+ `listOpenAIModels` sets `longRetention: false` for those ids so Prism does not
146
+ emit deprecated `prompt_cache_retention` for them. Breakpoint helpers are not
147
+ shipped in this package yet — hosts may pass `prompt_cache_options` through
148
+ `compat` / `extra` when needed.
129
149
  - Cache accounting is preserved in normalized `Usage`: OpenAI
130
150
  `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
131
- Responses does not report a cache-write token field.
151
+ Responses does not report a cache-write token field on older models.
132
152
  - Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
133
153
  are applied after caller `ProviderRequestOptions.headers` so caller config
134
154
  cannot replace credentials, content type, or the session request id; non-owned
135
155
  caller headers are kept.
136
156
 
157
+ ### Model discovery
158
+
159
+ - `listOpenAIModels({ apiKey, fetch, baseUrl, signal, headers })` calls official
160
+ [`GET /models`](https://developers.openai.com/api/reference/resources/models/methods/list)
161
+ and maps sparse `{ id, created, owned_by }` entries to `ModelConfig` with
162
+ `cache.kind: "openai_key"` and heuristic `longRetention` / `capabilities.reasoning`.
163
+ - `createOpenAIProviderPackage` never calls discovery; pass results via `models:`.
164
+ - Codex subscription models are **not** listed by `api.openai.com` — keep using
165
+ featured `openAICodexModels` or `codexModels:` override.
166
+ - Static `openAIModels` / `openAICodexModels` are offline bootstrap / featured aliases only.
167
+
168
+ ### Reasoning
169
+
170
+ Official: [Reasoning models](https://developers.openai.com/api/docs/guides/reasoning).
171
+
172
+ - Body field is top-level `reasoning: { effort, summary?, mode?, context? }`.
173
+ - Model defaults: `ModelConfig.compat.reasoning`; per-turn override:
174
+ `ProviderRequestOptions.compat.reasoning` (shallow-merged; request wins).
175
+ - Portable helper: `applyThinkingLevel(options, level, "openai_reasoning")`.
176
+ - Streaming tool args follow official
177
+ `response.output_item.added` + `response.function_call_arguments.delta`
178
+ (string `delta`), not Chat Completions object deltas.
179
+
137
180
  ## Security and performance notes
138
181
 
139
182
  - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
@@ -2,27 +2,41 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-provider-opencode-go` provides explicit, side-effect-free setup for the
6
- OpenCode Go API-key provider using Prism model metadata and
7
- OpenAI-compatible/Anthropic-compatible routes with `x-opencode-session`
8
- cache/session headers.
5
+ `@arnilo/prism-provider-opencode-go` provides explicit, side-effect-free setup for
6
+ [OpenCode Go](https://opencode.ai/docs/go/) — a low-cost subscription gateway for
7
+ open coding models. The package dual-routes by `ModelConfig.compat.route`:
9
8
 
10
- The package registers a provider, default model metadata, and an `api_key` auth
11
- method through `createExtensionKernel().load([...])`.
9
+ | Route | Endpoint | Official model families |
10
+ | --- | --- | --- |
11
+ | `"openai"` (default) | `POST {baseUrl}/chat/completions` | Grok, GLM, Kimi, MiMo, DeepSeek |
12
+ | `"anthropic"` | `POST {baseUrl}/messages` | MiniMax, Qwen |
13
+
14
+ Default base URL is the official Go API root:
15
+
16
+ ```txt
17
+ https://opencode.ai/zen/go/v1
18
+ ```
19
+
20
+ Session stickiness uses the sanitized `x-opencode-session` header. Anthropic-route
21
+ models may emit selected `cache_control` breakpoints; OpenAI-route models use
22
+ implicit caching and never receive Anthropic cache fields.
12
23
 
13
24
  ## When to use it
14
25
 
15
- Use it when a host app wants to run an OpenAI-compatible or Anthropic-compatible
16
- OpenCode Go endpoint through Prism's `AgentSession` runtime with per-request
17
- session/cache headers.
26
+ Use it when a host app wants OpenCode Go models through Prism's `AgentSession`
27
+ runtime with dual-route serialization, per-request session headers, and optional
28
+ caller-gated model discovery.
18
29
 
19
- Do not use it for automatic credential discovery, catalog fetches, or
30
+ Do not use it for automatic credential discovery, setup-time catalog fetches, or
20
31
  real-network tests.
21
32
 
22
33
  ## Inputs / request
23
34
 
24
35
  ```ts
25
- import { createOpenCodeGoProviderPackage } from "@arnilo/prism-provider-opencode-go";
36
+ import {
37
+ createOpenCodeGoProviderPackage,
38
+ listOpenCodeGoModels,
39
+ } from "@arnilo/prism-provider-opencode-go";
26
40
 
27
41
  createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): ProviderPackage
28
42
  ```
@@ -31,57 +45,126 @@ createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): Prov
31
45
  | --- | --- | --- |
32
46
  | `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
33
47
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
34
- | `baseUrl` | `string` | Overrides the OpenCode Go base URL. |
35
- | `models` | `readonly ModelConfig[]` | Overrides `openCodeGoModels` defaults. |
48
+ | `baseUrl` | `string` | Overrides official `https://opencode.ai/zen/go/v1`. |
49
+ | `models` | `readonly ModelConfig[]` | Overrides featured `openCodeGoModels` defaults. |
36
50
 
37
51
  `ProviderRequest.options.cacheKey` (falling back to `sessionId`) maps to the
38
- `x-opencode-session` header; the Anthropic-compatible route accepts
39
- `cache_control` breakpoints; `cacheRetention` maps to cache retention.
52
+ `x-opencode-session` header. Anthropic-route `cache_control` breakpoints and
53
+ `cacheRetention` map as documented below.
40
54
 
41
55
  ## Outputs / response / events
42
56
 
43
57
  | Surface | Behavior |
44
58
  | --- | --- |
45
59
  | Provider stream | Prism text, thinking, tool-call delta/final, `usage`, `done`, redacted `error`. |
46
- | Session/cache | `x-opencode-session` and cache headers added before `generate()`. |
60
+ | OpenAI thinking | `delta.reasoning_content` → thinking deltas; replay via `reasoning_content` when `preserveThinking`. |
61
+ | Anthropic thinking | `thinking_delta` → thinking deltas; replay via Anthropic thinking blocks when `preserveThinking`. |
62
+ | Session/cache | `x-opencode-session` + route-specific cache markers. |
47
63
  | Auth method | `api_key` for `opencode-go`, credential name `apiKey`. |
48
64
 
49
65
  ## Request/response example
50
66
 
51
- Example headers added before fetch:
52
-
53
67
  ```json
54
68
  {
55
69
  "Authorization": "Bearer <resolved-key>",
70
+ "content-type": "application/json",
56
71
  "x-opencode-session": "<ProviderRequest.options.cacheKey ?? sessionId>"
57
72
  }
58
73
  ```
59
74
 
75
+ OpenAI-route body (thinking passthrough + preserved reasoning):
76
+
77
+ ```json
78
+ {
79
+ "model": "kimi-k3",
80
+ "stream": true,
81
+ "stream_options": { "include_usage": true },
82
+ "reasoning_effort": "high",
83
+ "messages": [
84
+ {
85
+ "role": "assistant",
86
+ "content": "calling",
87
+ "tool_calls": [{ "id": "call_1", "type": "function", "function": { "name": "lookup", "arguments": "{\"q\":\"x\"}" } }],
88
+ "reasoning_content": "plan the lookup"
89
+ }
90
+ ]
91
+ }
92
+ ```
93
+
60
94
  ## Implementation example
61
95
 
62
96
  ```ts
63
97
  import { createExtensionKernel } from "@arnilo/prism";
64
- import { createOpenCodeGoProviderPackage } from "@arnilo/prism-provider-opencode-go";
98
+ import {
99
+ createOpenCodeGoProviderPackage,
100
+ listOpenCodeGoModels,
101
+ openCodeGoModels,
102
+ } from "@arnilo/prism-provider-opencode-go";
65
103
 
66
104
  const kernel = createExtensionKernel();
67
- await kernel.load([createOpenCodeGoProviderPackage({ apiKey: "fake-opencode-key" })]);
105
+ await kernel.load([createOpenCodeGoProviderPackage({ apiKey: process.env.OPENCODE_API_KEY })]);
68
106
  ```
69
107
 
70
- Override model metadata:
108
+ Caller-gated live catalog (never runs during package setup):
71
109
 
72
110
  ```ts
73
- import { createOpenCodeGoProviderPackage, openCodeGoModels } from "@arnilo/prism-provider-opencode-go";
111
+ const models = await listOpenCodeGoModels({ apiKey: process.env.OPENCODE_API_KEY });
112
+ await kernel.load([createOpenCodeGoProviderPackage({ apiKey: process.env.OPENCODE_API_KEY, models })]);
113
+ ```
114
+
115
+ Offline bootstrap with featured docs-verified aliases:
74
116
 
117
+ ```ts
75
118
  await kernel.load([
76
119
  createOpenCodeGoProviderPackage({ apiKey: "fake", models: openCodeGoModels }),
77
120
  ]);
78
121
  ```
79
122
 
123
+ ## Featured models and routes
124
+
125
+ Featured `openCodeGoModels` mirrors the official Go docs list (open coding models
126
+ only — **not** Zen GPT/Claude ids). Route selection follows the official endpoint
127
+ table; Pi secondary metadata is used only for context/output limits when docs omit them.
128
+
129
+ | Model ID | Route | Cache kind |
130
+ | --- | --- | --- |
131
+ | `grok-4.5`, `glm-5.2`, `glm-5.1`, `kimi-k3`, `kimi-k2.7-code`, `kimi-k2.6`, `mimo-v2.5`, `mimo-v2.5-pro`, `deepseek-v4-pro`, `deepseek-v4-flash` | `openai` | `implicit` |
132
+ | `minimax-m3`, `minimax-m2.7`, `minimax-m2.5`, `qwen3.7-max`, `qwen3.7-plus`, `qwen3.6-plus` | `anthropic` | `cache_control` |
133
+
134
+ ## Model discovery
135
+
136
+ Official list endpoint (sparse OpenAI-compatible shape):
137
+
138
+ ```txt
139
+ GET https://opencode.ai/zen/go/v1/models
140
+ ```
141
+
142
+ `listOpenCodeGoModels({ apiKey?, fetch?, baseUrl?, signal?, headers? })` maps each
143
+ `{ id, owned_by }` entry to `ModelConfig` with route/cache heuristics from the docs
144
+ endpoint table. Featured metadata (pricing/limits/thinking defaults) is applied when
145
+ the id matches `openCodeGoModels`. Discovery is **caller-gated** — setup performs
146
+ zero fetches.
147
+
148
+ ## Thinking / reasoning
149
+
150
+ OpenCode Go does not document gateway-owned thinking fields; Prism forwards
151
+ upstream-compatible compat and preserves prior reasoning for tool-call continuity:
152
+
153
+ | Surface | Behavior |
154
+ | --- | --- |
155
+ | OpenAI route stream | `reasoning_content` → thinking deltas |
156
+ | OpenAI route replay | thinking blocks → top-level `reasoning_content` when `preserveThinking` (default for reasoning models); never folded into text |
157
+ | OpenAI route body | optional `thinking` / `reasoning_effort` / `reasoning` from model + per-turn `options.compat` (request wins) |
158
+ | Anthropic route stream | `thinking_delta` → thinking deltas |
159
+ | Anthropic route replay | thinking blocks with optional `signature` when `preserveThinking` |
160
+
161
+ Owned compat keys (`route`, `thinking`, `reasoning`, `reasoning_effort`,
162
+ `preserveThinking`) are stripped before opaque compat spread so resolved values win.
163
+
80
164
  ## Extension and configuration notes
81
165
 
82
166
  - Hosts choose base URL, model list, credential source, and `fetch` impl.
83
- - The serializer is inherited from the OpenAI-compatible route; Anthropic-compatible
84
- routes preserve `tool_use`/`tool_result` blocks.
167
+ - Route selection is explicit via `compat.route` (`"anthropic"` or default `"openai"`).
85
168
  - Package contributes models via the extension `api` and an `api_key` auth method.
86
169
 
87
170
  ### Cache and session behavior
@@ -114,19 +197,24 @@ await kernel.load([
114
197
  - No network calls during import, setup, build, or default tests.
115
198
  - No automatic environment, file, keychain, or shell credential lookup.
116
199
  - API keys are resolved per request from caller-supplied values or resolvers and
117
- redacted from errors.
200
+ redacted from errors (including discovery failures).
118
201
  - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
119
202
  provider-owned headers (`content-type`, `x-opencode-session`, `authorization`)
120
203
  are applied last and cannot be overridden by caller headers.
121
- - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
122
- provider-specific env names; default tests are network-free.
204
+ - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus `OPENCODE_API_KEY`;
205
+ default tests are network-free.
206
+
207
+ ## Official evidence
208
+
209
+ - [OpenCode Go](https://opencode.ai/docs/go/) — model list, dual endpoints, pricing/usage, `GET /zen/go/v1/models`
210
+ - Pi secondary (ids/limits only): `packages/ai/src/providers/opencode-go.ts`, `opencode-go.models.ts`
123
211
 
124
212
  ## Related APIs
125
213
 
126
214
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
127
- `ModelConfig`, request/cache policies.
215
+ `ModelConfig`, discovery contract, request/cache policies.
216
+ - [Thinking and reasoning](../thinking-and-reasoning.md): per-turn `ThinkingLevel` → compat families.
128
217
  - [Credentials and redaction](../credentials-and-redaction.md):
129
218
  `resolveCredentialValue`, `redactSecrets`.
130
- - [OpenAI-compatible provider](openai-compatible.md): underlying Chat Completions
131
- adapter.
219
+ - [Provider caching](../provider-caching.md): route-specific OpenCode Go cache matrix.
132
220
  - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -3,12 +3,15 @@
3
3
  ## What it does
4
4
 
5
5
  `@arnilo/prism-provider-openrouter` provides explicit, side-effect-free setup for the
6
- OpenRouter API-key provider with app-controlled model catalog and per-model cache
7
- policy/routing overrides.
6
+ OpenRouter API-key provider with **app-controlled** model registration, routing
7
+ passthrough, official `reasoning` controls, and Anthropic-style `cache_control`
8
+ (plus sticky `session_id` routing).
8
9
 
9
10
  The package registers a provider, caller-supplied model metadata, and an
10
- `api_key` auth method through `createExtensionKernel().load([...])`. Apps control the model catalog instead of
11
- accepting a fetched or hard-coded one.
11
+ `api_key` auth method through `createExtensionKernel().load([...])`. There is
12
+ **no bundled mega-catalog**. Optional `listOpenRouterModels()` lets hosts fetch
13
+ the live official catalog and pass a filtered subset via `models:` — setup
14
+ itself never fetches.
12
15
 
13
16
  ## When to use it
14
17
 
@@ -16,16 +19,22 @@ Use it when a host app wants OpenRouter routing passthrough, reasoning controls,
16
19
  and per-model cache policy through Prism's `AgentSession` runtime, and needs to
17
20
  override cache behavior per model rather than accept a single hard-coded policy.
18
21
 
19
- Do not use it for catalog fetches, automatic credential discovery, or
20
- real-network tests.
22
+ Do not use it for automatic catalog fetch during setup, automatic credential
23
+ discovery, or real-network tests in CI defaults.
21
24
 
22
25
  ## Inputs / request
23
26
 
24
27
  ```ts
25
- import { createOpenRouterProviderPackage, defineOpenRouterModel } from "@arnilo/prism-provider-openrouter";
28
+ import {
29
+ createOpenRouterProviderPackage,
30
+ defineOpenRouterModel,
31
+ listOpenRouterModels,
32
+ } from "@arnilo/prism-provider-openrouter";
26
33
 
27
34
  createOpenRouterProviderPackage(options: OpenRouterProviderPackageOptions): ProviderPackage
28
- defineOpenRouterModel(config: OpenRouterModelConfig): OpenRouterModelConfig
35
+ defineOpenRouterModel(config: OpenRouterModelConfig): ModelConfig
36
+ listOpenRouterModels(options?: ListOpenRouterModelsOptions): Promise<ModelConfig[]>
37
+ mapOpenRouterModel(entry: OpenRouterModelEntry): ModelConfig
29
38
  ```
30
39
 
31
40
  | Field | Type | Purpose |
@@ -37,25 +46,32 @@ defineOpenRouterModel(config: OpenRouterModelConfig): OpenRouterModelConfig
37
46
  | `appTitle` | `string` | App title for the `X-Title` attribution header. |
38
47
  | `models` | `readonly ModelConfig[]` | App-supplied model catalog (no default fetch). |
39
48
 
40
- `OpenRouterModelConfig.compat.openRouterRouting` controls routing order,
41
- `data_collection`, reasoning, and per-model cache policy overrides.
49
+ `OpenRouterModelConfig.compat.openRouterRouting` controls routing order /
50
+ `data_collection`. `compat.reasoning` carries the official OpenRouter
51
+ `reasoning` object (`effort`, `max_tokens`, `exclude`, …). Per-turn
52
+ `providerOptions.compat.reasoning` merges over model defaults (request wins
53
+ key-by-key). `compat.preserveThinking` replays assistant thinking as body
54
+ `reasoning` for tool-call continuity.
42
55
 
43
56
  ## Outputs / response / events
44
57
 
45
58
  | Surface | Behavior |
46
59
  | --- | --- |
47
- | Provider stream | Prism text, thinking, tool-call delta/final, `usage` (with cache read/write mapped), `done`, redacted `error`. |
60
+ | Provider stream | Prism text, thinking (`delta.reasoning` / `reasoning_content`), tool-call delta/final, `usage` (with cache read/write mapped), `done`, redacted `error`. |
48
61
  | Attribution | `HTTP-Referer`/`X-Title` headers sent only when `appUrl`/`appTitle` are supplied. |
49
62
  | Auth method | `api_key` for `openrouter`, credential name `apiKey`. |
50
63
 
51
64
  ## Request/response example
52
65
 
53
- Per-model routing override:
66
+ Per-model routing + reasoning override:
54
67
 
55
68
  ```json
56
69
  {
57
70
  "model": "anthropic/claude-sonnet-4",
58
- "routing": { "order": ["anthropic"], "data_collection": "deny" }
71
+ "provider": { "order": ["anthropic"], "data_collection": "deny" },
72
+ "reasoning": { "effort": "high" },
73
+ "session_id": "session-with-spaces",
74
+ "cache_control": { "type": "ephemeral" }
59
75
  }
60
76
  ```
61
77
 
@@ -63,58 +79,107 @@ Per-model routing override:
63
79
 
64
80
  ```ts
65
81
  import { createExtensionKernel } from "@arnilo/prism";
66
- import { createOpenRouterProviderPackage, defineOpenRouterModel } from "@arnilo/prism-provider-openrouter";
82
+ import {
83
+ createOpenRouterProviderPackage,
84
+ defineOpenRouterModel,
85
+ listOpenRouterModels,
86
+ } from "@arnilo/prism-provider-openrouter";
67
87
 
88
+ // App-controlled registration (default — no fetch):
68
89
  const sonnet = defineOpenRouterModel({
69
90
  model: "anthropic/claude-sonnet-4",
70
- compat: { openRouterRouting: { order: ["anthropic"], data_collection: "deny" } },
91
+ compat: {
92
+ openRouterRouting: { order: ["anthropic"], data_collection: "deny" },
93
+ openRouterCache: true,
94
+ reasoning: { effort: "medium" },
95
+ },
71
96
  });
72
97
 
98
+ // Optional live discovery — caller-gated, never run by setup:
99
+ const live = await listOpenRouterModels({ apiKey: process.env.OPENROUTER_API_KEY });
100
+ const filtered = live.filter((m) => m.model.startsWith("anthropic/"));
101
+
73
102
  const kernel = createExtensionKernel();
74
103
  await kernel.load([
75
- createOpenRouterProviderPackage({ apiKey: "fake-openrouter-key", models: [sonnet] }),
104
+ createOpenRouterProviderPackage({
105
+ apiKey: process.env.OPENROUTER_API_KEY,
106
+ models: filtered.length ? filtered : [sonnet],
107
+ }),
76
108
  ]);
77
109
  ```
78
110
 
79
111
  ## Extension and configuration notes
80
112
 
81
113
  - Apps supply the model catalog via `models`; no catalog is fetched during setup.
114
+ - `listOpenRouterModels()` is the official `GET https://openrouter.ai/api/v1/models`
115
+ helper (auth optional for the public catalog). Map pricing/context/modalities/
116
+ reasoning metadata into `ModelConfig`; hosts still decide what to register.
82
117
  - `defineOpenRouterModel` lets apps override cache policy and routing per model.
83
118
  - Hosts choose base URL, attribution, credential source, and `fetch` impl.
84
119
  - Package contributes models and an `api_key` auth method.
85
120
 
121
+ ### Reasoning
122
+
123
+ - Body field is the official OpenRouter `reasoning` object
124
+ (`effort`: `max`/`xhigh`/`high`/`medium`/`low`/`minimal`/`none`, plus
125
+ `max_tokens`, `exclude`, `enabled`, `context`, `mode` as documented).
126
+ - Model `compat.reasoning` defaults merge with per-turn `options.compat.reasoning`
127
+ (request keys win). Task 4 `applyThinkingLevel(..., "openai_reasoning")` writes
128
+ `{ reasoning: { effort } }` into that path.
129
+ - Owned compat keys (`reasoning`, `openRouterRouting`, `openRouterCache`,
130
+ `preserveThinking`) are stripped from opaque compat spreads so resolved values
131
+ cannot be overwritten accidentally.
132
+ - When `preserveThinking` is enabled (default for reasoning-capable models),
133
+ assistant `thinking` blocks replay as top-level `reasoning` — not folded into
134
+ text — matching OpenRouter's tool-call continuity guidance.
135
+
86
136
  ### Cache and session behavior
87
137
 
88
138
  - `session_id` (request body) and the `X-Session-Id` header are derived from
89
139
  `ProviderRequestOptions.cacheKey` (falling back to `sessionId`) and sanitized
90
140
  + clamped to 256 characters via the shared `sanitizeCacheKey()` helper.
91
- Session ids route requests and identify conversations; never credentials or
92
- raw prompts.
93
- - Anthropic-style `cache_control: { type: "ephemeral" }` markers are applied only
94
- to the Prism `PromptCacheBreakpoint` locations the caller selects via
95
- `ProviderRequestOptions.cache.breakpoints` (resolved with the shared
96
- `applyCacheControl()` helper), and only on the last content block of each
97
- selected message — not to every content block of every message. With no
98
- breakpoints, no markers are emitted and the provider relies on implicit prefix
99
- caching where available.
141
+ OpenRouter uses this for provider sticky routing to maximize cache hits.
142
+ - **Automatic caching** (no breakpoints): when caching is enabled for an
143
+ explicit `cache_control` model (or `compat.openRouterCache` /
144
+ `cache.mode: "on"`), Prism emits a top-level
145
+ `cache_control: { type: "ephemeral" }` per OpenRouter's Anthropic automatic
146
+ caching docs. Note: top-level `cache_control` can exclude some backends
147
+ (e.g. Bedrock/Vertex) from routing.
148
+ - **Explicit breakpoints**: Anthropic-style markers are applied only to the
149
+ Prism `PromptCacheBreakpoint` locations the caller selects via
150
+ `ProviderRequestOptions.cache.breakpoints` (last content block of each
151
+ selected message). When breakpoints are present, top-level automatic
152
+ `cache_control` is omitted.
100
153
  - Caching is enabled unless disabled (`cacheRetention: "none"` /
101
154
  `cache.mode: "off"`) and the model opts in via `ModelConfig.cache.kind`
102
155
  (`"cache_control"`) or the legacy `compat.openRouterCache: true` flag.
103
156
  - `cacheRetention: "long"` (or `cache.retention: "long"`) emits
104
- `cache_control: { type: "ephemeral", ttl: "1h" }` markers when the model allows
105
- long retention (`ModelConfig.cache.longRetention !== false`); otherwise the
106
- default 5-minute ephemeral window applies.
107
- - Usage accounting is preserved: OpenRouter `prompt_tokens_details.cached_tokens`
108
- maps to `Usage.cacheReadTokens` and `prompt_tokens_details.cache_write_tokens`
109
- maps to `Usage.cacheWriteTokens`.
157
+ `ttl: "1h"` on markers / top-level automatic control when the model allows
158
+ long retention (`ModelConfig.cache.longRetention !== false`).
159
+ - Usage accounting: OpenRouter `prompt_tokens_details.cached_tokens` →
160
+ `Usage.cacheReadTokens`; `prompt_tokens_details.cache_write_tokens` →
161
+ `Usage.cacheWriteTokens`.
162
+
163
+ ### Model discovery
164
+
165
+ ```ts
166
+ const models = await listOpenRouterModels({ apiKey, fetch, signal, baseUrl });
167
+ createOpenRouterProviderPackage({ apiKey, models: models.filter(...) });
168
+ ```
169
+
170
+ `mapOpenRouterModel` converts official per-token USD pricing to
171
+ `ModelCost` with `unit: "per_million_tokens"`, infers `cache.kind` (`cache_control`
172
+ for Anthropic/Qwen/Gemini families with cache pricing; otherwise `implicit` when
173
+ cache-read pricing exists), and seeds `compat.reasoning.effort` from
174
+ `reasoning.default_effort` when present.
110
175
 
111
176
  ## Security and performance notes
112
177
 
113
178
  - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
114
- - No catalog fetch during setup; no automatic environment, file, keychain, or shell
115
- credential lookup.
179
+ - No catalog fetch during setup; discovery is caller-gated and bounded.
180
+ - No automatic environment, file, keychain, or shell credential lookup.
116
181
  - API keys are resolved per request from caller-supplied values or resolvers and
117
- redacted from errors.
182
+ redacted from errors (including discovery failures).
118
183
  - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
119
184
  but OpenRouter-owned headers are applied last: `Authorization`,
120
185
  `Content-Type`, `X-Session-Id`, `HTTP-Referer`, and `X-Title` cannot be
@@ -127,9 +192,14 @@ await kernel.load([
127
192
  ## Related APIs
128
193
 
129
194
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
130
- `ModelConfig`/`compat`, cache policy, request policies.
195
+ `ModelConfig`/`compat`, cache policy, caller-gated discovery.
196
+ - [Thinking and reasoning](../thinking-and-reasoning.md): `applyThinkingLevel`
197
+ / `openai_reasoning` family for OpenRouter.
131
198
  - [Credentials and redaction](../credentials-and-redaction.md):
132
199
  `resolveCredentialValue`, `redactSecrets`.
133
200
  - [Provider layer](../provider-layer.md): `ProviderRequest.options` and usage
134
201
  mapping.
135
202
  - [Provider conformance](../provider-conformance.md): network-free adapter tests.
203
+ - Official: [Models API](https://openrouter.ai/docs/api/api-reference/models/get-models),
204
+ [Prompt caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching),
205
+ [Reasoning tokens](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens).