@arnilo/prism 0.0.5 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/CHANGELOG.md +28 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +26 -16
  4. package/dist/agents.js +2 -3
  5. package/dist/contracts.d.ts +2 -0
  6. package/dist/ids.d.ts +2 -0
  7. package/dist/ids.js +6 -0
  8. package/dist/index.d.ts +5 -1
  9. package/dist/index.js +3 -1
  10. package/dist/session-stores.js +2 -3
  11. package/dist/testing/persistence-schema.d.ts +45 -7
  12. package/dist/testing/persistence-schema.js +138 -24
  13. package/dist/thinking.d.ts +42 -0
  14. package/dist/thinking.js +92 -0
  15. package/dist/tools.js +2 -3
  16. package/dist/use-case-model.d.ts +63 -0
  17. package/dist/use-case-model.js +52 -0
  18. package/docs/a2a.md +4 -2
  19. package/docs/agent-events.md +10 -15
  20. package/docs/agent-loops.md +11 -8
  21. package/docs/coding-agent-tools.md +33 -12
  22. package/docs/coding-security.md +2 -2
  23. package/docs/compaction-llm.md +17 -7
  24. package/docs/compaction-observational-memory.md +28 -4
  25. package/docs/credential-storage.md +58 -9
  26. package/docs/credentials-and-redaction.md +1 -1
  27. package/docs/database-persistence.md +8 -3
  28. package/docs/host-security.md +10 -6
  29. package/docs/index.md +23 -20
  30. package/docs/mcp-tools.md +26 -10
  31. package/docs/migration.md +146 -2
  32. package/docs/node-filesystem-config.md +1 -0
  33. package/docs/node-jsonl-session-store.md +5 -4
  34. package/docs/postgres-persistence.md +3 -3
  35. package/docs/provider-caching.md +16 -4
  36. package/docs/provider-conformance.md +39 -1
  37. package/docs/provider-packages.md +60 -3
  38. package/docs/providers/ai-sdk.md +36 -0
  39. package/docs/providers/kimi.md +124 -61
  40. package/docs/providers/neuralwatt.md +19 -13
  41. package/docs/providers/openai.md +56 -13
  42. package/docs/providers/opencode-go.md +118 -30
  43. package/docs/providers/openrouter.md +105 -35
  44. package/docs/providers/zai.md +94 -45
  45. package/docs/release-and-install.md +47 -49
  46. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  47. package/docs/runs-and-usage.md +1 -1
  48. package/docs/sqlite-persistence.md +2 -2
  49. package/docs/structured-output.md +1 -1
  50. package/docs/thinking-and-reasoning.md +98 -0
  51. package/docs/tool-execution-primitives.md +3 -3
  52. package/docs/tools.md +15 -0
  53. package/docs/use-case-model-selection.md +109 -0
  54. package/docs/workflow-orchestration-primitives.md +1 -0
  55. package/docs/workflows.md +17 -10
  56. package/docs/working-and-semantic-memory.md +1 -0
  57. package/package.json +2 -2
@@ -2,132 +2,195 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-provider-kimi` provides explicit, side-effect-free setup for Kimi For
6
- Coding using an Anthropic-compatible `/messages` endpoint with
7
- `User-Agent: KimiCLI/1.5` (unless overridden). Moonshot/Open Platform model
8
- metadata is optional.
5
+ `@arnilo/prism-provider-kimi` provides two distinct, side-effect-free routes:
9
6
 
10
- The package registers the `kimi-coding` provider, default Kimi Coding model
11
- metadata, and an `api_key` auth method through `createExtensionKernel().load([...])`.
7
+ 1. **Kimi For Coding** (default) — Anthropic-compatible `POST /messages` on
8
+ `https://api.kimi.com/coding` with `User-Agent: KimiCLI/1.5` (unless overridden).
9
+ 2. **Moonshot Open Platform** (opt-in) — OpenAI-compatible `POST /chat/completions`
10
+ on `https://api.moonshot.ai/v1` (or `api.moonshot.cn/v1`), registered only when
11
+ `includeMoonshotModels: true`.
12
+
13
+ Official model ids differ by route. Coding uses `kimi-for-coding`,
14
+ `kimi-for-coding-highspeed`, and `k3`. Open Platform uses `kimi-k2.7-code`,
15
+ `kimi-k3`, and related catalog ids. Pi's `k2p7` alias is **not** used.
16
+
17
+ Caller-gated discovery via `listKimiModels()` hits the official Moonshot
18
+ `GET /v1/models` endpoint. Package setup never fetches.
12
19
 
13
20
  ## When to use it
14
21
 
15
- Use it when a host app wants the Kimi For Coding endpoint through Prism's
16
- `AgentSession` runtime with Kimi-specific serializer behavior.
22
+ Use it when a host app wants Kimi For Coding and/or Moonshot Open Platform through
23
+ Prism's `AgentSession` runtime with Kimi-specific serializers, thinking controls,
24
+ and cache policy.
17
25
 
18
- Do not use it for Moonshot Open Platform default registration, automatic
19
- credential discovery, catalog fetches, or real-network tests.
26
+ Do not use it for automatic credential discovery, setup-time catalog fetches, or
27
+ real-network tests (live tests stay opt-in).
20
28
 
21
29
  ## Inputs / request
22
30
 
23
31
  ```ts
24
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
32
+ import {
33
+ createKimiProviderPackage,
34
+ listKimiModels,
35
+ defineKimiModel,
36
+ } from "@arnilo/prism-provider-kimi";
25
37
 
26
38
  createKimiProviderPackage(options: KimiProviderPackageOptions): ProviderPackage
27
- defineKimiModel(config: KimiModelConfig): KimiModelConfig
39
+ listKimiModels(options?: ListKimiModelsOptions): Promise<ModelConfig[]>
40
+ defineKimiModel(config: KimiModelConfig): ModelConfig
28
41
  ```
29
42
 
30
43
  | Field | Type | Purpose |
31
44
  | --- | --- | --- |
32
- | `kimiApiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source for Kimi. |
45
+ | `kimiApiKey` | `CredentialValueSource` | Kimi For Coding API key. |
46
+ | `moonshotApiKey` | `CredentialValueSource` | Moonshot Open Platform API key (not interchangeable with Coding keys). |
33
47
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
34
- | `baseUrl` | `string` | Overrides the Kimi base URL. |
35
- | `id` | `string` | Overrides the provider id (default `kimi-coding`). |
36
- | `userAgent` | `string` | Overrides `User-Agent: KimiCLI/1.5`. |
37
- | `models` | `readonly ModelConfig[]` | Overrides `kimiCodingModels` defaults. |
38
- | `includeMoonshotModels` | `boolean` | Registers Moonshot models when `true` (default off). |
39
- | `moonshotModels` | `readonly ModelConfig[]` | Overrides `moonshotKimiModels` when included. |
48
+ | `baseUrl` | `string` | Overrides the Coding base URL. |
49
+ | `moonshotBaseUrl` | `string` | Overrides Moonshot base URL (default `https://api.moonshot.ai/v1`). |
50
+ | `id` / `moonshotId` | `string` | Provider ids (defaults `kimi-coding` / `moonshot`). |
51
+ | `userAgent` | `string` | Overrides Coding `User-Agent: KimiCLI/1.5`. |
52
+ | `models` | `readonly ModelConfig[]` | Overrides featured Coding models. |
53
+ | `includeMoonshotModels` | `boolean` | Registers callable Moonshot provider + models when `true`. |
54
+ | `moonshotModels` | `readonly ModelConfig[]` | Overrides featured Moonshot models when included. |
40
55
 
41
56
  ## Outputs / response / events
42
57
 
43
58
  | Surface | Behavior |
44
59
  | --- | --- |
45
- | Provider stream | Prism text, thinking (preserved only when `model.compat.preserveThinking` is true, otherwise downgraded to text), tool-call delta/final, `usage`, `done`, redacted `error`. |
46
- | Block preservation | Text, thinking, assistant `tool_call` → `tool_use`, `tool_result` → `tool_result`, images when `capabilities.input` includes `"image"`. |
47
- | Auth method | `api_key` for `kimi-coding`, credential name `apiKey`. |
60
+ | Coding stream | Prism text, thinking deltas, tool-call delta/final, `usage` (`cache_read_input_tokens` / `cache_creation_input_tokens`), `done`, redacted `error`. |
61
+ | Moonshot stream | Same, with `delta.reasoning_content` → thinking; OpenAI-style usage cache details when present. |
62
+ | Block preservation | Coding: Anthropic `thinking` / `tool_use` / `tool_result`. Moonshot: `reasoning_content` on assistant replay when `preserveThinking`. |
63
+ | Auth methods | `api_key` for `kimi-coding`; also `moonshot` when opted in. |
48
64
 
49
65
  Unsupported block placements or unclaimed images fail before fetch.
50
66
 
67
+ ## Route differences
68
+
69
+ | | Kimi For Coding | Moonshot Open Platform |
70
+ | --- | --- | --- |
71
+ | Base URL | `https://api.kimi.com/coding` | `https://api.moonshot.ai/v1` (or `.cn`) |
72
+ | Wire API | Anthropic `/messages` | OpenAI `/chat/completions` |
73
+ | Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k3` (+ discovery) |
74
+ | Discovery | No public list API — curated featured aliases | Official `GET /v1/models` via `listKimiModels()` |
75
+ | Cache | Implicit by default; opt-in Anthropic `cache_control` | Implicit only — never emits Anthropic `cache_control` |
76
+ | Thinking | Block replay + body `thinking` / `reasoning_effort` | `reasoning_content` replay + body `thinking` / `reasoning_effort` |
77
+
78
+ The Anthropic `/messages` request/response contract for Kimi remains under-documented
79
+ upstream ([MoonshotAI/Kimi-K2#129](https://github.com/MoonshotAI/Kimi-K2/issues/129));
80
+ Prism treats official Chat Completions thinking fields as best-effort passthrough on
81
+ the Coding route.
82
+
83
+ ## Thinking / reasoning
84
+
85
+ Official fields (Open Platform docs; Coding docs for `k3` effort mapping):
86
+
87
+ | Model family | Official control | Prism `compat` |
88
+ | --- | --- | --- |
89
+ | K3 / Coding `k3` | top-level `reasoning_effort` (`max` on Open Platform; Coding also `low`/`high`) | `compat.reasoning_effort` — use Task 4 family `reasoning_effort` |
90
+ | K2.7-code / Coding | thinking always on; Preserved Thinking always on | omit `thinking` by default; `preserveThinking: true` for replay; do not send `disabled` |
91
+ | K2.6 / K2.5 | `thinking.type` enabled/disabled; K2.6 optional `keep: "all"` | `compat.thinking` — Task 4 family `thinking_type` |
92
+
93
+ Per-turn `ProviderRequestOptions.compat` wins over `ModelConfig.compat`. Helpers:
94
+ `kimiThinking`, `kimiReasoningEffort`, `kimiPreserveThinking`.
95
+
51
96
  ## Request/response example
52
97
 
53
- Example request (Anthropic-compatible `/messages` shape):
98
+ Coding (Anthropic-compatible `/messages`):
54
99
 
55
100
  ```json
56
101
  {
57
- "model": "kimi-latest",
58
- "messages": [{ "role": "user", "content": "Hello" }],
102
+ "model": "kimi-for-coding",
103
+ "messages": [{ "role": "user", "content": [{ "type": "text", "text": "Hello" }] }],
59
104
  "stream": true
60
105
  }
61
106
  ```
62
107
 
108
+ Moonshot (Chat Completions):
109
+
110
+ ```json
111
+ {
112
+ "model": "kimi-k3",
113
+ "messages": [{ "role": "user", "content": "Hello" }],
114
+ "stream": true,
115
+ "reasoning_effort": "max"
116
+ }
117
+ ```
118
+
63
119
  ## Implementation example
64
120
 
65
121
  ```ts
66
122
  import { createExtensionKernel } from "@arnilo/prism";
67
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
123
+ import {
124
+ createKimiProviderPackage,
125
+ listKimiModels,
126
+ } from "@arnilo/prism-provider-kimi";
68
127
 
69
128
  const kernel = createExtensionKernel();
70
129
  await kernel.load([
71
- createKimiProviderPackage({ kimiApiKey: "fake-kimi-key", includeMoonshotModels: false }),
130
+ createKimiProviderPackage({ kimiApiKey: "fake-kimi-key" }),
72
131
  ]);
73
- ```
74
-
75
- Register Moonshot/Open Platform metadata explicitly:
76
132
 
77
- ```ts
78
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
133
+ // Opt-in Moonshot Open Platform (callable provider + featured models)
134
+ await kernel.load([
135
+ createKimiProviderPackage({
136
+ kimiApiKey: "fake-kimi-key",
137
+ includeMoonshotModels: true,
138
+ moonshotApiKey: "fake-moonshot-key",
139
+ }),
140
+ ]);
79
141
 
142
+ // Caller-gated discovery — never runs during setup
143
+ const latest = await listKimiModels({ apiKey: "fake-moonshot-key", fetch });
80
144
  await kernel.load([
81
- createKimiProviderPackage({ kimiApiKey: "fake", includeMoonshotModels: true }),
145
+ createKimiProviderPackage({
146
+ includeMoonshotModels: true,
147
+ moonshotApiKey: "fake-moonshot-key",
148
+ moonshotModels: latest.filter((m) => m.model.startsWith("kimi-")),
149
+ }),
82
150
  ]);
83
151
  ```
84
152
 
85
153
  ## Extension and configuration notes
86
154
 
87
- - Hosts choose base URL, provider id, `User-Agent`, model list, credential source,
155
+ - Hosts choose base URLs, provider ids, `User-Agent`, model lists, credential sources,
88
156
  and `fetch` impl.
89
- - Moonshot/Open Platform metadata is registered only with
90
- `includeMoonshotModels: true`; it is not core behavior.
91
- - Package contributes models via the extension `api` and an `api_key` auth method.
157
+ - Moonshot is registered only with `includeMoonshotModels: true` (provider + models + auth).
158
+ - Featured catalogs are offline bootstrap only; refresh Open Platform via `listKimiModels()`.
92
159
 
93
160
  ### Cache behavior
94
161
 
95
- - Default catalog models (e.g. `kimi-k2.7-code` on the Anthropic-compatible
96
- `/messages` route) use **implicit caching** and send no explicit `cache_control`
97
- fields. `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention` have no
98
- effect on the request body unless the model opts in.
99
- - Hosts may opt a model into Anthropic-style `cache_control` by declaring
100
- `ModelConfig.cache.kind: "cache_control"` on the Anthropic route. When opted in,
101
- `cache_control: { type: "ephemeral" }` markers are applied only to the
102
- caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
103
- shared `applyCacheControl()` helper) on the last content block of each selected
104
- message — not to every block. `cacheRetention: "long"` adds `ttl: "1h"` when the
105
- model allows long retention (`ModelConfig.cache.longRetention !== false`).
106
- - The Moonshot Open Platform route (`compat.route: "openai"`) never receives
107
- Anthropic `cache_control` fields.
108
- - Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
109
- `Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
110
- `Usage.cacheWriteTokens`.
162
+ - Default Coding catalog models use **implicit caching** and send no explicit
163
+ `cache_control` fields unless the model opts in with
164
+ `ModelConfig.cache.kind: "cache_control"`.
165
+ - When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
166
+ caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
167
+ block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
168
+ the model allows long retention.
169
+ - The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
170
+ - Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
171
+ `cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
111
172
 
112
173
  ## Security and performance notes
113
174
 
114
- - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
175
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
176
+ helpers (`readSseData`, `readBoundedResponseText`).
115
177
  - No network calls during import, setup, build, or default tests.
116
178
  - No automatic environment, file, keychain, or shell credential lookup.
117
- - Kimi credentials are resolved per request from caller-supplied values or resolvers
118
- and redacted from errors.
179
+ - Credentials are resolved per request from caller-supplied values or resolvers
180
+ and redacted from errors (including discovery failures).
119
181
  - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
120
182
  but provider-owned headers (`content-type`, `user-agent`, `authorization`)
121
183
  are applied last and cannot be overridden by caller headers.
122
- - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
123
- provider-specific env names; default tests are network-free.
184
+ - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus provider-specific
185
+ env names; default tests are network-free.
124
186
 
125
187
  ## Related APIs
126
188
 
127
189
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
128
- `ModelConfig`/`compat`, Anthropic-compatible routes.
190
+ caller-gated discovery, Anthropic/OpenAI routes.
191
+ - [Thinking and reasoning](../thinking-and-reasoning.md): portable `ThinkingLevel`
192
+ helpers and Kimi family mapping.
129
193
  - [Credentials and redaction](../credentials-and-redaction.md):
130
194
  `resolveCredentialValue`, `redactSecrets`.
131
- - [Provider layer](../provider-layer.md): `ProviderRequest.options` and usage
132
- mapping.
195
+ - [Provider caching](../provider-caching.md): explicit/implicit matrix.
133
196
  - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -9,7 +9,7 @@ and implicit prefix caching.
9
9
 
10
10
  The package registers a provider, default model metadata for the featured NeuralWatt
11
11
  aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
12
- `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
12
+ `gemma-4-31b`, `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
13
13
  `qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
14
14
  method through `createExtensionKernel().load([...])`.
15
15
 
@@ -79,13 +79,15 @@ The endpoint is rate-limited to **1 request per second per customer** (429 with
79
79
  from `generate()` or package setup; the caller owns throttling.
80
80
 
81
81
  NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
82
- / `extra` escape hatches: `compat.reasoning_effort` (`"low" | "medium" | "high"`),
83
- `compat.thinking_token_budget`, `compat.chat_template_kwargs` (including `enable_thinking`),
84
- `compat.preserve_thinking`, `compat.clear_thinking`, and `compat.tool_choice`.
85
- `preserve_thinking: true` keeps prior assistant reasoning in request history so
86
- multi-turn reasoning continues with the earlier chain of thought; `clear_thinking:
87
- true` drops it for the next turn, resetting the chain. `options.extra` spreads after
88
- `compat` so per-call values and overrides win.
82
+ / `extra` escape hatches: `compat.reasoning_effort` (OpenAI-style scale; GLM-5.2 defaults
83
+ unset to `max` per official docs), `compat.thinking_token_budget`,
84
+ `compat.chat_template_kwargs` (including `enable_thinking`, `preserve_thinking`, and
85
+ `clear_thinking` per official gateway docs), and `compat.tool_choice`.
86
+ `preserve_thinking` / `clear_thinking` compat flags are routed into `chat_template_kwargs`
87
+ on the wire (not top-level body fields). Prism also uses these flags for client-side
88
+ message serialization: `preserve_thinking: true` keeps prior assistant reasoning in
89
+ request history; `clear_thinking: true` drops it for the next turn. `options.extra` spreads
90
+ after resolved fields so per-call values and overrides win.
89
91
 
90
92
  ## Outputs / response / events
91
93
 
@@ -112,7 +114,7 @@ Example request body (OpenAI-compatible Chat Completions shape):
112
114
  "messages": [{ "role": "user", "content": "Hello" }],
113
115
  "stream": true,
114
116
  "stream_options": { "include_usage": true },
115
- "reasoning_effort": "medium",
117
+ "reasoning_effort": "max",
116
118
  "thinking_token_budget": 8192
117
119
  }
118
120
  ```
@@ -192,6 +194,7 @@ validation. The caller owns throttling/caching — the helper makes one explicit
192
194
  | `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
193
195
  | `glm-5.2-short` | 195K | Tools, reasoning |
194
196
  | `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
197
+ | `gemma-4-31b` | 256K | Tools, vision, JSON mode |
195
198
  | `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
196
199
  | `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
197
200
  | `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
@@ -255,15 +258,18 @@ so multi-turn sessions continue the earlier chain of thought:
255
258
 
256
259
  - Prior `thinking` content blocks on an assistant message are serialized under a
257
260
  `reasoning_content` field on that message (matching the streaming
258
- `delta.reasoning_content` field). They are **not** flattened into text `content`, so
261
+ `delta.reasoning_content` field; the gateway also accepts `reasoning` as an alias).
262
+ They are **not** flattened into text `content`, so
259
263
  the model sees reasoning and answer as distinct.
260
264
  - Preservation is gated on `model.capabilities.reasoning === true` **or**
261
265
  `compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
262
266
  field and prior `thinking` blocks are dropped — they never leak into text content for
263
267
  providers/models that do not support reasoning.
264
- - `compat.clear_thinking: true` drops prior reasoning for the next turn even on
265
- reasoning-capable models, resetting the chain of thought. `clear_thinking` takes
266
- precedence over `preserve_thinking`.
268
+ - `compat.preserve_thinking` / `compat.clear_thinking` map into `chat_template_kwargs`
269
+ on the request body per official NeuralWatt docs (Kimi K2.6 `preserve_thinking`, GLM
270
+ `clear_thinking: false` for full-history). Prism also uses `clear_thinking: true` to
271
+ drop prior reasoning client-side even on reasoning-capable models; `clear_thinking`
272
+ takes precedence over `preserve_thinking`.
267
273
  - The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
268
274
  reasoning.
269
275
 
@@ -36,16 +36,22 @@ createOpenAIProviderPackage(options: OpenAIProviderPackageOptions): ProviderPack
36
36
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
37
37
  | `baseUrl` | `string` | Overrides `https://api.openai.com/v1`. |
38
38
  | `codexBaseUrl` | `string` | Overrides `https://chatgpt.com/backend-api/codex`. |
39
+ | `models` | `readonly ModelConfig[]` | Optional override for registered OpenAI Responses models (defaults to featured `openAIModels`). |
40
+ | `codexModels` | `readonly ModelConfig[]` | Optional override for registered Codex models (defaults to featured `openAICodexModels`). |
39
41
 
40
42
  `ProviderRequest.options.sessionId`, `cacheKey`, `cacheRetention`, `headers`,
41
- `compat`, and `extra` map to request headers/payload fields.
43
+ `compat`, and `extra` map to request headers/payload fields. Per-turn reasoning
44
+ uses official Responses `reasoning: { effort, summary? }` via
45
+ `ModelConfig.compat.reasoning` defaults merged with
46
+ `ProviderRequestOptions.compat.reasoning` (request wins). Prefer
47
+ `applyThinkingLevel(..., "openai_reasoning")` from `@arnilo/prism`.
42
48
 
43
49
  ## Outputs / response / events
44
50
 
45
51
  | Surface | Behavior |
46
52
  | --- | --- |
47
53
  | Provider stream | Prism text, thinking (downgraded to text), `tool_call` deltas/finals, `usage`, `done`, redacted `error` events. |
48
- | Block preservation | Text, thinking (downgraded), assistant `tool_call` → `function_call` input items, `tool_result` → `function_call_output` input items, images when `capabilities.input` includes `"image"`. |
54
+ | Block preservation | User/system text → `input_text`; assistant text → `output_text`; assistant `tool_call` → top-level `function_call` with `call_id`; `tool_result` → top-level `function_call_output`; images/files/audio when declared on the model. Bare thinking without an encrypted Responses reasoning item is omitted on replay. |
49
55
  | Auth methods | `api_key` for `openai`; `oauth` for `openai-codex`. |
50
56
 
51
57
  Unsupported block placements or unclaimed images fail before `fetch`.
@@ -56,10 +62,17 @@ Responses request body (Codex subscription shape, abbreviated):
56
62
 
57
63
  ```json
58
64
  {
59
- "model": "gpt-5-codex",
60
- "instructions": "You are a coding agent.",
61
- "input": [{ "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] }],
62
- "stream": true
65
+ "model": "gpt-5.1",
66
+ "input": [
67
+ { "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] },
68
+ { "role": "assistant", "content": [{ "type": "output_text", "text": "Calling lookup" }] },
69
+ { "type": "function_call", "call_id": "call_1", "name": "lookup", "arguments": "{\"q\":\"x\"}" },
70
+ { "type": "function_call_output", "call_id": "call_1", "output": "{\"ok\":true}" }
71
+ ],
72
+ "reasoning": { "effort": "high" },
73
+ "prompt_cache_key": "session-1",
74
+ "stream": true,
75
+ "store": false
63
76
  }
64
77
  ```
65
78
 
@@ -73,13 +86,13 @@ https://auth.openai.com/authorize?response_type=code&client_id=...&code_challeng
73
86
 
74
87
  ```ts
75
88
  import { createExtensionKernel, createEnvCredentialResolver } from "@arnilo/prism";
76
- import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
89
+ import { createOpenAIProviderPackage, listOpenAIModels } from "@arnilo/prism-provider-openai";
77
90
 
91
+ const apiKey = createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" });
92
+ const models = await listOpenAIModels({ apiKey }); // caller-gated; never runs during setup
78
93
  const kernel = createExtensionKernel();
79
94
  await kernel.load([
80
- createOpenAIProviderPackage({
81
- apiKey: createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" }),
82
- }),
95
+ createOpenAIProviderPackage({ apiKey, models }),
83
96
  ]);
84
97
  ```
85
98
 
@@ -115,25 +128,55 @@ const challenge = computeS256Challenge(verifier);
115
128
 
116
129
  ### Cache behavior
117
130
 
131
+ Official: [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching).
132
+
118
133
  - `prompt_cache_key` is derived from `ProviderRequestOptions.cacheKey` (falling
119
134
  back to `sessionId`) and sanitized + clamped to 64 characters via the shared
120
135
  `sanitizeCacheKey()` helper. Cache keys are session/customer identifiers only;
121
136
  never credentials or raw prompts.
122
- - `prompt_cache_retention` accepts only `"24h"` on the OpenAI Responses API
137
+ - `prompt_cache_retention` accepts only `"24h"` on pre-GPT-5.6 Responses models
123
138
  (extended caching). Prism `cacheRetention: "short"` and `"none"` omit the field
124
139
  so default automatic/implicit caching applies and no invalid literal is sent.
125
140
  `cacheRetention: "long"` maps to `prompt_cache_retention: "24h"` only when the
126
141
  model declares `ModelConfig.cache.longRetention === true`; models without that
127
- metadata omit the field. The catalog `gpt-5.1` model declares
142
+ metadata omit the field. Featured `gpt-5.1` declares
128
143
  `cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
144
+ - GPT-5.6+ official docs prefer `prompt_cache_options` / explicit breakpoints;
145
+ `listOpenAIModels` sets `longRetention: false` for those ids so Prism does not
146
+ emit deprecated `prompt_cache_retention` for them. Breakpoint helpers are not
147
+ shipped in this package yet — hosts may pass `prompt_cache_options` through
148
+ `compat` / `extra` when needed.
129
149
  - Cache accounting is preserved in normalized `Usage`: OpenAI
130
150
  `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
131
- Responses does not report a cache-write token field.
151
+ Responses does not report a cache-write token field on older models.
132
152
  - Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
133
153
  are applied after caller `ProviderRequestOptions.headers` so caller config
134
154
  cannot replace credentials, content type, or the session request id; non-owned
135
155
  caller headers are kept.
136
156
 
157
+ ### Model discovery
158
+
159
+ - `listOpenAIModels({ apiKey, fetch, baseUrl, signal, headers })` calls official
160
+ [`GET /models`](https://developers.openai.com/api/reference/resources/models/methods/list)
161
+ and maps sparse `{ id, created, owned_by }` entries to `ModelConfig` with
162
+ `cache.kind: "openai_key"` and heuristic `longRetention` / `capabilities.reasoning`.
163
+ - `createOpenAIProviderPackage` never calls discovery; pass results via `models:`.
164
+ - Codex subscription models are **not** listed by `api.openai.com` — keep using
165
+ featured `openAICodexModels` or `codexModels:` override.
166
+ - Static `openAIModels` / `openAICodexModels` are offline bootstrap / featured aliases only.
167
+
168
+ ### Reasoning
169
+
170
+ Official: [Reasoning models](https://developers.openai.com/api/docs/guides/reasoning).
171
+
172
+ - Body field is top-level `reasoning: { effort, summary?, mode?, context? }`.
173
+ - Model defaults: `ModelConfig.compat.reasoning`; per-turn override:
174
+ `ProviderRequestOptions.compat.reasoning` (shallow-merged; request wins).
175
+ - Portable helper: `applyThinkingLevel(options, level, "openai_reasoning")`.
176
+ - Streaming tool args follow official
177
+ `response.output_item.added` + `response.function_call_arguments.delta`
178
+ (string `delta`), not Chat Completions object deltas.
179
+
137
180
  ## Security and performance notes
138
181
 
139
182
  - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).