@arnilo/prism 0.0.5 → 0.0.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +28 -1
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +26 -16
- package/dist/agents.js +2 -3
- package/dist/contracts.d.ts +2 -0
- package/dist/ids.d.ts +2 -0
- package/dist/ids.js +6 -0
- package/dist/index.d.ts +5 -1
- package/dist/index.js +3 -1
- package/dist/session-stores.js +2 -3
- package/dist/testing/persistence-schema.d.ts +45 -7
- package/dist/testing/persistence-schema.js +138 -24
- package/dist/thinking.d.ts +42 -0
- package/dist/thinking.js +92 -0
- package/dist/tools.js +2 -3
- package/dist/use-case-model.d.ts +63 -0
- package/dist/use-case-model.js +52 -0
- package/docs/a2a.md +4 -2
- package/docs/agent-events.md +10 -15
- package/docs/agent-loops.md +11 -8
- package/docs/coding-agent-tools.md +33 -12
- package/docs/coding-security.md +2 -2
- package/docs/compaction-llm.md +17 -7
- package/docs/compaction-observational-memory.md +28 -4
- package/docs/credential-storage.md +58 -9
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/database-persistence.md +8 -3
- package/docs/host-security.md +10 -6
- package/docs/index.md +23 -20
- package/docs/mcp-tools.md +26 -10
- package/docs/migration.md +146 -2
- package/docs/node-filesystem-config.md +1 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/postgres-persistence.md +3 -3
- package/docs/provider-caching.md +16 -4
- package/docs/provider-conformance.md +39 -1
- package/docs/provider-packages.md +60 -3
- package/docs/providers/ai-sdk.md +36 -0
- package/docs/providers/kimi.md +124 -61
- package/docs/providers/neuralwatt.md +19 -13
- package/docs/providers/openai.md +56 -13
- package/docs/providers/opencode-go.md +118 -30
- package/docs/providers/openrouter.md +105 -35
- package/docs/providers/zai.md +94 -45
- package/docs/release-and-install.md +47 -49
- package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
- package/docs/runs-and-usage.md +1 -1
- package/docs/sqlite-persistence.md +2 -2
- package/docs/structured-output.md +1 -1
- package/docs/thinking-and-reasoning.md +98 -0
- package/docs/tool-execution-primitives.md +3 -3
- package/docs/tools.md +15 -0
- package/docs/use-case-model-selection.md +109 -0
- package/docs/workflow-orchestration-primitives.md +1 -0
- package/docs/workflows.md +17 -10
- package/docs/working-and-semantic-memory.md +1 -0
- package/package.json +2 -2
package/docs/providers/kimi.md
CHANGED
|
@@ -2,132 +2,195 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-provider-kimi` provides
|
|
6
|
-
Coding using an Anthropic-compatible `/messages` endpoint with
|
|
7
|
-
`User-Agent: KimiCLI/1.5` (unless overridden). Moonshot/Open Platform model
|
|
8
|
-
metadata is optional.
|
|
5
|
+
`@arnilo/prism-provider-kimi` provides two distinct, side-effect-free routes:
|
|
9
6
|
|
|
10
|
-
|
|
11
|
-
|
|
7
|
+
1. **Kimi For Coding** (default) — Anthropic-compatible `POST /messages` on
|
|
8
|
+
`https://api.kimi.com/coding` with `User-Agent: KimiCLI/1.5` (unless overridden).
|
|
9
|
+
2. **Moonshot Open Platform** (opt-in) — OpenAI-compatible `POST /chat/completions`
|
|
10
|
+
on `https://api.moonshot.ai/v1` (or `api.moonshot.cn/v1`), registered only when
|
|
11
|
+
`includeMoonshotModels: true`.
|
|
12
|
+
|
|
13
|
+
Official model ids differ by route. Coding uses `kimi-for-coding`,
|
|
14
|
+
`kimi-for-coding-highspeed`, and `k3`. Open Platform uses `kimi-k2.7-code`,
|
|
15
|
+
`kimi-k3`, and related catalog ids. Pi's `k2p7` alias is **not** used.
|
|
16
|
+
|
|
17
|
+
Caller-gated discovery via `listKimiModels()` hits the official Moonshot
|
|
18
|
+
`GET /v1/models` endpoint. Package setup never fetches.
|
|
12
19
|
|
|
13
20
|
## When to use it
|
|
14
21
|
|
|
15
|
-
Use it when a host app wants
|
|
16
|
-
`AgentSession` runtime with Kimi-specific
|
|
22
|
+
Use it when a host app wants Kimi For Coding and/or Moonshot Open Platform through
|
|
23
|
+
Prism's `AgentSession` runtime with Kimi-specific serializers, thinking controls,
|
|
24
|
+
and cache policy.
|
|
17
25
|
|
|
18
|
-
Do not use it for
|
|
19
|
-
|
|
26
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches, or
|
|
27
|
+
real-network tests (live tests stay opt-in).
|
|
20
28
|
|
|
21
29
|
## Inputs / request
|
|
22
30
|
|
|
23
31
|
```ts
|
|
24
|
-
import {
|
|
32
|
+
import {
|
|
33
|
+
createKimiProviderPackage,
|
|
34
|
+
listKimiModels,
|
|
35
|
+
defineKimiModel,
|
|
36
|
+
} from "@arnilo/prism-provider-kimi";
|
|
25
37
|
|
|
26
38
|
createKimiProviderPackage(options: KimiProviderPackageOptions): ProviderPackage
|
|
27
|
-
|
|
39
|
+
listKimiModels(options?: ListKimiModelsOptions): Promise<ModelConfig[]>
|
|
40
|
+
defineKimiModel(config: KimiModelConfig): ModelConfig
|
|
28
41
|
```
|
|
29
42
|
|
|
30
43
|
| Field | Type | Purpose |
|
|
31
44
|
| --- | --- | --- |
|
|
32
|
-
| `kimiApiKey` | `CredentialValueSource` |
|
|
45
|
+
| `kimiApiKey` | `CredentialValueSource` | Kimi For Coding API key. |
|
|
46
|
+
| `moonshotApiKey` | `CredentialValueSource` | Moonshot Open Platform API key (not interchangeable with Coding keys). |
|
|
33
47
|
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
34
|
-
| `baseUrl` | `string` | Overrides the
|
|
35
|
-
| `
|
|
36
|
-
| `
|
|
37
|
-
| `
|
|
38
|
-
| `
|
|
39
|
-
| `
|
|
48
|
+
| `baseUrl` | `string` | Overrides the Coding base URL. |
|
|
49
|
+
| `moonshotBaseUrl` | `string` | Overrides Moonshot base URL (default `https://api.moonshot.ai/v1`). |
|
|
50
|
+
| `id` / `moonshotId` | `string` | Provider ids (defaults `kimi-coding` / `moonshot`). |
|
|
51
|
+
| `userAgent` | `string` | Overrides Coding `User-Agent: KimiCLI/1.5`. |
|
|
52
|
+
| `models` | `readonly ModelConfig[]` | Overrides featured Coding models. |
|
|
53
|
+
| `includeMoonshotModels` | `boolean` | Registers callable Moonshot provider + models when `true`. |
|
|
54
|
+
| `moonshotModels` | `readonly ModelConfig[]` | Overrides featured Moonshot models when included. |
|
|
40
55
|
|
|
41
56
|
## Outputs / response / events
|
|
42
57
|
|
|
43
58
|
| Surface | Behavior |
|
|
44
59
|
| --- | --- |
|
|
45
|
-
|
|
|
46
|
-
|
|
|
47
|
-
|
|
|
60
|
+
| Coding stream | Prism text, thinking deltas, tool-call delta/final, `usage` (`cache_read_input_tokens` / `cache_creation_input_tokens`), `done`, redacted `error`. |
|
|
61
|
+
| Moonshot stream | Same, with `delta.reasoning_content` → thinking; OpenAI-style usage cache details when present. |
|
|
62
|
+
| Block preservation | Coding: Anthropic `thinking` / `tool_use` / `tool_result`. Moonshot: `reasoning_content` on assistant replay when `preserveThinking`. |
|
|
63
|
+
| Auth methods | `api_key` for `kimi-coding`; also `moonshot` when opted in. |
|
|
48
64
|
|
|
49
65
|
Unsupported block placements or unclaimed images fail before fetch.
|
|
50
66
|
|
|
67
|
+
## Route differences
|
|
68
|
+
|
|
69
|
+
| | Kimi For Coding | Moonshot Open Platform |
|
|
70
|
+
| --- | --- | --- |
|
|
71
|
+
| Base URL | `https://api.kimi.com/coding` | `https://api.moonshot.ai/v1` (or `.cn`) |
|
|
72
|
+
| Wire API | Anthropic `/messages` | OpenAI `/chat/completions` |
|
|
73
|
+
| Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k3` (+ discovery) |
|
|
74
|
+
| Discovery | No public list API — curated featured aliases | Official `GET /v1/models` via `listKimiModels()` |
|
|
75
|
+
| Cache | Implicit by default; opt-in Anthropic `cache_control` | Implicit only — never emits Anthropic `cache_control` |
|
|
76
|
+
| Thinking | Block replay + body `thinking` / `reasoning_effort` | `reasoning_content` replay + body `thinking` / `reasoning_effort` |
|
|
77
|
+
|
|
78
|
+
The Anthropic `/messages` request/response contract for Kimi remains under-documented
|
|
79
|
+
upstream ([MoonshotAI/Kimi-K2#129](https://github.com/MoonshotAI/Kimi-K2/issues/129));
|
|
80
|
+
Prism treats official Chat Completions thinking fields as best-effort passthrough on
|
|
81
|
+
the Coding route.
|
|
82
|
+
|
|
83
|
+
## Thinking / reasoning
|
|
84
|
+
|
|
85
|
+
Official fields (Open Platform docs; Coding docs for `k3` effort mapping):
|
|
86
|
+
|
|
87
|
+
| Model family | Official control | Prism `compat` |
|
|
88
|
+
| --- | --- | --- |
|
|
89
|
+
| K3 / Coding `k3` | top-level `reasoning_effort` (`max` on Open Platform; Coding also `low`/`high`) | `compat.reasoning_effort` — use Task 4 family `reasoning_effort` |
|
|
90
|
+
| K2.7-code / Coding | thinking always on; Preserved Thinking always on | omit `thinking` by default; `preserveThinking: true` for replay; do not send `disabled` |
|
|
91
|
+
| K2.6 / K2.5 | `thinking.type` enabled/disabled; K2.6 optional `keep: "all"` | `compat.thinking` — Task 4 family `thinking_type` |
|
|
92
|
+
|
|
93
|
+
Per-turn `ProviderRequestOptions.compat` wins over `ModelConfig.compat`. Helpers:
|
|
94
|
+
`kimiThinking`, `kimiReasoningEffort`, `kimiPreserveThinking`.
|
|
95
|
+
|
|
51
96
|
## Request/response example
|
|
52
97
|
|
|
53
|
-
|
|
98
|
+
Coding (Anthropic-compatible `/messages`):
|
|
54
99
|
|
|
55
100
|
```json
|
|
56
101
|
{
|
|
57
|
-
"model": "kimi-
|
|
58
|
-
"messages": [{ "role": "user", "content": "Hello" }],
|
|
102
|
+
"model": "kimi-for-coding",
|
|
103
|
+
"messages": [{ "role": "user", "content": [{ "type": "text", "text": "Hello" }] }],
|
|
59
104
|
"stream": true
|
|
60
105
|
}
|
|
61
106
|
```
|
|
62
107
|
|
|
108
|
+
Moonshot (Chat Completions):
|
|
109
|
+
|
|
110
|
+
```json
|
|
111
|
+
{
|
|
112
|
+
"model": "kimi-k3",
|
|
113
|
+
"messages": [{ "role": "user", "content": "Hello" }],
|
|
114
|
+
"stream": true,
|
|
115
|
+
"reasoning_effort": "max"
|
|
116
|
+
}
|
|
117
|
+
```
|
|
118
|
+
|
|
63
119
|
## Implementation example
|
|
64
120
|
|
|
65
121
|
```ts
|
|
66
122
|
import { createExtensionKernel } from "@arnilo/prism";
|
|
67
|
-
import {
|
|
123
|
+
import {
|
|
124
|
+
createKimiProviderPackage,
|
|
125
|
+
listKimiModels,
|
|
126
|
+
} from "@arnilo/prism-provider-kimi";
|
|
68
127
|
|
|
69
128
|
const kernel = createExtensionKernel();
|
|
70
129
|
await kernel.load([
|
|
71
|
-
createKimiProviderPackage({ kimiApiKey: "fake-kimi-key"
|
|
130
|
+
createKimiProviderPackage({ kimiApiKey: "fake-kimi-key" }),
|
|
72
131
|
]);
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
Register Moonshot/Open Platform metadata explicitly:
|
|
76
132
|
|
|
77
|
-
|
|
78
|
-
|
|
133
|
+
// Opt-in Moonshot Open Platform (callable provider + featured models)
|
|
134
|
+
await kernel.load([
|
|
135
|
+
createKimiProviderPackage({
|
|
136
|
+
kimiApiKey: "fake-kimi-key",
|
|
137
|
+
includeMoonshotModels: true,
|
|
138
|
+
moonshotApiKey: "fake-moonshot-key",
|
|
139
|
+
}),
|
|
140
|
+
]);
|
|
79
141
|
|
|
142
|
+
// Caller-gated discovery — never runs during setup
|
|
143
|
+
const latest = await listKimiModels({ apiKey: "fake-moonshot-key", fetch });
|
|
80
144
|
await kernel.load([
|
|
81
|
-
createKimiProviderPackage({
|
|
145
|
+
createKimiProviderPackage({
|
|
146
|
+
includeMoonshotModels: true,
|
|
147
|
+
moonshotApiKey: "fake-moonshot-key",
|
|
148
|
+
moonshotModels: latest.filter((m) => m.model.startsWith("kimi-")),
|
|
149
|
+
}),
|
|
82
150
|
]);
|
|
83
151
|
```
|
|
84
152
|
|
|
85
153
|
## Extension and configuration notes
|
|
86
154
|
|
|
87
|
-
- Hosts choose base
|
|
155
|
+
- Hosts choose base URLs, provider ids, `User-Agent`, model lists, credential sources,
|
|
88
156
|
and `fetch` impl.
|
|
89
|
-
- Moonshot
|
|
90
|
-
|
|
91
|
-
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
157
|
+
- Moonshot is registered only with `includeMoonshotModels: true` (provider + models + auth).
|
|
158
|
+
- Featured catalogs are offline bootstrap only; refresh Open Platform via `listKimiModels()`.
|
|
92
159
|
|
|
93
160
|
### Cache behavior
|
|
94
161
|
|
|
95
|
-
- Default catalog models
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
-
|
|
100
|
-
`
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
model allows long retention (`ModelConfig.cache.longRetention !== false`).
|
|
106
|
-
- The Moonshot Open Platform route (`compat.route: "openai"`) never receives
|
|
107
|
-
Anthropic `cache_control` fields.
|
|
108
|
-
- Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
|
|
109
|
-
`Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
|
|
110
|
-
`Usage.cacheWriteTokens`.
|
|
162
|
+
- Default Coding catalog models use **implicit caching** and send no explicit
|
|
163
|
+
`cache_control` fields unless the model opts in with
|
|
164
|
+
`ModelConfig.cache.kind: "cache_control"`.
|
|
165
|
+
- When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
|
|
166
|
+
caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
|
|
167
|
+
block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
|
|
168
|
+
the model allows long retention.
|
|
169
|
+
- The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
|
|
170
|
+
- Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
|
|
171
|
+
`cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
|
|
111
172
|
|
|
112
173
|
## Security and performance notes
|
|
113
174
|
|
|
114
|
-
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
|
|
175
|
+
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
|
|
176
|
+
helpers (`readSseData`, `readBoundedResponseText`).
|
|
115
177
|
- No network calls during import, setup, build, or default tests.
|
|
116
178
|
- No automatic environment, file, keychain, or shell credential lookup.
|
|
117
|
-
-
|
|
118
|
-
and redacted from errors.
|
|
179
|
+
- Credentials are resolved per request from caller-supplied values or resolvers
|
|
180
|
+
and redacted from errors (including discovery failures).
|
|
119
181
|
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
|
|
120
182
|
but provider-owned headers (`content-type`, `user-agent`, `authorization`)
|
|
121
183
|
are applied last and cannot be overridden by caller headers.
|
|
122
|
-
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus
|
|
123
|
-
|
|
184
|
+
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus provider-specific
|
|
185
|
+
env names; default tests are network-free.
|
|
124
186
|
|
|
125
187
|
## Related APIs
|
|
126
188
|
|
|
127
189
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
128
|
-
|
|
190
|
+
caller-gated discovery, Anthropic/OpenAI routes.
|
|
191
|
+
- [Thinking and reasoning](../thinking-and-reasoning.md): portable `ThinkingLevel`
|
|
192
|
+
helpers and Kimi family mapping.
|
|
129
193
|
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
130
194
|
`resolveCredentialValue`, `redactSecrets`.
|
|
131
|
-
- [Provider
|
|
132
|
-
mapping.
|
|
195
|
+
- [Provider caching](../provider-caching.md): explicit/implicit matrix.
|
|
133
196
|
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
@@ -9,7 +9,7 @@ and implicit prefix caching.
|
|
|
9
9
|
|
|
10
10
|
The package registers a provider, default model metadata for the featured NeuralWatt
|
|
11
11
|
aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
|
|
12
|
-
`kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
|
|
12
|
+
`gemma-4-31b`, `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
|
|
13
13
|
`qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
|
|
14
14
|
method through `createExtensionKernel().load([...])`.
|
|
15
15
|
|
|
@@ -79,13 +79,15 @@ The endpoint is rate-limited to **1 request per second per customer** (429 with
|
|
|
79
79
|
from `generate()` or package setup; the caller owns throttling.
|
|
80
80
|
|
|
81
81
|
NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
|
|
82
|
-
/ `extra` escape hatches: `compat.reasoning_effort` (
|
|
83
|
-
|
|
84
|
-
`compat.
|
|
85
|
-
`
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
`
|
|
82
|
+
/ `extra` escape hatches: `compat.reasoning_effort` (OpenAI-style scale; GLM-5.2 defaults
|
|
83
|
+
unset to `max` per official docs), `compat.thinking_token_budget`,
|
|
84
|
+
`compat.chat_template_kwargs` (including `enable_thinking`, `preserve_thinking`, and
|
|
85
|
+
`clear_thinking` per official gateway docs), and `compat.tool_choice`.
|
|
86
|
+
`preserve_thinking` / `clear_thinking` compat flags are routed into `chat_template_kwargs`
|
|
87
|
+
on the wire (not top-level body fields). Prism also uses these flags for client-side
|
|
88
|
+
message serialization: `preserve_thinking: true` keeps prior assistant reasoning in
|
|
89
|
+
request history; `clear_thinking: true` drops it for the next turn. `options.extra` spreads
|
|
90
|
+
after resolved fields so per-call values and overrides win.
|
|
89
91
|
|
|
90
92
|
## Outputs / response / events
|
|
91
93
|
|
|
@@ -112,7 +114,7 @@ Example request body (OpenAI-compatible Chat Completions shape):
|
|
|
112
114
|
"messages": [{ "role": "user", "content": "Hello" }],
|
|
113
115
|
"stream": true,
|
|
114
116
|
"stream_options": { "include_usage": true },
|
|
115
|
-
"reasoning_effort": "
|
|
117
|
+
"reasoning_effort": "max",
|
|
116
118
|
"thinking_token_budget": 8192
|
|
117
119
|
}
|
|
118
120
|
```
|
|
@@ -192,6 +194,7 @@ validation. The caller owns throttling/caching — the helper makes one explicit
|
|
|
192
194
|
| `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
|
|
193
195
|
| `glm-5.2-short` | 195K | Tools, reasoning |
|
|
194
196
|
| `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
|
|
197
|
+
| `gemma-4-31b` | 256K | Tools, vision, JSON mode |
|
|
195
198
|
| `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
|
|
196
199
|
| `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
|
|
197
200
|
| `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
|
|
@@ -255,15 +258,18 @@ so multi-turn sessions continue the earlier chain of thought:
|
|
|
255
258
|
|
|
256
259
|
- Prior `thinking` content blocks on an assistant message are serialized under a
|
|
257
260
|
`reasoning_content` field on that message (matching the streaming
|
|
258
|
-
`delta.reasoning_content` field
|
|
261
|
+
`delta.reasoning_content` field; the gateway also accepts `reasoning` as an alias).
|
|
262
|
+
They are **not** flattened into text `content`, so
|
|
259
263
|
the model sees reasoning and answer as distinct.
|
|
260
264
|
- Preservation is gated on `model.capabilities.reasoning === true` **or**
|
|
261
265
|
`compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
|
|
262
266
|
field and prior `thinking` blocks are dropped — they never leak into text content for
|
|
263
267
|
providers/models that do not support reasoning.
|
|
264
|
-
- `compat.
|
|
265
|
-
|
|
266
|
-
|
|
268
|
+
- `compat.preserve_thinking` / `compat.clear_thinking` map into `chat_template_kwargs`
|
|
269
|
+
on the request body per official NeuralWatt docs (Kimi K2.6 `preserve_thinking`, GLM
|
|
270
|
+
`clear_thinking: false` for full-history). Prism also uses `clear_thinking: true` to
|
|
271
|
+
drop prior reasoning client-side even on reasoning-capable models; `clear_thinking`
|
|
272
|
+
takes precedence over `preserve_thinking`.
|
|
267
273
|
- The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
|
|
268
274
|
reasoning.
|
|
269
275
|
|
package/docs/providers/openai.md
CHANGED
|
@@ -36,16 +36,22 @@ createOpenAIProviderPackage(options: OpenAIProviderPackageOptions): ProviderPack
|
|
|
36
36
|
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
37
37
|
| `baseUrl` | `string` | Overrides `https://api.openai.com/v1`. |
|
|
38
38
|
| `codexBaseUrl` | `string` | Overrides `https://chatgpt.com/backend-api/codex`. |
|
|
39
|
+
| `models` | `readonly ModelConfig[]` | Optional override for registered OpenAI Responses models (defaults to featured `openAIModels`). |
|
|
40
|
+
| `codexModels` | `readonly ModelConfig[]` | Optional override for registered Codex models (defaults to featured `openAICodexModels`). |
|
|
39
41
|
|
|
40
42
|
`ProviderRequest.options.sessionId`, `cacheKey`, `cacheRetention`, `headers`,
|
|
41
|
-
`compat`, and `extra` map to request headers/payload fields.
|
|
43
|
+
`compat`, and `extra` map to request headers/payload fields. Per-turn reasoning
|
|
44
|
+
uses official Responses `reasoning: { effort, summary? }` via
|
|
45
|
+
`ModelConfig.compat.reasoning` defaults merged with
|
|
46
|
+
`ProviderRequestOptions.compat.reasoning` (request wins). Prefer
|
|
47
|
+
`applyThinkingLevel(..., "openai_reasoning")` from `@arnilo/prism`.
|
|
42
48
|
|
|
43
49
|
## Outputs / response / events
|
|
44
50
|
|
|
45
51
|
| Surface | Behavior |
|
|
46
52
|
| --- | --- |
|
|
47
53
|
| Provider stream | Prism text, thinking (downgraded to text), `tool_call` deltas/finals, `usage`, `done`, redacted `error` events. |
|
|
48
|
-
| Block preservation |
|
|
54
|
+
| Block preservation | User/system text → `input_text`; assistant text → `output_text`; assistant `tool_call` → top-level `function_call` with `call_id`; `tool_result` → top-level `function_call_output`; images/files/audio when declared on the model. Bare thinking without an encrypted Responses reasoning item is omitted on replay. |
|
|
49
55
|
| Auth methods | `api_key` for `openai`; `oauth` for `openai-codex`. |
|
|
50
56
|
|
|
51
57
|
Unsupported block placements or unclaimed images fail before `fetch`.
|
|
@@ -56,10 +62,17 @@ Responses request body (Codex subscription shape, abbreviated):
|
|
|
56
62
|
|
|
57
63
|
```json
|
|
58
64
|
{
|
|
59
|
-
"model": "gpt-5
|
|
60
|
-
"
|
|
61
|
-
|
|
62
|
-
|
|
65
|
+
"model": "gpt-5.1",
|
|
66
|
+
"input": [
|
|
67
|
+
{ "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] },
|
|
68
|
+
{ "role": "assistant", "content": [{ "type": "output_text", "text": "Calling lookup" }] },
|
|
69
|
+
{ "type": "function_call", "call_id": "call_1", "name": "lookup", "arguments": "{\"q\":\"x\"}" },
|
|
70
|
+
{ "type": "function_call_output", "call_id": "call_1", "output": "{\"ok\":true}" }
|
|
71
|
+
],
|
|
72
|
+
"reasoning": { "effort": "high" },
|
|
73
|
+
"prompt_cache_key": "session-1",
|
|
74
|
+
"stream": true,
|
|
75
|
+
"store": false
|
|
63
76
|
}
|
|
64
77
|
```
|
|
65
78
|
|
|
@@ -73,13 +86,13 @@ https://auth.openai.com/authorize?response_type=code&client_id=...&code_challeng
|
|
|
73
86
|
|
|
74
87
|
```ts
|
|
75
88
|
import { createExtensionKernel, createEnvCredentialResolver } from "@arnilo/prism";
|
|
76
|
-
import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
|
|
89
|
+
import { createOpenAIProviderPackage, listOpenAIModels } from "@arnilo/prism-provider-openai";
|
|
77
90
|
|
|
91
|
+
const apiKey = createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" });
|
|
92
|
+
const models = await listOpenAIModels({ apiKey }); // caller-gated; never runs during setup
|
|
78
93
|
const kernel = createExtensionKernel();
|
|
79
94
|
await kernel.load([
|
|
80
|
-
createOpenAIProviderPackage({
|
|
81
|
-
apiKey: createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" }),
|
|
82
|
-
}),
|
|
95
|
+
createOpenAIProviderPackage({ apiKey, models }),
|
|
83
96
|
]);
|
|
84
97
|
```
|
|
85
98
|
|
|
@@ -115,25 +128,55 @@ const challenge = computeS256Challenge(verifier);
|
|
|
115
128
|
|
|
116
129
|
### Cache behavior
|
|
117
130
|
|
|
131
|
+
Official: [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching).
|
|
132
|
+
|
|
118
133
|
- `prompt_cache_key` is derived from `ProviderRequestOptions.cacheKey` (falling
|
|
119
134
|
back to `sessionId`) and sanitized + clamped to 64 characters via the shared
|
|
120
135
|
`sanitizeCacheKey()` helper. Cache keys are session/customer identifiers only;
|
|
121
136
|
never credentials or raw prompts.
|
|
122
|
-
- `prompt_cache_retention` accepts only `"24h"` on
|
|
137
|
+
- `prompt_cache_retention` accepts only `"24h"` on pre-GPT-5.6 Responses models
|
|
123
138
|
(extended caching). Prism `cacheRetention: "short"` and `"none"` omit the field
|
|
124
139
|
so default automatic/implicit caching applies and no invalid literal is sent.
|
|
125
140
|
`cacheRetention: "long"` maps to `prompt_cache_retention: "24h"` only when the
|
|
126
141
|
model declares `ModelConfig.cache.longRetention === true`; models without that
|
|
127
|
-
metadata omit the field.
|
|
142
|
+
metadata omit the field. Featured `gpt-5.1` declares
|
|
128
143
|
`cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
|
|
144
|
+
- GPT-5.6+ official docs prefer `prompt_cache_options` / explicit breakpoints;
|
|
145
|
+
`listOpenAIModels` sets `longRetention: false` for those ids so Prism does not
|
|
146
|
+
emit deprecated `prompt_cache_retention` for them. Breakpoint helpers are not
|
|
147
|
+
shipped in this package yet — hosts may pass `prompt_cache_options` through
|
|
148
|
+
`compat` / `extra` when needed.
|
|
129
149
|
- Cache accounting is preserved in normalized `Usage`: OpenAI
|
|
130
150
|
`input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
|
|
131
|
-
Responses does not report a cache-write token field.
|
|
151
|
+
Responses does not report a cache-write token field on older models.
|
|
132
152
|
- Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
|
|
133
153
|
are applied after caller `ProviderRequestOptions.headers` so caller config
|
|
134
154
|
cannot replace credentials, content type, or the session request id; non-owned
|
|
135
155
|
caller headers are kept.
|
|
136
156
|
|
|
157
|
+
### Model discovery
|
|
158
|
+
|
|
159
|
+
- `listOpenAIModels({ apiKey, fetch, baseUrl, signal, headers })` calls official
|
|
160
|
+
[`GET /models`](https://developers.openai.com/api/reference/resources/models/methods/list)
|
|
161
|
+
and maps sparse `{ id, created, owned_by }` entries to `ModelConfig` with
|
|
162
|
+
`cache.kind: "openai_key"` and heuristic `longRetention` / `capabilities.reasoning`.
|
|
163
|
+
- `createOpenAIProviderPackage` never calls discovery; pass results via `models:`.
|
|
164
|
+
- Codex subscription models are **not** listed by `api.openai.com` — keep using
|
|
165
|
+
featured `openAICodexModels` or `codexModels:` override.
|
|
166
|
+
- Static `openAIModels` / `openAICodexModels` are offline bootstrap / featured aliases only.
|
|
167
|
+
|
|
168
|
+
### Reasoning
|
|
169
|
+
|
|
170
|
+
Official: [Reasoning models](https://developers.openai.com/api/docs/guides/reasoning).
|
|
171
|
+
|
|
172
|
+
- Body field is top-level `reasoning: { effort, summary?, mode?, context? }`.
|
|
173
|
+
- Model defaults: `ModelConfig.compat.reasoning`; per-turn override:
|
|
174
|
+
`ProviderRequestOptions.compat.reasoning` (shallow-merged; request wins).
|
|
175
|
+
- Portable helper: `applyThinkingLevel(options, level, "openai_reasoning")`.
|
|
176
|
+
- Streaming tool args follow official
|
|
177
|
+
`response.output_item.added` + `response.function_call_arguments.delta`
|
|
178
|
+
(string `delta`), not Chat Completions object deltas.
|
|
179
|
+
|
|
137
180
|
## Security and performance notes
|
|
138
181
|
|
|
139
182
|
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
|