@arnilo/prism 0.0.4 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/CHANGELOG.md +46 -1
  2. package/README.md +34 -10
  3. package/dist/agent-loops.d.ts +1 -0
  4. package/dist/agent-loops.js +26 -16
  5. package/dist/agents.js +147 -21
  6. package/dist/cli-init.d.ts +41 -0
  7. package/dist/cli-init.js +390 -0
  8. package/dist/cli-runner.d.ts +7 -1
  9. package/dist/cli-runner.js +13 -1
  10. package/dist/content.d.ts +19 -0
  11. package/dist/content.js +197 -69
  12. package/dist/contracts.d.ts +96 -9
  13. package/dist/contracts.js +8 -0
  14. package/dist/feedback.d.ts +48 -0
  15. package/dist/feedback.js +230 -0
  16. package/dist/ids.d.ts +2 -0
  17. package/dist/ids.js +6 -0
  18. package/dist/index.d.ts +10 -4
  19. package/dist/index.js +6 -3
  20. package/dist/providers/media.d.ts +3 -1
  21. package/dist/providers/media.js +11 -1
  22. package/dist/session-stores.js +2 -3
  23. package/dist/testing/feedback.d.ts +6 -0
  24. package/dist/testing/feedback.js +37 -0
  25. package/dist/testing/persistence-schema.d.ts +48 -10
  26. package/dist/testing/persistence-schema.js +166 -22
  27. package/dist/testing/run-ledger-conformance.js +7 -1
  28. package/dist/thinking.d.ts +42 -0
  29. package/dist/thinking.js +92 -0
  30. package/dist/tools.js +2 -3
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +75 -0
  34. package/docs/agent-events.md +14 -21
  35. package/docs/agent-loops.md +12 -9
  36. package/docs/agent-session-runtime.md +14 -16
  37. package/docs/cli-rpc.md +35 -7
  38. package/docs/coding-agent-tools.md +35 -14
  39. package/docs/coding-security.md +7 -3
  40. package/docs/compaction-llm.md +17 -7
  41. package/docs/compaction-observational-memory.md +30 -4
  42. package/docs/context-and-skills.md +1 -0
  43. package/docs/credential-storage.md +58 -9
  44. package/docs/credentials-and-redaction.md +3 -3
  45. package/docs/database-persistence.md +17 -9
  46. package/docs/evaluations.md +122 -0
  47. package/docs/extensions.md +2 -2
  48. package/docs/host-security.md +26 -5
  49. package/docs/index.md +43 -28
  50. package/docs/mcp-tools.md +74 -13
  51. package/docs/migration.md +177 -3
  52. package/docs/multimodal-content.md +14 -6
  53. package/docs/node-filesystem-config.md +1 -0
  54. package/docs/node-jsonl-session-store.md +5 -4
  55. package/docs/observability.md +14 -6
  56. package/docs/performance.md +209 -0
  57. package/docs/postgres-persistence.md +8 -6
  58. package/docs/provider-caching.md +16 -4
  59. package/docs/provider-conformance.md +40 -1
  60. package/docs/provider-packages.md +62 -3
  61. package/docs/providers/ai-sdk.md +149 -0
  62. package/docs/providers/kimi.md +124 -61
  63. package/docs/providers/neuralwatt.md +19 -13
  64. package/docs/providers/openai.md +56 -13
  65. package/docs/providers/opencode-go.md +118 -30
  66. package/docs/providers/openrouter.md +105 -35
  67. package/docs/providers/zai.md +94 -45
  68. package/docs/public-contracts.md +6 -5
  69. package/docs/rag.md +113 -0
  70. package/docs/release-and-install.md +100 -79
  71. package/docs/review-coverage-2026-07-15.md +193 -0
  72. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  73. package/docs/runs-and-usage.md +42 -5
  74. package/docs/server.md +139 -0
  75. package/docs/settings-auth-trust-security.md +5 -5
  76. package/docs/sqlite-persistence.md +6 -5
  77. package/docs/structured-output.md +1 -1
  78. package/docs/supervisors.md +71 -0
  79. package/docs/thinking-and-reasoning.md +98 -0
  80. package/docs/tool-execution-primitives.md +3 -3
  81. package/docs/tools.md +15 -0
  82. package/docs/use-case-model-selection.md +109 -0
  83. package/docs/workflow-orchestration-primitives.md +20 -3
  84. package/docs/workflows.md +114 -33
  85. package/docs/working-and-semantic-memory.md +170 -0
  86. package/package.json +13 -3
  87. package/templates/init/README.md.tmpl +28 -0
  88. package/templates/init/env.example.tmpl +1 -0
  89. package/templates/init/gitignore.tmpl +11 -0
  90. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  91. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  92. package/templates/init/package.json.tmpl +22 -0
  93. package/templates/init/providers.json +76 -0
  94. package/templates/init/src/agent.ts.tmpl +10 -0
  95. package/templates/init/src/index.ts.tmpl +12 -0
  96. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  97. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -0,0 +1,149 @@
1
+ # AI SDK provider adapter
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-provider-ai-sdk` adapts a host-supplied AI SDK `LanguageModelV4` into a Prism `AIProvider`. It maps Prism messages, tools, and structured-output options into `doStream` call options, then translates stream parts into Prism provider events incrementally.
6
+
7
+ Supported specification: `@ai-sdk/provider` **v4** (`specificationVersion: "v4"`). Core `@arnilo/prism` does not depend on the AI SDK.
8
+
9
+ ## When to use it
10
+
11
+ Use this package when a host already creates AI SDK language models and wants them inside Prism agent/session loops without adding another first-party HTTP provider.
12
+
13
+ Do not use it as a credential store, model catalog, or high-level `streamText`/`generateText` replacement. Hosts keep owning credentials inside the supplied model.
14
+
15
+ ## Inputs / request
16
+
17
+ ```ts
18
+ import { createAiSdkProvider } from "@arnilo/prism-provider-ai-sdk";
19
+
20
+ createAiSdkProvider(options: {
21
+ model: LanguageModelV4;
22
+ id?: string;
23
+ }): AIProvider
24
+ ```
25
+
26
+ | Field | Type | Purpose |
27
+ | --- | --- | --- |
28
+ | `model` | `LanguageModelV4` | Host-owned AI SDK language model. |
29
+ | `id` | `string` | Prism provider id. Defaults to `ai-sdk:<model.provider>` or `ai-sdk`. |
30
+
31
+ Mapped request surfaces:
32
+
33
+ | Prism | AI SDK |
34
+ | --- | --- |
35
+ | `messages` | `LanguageModelV4Prompt` |
36
+ | `tools` | `LanguageModelV4FunctionTool[]` with JSON Schema `inputSchema` |
37
+ | `options.structuredOutput` | `responseFormat: { type: "json", name, schema }` |
38
+ | `model.parameters` | `maxOutputTokens`, `temperature`, `topP`, `topK`, penalties, `seed`, `stopSequences` |
39
+ | `request.signal` | `abortSignal` (always wins over adapter options) |
40
+ | `options.headers` | extension headers only; model owns auth |
41
+
42
+ Unsupported content fails before `doStream` (for example unresolved `resourceUri`, audio/file/document without declared capability, `tool_call_delta` in history, non-text system content).
43
+
44
+ ## Outputs / response / events
45
+
46
+ | AI SDK stream part | Prism event |
47
+ | --- | --- |
48
+ | `text-delta` | `content_delta` text |
49
+ | `reasoning-delta` | `content_delta` thinking |
50
+ | `tool-input-start` / `tool-input-delta` | `tool_call_delta` |
51
+ | `tool-call` (client-executed) | `tool_call` |
52
+ | `finish` usage | `usage` then `done` |
53
+ | `error` / thrown / abort | redacted `error` |
54
+
55
+ `finish.usage.inputTokens.cacheRead` / `cacheWrite` map to Prism `Usage.cacheReadTokens` / `cacheWriteTokens`. The adapter does not invent cache request fields; prompt caching is owned by the host `LanguageModelV4` and its upstream provider.
56
+
57
+ Provider-executed tool calls, files/sources/custom parts, warnings, and raw chunks are ignored rather than silently converted into unsupported Prism content.
58
+
59
+ ## Request/response example
60
+
61
+ ```json
62
+ {
63
+ "prompt": [{ "role": "user", "content": [{ "type": "text", "text": "hello" }] }],
64
+ "tools": [{ "type": "function", "name": "echo", "inputSchema": { "type": "object" } }],
65
+ "responseFormat": {
66
+ "type": "json",
67
+ "name": "Answer",
68
+ "schema": { "type": "object", "properties": { "ok": { "type": "boolean" } } }
69
+ }
70
+ }
71
+ ```
72
+
73
+ ## Implementation example
74
+
75
+ ```ts
76
+ import { createAgent } from "@arnilo/prism";
77
+ import { createAiSdkProvider } from "@arnilo/prism-provider-ai-sdk";
78
+
79
+ const provider = createAiSdkProvider({ model: hostCreatedLanguageModelV4 });
80
+
81
+ const agent = createAgent({
82
+ provider,
83
+ model: {
84
+ provider: provider.id,
85
+ model: hostCreatedLanguageModelV4.modelId,
86
+ capabilities: { tools: true, streaming: true, structuredOutput: true, input: ["text"] },
87
+ },
88
+ });
89
+
90
+ const result = await agent.createSession().run("Summarize this");
91
+ console.log(result.text);
92
+ ```
93
+
94
+ ## Model catalog and discovery
95
+
96
+ There is **no Prism-side model catalog** and **no `list*Models()` export** by design. Hosts supply a ready-made `LanguageModelV4` instance (typically from `@ai-sdk/openai`, `@ai-sdk/anthropic`, AI Gateway, or a custom provider) and register a matching `ModelConfig` for capabilities/limits.
97
+
98
+ Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials.
99
+
100
+ ## Prompt caching
101
+
102
+ The adapter is **host-owned for request caching**. It does not emit `cache_control`, `prompt_cache_key`, `cacheKey`, or `cacheRetention` on AI SDK call options. Hosts configure caching on the underlying AI SDK model/provider (for example via AI SDK `providerOptions` on the model factory or per-call options forwarded through `options.compat` / `options.extra` → `providerOptions.prism`).
103
+
104
+ When the host model reports cache accounting on the `finish` stream part, Prism maps official AI SDK v4 usage fields:
105
+
106
+ | AI SDK `LanguageModelV4Usage` | Prism `Usage` |
107
+ | --- | --- |
108
+ | `inputTokens.cacheRead` | `cacheReadTokens` |
109
+ | `inputTokens.cacheWrite` | `cacheWriteTokens` |
110
+ | `inputTokens.total` | `inputTokens` |
111
+ | `outputTokens.total` | `outputTokens` |
112
+
113
+ See [Provider caching](../provider-caching.md) for the cross-provider matrix.
114
+
115
+ ## Thinking and reasoning
116
+
117
+ Reasoning effort, budgets, and provider-specific thinking controls are **host-model-owned**. Prism does not map `ThinkingLevel` into AI SDK call options (`thinkingFamilyForModel` → `noop`). Hosts configure reasoning on the AI SDK model (for example OpenAI `reasoning.effort` via AI SDK `providerOptions`) and may pass per-turn overrides through `ProviderRequestOptions.compat` / `extra`, which the adapter forwards as `providerOptions.prism`.
118
+
119
+ Stream mapping:
120
+
121
+ | Direction | Mapping |
122
+ | --- | --- |
123
+ | AI SDK `reasoning-delta` → Prism | `content_delta` thinking |
124
+ | Prism `thinking` blocks → AI SDK prompt | `{ type: "reasoning", text }` on assistant messages |
125
+
126
+ Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); [Language Model Specification V4](https://github.com/vercel/ai/tree/main/packages/provider/src/language-model/v4); `@ai-sdk/provider` `LanguageModelV4Usage` (`inputTokens.cacheRead` / `cacheWrite`).
127
+
128
+ ## Extension and configuration notes
129
+
130
+ - Peer dependency: `@ai-sdk/provider@^4.0.0`. Upgrade policy tracks one specification major at a time.
131
+ - First-party HTTP providers remain independent; this adapter is available directly, through `@arnilo/prism-providers`, or through `@arnilo/prism-all`. Installation does not select a model or invoke AI SDK.
132
+ - `options.compat` / `options.extra` pass through as AI SDK `providerOptions.prism`.
133
+ - Export helpers `toAiSdkCallOptions`, `toAiSdkPrompt`, and `mapAiSdkStream` for tests and custom hosts.
134
+
135
+ ## Security and performance notes
136
+
137
+ - Host credentials stay inside the supplied AI SDK model. The adapter never reads env keys or credential stores.
138
+ - Abort and resource limits come from Prism `request.signal`; adapter options cannot replace that bound.
139
+ - Stream parts are translated incrementally with no full-response buffering and no duplicate model call.
140
+ - Unsupported content fails closed before model invocation. Errors use Prism `providerError` redaction.
141
+ - Provider metadata/warnings are not emitted as prompt or tool content.
142
+
143
+ ## Related APIs
144
+
145
+ - [Provider packages](../provider-packages.md)
146
+ - [Provider conformance](../provider-conformance.md)
147
+ - [Provider layer](../provider-layer.md)
148
+ - [Structured output](../structured-output.md)
149
+ - [Agent session runtime](../agent-session-runtime.md)
@@ -2,132 +2,195 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-provider-kimi` provides explicit, side-effect-free setup for Kimi For
6
- Coding using an Anthropic-compatible `/messages` endpoint with
7
- `User-Agent: KimiCLI/1.5` (unless overridden). Moonshot/Open Platform model
8
- metadata is optional.
5
+ `@arnilo/prism-provider-kimi` provides two distinct, side-effect-free routes:
9
6
 
10
- The package registers the `kimi-coding` provider, default Kimi Coding model
11
- metadata, and an `api_key` auth method through `createExtensionKernel().load([...])`.
7
+ 1. **Kimi For Coding** (default) — Anthropic-compatible `POST /messages` on
8
+ `https://api.kimi.com/coding` with `User-Agent: KimiCLI/1.5` (unless overridden).
9
+ 2. **Moonshot Open Platform** (opt-in) — OpenAI-compatible `POST /chat/completions`
10
+ on `https://api.moonshot.ai/v1` (or `api.moonshot.cn/v1`), registered only when
11
+ `includeMoonshotModels: true`.
12
+
13
+ Official model ids differ by route. Coding uses `kimi-for-coding`,
14
+ `kimi-for-coding-highspeed`, and `k3`. Open Platform uses `kimi-k2.7-code`,
15
+ `kimi-k3`, and related catalog ids. Pi's `k2p7` alias is **not** used.
16
+
17
+ Caller-gated discovery via `listKimiModels()` hits the official Moonshot
18
+ `GET /v1/models` endpoint. Package setup never fetches.
12
19
 
13
20
  ## When to use it
14
21
 
15
- Use it when a host app wants the Kimi For Coding endpoint through Prism's
16
- `AgentSession` runtime with Kimi-specific serializer behavior.
22
+ Use it when a host app wants Kimi For Coding and/or Moonshot Open Platform through
23
+ Prism's `AgentSession` runtime with Kimi-specific serializers, thinking controls,
24
+ and cache policy.
17
25
 
18
- Do not use it for Moonshot Open Platform default registration, automatic
19
- credential discovery, catalog fetches, or real-network tests.
26
+ Do not use it for automatic credential discovery, setup-time catalog fetches, or
27
+ real-network tests (live tests stay opt-in).
20
28
 
21
29
  ## Inputs / request
22
30
 
23
31
  ```ts
24
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
32
+ import {
33
+ createKimiProviderPackage,
34
+ listKimiModels,
35
+ defineKimiModel,
36
+ } from "@arnilo/prism-provider-kimi";
25
37
 
26
38
  createKimiProviderPackage(options: KimiProviderPackageOptions): ProviderPackage
27
- defineKimiModel(config: KimiModelConfig): KimiModelConfig
39
+ listKimiModels(options?: ListKimiModelsOptions): Promise<ModelConfig[]>
40
+ defineKimiModel(config: KimiModelConfig): ModelConfig
28
41
  ```
29
42
 
30
43
  | Field | Type | Purpose |
31
44
  | --- | --- | --- |
32
- | `kimiApiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source for Kimi. |
45
+ | `kimiApiKey` | `CredentialValueSource` | Kimi For Coding API key. |
46
+ | `moonshotApiKey` | `CredentialValueSource` | Moonshot Open Platform API key (not interchangeable with Coding keys). |
33
47
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
34
- | `baseUrl` | `string` | Overrides the Kimi base URL. |
35
- | `id` | `string` | Overrides the provider id (default `kimi-coding`). |
36
- | `userAgent` | `string` | Overrides `User-Agent: KimiCLI/1.5`. |
37
- | `models` | `readonly ModelConfig[]` | Overrides `kimiCodingModels` defaults. |
38
- | `includeMoonshotModels` | `boolean` | Registers Moonshot models when `true` (default off). |
39
- | `moonshotModels` | `readonly ModelConfig[]` | Overrides `moonshotKimiModels` when included. |
48
+ | `baseUrl` | `string` | Overrides the Coding base URL. |
49
+ | `moonshotBaseUrl` | `string` | Overrides Moonshot base URL (default `https://api.moonshot.ai/v1`). |
50
+ | `id` / `moonshotId` | `string` | Provider ids (defaults `kimi-coding` / `moonshot`). |
51
+ | `userAgent` | `string` | Overrides Coding `User-Agent: KimiCLI/1.5`. |
52
+ | `models` | `readonly ModelConfig[]` | Overrides featured Coding models. |
53
+ | `includeMoonshotModels` | `boolean` | Registers callable Moonshot provider + models when `true`. |
54
+ | `moonshotModels` | `readonly ModelConfig[]` | Overrides featured Moonshot models when included. |
40
55
 
41
56
  ## Outputs / response / events
42
57
 
43
58
  | Surface | Behavior |
44
59
  | --- | --- |
45
- | Provider stream | Prism text, thinking (preserved only when `model.compat.preserveThinking` is true, otherwise downgraded to text), tool-call delta/final, `usage`, `done`, redacted `error`. |
46
- | Block preservation | Text, thinking, assistant `tool_call` → `tool_use`, `tool_result` → `tool_result`, images when `capabilities.input` includes `"image"`. |
47
- | Auth method | `api_key` for `kimi-coding`, credential name `apiKey`. |
60
+ | Coding stream | Prism text, thinking deltas, tool-call delta/final, `usage` (`cache_read_input_tokens` / `cache_creation_input_tokens`), `done`, redacted `error`. |
61
+ | Moonshot stream | Same, with `delta.reasoning_content` → thinking; OpenAI-style usage cache details when present. |
62
+ | Block preservation | Coding: Anthropic `thinking` / `tool_use` / `tool_result`. Moonshot: `reasoning_content` on assistant replay when `preserveThinking`. |
63
+ | Auth methods | `api_key` for `kimi-coding`; also `moonshot` when opted in. |
48
64
 
49
65
  Unsupported block placements or unclaimed images fail before fetch.
50
66
 
67
+ ## Route differences
68
+
69
+ | | Kimi For Coding | Moonshot Open Platform |
70
+ | --- | --- | --- |
71
+ | Base URL | `https://api.kimi.com/coding` | `https://api.moonshot.ai/v1` (or `.cn`) |
72
+ | Wire API | Anthropic `/messages` | OpenAI `/chat/completions` |
73
+ | Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k3` (+ discovery) |
74
+ | Discovery | No public list API — curated featured aliases | Official `GET /v1/models` via `listKimiModels()` |
75
+ | Cache | Implicit by default; opt-in Anthropic `cache_control` | Implicit only — never emits Anthropic `cache_control` |
76
+ | Thinking | Block replay + body `thinking` / `reasoning_effort` | `reasoning_content` replay + body `thinking` / `reasoning_effort` |
77
+
78
+ The Anthropic `/messages` request/response contract for Kimi remains under-documented
79
+ upstream ([MoonshotAI/Kimi-K2#129](https://github.com/MoonshotAI/Kimi-K2/issues/129));
80
+ Prism treats official Chat Completions thinking fields as best-effort passthrough on
81
+ the Coding route.
82
+
83
+ ## Thinking / reasoning
84
+
85
+ Official fields (Open Platform docs; Coding docs for `k3` effort mapping):
86
+
87
+ | Model family | Official control | Prism `compat` |
88
+ | --- | --- | --- |
89
+ | K3 / Coding `k3` | top-level `reasoning_effort` (`max` on Open Platform; Coding also `low`/`high`) | `compat.reasoning_effort` — use Task 4 family `reasoning_effort` |
90
+ | K2.7-code / Coding | thinking always on; Preserved Thinking always on | omit `thinking` by default; `preserveThinking: true` for replay; do not send `disabled` |
91
+ | K2.6 / K2.5 | `thinking.type` enabled/disabled; K2.6 optional `keep: "all"` | `compat.thinking` — Task 4 family `thinking_type` |
92
+
93
+ Per-turn `ProviderRequestOptions.compat` wins over `ModelConfig.compat`. Helpers:
94
+ `kimiThinking`, `kimiReasoningEffort`, `kimiPreserveThinking`.
95
+
51
96
  ## Request/response example
52
97
 
53
- Example request (Anthropic-compatible `/messages` shape):
98
+ Coding (Anthropic-compatible `/messages`):
54
99
 
55
100
  ```json
56
101
  {
57
- "model": "kimi-latest",
58
- "messages": [{ "role": "user", "content": "Hello" }],
102
+ "model": "kimi-for-coding",
103
+ "messages": [{ "role": "user", "content": [{ "type": "text", "text": "Hello" }] }],
59
104
  "stream": true
60
105
  }
61
106
  ```
62
107
 
108
+ Moonshot (Chat Completions):
109
+
110
+ ```json
111
+ {
112
+ "model": "kimi-k3",
113
+ "messages": [{ "role": "user", "content": "Hello" }],
114
+ "stream": true,
115
+ "reasoning_effort": "max"
116
+ }
117
+ ```
118
+
63
119
  ## Implementation example
64
120
 
65
121
  ```ts
66
122
  import { createExtensionKernel } from "@arnilo/prism";
67
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
123
+ import {
124
+ createKimiProviderPackage,
125
+ listKimiModels,
126
+ } from "@arnilo/prism-provider-kimi";
68
127
 
69
128
  const kernel = createExtensionKernel();
70
129
  await kernel.load([
71
- createKimiProviderPackage({ kimiApiKey: "fake-kimi-key", includeMoonshotModels: false }),
130
+ createKimiProviderPackage({ kimiApiKey: "fake-kimi-key" }),
72
131
  ]);
73
- ```
74
-
75
- Register Moonshot/Open Platform metadata explicitly:
76
132
 
77
- ```ts
78
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
133
+ // Opt-in Moonshot Open Platform (callable provider + featured models)
134
+ await kernel.load([
135
+ createKimiProviderPackage({
136
+ kimiApiKey: "fake-kimi-key",
137
+ includeMoonshotModels: true,
138
+ moonshotApiKey: "fake-moonshot-key",
139
+ }),
140
+ ]);
79
141
 
142
+ // Caller-gated discovery — never runs during setup
143
+ const latest = await listKimiModels({ apiKey: "fake-moonshot-key", fetch });
80
144
  await kernel.load([
81
- createKimiProviderPackage({ kimiApiKey: "fake", includeMoonshotModels: true }),
145
+ createKimiProviderPackage({
146
+ includeMoonshotModels: true,
147
+ moonshotApiKey: "fake-moonshot-key",
148
+ moonshotModels: latest.filter((m) => m.model.startsWith("kimi-")),
149
+ }),
82
150
  ]);
83
151
  ```
84
152
 
85
153
  ## Extension and configuration notes
86
154
 
87
- - Hosts choose base URL, provider id, `User-Agent`, model list, credential source,
155
+ - Hosts choose base URLs, provider ids, `User-Agent`, model lists, credential sources,
88
156
  and `fetch` impl.
89
- - Moonshot/Open Platform metadata is registered only with
90
- `includeMoonshotModels: true`; it is not core behavior.
91
- - Package contributes models via the extension `api` and an `api_key` auth method.
157
+ - Moonshot is registered only with `includeMoonshotModels: true` (provider + models + auth).
158
+ - Featured catalogs are offline bootstrap only; refresh Open Platform via `listKimiModels()`.
92
159
 
93
160
  ### Cache behavior
94
161
 
95
- - Default catalog models (e.g. `kimi-k2.7-code` on the Anthropic-compatible
96
- `/messages` route) use **implicit caching** and send no explicit `cache_control`
97
- fields. `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention` have no
98
- effect on the request body unless the model opts in.
99
- - Hosts may opt a model into Anthropic-style `cache_control` by declaring
100
- `ModelConfig.cache.kind: "cache_control"` on the Anthropic route. When opted in,
101
- `cache_control: { type: "ephemeral" }` markers are applied only to the
102
- caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
103
- shared `applyCacheControl()` helper) on the last content block of each selected
104
- message — not to every block. `cacheRetention: "long"` adds `ttl: "1h"` when the
105
- model allows long retention (`ModelConfig.cache.longRetention !== false`).
106
- - The Moonshot Open Platform route (`compat.route: "openai"`) never receives
107
- Anthropic `cache_control` fields.
108
- - Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
109
- `Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
110
- `Usage.cacheWriteTokens`.
162
+ - Default Coding catalog models use **implicit caching** and send no explicit
163
+ `cache_control` fields unless the model opts in with
164
+ `ModelConfig.cache.kind: "cache_control"`.
165
+ - When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
166
+ caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
167
+ block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
168
+ the model allows long retention.
169
+ - The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
170
+ - Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
171
+ `cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
111
172
 
112
173
  ## Security and performance notes
113
174
 
114
- - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
175
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
176
+ helpers (`readSseData`, `readBoundedResponseText`).
115
177
  - No network calls during import, setup, build, or default tests.
116
178
  - No automatic environment, file, keychain, or shell credential lookup.
117
- - Kimi credentials are resolved per request from caller-supplied values or resolvers
118
- and redacted from errors.
179
+ - Credentials are resolved per request from caller-supplied values or resolvers
180
+ and redacted from errors (including discovery failures).
119
181
  - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
120
182
  but provider-owned headers (`content-type`, `user-agent`, `authorization`)
121
183
  are applied last and cannot be overridden by caller headers.
122
- - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
123
- provider-specific env names; default tests are network-free.
184
+ - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus provider-specific
185
+ env names; default tests are network-free.
124
186
 
125
187
  ## Related APIs
126
188
 
127
189
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
128
- `ModelConfig`/`compat`, Anthropic-compatible routes.
190
+ caller-gated discovery, Anthropic/OpenAI routes.
191
+ - [Thinking and reasoning](../thinking-and-reasoning.md): portable `ThinkingLevel`
192
+ helpers and Kimi family mapping.
129
193
  - [Credentials and redaction](../credentials-and-redaction.md):
130
194
  `resolveCredentialValue`, `redactSecrets`.
131
- - [Provider layer](../provider-layer.md): `ProviderRequest.options` and usage
132
- mapping.
195
+ - [Provider caching](../provider-caching.md): explicit/implicit matrix.
133
196
  - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -9,7 +9,7 @@ and implicit prefix caching.
9
9
 
10
10
  The package registers a provider, default model metadata for the featured NeuralWatt
11
11
  aliases (`glm-5.2`, `glm-5.2-fast`, `glm-5.2-short`, `glm-5.2-short-fast`,
12
- `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
12
+ `gemma-4-31b`, `kimi-k2.6`, `kimi-k2.6-fast`, `kimi-k2.7-code`, `qwen3.5-397b`,
13
13
  `qwen3.5-397b-fast`, `qwen3.6-35b`, `qwen3.6-35b-fast`), and an `api_key` auth
14
14
  method through `createExtensionKernel().load([...])`.
15
15
 
@@ -79,13 +79,15 @@ The endpoint is rate-limited to **1 request per second per customer** (429 with
79
79
  from `generate()` or package setup; the caller owns throttling.
80
80
 
81
81
  NeuralWatt-specific request fields flow through the generic `ProviderRequestOptions.compat`
82
- / `extra` escape hatches: `compat.reasoning_effort` (`"low" | "medium" | "high"`),
83
- `compat.thinking_token_budget`, `compat.chat_template_kwargs` (including `enable_thinking`),
84
- `compat.preserve_thinking`, `compat.clear_thinking`, and `compat.tool_choice`.
85
- `preserve_thinking: true` keeps prior assistant reasoning in request history so
86
- multi-turn reasoning continues with the earlier chain of thought; `clear_thinking:
87
- true` drops it for the next turn, resetting the chain. `options.extra` spreads after
88
- `compat` so per-call values and overrides win.
82
+ / `extra` escape hatches: `compat.reasoning_effort` (OpenAI-style scale; GLM-5.2 defaults
83
+ unset to `max` per official docs), `compat.thinking_token_budget`,
84
+ `compat.chat_template_kwargs` (including `enable_thinking`, `preserve_thinking`, and
85
+ `clear_thinking` per official gateway docs), and `compat.tool_choice`.
86
+ `preserve_thinking` / `clear_thinking` compat flags are routed into `chat_template_kwargs`
87
+ on the wire (not top-level body fields). Prism also uses these flags for client-side
88
+ message serialization: `preserve_thinking: true` keeps prior assistant reasoning in
89
+ request history; `clear_thinking: true` drops it for the next turn. `options.extra` spreads
90
+ after resolved fields so per-call values and overrides win.
89
91
 
90
92
  ## Outputs / response / events
91
93
 
@@ -112,7 +114,7 @@ Example request body (OpenAI-compatible Chat Completions shape):
112
114
  "messages": [{ "role": "user", "content": "Hello" }],
113
115
  "stream": true,
114
116
  "stream_options": { "include_usage": true },
115
- "reasoning_effort": "medium",
117
+ "reasoning_effort": "max",
116
118
  "thinking_token_budget": 8192
117
119
  }
118
120
  ```
@@ -192,6 +194,7 @@ validation. The caller owns throttling/caching — the helper makes one explicit
192
194
  | `glm-5.2-fast` | 1024K | Tools, fast/no reasoning |
193
195
  | `glm-5.2-short` | 195K | Tools, reasoning |
194
196
  | `glm-5.2-short-fast` | 195K | Tools, fast/no reasoning |
197
+ | `gemma-4-31b` | 256K | Tools, vision, JSON mode |
195
198
  | `kimi-k2.6` | 256K | Tools, reasoning, vision, JSON mode |
196
199
  | `kimi-k2.6-fast` | 256K | Tools, vision, JSON mode, fast/no reasoning |
197
200
  | `kimi-k2.7-code` | 256K | Tools, reasoning, vision, JSON mode |
@@ -255,15 +258,18 @@ so multi-turn sessions continue the earlier chain of thought:
255
258
 
256
259
  - Prior `thinking` content blocks on an assistant message are serialized under a
257
260
  `reasoning_content` field on that message (matching the streaming
258
- `delta.reasoning_content` field). They are **not** flattened into text `content`, so
261
+ `delta.reasoning_content` field; the gateway also accepts `reasoning` as an alias).
262
+ They are **not** flattened into text `content`, so
259
263
  the model sees reasoning and answer as distinct.
260
264
  - Preservation is gated on `model.capabilities.reasoning === true` **or**
261
265
  `compat.preserve_thinking: true`. Non-reasoning models receive no `reasoning_content`
262
266
  field and prior `thinking` blocks are dropped — they never leak into text content for
263
267
  providers/models that do not support reasoning.
264
- - `compat.clear_thinking: true` drops prior reasoning for the next turn even on
265
- reasoning-capable models, resetting the chain of thought. `clear_thinking` takes
266
- precedence over `preserve_thinking`.
268
+ - `compat.preserve_thinking` / `compat.clear_thinking` map into `chat_template_kwargs`
269
+ on the request body per official NeuralWatt docs (Kimi K2.6 `preserve_thinking`, GLM
270
+ `clear_thinking: false` for full-history). Prism also uses `clear_thinking: true` to
271
+ drop prior reasoning client-side even on reasoning-capable models; `clear_thinking`
272
+ takes precedence over `preserve_thinking`.
267
273
  - The provider only echoes caller-provided `thinking` blocks; it never synthesizes new
268
274
  reasoning.
269
275
 
@@ -36,16 +36,22 @@ createOpenAIProviderPackage(options: OpenAIProviderPackageOptions): ProviderPack
36
36
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
37
37
  | `baseUrl` | `string` | Overrides `https://api.openai.com/v1`. |
38
38
  | `codexBaseUrl` | `string` | Overrides `https://chatgpt.com/backend-api/codex`. |
39
+ | `models` | `readonly ModelConfig[]` | Optional override for registered OpenAI Responses models (defaults to featured `openAIModels`). |
40
+ | `codexModels` | `readonly ModelConfig[]` | Optional override for registered Codex models (defaults to featured `openAICodexModels`). |
39
41
 
40
42
  `ProviderRequest.options.sessionId`, `cacheKey`, `cacheRetention`, `headers`,
41
- `compat`, and `extra` map to request headers/payload fields.
43
+ `compat`, and `extra` map to request headers/payload fields. Per-turn reasoning
44
+ uses official Responses `reasoning: { effort, summary? }` via
45
+ `ModelConfig.compat.reasoning` defaults merged with
46
+ `ProviderRequestOptions.compat.reasoning` (request wins). Prefer
47
+ `applyThinkingLevel(..., "openai_reasoning")` from `@arnilo/prism`.
42
48
 
43
49
  ## Outputs / response / events
44
50
 
45
51
  | Surface | Behavior |
46
52
  | --- | --- |
47
53
  | Provider stream | Prism text, thinking (downgraded to text), `tool_call` deltas/finals, `usage`, `done`, redacted `error` events. |
48
- | Block preservation | Text, thinking (downgraded), assistant `tool_call` → `function_call` input items, `tool_result` → `function_call_output` input items, images when `capabilities.input` includes `"image"`. |
54
+ | Block preservation | User/system text → `input_text`; assistant text → `output_text`; assistant `tool_call` → top-level `function_call` with `call_id`; `tool_result` → top-level `function_call_output`; images/files/audio when declared on the model. Bare thinking without an encrypted Responses reasoning item is omitted on replay. |
49
55
  | Auth methods | `api_key` for `openai`; `oauth` for `openai-codex`. |
50
56
 
51
57
  Unsupported block placements or unclaimed images fail before `fetch`.
@@ -56,10 +62,17 @@ Responses request body (Codex subscription shape, abbreviated):
56
62
 
57
63
  ```json
58
64
  {
59
- "model": "gpt-5-codex",
60
- "instructions": "You are a coding agent.",
61
- "input": [{ "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] }],
62
- "stream": true
65
+ "model": "gpt-5.1",
66
+ "input": [
67
+ { "role": "user", "content": [{ "type": "input_text", "text": "Hello" }] },
68
+ { "role": "assistant", "content": [{ "type": "output_text", "text": "Calling lookup" }] },
69
+ { "type": "function_call", "call_id": "call_1", "name": "lookup", "arguments": "{\"q\":\"x\"}" },
70
+ { "type": "function_call_output", "call_id": "call_1", "output": "{\"ok\":true}" }
71
+ ],
72
+ "reasoning": { "effort": "high" },
73
+ "prompt_cache_key": "session-1",
74
+ "stream": true,
75
+ "store": false
63
76
  }
64
77
  ```
65
78
 
@@ -73,13 +86,13 @@ https://auth.openai.com/authorize?response_type=code&client_id=...&code_challeng
73
86
 
74
87
  ```ts
75
88
  import { createExtensionKernel, createEnvCredentialResolver } from "@arnilo/prism";
76
- import { createOpenAIProviderPackage } from "@arnilo/prism-provider-openai";
89
+ import { createOpenAIProviderPackage, listOpenAIModels } from "@arnilo/prism-provider-openai";
77
90
 
91
+ const apiKey = createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" });
92
+ const models = await listOpenAIModels({ apiKey }); // caller-gated; never runs during setup
78
93
  const kernel = createExtensionKernel();
79
94
  await kernel.load([
80
- createOpenAIProviderPackage({
81
- apiKey: createEnvCredentialResolver({ OPENAI_API_KEY: "fake" }, { openai: "OPENAI_API_KEY" }),
82
- }),
95
+ createOpenAIProviderPackage({ apiKey, models }),
83
96
  ]);
84
97
  ```
85
98
 
@@ -115,25 +128,55 @@ const challenge = computeS256Challenge(verifier);
115
128
 
116
129
  ### Cache behavior
117
130
 
131
+ Official: [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching).
132
+
118
133
  - `prompt_cache_key` is derived from `ProviderRequestOptions.cacheKey` (falling
119
134
  back to `sessionId`) and sanitized + clamped to 64 characters via the shared
120
135
  `sanitizeCacheKey()` helper. Cache keys are session/customer identifiers only;
121
136
  never credentials or raw prompts.
122
- - `prompt_cache_retention` accepts only `"24h"` on the OpenAI Responses API
137
+ - `prompt_cache_retention` accepts only `"24h"` on pre-GPT-5.6 Responses models
123
138
  (extended caching). Prism `cacheRetention: "short"` and `"none"` omit the field
124
139
  so default automatic/implicit caching applies and no invalid literal is sent.
125
140
  `cacheRetention: "long"` maps to `prompt_cache_retention: "24h"` only when the
126
141
  model declares `ModelConfig.cache.longRetention === true`; models without that
127
- metadata omit the field. The catalog `gpt-5.1` model declares
142
+ metadata omit the field. Featured `gpt-5.1` declares
128
143
  `cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
144
+ - GPT-5.6+ official docs prefer `prompt_cache_options` / explicit breakpoints;
145
+ `listOpenAIModels` sets `longRetention: false` for those ids so Prism does not
146
+ emit deprecated `prompt_cache_retention` for them. Breakpoint helpers are not
147
+ shipped in this package yet — hosts may pass `prompt_cache_options` through
148
+ `compat` / `extra` when needed.
129
149
  - Cache accounting is preserved in normalized `Usage`: OpenAI
130
150
  `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
131
- Responses does not report a cache-write token field.
151
+ Responses does not report a cache-write token field on older models.
132
152
  - Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
133
153
  are applied after caller `ProviderRequestOptions.headers` so caller config
134
154
  cannot replace credentials, content type, or the session request id; non-owned
135
155
  caller headers are kept.
136
156
 
157
+ ### Model discovery
158
+
159
+ - `listOpenAIModels({ apiKey, fetch, baseUrl, signal, headers })` calls official
160
+ [`GET /models`](https://developers.openai.com/api/reference/resources/models/methods/list)
161
+ and maps sparse `{ id, created, owned_by }` entries to `ModelConfig` with
162
+ `cache.kind: "openai_key"` and heuristic `longRetention` / `capabilities.reasoning`.
163
+ - `createOpenAIProviderPackage` never calls discovery; pass results via `models:`.
164
+ - Codex subscription models are **not** listed by `api.openai.com` — keep using
165
+ featured `openAICodexModels` or `codexModels:` override.
166
+ - Static `openAIModels` / `openAICodexModels` are offline bootstrap / featured aliases only.
167
+
168
+ ### Reasoning
169
+
170
+ Official: [Reasoning models](https://developers.openai.com/api/docs/guides/reasoning).
171
+
172
+ - Body field is top-level `reasoning: { effort, summary?, mode?, context? }`.
173
+ - Model defaults: `ModelConfig.compat.reasoning`; per-turn override:
174
+ `ProviderRequestOptions.compat.reasoning` (shallow-merged; request wins).
175
+ - Portable helper: `applyThinkingLevel(options, level, "openai_reasoning")`.
176
+ - Streaming tool args follow official
177
+ `response.output_item.added` + `response.function_call_arguments.delta`
178
+ (string `delta`), not Chat Completions object deltas.
179
+
137
180
  ## Security and performance notes
138
181
 
139
182
  - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).