@arnilo/prism 0.0.5 → 0.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/CHANGELOG.md +39 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +27 -16
  4. package/dist/agent-run-lifecycle.d.ts +28 -0
  5. package/dist/agent-run-lifecycle.js +33 -0
  6. package/dist/agent-run-state.d.ts +53 -0
  7. package/dist/agent-run-state.js +127 -0
  8. package/dist/agents.d.ts +3 -1
  9. package/dist/agents.js +337 -46
  10. package/dist/contracts.d.ts +205 -3
  11. package/dist/contracts.js +4 -0
  12. package/dist/guardrails.d.ts +25 -0
  13. package/dist/guardrails.js +133 -0
  14. package/dist/ids.d.ts +2 -0
  15. package/dist/ids.js +6 -0
  16. package/dist/index.d.ts +17 -3
  17. package/dist/index.js +10 -3
  18. package/dist/input.js +2 -0
  19. package/dist/resources.js +2 -1
  20. package/dist/run-limits.d.ts +34 -0
  21. package/dist/run-limits.js +163 -0
  22. package/dist/secure-agent.d.ts +3 -0
  23. package/dist/secure-agent.js +63 -0
  24. package/dist/session-stores.js +2 -3
  25. package/dist/testing/persistence-schema.d.ts +45 -7
  26. package/dist/testing/persistence-schema.js +138 -24
  27. package/dist/thinking.d.ts +42 -0
  28. package/dist/thinking.js +92 -0
  29. package/dist/tools.d.ts +10 -2
  30. package/dist/tools.js +56 -7
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +4 -2
  34. package/docs/agent-events.md +23 -16
  35. package/docs/agent-loops.md +19 -8
  36. package/docs/agent-session-runtime.md +33 -1
  37. package/docs/coding-agent-tools.md +33 -12
  38. package/docs/coding-security.md +2 -2
  39. package/docs/compaction-llm.md +17 -7
  40. package/docs/compaction-observational-memory.md +28 -4
  41. package/docs/credential-storage.md +58 -9
  42. package/docs/credentials-and-redaction.md +1 -1
  43. package/docs/database-persistence.md +8 -3
  44. package/docs/guardrails.md +75 -0
  45. package/docs/host-security.md +16 -8
  46. package/docs/index.md +26 -22
  47. package/docs/mcp-tools.md +32 -12
  48. package/docs/migration.md +164 -2
  49. package/docs/node-filesystem-config.md +1 -0
  50. package/docs/node-jsonl-session-store.md +5 -4
  51. package/docs/postgres-persistence.md +3 -3
  52. package/docs/provider-caching.md +16 -4
  53. package/docs/provider-conformance.md +39 -1
  54. package/docs/provider-packages.md +60 -3
  55. package/docs/providers/ai-sdk.md +36 -0
  56. package/docs/providers/kimi.md +124 -61
  57. package/docs/providers/neuralwatt.md +19 -13
  58. package/docs/providers/openai.md +56 -13
  59. package/docs/providers/opencode-go.md +118 -30
  60. package/docs/providers/openrouter.md +105 -35
  61. package/docs/providers/zai.md +94 -45
  62. package/docs/release-and-install.md +47 -49
  63. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  64. package/docs/runs-and-usage.md +30 -3
  65. package/docs/server.md +5 -2
  66. package/docs/sqlite-persistence.md +2 -2
  67. package/docs/structured-output.md +1 -1
  68. package/docs/thinking-and-reasoning.md +98 -0
  69. package/docs/tool-execution-primitives.md +3 -3
  70. package/docs/tools.md +21 -1
  71. package/docs/use-case-model-selection.md +109 -0
  72. package/docs/workflow-orchestration-primitives.md +1 -0
  73. package/docs/workflows.md +18 -10
  74. package/docs/working-and-semantic-memory.md +1 -0
  75. package/package.json +2 -2
@@ -24,7 +24,7 @@ Use this package when you need server-backed persistence with pooled connections
24
24
  - managed cloud databases (RDS, Cloud SQL, Neon, Supabase, etc.)
25
25
  - CI integration tests against a real PostgreSQL service
26
26
 
27
- Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests.
27
+ Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests. This adapter stores sessions/runs, not semantic vectors; use the separate [`@arnilo/prism-memory` pgvector path](working-and-semantic-memory.md), which rejects non-finite vectors before SQL, when vector recall is needed.
28
28
 
29
29
  ## Inputs / request
30
30
 
@@ -59,7 +59,7 @@ Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backu
59
59
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
60
60
  | `close()` | Ends the pool when the adapter created it from `connectionString`. |
61
61
 
62
- Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks.
62
+ Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks. While holding that lock, startup verifies ordered contract name/version/SHA-256 rows and full schema-v3 `information_schema`/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
63
63
 
64
64
  ## Request/response example
65
65
 
@@ -127,7 +127,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
127
127
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
128
128
  - **Bounded pool.** Adapter-owned pools default to `max: 10`. Hosts with heavy concurrency should supply their own pool sizing.
129
129
  - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid sequential scans.
130
- - **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once.
130
+ - **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once. Startup catalog reads are bounded metadata queries, not application-row scans.
131
131
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns participate in query filters; hosts must still scope writes correctly.
132
132
  - **Benchmark target.** Indexed append + paginated branch read on a warm pool should stay under **50 ms p95** for local/CI-sized datasets (≤100k entries per session); measure with your pool size and hardware before production sizing.
133
133
 
@@ -21,6 +21,7 @@ Use this page when a host or provider package needs to:
21
21
  - Carry a stable cache key across turns without putting provider-specific fields in core.
22
22
  - Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
23
23
  - Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
24
+ - Understand when **caller-gated model discovery** may fill `ModelConfig.cache` / `ModelConfig.cost` from a live `/models` response (see [Discovery and live cache/cost metadata](#discovery-and-live-cache-cost-metadata)).
24
25
 
25
26
  Do not use cache keys for credentials, bearer tokens, API keys, OAuth tokens, user secrets, or raw private prompts.
26
27
 
@@ -145,21 +146,23 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
145
146
  | Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
146
147
  | --- | --- | --- | --- | --- |
147
148
  | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
148
- | `@arnilo/prism-provider-openrouter` | `cache_control` | Applies `cache_control` markers only to caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. | Breakpoint-stable prefixes can be reused by upstream providers. | Best-effort only; no marker is added to every block. |
149
+ | `@arnilo/prism-provider-openrouter` | `cache_control` | Top-level automatic `cache_control` when enabled without breakpoints; otherwise markers only on caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable / automatic prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
149
150
  | `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
150
151
  | `@arnilo/prism-provider-zai` | `implicit` | No explicit cache payload; GLM context caching is automatic. | Resend unchanged prior history for implicit context-cache reuse. | Best-effort only; cache options do not force hits. |
151
152
  | `@arnilo/prism-provider-kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
152
153
  | `@arnilo/prism-provider-neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
154
+ | `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
153
155
 
154
156
  Detailed first-party provider notes:
155
157
 
156
- - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
158
+ - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. GPT-5.6+ official docs use `prompt_cache_options` / breakpoints instead of retention — `listOpenAIModels` sets `longRetention: false` for those ids. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
157
159
  - OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
158
- - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars; applies Anthropic-style `cache_control` markers only to caller-selected `cache.breakpoints` (last content block of each selected message), not every block; `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
159
- - OpenCode Go (`@arnilo/prism-provider-opencode-go`): `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route sends none. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`.
160
+ - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing; with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
161
+ - OpenCode Go (`@arnilo/prism-provider-opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
160
162
  - Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
161
163
  - NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
162
164
  - Kimi (`@arnilo/prism-provider-kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
165
+ - AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
163
166
 
164
167
  ### NeuralWatt cache-aware limiter
165
168
 
@@ -185,6 +188,15 @@ sessions differently from one-shot chat:
185
188
  See [NeuralWatt provider](providers/neuralwatt.md) for the package-level cache,
186
189
  usage, and retry details.
187
190
 
191
+ ## Discovery and live cache/cost metadata
192
+
193
+ Caller-gated `list*Models()` helpers (see [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery)) may map official list-models fields onto `ModelConfig`:
194
+
195
+ - `cache` — when the provider documents cache kind / long-retention / breakpoint support in model metadata (otherwise keep the package's known default, e.g. NeuralWatt/Z.AI `implicit`, OpenAI `openai_key`).
196
+ - `cost` — when the provider documents per-token or per-million rates, including cache-read rates such as NeuralWatt `cached_input_per_million`.
197
+
198
+ Static featured catalogs remain offline bootstrap and must **not** invent pricing or cache capabilities the official docs do not state. Discovery is never invoked by `create*ProviderPackage()`; hosts that want live `cost`/`cache` pass the returned models into package `models:` (or register them themselves).
199
+
188
200
  ## Security and performance notes
189
201
 
190
202
  - Cache hints are best-effort and do not guarantee cache hits.
@@ -133,6 +133,44 @@ await assertProviderStreamConforms({
133
133
  });
134
134
  ```
135
135
 
136
+ ## Model discovery checklist
137
+
138
+ Every first-party package that ships (or plans) a `list*Models()` helper must keep setup network-free. Add these assertions in the package suite (pattern from NeuralWatt):
139
+
140
+ 1. **`*_provider_setup_does_not_call_model_discovery`** — inject a counting `fetch` into `create*ProviderPackage({ fetch })`, run `setup`, assert `calls === 0`.
141
+ 2. **`list_*_models_maps_fixture_…`** — fixture response maps to `ModelConfig` (`id` → `model`, documented capabilities/limits/cost/cache); no credentials in returned objects.
142
+ 3. **`list_*_models_forwards_auth_abort_baseurl`** (as applicable) — Authorization owned by helper when key present; auth omitted when optional and unset; `signal` / `baseUrl` forwarded.
143
+ 4. **`list_*_models_redacts_token_in_errors`** — non-OK bodies use `readBoundedResponseText` + `redactSecrets`; secret canaries absent from thrown messages.
144
+ 5. **Malformed payload rejects** — missing `data` array (or provider-equivalent) throws a clear discovery error.
145
+
146
+ OpenRouter stays app-registration-first: an optional list helper must still not run during setup. AI SDK has no discovery export. Packages without a public list API document curated official-doc refresh instead of inventing a fake helper.
147
+
148
+ Canonical contract: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery).
149
+
150
+ ## Thinking / reasoning checklist
151
+
152
+ Every first-party package that exposes thinking or reasoning controls should cover:
153
+
154
+ 1. **Model default** — `ModelConfig.compat` (or documented capability) sets the official wire field when no per-turn override is present.
155
+ 2. **Per-turn override wins** — `ProviderRequestOptions.compat` via `mergeProviderRequestOptions` / `applyThinkingLevel` overrides the model default.
156
+ 3. **Shared family mapping** — effort levels from `ThinkingLevel` land in the package's recommended family (`openai_reasoning` / `reasoning_effort` / `thinking_type` / `noop`) per [Thinking and reasoning](thinking-and-reasoning.md).
157
+ 4. **Non-reasoning / noop** — applying a level with `noop` (or omitting compat) must not invent unsupported body fields.
158
+ 5. **No inert `extra.thinkingLevel`** — package code must not rely on `options.extra.thinkingLevel` for wire mapping.
159
+
160
+ Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md).
161
+
162
+ ## AI SDK adapter checklist
163
+
164
+ `@arnilo/prism-provider-ai-sdk` is a host-owned `LanguageModelV4` bridge. It does not participate in the discovery or thinking/reasoning checklists above. Cover instead:
165
+
166
+ 1. **No catalog / no setup fetch** — package exports no `list*Models()`; `createAiSdkProvider` wraps a host model only.
167
+ 2. **Specification gate** — rejects non-v4 models (`specificationVersion !== "v4"` or missing `doStream`).
168
+ 3. **Cache usage mapping** — `finish.usage.inputTokens.cacheRead`/`cacheWrite` map to `Usage.cacheReadTokens`/`cacheWriteTokens`; adapter does not emit cache request fields.
169
+ 4. **Reasoning stream mapping** — `reasoning-delta` → thinking deltas; assistant `thinking` blocks replay as AI SDK `reasoning` prompt parts.
170
+ 5. **Host-owned controls** — `options.compat` / `options.extra` forward as `providerOptions.prism`; reasoning effort stays on the host model.
171
+
172
+ Canonical contract: [AI SDK provider adapter](providers/ai-sdk.md).
173
+
136
174
  ## Extension and configuration notes
137
175
 
138
176
  The helpers are a testing subpath only. Provider packages can use them with their own mocked fetch/transport or `createMockProvider()`. Live provider tests should stay opt-in and env-gated outside Prism's default test suite.
@@ -148,7 +186,7 @@ The helpers are a testing subpath only. Provider packages can use them with thei
148
186
  ## Related APIs
149
187
 
150
188
  - [Provider layer](provider-layer.md): `AIProvider`, provider events, and mock provider.
151
- - [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters.
189
+ - [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters; includes the caller-gated discovery contract and setup zero-fetch rule.
152
190
  - [AI SDK provider adapter](providers/ai-sdk.md): optional `LanguageModelV4` bridge tested with a fake AI SDK model.
153
191
  - [OpenAI-compatible provider](providers/openai-compatible.md): optional provider adapter tested with mocked streams.
154
192
  - [Public contracts](public-contracts.md): provider request/event/usage contracts.
@@ -67,7 +67,7 @@ Phase 6 also adds optional [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md
67
67
 
68
68
  Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
69
69
 
70
- These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers static OpenCode Go metadata and package-local OpenAI/Anthropic-compatible routes from caller-supplied credentials only. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/reasoning/cache passthrough and no setup catalog fetch. `@arnilo/prism-provider-zai` now registers static GLM metadata with Z.AI thinking/reasoning/tool-stream request mapping. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default and optional Moonshot metadata only when requested. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
70
+ These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only, with optional `models`/`codexModels` overrides and an opt-in `listOpenAIModels()` helper for official `GET /models` discovery. `@arnilo/prism-provider-opencode-go` now registers docs-verified OpenCode Go open coding models with dual OpenAI/Anthropic routes (`compat.route`), official default base `https://opencode.ai/zen/go/v1`, `reasoning_content`/thinking preserve, and an opt-in `listOpenCodeGoModels()` helper for official `GET /zen/go/v1/models`. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/`reasoning`/cache passthrough, assistant `reasoning` replay, optional top-level automatic `cache_control`, and an opt-in `listOpenRouterModels()` helper for official `GET /api/v1/models` (setup still never fetches). `@arnilo/prism-provider-zai` now registers featured GLM-5.x/4.x metadata with official `thinking`/`reasoning_effort`/`tool_stream`/`clear_thinking` mapping, Preserved Thinking `reasoning_content` replay, implicit context caching, and an opt-in `listZaiModels()` helper for OpenAI-compatible `GET /models`. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default, optional callable Moonshot Open Platform Chat Completions when `includeMoonshotModels` is requested, official Coding/Open Platform featured ids, thinking/`reasoning_effort` compat mapping, and an opt-in `listKimiModels()` helper for Moonshot `GET /v1/models`. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
71
71
 
72
72
  ### First-party cache behavior
73
73
 
@@ -75,14 +75,71 @@ Every first-party provider package hardens prompt-cache behavior so it cannot em
75
75
 
76
76
  - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
77
77
  - **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
78
- - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; `cache_control` markers applied only to caller-selected `cache.breakpoints` (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
79
- - **OpenCode Go**: `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; the Anthropic route applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`), the OpenAI route sends none. Per-route usage mapping.
78
+ - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
79
+ - **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
80
80
  - **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
81
81
  - **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
82
82
  - **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
83
83
 
84
84
  See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
85
85
 
86
+ ## Caller-gated model discovery
87
+
88
+ First-party packages keep `create*ProviderPackage()` network-free. Latest models come from **caller-gated** `list*Models()` helpers that hosts invoke explicitly and then pass back via `models:` (or register themselves). Plan 015's "no setup catalog fetch" rule still holds; Plan 067 adds on-demand discovery without hidden latency.
89
+
90
+ ### Contract
91
+
92
+ ```ts
93
+ export async function listExampleModels(options: {
94
+ apiKey?: CredentialValueSource;
95
+ fetch?: typeof fetch;
96
+ baseUrl?: string;
97
+ signal?: AbortSignal;
98
+ headers?: Readonly<Record<string, string>>;
99
+ }): Promise<ModelConfig[]> {
100
+ // GET {baseUrl}/models — never called from create*ProviderPackage()
101
+ }
102
+ ```
103
+
104
+ | Rule | Requirement |
105
+ | --- | --- |
106
+ | Setup | `create*ProviderPackage().setup` performs **zero** fetches / discovery calls |
107
+ | Shape | Package-local `list*Models(options) → Promise<ModelConfig[]>` + optional `map*Model(entry)` |
108
+ | Injectables | `fetch`, `baseUrl`, `signal`, optional `apiKey` / `headers` |
109
+ | Transport | Error bodies via `@arnilo/prism/providers/transport` `readBoundedResponseText`; credentials via `resolveCredentialValue` + `redactSecrets` |
110
+ | Return | `ModelConfig[]` only — never embed API keys, tokens, or auth headers in returned metadata |
111
+ | Static catalog | Featured aliases / offline bootstrap only; may omit live pricing until discovery fills `cost` / `cache` |
112
+ | Core | Prefer package-local helpers. Do **not** add a core model-discovery registry. Extract a shared HTTP/list helper only when ≥2 packages share identical parsing |
113
+
114
+ Template: [`listNeuralWattModels`](providers/neuralwatt.md) in `@arnilo/prism-provider-neuralwatt`.
115
+
116
+ ### Per-package policy
117
+
118
+ | Package | Discovery helper | Setup catalog | Notes |
119
+ | --- | --- | --- | --- |
120
+ | OpenAI | **`listOpenAIModels` (exists)** | Featured Responses/Codex aliases; factory accepts `models?` / `codexModels?` | Official `GET /v1/models`; Codex not listed by api.openai.com |
121
+ | Kimi | **`listKimiModels`** (Moonshot `GET /v1/models`) | Featured Coding ids + optional callable Moonshot | Official Moonshot/Kimi list-models; Coding curated |
122
+ | Z.AI | **`listZaiModels`** (OpenAI-compatible `GET /models`) + curated featured refresh | Featured GLM-5.2…4.5 aliases | No first-class docs.z.ai list page; discovery is best-effort; featured set from Chat Completions enum / overview |
123
+ | OpenRouter | **`listOpenRouterModels`** (official `GET /api/v1/models`) | **App-controlled** `models:` only — no bundled mega-catalog | Helper feeds host registration; setup still does not fetch |
124
+ | OpenCode Go | **`listOpenCodeGoModels`** (official `GET /zen/go/v1/models`) | Featured dual-route official Go aliases | Official Go docs endpoint table + sparse list API |
125
+ | NeuralWatt | **`listNeuralWattModels` (exists)** | Featured aliases without guessed pricing | Auth optional for public models |
126
+ | AI SDK | None | Host-owned `LanguageModelV4` | No Prism-side catalog by design |
127
+
128
+ Host pattern:
129
+
130
+ ```ts
131
+ const models = await listNeuralWattModels({ apiKey, fetch });
132
+ await kernel.load([createNeuralWattProviderPackage({ apiKey, models })]);
133
+ ```
134
+
135
+ Discovery may populate `ModelConfig.cache` and `ModelConfig.cost` from live metadata when the provider documents those fields; see [Provider caching](provider-caching.md#discovery-and-live-cache-cost-metadata). Package authors: include the [setup zero-fetch checklist](provider-conformance.md#model-discovery-checklist) in every first-party suite that ships or plans a `list*Models` helper.
136
+
137
+ ## Per-turn thinking / reasoning
138
+
139
+ Hosts set effort with portable helpers from `@arnilo/prism` (`applyThinkingLevel`, `thinkingCompatFor`) that write official fields into `ProviderRequestOptions.compat`. Model defaults stay on `ModelConfig.compat`; per-turn patches win via `mergeProviderRequestOptions`. Providers keep reading `options.compat` / `model.compat` — do not invent a parallel options tree or put effort only in `extra`.
140
+
141
+ Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md). Package-local knobs (NeuralWatt budgets, Z.AI `tool_stream`, Kimi keep/all) remain on `compat` beside the shared families.
142
+
86
143
  ## Third-party provider packaging
87
144
 
88
145
  A third party ships their own providers the same way Prism ships first-party
@@ -52,6 +52,8 @@ Unsupported content fails before `doStream` (for example unresolved `resourceUri
52
52
  | `finish` usage | `usage` then `done` |
53
53
  | `error` / thrown / abort | redacted `error` |
54
54
 
55
+ `finish.usage.inputTokens.cacheRead` / `cacheWrite` map to Prism `Usage.cacheReadTokens` / `cacheWriteTokens`. The adapter does not invent cache request fields; prompt caching is owned by the host `LanguageModelV4` and its upstream provider.
56
+
55
57
  Provider-executed tool calls, files/sources/custom parts, warnings, and raw chunks are ignored rather than silently converted into unsupported Prism content.
56
58
 
57
59
  ## Request/response example
@@ -89,6 +91,40 @@ const result = await agent.createSession().run("Summarize this");
89
91
  console.log(result.text);
90
92
  ```
91
93
 
94
+ ## Model catalog and discovery
95
+
96
+ There is **no Prism-side model catalog** and **no `list*Models()` export** by design. Hosts supply a ready-made `LanguageModelV4` instance (typically from `@ai-sdk/openai`, `@ai-sdk/anthropic`, AI Gateway, or a custom provider) and register a matching `ModelConfig` for capabilities/limits.
97
+
98
+ Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials.
99
+
100
+ ## Prompt caching
101
+
102
+ The adapter is **host-owned for request caching**. It does not emit `cache_control`, `prompt_cache_key`, `cacheKey`, or `cacheRetention` on AI SDK call options. Hosts configure caching on the underlying AI SDK model/provider (for example via AI SDK `providerOptions` on the model factory or per-call options forwarded through `options.compat` / `options.extra` → `providerOptions.prism`).
103
+
104
+ When the host model reports cache accounting on the `finish` stream part, Prism maps official AI SDK v4 usage fields:
105
+
106
+ | AI SDK `LanguageModelV4Usage` | Prism `Usage` |
107
+ | --- | --- |
108
+ | `inputTokens.cacheRead` | `cacheReadTokens` |
109
+ | `inputTokens.cacheWrite` | `cacheWriteTokens` |
110
+ | `inputTokens.total` | `inputTokens` |
111
+ | `outputTokens.total` | `outputTokens` |
112
+
113
+ See [Provider caching](../provider-caching.md) for the cross-provider matrix.
114
+
115
+ ## Thinking and reasoning
116
+
117
+ Reasoning effort, budgets, and provider-specific thinking controls are **host-model-owned**. Prism does not map `ThinkingLevel` into AI SDK call options (`thinkingFamilyForModel` → `noop`). Hosts configure reasoning on the AI SDK model (for example OpenAI `reasoning.effort` via AI SDK `providerOptions`) and may pass per-turn overrides through `ProviderRequestOptions.compat` / `extra`, which the adapter forwards as `providerOptions.prism`.
118
+
119
+ Stream mapping:
120
+
121
+ | Direction | Mapping |
122
+ | --- | --- |
123
+ | AI SDK `reasoning-delta` → Prism | `content_delta` thinking |
124
+ | Prism `thinking` blocks → AI SDK prompt | `{ type: "reasoning", text }` on assistant messages |
125
+
126
+ Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); [Language Model Specification V4](https://github.com/vercel/ai/tree/main/packages/provider/src/language-model/v4); `@ai-sdk/provider` `LanguageModelV4Usage` (`inputTokens.cacheRead` / `cacheWrite`).
127
+
92
128
  ## Extension and configuration notes
93
129
 
94
130
  - Peer dependency: `@ai-sdk/provider@^4.0.0`. Upgrade policy tracks one specification major at a time.
@@ -2,132 +2,195 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-provider-kimi` provides explicit, side-effect-free setup for Kimi For
6
- Coding using an Anthropic-compatible `/messages` endpoint with
7
- `User-Agent: KimiCLI/1.5` (unless overridden). Moonshot/Open Platform model
8
- metadata is optional.
5
+ `@arnilo/prism-provider-kimi` provides two distinct, side-effect-free routes:
9
6
 
10
- The package registers the `kimi-coding` provider, default Kimi Coding model
11
- metadata, and an `api_key` auth method through `createExtensionKernel().load([...])`.
7
+ 1. **Kimi For Coding** (default) — Anthropic-compatible `POST /messages` on
8
+ `https://api.kimi.com/coding` with `User-Agent: KimiCLI/1.5` (unless overridden).
9
+ 2. **Moonshot Open Platform** (opt-in) — OpenAI-compatible `POST /chat/completions`
10
+ on `https://api.moonshot.ai/v1` (or `api.moonshot.cn/v1`), registered only when
11
+ `includeMoonshotModels: true`.
12
+
13
+ Official model ids differ by route. Coding uses `kimi-for-coding`,
14
+ `kimi-for-coding-highspeed`, and `k3`. Open Platform uses `kimi-k2.7-code`,
15
+ `kimi-k3`, and related catalog ids. Pi's `k2p7` alias is **not** used.
16
+
17
+ Caller-gated discovery via `listKimiModels()` hits the official Moonshot
18
+ `GET /v1/models` endpoint. Package setup never fetches.
12
19
 
13
20
  ## When to use it
14
21
 
15
- Use it when a host app wants the Kimi For Coding endpoint through Prism's
16
- `AgentSession` runtime with Kimi-specific serializer behavior.
22
+ Use it when a host app wants Kimi For Coding and/or Moonshot Open Platform through
23
+ Prism's `AgentSession` runtime with Kimi-specific serializers, thinking controls,
24
+ and cache policy.
17
25
 
18
- Do not use it for Moonshot Open Platform default registration, automatic
19
- credential discovery, catalog fetches, or real-network tests.
26
+ Do not use it for automatic credential discovery, setup-time catalog fetches, or
27
+ real-network tests (live tests stay opt-in).
20
28
 
21
29
  ## Inputs / request
22
30
 
23
31
  ```ts
24
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
32
+ import {
33
+ createKimiProviderPackage,
34
+ listKimiModels,
35
+ defineKimiModel,
36
+ } from "@arnilo/prism-provider-kimi";
25
37
 
26
38
  createKimiProviderPackage(options: KimiProviderPackageOptions): ProviderPackage
27
- defineKimiModel(config: KimiModelConfig): KimiModelConfig
39
+ listKimiModels(options?: ListKimiModelsOptions): Promise<ModelConfig[]>
40
+ defineKimiModel(config: KimiModelConfig): ModelConfig
28
41
  ```
29
42
 
30
43
  | Field | Type | Purpose |
31
44
  | --- | --- | --- |
32
- | `kimiApiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source for Kimi. |
45
+ | `kimiApiKey` | `CredentialValueSource` | Kimi For Coding API key. |
46
+ | `moonshotApiKey` | `CredentialValueSource` | Moonshot Open Platform API key (not interchangeable with Coding keys). |
33
47
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
34
- | `baseUrl` | `string` | Overrides the Kimi base URL. |
35
- | `id` | `string` | Overrides the provider id (default `kimi-coding`). |
36
- | `userAgent` | `string` | Overrides `User-Agent: KimiCLI/1.5`. |
37
- | `models` | `readonly ModelConfig[]` | Overrides `kimiCodingModels` defaults. |
38
- | `includeMoonshotModels` | `boolean` | Registers Moonshot models when `true` (default off). |
39
- | `moonshotModels` | `readonly ModelConfig[]` | Overrides `moonshotKimiModels` when included. |
48
+ | `baseUrl` | `string` | Overrides the Coding base URL. |
49
+ | `moonshotBaseUrl` | `string` | Overrides Moonshot base URL (default `https://api.moonshot.ai/v1`). |
50
+ | `id` / `moonshotId` | `string` | Provider ids (defaults `kimi-coding` / `moonshot`). |
51
+ | `userAgent` | `string` | Overrides Coding `User-Agent: KimiCLI/1.5`. |
52
+ | `models` | `readonly ModelConfig[]` | Overrides featured Coding models. |
53
+ | `includeMoonshotModels` | `boolean` | Registers callable Moonshot provider + models when `true`. |
54
+ | `moonshotModels` | `readonly ModelConfig[]` | Overrides featured Moonshot models when included. |
40
55
 
41
56
  ## Outputs / response / events
42
57
 
43
58
  | Surface | Behavior |
44
59
  | --- | --- |
45
- | Provider stream | Prism text, thinking (preserved only when `model.compat.preserveThinking` is true, otherwise downgraded to text), tool-call delta/final, `usage`, `done`, redacted `error`. |
46
- | Block preservation | Text, thinking, assistant `tool_call` → `tool_use`, `tool_result` → `tool_result`, images when `capabilities.input` includes `"image"`. |
47
- | Auth method | `api_key` for `kimi-coding`, credential name `apiKey`. |
60
+ | Coding stream | Prism text, thinking deltas, tool-call delta/final, `usage` (`cache_read_input_tokens` / `cache_creation_input_tokens`), `done`, redacted `error`. |
61
+ | Moonshot stream | Same, with `delta.reasoning_content` → thinking; OpenAI-style usage cache details when present. |
62
+ | Block preservation | Coding: Anthropic `thinking` / `tool_use` / `tool_result`. Moonshot: `reasoning_content` on assistant replay when `preserveThinking`. |
63
+ | Auth methods | `api_key` for `kimi-coding`; also `moonshot` when opted in. |
48
64
 
49
65
  Unsupported block placements or unclaimed images fail before fetch.
50
66
 
67
+ ## Route differences
68
+
69
+ | | Kimi For Coding | Moonshot Open Platform |
70
+ | --- | --- | --- |
71
+ | Base URL | `https://api.kimi.com/coding` | `https://api.moonshot.ai/v1` (or `.cn`) |
72
+ | Wire API | Anthropic `/messages` | OpenAI `/chat/completions` |
73
+ | Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k3` (+ discovery) |
74
+ | Discovery | No public list API — curated featured aliases | Official `GET /v1/models` via `listKimiModels()` |
75
+ | Cache | Implicit by default; opt-in Anthropic `cache_control` | Implicit only — never emits Anthropic `cache_control` |
76
+ | Thinking | Block replay + body `thinking` / `reasoning_effort` | `reasoning_content` replay + body `thinking` / `reasoning_effort` |
77
+
78
+ The Anthropic `/messages` request/response contract for Kimi remains under-documented
79
+ upstream ([MoonshotAI/Kimi-K2#129](https://github.com/MoonshotAI/Kimi-K2/issues/129));
80
+ Prism treats official Chat Completions thinking fields as best-effort passthrough on
81
+ the Coding route.
82
+
83
+ ## Thinking / reasoning
84
+
85
+ Official fields (Open Platform docs; Coding docs for `k3` effort mapping):
86
+
87
+ | Model family | Official control | Prism `compat` |
88
+ | --- | --- | --- |
89
+ | K3 / Coding `k3` | top-level `reasoning_effort` (`max` on Open Platform; Coding also `low`/`high`) | `compat.reasoning_effort` — use Task 4 family `reasoning_effort` |
90
+ | K2.7-code / Coding | thinking always on; Preserved Thinking always on | omit `thinking` by default; `preserveThinking: true` for replay; do not send `disabled` |
91
+ | K2.6 / K2.5 | `thinking.type` enabled/disabled; K2.6 optional `keep: "all"` | `compat.thinking` — Task 4 family `thinking_type` |
92
+
93
+ Per-turn `ProviderRequestOptions.compat` wins over `ModelConfig.compat`. Helpers:
94
+ `kimiThinking`, `kimiReasoningEffort`, `kimiPreserveThinking`.
95
+
51
96
  ## Request/response example
52
97
 
53
- Example request (Anthropic-compatible `/messages` shape):
98
+ Coding (Anthropic-compatible `/messages`):
54
99
 
55
100
  ```json
56
101
  {
57
- "model": "kimi-latest",
58
- "messages": [{ "role": "user", "content": "Hello" }],
102
+ "model": "kimi-for-coding",
103
+ "messages": [{ "role": "user", "content": [{ "type": "text", "text": "Hello" }] }],
59
104
  "stream": true
60
105
  }
61
106
  ```
62
107
 
108
+ Moonshot (Chat Completions):
109
+
110
+ ```json
111
+ {
112
+ "model": "kimi-k3",
113
+ "messages": [{ "role": "user", "content": "Hello" }],
114
+ "stream": true,
115
+ "reasoning_effort": "max"
116
+ }
117
+ ```
118
+
63
119
  ## Implementation example
64
120
 
65
121
  ```ts
66
122
  import { createExtensionKernel } from "@arnilo/prism";
67
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
123
+ import {
124
+ createKimiProviderPackage,
125
+ listKimiModels,
126
+ } from "@arnilo/prism-provider-kimi";
68
127
 
69
128
  const kernel = createExtensionKernel();
70
129
  await kernel.load([
71
- createKimiProviderPackage({ kimiApiKey: "fake-kimi-key", includeMoonshotModels: false }),
130
+ createKimiProviderPackage({ kimiApiKey: "fake-kimi-key" }),
72
131
  ]);
73
- ```
74
-
75
- Register Moonshot/Open Platform metadata explicitly:
76
132
 
77
- ```ts
78
- import { createKimiProviderPackage } from "@arnilo/prism-provider-kimi";
133
+ // Opt-in Moonshot Open Platform (callable provider + featured models)
134
+ await kernel.load([
135
+ createKimiProviderPackage({
136
+ kimiApiKey: "fake-kimi-key",
137
+ includeMoonshotModels: true,
138
+ moonshotApiKey: "fake-moonshot-key",
139
+ }),
140
+ ]);
79
141
 
142
+ // Caller-gated discovery — never runs during setup
143
+ const latest = await listKimiModels({ apiKey: "fake-moonshot-key", fetch });
80
144
  await kernel.load([
81
- createKimiProviderPackage({ kimiApiKey: "fake", includeMoonshotModels: true }),
145
+ createKimiProviderPackage({
146
+ includeMoonshotModels: true,
147
+ moonshotApiKey: "fake-moonshot-key",
148
+ moonshotModels: latest.filter((m) => m.model.startsWith("kimi-")),
149
+ }),
82
150
  ]);
83
151
  ```
84
152
 
85
153
  ## Extension and configuration notes
86
154
 
87
- - Hosts choose base URL, provider id, `User-Agent`, model list, credential source,
155
+ - Hosts choose base URLs, provider ids, `User-Agent`, model lists, credential sources,
88
156
  and `fetch` impl.
89
- - Moonshot/Open Platform metadata is registered only with
90
- `includeMoonshotModels: true`; it is not core behavior.
91
- - Package contributes models via the extension `api` and an `api_key` auth method.
157
+ - Moonshot is registered only with `includeMoonshotModels: true` (provider + models + auth).
158
+ - Featured catalogs are offline bootstrap only; refresh Open Platform via `listKimiModels()`.
92
159
 
93
160
  ### Cache behavior
94
161
 
95
- - Default catalog models (e.g. `kimi-k2.7-code` on the Anthropic-compatible
96
- `/messages` route) use **implicit caching** and send no explicit `cache_control`
97
- fields. `ProviderRequestOptions.cache` / `cacheKey` / `cacheRetention` have no
98
- effect on the request body unless the model opts in.
99
- - Hosts may opt a model into Anthropic-style `cache_control` by declaring
100
- `ModelConfig.cache.kind: "cache_control"` on the Anthropic route. When opted in,
101
- `cache_control: { type: "ephemeral" }` markers are applied only to the
102
- caller-selected `ProviderRequestOptions.cache.breakpoints` (resolved with the
103
- shared `applyCacheControl()` helper) on the last content block of each selected
104
- message — not to every block. `cacheRetention: "long"` adds `ttl: "1h"` when the
105
- model allows long retention (`ModelConfig.cache.longRetention !== false`).
106
- - The Moonshot Open Platform route (`compat.route: "openai"`) never receives
107
- Anthropic `cache_control` fields.
108
- - Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
109
- `Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
110
- `Usage.cacheWriteTokens`.
162
+ - Default Coding catalog models use **implicit caching** and send no explicit
163
+ `cache_control` fields unless the model opts in with
164
+ `ModelConfig.cache.kind: "cache_control"`.
165
+ - When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
166
+ caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
167
+ block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
168
+ the model allows long retention.
169
+ - The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
170
+ - Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
171
+ `cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
111
172
 
112
173
  ## Security and performance notes
113
174
 
114
- - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
175
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
176
+ helpers (`readSseData`, `readBoundedResponseText`).
115
177
  - No network calls during import, setup, build, or default tests.
116
178
  - No automatic environment, file, keychain, or shell credential lookup.
117
- - Kimi credentials are resolved per request from caller-supplied values or resolvers
118
- and redacted from errors.
179
+ - Credentials are resolved per request from caller-supplied values or resolvers
180
+ and redacted from errors (including discovery failures).
119
181
  - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
120
182
  but provider-owned headers (`content-type`, `user-agent`, `authorization`)
121
183
  are applied last and cannot be overridden by caller headers.
122
- - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
123
- provider-specific env names; default tests are network-free.
184
+ - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus provider-specific
185
+ env names; default tests are network-free.
124
186
 
125
187
  ## Related APIs
126
188
 
127
189
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
128
- `ModelConfig`/`compat`, Anthropic-compatible routes.
190
+ caller-gated discovery, Anthropic/OpenAI routes.
191
+ - [Thinking and reasoning](../thinking-and-reasoning.md): portable `ThinkingLevel`
192
+ helpers and Kimi family mapping.
129
193
  - [Credentials and redaction](../credentials-and-redaction.md):
130
194
  `resolveCredentialValue`, `redactSecrets`.
131
- - [Provider layer](../provider-layer.md): `ProviderRequest.options` and usage
132
- mapping.
195
+ - [Provider caching](../provider-caching.md): explicit/implicit matrix.
133
196
  - [Provider conformance](../provider-conformance.md): network-free adapter tests.