@arnilo/prism 0.0.5 → 0.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -1
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +27 -16
- package/dist/agent-run-lifecycle.d.ts +28 -0
- package/dist/agent-run-lifecycle.js +33 -0
- package/dist/agent-run-state.d.ts +53 -0
- package/dist/agent-run-state.js +127 -0
- package/dist/agents.d.ts +3 -1
- package/dist/agents.js +337 -46
- package/dist/contracts.d.ts +205 -3
- package/dist/contracts.js +4 -0
- package/dist/guardrails.d.ts +25 -0
- package/dist/guardrails.js +133 -0
- package/dist/ids.d.ts +2 -0
- package/dist/ids.js +6 -0
- package/dist/index.d.ts +17 -3
- package/dist/index.js +10 -3
- package/dist/input.js +2 -0
- package/dist/resources.js +2 -1
- package/dist/run-limits.d.ts +34 -0
- package/dist/run-limits.js +163 -0
- package/dist/secure-agent.d.ts +3 -0
- package/dist/secure-agent.js +63 -0
- package/dist/session-stores.js +2 -3
- package/dist/testing/persistence-schema.d.ts +45 -7
- package/dist/testing/persistence-schema.js +138 -24
- package/dist/thinking.d.ts +42 -0
- package/dist/thinking.js +92 -0
- package/dist/tools.d.ts +10 -2
- package/dist/tools.js +56 -7
- package/dist/use-case-model.d.ts +63 -0
- package/dist/use-case-model.js +52 -0
- package/docs/a2a.md +4 -2
- package/docs/agent-events.md +23 -16
- package/docs/agent-loops.md +19 -8
- package/docs/agent-session-runtime.md +33 -1
- package/docs/coding-agent-tools.md +33 -12
- package/docs/coding-security.md +2 -2
- package/docs/compaction-llm.md +17 -7
- package/docs/compaction-observational-memory.md +28 -4
- package/docs/credential-storage.md +58 -9
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/database-persistence.md +8 -3
- package/docs/guardrails.md +75 -0
- package/docs/host-security.md +16 -8
- package/docs/index.md +26 -22
- package/docs/mcp-tools.md +32 -12
- package/docs/migration.md +164 -2
- package/docs/node-filesystem-config.md +1 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/postgres-persistence.md +3 -3
- package/docs/provider-caching.md +16 -4
- package/docs/provider-conformance.md +39 -1
- package/docs/provider-packages.md +60 -3
- package/docs/providers/ai-sdk.md +36 -0
- package/docs/providers/kimi.md +124 -61
- package/docs/providers/neuralwatt.md +19 -13
- package/docs/providers/openai.md +56 -13
- package/docs/providers/opencode-go.md +118 -30
- package/docs/providers/openrouter.md +105 -35
- package/docs/providers/zai.md +94 -45
- package/docs/release-and-install.md +47 -49
- package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
- package/docs/runs-and-usage.md +30 -3
- package/docs/server.md +5 -2
- package/docs/sqlite-persistence.md +2 -2
- package/docs/structured-output.md +1 -1
- package/docs/thinking-and-reasoning.md +98 -0
- package/docs/tool-execution-primitives.md +3 -3
- package/docs/tools.md +21 -1
- package/docs/use-case-model-selection.md +109 -0
- package/docs/workflow-orchestration-primitives.md +1 -0
- package/docs/workflows.md +18 -10
- package/docs/working-and-semantic-memory.md +1 -0
- package/package.json +2 -2
|
@@ -24,7 +24,7 @@ Use this package when you need server-backed persistence with pooled connections
|
|
|
24
24
|
- managed cloud databases (RDS, Cloud SQL, Neon, Supabase, etc.)
|
|
25
25
|
- CI integration tests against a real PostgreSQL service
|
|
26
26
|
|
|
27
|
-
Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests.
|
|
27
|
+
Prefer [`@arnilo/prism-session-store-sqlite`](sqlite-persistence.md) for local CLI tools, single-writer workloads, and network-free default tests. This adapter stores sessions/runs, not semantic vectors; use the separate [`@arnilo/prism-memory` pgvector path](working-and-semantic-memory.md), which rejects non-finite vectors before SQL, when vector recall is needed.
|
|
28
28
|
|
|
29
29
|
## Inputs / request
|
|
30
30
|
|
|
@@ -59,7 +59,7 @@ Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backu
|
|
|
59
59
|
| `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
|
|
60
60
|
| `close()` | Ends the pool when the adapter created it from `connectionString`. |
|
|
61
61
|
|
|
62
|
-
Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks.
|
|
62
|
+
Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks. While holding that lock, startup verifies ordered contract name/version/SHA-256 rows and full schema-v3 `information_schema`/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
|
|
63
63
|
|
|
64
64
|
## Request/response example
|
|
65
65
|
|
|
@@ -127,7 +127,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
|
|
|
127
127
|
- **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
|
|
128
128
|
- **Bounded pool.** Adapter-owned pools default to `max: 10`. Hosts with heavy concurrency should supply their own pool sizing.
|
|
129
129
|
- **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid sequential scans.
|
|
130
|
-
- **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once.
|
|
130
|
+
- **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once. Startup catalog reads are bounded metadata queries, not application-row scans.
|
|
131
131
|
- **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns participate in query filters; hosts must still scope writes correctly.
|
|
132
132
|
- **Benchmark target.** Indexed append + paginated branch read on a warm pool should stay under **50 ms p95** for local/CI-sized datasets (≤100k entries per session); measure with your pool size and hardware before production sizing.
|
|
133
133
|
|
package/docs/provider-caching.md
CHANGED
|
@@ -21,6 +21,7 @@ Use this page when a host or provider package needs to:
|
|
|
21
21
|
- Carry a stable cache key across turns without putting provider-specific fields in core.
|
|
22
22
|
- Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
|
|
23
23
|
- Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
|
|
24
|
+
- Understand when **caller-gated model discovery** may fill `ModelConfig.cache` / `ModelConfig.cost` from a live `/models` response (see [Discovery and live cache/cost metadata](#discovery-and-live-cache-cost-metadata)).
|
|
24
25
|
|
|
25
26
|
Do not use cache keys for credentials, bearer tokens, API keys, OAuth tokens, user secrets, or raw private prompts.
|
|
26
27
|
|
|
@@ -145,21 +146,23 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
|
|
|
145
146
|
| Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
|
|
146
147
|
| --- | --- | --- | --- | --- |
|
|
147
148
|
| `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
|
|
148
|
-
| `@arnilo/prism-provider-openrouter` | `cache_control` |
|
|
149
|
+
| `@arnilo/prism-provider-openrouter` | `cache_control` | Top-level automatic `cache_control` when enabled without breakpoints; otherwise markers only on caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable / automatic prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
|
|
149
150
|
| `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
|
|
150
151
|
| `@arnilo/prism-provider-zai` | `implicit` | No explicit cache payload; GLM context caching is automatic. | Resend unchanged prior history for implicit context-cache reuse. | Best-effort only; cache options do not force hits. |
|
|
151
152
|
| `@arnilo/prism-provider-kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
|
|
152
153
|
| `@arnilo/prism-provider-neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
|
|
154
|
+
| `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
|
|
153
155
|
|
|
154
156
|
Detailed first-party provider notes:
|
|
155
157
|
|
|
156
|
-
- OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
158
|
+
- OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. GPT-5.6+ official docs use `prompt_cache_options` / breakpoints instead of retention — `listOpenAIModels` sets `longRetention: false` for those ids. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
157
159
|
- OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
|
|
158
|
-
- OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars; applies Anthropic-style
|
|
159
|
-
- OpenCode Go (`@arnilo/prism-provider-opencode-go`): `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route sends none
|
|
160
|
+
- OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing; with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
|
|
161
|
+
- OpenCode Go (`@arnilo/prism-provider-opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
|
|
160
162
|
- Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
161
163
|
- NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
|
|
162
164
|
- Kimi (`@arnilo/prism-provider-kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
165
|
+
- AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
|
|
163
166
|
|
|
164
167
|
### NeuralWatt cache-aware limiter
|
|
165
168
|
|
|
@@ -185,6 +188,15 @@ sessions differently from one-shot chat:
|
|
|
185
188
|
See [NeuralWatt provider](providers/neuralwatt.md) for the package-level cache,
|
|
186
189
|
usage, and retry details.
|
|
187
190
|
|
|
191
|
+
## Discovery and live cache/cost metadata
|
|
192
|
+
|
|
193
|
+
Caller-gated `list*Models()` helpers (see [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery)) may map official list-models fields onto `ModelConfig`:
|
|
194
|
+
|
|
195
|
+
- `cache` — when the provider documents cache kind / long-retention / breakpoint support in model metadata (otherwise keep the package's known default, e.g. NeuralWatt/Z.AI `implicit`, OpenAI `openai_key`).
|
|
196
|
+
- `cost` — when the provider documents per-token or per-million rates, including cache-read rates such as NeuralWatt `cached_input_per_million`.
|
|
197
|
+
|
|
198
|
+
Static featured catalogs remain offline bootstrap and must **not** invent pricing or cache capabilities the official docs do not state. Discovery is never invoked by `create*ProviderPackage()`; hosts that want live `cost`/`cache` pass the returned models into package `models:` (or register them themselves).
|
|
199
|
+
|
|
188
200
|
## Security and performance notes
|
|
189
201
|
|
|
190
202
|
- Cache hints are best-effort and do not guarantee cache hits.
|
|
@@ -133,6 +133,44 @@ await assertProviderStreamConforms({
|
|
|
133
133
|
});
|
|
134
134
|
```
|
|
135
135
|
|
|
136
|
+
## Model discovery checklist
|
|
137
|
+
|
|
138
|
+
Every first-party package that ships (or plans) a `list*Models()` helper must keep setup network-free. Add these assertions in the package suite (pattern from NeuralWatt):
|
|
139
|
+
|
|
140
|
+
1. **`*_provider_setup_does_not_call_model_discovery`** — inject a counting `fetch` into `create*ProviderPackage({ fetch })`, run `setup`, assert `calls === 0`.
|
|
141
|
+
2. **`list_*_models_maps_fixture_…`** — fixture response maps to `ModelConfig` (`id` → `model`, documented capabilities/limits/cost/cache); no credentials in returned objects.
|
|
142
|
+
3. **`list_*_models_forwards_auth_abort_baseurl`** (as applicable) — Authorization owned by helper when key present; auth omitted when optional and unset; `signal` / `baseUrl` forwarded.
|
|
143
|
+
4. **`list_*_models_redacts_token_in_errors`** — non-OK bodies use `readBoundedResponseText` + `redactSecrets`; secret canaries absent from thrown messages.
|
|
144
|
+
5. **Malformed payload rejects** — missing `data` array (or provider-equivalent) throws a clear discovery error.
|
|
145
|
+
|
|
146
|
+
OpenRouter stays app-registration-first: an optional list helper must still not run during setup. AI SDK has no discovery export. Packages without a public list API document curated official-doc refresh instead of inventing a fake helper.
|
|
147
|
+
|
|
148
|
+
Canonical contract: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery).
|
|
149
|
+
|
|
150
|
+
## Thinking / reasoning checklist
|
|
151
|
+
|
|
152
|
+
Every first-party package that exposes thinking or reasoning controls should cover:
|
|
153
|
+
|
|
154
|
+
1. **Model default** — `ModelConfig.compat` (or documented capability) sets the official wire field when no per-turn override is present.
|
|
155
|
+
2. **Per-turn override wins** — `ProviderRequestOptions.compat` via `mergeProviderRequestOptions` / `applyThinkingLevel` overrides the model default.
|
|
156
|
+
3. **Shared family mapping** — effort levels from `ThinkingLevel` land in the package's recommended family (`openai_reasoning` / `reasoning_effort` / `thinking_type` / `noop`) per [Thinking and reasoning](thinking-and-reasoning.md).
|
|
157
|
+
4. **Non-reasoning / noop** — applying a level with `noop` (or omitting compat) must not invent unsupported body fields.
|
|
158
|
+
5. **No inert `extra.thinkingLevel`** — package code must not rely on `options.extra.thinkingLevel` for wire mapping.
|
|
159
|
+
|
|
160
|
+
Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md).
|
|
161
|
+
|
|
162
|
+
## AI SDK adapter checklist
|
|
163
|
+
|
|
164
|
+
`@arnilo/prism-provider-ai-sdk` is a host-owned `LanguageModelV4` bridge. It does not participate in the discovery or thinking/reasoning checklists above. Cover instead:
|
|
165
|
+
|
|
166
|
+
1. **No catalog / no setup fetch** — package exports no `list*Models()`; `createAiSdkProvider` wraps a host model only.
|
|
167
|
+
2. **Specification gate** — rejects non-v4 models (`specificationVersion !== "v4"` or missing `doStream`).
|
|
168
|
+
3. **Cache usage mapping** — `finish.usage.inputTokens.cacheRead`/`cacheWrite` map to `Usage.cacheReadTokens`/`cacheWriteTokens`; adapter does not emit cache request fields.
|
|
169
|
+
4. **Reasoning stream mapping** — `reasoning-delta` → thinking deltas; assistant `thinking` blocks replay as AI SDK `reasoning` prompt parts.
|
|
170
|
+
5. **Host-owned controls** — `options.compat` / `options.extra` forward as `providerOptions.prism`; reasoning effort stays on the host model.
|
|
171
|
+
|
|
172
|
+
Canonical contract: [AI SDK provider adapter](providers/ai-sdk.md).
|
|
173
|
+
|
|
136
174
|
## Extension and configuration notes
|
|
137
175
|
|
|
138
176
|
The helpers are a testing subpath only. Provider packages can use them with their own mocked fetch/transport or `createMockProvider()`. Live provider tests should stay opt-in and env-gated outside Prism's default test suite.
|
|
@@ -148,7 +186,7 @@ The helpers are a testing subpath only. Provider packages can use them with thei
|
|
|
148
186
|
## Related APIs
|
|
149
187
|
|
|
150
188
|
- [Provider layer](provider-layer.md): `AIProvider`, provider events, and mock provider.
|
|
151
|
-
- [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters.
|
|
189
|
+
- [Provider packages](provider-packages.md): package authors can use conformance helpers for adapters; includes the caller-gated discovery contract and setup zero-fetch rule.
|
|
152
190
|
- [AI SDK provider adapter](providers/ai-sdk.md): optional `LanguageModelV4` bridge tested with a fake AI SDK model.
|
|
153
191
|
- [OpenAI-compatible provider](providers/openai-compatible.md): optional provider adapter tested with mocked streams.
|
|
154
192
|
- [Public contracts](public-contracts.md): provider request/event/usage contracts.
|
|
@@ -67,7 +67,7 @@ Phase 6 also adds optional [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md
|
|
|
67
67
|
|
|
68
68
|
Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, or `OPENCODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification.
|
|
69
69
|
|
|
70
|
-
These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only. `@arnilo/prism-provider-opencode-go` now registers
|
|
70
|
+
These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-provider-openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only, with optional `models`/`codexModels` overrides and an opt-in `listOpenAIModels()` helper for official `GET /models` discovery. `@arnilo/prism-provider-opencode-go` now registers docs-verified OpenCode Go open coding models with dual OpenAI/Anthropic routes (`compat.route`), official default base `https://opencode.ai/zen/go/v1`, `reasoning_content`/thinking preserve, and an opt-in `listOpenCodeGoModels()` helper for official `GET /zen/go/v1/models`. `@arnilo/prism-provider-openrouter` now registers an app-controlled OpenRouter catalog with routing/`reasoning`/cache passthrough, assistant `reasoning` replay, optional top-level automatic `cache_control`, and an opt-in `listOpenRouterModels()` helper for official `GET /api/v1/models` (setup still never fetches). `@arnilo/prism-provider-zai` now registers featured GLM-5.x/4.x metadata with official `thinking`/`reasoning_effort`/`tool_stream`/`clear_thinking` mapping, Preserved Thinking `reasoning_content` replay, implicit context caching, and an opt-in `listZaiModels()` helper for OpenAI-compatible `GET /models`. `@arnilo/prism-provider-kimi` now registers Kimi Coding Anthropic-compatible behavior by default, optional callable Moonshot Open Platform Chat Completions when `includeMoonshotModels` is requested, official Coding/Open Platform featured ids, thinking/`reasoning_effort` compat mapping, and an opt-in `listKimiModels()` helper for Moonshot `GET /v1/models`. `@arnilo/prism-provider-neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation.
|
|
71
71
|
|
|
72
72
|
### First-party cache behavior
|
|
73
73
|
|
|
@@ -75,14 +75,71 @@ Every first-party provider package hardens prompt-cache behavior so it cannot em
|
|
|
75
75
|
|
|
76
76
|
- **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
77
77
|
- **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
78
|
-
- **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; `cache_control
|
|
79
|
-
- **OpenCode Go**: `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars;
|
|
78
|
+
- **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
|
|
79
|
+
- **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
|
|
80
80
|
- **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
|
|
81
81
|
- **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
|
|
82
82
|
- **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
|
|
83
83
|
|
|
84
84
|
See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
|
|
85
85
|
|
|
86
|
+
## Caller-gated model discovery
|
|
87
|
+
|
|
88
|
+
First-party packages keep `create*ProviderPackage()` network-free. Latest models come from **caller-gated** `list*Models()` helpers that hosts invoke explicitly and then pass back via `models:` (or register themselves). Plan 015's "no setup catalog fetch" rule still holds; Plan 067 adds on-demand discovery without hidden latency.
|
|
89
|
+
|
|
90
|
+
### Contract
|
|
91
|
+
|
|
92
|
+
```ts
|
|
93
|
+
export async function listExampleModels(options: {
|
|
94
|
+
apiKey?: CredentialValueSource;
|
|
95
|
+
fetch?: typeof fetch;
|
|
96
|
+
baseUrl?: string;
|
|
97
|
+
signal?: AbortSignal;
|
|
98
|
+
headers?: Readonly<Record<string, string>>;
|
|
99
|
+
}): Promise<ModelConfig[]> {
|
|
100
|
+
// GET {baseUrl}/models — never called from create*ProviderPackage()
|
|
101
|
+
}
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
| Rule | Requirement |
|
|
105
|
+
| --- | --- |
|
|
106
|
+
| Setup | `create*ProviderPackage().setup` performs **zero** fetches / discovery calls |
|
|
107
|
+
| Shape | Package-local `list*Models(options) → Promise<ModelConfig[]>` + optional `map*Model(entry)` |
|
|
108
|
+
| Injectables | `fetch`, `baseUrl`, `signal`, optional `apiKey` / `headers` |
|
|
109
|
+
| Transport | Error bodies via `@arnilo/prism/providers/transport` `readBoundedResponseText`; credentials via `resolveCredentialValue` + `redactSecrets` |
|
|
110
|
+
| Return | `ModelConfig[]` only — never embed API keys, tokens, or auth headers in returned metadata |
|
|
111
|
+
| Static catalog | Featured aliases / offline bootstrap only; may omit live pricing until discovery fills `cost` / `cache` |
|
|
112
|
+
| Core | Prefer package-local helpers. Do **not** add a core model-discovery registry. Extract a shared HTTP/list helper only when ≥2 packages share identical parsing |
|
|
113
|
+
|
|
114
|
+
Template: [`listNeuralWattModels`](providers/neuralwatt.md) in `@arnilo/prism-provider-neuralwatt`.
|
|
115
|
+
|
|
116
|
+
### Per-package policy
|
|
117
|
+
|
|
118
|
+
| Package | Discovery helper | Setup catalog | Notes |
|
|
119
|
+
| --- | --- | --- | --- |
|
|
120
|
+
| OpenAI | **`listOpenAIModels` (exists)** | Featured Responses/Codex aliases; factory accepts `models?` / `codexModels?` | Official `GET /v1/models`; Codex not listed by api.openai.com |
|
|
121
|
+
| Kimi | **`listKimiModels`** (Moonshot `GET /v1/models`) | Featured Coding ids + optional callable Moonshot | Official Moonshot/Kimi list-models; Coding curated |
|
|
122
|
+
| Z.AI | **`listZaiModels`** (OpenAI-compatible `GET /models`) + curated featured refresh | Featured GLM-5.2…4.5 aliases | No first-class docs.z.ai list page; discovery is best-effort; featured set from Chat Completions enum / overview |
|
|
123
|
+
| OpenRouter | **`listOpenRouterModels`** (official `GET /api/v1/models`) | **App-controlled** `models:` only — no bundled mega-catalog | Helper feeds host registration; setup still does not fetch |
|
|
124
|
+
| OpenCode Go | **`listOpenCodeGoModels`** (official `GET /zen/go/v1/models`) | Featured dual-route official Go aliases | Official Go docs endpoint table + sparse list API |
|
|
125
|
+
| NeuralWatt | **`listNeuralWattModels` (exists)** | Featured aliases without guessed pricing | Auth optional for public models |
|
|
126
|
+
| AI SDK | None | Host-owned `LanguageModelV4` | No Prism-side catalog by design |
|
|
127
|
+
|
|
128
|
+
Host pattern:
|
|
129
|
+
|
|
130
|
+
```ts
|
|
131
|
+
const models = await listNeuralWattModels({ apiKey, fetch });
|
|
132
|
+
await kernel.load([createNeuralWattProviderPackage({ apiKey, models })]);
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
Discovery may populate `ModelConfig.cache` and `ModelConfig.cost` from live metadata when the provider documents those fields; see [Provider caching](provider-caching.md#discovery-and-live-cache-cost-metadata). Package authors: include the [setup zero-fetch checklist](provider-conformance.md#model-discovery-checklist) in every first-party suite that ships or plans a `list*Models` helper.
|
|
136
|
+
|
|
137
|
+
## Per-turn thinking / reasoning
|
|
138
|
+
|
|
139
|
+
Hosts set effort with portable helpers from `@arnilo/prism` (`applyThinkingLevel`, `thinkingCompatFor`) that write official fields into `ProviderRequestOptions.compat`. Model defaults stay on `ModelConfig.compat`; per-turn patches win via `mergeProviderRequestOptions`. Providers keep reading `options.compat` / `model.compat` — do not invent a parallel options tree or put effort only in `extra`.
|
|
140
|
+
|
|
141
|
+
Canonical contract: [Thinking and reasoning](thinking-and-reasoning.md). Package-local knobs (NeuralWatt budgets, Z.AI `tool_stream`, Kimi keep/all) remain on `compat` beside the shared families.
|
|
142
|
+
|
|
86
143
|
## Third-party provider packaging
|
|
87
144
|
|
|
88
145
|
A third party ships their own providers the same way Prism ships first-party
|
package/docs/providers/ai-sdk.md
CHANGED
|
@@ -52,6 +52,8 @@ Unsupported content fails before `doStream` (for example unresolved `resourceUri
|
|
|
52
52
|
| `finish` usage | `usage` then `done` |
|
|
53
53
|
| `error` / thrown / abort | redacted `error` |
|
|
54
54
|
|
|
55
|
+
`finish.usage.inputTokens.cacheRead` / `cacheWrite` map to Prism `Usage.cacheReadTokens` / `cacheWriteTokens`. The adapter does not invent cache request fields; prompt caching is owned by the host `LanguageModelV4` and its upstream provider.
|
|
56
|
+
|
|
55
57
|
Provider-executed tool calls, files/sources/custom parts, warnings, and raw chunks are ignored rather than silently converted into unsupported Prism content.
|
|
56
58
|
|
|
57
59
|
## Request/response example
|
|
@@ -89,6 +91,40 @@ const result = await agent.createSession().run("Summarize this");
|
|
|
89
91
|
console.log(result.text);
|
|
90
92
|
```
|
|
91
93
|
|
|
94
|
+
## Model catalog and discovery
|
|
95
|
+
|
|
96
|
+
There is **no Prism-side model catalog** and **no `list*Models()` export** by design. Hosts supply a ready-made `LanguageModelV4` instance (typically from `@ai-sdk/openai`, `@ai-sdk/anthropic`, AI Gateway, or a custom provider) and register a matching `ModelConfig` for capabilities/limits.
|
|
97
|
+
|
|
98
|
+
Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials.
|
|
99
|
+
|
|
100
|
+
## Prompt caching
|
|
101
|
+
|
|
102
|
+
The adapter is **host-owned for request caching**. It does not emit `cache_control`, `prompt_cache_key`, `cacheKey`, or `cacheRetention` on AI SDK call options. Hosts configure caching on the underlying AI SDK model/provider (for example via AI SDK `providerOptions` on the model factory or per-call options forwarded through `options.compat` / `options.extra` → `providerOptions.prism`).
|
|
103
|
+
|
|
104
|
+
When the host model reports cache accounting on the `finish` stream part, Prism maps official AI SDK v4 usage fields:
|
|
105
|
+
|
|
106
|
+
| AI SDK `LanguageModelV4Usage` | Prism `Usage` |
|
|
107
|
+
| --- | --- |
|
|
108
|
+
| `inputTokens.cacheRead` | `cacheReadTokens` |
|
|
109
|
+
| `inputTokens.cacheWrite` | `cacheWriteTokens` |
|
|
110
|
+
| `inputTokens.total` | `inputTokens` |
|
|
111
|
+
| `outputTokens.total` | `outputTokens` |
|
|
112
|
+
|
|
113
|
+
See [Provider caching](../provider-caching.md) for the cross-provider matrix.
|
|
114
|
+
|
|
115
|
+
## Thinking and reasoning
|
|
116
|
+
|
|
117
|
+
Reasoning effort, budgets, and provider-specific thinking controls are **host-model-owned**. Prism does not map `ThinkingLevel` into AI SDK call options (`thinkingFamilyForModel` → `noop`). Hosts configure reasoning on the AI SDK model (for example OpenAI `reasoning.effort` via AI SDK `providerOptions`) and may pass per-turn overrides through `ProviderRequestOptions.compat` / `extra`, which the adapter forwards as `providerOptions.prism`.
|
|
118
|
+
|
|
119
|
+
Stream mapping:
|
|
120
|
+
|
|
121
|
+
| Direction | Mapping |
|
|
122
|
+
| --- | --- |
|
|
123
|
+
| AI SDK `reasoning-delta` → Prism | `content_delta` thinking |
|
|
124
|
+
| Prism `thinking` blocks → AI SDK prompt | `{ type: "reasoning", text }` on assistant messages |
|
|
125
|
+
|
|
126
|
+
Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); [Language Model Specification V4](https://github.com/vercel/ai/tree/main/packages/provider/src/language-model/v4); `@ai-sdk/provider` `LanguageModelV4Usage` (`inputTokens.cacheRead` / `cacheWrite`).
|
|
127
|
+
|
|
92
128
|
## Extension and configuration notes
|
|
93
129
|
|
|
94
130
|
- Peer dependency: `@ai-sdk/provider@^4.0.0`. Upgrade policy tracks one specification major at a time.
|
package/docs/providers/kimi.md
CHANGED
|
@@ -2,132 +2,195 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-provider-kimi` provides
|
|
6
|
-
Coding using an Anthropic-compatible `/messages` endpoint with
|
|
7
|
-
`User-Agent: KimiCLI/1.5` (unless overridden). Moonshot/Open Platform model
|
|
8
|
-
metadata is optional.
|
|
5
|
+
`@arnilo/prism-provider-kimi` provides two distinct, side-effect-free routes:
|
|
9
6
|
|
|
10
|
-
|
|
11
|
-
|
|
7
|
+
1. **Kimi For Coding** (default) — Anthropic-compatible `POST /messages` on
|
|
8
|
+
`https://api.kimi.com/coding` with `User-Agent: KimiCLI/1.5` (unless overridden).
|
|
9
|
+
2. **Moonshot Open Platform** (opt-in) — OpenAI-compatible `POST /chat/completions`
|
|
10
|
+
on `https://api.moonshot.ai/v1` (or `api.moonshot.cn/v1`), registered only when
|
|
11
|
+
`includeMoonshotModels: true`.
|
|
12
|
+
|
|
13
|
+
Official model ids differ by route. Coding uses `kimi-for-coding`,
|
|
14
|
+
`kimi-for-coding-highspeed`, and `k3`. Open Platform uses `kimi-k2.7-code`,
|
|
15
|
+
`kimi-k3`, and related catalog ids. Pi's `k2p7` alias is **not** used.
|
|
16
|
+
|
|
17
|
+
Caller-gated discovery via `listKimiModels()` hits the official Moonshot
|
|
18
|
+
`GET /v1/models` endpoint. Package setup never fetches.
|
|
12
19
|
|
|
13
20
|
## When to use it
|
|
14
21
|
|
|
15
|
-
Use it when a host app wants
|
|
16
|
-
`AgentSession` runtime with Kimi-specific
|
|
22
|
+
Use it when a host app wants Kimi For Coding and/or Moonshot Open Platform through
|
|
23
|
+
Prism's `AgentSession` runtime with Kimi-specific serializers, thinking controls,
|
|
24
|
+
and cache policy.
|
|
17
25
|
|
|
18
|
-
Do not use it for
|
|
19
|
-
|
|
26
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches, or
|
|
27
|
+
real-network tests (live tests stay opt-in).
|
|
20
28
|
|
|
21
29
|
## Inputs / request
|
|
22
30
|
|
|
23
31
|
```ts
|
|
24
|
-
import {
|
|
32
|
+
import {
|
|
33
|
+
createKimiProviderPackage,
|
|
34
|
+
listKimiModels,
|
|
35
|
+
defineKimiModel,
|
|
36
|
+
} from "@arnilo/prism-provider-kimi";
|
|
25
37
|
|
|
26
38
|
createKimiProviderPackage(options: KimiProviderPackageOptions): ProviderPackage
|
|
27
|
-
|
|
39
|
+
listKimiModels(options?: ListKimiModelsOptions): Promise<ModelConfig[]>
|
|
40
|
+
defineKimiModel(config: KimiModelConfig): ModelConfig
|
|
28
41
|
```
|
|
29
42
|
|
|
30
43
|
| Field | Type | Purpose |
|
|
31
44
|
| --- | --- | --- |
|
|
32
|
-
| `kimiApiKey` | `CredentialValueSource` |
|
|
45
|
+
| `kimiApiKey` | `CredentialValueSource` | Kimi For Coding API key. |
|
|
46
|
+
| `moonshotApiKey` | `CredentialValueSource` | Moonshot Open Platform API key (not interchangeable with Coding keys). |
|
|
33
47
|
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
34
|
-
| `baseUrl` | `string` | Overrides the
|
|
35
|
-
| `
|
|
36
|
-
| `
|
|
37
|
-
| `
|
|
38
|
-
| `
|
|
39
|
-
| `
|
|
48
|
+
| `baseUrl` | `string` | Overrides the Coding base URL. |
|
|
49
|
+
| `moonshotBaseUrl` | `string` | Overrides Moonshot base URL (default `https://api.moonshot.ai/v1`). |
|
|
50
|
+
| `id` / `moonshotId` | `string` | Provider ids (defaults `kimi-coding` / `moonshot`). |
|
|
51
|
+
| `userAgent` | `string` | Overrides Coding `User-Agent: KimiCLI/1.5`. |
|
|
52
|
+
| `models` | `readonly ModelConfig[]` | Overrides featured Coding models. |
|
|
53
|
+
| `includeMoonshotModels` | `boolean` | Registers callable Moonshot provider + models when `true`. |
|
|
54
|
+
| `moonshotModels` | `readonly ModelConfig[]` | Overrides featured Moonshot models when included. |
|
|
40
55
|
|
|
41
56
|
## Outputs / response / events
|
|
42
57
|
|
|
43
58
|
| Surface | Behavior |
|
|
44
59
|
| --- | --- |
|
|
45
|
-
|
|
|
46
|
-
|
|
|
47
|
-
|
|
|
60
|
+
| Coding stream | Prism text, thinking deltas, tool-call delta/final, `usage` (`cache_read_input_tokens` / `cache_creation_input_tokens`), `done`, redacted `error`. |
|
|
61
|
+
| Moonshot stream | Same, with `delta.reasoning_content` → thinking; OpenAI-style usage cache details when present. |
|
|
62
|
+
| Block preservation | Coding: Anthropic `thinking` / `tool_use` / `tool_result`. Moonshot: `reasoning_content` on assistant replay when `preserveThinking`. |
|
|
63
|
+
| Auth methods | `api_key` for `kimi-coding`; also `moonshot` when opted in. |
|
|
48
64
|
|
|
49
65
|
Unsupported block placements or unclaimed images fail before fetch.
|
|
50
66
|
|
|
67
|
+
## Route differences
|
|
68
|
+
|
|
69
|
+
| | Kimi For Coding | Moonshot Open Platform |
|
|
70
|
+
| --- | --- | --- |
|
|
71
|
+
| Base URL | `https://api.kimi.com/coding` | `https://api.moonshot.ai/v1` (or `.cn`) |
|
|
72
|
+
| Wire API | Anthropic `/messages` | OpenAI `/chat/completions` |
|
|
73
|
+
| Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k3` (+ discovery) |
|
|
74
|
+
| Discovery | No public list API — curated featured aliases | Official `GET /v1/models` via `listKimiModels()` |
|
|
75
|
+
| Cache | Implicit by default; opt-in Anthropic `cache_control` | Implicit only — never emits Anthropic `cache_control` |
|
|
76
|
+
| Thinking | Block replay + body `thinking` / `reasoning_effort` | `reasoning_content` replay + body `thinking` / `reasoning_effort` |
|
|
77
|
+
|
|
78
|
+
The Anthropic `/messages` request/response contract for Kimi remains under-documented
|
|
79
|
+
upstream ([MoonshotAI/Kimi-K2#129](https://github.com/MoonshotAI/Kimi-K2/issues/129));
|
|
80
|
+
Prism treats official Chat Completions thinking fields as best-effort passthrough on
|
|
81
|
+
the Coding route.
|
|
82
|
+
|
|
83
|
+
## Thinking / reasoning
|
|
84
|
+
|
|
85
|
+
Official fields (Open Platform docs; Coding docs for `k3` effort mapping):
|
|
86
|
+
|
|
87
|
+
| Model family | Official control | Prism `compat` |
|
|
88
|
+
| --- | --- | --- |
|
|
89
|
+
| K3 / Coding `k3` | top-level `reasoning_effort` (`max` on Open Platform; Coding also `low`/`high`) | `compat.reasoning_effort` — use Task 4 family `reasoning_effort` |
|
|
90
|
+
| K2.7-code / Coding | thinking always on; Preserved Thinking always on | omit `thinking` by default; `preserveThinking: true` for replay; do not send `disabled` |
|
|
91
|
+
| K2.6 / K2.5 | `thinking.type` enabled/disabled; K2.6 optional `keep: "all"` | `compat.thinking` — Task 4 family `thinking_type` |
|
|
92
|
+
|
|
93
|
+
Per-turn `ProviderRequestOptions.compat` wins over `ModelConfig.compat`. Helpers:
|
|
94
|
+
`kimiThinking`, `kimiReasoningEffort`, `kimiPreserveThinking`.
|
|
95
|
+
|
|
51
96
|
## Request/response example
|
|
52
97
|
|
|
53
|
-
|
|
98
|
+
Coding (Anthropic-compatible `/messages`):
|
|
54
99
|
|
|
55
100
|
```json
|
|
56
101
|
{
|
|
57
|
-
"model": "kimi-
|
|
58
|
-
"messages": [{ "role": "user", "content": "Hello" }],
|
|
102
|
+
"model": "kimi-for-coding",
|
|
103
|
+
"messages": [{ "role": "user", "content": [{ "type": "text", "text": "Hello" }] }],
|
|
59
104
|
"stream": true
|
|
60
105
|
}
|
|
61
106
|
```
|
|
62
107
|
|
|
108
|
+
Moonshot (Chat Completions):
|
|
109
|
+
|
|
110
|
+
```json
|
|
111
|
+
{
|
|
112
|
+
"model": "kimi-k3",
|
|
113
|
+
"messages": [{ "role": "user", "content": "Hello" }],
|
|
114
|
+
"stream": true,
|
|
115
|
+
"reasoning_effort": "max"
|
|
116
|
+
}
|
|
117
|
+
```
|
|
118
|
+
|
|
63
119
|
## Implementation example
|
|
64
120
|
|
|
65
121
|
```ts
|
|
66
122
|
import { createExtensionKernel } from "@arnilo/prism";
|
|
67
|
-
import {
|
|
123
|
+
import {
|
|
124
|
+
createKimiProviderPackage,
|
|
125
|
+
listKimiModels,
|
|
126
|
+
} from "@arnilo/prism-provider-kimi";
|
|
68
127
|
|
|
69
128
|
const kernel = createExtensionKernel();
|
|
70
129
|
await kernel.load([
|
|
71
|
-
createKimiProviderPackage({ kimiApiKey: "fake-kimi-key"
|
|
130
|
+
createKimiProviderPackage({ kimiApiKey: "fake-kimi-key" }),
|
|
72
131
|
]);
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
Register Moonshot/Open Platform metadata explicitly:
|
|
76
132
|
|
|
77
|
-
|
|
78
|
-
|
|
133
|
+
// Opt-in Moonshot Open Platform (callable provider + featured models)
|
|
134
|
+
await kernel.load([
|
|
135
|
+
createKimiProviderPackage({
|
|
136
|
+
kimiApiKey: "fake-kimi-key",
|
|
137
|
+
includeMoonshotModels: true,
|
|
138
|
+
moonshotApiKey: "fake-moonshot-key",
|
|
139
|
+
}),
|
|
140
|
+
]);
|
|
79
141
|
|
|
142
|
+
// Caller-gated discovery — never runs during setup
|
|
143
|
+
const latest = await listKimiModels({ apiKey: "fake-moonshot-key", fetch });
|
|
80
144
|
await kernel.load([
|
|
81
|
-
createKimiProviderPackage({
|
|
145
|
+
createKimiProviderPackage({
|
|
146
|
+
includeMoonshotModels: true,
|
|
147
|
+
moonshotApiKey: "fake-moonshot-key",
|
|
148
|
+
moonshotModels: latest.filter((m) => m.model.startsWith("kimi-")),
|
|
149
|
+
}),
|
|
82
150
|
]);
|
|
83
151
|
```
|
|
84
152
|
|
|
85
153
|
## Extension and configuration notes
|
|
86
154
|
|
|
87
|
-
- Hosts choose base
|
|
155
|
+
- Hosts choose base URLs, provider ids, `User-Agent`, model lists, credential sources,
|
|
88
156
|
and `fetch` impl.
|
|
89
|
-
- Moonshot
|
|
90
|
-
|
|
91
|
-
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
157
|
+
- Moonshot is registered only with `includeMoonshotModels: true` (provider + models + auth).
|
|
158
|
+
- Featured catalogs are offline bootstrap only; refresh Open Platform via `listKimiModels()`.
|
|
92
159
|
|
|
93
160
|
### Cache behavior
|
|
94
161
|
|
|
95
|
-
- Default catalog models
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
-
|
|
100
|
-
`
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
model allows long retention (`ModelConfig.cache.longRetention !== false`).
|
|
106
|
-
- The Moonshot Open Platform route (`compat.route: "openai"`) never receives
|
|
107
|
-
Anthropic `cache_control` fields.
|
|
108
|
-
- Usage accounting is preserved: Anthropic-route `cache_read_input_tokens` maps to
|
|
109
|
-
`Usage.cacheReadTokens` and `cache_creation_input_tokens` maps to
|
|
110
|
-
`Usage.cacheWriteTokens`.
|
|
162
|
+
- Default Coding catalog models use **implicit caching** and send no explicit
|
|
163
|
+
`cache_control` fields unless the model opts in with
|
|
164
|
+
`ModelConfig.cache.kind: "cache_control"`.
|
|
165
|
+
- When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
|
|
166
|
+
caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
|
|
167
|
+
block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
|
|
168
|
+
the model allows long retention.
|
|
169
|
+
- The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
|
|
170
|
+
- Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
|
|
171
|
+
`cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
|
|
111
172
|
|
|
112
173
|
## Security and performance notes
|
|
113
174
|
|
|
114
|
-
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
|
|
175
|
+
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
|
|
176
|
+
helpers (`readSseData`, `readBoundedResponseText`).
|
|
115
177
|
- No network calls during import, setup, build, or default tests.
|
|
116
178
|
- No automatic environment, file, keychain, or shell credential lookup.
|
|
117
|
-
-
|
|
118
|
-
and redacted from errors.
|
|
179
|
+
- Credentials are resolved per request from caller-supplied values or resolvers
|
|
180
|
+
and redacted from errors (including discovery failures).
|
|
119
181
|
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
|
|
120
182
|
but provider-owned headers (`content-type`, `user-agent`, `authorization`)
|
|
121
183
|
are applied last and cannot be overridden by caller headers.
|
|
122
|
-
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus
|
|
123
|
-
|
|
184
|
+
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus provider-specific
|
|
185
|
+
env names; default tests are network-free.
|
|
124
186
|
|
|
125
187
|
## Related APIs
|
|
126
188
|
|
|
127
189
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
128
|
-
|
|
190
|
+
caller-gated discovery, Anthropic/OpenAI routes.
|
|
191
|
+
- [Thinking and reasoning](../thinking-and-reasoning.md): portable `ThinkingLevel`
|
|
192
|
+
helpers and Kimi family mapping.
|
|
129
193
|
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
130
194
|
`resolveCredentialValue`, `redactSecrets`.
|
|
131
|
-
- [Provider
|
|
132
|
-
mapping.
|
|
195
|
+
- [Provider caching](../provider-caching.md): explicit/implicit matrix.
|
|
133
196
|
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|