@arnilo/prism 0.4.0 → 0.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +41 -1
- package/README.md +23 -20
- package/dist/agent-run-state.d.ts +1 -2
- package/dist/agent-run-state.js +0 -3
- package/dist/agent-session/session/assemble.d.ts +6 -0
- package/dist/agent-session/session/assemble.js +391 -0
- package/dist/agent-session/session/persist.d.ts +28 -0
- package/dist/agent-session/session/persist.js +166 -0
- package/dist/agent-session/session/provider-round.d.ts +6 -0
- package/dist/agent-session/session/provider-round.js +231 -0
- package/dist/agent-session/session/tool-round.d.ts +31 -0
- package/dist/agent-session/session/tool-round.js +473 -0
- package/dist/agent-session/session/types.d.ts +115 -0
- package/dist/agent-session/session/types.js +5 -0
- package/dist/agent-session/session.d.ts +49 -43
- package/dist/agent-session/session.js +24 -1180
- package/dist/capture.d.ts +63 -0
- package/dist/capture.js +67 -0
- package/dist/cli-init.d.ts +18 -2
- package/dist/cli-init.js +2 -7
- package/dist/cli-runner.d.ts +2 -2
- package/dist/cli-runner.js +45 -9
- package/dist/content.d.ts +3 -3
- package/dist/content.js +3 -1
- package/dist/contracts-core/agent.d.ts +4 -0
- package/dist/contracts-core/batch.d.ts +97 -0
- package/dist/contracts-core/batch.js +65 -0
- package/dist/contracts-core/content.d.ts +72 -1
- package/dist/contracts-core/embeddings.d.ts +30 -0
- package/dist/contracts-core/embeddings.js +17 -0
- package/dist/contracts-core/images.d.ts +60 -0
- package/dist/contracts-core/images.js +17 -0
- package/dist/contracts-core/moderation.d.ts +46 -0
- package/dist/contracts-core/moderation.js +34 -0
- package/dist/contracts-core/speech.d.ts +39 -0
- package/dist/contracts-core/speech.js +17 -0
- package/dist/contracts-core/transcription.d.ts +48 -0
- package/dist/contracts-core/transcription.js +17 -0
- package/dist/contracts-core/video.d.ts +61 -0
- package/dist/contracts-core/video.js +17 -0
- package/dist/contracts-core.d.ts +7 -0
- package/dist/contracts-core.js +7 -0
- package/dist/contracts-protocol.d.ts +2 -0
- package/dist/index.d.ts +7 -5
- package/dist/index.js +5 -4
- package/dist/input.js +3 -2
- package/dist/node/agent-definitions.d.ts +1 -8
- package/dist/node/agent-definitions.js +0 -34
- package/dist/node/settings.d.ts +0 -1
- package/dist/node/settings.js +0 -5
- package/dist/pinned-fetch.js +29 -3
- package/dist/provider-events.js +3 -4
- package/dist/provider-request-policy.d.ts +15 -0
- package/dist/provider-request-policy.js +52 -0
- package/dist/providers/media.d.ts +1 -2
- package/dist/providers/media.js +1 -4
- package/dist/rpc.d.ts +1 -1
- package/dist/rpc.js +4 -4
- package/dist/testing/provider-conformance.d.ts +114 -5
- package/dist/testing/provider-conformance.js +342 -0
- package/dist/testing/tool-effect-store-conformance.d.ts +0 -1
- package/dist/testing/tool-effect-store-conformance.js +0 -3
- package/dist/thinking.d.ts +48 -9
- package/dist/thinking.js +134 -8
- package/docs/0.1.0-readiness.md +3 -3
- package/docs/a2a.md +2 -2
- package/docs/acp.md +3 -3
- package/docs/ag-ui-adoption.md +1 -1
- package/docs/ag-ui.md +1 -2
- package/docs/agent-definitions.md +1 -1
- package/docs/agent-events.md +5 -5
- package/docs/agent-identity.md +13 -2
- package/docs/agent-session-runtime.md +2 -1
- package/docs/audit-export.md +3 -3
- package/docs/batch-jobs.md +120 -0
- package/docs/cli-rpc.md +20 -9
- package/docs/coding-agent-tools.md +19 -19
- package/docs/coding-review-and-diagnostics.md +2 -2
- package/docs/coding-security.md +4 -4
- package/docs/coding-workspaces.md +2 -2
- package/docs/compaction-llm.md +2 -0
- package/docs/compaction-observational-memory.md +3 -0
- package/docs/computer-use-linux.md +13 -2
- package/docs/context-and-skills.md +1 -1
- package/docs/conversations.md +4 -4
- package/docs/credential-storage.md +11 -7
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/data-classification.md +1 -1
- package/docs/database-persistence.md +4 -4
- package/docs/dev-inspector.md +6 -6
- package/docs/device-adapters.md +2 -2
- package/docs/diagrams.md +1 -1
- package/docs/document-reader.md +6 -6
- package/docs/documents.md +5 -4
- package/docs/embeddings.md +112 -0
- package/docs/enterprise-postgres-state.md +7 -7
- package/docs/evaluations.md +8 -8
- package/docs/extensions.md +3 -3
- package/docs/forge-integration.md +3 -3
- package/docs/graft.md +2 -2
- package/docs/guardrails.md +1 -1
- package/docs/host-security.md +15 -15
- package/docs/image-generation.md +129 -0
- package/docs/impeccable.md +5 -3
- package/docs/index.md +64 -36
- package/docs/indexed-code-search.md +2 -2
- package/docs/input-and-prompt-assembly.md +1 -1
- package/docs/language-intelligence.md +4 -4
- package/docs/live-testing.md +126 -0
- package/docs/mcp-tools.md +43 -12
- package/docs/middleware-hooks.md +1 -1
- package/docs/migrate-to-0.4.md +3 -3
- package/docs/migrate-to-0.5.md +144 -0
- package/docs/migration.md +33 -1
- package/docs/model-registry.md +38 -0
- package/docs/model-routing.md +5 -5
- package/docs/moderation.md +117 -0
- package/docs/multi-agent-patterns.md +4 -4
- package/docs/multimodal-content.md +26 -2
- package/docs/obscura.md +2 -2
- package/docs/observability.md +32 -7
- package/docs/openapi-tools.md +13 -3
- package/docs/operations.md +11 -0
- package/docs/performance.md +7 -7
- package/docs/persistence-credentials-multimodality-primitives.md +6 -6
- package/docs/policy-and-audit.md +17 -7
- package/docs/ponytail.md +1 -1
- package/docs/postgres-persistence.md +5 -5
- package/docs/process-sessions.md +2 -2
- package/docs/prompt-registry.md +7 -7
- package/docs/provider-caching.md +8 -2
- package/docs/provider-conformance.md +23 -1
- package/docs/provider-packages.md +49 -17
- package/docs/provider-primitives.md +1 -1
- package/docs/provider-request-policies.md +19 -6
- package/docs/providers/ai-sdk.md +27 -3
- package/docs/providers/alibaba.md +17 -1
- package/docs/providers/anthropic.md +16 -0
- package/docs/providers/azure.md +29 -1
- package/docs/providers/bedrock.md +27 -0
- package/docs/providers/clinepass.md +16 -0
- package/docs/providers/commandcode.md +265 -0
- package/docs/providers/deepseek.md +16 -0
- package/docs/providers/google.md +16 -0
- package/docs/providers/hyper.md +296 -0
- package/docs/providers/kimi.md +16 -0
- package/docs/providers/neuralwatt.md +16 -0
- package/docs/providers/ollama.md +27 -0
- package/docs/providers/openai-compatible.md +16 -0
- package/docs/providers/openai.md +16 -0
- package/docs/providers/opencode-go.md +16 -0
- package/docs/providers/openrouter.md +17 -1
- package/docs/providers/vertex.md +28 -0
- package/docs/providers/xai.md +16 -0
- package/docs/providers/zai.md +16 -0
- package/docs/public-contracts.md +1 -1
- package/docs/rag.md +26 -4
- package/docs/release-and-install.md +103 -46
- package/docs/resource-loading.md +1 -1
- package/docs/runs-and-usage.md +14 -2
- package/docs/server.md +5 -5
- package/docs/settings-auth-trust-security.md +7 -5
- package/docs/sheets.md +2 -2
- package/docs/speech.md +126 -0
- package/docs/sqlite-persistence.md +4 -4
- package/docs/supervisors.md +3 -3
- package/docs/thinking-and-reasoning.md +99 -61
- package/docs/tool-conformance.md +1 -1
- package/docs/tool-execution-primitives.md +8 -8
- package/docs/tools.md +4 -4
- package/docs/use-case-model-selection.md +1 -1
- package/docs/web-tools.md +1 -1
- package/docs/wiki.md +1 -1
- package/docs/work-artifacts-and-review.md +17 -6
- package/docs/work-connectors.md +4 -4
- package/docs/work-tools.md +5 -5
- package/docs/workflow-orchestration-primitives.md +11 -11
- package/docs/workflows.md +5 -5
- package/package.json +11 -8
- package/templates/init/providers.json +24 -8
- package/docs/antigravity-agent.md +0 -207
package/docs/process-sessions.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`createProcessSessions` is an optional host-activated registry in `@arnilo/prism-coding-agent` for **long-running** child processes: start, cursor-paged output, input, wait, signal/kill, and release (detach). Sessions have bounded lifetime (sweep on registry access — no import-time timers), ownership/identity attribution, durable metadata (command fingerprint without env/secrets), and typed `CodingProcessEvent`s via a host callback. Reuses `ExecutionPolicy`, `killProcessTree`, and `OutputAccumulator` (including spill + `readRaw` cursor paging). Optional duck-typed `sandbox` backend uses `startProcess` when present; one-shot adapters fail closed. Nothing spawns on import or construction.
|
|
5
|
+
`createProcessSessions` is an optional host-activated registry in `@arnilo/prism-coding-tools/agent` for **long-running** child processes: start, cursor-paged output, input, wait, signal/kill, and release (detach). Sessions have bounded lifetime (sweep on registry access — no import-time timers), ownership/identity attribution, durable metadata (command fingerprint without env/secrets), and typed `CodingProcessEvent`s via a host callback. Reuses `ExecutionPolicy`, `killProcessTree`, and `OutputAccumulator` (including spill + `readRaw` cursor paging). Optional duck-typed `sandbox` backend uses `startProcess` when present; one-shot adapters fail closed. Nothing spawns on import or construction.
|
|
6
6
|
|
|
7
7
|
| Export | Purpose |
|
|
8
8
|
| --- | --- |
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
Use when a host needs attachable long-running processes (watch modes, language servers, interactive CLIs) that one-shot `shell` cannot model. Do not use as a job-control language. `pty: true` (host-selected PTY) requires a `ptyBackend` passed to `createProcessSessions`; without one it fails closed before spawn with `ERR_PRISM_PROCESS_PTY_UNSUPPORTED`. Pass a sandbox with `startProcess` for contained long-running work; omit `sandbox` for native spawn.
|
|
21
21
|
|
|
22
22
|
```ts
|
|
23
|
-
import { createProcessSessions } from "@arnilo/prism-coding-agent";
|
|
23
|
+
import { createProcessSessions } from "@arnilo/prism-coding-tools/agent";
|
|
24
24
|
|
|
25
25
|
const sessions = createProcessSessions({ cwd: workspaceRoot, policy, sandbox, onEvent });
|
|
26
26
|
const p = await sessions.start({ command: "npm", args: ["test", "--", "--watch"] });
|
package/docs/prompt-registry.md
CHANGED
|
@@ -2,12 +2,12 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
The optional `@arnilo/prism-prompts` package stores prompt assets as immutable, content-hashed versions. It provides a memory store plus SQLite and PostgreSQL adapters. The registry returns prompt data; it does not compose system-prompt layers, evaluate prompt quality, discover files, or activate text.
|
|
5
|
+
The optional `@arnilo/prism-core/governance/prompts` package stores prompt assets as immutable, content-hashed versions. It provides a memory store plus SQLite and PostgreSQL adapters. The registry returns prompt data; it does not compose system-prompt layers, evaluate prompt quality, discover files, or activate text.
|
|
6
6
|
|
|
7
7
|
## Inputs / request
|
|
8
8
|
|
|
9
9
|
```ts
|
|
10
|
-
import { createMemoryPromptStore } from "@arnilo/prism-prompts";
|
|
10
|
+
import { createMemoryPromptStore } from "@arnilo/prism-core/governance/prompts";
|
|
11
11
|
|
|
12
12
|
const store = createMemoryPromptStore();
|
|
13
13
|
const version = await store.put({
|
|
@@ -57,10 +57,10 @@ The ref is opaque identity — name, version number, and the store's SHA-256 bod
|
|
|
57
57
|
## Durable adapters
|
|
58
58
|
|
|
59
59
|
```ts
|
|
60
|
-
import { createSqlitePromptStore } from "@arnilo/prism-prompts";
|
|
60
|
+
import { createSqlitePromptStore } from "@arnilo/prism-core/governance/prompts";
|
|
61
61
|
const sqlite = createSqlitePromptStore({ filename: "./prompts.db" });
|
|
62
62
|
|
|
63
|
-
import { createPostgresPromptStore } from "@arnilo/prism-prompts";
|
|
63
|
+
import { createPostgresPromptStore } from "@arnilo/prism-core/governance/prompts";
|
|
64
64
|
const postgres = await createPostgresPromptStore({
|
|
65
65
|
connectionString: process.env.DATABASE_URL,
|
|
66
66
|
schema: "prism",
|
|
@@ -74,7 +74,7 @@ SQLite uses `better-sqlite3`; PostgreSQL uses a caller-supplied or adapter-owned
|
|
|
74
74
|
`assertPromptPromotion` composes [evaluations](evaluations.md) with the store to answer one question — should this candidate version replace the baseline? It resolves both versions (read-only), runs them head-to-head over a dataset through `runComparison`, and returns a typed verdict. It never promotes anything, writes nothing, and never touches a live agent:
|
|
75
75
|
|
|
76
76
|
```ts
|
|
77
|
-
import { assertPromptPromotion } from "@arnilo/prism-prompts";
|
|
77
|
+
import { assertPromptPromotion } from "@arnilo/prism-core/governance/prompts";
|
|
78
78
|
|
|
79
79
|
const v = await assertPromptPromotion({
|
|
80
80
|
store,
|
|
@@ -90,13 +90,13 @@ const v = await assertPromptPromotion({
|
|
|
90
90
|
if (v.verdict === "promote") await store.put({ ...hostInput, body: v.candidate.body, labels: ["production"] });
|
|
91
91
|
```
|
|
92
92
|
|
|
93
|
-
The verdict carries `promote`/`hold`, per-scorer `wins/losses/ties/failures`, `winRate`, the raw `ComparisonReport`, a redacted bounded `reportJson` (`serializeEvaluationReport`), and `reasons` on hold. The default gate holds unless the candidate wins strictly more scored comparisons than the baseline; `minimumWinRate` and `thresholds` add stricter gates, and threshold equality passes. Requires the optional peer `@arnilo/prism-evals` (install it or the helper fails closed with `ERR_PRISM_PROMPT_EVALS_PEER`). Promotion itself stays a host decision: applying the verdict means `put`-ing a new version with labels — the helper never does.
|
|
93
|
+
The verdict carries `promote`/`hold`, per-scorer `wins/losses/ties/failures`, `winRate`, the raw `ComparisonReport`, a redacted bounded `reportJson` (`serializeEvaluationReport`), and `reasons` on hold. The default gate holds unless the candidate wins strictly more scored comparisons than the baseline; `minimumWinRate` and `thresholds` add stricter gates, and threshold equality passes. Requires the optional peer `@arnilo/prism-core/governance/evals` (install it or the helper fails closed with `ERR_PRISM_PROMPT_EVALS_PEER`). Promotion itself stays a host decision: applying the verdict means `put`-ing a new version with labels — the helper never does.
|
|
94
94
|
|
|
95
95
|
## Limits and security
|
|
96
96
|
|
|
97
97
|
Names, bodies, labels, metadata, cursors, pages, and diffs have finite defaults and hard caps. Prompt bodies are data: no evaluation, template execution, file discovery, or implicit layer injection occurs. Store body hashes are integrity checks; a durable row whose hash no longer matches its body fails closed. Never put credentials or provider clients in prompt metadata.
|
|
98
98
|
|
|
99
|
-
Threat model: the registry is **host-trusted data**. Anyone who can write versions into the store is inside the trust boundary — `put`, label management, and `assertPromptPromotion` verdicts are host operations, never agent-reachable surfaces. Untrusted prompt-injection defense stays at Prism's existing untrusted-content boundaries (tool results, attachments, and provider output), which the store neither bypasses nor weakens: a resolved body enters the system-prompt layer exactly like a host-authored constant. The optional `@arnilo/prism-evals` peer is only loaded by `assertPromptPromotion` and never makes the store itself depend on evaluation infrastructure.
|
|
99
|
+
Threat model: the registry is **host-trusted data**. Anyone who can write versions into the store is inside the trust boundary — `put`, label management, and `assertPromptPromotion` verdicts are host operations, never agent-reachable surfaces. Untrusted prompt-injection defense stays at Prism's existing untrusted-content boundaries (tool results, attachments, and provider output), which the store neither bypasses nor weakens: a resolved body enters the system-prompt layer exactly like a host-authored constant. The optional `@arnilo/prism-core/governance/evals` peer is only loaded by `assertPromptPromotion` and never makes the store itself depend on evaluation infrastructure.
|
|
100
100
|
|
|
101
101
|
## Related APIs
|
|
102
102
|
|
package/docs/provider-caching.md
CHANGED
|
@@ -12,6 +12,8 @@ Provider caching documents Prism's cache intent surface:
|
|
|
12
12
|
|
|
13
13
|
Cache hints are best-effort. They describe intent; providers decide whether their native API can use them. Prism does not guarantee cache hits.
|
|
14
14
|
|
|
15
|
+
**Kernel defaults (0.5.1).** `applyDefaultProviderRequestOptions` fills `cache.breakpoints` with `{ location: "system_prompt" }` and `{ location: "last_stable_message" }` plus `cacheRetention: "short"` when `model.cache.kind === "cache_control"` or `model.cache.explicitBreakpoints === true`, unless the host set `cache.mode: "off"`, `cacheRetention: "none"`, or a non-empty breakpoint list. Implicit / `none` / host-owned (Azure, Bedrock, Vertex, AI SDK) models get no Prism markers. Session and cache keys are correlation ids — Cache keys must never be credentials.
|
|
16
|
+
|
|
15
17
|
## When to use it
|
|
16
18
|
|
|
17
19
|
Use this page when a host or provider package needs to:
|
|
@@ -150,8 +152,10 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
|
|
|
150
152
|
| `@arnilo/prism-providers/openai` | `openai_key` | Sends sanitized `prompt_cache_key`; pre-5.6 models emit `prompt_cache_retention: "24h"` when `longRetention`; GPT-5.6+ models (`explicitBreakpoints`) map `cache.breakpoints`/`cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` + `prompt_cache_breakpoint` markers (≤4 writes). | Stable cache key + stable prefix can improve reuse; keep selected anchors stable. | Best-effort only; `"short"`/`"none"` omit retention; `"30m"` TTL is the default and never emitted. |
|
|
151
153
|
| `@arnilo/prism-providers/anthropic` | `cache_control` | Marks only selected Anthropic message anchors; `system_prompt` breakpoints emit native `system` text blocks with the marker; `"long"` maps to documented `ttl: "1h"`. | Keep selected anchors stable. | Best-effort; never stamp every block. |
|
|
152
154
|
| `@arnilo/prism-providers/google` | none | Sends no Prism cache marker. | Host/model may have upstream behavior. | Gemini cache controls are not mapped in this package. |
|
|
153
|
-
| `@arnilo/prism-providers/openrouter` | `cache_control` |
|
|
155
|
+
| `@arnilo/prism-providers/openrouter` | `cache_control` | Kernel defaults (agent session / helper) emit **per-message** `cache_control` on `system_prompt` + `last_stable_message`. Raw `generate` with empty breakpoints may still send top-level automatic `cache_control`. `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
|
|
154
156
|
| `@arnilo/prism-providers/opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
|
|
157
|
+
| `@arnilo/prism-providers/hyper` | route-specific implicit / `cache_control` | Chat route sends no markers (implicit prefix caching); `qwen3.6-*` messages route applies `cache_control` only to caller-selected `cache.breakpoints` (max 4); **no `ttl`** (undocumented). Opt-in responses route emits OpenAI-standard `prompt_cache_key` from hints only, never retention/options. | Keep selected Anthropic anchors and prior history stable. | Best-effort only; `402 billing_error` when Hypercredits run out. |
|
|
158
|
+
| `@arnilo/prism-providers/commandcode` | route-specific implicit / `cache_control` | Chat route sends no markers (implicit prefix caching); `claude-*` messages route applies `cache_control` only to caller-selected `cache.breakpoints` (max 4); **no `ttl`** (undocumented). GPT-5.6 tiers keep docs `cacheWrite` prices in `cost` but stay implicit until the live probe verifies `prompt_cache_key`. | Keep selected Anthropic anchors and prior history stable. | Best-effort only; GPT-5.6 explicit `prompt_cache_key` upgrade pending probe (plan 055 Task 9); OSS models bill at mean per-provider price; DeepSeek off-peak 17h/day, peak 2×. |
|
|
155
159
|
| `@arnilo/prism-providers/zai` | `implicit` | No explicit cache payload; GLM context caching is automatic. | Resend unchanged prior history for implicit context-cache reuse. | Best-effort only; cache options do not force hits. |
|
|
156
160
|
| `@arnilo/prism-providers/kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
|
|
157
161
|
| `@arnilo/prism-providers/neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
|
|
@@ -171,8 +175,10 @@ Detailed first-party provider notes:
|
|
|
171
175
|
- OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
|
|
172
176
|
- Anthropic (`@arnilo/prism-providers/anthropic`): `kind: "cache_control"`; selected Anthropic message anchors receive `cache_control` and eligible long retention maps to `ttl: "1h"`. A `system_prompt` breakpoint serializes `system` as native text blocks carrying the marker (shared `systemCacheControlField()` helper; plain joined string when unmarked). Cache read/create usage maps to normalized cache read/write tokens.
|
|
173
177
|
- Google (`@arnilo/prism-providers/google`): sends no Prism cache-control payload. Do not infer cache hits or cache token counts from absent Gemini fields.
|
|
174
|
-
- OpenRouter (`@arnilo/prism-providers/openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing (from `cache.key` ?? legacy `cacheKey` ?? `sessionId`); with
|
|
178
|
+
- OpenRouter (`@arnilo/prism-providers/openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing (from `cache.key` ?? legacy `cacheKey` ?? `sessionId`); kernel defaults supply `system_prompt` + `last_stable_message` breakpoints so agent-session requests use per-message markers (not top-level automatic). Raw `generate` with empty breakpoints still emits top-level automatic `cache_control: { type: "ephemeral" }`; `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
|
|
175
179
|
- OpenCode Go (`@arnilo/prism-providers/opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
|
|
180
|
+
- Hyper (`@arnilo/prism-providers/hyper`): `kind: "implicit"` on the chat route, `kind: "cache_control"` on the messages route. Model route selection follows the live catalog's pricing shape: models with explicit cache-write pricing (qwen3.6-*) default to the Anthropic route; the rest stay chat-route implicit with the write fee recorded in `cost.cacheWrite`. The messages route applies `cache_control: { type: "ephemeral" }` only to caller-selected `cache.breakpoints` (shared `applyCacheControl`, max 4) — never every block, and **never a `ttl`** (Hyper documents no TTL values). Chat route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens` (and `prompt_cache_hit_tokens`); messages route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. `parseHyperUsageCost` surfaces USD/cost/remaining Hypercredits from the OpenAI usage chunk for cost telemetry; `402 billing_error` is the drained-balance signal. Caller-gated `listHyperModels`; operator-gated `getHyperCredits`.
|
|
181
|
+
- Command Code (`@arnilo/prism-providers/commandcode`): `kind: "implicit"` on the chat route, `kind: "cache_control"` on the messages route. `claude-*` tiers (the only Anthropic-route models by server-enforced routing) default to `cache_control` with markers only on caller-selected `cache.breakpoints` max 4 — never every block, and **never a `ttl`** (the upstream TTL window is undocumented). GPT-5.6 sol/terra/luna keep the docs cache-write price in `cost.cacheWrite` but stay `implicit` until the live `prompt_cache_key` probe passes (plan 055 Task 9); all other chat-route models are implicit with `prompt_cache_hit_tokens`/`cached_tokens` mapped by the shared OpenAI usage mapping. Messages route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Billing caveats: OSS models bill at the mean per-provider price; DeepSeek off-peak rates (17h/day) with ~2× peak 01–04 & 06–10 UTC; deals (MiniMax M3, MiMo) auto-applied. Caller-gated `listCommandCodeModels`; optional ZDR (`zdr: true`) adds `x-cmd-zdr: 1` and may route to costlier upstreams or fail `422 cmd_zdr_no_providers`.
|
|
176
182
|
- Z.AI (`@arnilo/prism-providers/zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
177
183
|
- NeuralWatt (`@arnilo/prism-providers/neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
|
|
178
184
|
- Kimi (`@arnilo/prism-providers/kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
@@ -31,13 +31,15 @@ Offline conformance is mandatory for every package; credentialed probes are not
|
|
|
31
31
|
| Package | Required offline evidence | Restricted live evidence |
|
|
32
32
|
| --- | --- | --- |
|
|
33
33
|
| OpenAI | Responses serialization/stream ordering, provider-hosted authority, continuation cap/cursor, Realtime fake WebSocket caps | Standard API-key smoke; separate protected hosted-tool/Realtime entitlement probe |
|
|
34
|
-
| AI SDK | Exact 4.0.
|
|
34
|
+
| AI SDK | Exact 4.0.10/V4 gate (`4.0.3` and `4.0.4` also listed); every mapped stream part; authority, cache usage, redaction, unsupported mapping | Host-created V4 model only; no Prism credential fixture |
|
|
35
35
|
| Anthropic | Messages serialization, cache/thinking/tools, header/redaction/abort assertions | Protected `ANTHROPIC_API_KEY` smoke |
|
|
36
36
|
| Google | `generateContent` serialization, complete tool calls, media/abort/redaction assertions | Protected `GOOGLE_API_KEY` or `GEMINI_API_KEY` smoke |
|
|
37
37
|
| Kimi | Coding/Moonshot route fixtures, thinking/tool reconstruction, headers/redaction | Protected `KIMI_API_KEY` smoke |
|
|
38
38
|
| Z.AI | GLM thinking/tool-stream fixtures, implicit-cache usage, headers/redaction | Protected `ZAI_API_KEY` smoke |
|
|
39
39
|
| OpenRouter | routing/reasoning/cache-control fixture, stream/tool reconstruction, headers/redaction | Protected `OPENROUTER_API_KEY` smoke |
|
|
40
40
|
| OpenCode Go | OpenAI/Anthropic route fixture, completion proof, PDF/media boundary, headers/redaction | Protected `OPENCODE_API_KEY` smoke |
|
|
41
|
+
| Hyper | dual-route OpenAI/Anthropic fixtures, thinking/effort replay, message-delta usage at end, cache-breakpoint fixture, auth/redaction | Protected `HYPER_API_KEY` smoke |
|
|
42
|
+
| Command Code | dual-route OpenAI/Anthropic fixtures, thinking replay, no-`ttl` cache-breakpoint fixture, ZDR header ownership, 403/422 classification, auth/redaction | Protected `COMMAND_CODE_API_KEY` smoke |
|
|
41
43
|
| Alibaba | DashScope presets, Qwen thinking, image rejection/mapping, cache/usage fixture | Protected account/region host probe; no generic key fixture |
|
|
42
44
|
| Ollama | cloud/local preset, reasoning/image mapping, implicit-cache fixture | Protected cloud or host-local authenticated daemon probe; no daemon starts in tests |
|
|
43
45
|
| NeuralWatt | stream/retry/quota/telemetry fixtures, implicit-cache usage, headers/redaction | Protected `NEURALWATT_API_KEY` smoke |
|
|
@@ -47,6 +49,26 @@ Offline conformance is mandatory for every package; credentialed probes are not
|
|
|
47
49
|
|
|
48
50
|
All rows must retain bounded request/response fixtures, abort propagation, provider-owned-header precedence, and fake-secret leak assertions where the package surfaces those values. A successful fake transport proves Prism mapping, not account entitlement or vendor availability.
|
|
49
51
|
|
|
52
|
+
## Modality conformance matrix (plan 061)
|
|
53
|
+
|
|
54
|
+
One provider-neutral contract per modality, each with a capability flag, a typed error family, and an offline conformance runner from `@arnilo/prism/testing/provider-conformance`. Fake transports prove Prism mapping and typed-cap behavior only — never account entitlement or vendor availability.
|
|
55
|
+
|
|
56
|
+
| Contract / runner | Offline evidence (fake transport) | Live probe |
|
|
57
|
+
| --- | --- | --- |
|
|
58
|
+
| `EmbeddingsProvider` — `runEmbeddingsConformance` | batch-cap/empty-input typed errors, index-ordered vectors, dimensions truthfulness, usage | probe pending |
|
|
59
|
+
| `SpeechProvider` — `runSpeechConformance` | input/byte caps typed, audio bytes + mime provenance, first-chunk streaming | probe pending |
|
|
60
|
+
| `TranscriptionProvider` — `runTranscriptionConformance` | audio caps typed, SSE partials, terminal transcript | probe pending |
|
|
61
|
+
| `ImageGenerationProvider` — `runImageGenerationConformance` | prompt/byte caps typed, b64 provenance, edit route | probe pending |
|
|
62
|
+
| `VideoGenerationProvider` — `runVideoGenerationConformance` | submit/status lifecycle, terminal-only resolution, provenance | probe pending |
|
|
63
|
+
| `ModerationProvider` — `runModerationConformance` | empty/oversized typed errors, neutral category mapping + raw passthrough, score range [0,1], no local thresholds | probe pending |
|
|
64
|
+
| `BatchJobsProvider` — `runBatchJobsConformance` | empty/oversized submits typed, opaque ids, state union, `pollBatch` terminal resolution, paged results walk to exhaustion | probe pending |
|
|
65
|
+
|
|
66
|
+
Failure/cancel terminal transitions for long-running contracts (video, batch) are covered by provider fakes in the adapter suites; `pollBatch` surfaces typed `job_failed`/`job_cancelled`/`job_expired` errors. Live probes stay operator-gated like the package matrix above; ledger rows live in `docs/_evidence/modality-contracts-2026-09-03.md` marked "probe pending" until run.
|
|
67
|
+
|
|
68
|
+
## Nightly live-provider canary (plan 060)
|
|
69
|
+
|
|
70
|
+
`.github/workflows/canary-providers.yml` runs nightly (`0 3 * * *`, plus `workflow_dispatch`) against the `live-canaries` environment. Each matrix leg runs one provider's existing live suite (text completion + tool-call/structured legs, `assertNoSecretLeak` inside); legs skip without their key. The provider set is the `PRISM_CANARY_PROVIDERS` repo variable (JSON array, default `["openai"]`). Evidence is per-leg and aggregate job summaries (statuses only — responses are never logged); any failure opens a tracking issue and blocks nothing.
|
|
71
|
+
|
|
50
72
|
## Inputs / request
|
|
51
73
|
|
|
52
74
|
```ts
|
|
@@ -18,6 +18,36 @@ Use provider packages when a host wants to bundle model metadata, provider adapt
|
|
|
18
18
|
|
|
19
19
|
Do not use provider packages as a package manager, credential store, env loader, provider-specific cache implementation, or live integration runner.
|
|
20
20
|
|
|
21
|
+
### Provider inventory
|
|
22
|
+
|
|
23
|
+
<!-- generated:package-truth:providers begin -->
|
|
24
|
+
**20 provider adapters** — first-party adapters ship as `@arnilo/prism-providers/<adapter>` subpaths in one tarball (importing one never evaluates another):
|
|
25
|
+
|
|
26
|
+
| adapter package | version |
|
|
27
|
+
| --- | --- |
|
|
28
|
+
| `@arnilo/prism-providers/ai-sdk` | 0.5.1 |
|
|
29
|
+
| `@arnilo/prism-providers/alibaba` | 0.5.1 |
|
|
30
|
+
| `@arnilo/prism-providers/anthropic` | 0.5.1 |
|
|
31
|
+
| `@arnilo/prism-providers/azure` | 0.5.1 |
|
|
32
|
+
| `@arnilo/prism-providers/bedrock` | 0.5.1 |
|
|
33
|
+
| `@arnilo/prism-providers/clinepass` | 0.5.1 |
|
|
34
|
+
| `@arnilo/prism-providers/commandcode` | 0.5.1 |
|
|
35
|
+
| `@arnilo/prism-providers/deepseek` | 0.5.1 |
|
|
36
|
+
| `@arnilo/prism-providers/google` | 0.5.1 |
|
|
37
|
+
| `@arnilo/prism-providers/hyper` | 0.5.1 |
|
|
38
|
+
| `@arnilo/prism-providers/kimi` | 0.5.1 |
|
|
39
|
+
| `@arnilo/prism-providers/model-discovery` | 0.5.1 |
|
|
40
|
+
| `@arnilo/prism-providers/neuralwatt` | 0.5.1 |
|
|
41
|
+
| `@arnilo/prism-providers/ollama` | 0.5.1 |
|
|
42
|
+
| `@arnilo/prism-providers/openai` | 0.5.1 |
|
|
43
|
+
| `@arnilo/prism-providers/opencode-go` | 0.5.1 |
|
|
44
|
+
| `@arnilo/prism-providers/openrouter` | 0.5.1 |
|
|
45
|
+
| `@arnilo/prism-providers/vertex` | 0.5.1 |
|
|
46
|
+
| `@arnilo/prism-providers/xai` | 0.5.1 |
|
|
47
|
+
| `@arnilo/prism-providers/zai` | 0.5.1 |
|
|
48
|
+
<!-- generated:package-truth:providers end -->
|
|
49
|
+
|
|
50
|
+
|
|
21
51
|
### Subscription OAuth support matrix
|
|
22
52
|
|
|
23
53
|
| Package | 0.0.12 auth registration | Subscription OAuth boundary |
|
|
@@ -28,6 +58,8 @@ Do not use provider packages as a package manager, credential store, env loader,
|
|
|
28
58
|
| `@arnilo/prism-providers/clinepass` | `api_key` only | No Cline WorkOS / Cline OAuth store share. Host supplies `CLINE_API_KEY`. |
|
|
29
59
|
| `@arnilo/prism-providers/anthropic` | `api_key` only | No Claude Code/Claude.ai subscription OAuth, credential-file/setup-token import, or routing. [Anthropic requires product developers to use API keys or supported cloud providers](https://docs.anthropic.com/en/docs/claude-code/legal-and-compliance). |
|
|
30
60
|
| `@arnilo/prism-providers/google` | `api_key` only | No Gemini CLI OAuth or credential/token import. [Gemini CLI prohibits third-party OAuth piggybacking](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/tos-privacy.md); use Google AI Studio API keys. Vertex/ADC uses separate [`@arnilo/prism-providers/vertex`](providers/vertex.md). |
|
|
61
|
+
| `@arnilo/prism-providers/hyper` | `api_key` only | No subscription OAuth — Charm Hyper is pay-per-use Hypercredits; host supplies `HYPER_API_KEY` (keys start `sk-hyper-`). |
|
|
62
|
+
| `@arnilo/prism-providers/commandcode` | `api_key` only | No subscription OAuth — Command Code Go/GOAT/Pro/Max coding plans and the Provider plan all authenticate with the same Studio API key; host supplies `COMMAND_CODE_API_KEY`. |
|
|
31
63
|
| `@arnilo/prism-providers/azure` | host Entra token or Azure resource key | Workload identity via `credential` callback; endpoint host preserved ([docs](providers/azure.md)). |
|
|
32
64
|
| `@arnilo/prism-providers/bedrock` | host IAM/IRSA credentials | SigV4 over OpenAI-compatible Bedrock Runtime; region/PrivateLink preserved ([docs](providers/bedrock.md)). |
|
|
33
65
|
| `@arnilo/prism-providers/vertex` | host ADC / workload token | OpenAPI-compatible Vertex endpoint; separate from consumer Google package ([docs](providers/vertex.md)). |
|
|
@@ -58,18 +90,15 @@ export default defineProviderPackage({
|
|
|
58
90
|
|
|
59
91
|
`compat` is provider-owned inert JSON. Core does not branch on provider names or interpret vendor-specific fields.
|
|
60
92
|
|
|
61
|
-
Provider packages can also contribute auth descriptors and
|
|
93
|
+
Provider packages can also contribute auth descriptors and prompt layers without resolving credentials:
|
|
62
94
|
|
|
63
95
|
```ts
|
|
64
|
-
import { createSessionCachePolicy } from "@arnilo/prism";
|
|
65
|
-
|
|
66
96
|
api.registerAuthMethod({ provider: "demo", kind: "api_key", credentialName: "apiKey" });
|
|
67
97
|
api.registerAuthMethod({ provider: "demo", kind: "oauth", oauth: demoOAuthProvider });
|
|
68
|
-
api.registerProviderRequestPolicy(createSessionCachePolicy({ retention: "short" }));
|
|
69
98
|
api.registerSystemPromptContribution({ id: "demo-prompt", source: "package", mode: "append", text: "Use demo provider rules." });
|
|
70
99
|
```
|
|
71
100
|
|
|
72
|
-
|
|
101
|
+
Host picks provider + model + intent (messages, tools, optional `thinkingLevel` / cache retention). Prism constructs a valid wire request for every generate site it owns. Host request policies are overlays — they are never required for success. Hosts still own credentials, env objects, OAuth stores, extra headers, custom `cacheKey`, `compat`/`extra` overlays, and which prompt contributions become active. Request policies can overlay generic `ProviderRequest.options` such as `sessionId`, `cacheRetention`, `headers`, `compat`, and opaque `extra`; provider adapters map those options to provider payloads. Caller headers are extension headers only: provider adapters must apply provider-owned headers (auth, content type, session/cache/security, attribution) after caller headers so requests cannot override credentials or provider policy.
|
|
73
102
|
|
|
74
103
|
Provider request options: `ProviderRequestOptions` carries session/cache/header/compat/extra hints only. Timeouts are host-owned (`RunOptions.signal`/host abort controllers); retries are runtime-owned (`AgentConfig.retry`/`RunOptions.retry`). Provider-level timeout/retry hints were removed in 0.1.5. Provider packages should not add provider-specific retry loops unless the vendor protocol requires it and runtime retry cannot cover the failure mode.
|
|
75
104
|
|
|
@@ -83,7 +112,7 @@ Phase 12 adds explicit npm workspaces for [`@arnilo/prism-providers/openai`](pro
|
|
|
83
112
|
|
|
84
113
|
Phase 6 also adds optional [`@arnilo/prism-providers/ai-sdk`](providers/ai-sdk.md), which adapts a host-owned AI SDK `LanguageModelV4` to Prism's `AIProvider`. It joins `@arnilo/prism-providers` as the seventh adapter while remaining independent from the six HTTP implementations.
|
|
85
114
|
|
|
86
|
-
Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, `OPENCODE_API_KEY`, `DEEPSEEK_API_KEY`, `XAI_API_KEY`, or `
|
|
115
|
+
Provider live tests are real smoke tests gated by `PRISM_LIVE_PROVIDER_TESTS=1` plus the provider-specific API key (`OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `KIMI_API_KEY`, `ZAI_API_KEY`, `NEURALWATT_API_KEY`, `OPENCODE_API_KEY`, `DEEPSEEK_API_KEY`, `XAI_API_KEY`, `CLINE_API_KEY`, `HYPER_API_KEY`, or `COMMAND_CODE_API_KEY`). They cover text generation, tool-call loop behavior, abort/error paths where supported, and no-secret-leak assertions; they skip by default and never run in release verification. SuperGrok login is operator-only (`PRISM_LIVE_XAI_OAUTH=1`).
|
|
87
116
|
|
|
88
117
|
These workspaces still follow the same rule as external packages: no provider SDK dependency, catalog fetch, env scan, keychain/file credential lookup, shell auth command, OAuth login, or live provider call runs by default. `@arnilo/prism-providers/openai` now registers OpenAI Responses and OpenAI Codex providers from caller-supplied credentials only, with optional `models`/`codexModels` overrides and an opt-in `listOpenAIModels()` helper for official `GET /models` discovery. `@arnilo/prism-providers/opencode-go` now registers docs-verified OpenCode Go open coding models with dual OpenAI/Anthropic routes (`compat.route`), official default base `https://opencode.ai/zen/go/v1`, `reasoning_content`/thinking preserve, and an opt-in `listOpenCodeGoModels()` helper for official `GET /zen/go/v1/models`. `@arnilo/prism-providers/openrouter` now registers an app-controlled OpenRouter catalog with routing/`reasoning`/cache passthrough, assistant `reasoning` replay, optional top-level automatic `cache_control`, and an opt-in `listOpenRouterModels()` helper for official `GET /api/v1/models` (setup still never fetches). `@arnilo/prism-providers/zai` now registers featured GLM-5.x/4.x metadata with official `thinking`/`reasoning_effort`/`tool_stream`/`clear_thinking` mapping, Preserved Thinking `reasoning_content` replay, implicit context caching, and an opt-in `listZaiModels()` helper for OpenAI-compatible `GET /models`. `@arnilo/prism-providers/kimi` now registers Kimi Coding Anthropic-compatible behavior by default, optional callable Moonshot Open Platform Chat Completions when `includeMoonshotModels` is requested, official Coding/Open Platform featured ids, thinking/`reasoning_effort` compat mapping, and an opt-in `listKimiModels()` helper for Moonshot `GET /v1/models`. `@arnilo/prism-providers/neuralwatt` now registers static featured model metadata with NeuralWatt reasoning_effort/thinking_token_budget/chat_template_kwargs request mapping, SSE comment tolerance, an opt-in `listNeuralWattModels()` helper for explicit `/v1/models` discovery, `getNeuralWattQuota()` for on-demand account balance/usage/energy, `neuralWattEventsWithTelemetry()`/`mapNeuralWattTelemetry()` for `: energy`/`: cost` telemetry, and `classifyNeuralWattError()` for retry classification. None of these helpers run during package setup or generation. `@arnilo/prism-providers/deepseek` registers featured `deepseek-v4-flash` / `deepseek-v4-pro` with official `thinking` / `reasoning_effort`, tool-turn `reasoning_content` replay, implicit prefix cache, and caller-gated `listDeepSeekModels`. `@arnilo/prism-providers/xai` registers featured Completions (`grok-4.6` / `grok-4.3` / `grok-build-0.1`), `x-grok-conv-id`, `reasoning_content` replay, caller-gated `listXaiModels`, and host-invoked SuperGrok device-code OAuth against `auth.x.ai`. `@arnilo/prism-providers/clinepass` registers a static `cline-pass/*` catalog, stream-only Chat Completions, per-model `reasoning_effort` maps, and `api_key` only (no WorkOS, no `listClinePassModels`). `@arnilo/prism-providers/anthropic` registers native Anthropic Messages (`createAnthropicProviderPackage` / `listAnthropicModels`). `@arnilo/prism-providers/google` registers native Gemini `generateContent` streaming (`createGoogleProviderPackage` / `listGoogleModels`; Vertex identity stays in the separate package). Both follow the same zero-setup-network / host-owned credential / provider-owned-header rules; see [`docs/providers/anthropic.md`](providers/anthropic.md) and [`docs/providers/google.md`](providers/google.md).
|
|
89
118
|
|
|
@@ -94,13 +123,15 @@ Every package remains explicit, setup-zero-fetch, and late-credential-bound. `Mo
|
|
|
94
123
|
| Package | Protocol / model source | Content mapping | Stream, tools, and reasoning | Cache / canary |
|
|
95
124
|
| --- | --- | --- | --- | --- |
|
|
96
125
|
| OpenAI | Responses; featured or caller-gated `listOpenAIModels` | text, image, audio, file, document | Host and provider-hosted tools; 8-hop continuation; Realtime seam; Responses reasoning | `openai_key`; checked-in standard smoke + protected hosted/Realtime probe |
|
|
97
|
-
| AI SDK | Host `LanguageModelV4`; no Prism catalog | declared text/image/audio/file/document prompt parts (role-limited) | v4 mapping; provider-executed tool authority; host-owned reasoning | host-owned; exact 4.0.
|
|
126
|
+
| AI SDK | Host `LanguageModelV4`; no Prism catalog | declared text/image/audio/file/document prompt parts (role-limited) | v4 mapping; provider-executed tool authority; host-owned reasoning | host-owned; exact 4.0.10 matrix (`4.0.3` and `4.0.4` also listed); protected host integration |
|
|
98
127
|
| Anthropic | Messages; caller-gated list | text, image, PDF document/file | tool deltas, thinking | `cache_control`; protected API-key smoke |
|
|
99
128
|
| Google | Gemini `generateContent`; caller-gated list | text, image, audio, document/file | complete tool calls, thinking | no Prism cache marker; protected API-key smoke |
|
|
100
129
|
| Kimi | Coding Messages or opt-in Moonshot; caller-gated list | text, image, PDF document/file by route/model | tool deltas, route-native thinking replay | implicit / optional Anthropic markers; protected API-key smoke |
|
|
101
130
|
| Z.AI | OpenAI-compatible; caller-gated list | text, image | tool deltas, `reasoning_content` | implicit; protected API-key smoke |
|
|
102
131
|
| OpenRouter | OpenAI-compatible; host catalog + caller-gated list | text, image | tool deltas, reasoning replay/routing metadata | `cache_control`; protected API-key smoke |
|
|
103
132
|
| OpenCode Go | OpenAI or Anthropic route; caller-gated list | text/image OpenAI route; PDF document/file Anthropic route | tool deltas, route-native thinking | route-specific; protected API-key smoke |
|
|
133
|
+
| Hyper | OpenAI, Anthropic, or explicit-pass-through Responses route (`compat.route: "responses"` reuses the shared OpenAI Responses machinery); caller-gated list | text, image | tool deltas, route-native thinking, `reasoning_effort` | route-specific implicit / `cache_control`; protected API-key smoke |
|
|
134
|
+
| Command Code | OpenAI or Anthropic route; caller-gated list | text, image | tool deltas, route-native thinking | route-specific implicit / `cache_control`; protected API-key smoke |
|
|
104
135
|
| Alibaba | DashScope OpenAI-compatible; caller-gated list | text, image | tool deltas, Qwen thinking | implicit / optional markers; protected host probe |
|
|
105
136
|
| Ollama | Cloud/local OpenAI-compatible; caller-gated list | text, image | tool deltas, reasoning effort | implicit only; protected host/daemon probe |
|
|
106
137
|
| NeuralWatt | OpenAI-compatible; caller-gated list | text, image | tool deltas, reasoning and telemetry | implicit; protected API-key smoke |
|
|
@@ -119,6 +150,8 @@ Every first-party provider package hardens prompt-cache behavior so it cannot em
|
|
|
119
150
|
- **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
120
151
|
- **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
|
|
121
152
|
- **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
|
|
153
|
+
- **Hyper** (`kind: implicit` chat route / `cache_control` messages route): `qwen3.6-*` models default to the Anthropic route with `cache_control` markers applied only to selected breakpoints (max 4; **no `ttl`** — undocumented). DeepSeek/Kimi/GLM/Gemma/etc. stay on the chat route with implicit prefix caching; `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage. Third route: `compat.route: "responses"` opts into the Codex-style `/v1/responses` OpenAI-standard pass-through, reusing the OpenAI package's Responses machinery wholesale (stream/usage/continuation/media; errors labeled `Hyper …`); cache hints surface as OpenAI-standard `prompt_cache_key` there, never retention/options. Hypercredits: `parseHyperUsageCost` surfaces USD/cost/remaining from the usage chunk. Caller-gated `listHyperModels`, operator-gated `getHyperCredits`. Billing caveat: `402 billing_error` when Hypercredits run out.
|
|
154
|
+
- **Command Code** (`kind: implicit` chat route / `cache_control` messages route): `claude-*` tiers default to the Anthropic route with `cache_control` markers only on selected breakpoints (max 4; **no `ttl`** — undocumented); GPT-5.6/OSS models stay `implicit` with docs cache-write prices recorded in `cost.cacheWrite` (explicit `prompt_cache_key` upgrade pending live-probe verification). `prompt_cache_hit_tokens` chat-route and `cache_read_input_tokens`/`cache_creation_input_tokens` messages-route map to cache usage. ZDR opt-in (`zdr: true` → `x-cmd-zdr: 1`) may route to costlier upstreams or fail `422 cmd_zdr_no_providers`. Billing caveats: OSS models bill at the mean per-provider price; DeepSeek off-peak 17h/day, peak 2×; deals auto-applied. Caller-gated `listCommandCodeModels`.
|
|
122
155
|
- **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
|
|
123
156
|
- **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
|
|
124
157
|
- **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
|
|
@@ -258,10 +291,10 @@ Provider package manifest contribution and the generic request options a provide
|
|
|
258
291
|
|
|
259
292
|
## Implementation example
|
|
260
293
|
|
|
261
|
-
Wire a provider package with model metadata
|
|
294
|
+
Wire a provider package with model metadata through the extension kernel. Session correlation and cache defaults are kernel-constructed — do not register `createSessionCachePolicy` for request success:
|
|
262
295
|
|
|
263
296
|
```ts
|
|
264
|
-
import { createExtensionKernel, defineProviderPackage
|
|
297
|
+
import { createExtensionKernel, defineProviderPackage } from "@arnilo/prism";
|
|
265
298
|
|
|
266
299
|
const pkg = defineProviderPackage({
|
|
267
300
|
name: "demo-provider",
|
|
@@ -276,7 +309,6 @@ const pkg = defineProviderPackage({
|
|
|
276
309
|
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, currency: "USD" },
|
|
277
310
|
compat: { vendorSpecific: true },
|
|
278
311
|
});
|
|
279
|
-
api.registerProviderRequestPolicy(createSessionCachePolicy({ retention: "short" }));
|
|
280
312
|
},
|
|
281
313
|
});
|
|
282
314
|
|
|
@@ -286,13 +318,13 @@ await kernel.load([pkg]);
|
|
|
286
318
|
|
|
287
319
|
## Extension and configuration notes
|
|
288
320
|
|
|
289
|
-
- Hosts
|
|
290
|
-
|
|
291
|
-
them.
|
|
292
|
-
- `createSessionCachePolicy()`
|
|
293
|
-
|
|
321
|
+
- Hosts own credentials, env objects, OAuth stores, extra headers, custom `cacheKey`,
|
|
322
|
+
`compat`/`extra`, and which prompt contributions become active; the package only *declares*
|
|
323
|
+
them. Prism constructs session/cache/thinking options for owned generate sites.
|
|
324
|
+
- `createSessionCachePolicy()` is an optional host overlay (`provider_request` cache policy hook)
|
|
325
|
+
that sets generic `ProviderRequest.options`
|
|
294
326
|
(`sessionId`, `cacheKey`, `cacheRetention`, `headers`, opaque `extra`) before
|
|
295
|
-
`AIProvider.generate()`; provider adapters map those options to provider payloads.
|
|
327
|
+
`AIProvider.generate()`; provider adapters map those options to provider payloads. Not required for request success.
|
|
296
328
|
- `ModelConfig.cache` is the generic cache capability metadata documented in [Model registry](model-registry.md); `ModelConfig.compat` remains provider-owned inert JSON for behavior that has no generic field yet.
|
|
297
329
|
- `ModelConfig.compat` is provider-owned inert JSON: cache policy overrides,
|
|
298
330
|
reasoning/thinking formats, and provider-specific usage mapping live there
|
|
@@ -310,7 +342,7 @@ await kernel.load([pkg]);
|
|
|
310
342
|
- Provider-specific behavior belongs in provider packages, not Prism core.
|
|
311
343
|
- Adapter serializers should preserve Prism content blocks (text, thinking, tool_call, tool_result, and image when the model declares image input) in provider-native request shape, or fail explicitly when a block is unsupported.
|
|
312
344
|
- Adapter header merging must put caller-supplied `ProviderRequest.options.headers` first and provider-owned headers last. Caller headers may add non-owned headers, but cannot replace resolved credentials, content type, session/cache/security headers, or provider attribution headers.
|
|
313
|
-
- For allow-list/residency/budget/circuit selection before resolve, use optional `@arnilo/prism-model-router` over `createProviderResolver` — do not fork provider packages for governance.
|
|
345
|
+
- For allow-list/residency/budget/circuit selection before resolve, use optional `@arnilo/prism-core/governance/model-router` over `createProviderResolver` — do not fork provider packages for governance.
|
|
314
346
|
- Enterprise cloud adapters (`azure` / `bedrock` / `vertex`) stay separate from consumer Anthropic/Google packages and authenticate only through host credential callbacks.
|
|
315
347
|
|
|
316
348
|
## Manifest declarations
|
|
@@ -251,7 +251,7 @@ export interface ProviderTurnMetadata {
|
|
|
251
251
|
// - provider_turn_finished { sessionId, runId, turn, metadata, usage?, error? }
|
|
252
252
|
```
|
|
253
253
|
|
|
254
|
-
Optional package `@arnilo/prism-observability
|
|
254
|
+
Optional package `@arnilo/prism-core/governance/observability` subscribes via middleware + agent events. **Default:** content redacted/absent; high-cardinality IDs are span attributes, not metric labels. NeuralWatt `neuralwatt:telemetry` events remain package-local; the adapter may forward numeric energy/cost fields.
|
|
255
255
|
|
|
256
256
|
## Migration conformance fixtures
|
|
257
257
|
|
|
@@ -6,15 +6,17 @@ Provider request policies are small host/package hooks that can adjust `Provider
|
|
|
6
6
|
|
|
7
7
|
Public helpers:
|
|
8
8
|
|
|
9
|
+
- `applyDefaultProviderRequestOptions(request, { sessionId, thinkingLevel })` is the kernel constructor. Fill-if-missing `sessionId` / `cacheKey`; default cache breakpoints + `cacheRetention: "short"` when `model.cache.kind` is `cache_control` or `explicitBreakpoints` is true; optional `thinkingLevel` patches via `applyThinkingLevelForModel` after those fills. Host values win. Agent sessions, observational-memory workers, and LLM compaction already call it.
|
|
10
|
+
- `ProviderRequirementError` (`ERR_PRISM_PROVIDER_REQUIREMENT`) is the fail-fast typed error adapters throw when a mandatory option is still missing (OpenCode Go `sessionId` → `x-opencode-session`). Thrown before fetch; messages contain no request bodies or secrets.
|
|
9
11
|
- `createProviderRequestPolicyChain(policies)` runs policies in order.
|
|
10
|
-
- `createSessionCachePolicy(options)` sets legacy `cacheKey` / `cacheRetention` aliases from `sessionId`.
|
|
12
|
+
- `createSessionCachePolicy(options)` is a host overlay that sets legacy `cacheKey` / `cacheRetention` aliases from `sessionId`. Not required for request success.
|
|
11
13
|
- `mergeProviderRequestOptions(base, patch)` merges request options, including structured `cache` hints.
|
|
12
14
|
|
|
13
15
|
## When to use it
|
|
14
16
|
|
|
15
|
-
Use
|
|
17
|
+
Use `createAgent({ thinkingLevel })` / `session.run(input, { thinkingLevel })` for session intent. Use `applyDefaultProviderRequestOptions` on custom `provider.generate` sites. Use provider request policies only as overlays (custom `cacheKey`, extra headers, `compat`/`extra`) — never to make a request valid.
|
|
16
18
|
|
|
17
|
-
Do not use request policies to resolve credentials, read env vars, perform OAuth refresh, fetch model lists, or override provider-owned auth/session/security headers.
|
|
19
|
+
Do not use request policies to resolve credentials, read env vars, perform OAuth refresh, fetch model lists, or override provider-owned auth/session/security headers. Session and cache keys are correlation ids, never secrets.
|
|
18
20
|
|
|
19
21
|
## Inputs / request
|
|
20
22
|
|
|
@@ -24,9 +26,11 @@ import type { ProviderRequestPolicy, ProviderRequestPolicyContext, ProviderReque
|
|
|
24
26
|
|
|
25
27
|
| API | Input | Purpose |
|
|
26
28
|
| --- | --- | --- |
|
|
29
|
+
| `applyDefaultProviderRequestOptions(request, ctx)` | `{ sessionId?, thinkingLevel? }` | Fill-if-missing `sessionId` / `cacheKey`; cache defaults from `model.cache`; thinking patch. |
|
|
30
|
+
| `ProviderRequirementError` | `message`, `{ requirement, providerId? }` | Typed missing-requirement error. |
|
|
27
31
|
| `ProviderRequestPolicy.apply(context)` | `{ sessionId?, request, options? }` | Returns a patched request or options. |
|
|
28
32
|
| `createProviderRequestPolicyChain(policies)` | ordered policies | Applies patches in order. |
|
|
29
|
-
| `createSessionCachePolicy({ retention?, cacheKey? })` | optional cache
|
|
33
|
+
| `createSessionCachePolicy({ retention?, cacheKey? })` | optional cache overlay | Sets legacy aliases after kernel defaults. |
|
|
30
34
|
| `mergeProviderRequestOptions(base, patch)` | two option bags | Shallow merges scalars and structurally merges `cache`. |
|
|
31
35
|
|
|
32
36
|
`mergeProviderRequestOptions()` behavior:
|
|
@@ -62,12 +66,21 @@ No agent events are emitted by the policy chain itself.
|
|
|
62
66
|
|
|
63
67
|
```ts
|
|
64
68
|
import {
|
|
69
|
+
applyDefaultProviderRequestOptions,
|
|
65
70
|
createProviderRequestPolicyChain,
|
|
66
71
|
createSessionCachePolicy,
|
|
67
72
|
mergeProviderRequestOptions,
|
|
73
|
+
ProviderRequirementError,
|
|
68
74
|
type ProviderRequestPolicy,
|
|
69
75
|
} from "@arnilo/prism";
|
|
70
76
|
|
|
77
|
+
const stamped = applyDefaultProviderRequestOptions(request, {
|
|
78
|
+
sessionId: session.id,
|
|
79
|
+
thinkingLevel: "low",
|
|
80
|
+
});
|
|
81
|
+
// stamped.options.sessionId === request.options?.sessionId ?? session.id
|
|
82
|
+
// cache_control models also get default cache.breakpoints unless the host set mode/off / retention/none / explicit breakpoints
|
|
83
|
+
|
|
71
84
|
const structuredCache: ProviderRequestPolicy = {
|
|
72
85
|
name: "demo.structured-cache",
|
|
73
86
|
apply({ request }) {
|
|
@@ -93,7 +106,7 @@ const chain = createProviderRequestPolicyChain([
|
|
|
93
106
|
|
|
94
107
|
## Extension and configuration notes
|
|
95
108
|
|
|
96
|
-
Provider packages can register request policies during `defineProviderPackage().setup(api)`. Hosts decide which
|
|
109
|
+
Provider packages can register request policies during `defineProviderPackage().setup(api)`. Hosts decide which overlay policies load and in which order. Prism has no hidden provider request policy registry. Kernel construction (`applyDefaultProviderRequestOptions`) already fills session/cache/thinking from `model.cache` and run intent — package-registered policies are never auto-activated and never required for success.
|
|
97
110
|
|
|
98
111
|
Policy output should stay generic: use `ProviderRequestOptions.cache`, `headers`, `compat`, and `extra` instead of provider-name branches in core.
|
|
99
112
|
|
|
@@ -104,7 +117,7 @@ Policy output should stay generic: use `ProviderRequestOptions.cache`, `headers`
|
|
|
104
117
|
- Cache keys must never be credentials.
|
|
105
118
|
- Policy chains are O(number of policies) plus option merge cost.
|
|
106
119
|
- Policies should be pure and synchronous unless the host explicitly accepts async work.
|
|
107
|
-
- Optional `@arnilo/prism-model-router` returns a `ProviderRequestPolicy` that strips `openRouterRouting` unless governance allows it — chain it with other policies.
|
|
120
|
+
- Optional `@arnilo/prism-core/governance/model-router` returns a `ProviderRequestPolicy` that strips `openRouterRouting` unless governance allows it — chain it with other policies.
|
|
108
121
|
|
|
109
122
|
## Related APIs
|
|
110
123
|
|
package/docs/providers/ai-sdk.md
CHANGED
|
@@ -12,6 +12,7 @@ Core `@arnilo/prism` does not depend on the AI SDK.
|
|
|
12
12
|
| --- | --- | --- |
|
|
13
13
|
| `4.0.3` | `LanguageModelV4`, `specificationVersion: "v4"` | Supported and offline-tested |
|
|
14
14
|
| `4.0.4` | `LanguageModelV4`, `specificationVersion: "v4"` | Supported and offline-tested |
|
|
15
|
+
| `4.0.10` | `LanguageModelV4`, `specificationVersion: "v4"` | Current peer; supported and offline-tested |
|
|
15
16
|
|
|
16
17
|
The peer dependency is intentionally exact. `createAiSdkProvider()` reads its resolved `@ai-sdk/provider/package.json` version during setup and throws typed `AiSdkProviderError { code: "unsupported_version" }` for an unlisted version; it does not infer compatibility from a matching `"v4"` string.
|
|
17
18
|
|
|
@@ -131,7 +132,7 @@ See [Provider caching](../provider-caching.md) for the cross-provider matrix.
|
|
|
131
132
|
|
|
132
133
|
## Thinking and reasoning
|
|
133
134
|
|
|
134
|
-
Reasoning effort, budgets, and provider-specific thinking controls are **host-model-owned**. Prism does not map `ThinkingLevel` into AI SDK call options
|
|
135
|
+
Reasoning effort, budgets, and provider-specific thinking controls are **host-model-owned**. Prism does not map `ThinkingLevel` into AI SDK call options: the deliberate `noop` compat family (stamped `compat.thinkingFamily: "noop"` on AI SDK models, and the `thinkingFamilyForModel` fallback) leaves the request options unchanged. Hosts configure reasoning on the AI SDK model (for example OpenAI `reasoning.effort` via AI SDK `providerOptions`) and may pass per-turn overrides through `ProviderRequestOptions.compat` / `extra`, which the adapter forwards as `providerOptions.prism`.
|
|
135
136
|
|
|
136
137
|
Stream mapping:
|
|
137
138
|
|
|
@@ -144,11 +145,23 @@ Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/provi
|
|
|
144
145
|
|
|
145
146
|
## Extension and configuration notes
|
|
146
147
|
|
|
147
|
-
- Peer dependency: `@ai-sdk/provider@4.0.
|
|
148
|
-
- First-party HTTP providers remain independent; this adapter is available directly
|
|
148
|
+
- Peer dependency: `@ai-sdk/provider@4.0.10` (matrix also lists `4.0.3` and `4.0.4`). Upgrade policy adds a matrix row and offline conformance fixture before accepting any new version.
|
|
149
|
+
- First-party HTTP providers remain independent; this adapter is available directly or through `@arnilo/prism-providers`. Installation does not select a model or invoke AI SDK.
|
|
149
150
|
- `options.compat` / `options.extra` pass through as AI SDK `providerOptions.prism`.
|
|
150
151
|
- Export helpers `toAiSdkCallOptions`, `toAiSdkPrompt`, and `mapAiSdkStream` for tests and custom hosts.
|
|
151
152
|
|
|
153
|
+
## Request construction (0.5.1)
|
|
154
|
+
|
|
155
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
156
|
+
|
|
157
|
+
| | |
|
|
158
|
+
| --- | --- |
|
|
159
|
+
| P1 session wire | none |
|
|
160
|
+
| Mandatory | no |
|
|
161
|
+
| P2 default cache | host-owned, no Prism cache fields |
|
|
162
|
+
|
|
163
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
164
|
+
|
|
152
165
|
## Security and performance notes
|
|
153
166
|
|
|
154
167
|
- Host credentials stay inside the supplied AI SDK model. The adapter never reads env keys or credential stores.
|
|
@@ -157,6 +170,17 @@ Official evidence: [Custom providers / LanguageModelV4](https://ai-sdk.dev/provi
|
|
|
157
170
|
- Unsupported content and stream parts fail closed before/at mapping; `structuredOutput.strict` is rejected because V4 cannot carry it.
|
|
158
171
|
- Pass `redactor` for direct use; agent runs apply their active redactor. Provider metadata/warnings never become prompt, tool, event, or telemetry content.
|
|
159
172
|
|
|
173
|
+
## Live probe
|
|
174
|
+
|
|
175
|
+
Runs the adapter over the real `@ai-sdk/openai` provider package (dev dependency) to prove the `LanguageModelV4` mapping on a genuine AI SDK wire:
|
|
176
|
+
|
|
177
|
+
```bash
|
|
178
|
+
PRISM_LIVE_PROVIDER_TESTS=1 OPENAI_API_KEY=... \
|
|
179
|
+
node --test packages/prism-providers/dist/ai-sdk/__tests__/live.test.js
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
`PRISM_LIVE_AISDK_MODEL` overrides the probed model (default `gpt-5.1`). Without a key the suite skips.
|
|
183
|
+
|
|
160
184
|
## Related APIs
|
|
161
185
|
|
|
162
186
|
- [Provider packages](../provider-packages.md)
|
|
@@ -157,7 +157,7 @@ reranker. The verified compatible route is workspace-dedicated only:
|
|
|
157
157
|
(`qwen3-rerank`, ≤500 documents, 4,000 tokens/item; base path `compatible-api/v1`,
|
|
158
158
|
not `compatible-mode/v1`). A future `createAlibabaReranker` over that route is
|
|
159
159
|
demand-gated: implement when a caller supplies a workspace-dedicated `baseUrl` and
|
|
160
|
-
needs rerank (structural `Reranker` shape from `@arnilo/prism-rag`, no new
|
|
160
|
+
needs rerank (structural `Reranker` shape from `@arnilo/prism-memory/rag`, no new
|
|
161
161
|
dependency). Multimodal rerank (`qwen3-vl-rerank`) is native-only and stays out.
|
|
162
162
|
|
|
163
163
|
## Outputs / response / events
|
|
@@ -245,6 +245,18 @@ await kernel.load([
|
|
|
245
245
|
- Usage accounting: `cached_tokens` → `Usage.cacheReadTokens`,
|
|
246
246
|
`cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
|
|
247
247
|
|
|
248
|
+
## Request construction (0.5.1)
|
|
249
|
+
|
|
250
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
251
|
+
|
|
252
|
+
| | |
|
|
253
|
+
| --- | --- |
|
|
254
|
+
| P1 session wire | none |
|
|
255
|
+
| Mandatory | no |
|
|
256
|
+
| P2 default cache | only if `model.cache.kind === "cache_control"` |
|
|
257
|
+
|
|
258
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
259
|
+
|
|
248
260
|
## Security and performance notes
|
|
249
261
|
|
|
250
262
|
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
|
|
@@ -264,6 +276,10 @@ await kernel.load([
|
|
|
264
276
|
- Model discovery is caller-gated and never invoked in the provider hot path.
|
|
265
277
|
- Live tests stay opt-in; default tests are network-free.
|
|
266
278
|
|
|
279
|
+
## Thinking and reasoning
|
|
280
|
+
|
|
281
|
+
Qwen hybrid models route through the `thinking_type` family: `none` maps to `enable_thinking: false`, any other declared level to `true` — there are no effort levels upstream, so declared `capabilities.thinkingLevels` for hybrid models are the on/off set (`none`–`max` used as a toggle vocabulary). Thinking-only models (`qwq*`, `*-thinking`) always think and never receive `enable_thinking: false`; `thinking_budget` stays a package-local passthrough. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
282
|
+
|
|
267
283
|
## Related APIs
|
|
268
284
|
|
|
269
285
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
@@ -77,6 +77,18 @@ api.registerProviderPackage(createAnthropicProviderPackage({ apiKey: hostKey, mo
|
|
|
77
77
|
- Live smoke: `PRISM_LIVE_PROVIDER_TESTS=1` + `ANTHROPIC_API_KEY`.
|
|
78
78
|
- Anthropic says OAuth is for purchasers' ordinary Claude Code/native-app use; developers building products must use Claude Console API keys or a supported cloud provider and may not offer Claude.ai login or route Free/Pro/Max credentials ([legal and compliance](https://docs.anthropic.com/en/docs/claude-code/legal-and-compliance)). Prism therefore has no Anthropic subscription OAuth API or token-import shortcut.
|
|
79
79
|
|
|
80
|
+
## Request construction (0.5.1)
|
|
81
|
+
|
|
82
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
83
|
+
|
|
84
|
+
| | |
|
|
85
|
+
| --- | --- |
|
|
86
|
+
| P1 session wire | `x-client-request-id` from `sessionId` |
|
|
87
|
+
| Mandatory | no |
|
|
88
|
+
| P2 default cache | `cache_control` on `system_prompt` + `last_stable_message`; `cacheRetention: "short"` |
|
|
89
|
+
|
|
90
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
91
|
+
|
|
80
92
|
## Security and performance notes
|
|
81
93
|
|
|
82
94
|
- No network during import/setup/default tests; credentials host-owned and late-bound.
|
|
@@ -84,6 +96,10 @@ api.registerProviderPackage(createAnthropicProviderPackage({ apiKey: hostKey, mo
|
|
|
84
96
|
- Media/SSRF bounds reuse `@arnilo/prism/providers/media` / transport helpers.
|
|
85
97
|
- Offline conformance: `@arnilo/prism/testing/provider-conformance`.
|
|
86
98
|
|
|
99
|
+
## Thinking and reasoning
|
|
100
|
+
|
|
101
|
+
Anthropic models route through the `output_config_effort` family: the adapter merges `compat.output_config.effort` (from `applyThinkingLevelForModel`) and the provider emits `output_config: { effort }` on Messages bodies. Declared levels per generation (from `capabilities.thinkingLevels`): Opus 4.8/4.7, Sonnet 5, Fable/Mythos 5, Opus 5 accept `low`–`max` incl. `xhigh`; Mythos Preview, Opus 4.6, Sonnet 4.6 accept `low/medium/high/max`; Opus 4.5 accepts `low/medium/high`; Haiku 4.5 declares `none`–`high` (upstream effort support undocumented — live-probe pending). Undeclared levels snap to the nearest declared level (ladder distance, ties up); values below the minimum snap up. Thinking type: `adaptive` on 4.6+/Sonnet 5/Fable/Mythos (bare `enabled` maps to `adaptive`); legacy 4.5 models get `enabled` plus a `budget_tokens` default of 10000 when absent. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
102
|
+
|
|
87
103
|
## Related APIs
|
|
88
104
|
|
|
89
105
|
- [Provider packages](../provider-packages.md): package setup + discovery contract.
|
package/docs/providers/azure.md
CHANGED
|
@@ -56,7 +56,19 @@ Opt-in live canaries: inject real `fetch` + host credential behind host CI secre
|
|
|
56
56
|
|
|
57
57
|
## Extension and configuration notes
|
|
58
58
|
|
|
59
|
-
Register via `createExtensionKernel().load([createAzureOpenAIProviderPackage(...)])`. Pair with `@arnilo/prism-model-router` for residency allow-lists on Azure regions/endpoints.
|
|
59
|
+
Register via `createExtensionKernel().load([createAzureOpenAIProviderPackage(...)])`. Pair with `@arnilo/prism-core/governance/model-router` for residency allow-lists on Azure regions/endpoints.
|
|
60
|
+
|
|
61
|
+
## Request construction (0.5.1)
|
|
62
|
+
|
|
63
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
64
|
+
|
|
65
|
+
| | |
|
|
66
|
+
| --- | --- |
|
|
67
|
+
| P1 session wire | none |
|
|
68
|
+
| Mandatory | no |
|
|
69
|
+
| P2 default cache | host-owned, no Prism cache fields |
|
|
70
|
+
|
|
71
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
60
72
|
|
|
61
73
|
## Security and performance notes
|
|
62
74
|
|
|
@@ -66,6 +78,22 @@ Register via `createExtensionKernel().load([createAzureOpenAIProviderPackage(...
|
|
|
66
78
|
- No Azure SDK dependency.
|
|
67
79
|
- Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; Azure cache policy stays host-owned, so no cache wire fields (`cache_control`, `prompt_cache_*`) are emitted even when the request carries Prism cache hints — only upstream-reported `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
68
80
|
|
|
81
|
+
## Live probe
|
|
82
|
+
|
|
83
|
+
Opt-in smoke over your real Azure OpenAI resource:
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
PRISM_LIVE_PROVIDER_TESTS=1 AZURE_OPENAI_ENDPOINT=https://<resource>.openai.azure.com \
|
|
87
|
+
AZURE_OPENAI_API_KEY=... PRISM_LIVE_AZURE_MODEL=<deployment-name> \
|
|
88
|
+
node --test packages/prism-providers/dist/azure/__tests__/live.test.js
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
The deployment name is the model knob (`AZURE_OPENAI_DEPLOYMENT` also works). Without any of these variables the suite skips — it never fails.
|
|
92
|
+
|
|
93
|
+
## Thinking and reasoning
|
|
94
|
+
|
|
95
|
+
Azure OpenAI deployments use the OpenAI-compatible wire: the provider forwards sanitized thinking compat (`reasoning_effort` with `effort`/`reasoningEffort` aliases, or a `reasoning` object with `effort` + preserved `summary`) via its body builder, snapping effort to the model's declared levels (gpt-5.1 → `none/low/medium/high`, ties up). Unrecognized compat keys are dropped. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
96
|
+
|
|
69
97
|
## Related APIs
|
|
70
98
|
|
|
71
99
|
- [OpenAI-compatible provider](openai-compatible.md)
|