@arnilo/prism 0.4.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +33 -1
- package/README.md +23 -20
- package/dist/agent-run-state.d.ts +1 -2
- package/dist/agent-run-state.js +0 -3
- package/dist/agent-session/session/assemble.d.ts +6 -0
- package/dist/agent-session/session/assemble.js +391 -0
- package/dist/agent-session/session/persist.d.ts +28 -0
- package/dist/agent-session/session/persist.js +166 -0
- package/dist/agent-session/session/provider-round.d.ts +6 -0
- package/dist/agent-session/session/provider-round.js +231 -0
- package/dist/agent-session/session/tool-round.d.ts +31 -0
- package/dist/agent-session/session/tool-round.js +473 -0
- package/dist/agent-session/session/types.d.ts +115 -0
- package/dist/agent-session/session/types.js +5 -0
- package/dist/agent-session/session.d.ts +49 -43
- package/dist/agent-session/session.js +11 -1177
- package/dist/capture.d.ts +63 -0
- package/dist/capture.js +67 -0
- package/dist/cli-init.d.ts +18 -2
- package/dist/cli-init.js +2 -7
- package/dist/cli-runner.d.ts +2 -2
- package/dist/cli-runner.js +45 -9
- package/dist/content.d.ts +3 -3
- package/dist/content.js +3 -1
- package/dist/contracts-core/agent.d.ts +2 -0
- package/dist/contracts-core/batch.d.ts +97 -0
- package/dist/contracts-core/batch.js +65 -0
- package/dist/contracts-core/content.d.ts +72 -1
- package/dist/contracts-core/embeddings.d.ts +30 -0
- package/dist/contracts-core/embeddings.js +17 -0
- package/dist/contracts-core/images.d.ts +60 -0
- package/dist/contracts-core/images.js +17 -0
- package/dist/contracts-core/moderation.d.ts +46 -0
- package/dist/contracts-core/moderation.js +34 -0
- package/dist/contracts-core/speech.d.ts +39 -0
- package/dist/contracts-core/speech.js +17 -0
- package/dist/contracts-core/transcription.d.ts +48 -0
- package/dist/contracts-core/transcription.js +17 -0
- package/dist/contracts-core/video.d.ts +61 -0
- package/dist/contracts-core/video.js +17 -0
- package/dist/contracts-core.d.ts +7 -0
- package/dist/contracts-core.js +7 -0
- package/dist/index.d.ts +5 -3
- package/dist/index.js +4 -3
- package/dist/node/agent-definitions.d.ts +1 -8
- package/dist/node/agent-definitions.js +0 -34
- package/dist/node/settings.d.ts +0 -1
- package/dist/node/settings.js +0 -5
- package/dist/pinned-fetch.js +29 -3
- package/dist/provider-events.js +3 -4
- package/dist/providers/media.d.ts +1 -2
- package/dist/providers/media.js +1 -4
- package/dist/rpc.d.ts +1 -1
- package/dist/rpc.js +4 -4
- package/dist/testing/provider-conformance.d.ts +114 -5
- package/dist/testing/provider-conformance.js +342 -0
- package/dist/testing/tool-effect-store-conformance.d.ts +0 -1
- package/dist/testing/tool-effect-store-conformance.js +0 -3
- package/dist/thinking.d.ts +48 -9
- package/dist/thinking.js +134 -8
- package/docs/0.1.0-readiness.md +3 -3
- package/docs/a2a.md +2 -2
- package/docs/acp.md +3 -3
- package/docs/ag-ui-adoption.md +1 -1
- package/docs/ag-ui.md +1 -2
- package/docs/agent-definitions.md +1 -1
- package/docs/agent-events.md +5 -5
- package/docs/agent-identity.md +13 -2
- package/docs/audit-export.md +3 -3
- package/docs/batch-jobs.md +120 -0
- package/docs/cli-rpc.md +20 -9
- package/docs/coding-agent-tools.md +19 -19
- package/docs/coding-review-and-diagnostics.md +2 -2
- package/docs/coding-security.md +4 -4
- package/docs/coding-workspaces.md +2 -2
- package/docs/computer-use-linux.md +13 -2
- package/docs/context-and-skills.md +1 -1
- package/docs/conversations.md +4 -4
- package/docs/credential-storage.md +11 -7
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/data-classification.md +1 -1
- package/docs/database-persistence.md +4 -4
- package/docs/dev-inspector.md +6 -6
- package/docs/device-adapters.md +2 -2
- package/docs/diagrams.md +1 -1
- package/docs/document-reader.md +6 -6
- package/docs/documents.md +5 -4
- package/docs/embeddings.md +112 -0
- package/docs/enterprise-postgres-state.md +7 -7
- package/docs/evaluations.md +8 -8
- package/docs/extensions.md +3 -3
- package/docs/forge-integration.md +3 -3
- package/docs/graft.md +2 -2
- package/docs/guardrails.md +1 -1
- package/docs/host-security.md +15 -15
- package/docs/image-generation.md +129 -0
- package/docs/impeccable.md +5 -3
- package/docs/index.md +60 -33
- package/docs/indexed-code-search.md +2 -2
- package/docs/language-intelligence.md +4 -4
- package/docs/live-testing.md +126 -0
- package/docs/mcp-tools.md +43 -12
- package/docs/middleware-hooks.md +1 -1
- package/docs/migrate-to-0.4.md +3 -3
- package/docs/migrate-to-0.5.md +122 -0
- package/docs/migration.md +29 -1
- package/docs/model-registry.md +38 -0
- package/docs/model-routing.md +5 -5
- package/docs/moderation.md +117 -0
- package/docs/multi-agent-patterns.md +4 -4
- package/docs/multimodal-content.md +26 -2
- package/docs/obscura.md +2 -2
- package/docs/observability.md +32 -7
- package/docs/openapi-tools.md +13 -3
- package/docs/operations.md +11 -0
- package/docs/performance.md +7 -7
- package/docs/persistence-credentials-multimodality-primitives.md +6 -6
- package/docs/policy-and-audit.md +17 -7
- package/docs/ponytail.md +1 -1
- package/docs/postgres-persistence.md +5 -5
- package/docs/process-sessions.md +2 -2
- package/docs/prompt-registry.md +7 -7
- package/docs/provider-caching.md +4 -0
- package/docs/provider-conformance.md +23 -1
- package/docs/provider-packages.md +39 -3
- package/docs/provider-primitives.md +1 -1
- package/docs/provider-request-policies.md +1 -1
- package/docs/providers/ai-sdk.md +15 -3
- package/docs/providers/alibaba.md +5 -1
- package/docs/providers/anthropic.md +4 -0
- package/docs/providers/azure.md +17 -1
- package/docs/providers/bedrock.md +15 -0
- package/docs/providers/clinepass.md +4 -0
- package/docs/providers/commandcode.md +253 -0
- package/docs/providers/deepseek.md +4 -0
- package/docs/providers/google.md +4 -0
- package/docs/providers/hyper.md +284 -0
- package/docs/providers/kimi.md +4 -0
- package/docs/providers/neuralwatt.md +4 -0
- package/docs/providers/ollama.md +15 -0
- package/docs/providers/openai-compatible.md +4 -0
- package/docs/providers/openai.md +4 -0
- package/docs/providers/opencode-go.md +4 -0
- package/docs/providers/openrouter.md +5 -1
- package/docs/providers/vertex.md +16 -0
- package/docs/providers/xai.md +4 -0
- package/docs/providers/zai.md +4 -0
- package/docs/rag.md +26 -4
- package/docs/release-and-install.md +103 -46
- package/docs/resource-loading.md +1 -1
- package/docs/runs-and-usage.md +14 -2
- package/docs/server.md +5 -5
- package/docs/settings-auth-trust-security.md +7 -5
- package/docs/sheets.md +2 -2
- package/docs/speech.md +126 -0
- package/docs/sqlite-persistence.md +4 -4
- package/docs/supervisors.md +3 -3
- package/docs/thinking-and-reasoning.md +93 -60
- package/docs/tool-conformance.md +1 -1
- package/docs/tool-execution-primitives.md +8 -8
- package/docs/tools.md +4 -4
- package/docs/web-tools.md +1 -1
- package/docs/wiki.md +1 -1
- package/docs/work-artifacts-and-review.md +17 -6
- package/docs/work-connectors.md +4 -4
- package/docs/work-tools.md +5 -5
- package/docs/workflow-orchestration-primitives.md +11 -11
- package/docs/workflows.md +5 -5
- package/package.json +11 -8
- package/templates/init/providers.json +24 -8
- package/docs/antigravity-agent.md +0 -207
|
@@ -0,0 +1,253 @@
|
|
|
1
|
+
# Command Code provider package
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-providers/commandcode` provides explicit, side-effect-free setup
|
|
6
|
+
for the [Command Code Provider API](https://commandcode.ai/docs/provider) — an
|
|
7
|
+
aggregator exposing every top commercial and open model through OpenAI- and
|
|
8
|
+
Anthropic-compatible endpoints, billed at cost with deals auto-applied. The
|
|
9
|
+
package dual-routes by `ModelConfig.compat.route`:
|
|
10
|
+
|
|
11
|
+
| Route | Endpoint | Official model families |
|
|
12
|
+
| --- | --- | --- |
|
|
13
|
+
| `"openai"` (default) | `POST {baseUrl}/chat/completions` | everything except `claude-*` (GPT-5.6, DeepSeek, Kimi, GLM, MiniMax, Qwen, MiMo, Gemini flash, Grok) |
|
|
14
|
+
| `"anthropic"` | `POST {baseUrl}/messages` | `claude-*` tiers (Opus/Sonnet/Fable/Haiku) |
|
|
15
|
+
|
|
16
|
+
Default base URL is the official Provider API root:
|
|
17
|
+
|
|
18
|
+
```txt
|
|
19
|
+
https://api.commandcode.ai/provider/v1
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
Authentication: `Authorization: Bearer <key>` on the chat route,
|
|
23
|
+
`x-api-key` + `anthropic-version: 2023-06-01` on the messages route (Claude
|
|
24
|
+
Code compatibility). The same key authenticates the CLI and the API.
|
|
25
|
+
|
|
26
|
+
## When to use it
|
|
27
|
+
|
|
28
|
+
Use it when a host app wants Command Code models through Prism's `AgentSession`
|
|
29
|
+
runtime: dual-route serialization, `cache_control` breakpoints on Claude
|
|
30
|
+
models, implicit caching elsewhere, reasoning replay, caller-gated model
|
|
31
|
+
discovery, and optional zero-data-retention (`zdr: true` → provider-owned
|
|
32
|
+
`x-cmd-zdr: 1`, which routes only through ZDR-capable upstreams).
|
|
33
|
+
|
|
34
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches,
|
|
35
|
+
or real-network tests (live probes are operator-gated, see below).
|
|
36
|
+
|
|
37
|
+
## Inputs / request
|
|
38
|
+
|
|
39
|
+
```ts
|
|
40
|
+
import {
|
|
41
|
+
createCommandCodeProviderPackage,
|
|
42
|
+
listCommandCodeModels,
|
|
43
|
+
} from "@arnilo/prism-providers/commandcode";
|
|
44
|
+
|
|
45
|
+
createCommandCodeProviderPackage(options: CommandCodeProviderPackageOptions): ProviderPackage
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
| Field | Type | Purpose |
|
|
49
|
+
| --- | --- | --- |
|
|
50
|
+
| `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
|
|
51
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
52
|
+
| `baseUrl` | `string` | Overrides official `https://api.commandcode.ai/provider/v1`. |
|
|
53
|
+
| `models` | `readonly ModelConfig[]` | Overrides featured `commandCodeModels` defaults. |
|
|
54
|
+
| `zdr` | `boolean` | Enforce zero data retention (`x-cmd-zdr: 1`). May route to costlier upstreams or fail `422 cmd_zdr_no_providers`. |
|
|
55
|
+
|
|
56
|
+
`ProviderRequest.options.cache.breakpoints` select messages-route
|
|
57
|
+
`cache_control` markers for Claude models (max 4, no `ttl`).
|
|
58
|
+
|
|
59
|
+
## Outputs / response / events
|
|
60
|
+
|
|
61
|
+
| Surface | Behavior |
|
|
62
|
+
| --- | --- |
|
|
63
|
+
| Provider stream | Prism text, thinking, tool-call delta/final, `usage`, `done`, redacted `error`. |
|
|
64
|
+
| Stream completion | `done` only on completion evidence — chat route: `[DONE]` marker plus terminal `finish_reason`; messages route: `message_stop`. Truncated streams end with terminal `error`. |
|
|
65
|
+
| OpenAI thinking | `delta.reasoning_content` → thinking deltas; replay via top-level `reasoning_content` when `preserveThinking` (default), never folded into text. |
|
|
66
|
+
| Anthropic thinking | `thinking_delta` → thinking deltas; replay via Anthropic thinking blocks when `preserveThinking`. |
|
|
67
|
+
| Usage | Standard tokens + cache read/write per route; mapped through the shared OpenAI/Anthropic usage mappings. |
|
|
68
|
+
| Auth method | `api_key` for `commandcode`, credential name `apiKey`. |
|
|
69
|
+
|
|
70
|
+
## Request/response example
|
|
71
|
+
|
|
72
|
+
```json
|
|
73
|
+
{
|
|
74
|
+
"authorization": "Bearer cmd_…",
|
|
75
|
+
"content-type": "application/json"
|
|
76
|
+
}
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
Messages route instead sends provider-owned `x-api-key: <key>` and
|
|
80
|
+
`anthropic-version: 2023-06-01`. All provider-owned headers are applied after
|
|
81
|
+
caller headers and cannot be overridden.
|
|
82
|
+
|
|
83
|
+
Chat-route body (thinking passthrough + preserved reasoning):
|
|
84
|
+
|
|
85
|
+
```json
|
|
86
|
+
{
|
|
87
|
+
"model": "Qwen/Qwen3.8-Flash",
|
|
88
|
+
"stream": true,
|
|
89
|
+
"stream_options": { "include_usage": true },
|
|
90
|
+
"max_tokens": 512,
|
|
91
|
+
"messages": [
|
|
92
|
+
{
|
|
93
|
+
"role": "assistant",
|
|
94
|
+
"tool_calls": [{ "id": "call_1", "type": "function", "function": { "name": "lookup", "arguments": "{}" } }],
|
|
95
|
+
"reasoning_content": "plan the lookup"
|
|
96
|
+
}
|
|
97
|
+
]
|
|
98
|
+
}
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
## Implementation example
|
|
102
|
+
|
|
103
|
+
```ts
|
|
104
|
+
import { createExtensionKernel } from "@arnilo/prism";
|
|
105
|
+
import { createCommandCodeProviderPackage } from "@arnilo/prism-providers/commandcode";
|
|
106
|
+
|
|
107
|
+
const kernel = createExtensionKernel();
|
|
108
|
+
await kernel.load([createCommandCodeProviderPackage({ apiKey: process.env.COMMAND_CODE_API_KEY })]);
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Caller-gated live catalog (never runs during package setup):
|
|
112
|
+
|
|
113
|
+
```ts
|
|
114
|
+
const models = await listCommandCodeModels({ fetch }); // public endpoint, no auth needed
|
|
115
|
+
await kernel.load([createCommandCodeProviderPackage({ apiKey: process.env.COMMAND_CODE_API_KEY, models })]);
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
## Featured models and routes
|
|
119
|
+
|
|
120
|
+
Featured `commandCodeModels` is a curated 38-model bootstrap catalog: ids and
|
|
121
|
+
context windows from the live `GET /provider/v1/models` snapshot (2026-09, 67
|
|
122
|
+
ids), USD-per-million-token pricing from the docs table
|
|
123
|
+
(<https://commandcode.ai/docs/resources/pricing-limits>). `compat.pricing_source`
|
|
124
|
+
records caveats: open-source models bill at the **mean per-provider price**;
|
|
125
|
+
DeepSeek rates are **off-peak** (17h/day; peak 2× during 01–04 & 06–10 UTC);
|
|
126
|
+
deals (MiniMax M3 −50%, MiMo −98/99%) are already applied upstream. Custom
|
|
127
|
+
pricing metadata is always stripped before the wire.
|
|
128
|
+
|
|
129
|
+
| Model family | Route | Cache kind |
|
|
130
|
+
| --- | --- | --- |
|
|
131
|
+
| `claude-opus-5/4-8/4-7`, `claude-sonnet-5/4-6`, `claude-fable-5-1/5`, `claude-haiku-4-5` | `anthropic` | `cache_control` (max 4 breakpoints, no `ttl` — undocumented) |
|
|
132
|
+
| `gpt-5.6-sol/terra/luna` | `openai` | `implicit` (docs cache-write price recorded in `cost.cacheWrite`; explicit-key upgrade gated on live probe — Task 9) |
|
|
133
|
+
| `deepseek/*`, Kimi, GLM, MiniMax, Qwen, MiMo, Gemini flash, Grok | `openai` | `implicit` |
|
|
134
|
+
|
|
135
|
+
## Model discovery
|
|
136
|
+
|
|
137
|
+
```txt
|
|
138
|
+
GET https://api.commandcode.ai/provider/v1/models
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
Public endpoint — works without authentication and emits no auth header when no
|
|
142
|
+
key resolves. `listCommandCodeModels({ fetch?, baseUrl?, apiKey?, signal?,
|
|
143
|
+
headers? })` maps each `{ id, name, context_length }` entry to `ModelConfig`:
|
|
144
|
+
route from id (`claude-*` → anthropic), context window from the endpoint, and
|
|
145
|
+
featured docs metadata (cost/cache kind) applied when the id matches a curated
|
|
146
|
+
entry. The endpoint carries no pricing or capabilities — unknown ids get
|
|
147
|
+
route-derived cache kind and no cost. Discovery is **caller-gated** — setup
|
|
148
|
+
performs zero fetches.
|
|
149
|
+
|
|
150
|
+
## Thinking / reasoning
|
|
151
|
+
|
|
152
|
+
| Surface | Behavior |
|
|
153
|
+
| --- | --- |
|
|
154
|
+
| OpenAI route stream | `reasoning_content` → thinking deltas |
|
|
155
|
+
| OpenAI route replay | thinking blocks → top-level `reasoning_content` when `preserveThinking`; never folded into text |
|
|
156
|
+
| Anthropic route stream | `thinking_delta` → thinking deltas |
|
|
157
|
+
| Anthropic route replay | thinking blocks when `preserveThinking` |
|
|
158
|
+
|
|
159
|
+
Owned compat keys (`route`, `preserveThinking`, `pricing_source`) are stripped
|
|
160
|
+
before opaque compat spread so resolved values win.
|
|
161
|
+
|
|
162
|
+
## Extension and configuration notes
|
|
163
|
+
|
|
164
|
+
- Hosts choose base URL, model list, credential source, `fetch` impl, and ZDR.
|
|
165
|
+
- Route selection is explicit via `compat.route` (`"anthropic"` for `claude-*`
|
|
166
|
+
ids, default `"openai"`). Sending a model to the wrong endpoint 400s.
|
|
167
|
+
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
168
|
+
|
|
169
|
+
### Cache and session behavior
|
|
170
|
+
|
|
171
|
+
- The chat route sends **no** `cache_control` fields; it relies on OpenAI-style
|
|
172
|
+
implicit caching (upstream provider behavior, passed through). Read tokens
|
|
173
|
+
map from `prompt_tokens_details.cached_tokens` / `cache_write_tokens` /
|
|
174
|
+
`prompt_cache_hit_tokens`.
|
|
175
|
+
- The messages route applies `cache_control: { type: "ephemeral" }` markers
|
|
176
|
+
only to caller-selected `cache.breakpoints` (shared `applyCacheControl()`
|
|
177
|
+
helper) on the last content block of each selected message. A
|
|
178
|
+
`system_prompt` breakpoint serializes `system` as marked text blocks (plain
|
|
179
|
+
string otherwise). Caching is enabled unless disabled
|
|
180
|
+
(`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
|
|
181
|
+
`ModelConfig.cache.kind: "cache_control"`.
|
|
182
|
+
- **No `ttl` is ever emitted**: the upstream TTL window is undocumented on the
|
|
183
|
+
Provider API; `cacheRetention: "long"` must not produce a marker the gateway
|
|
184
|
+
may reject.
|
|
185
|
+
- Usage accounting per route: chat route maps
|
|
186
|
+
`prompt_tokens_details.cached_tokens`/`cache_write_tokens` (and
|
|
187
|
+
`prompt_cache_hit_tokens`); messages route maps
|
|
188
|
+
`cache_read_input_tokens`/`cache_creation_input_tokens`.
|
|
189
|
+
- Session identity is simple: no session header is emitted (undocumented).
|
|
190
|
+
|
|
191
|
+
### Live-verified mapping (findings ledger)
|
|
192
|
+
|
|
193
|
+
The following claims are encoded as operator-gated probes in
|
|
194
|
+
`packages/prism-providers/src/commandcode/__tests__/live.test.ts`. Each probe's
|
|
195
|
+
assertion encodes the documented claim, so a probe failure **is** the finding;
|
|
196
|
+
record the outcome here and adjust the mapping. Status: **pending operator
|
|
197
|
+
run** (no key in CI):
|
|
198
|
+
|
|
199
|
+
| # | Claim (documented) | Probe | Status |
|
|
200
|
+
| --- | --- | --- | --- |
|
|
201
|
+
| 1 | Warm chat-route replay reports cached tokens (implicit caching passes through) | `live_chat_route_reports_cached_tokens_on_warm_prefix_replay` | pending |
|
|
202
|
+
| 2 | `cache_control` on messages reports `cache_creation_input_tokens` on the creating call | `live_messages_route_cache_control_reports_creation_and_read_tokens` | pending |
|
|
203
|
+
| 3 | Same-prefix warm replay reads the created cache entry (TTL ≥ one request) | same probe (warm leg) | pending |
|
|
204
|
+
| 4 | OpenAI `prompt_cache_key` is accepted and honored for GPT-5.6 (explicit caching) — decides Task 9 | `live_gpt56_prompt_cache_key_passthrough_probe` | pending — Task 9 closed gated: **pass** → upgrade `gpt-5.6-*` to the OpenAI explicit mapping (`promptCacheKey`/`promptCacheOptions`/`applyPromptCacheBreakpoints` via `@arnilo/prism-providers/openai`, verified exported); **fail** (400/ignored/warm replay shows no cached tokens) → verified-negative, keep `implicit`, record here |
|
|
205
|
+
| 5 | OpenAI `reasoning_effort` is accepted (200) on the chat route | `live_reasoning_effort_is_accepted_on_chat_route` | pending |
|
|
206
|
+
| 6 | ZDR requests route (done) or fail `422 cmd_zdr_no_providers` when no ZDR-capable upstream exists | `live_zdr_route_probe_is_opt_in_and_routable` | pending |
|
|
207
|
+
|
|
208
|
+
Run the gate:
|
|
209
|
+
|
|
210
|
+
```sh
|
|
211
|
+
PRISM_LIVE_PROVIDER_TESTS=1 COMMAND_CODE_API_KEY=cmd_... \
|
|
212
|
+
npm run test --workspace=@arnilo/prism-providers/commandcode
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
## Security and performance notes
|
|
216
|
+
|
|
217
|
+
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers.
|
|
218
|
+
- No network calls during import, setup, build, or default tests.
|
|
219
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
220
|
+
- API keys are resolved per request from caller-supplied values or resolvers
|
|
221
|
+
and redacted from errors (upstream error bodies may carry the upstream
|
|
222
|
+
provider's message — always redacted).
|
|
223
|
+
- `403 upgrade_required` (Go plan — no API access) and
|
|
224
|
+
`422 cmd_zdr_no_providers` are non-retryable; `429` and `5xx` are retryable
|
|
225
|
+
with `retry-after` surfaced as `retry_after_ms`; `400/401/422` are
|
|
226
|
+
non-retryable.
|
|
227
|
+
- Caller headers cannot override provider-owned headers (`content-type`,
|
|
228
|
+
`authorization` on chat, `x-api-key`/`anthropic-version` on messages,
|
|
229
|
+
`x-cmd-zdr` when ZDR is opted in).
|
|
230
|
+
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus
|
|
231
|
+
`COMMAND_CODE_API_KEY`; default tests are network-free.
|
|
232
|
+
|
|
233
|
+
## Official evidence
|
|
234
|
+
|
|
235
|
+
- Command Code Provider API docs: `https://commandcode.ai/docs/provider`
|
|
236
|
+
- Pricing & limits (per-model USD, deals, off-peak): `https://commandcode.ai/docs/resources/pricing-limits`
|
|
237
|
+
- Live `GET https://api.commandcode.ai/provider/v1/models` snapshot (2026-09) — 67 ids, context windows
|
|
238
|
+
- Probe ledger above pending operator-gated live run; row 4 decides plan 055 Task 9
|
|
239
|
+
(explicit GPT-5.6 caching upgrade) — see `plans/055-First-Class-Hyper-And-Command-Code-Providers.md`
|
|
240
|
+
|
|
241
|
+
## Thinking and reasoning
|
|
242
|
+
|
|
243
|
+
Command Code models carry provenance-commented level tables: `claude-*` → `output_config_effort` (Anthropic-route effort sets mirroring the native Anthropic package), `gpt-5.6*` → `openai_reasoning` (`none`–`xhigh`), `deepseek-v4`/`kimi-k3`/`glm-5.3` → `reasoning_effort` (`low/high/max`), `glm-5.2` → `reasoning_effort` (`low`–`max`), Kimi-K2.x/MiniMax/Qwen → `thinking_type`, gemini-3.x → `noop` (the gateway chat route has no `thinking_level` wire, so no levels are declared), mimo/unknown → passthrough. Effort snaps to declared sets on both routes. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
244
|
+
|
|
245
|
+
## Related APIs
|
|
246
|
+
|
|
247
|
+
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
248
|
+
`ModelConfig`, discovery contract, request/cache policies.
|
|
249
|
+
- [Thinking and reasoning](../thinking-and-reasoning.md): per-turn `ThinkingLevel` → compat families.
|
|
250
|
+
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
251
|
+
`resolveCredentialValue`, `redactSecrets`.
|
|
252
|
+
- [Provider caching](../provider-caching.md): per-provider cache behavior matrix.
|
|
253
|
+
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
@@ -129,6 +129,10 @@ await session.prompt("Plan the refactor", {
|
|
|
129
129
|
- One POST per generate. No provider retry loop.
|
|
130
130
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus `DEEPSEEK_API_KEY`.
|
|
131
131
|
|
|
132
|
+
## Thinking and reasoning
|
|
133
|
+
|
|
134
|
+
DeepSeek models declare `low/high/max` and stamp `reasoning_effort`; the wire table maps `medium`/`xhigh`→`high`, `none`/`minimal` stop thinking. `thinking.type: "enabled"/"disabled"` stays available (thinking on by default, `high`); a request-level `reasoning_effort: none` stops thinking only when no explicit `thinking` switch was sent. Tool turns must replay `reasoning_content` or the API returns 400. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
135
|
+
|
|
132
136
|
## Related APIs
|
|
133
137
|
|
|
134
138
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
package/docs/providers/google.md
CHANGED
|
@@ -80,6 +80,10 @@ api.registerProviderPackage(createGoogleProviderPackage({ apiKey: hostKey, model
|
|
|
80
80
|
- Media bounds reuse shared provider media helpers; tool args arrive complete per chunk (no partial JSON reconstruction required).
|
|
81
81
|
- Offline conformance: `@arnilo/prism/testing/provider-conformance`.
|
|
82
82
|
|
|
83
|
+
## Thinking and reasoning
|
|
84
|
+
|
|
85
|
+
Google models route through the `google` family: the adapter merges `compat.thinkingLevel` and the provider emits `generationConfig.thinkingConfig`. Gemini 3.x models use `thinkingLevel` with declared per-model sets: 3.6/3.5-flash and 3-flash-preview accept `minimal`–`high`; 3.1-pro accepts `low/medium/high` (default `high`); 3-pro accepts `low/high`. Gemini 2.5 models are budget-only (`compat.thinkingBudgetRange`): 2.5-pro `128–32768` (cannot disable), 2.5-flash/flash-lite `0–24576` (`0` disables). `none` on a budget-only model maps to the range minimum (`thinkingBudget: 0` where disabling is supported, `128` where not); non-none levels are dropped on budget-only models. Declared level sets snap via nearest-declared (ties up), so `none`/`minimal` on 3.1-pro snap up to `low`. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
86
|
+
|
|
83
87
|
## Related APIs
|
|
84
88
|
|
|
85
89
|
- [Google Vertex AI](vertex.md): enterprise ADC/workload-identity package (separate from this consumer API-key package).
|
|
@@ -0,0 +1,284 @@
|
|
|
1
|
+
# Hyper provider package
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-providers/hyper` provides explicit, side-effect-free setup for
|
|
6
|
+
[Charm Hyper](https://hyper.charm.land) — a pay-per-use reasoning-model gateway
|
|
7
|
+
billed in Hypercredits (1 HC = $0.05). The package routes by `ModelConfig.compat.route`:
|
|
8
|
+
|
|
9
|
+
| Route | Endpoint | Official model families |
|
|
10
|
+
| --- | --- | --- |
|
|
11
|
+
| `"openai"` (default) | `POST {baseUrl}/chat/completions` | most models (DeepSeek, Kimi, GLM, Gemma, …) |
|
|
12
|
+
| `"anthropic"` | `POST {baseUrl}/messages` | `qwen3.6-*` (Anthropic-shaped explicit-write cache pricing) |
|
|
13
|
+
| `"responses"` (explicit opt-in) | `POST {baseUrl}/responses` | OpenAI-standard pass-through; hosts bring Responses-shaped model metadata (Codex-style clients) |
|
|
14
|
+
|
|
15
|
+
Default base URL is the official API root:
|
|
16
|
+
|
|
17
|
+
```txt
|
|
18
|
+
https://hyper.charm.land/v1
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
Authentication is `Authorization: Bearer <key>` on every route; the messages
|
|
22
|
+
route additionally sends provider-owned `x-api-key` + `anthropic-version:
|
|
23
|
+
2023-06-01` headers (Claude Code compatibility). API keys start with
|
|
24
|
+
`sk-hyper-`. The responses route reuses the OpenAI package's Responses
|
|
25
|
+
machinery wholesale — body serialization, stream events, usage mapping,
|
|
26
|
+
continuation cursors, and media handling — with Hyper's base URL and auth, so
|
|
27
|
+
its wire behavior matches the OpenAI-standard pass-through Charm documents.
|
|
28
|
+
Errors on that route are labeled `Hyper …` (e.g. `Hyper request failed: 429 …`).
|
|
29
|
+
|
|
30
|
+
## When to use it
|
|
31
|
+
|
|
32
|
+
Use it when a host app wants Hyper models through Prism's `AgentSession`
|
|
33
|
+
runtime with dual-route serialization, reasoning-content replay, cache-hint
|
|
34
|
+
breakpoints, and caller-gated model discovery and credit checks.
|
|
35
|
+
|
|
36
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches,
|
|
37
|
+
or real-network tests (live probes are operator-gated, see below).
|
|
38
|
+
|
|
39
|
+
## Inputs / request
|
|
40
|
+
|
|
41
|
+
```ts
|
|
42
|
+
import {
|
|
43
|
+
createHyperProviderPackage,
|
|
44
|
+
getHyperCredits,
|
|
45
|
+
listHyperModels,
|
|
46
|
+
} from "@arnilo/prism-providers/hyper";
|
|
47
|
+
|
|
48
|
+
createHyperProviderPackage(options: HyperProviderPackageOptions): ProviderPackage
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
| Field | Type | Purpose |
|
|
52
|
+
| --- | --- | --- |
|
|
53
|
+
| `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
|
|
54
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
55
|
+
| `baseUrl` | `string` | Overrides official `https://hyper.charm.land/v1`. |
|
|
56
|
+
| `models` | `readonly ModelConfig[]` | Overrides featured `hyperModels` defaults. |
|
|
57
|
+
|
|
58
|
+
`ProviderRequest.options.cache.breakpoints` select Anthropic-route
|
|
59
|
+
`cache_control` markers as documented below; `options.compat.reasoning_effort`
|
|
60
|
+
selects the per-request reasoning effort (clamped to the model's documented
|
|
61
|
+
`effortLevels`).
|
|
62
|
+
|
|
63
|
+
## Outputs / response / events
|
|
64
|
+
|
|
65
|
+
| Surface | Behavior |
|
|
66
|
+
| --- | --- |
|
|
67
|
+
| Provider stream | Prism text, thinking, tool-call delta/final, `usage`, `done`, redacted `error`. |
|
|
68
|
+
| Stream completion | `done` only on completion evidence — OpenAI route: `[DONE]` marker plus terminal `finish_reason`; Anthropic route: `message_stop`; Responses route: `response.completed` (or hop-cap/duplicate-cursor failure). Truncated streams end with terminal `error`; partial output never surfaces as `succeeded`. |
|
|
69
|
+
| OpenAI thinking | `delta.reasoning_content` → thinking deltas; replay via top-level `reasoning_content` when `preserveThinking` (default), never folded into text. |
|
|
70
|
+
| Anthropic thinking | `thinking_delta` → thinking deltas; replay via Anthropic thinking blocks when `preserveThinking`. |
|
|
71
|
+
| Usage | Standard tokens + cache read/write per route; `cost.usd`/`cost.hypercredits`/`remaining.hypercredits` available via `parseHyperUsageCost(wireUsage)` (NeuralWatt pattern) for hosts that surface cost telemetry. Cost/remaining fields are chat-route only; the responses route is a standard OpenAI pass-through (`input_tokens`/`output_tokens`/`total_tokens` + `input_tokens_details.cached_tokens`/`cache_write_tokens`) mapped by the shared Responses machinery. |
|
|
72
|
+
| Auth method | `api_key` for `hyper`, credential name `apiKey`. |
|
|
73
|
+
|
|
74
|
+
## Request/response example
|
|
75
|
+
|
|
76
|
+
```json
|
|
77
|
+
{
|
|
78
|
+
"Authorization": "Bearer sk-hyper-…",
|
|
79
|
+
"content-type": "application/json"
|
|
80
|
+
}
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
Messages route adds provider-owned `x-api-key: <key>` and
|
|
84
|
+
`anthropic-version: 2023-06-01`; Bearer-only authentication is also accepted
|
|
85
|
+
there (Claude Code compatibility). All provider-owned headers are applied after
|
|
86
|
+
caller headers and cannot be overridden.
|
|
87
|
+
|
|
88
|
+
Chat-route body shape (thinking passthrough + preserved reasoning):
|
|
89
|
+
|
|
90
|
+
```json
|
|
91
|
+
{
|
|
92
|
+
"model": "deepseek-v4-pro",
|
|
93
|
+
"stream": true,
|
|
94
|
+
"stream_options": { "include_usage": true },
|
|
95
|
+
"reasoning_effort": "high",
|
|
96
|
+
"messages": [
|
|
97
|
+
{
|
|
98
|
+
"role": "assistant",
|
|
99
|
+
"tool_calls": [{ "id": "call_1", "type": "function", "function": { "name": "lookup", "arguments": "{}" } }],
|
|
100
|
+
"reasoning_content": "plan the lookup"
|
|
101
|
+
}
|
|
102
|
+
]
|
|
103
|
+
}
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
Responses-route body (OpenAI-standard pass-through, shared Responses machinery):
|
|
107
|
+
|
|
108
|
+
```json
|
|
109
|
+
{
|
|
110
|
+
"model": "deepseek-v4-pro",
|
|
111
|
+
"input": [{ "role": "user", "content": [{ "type": "input_text", "text": "hi" }] }],
|
|
112
|
+
"tools": [{ "type": "function", "name": "lookup", "parameters": {} }],
|
|
113
|
+
"stream": true,
|
|
114
|
+
"store": false,
|
|
115
|
+
"reasoning": { "effort": "high" }
|
|
116
|
+
}
|
|
117
|
+
```
|
|
118
|
+
Cache hints (`options.cacheKey`/`sessionId`) surface as the OpenAI-standard
|
|
119
|
+
`sanitized prompt_cache_key` on this route only; `prompt_cache_retention`/`prompt_cache_options`
|
|
120
|
+
are never emitted for implicit Hyper models (no documented 24h/`explicitBreakpoints` support).
|
|
121
|
+
|
|
122
|
+
## Implementation example
|
|
123
|
+
|
|
124
|
+
```ts
|
|
125
|
+
import { createExtensionKernel } from "@arnilo/prism";
|
|
126
|
+
import {
|
|
127
|
+
createHyperProviderPackage,
|
|
128
|
+
getHyperCredits,
|
|
129
|
+
listHyperModels,
|
|
130
|
+
} from "@arnilo/prism-providers/hyper";
|
|
131
|
+
|
|
132
|
+
const kernel = createExtensionKernel();
|
|
133
|
+
await kernel.load([createHyperProviderPackage({ apiKey: process.env.HYPER_API_KEY })]);
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
Caller-gated live catalog (never runs during package setup):
|
|
137
|
+
|
|
138
|
+
```ts
|
|
139
|
+
const models = await listHyperModels({ fetch }); // public endpoint, no auth needed
|
|
140
|
+
await kernel.load([createHyperProviderPackage({ apiKey: process.env.HYPER_API_KEY, models })]);
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
Optional credit display (hosts poll on their own schedule; never called from
|
|
144
|
+
`generate()`):
|
|
145
|
+
|
|
146
|
+
```ts
|
|
147
|
+
const { balance } = await getHyperCredits({ apiKey: process.env.HYPER_API_KEY });
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
## Featured models and routes
|
|
151
|
+
|
|
152
|
+
Featured `hyperModels` mirrors the official `/v1/models` catalog (31 models,
|
|
153
|
+
2026-07 snapshot), with limits, vision capability, documented `effort_levels`,
|
|
154
|
+
and per-million-token pricing (input/output/cache-hit) captured in metadata.
|
|
155
|
+
Route selection follows the observed pricing shape: models whose live catalog
|
|
156
|
+
entry prices explicit cache writes (cache_create > 0, with hit pricing) are
|
|
157
|
+
Anthropic-route `cache_control`; models with implicit write pricing
|
|
158
|
+
(cache_create = 0, no read charge) stay chat-route `implicit` with the write
|
|
159
|
+
fee recorded in `cost.cacheWrite`.
|
|
160
|
+
|
|
161
|
+
| Model family | Route | Cache kind |
|
|
162
|
+
| --- | --- | --- |
|
|
163
|
+
| `deepseek-v4-pro`, `deepseek-v4-pro-0813`, `deepseek-v4-flash` | `openai` | `implicit` |
|
|
164
|
+
| `kimi-k3`, `kimi-k2.7`, `kimi-k2.5`, `glm-5.3-flash`, `glm-5.1`, `gemma-4-fast`, `gpt-oss-120b`, `llama-*`, `minimax-m2.7`, `qwen3-coder`, `qwen3-next`, `qwen3.7-*` | `openai` | `implicit` |
|
|
165
|
+
| `qwen3.6-plus`, `qwen3.6-flash` | `anthropic` | `cache_control` (max 4 breakpoints, no `ttl` — undocumented) |
|
|
166
|
+
|
|
167
|
+
## Model discovery
|
|
168
|
+
|
|
169
|
+
```txt
|
|
170
|
+
GET https://hyper.charm.land/v1/models
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
Public endpoint — works without authentication and emits no auth header when no
|
|
174
|
+
key resolves. `listHyperModels({ fetch?, baseUrl?, apiKey?, signal?, headers? })`
|
|
175
|
+
maps each `{ id, context_window, max_output_tokens, capabilities.vision,
|
|
176
|
+
reasoning.effort_levels, pricing{cache_create, cache_hit} }` entry to
|
|
177
|
+
`ModelConfig` (route from pricing shape via `routeForHyperModel`). Discovery is
|
|
178
|
+
**caller-gated** — setup performs zero fetches.
|
|
179
|
+
|
|
180
|
+
## Thinking / reasoning
|
|
181
|
+
|
|
182
|
+
| Surface | Behavior |
|
|
183
|
+
| --- | --- |
|
|
184
|
+
| OpenAI route stream | `reasoning_content` → thinking deltas |
|
|
185
|
+
| OpenAI route replay | thinking blocks → top-level `reasoning_content` when `preserveThinking`; never folded into text |
|
|
186
|
+
| OpenAI route body | `reasoning_effort` from model default or `options.compat` (request wins), clamped to the model's documented `effortLevels`; invalid values are dropped |
|
|
187
|
+
| Anthropic route stream | `thinking_delta` → thinking deltas |
|
|
188
|
+
| Anthropic route replay | thinking blocks when `preserveThinking` |
|
|
189
|
+
|
|
190
|
+
Owned compat keys (`route`, `preserveThinking`, `reasoning_effort`,
|
|
191
|
+
`effortLevels`) are stripped before opaque compat spread so resolved values win.
|
|
192
|
+
|
|
193
|
+
## Extension and configuration notes
|
|
194
|
+
|
|
195
|
+
- Hosts choose base URL, model list, credential source, and `fetch` impl.
|
|
196
|
+
- Route selection is explicit via `compat.route` (`"anthropic"` or `"responses"`; default `"openai"`). The responses route is never auto-derived — hosts opt in with Responses-shaped model metadata (Codex-style clients), and featured models stay `openai`/`anthropic`.
|
|
197
|
+
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
198
|
+
|
|
199
|
+
### Cache and session behavior
|
|
200
|
+
|
|
201
|
+
- The chat route sends **no** Anthropic `cache_control` fields; it relies on
|
|
202
|
+
OpenAI-style implicit caching. Read tokens map from
|
|
203
|
+
`prompt_tokens_details.cached_tokens` / `cache_write_tokens` /
|
|
204
|
+
`prompt_cache_hit_tokens` (the shared OpenAI usage mapping covers both field
|
|
205
|
+
spellings).
|
|
206
|
+
- The Anthropic route applies `cache_control: { type: "ephemeral" }` markers
|
|
207
|
+
only to the caller-selected `cache.breakpoints` (shared `applyCacheControl()`
|
|
208
|
+
helper) on the last content block of each selected message — not to every
|
|
209
|
+
block. A `system_prompt` breakpoint serializes `system` as marked text blocks
|
|
210
|
+
(plain string otherwise). Caching is enabled unless disabled
|
|
211
|
+
(`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
|
|
212
|
+
`ModelConfig.cache.kind: "cache_control"`.
|
|
213
|
+
- The responses route carries OpenAI-standard `prompt_cache_key` only when the
|
|
214
|
+
caller supplies cache hints (`cacheKey`/`sessionId`, sanitized + clamped to
|
|
215
|
+
64 chars by the shared helper); no `prompt_cache_retention`/`prompt_cache_options`
|
|
216
|
+
(implicit models, no documented 24h/explicit modes).
|
|
217
|
+
- **No `ttl` is ever emitted**: Hyper does not document `cache_control` TTL
|
|
218
|
+
values; `cacheRetention: "long"` must not produce a marker Hyper may reject.
|
|
219
|
+
Re-verify against live behavior before emitting TTLs.
|
|
220
|
+
- Usage accounting per route: chat route maps
|
|
221
|
+
`prompt_tokens_details.cached_tokens`/`cache_write_tokens` (and
|
|
222
|
+
`prompt_cache_hit_tokens`); messages route maps
|
|
223
|
+
`cache_read_input_tokens`/`cache_creation_input_tokens`.
|
|
224
|
+
- Session identity is simple: no session header is emitted (unlike OpenCode Go).
|
|
225
|
+
|
|
226
|
+
### Live-verified mapping (findings ledger)
|
|
227
|
+
|
|
228
|
+
The following claims are encoded as operator-gated probes in
|
|
229
|
+
`packages/prism-providers/src/hyper/__tests__/live.test.ts`. Each probe's
|
|
230
|
+
assertion encodes the documented claim, so a probe failure **is** the finding;
|
|
231
|
+
record the outcome here and adjust the mapping. Status: **pending operator
|
|
232
|
+
run** (no key in CI):
|
|
233
|
+
|
|
234
|
+
| # | Claim (documented) | Probe | Status |
|
|
235
|
+
| --- | --- | --- | --- |
|
|
236
|
+
| 1 | Warm chat-route replay reports cached tokens (`cached_tokens`/`prompt_cache_hit_tokens` → `cacheReadTokens`) | `live_chat_route_reports_cached_tokens_on_warm_prefix_replay` | pending |
|
|
237
|
+
| 2 | `cache_control` on messages reports `cache_creation_input_tokens` on the creating call | `live_messages_route_cache_control_reports_creation_and_read_tokens` | pending |
|
|
238
|
+
| 3 | Same-prefix warm replay reads the created cache entry (TTL ≥ one request) | same probe (warm leg) | pending |
|
|
239
|
+
| 4 | `reasoning_effort` from the model's documented `effortLevels` is accepted (HTTP 200) | `live_reasoning_effort_is_accepted_on_chat_route` | pending |
|
|
240
|
+
|
|
241
|
+
Run the gate:
|
|
242
|
+
|
|
243
|
+
```sh
|
|
244
|
+
PRISM_LIVE_PROVIDER_TESTS=1 HYPER_API_KEY=sk-hyper-... \
|
|
245
|
+
npm run test --workspace=@arnilo/prism-providers/hyper
|
|
246
|
+
```
|
|
247
|
+
|
|
248
|
+
## Security and performance notes
|
|
249
|
+
|
|
250
|
+
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers.
|
|
251
|
+
- No network calls during import, setup, build, or default tests.
|
|
252
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
253
|
+
- API keys are resolved per request from caller-supplied values or resolvers
|
|
254
|
+
and redacted from errors (including discovery and credits failures).
|
|
255
|
+
- `402` (insufficient Hypercredits) surfaces non-retryable `billing_error`;
|
|
256
|
+
`429` and `5xx` are retryable with `retry-after` surfaced as
|
|
257
|
+
`retry_after_ms`; `400/401/403/404` are non-retryable.
|
|
258
|
+
- Caller headers cannot override provider-owned headers (`content-type`,
|
|
259
|
+
`authorization`, and on the messages route `x-api-key`/`anthropic-version`).
|
|
260
|
+
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus
|
|
261
|
+
`HYPER_API_KEY`; default tests are network-free.
|
|
262
|
+
|
|
263
|
+
## Official evidence
|
|
264
|
+
|
|
265
|
+
- Hyper API docs: `https://hyper.charm.land/docs/api/{authentication,list-models,openai-chat-completions,openai-responses,anthropic-messages,credits}.html`
|
|
266
|
+
- Hyper model catalog: `https://hyper.charm.land/docs/models.html`, `https://hyper.charm.land/faq`
|
|
267
|
+
- Live `GET https://hyper.charm.land/v1/models` snapshot (2026-07) — pricing/limits in the static catalog
|
|
268
|
+
- Probe ledger above pending operator-gated live run
|
|
269
|
+
- Intelligent-routing re-check (2026-09): roadmap-only, no documented controls — see
|
|
270
|
+
`../_evidence/phase55-hyper-intelligent-routing.md`
|
|
271
|
+
|
|
272
|
+
## Thinking and reasoning
|
|
273
|
+
|
|
274
|
+
Hyper models derive `capabilities.thinkingLevels` from `compat.effortLevels` (the live `/v1/models` `reasoning.effort_levels` list) and stamp `reasoning_effort`. `hyperReasoningEffort` snaps to the declared set instead of dropping out-of-set values (`max`↔`xhigh` on deepseek-v4-flash); undeclared models and opaque values pass through. The Anthropic route emits resolved `output_config.effort` (snapped) instead of leaking raw `reasoning_effort`. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
275
|
+
|
|
276
|
+
## Related APIs
|
|
277
|
+
|
|
278
|
+
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
279
|
+
`ModelConfig`, discovery contract, request/cache policies.
|
|
280
|
+
- [Thinking and reasoning](../thinking-and-reasoning.md): per-turn `ThinkingLevel` → compat families.
|
|
281
|
+
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
282
|
+
`resolveCredentialValue`, `redactSecrets`.
|
|
283
|
+
- [Provider caching](../provider-caching.md): per-provider cache behavior matrix.
|
|
284
|
+
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
package/docs/providers/kimi.md
CHANGED
|
@@ -199,6 +199,10 @@ await kernel.load([
|
|
|
199
199
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus provider-specific
|
|
200
200
|
env names; default tests are network-free.
|
|
201
201
|
|
|
202
|
+
## Thinking and reasoning
|
|
203
|
+
|
|
204
|
+
Kimi models are family-stamped by id: K3 (`kimiThinkingFamily`) routes through `reasoning_effort` with declared levels `low/high/max` — other portable levels snap (`medium`→`high`, `xhigh`→`max`, `none`/`minimal`→`low`); unknown effort strings pass through for forward compatibility. K2.x models route through `thinking_type` (on/off toggle; K2.7-code thinking is always on). Do not send conflicting `thinking` + `reasoning_effort`. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
205
|
+
|
|
202
206
|
## Related APIs
|
|
203
207
|
|
|
204
208
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
@@ -380,6 +380,10 @@ const decision = classifyNeuralWattError({ status: 429, headers: { "retry-after"
|
|
|
380
380
|
- Live tests stay opt-in behind `NEURALWATT_API_KEY` (plus `PRISM_LIVE_PROVIDER_TESTS=1`);
|
|
381
381
|
default tests are network-free.
|
|
382
382
|
|
|
383
|
+
## Thinking and reasoning
|
|
384
|
+
|
|
385
|
+
NeuralWatt reasoning models declare `low/medium/high/max` and snap `reasoning_effort` to that set; non-reasoning models (`-fast`, gemma) declare nothing. `thinking_token_budget` and `chat_template_kwargs` (`preserve_thinking`/`clear_thinking`) stay package-local. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
386
|
+
|
|
383
387
|
## Related APIs
|
|
384
388
|
|
|
385
389
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
package/docs/providers/ollama.md
CHANGED
|
@@ -155,6 +155,21 @@ await kernel.load([
|
|
|
155
155
|
- Model discovery is caller-gated and never invoked in the provider hot path.
|
|
156
156
|
- Live tests stay opt-in; default tests are network-free.
|
|
157
157
|
|
|
158
|
+
## Live probe
|
|
159
|
+
|
|
160
|
+
The only credential-free live suite: points at a real Ollama server (local `ollama serve` or Ollama Cloud):
|
|
161
|
+
|
|
162
|
+
```bash
|
|
163
|
+
PRISM_LIVE_PROVIDER_TESTS=1 OLLAMA_BASE_URL=http://localhost:11434 \
|
|
164
|
+
node --test packages/prism-providers/dist/ollama/__tests__/live.test.js
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
A health gate lists served models and probes the first (`PRISM_LIVE_OLLAMA_MODEL` to pin one). No server or no pulled models → skip.
|
|
168
|
+
|
|
169
|
+
## Thinking and reasoning
|
|
170
|
+
|
|
171
|
+
Ollama models stamp `reasoning_effort` and snap to declared levels: `gpt-oss*` declares `low/medium/high`; other ids declare nothing and pass effort through verbatim. The native `think` field has a disjoint value set and is never emitted alongside `reasoning_effort`. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
172
|
+
|
|
158
173
|
## Related APIs
|
|
159
174
|
|
|
160
175
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
@@ -150,6 +150,10 @@ const provider = createOpenAICompatibleProvider({
|
|
|
150
150
|
- Tests should use injected `fetch` and never make real network calls.
|
|
151
151
|
- Tool-call arguments are accumulated as streamed text, parsed with `parseJsonObjectArguments` when the final tool call is emitted; empty argument text yields `{}`, malformed JSON yields an `error` event.
|
|
152
152
|
|
|
153
|
+
## Thinking and reasoning
|
|
154
|
+
|
|
155
|
+
The shared OpenAI-compatible base (`createOpenAICompatibleProvider`) does **not** spread `compat` onto request bodies — packages that want thinking forwarded wire a `buildBodyExtra`/`transformBody` hook (Azure, Vertex, and Bedrock use the shared sanitized forwarder: `reasoning_effort` + aliases or a `reasoning` object, effort snapped to declared levels). Host-owned adapters should do the same or accept the no-op. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
156
|
+
|
|
153
157
|
## Related APIs
|
|
154
158
|
|
|
155
159
|
- [Provider layer](../provider-layer.md): registries, provider events, tool-call helpers, and mock provider.
|
package/docs/providers/openai.md
CHANGED
|
@@ -223,6 +223,10 @@ Official: [Reasoning models](https://developers.openai.com/api/docs/guides/reaso
|
|
|
223
223
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
|
|
224
224
|
provider-specific env names; default `npm test` is network-free.
|
|
225
225
|
|
|
226
|
+
## Thinking and reasoning
|
|
227
|
+
|
|
228
|
+
OpenAI models route through the `openai_reasoning` family: the adapter merges `compat.reasoning.effort` and Responses bodies carry `reasoning.effort`. Declared levels (`capabilities.thinkingLevels`): gpt-5.1 family `none/low/medium/high` (default `none`); gpt-5.2 family `none`–`xhigh` (default `medium`); gpt-5.x/o1/o3/o4 families `minimal`–`high` (default `medium`). Unknown model ids declare nothing and pass through untouched. `resolveOpenAIReasoning` snaps a merged effort to the declared set (nearest by ladder distance, ties up; below-minimum snaps up); an existing `reasoning.summary` is preserved. There is no upstream API to enumerate effort values — the tables are doc-pinned in [the evidence matrix](../_evidence/thinking-coverage-2026-09-05.md). See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
229
|
+
|
|
226
230
|
## Related APIs
|
|
227
231
|
|
|
228
232
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`, auth
|
|
@@ -252,6 +252,10 @@ Owned compat keys (`route`, `thinking`, `reasoning`, `reasoning_effort`,
|
|
|
252
252
|
- [OpenCode Go](https://opencode.ai/docs/go/) — model list, dual endpoints, pricing/usage, `GET /zen/go/v1/models`
|
|
253
253
|
- Pi secondary (ids/limits only): `packages/ai/src/providers/opencode-go.ts`, `opencode-go.models.ts`
|
|
254
254
|
|
|
255
|
+
## Thinking and reasoning
|
|
256
|
+
|
|
257
|
+
OpenCode-Go models carry level tables: `kimi-k3`/`deepseek-v4`/`glm-5.3` → `reasoning_effort` (`low/high/max`), `glm-5.2` → `reasoning_effort` (`low`–`max`), `grok-4.6` → `low/medium/high/xhigh`, `grok-4.5` → `low/medium/high`, Kimi-K2.x/MiniMax/Qwen → `thinking_type`; mimo/unknown declare nothing and pass through. The Anthropic route uses the shared serializer with resolved thinking + `output_config.effort` hooks; the OpenAI route snaps effort to declared sets. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
258
|
+
|
|
255
259
|
## Related APIs
|
|
256
260
|
|
|
257
261
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
@@ -189,7 +189,11 @@ cache-read pricing exists), and seeds `compat.reasoning.effort` from
|
|
|
189
189
|
hidden app identity.
|
|
190
190
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
|
|
191
191
|
provider-specific env names; default tests are network-free.
|
|
192
|
-
- Enterprise hosts that must gate `compat.openRouterRouting` should wrap selection with `@arnilo/prism-model-router` (`allowOpenRouterRouting`); the OpenRouter adapter itself still passthroughs routing when present on the request.
|
|
192
|
+
- Enterprise hosts that must gate `compat.openRouterRouting` should wrap selection with `@arnilo/prism-core/governance/model-router` (`allowOpenRouterRouting`); the OpenRouter adapter itself still passthroughs routing when present on the request.
|
|
193
|
+
|
|
194
|
+
## Thinking and reasoning
|
|
195
|
+
|
|
196
|
+
OpenRouter models route through the `openai_reasoning` family with **API-derived** levels: `mapOpenRouterModel` reads the models API `reasoning.supported_efforts` into `capabilities.thinkingLevels` (a `mandatory` model excludes `none`), and `resolveOpenRouterReasoning` snaps a merged effort to that set. Models without reasoning metadata pass effort through verbatim (OpenRouter accepts the full ladder globally). See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
193
197
|
|
|
194
198
|
## Related APIs
|
|
195
199
|
|