@arnilo/prism 0.4.0 → 0.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +41 -1
- package/README.md +23 -20
- package/dist/agent-run-state.d.ts +1 -2
- package/dist/agent-run-state.js +0 -3
- package/dist/agent-session/session/assemble.d.ts +6 -0
- package/dist/agent-session/session/assemble.js +391 -0
- package/dist/agent-session/session/persist.d.ts +28 -0
- package/dist/agent-session/session/persist.js +166 -0
- package/dist/agent-session/session/provider-round.d.ts +6 -0
- package/dist/agent-session/session/provider-round.js +231 -0
- package/dist/agent-session/session/tool-round.d.ts +31 -0
- package/dist/agent-session/session/tool-round.js +473 -0
- package/dist/agent-session/session/types.d.ts +115 -0
- package/dist/agent-session/session/types.js +5 -0
- package/dist/agent-session/session.d.ts +49 -43
- package/dist/agent-session/session.js +24 -1180
- package/dist/capture.d.ts +63 -0
- package/dist/capture.js +67 -0
- package/dist/cli-init.d.ts +18 -2
- package/dist/cli-init.js +2 -7
- package/dist/cli-runner.d.ts +2 -2
- package/dist/cli-runner.js +45 -9
- package/dist/content.d.ts +3 -3
- package/dist/content.js +3 -1
- package/dist/contracts-core/agent.d.ts +4 -0
- package/dist/contracts-core/batch.d.ts +97 -0
- package/dist/contracts-core/batch.js +65 -0
- package/dist/contracts-core/content.d.ts +72 -1
- package/dist/contracts-core/embeddings.d.ts +30 -0
- package/dist/contracts-core/embeddings.js +17 -0
- package/dist/contracts-core/images.d.ts +60 -0
- package/dist/contracts-core/images.js +17 -0
- package/dist/contracts-core/moderation.d.ts +46 -0
- package/dist/contracts-core/moderation.js +34 -0
- package/dist/contracts-core/speech.d.ts +39 -0
- package/dist/contracts-core/speech.js +17 -0
- package/dist/contracts-core/transcription.d.ts +48 -0
- package/dist/contracts-core/transcription.js +17 -0
- package/dist/contracts-core/video.d.ts +61 -0
- package/dist/contracts-core/video.js +17 -0
- package/dist/contracts-core.d.ts +7 -0
- package/dist/contracts-core.js +7 -0
- package/dist/contracts-protocol.d.ts +2 -0
- package/dist/index.d.ts +7 -5
- package/dist/index.js +5 -4
- package/dist/input.js +3 -2
- package/dist/node/agent-definitions.d.ts +1 -8
- package/dist/node/agent-definitions.js +0 -34
- package/dist/node/settings.d.ts +0 -1
- package/dist/node/settings.js +0 -5
- package/dist/pinned-fetch.js +29 -3
- package/dist/provider-events.js +3 -4
- package/dist/provider-request-policy.d.ts +15 -0
- package/dist/provider-request-policy.js +52 -0
- package/dist/providers/media.d.ts +1 -2
- package/dist/providers/media.js +1 -4
- package/dist/rpc.d.ts +1 -1
- package/dist/rpc.js +4 -4
- package/dist/testing/provider-conformance.d.ts +114 -5
- package/dist/testing/provider-conformance.js +342 -0
- package/dist/testing/tool-effect-store-conformance.d.ts +0 -1
- package/dist/testing/tool-effect-store-conformance.js +0 -3
- package/dist/thinking.d.ts +48 -9
- package/dist/thinking.js +134 -8
- package/docs/0.1.0-readiness.md +3 -3
- package/docs/a2a.md +2 -2
- package/docs/acp.md +3 -3
- package/docs/ag-ui-adoption.md +1 -1
- package/docs/ag-ui.md +1 -2
- package/docs/agent-definitions.md +1 -1
- package/docs/agent-events.md +5 -5
- package/docs/agent-identity.md +13 -2
- package/docs/agent-session-runtime.md +2 -1
- package/docs/audit-export.md +3 -3
- package/docs/batch-jobs.md +120 -0
- package/docs/cli-rpc.md +20 -9
- package/docs/coding-agent-tools.md +19 -19
- package/docs/coding-review-and-diagnostics.md +2 -2
- package/docs/coding-security.md +4 -4
- package/docs/coding-workspaces.md +2 -2
- package/docs/compaction-llm.md +2 -0
- package/docs/compaction-observational-memory.md +3 -0
- package/docs/computer-use-linux.md +13 -2
- package/docs/context-and-skills.md +1 -1
- package/docs/conversations.md +4 -4
- package/docs/credential-storage.md +11 -7
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/data-classification.md +1 -1
- package/docs/database-persistence.md +4 -4
- package/docs/dev-inspector.md +6 -6
- package/docs/device-adapters.md +2 -2
- package/docs/diagrams.md +1 -1
- package/docs/document-reader.md +6 -6
- package/docs/documents.md +5 -4
- package/docs/embeddings.md +112 -0
- package/docs/enterprise-postgres-state.md +7 -7
- package/docs/evaluations.md +8 -8
- package/docs/extensions.md +3 -3
- package/docs/forge-integration.md +3 -3
- package/docs/graft.md +2 -2
- package/docs/guardrails.md +1 -1
- package/docs/host-security.md +15 -15
- package/docs/image-generation.md +129 -0
- package/docs/impeccable.md +5 -3
- package/docs/index.md +64 -36
- package/docs/indexed-code-search.md +2 -2
- package/docs/input-and-prompt-assembly.md +1 -1
- package/docs/language-intelligence.md +4 -4
- package/docs/live-testing.md +126 -0
- package/docs/mcp-tools.md +43 -12
- package/docs/middleware-hooks.md +1 -1
- package/docs/migrate-to-0.4.md +3 -3
- package/docs/migrate-to-0.5.md +144 -0
- package/docs/migration.md +33 -1
- package/docs/model-registry.md +38 -0
- package/docs/model-routing.md +5 -5
- package/docs/moderation.md +117 -0
- package/docs/multi-agent-patterns.md +4 -4
- package/docs/multimodal-content.md +26 -2
- package/docs/obscura.md +2 -2
- package/docs/observability.md +32 -7
- package/docs/openapi-tools.md +13 -3
- package/docs/operations.md +11 -0
- package/docs/performance.md +7 -7
- package/docs/persistence-credentials-multimodality-primitives.md +6 -6
- package/docs/policy-and-audit.md +17 -7
- package/docs/ponytail.md +1 -1
- package/docs/postgres-persistence.md +5 -5
- package/docs/process-sessions.md +2 -2
- package/docs/prompt-registry.md +7 -7
- package/docs/provider-caching.md +8 -2
- package/docs/provider-conformance.md +23 -1
- package/docs/provider-packages.md +49 -17
- package/docs/provider-primitives.md +1 -1
- package/docs/provider-request-policies.md +19 -6
- package/docs/providers/ai-sdk.md +27 -3
- package/docs/providers/alibaba.md +17 -1
- package/docs/providers/anthropic.md +16 -0
- package/docs/providers/azure.md +29 -1
- package/docs/providers/bedrock.md +27 -0
- package/docs/providers/clinepass.md +16 -0
- package/docs/providers/commandcode.md +265 -0
- package/docs/providers/deepseek.md +16 -0
- package/docs/providers/google.md +16 -0
- package/docs/providers/hyper.md +296 -0
- package/docs/providers/kimi.md +16 -0
- package/docs/providers/neuralwatt.md +16 -0
- package/docs/providers/ollama.md +27 -0
- package/docs/providers/openai-compatible.md +16 -0
- package/docs/providers/openai.md +16 -0
- package/docs/providers/opencode-go.md +16 -0
- package/docs/providers/openrouter.md +17 -1
- package/docs/providers/vertex.md +28 -0
- package/docs/providers/xai.md +16 -0
- package/docs/providers/zai.md +16 -0
- package/docs/public-contracts.md +1 -1
- package/docs/rag.md +26 -4
- package/docs/release-and-install.md +103 -46
- package/docs/resource-loading.md +1 -1
- package/docs/runs-and-usage.md +14 -2
- package/docs/server.md +5 -5
- package/docs/settings-auth-trust-security.md +7 -5
- package/docs/sheets.md +2 -2
- package/docs/speech.md +126 -0
- package/docs/sqlite-persistence.md +4 -4
- package/docs/supervisors.md +3 -3
- package/docs/thinking-and-reasoning.md +99 -61
- package/docs/tool-conformance.md +1 -1
- package/docs/tool-execution-primitives.md +8 -8
- package/docs/tools.md +4 -4
- package/docs/use-case-model-selection.md +1 -1
- package/docs/web-tools.md +1 -1
- package/docs/wiki.md +1 -1
- package/docs/work-artifacts-and-review.md +17 -6
- package/docs/work-connectors.md +4 -4
- package/docs/work-tools.md +5 -5
- package/docs/workflow-orchestration-primitives.md +11 -11
- package/docs/workflows.md +5 -5
- package/package.json +11 -8
- package/templates/init/providers.json +24 -8
- package/docs/antigravity-agent.md +0 -207
|
@@ -57,6 +57,18 @@ Live canaries stay opt-in behind host credentials; default tests are network-fre
|
|
|
57
57
|
|
|
58
58
|
Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hosts needing Converse-only models should supply a custom provider or AI SDK bridge.
|
|
59
59
|
|
|
60
|
+
## Request construction (0.5.1)
|
|
61
|
+
|
|
62
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
63
|
+
|
|
64
|
+
| | |
|
|
65
|
+
| --- | --- |
|
|
66
|
+
| P1 session wire | none |
|
|
67
|
+
| Mandatory | no |
|
|
68
|
+
| P2 default cache | host-owned, no Prism cache fields |
|
|
69
|
+
|
|
70
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
71
|
+
|
|
60
72
|
## Security and performance notes
|
|
61
73
|
|
|
62
74
|
- No AWS SDK; package-local SigV4 only for `bedrock` service.
|
|
@@ -66,6 +78,21 @@ Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hos
|
|
|
66
78
|
- Credential secrets are redacted from provider errors.
|
|
67
79
|
- No credential prefetch at import.
|
|
68
80
|
|
|
81
|
+
## Live probe
|
|
82
|
+
|
|
83
|
+
Opt-in smoke over real AWS Bedrock (package-local SigV4, static keys or session token):
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
PRISM_LIVE_PROVIDER_TESTS=1 AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=us-east-1 \
|
|
87
|
+
node --test packages/prism-providers/dist/bedrock/__tests__/live.test.js
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
`PRISM_LIVE_BEDROCK_MODEL` overrides the probed model (default `us.anthropic.claude-haiku-4-5-20251001-v1:0`). Without credentials the suite skips.
|
|
91
|
+
|
|
92
|
+
## Thinking and reasoning
|
|
93
|
+
|
|
94
|
+
Bedrock OpenAI-compat chat expects snake_case `reasoning_effort` (with `effort`/`reasoningEffort` aliases) or a sanitized `reasoning` object. OpenAI-family models on Bedrock snap effort to their declared levels (gpt-5.1 → `none/low/medium/high`); non-OpenAI models pass through untouched. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
95
|
+
|
|
69
96
|
## Related APIs
|
|
70
97
|
|
|
71
98
|
- [OpenAI-compatible provider](openai-compatible.md)
|
|
@@ -98,6 +98,18 @@ await session.prompt("Plan the refactor", {
|
|
|
98
98
|
- Multi-backend gateway: key compat off `api.cline.bot`, not the upstream vendor.
|
|
99
99
|
- Reference USD-per-million costs are catalog metadata; ClinePass itself is a subscription.
|
|
100
100
|
|
|
101
|
+
## Request construction (0.5.1)
|
|
102
|
+
|
|
103
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
104
|
+
|
|
105
|
+
| | |
|
|
106
|
+
| --- | --- |
|
|
107
|
+
| P1 session wire | none; strips `cache_control` |
|
|
108
|
+
| Mandatory | no |
|
|
109
|
+
| P2 default cache | implicit, no markers |
|
|
110
|
+
|
|
111
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
112
|
+
|
|
101
113
|
## Security and performance notes
|
|
102
114
|
|
|
103
115
|
- No network on import, setup, build, or default tests.
|
|
@@ -106,6 +118,10 @@ await session.prompt("Plan the refactor", {
|
|
|
106
118
|
- Provider-owned headers win. One POST per generate. Bounded error bodies.
|
|
107
119
|
- Live tests: `PRISM_LIVE_PROVIDER_TESTS=1` plus `CLINE_API_KEY`.
|
|
108
120
|
|
|
121
|
+
## Thinking and reasoning
|
|
122
|
+
|
|
123
|
+
ClinePass routes through `reasoning_effort` with per-model slot maps (`compat.thinkingLevelMap`) as wire authority; declared `capabilities.thinkingLevels` mirror each map's portable slots (e.g. GLM: `none/low/medium/high/xhigh`). Portable `max` never reaches the wire (upstream 500s) — the map sends `high`; unsupported slots omit the field. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
124
|
+
|
|
109
125
|
## Related APIs
|
|
110
126
|
|
|
111
127
|
- [Provider packages](../provider-packages.md)
|
|
@@ -0,0 +1,265 @@
|
|
|
1
|
+
# Command Code provider package
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-providers/commandcode` provides explicit, side-effect-free setup
|
|
6
|
+
for the [Command Code Provider API](https://commandcode.ai/docs/provider) — an
|
|
7
|
+
aggregator exposing every top commercial and open model through OpenAI- and
|
|
8
|
+
Anthropic-compatible endpoints, billed at cost with deals auto-applied. The
|
|
9
|
+
package dual-routes by `ModelConfig.compat.route`:
|
|
10
|
+
|
|
11
|
+
| Route | Endpoint | Official model families |
|
|
12
|
+
| --- | --- | --- |
|
|
13
|
+
| `"openai"` (default) | `POST {baseUrl}/chat/completions` | everything except `claude-*` (GPT-5.6, DeepSeek, Kimi, GLM, MiniMax, Qwen, MiMo, Gemini flash, Grok) |
|
|
14
|
+
| `"anthropic"` | `POST {baseUrl}/messages` | `claude-*` tiers (Opus/Sonnet/Fable/Haiku) |
|
|
15
|
+
|
|
16
|
+
Default base URL is the official Provider API root:
|
|
17
|
+
|
|
18
|
+
```txt
|
|
19
|
+
https://api.commandcode.ai/provider/v1
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
Authentication: `Authorization: Bearer <key>` on the chat route,
|
|
23
|
+
`x-api-key` + `anthropic-version: 2023-06-01` on the messages route (Claude
|
|
24
|
+
Code compatibility). The same key authenticates the CLI and the API.
|
|
25
|
+
|
|
26
|
+
## When to use it
|
|
27
|
+
|
|
28
|
+
Use it when a host app wants Command Code models through Prism's `AgentSession`
|
|
29
|
+
runtime: dual-route serialization, `cache_control` breakpoints on Claude
|
|
30
|
+
models, implicit caching elsewhere, reasoning replay, caller-gated model
|
|
31
|
+
discovery, and optional zero-data-retention (`zdr: true` → provider-owned
|
|
32
|
+
`x-cmd-zdr: 1`, which routes only through ZDR-capable upstreams).
|
|
33
|
+
|
|
34
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches,
|
|
35
|
+
or real-network tests (live probes are operator-gated, see below).
|
|
36
|
+
|
|
37
|
+
## Inputs / request
|
|
38
|
+
|
|
39
|
+
```ts
|
|
40
|
+
import {
|
|
41
|
+
createCommandCodeProviderPackage,
|
|
42
|
+
listCommandCodeModels,
|
|
43
|
+
} from "@arnilo/prism-providers/commandcode";
|
|
44
|
+
|
|
45
|
+
createCommandCodeProviderPackage(options: CommandCodeProviderPackageOptions): ProviderPackage
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
| Field | Type | Purpose |
|
|
49
|
+
| --- | --- | --- |
|
|
50
|
+
| `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
|
|
51
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
52
|
+
| `baseUrl` | `string` | Overrides official `https://api.commandcode.ai/provider/v1`. |
|
|
53
|
+
| `models` | `readonly ModelConfig[]` | Overrides featured `commandCodeModels` defaults. |
|
|
54
|
+
| `zdr` | `boolean` | Enforce zero data retention (`x-cmd-zdr: 1`). May route to costlier upstreams or fail `422 cmd_zdr_no_providers`. |
|
|
55
|
+
|
|
56
|
+
`ProviderRequest.options.cache.breakpoints` select messages-route
|
|
57
|
+
`cache_control` markers for Claude models (max 4, no `ttl`).
|
|
58
|
+
|
|
59
|
+
## Outputs / response / events
|
|
60
|
+
|
|
61
|
+
| Surface | Behavior |
|
|
62
|
+
| --- | --- |
|
|
63
|
+
| Provider stream | Prism text, thinking, tool-call delta/final, `usage`, `done`, redacted `error`. |
|
|
64
|
+
| Stream completion | `done` only on completion evidence — chat route: `[DONE]` marker plus terminal `finish_reason`; messages route: `message_stop`. Truncated streams end with terminal `error`. |
|
|
65
|
+
| OpenAI thinking | `delta.reasoning_content` → thinking deltas; replay via top-level `reasoning_content` when `preserveThinking` (default), never folded into text. |
|
|
66
|
+
| Anthropic thinking | `thinking_delta` → thinking deltas; replay via Anthropic thinking blocks when `preserveThinking`. |
|
|
67
|
+
| Usage | Standard tokens + cache read/write per route; mapped through the shared OpenAI/Anthropic usage mappings. |
|
|
68
|
+
| Auth method | `api_key` for `commandcode`, credential name `apiKey`. |
|
|
69
|
+
|
|
70
|
+
## Request/response example
|
|
71
|
+
|
|
72
|
+
```json
|
|
73
|
+
{
|
|
74
|
+
"authorization": "Bearer cmd_…",
|
|
75
|
+
"content-type": "application/json"
|
|
76
|
+
}
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
Messages route instead sends provider-owned `x-api-key: <key>` and
|
|
80
|
+
`anthropic-version: 2023-06-01`. All provider-owned headers are applied after
|
|
81
|
+
caller headers and cannot be overridden.
|
|
82
|
+
|
|
83
|
+
Chat-route body (thinking passthrough + preserved reasoning):
|
|
84
|
+
|
|
85
|
+
```json
|
|
86
|
+
{
|
|
87
|
+
"model": "Qwen/Qwen3.8-Flash",
|
|
88
|
+
"stream": true,
|
|
89
|
+
"stream_options": { "include_usage": true },
|
|
90
|
+
"max_tokens": 512,
|
|
91
|
+
"messages": [
|
|
92
|
+
{
|
|
93
|
+
"role": "assistant",
|
|
94
|
+
"tool_calls": [{ "id": "call_1", "type": "function", "function": { "name": "lookup", "arguments": "{}" } }],
|
|
95
|
+
"reasoning_content": "plan the lookup"
|
|
96
|
+
}
|
|
97
|
+
]
|
|
98
|
+
}
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
## Implementation example
|
|
102
|
+
|
|
103
|
+
```ts
|
|
104
|
+
import { createExtensionKernel } from "@arnilo/prism";
|
|
105
|
+
import { createCommandCodeProviderPackage } from "@arnilo/prism-providers/commandcode";
|
|
106
|
+
|
|
107
|
+
const kernel = createExtensionKernel();
|
|
108
|
+
await kernel.load([createCommandCodeProviderPackage({ apiKey: process.env.COMMAND_CODE_API_KEY })]);
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Caller-gated live catalog (never runs during package setup):
|
|
112
|
+
|
|
113
|
+
```ts
|
|
114
|
+
const models = await listCommandCodeModels({ fetch }); // public endpoint, no auth needed
|
|
115
|
+
await kernel.load([createCommandCodeProviderPackage({ apiKey: process.env.COMMAND_CODE_API_KEY, models })]);
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
## Featured models and routes
|
|
119
|
+
|
|
120
|
+
Featured `commandCodeModels` is a curated 38-model bootstrap catalog: ids and
|
|
121
|
+
context windows from the live `GET /provider/v1/models` snapshot (2026-09, 67
|
|
122
|
+
ids), USD-per-million-token pricing from the docs table
|
|
123
|
+
(<https://commandcode.ai/docs/resources/pricing-limits>). `compat.pricing_source`
|
|
124
|
+
records caveats: open-source models bill at the **mean per-provider price**;
|
|
125
|
+
DeepSeek rates are **off-peak** (17h/day; peak 2× during 01–04 & 06–10 UTC);
|
|
126
|
+
deals (MiniMax M3 −50%, MiMo −98/99%) are already applied upstream. Custom
|
|
127
|
+
pricing metadata is always stripped before the wire.
|
|
128
|
+
|
|
129
|
+
| Model family | Route | Cache kind |
|
|
130
|
+
| --- | --- | --- |
|
|
131
|
+
| `claude-opus-5/4-8/4-7`, `claude-sonnet-5/4-6`, `claude-fable-5-1/5`, `claude-haiku-4-5` | `anthropic` | `cache_control` (max 4 breakpoints, no `ttl` — undocumented) |
|
|
132
|
+
| `gpt-5.6-sol/terra/luna` | `openai` | `implicit` (docs cache-write price recorded in `cost.cacheWrite`; explicit-key upgrade gated on live probe — Task 9) |
|
|
133
|
+
| `deepseek/*`, Kimi, GLM, MiniMax, Qwen, MiMo, Gemini flash, Grok | `openai` | `implicit` |
|
|
134
|
+
|
|
135
|
+
## Model discovery
|
|
136
|
+
|
|
137
|
+
```txt
|
|
138
|
+
GET https://api.commandcode.ai/provider/v1/models
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
Public endpoint — works without authentication and emits no auth header when no
|
|
142
|
+
key resolves. `listCommandCodeModels({ fetch?, baseUrl?, apiKey?, signal?,
|
|
143
|
+
headers? })` maps each `{ id, name, context_length }` entry to `ModelConfig`:
|
|
144
|
+
route from id (`claude-*` → anthropic), context window from the endpoint, and
|
|
145
|
+
featured docs metadata (cost/cache kind) applied when the id matches a curated
|
|
146
|
+
entry. The endpoint carries no pricing or capabilities — unknown ids get
|
|
147
|
+
route-derived cache kind and no cost. Discovery is **caller-gated** — setup
|
|
148
|
+
performs zero fetches.
|
|
149
|
+
|
|
150
|
+
## Thinking / reasoning
|
|
151
|
+
|
|
152
|
+
| Surface | Behavior |
|
|
153
|
+
| --- | --- |
|
|
154
|
+
| OpenAI route stream | `reasoning_content` → thinking deltas |
|
|
155
|
+
| OpenAI route replay | thinking blocks → top-level `reasoning_content` when `preserveThinking`; never folded into text |
|
|
156
|
+
| Anthropic route stream | `thinking_delta` → thinking deltas |
|
|
157
|
+
| Anthropic route replay | thinking blocks when `preserveThinking` |
|
|
158
|
+
|
|
159
|
+
Owned compat keys (`route`, `preserveThinking`, `pricing_source`) are stripped
|
|
160
|
+
before opaque compat spread so resolved values win.
|
|
161
|
+
|
|
162
|
+
## Extension and configuration notes
|
|
163
|
+
|
|
164
|
+
- Hosts choose base URL, model list, credential source, `fetch` impl, and ZDR.
|
|
165
|
+
- Route selection is explicit via `compat.route` (`"anthropic"` for `claude-*`
|
|
166
|
+
ids, default `"openai"`). Sending a model to the wrong endpoint 400s.
|
|
167
|
+
- Package contributes models via the extension `api` and an `api_key` auth method.
|
|
168
|
+
|
|
169
|
+
### Cache and session behavior
|
|
170
|
+
|
|
171
|
+
- The chat route sends **no** `cache_control` fields; it relies on OpenAI-style
|
|
172
|
+
implicit caching (upstream provider behavior, passed through). Read tokens
|
|
173
|
+
map from `prompt_tokens_details.cached_tokens` / `cache_write_tokens` /
|
|
174
|
+
`prompt_cache_hit_tokens`.
|
|
175
|
+
- The messages route applies `cache_control: { type: "ephemeral" }` markers
|
|
176
|
+
only to caller-selected `cache.breakpoints` (shared `applyCacheControl()`
|
|
177
|
+
helper) on the last content block of each selected message. A
|
|
178
|
+
`system_prompt` breakpoint serializes `system` as marked text blocks (plain
|
|
179
|
+
string otherwise). Caching is enabled unless disabled
|
|
180
|
+
(`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
|
|
181
|
+
`ModelConfig.cache.kind: "cache_control"`.
|
|
182
|
+
- **No `ttl` is ever emitted**: the upstream TTL window is undocumented on the
|
|
183
|
+
Provider API; `cacheRetention: "long"` must not produce a marker the gateway
|
|
184
|
+
may reject.
|
|
185
|
+
- Usage accounting per route: chat route maps
|
|
186
|
+
`prompt_tokens_details.cached_tokens`/`cache_write_tokens` (and
|
|
187
|
+
`prompt_cache_hit_tokens`); messages route maps
|
|
188
|
+
`cache_read_input_tokens`/`cache_creation_input_tokens`.
|
|
189
|
+
- Session identity is simple: no session header is emitted (undocumented).
|
|
190
|
+
|
|
191
|
+
### Live-verified mapping (findings ledger)
|
|
192
|
+
|
|
193
|
+
The following claims are encoded as operator-gated probes in
|
|
194
|
+
`packages/prism-providers/src/commandcode/__tests__/live.test.ts`. Each probe's
|
|
195
|
+
assertion encodes the documented claim, so a probe failure **is** the finding;
|
|
196
|
+
record the outcome here and adjust the mapping. Status: **pending operator
|
|
197
|
+
run** (no key in CI):
|
|
198
|
+
|
|
199
|
+
| # | Claim (documented) | Probe | Status |
|
|
200
|
+
| --- | --- | --- | --- |
|
|
201
|
+
| 1 | Warm chat-route replay reports cached tokens (implicit caching passes through) | `live_chat_route_reports_cached_tokens_on_warm_prefix_replay` | pending |
|
|
202
|
+
| 2 | `cache_control` on messages reports `cache_creation_input_tokens` on the creating call | `live_messages_route_cache_control_reports_creation_and_read_tokens` | pending |
|
|
203
|
+
| 3 | Same-prefix warm replay reads the created cache entry (TTL ≥ one request) | same probe (warm leg) | pending |
|
|
204
|
+
| 4 | OpenAI `prompt_cache_key` is accepted and honored for GPT-5.6 (explicit caching) — decides Task 9 | `live_gpt56_prompt_cache_key_passthrough_probe` | pending — Task 9 closed gated: **pass** → upgrade `gpt-5.6-*` to the OpenAI explicit mapping (`promptCacheKey`/`promptCacheOptions`/`applyPromptCacheBreakpoints` via `@arnilo/prism-providers/openai`, verified exported); **fail** (400/ignored/warm replay shows no cached tokens) → verified-negative, keep `implicit`, record here |
|
|
205
|
+
| 5 | OpenAI `reasoning_effort` is accepted (200) on the chat route | `live_reasoning_effort_is_accepted_on_chat_route` | pending |
|
|
206
|
+
| 6 | ZDR requests route (done) or fail `422 cmd_zdr_no_providers` when no ZDR-capable upstream exists | `live_zdr_route_probe_is_opt_in_and_routable` | pending |
|
|
207
|
+
|
|
208
|
+
Run the gate:
|
|
209
|
+
|
|
210
|
+
```sh
|
|
211
|
+
PRISM_LIVE_PROVIDER_TESTS=1 COMMAND_CODE_API_KEY=cmd_... \
|
|
212
|
+
npm run test --workspace=@arnilo/prism-providers/commandcode
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
## Request construction (0.5.1)
|
|
216
|
+
|
|
217
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
218
|
+
|
|
219
|
+
| | |
|
|
220
|
+
| --- | --- |
|
|
221
|
+
| P1 session wire | none |
|
|
222
|
+
| Mandatory | no |
|
|
223
|
+
| P2 default cache | Anthropic: `cache_control`; OpenAI: implicit (no `prompt_cache_key`) |
|
|
224
|
+
|
|
225
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
226
|
+
|
|
227
|
+
## Security and performance notes
|
|
228
|
+
|
|
229
|
+
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers.
|
|
230
|
+
- No network calls during import, setup, build, or default tests.
|
|
231
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
232
|
+
- API keys are resolved per request from caller-supplied values or resolvers
|
|
233
|
+
and redacted from errors (upstream error bodies may carry the upstream
|
|
234
|
+
provider's message — always redacted).
|
|
235
|
+
- `403 upgrade_required` (Go plan — no API access) and
|
|
236
|
+
`422 cmd_zdr_no_providers` are non-retryable; `429` and `5xx` are retryable
|
|
237
|
+
with `retry-after` surfaced as `retry_after_ms`; `400/401/422` are
|
|
238
|
+
non-retryable.
|
|
239
|
+
- Caller headers cannot override provider-owned headers (`content-type`,
|
|
240
|
+
`authorization` on chat, `x-api-key`/`anthropic-version` on messages,
|
|
241
|
+
`x-cmd-zdr` when ZDR is opted in).
|
|
242
|
+
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus
|
|
243
|
+
`COMMAND_CODE_API_KEY`; default tests are network-free.
|
|
244
|
+
|
|
245
|
+
## Official evidence
|
|
246
|
+
|
|
247
|
+
- Command Code Provider API docs: `https://commandcode.ai/docs/provider`
|
|
248
|
+
- Pricing & limits (per-model USD, deals, off-peak): `https://commandcode.ai/docs/resources/pricing-limits`
|
|
249
|
+
- Live `GET https://api.commandcode.ai/provider/v1/models` snapshot (2026-09) — 67 ids, context windows
|
|
250
|
+
- Probe ledger above pending operator-gated live run; row 4 decides plan 055 Task 9
|
|
251
|
+
(explicit GPT-5.6 caching upgrade) — see `plans/055-First-Class-Hyper-And-Command-Code-Providers.md`
|
|
252
|
+
|
|
253
|
+
## Thinking and reasoning
|
|
254
|
+
|
|
255
|
+
Command Code models carry provenance-commented level tables: `claude-*` → `output_config_effort` (Anthropic-route effort sets mirroring the native Anthropic package), `gpt-5.6*` → `openai_reasoning` (`none`–`xhigh`), `deepseek-v4`/`kimi-k3`/`glm-5.3` → `reasoning_effort` (`low/high/max`), `glm-5.2` → `reasoning_effort` (`low`–`max`), Kimi-K2.x/MiniMax/Qwen → `thinking_type`, gemini-3.x → `noop` (the gateway chat route has no `thinking_level` wire, so no levels are declared), mimo/unknown → passthrough. Effort snaps to declared sets on both routes. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
256
|
+
|
|
257
|
+
## Related APIs
|
|
258
|
+
|
|
259
|
+
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
260
|
+
`ModelConfig`, discovery contract, request/cache policies.
|
|
261
|
+
- [Thinking and reasoning](../thinking-and-reasoning.md): per-turn `ThinkingLevel` → compat families.
|
|
262
|
+
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
263
|
+
`resolveCredentialValue`, `redactSecrets`.
|
|
264
|
+
- [Provider caching](../provider-caching.md): per-provider cache behavior matrix.
|
|
265
|
+
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
@@ -119,6 +119,18 @@ await session.prompt("Plan the refactor", {
|
|
|
119
119
|
- Tool-turn assistants must replay `reasoning_content` or the API returns 400.
|
|
120
120
|
Non-tool multi-turn may omit it (the API ignores it).
|
|
121
121
|
|
|
122
|
+
## Request construction (0.5.1)
|
|
123
|
+
|
|
124
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
125
|
+
|
|
126
|
+
| | |
|
|
127
|
+
| --- | --- |
|
|
128
|
+
| P1 session wire | none |
|
|
129
|
+
| Mandatory | no |
|
|
130
|
+
| P2 default cache | implicit, no markers |
|
|
131
|
+
|
|
132
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
133
|
+
|
|
122
134
|
## Security and performance notes
|
|
123
135
|
|
|
124
136
|
- SSE streams and HTTP error bodies use bounded transport helpers.
|
|
@@ -129,6 +141,10 @@ await session.prompt("Plan the refactor", {
|
|
|
129
141
|
- One POST per generate. No provider retry loop.
|
|
130
142
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus `DEEPSEEK_API_KEY`.
|
|
131
143
|
|
|
144
|
+
## Thinking and reasoning
|
|
145
|
+
|
|
146
|
+
DeepSeek models declare `low/high/max` and stamp `reasoning_effort`; the wire table maps `medium`/`xhigh`→`high`, `none`/`minimal` stop thinking. `thinking.type: "enabled"/"disabled"` stays available (thinking on by default, `high`); a request-level `reasoning_effort: none` stops thinking only when no explicit `thinking` switch was sent. Tool turns must replay `reasoning_content` or the API returns 400. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
147
|
+
|
|
132
148
|
## Related APIs
|
|
133
149
|
|
|
134
150
|
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
package/docs/providers/google.md
CHANGED
|
@@ -73,6 +73,18 @@ api.registerProviderPackage(createGoogleProviderPackage({ apiKey: hostKey, model
|
|
|
73
73
|
- Vertex / enterprise identity stays out of 0.0.11.
|
|
74
74
|
- Gemini CLI says third-party software accessing its backend through Gemini CLI OAuth violates applicable terms, and its FAQ directs third-party coding agents to Vertex AI or Google AI Studio API keys ([terms](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/tos-privacy.md), [FAQ](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/faq.md)). Prism therefore has no Gemini CLI OAuth API or token-import shortcut.
|
|
75
75
|
|
|
76
|
+
## Request construction (0.5.1)
|
|
77
|
+
|
|
78
|
+
Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
|
|
79
|
+
|
|
80
|
+
| | |
|
|
81
|
+
| --- | --- |
|
|
82
|
+
| P1 session wire | `x-client-request-id` from `sessionId` |
|
|
83
|
+
| Mandatory | no |
|
|
84
|
+
| P2 default cache | none (no Prism cache markers) |
|
|
85
|
+
|
|
86
|
+
See [Provider request policies](../provider-request-policies.md).
|
|
87
|
+
|
|
76
88
|
## Security and performance notes
|
|
77
89
|
|
|
78
90
|
- No network during import/setup/default tests; credentials host-owned and late-bound.
|
|
@@ -80,6 +92,10 @@ api.registerProviderPackage(createGoogleProviderPackage({ apiKey: hostKey, model
|
|
|
80
92
|
- Media bounds reuse shared provider media helpers; tool args arrive complete per chunk (no partial JSON reconstruction required).
|
|
81
93
|
- Offline conformance: `@arnilo/prism/testing/provider-conformance`.
|
|
82
94
|
|
|
95
|
+
## Thinking and reasoning
|
|
96
|
+
|
|
97
|
+
Google models route through the `google` family: the adapter merges `compat.thinkingLevel` and the provider emits `generationConfig.thinkingConfig`. Gemini 3.x models use `thinkingLevel` with declared per-model sets: 3.6/3.5-flash and 3-flash-preview accept `minimal`–`high`; 3.1-pro accepts `low/medium/high` (default `high`); 3-pro accepts `low/high`. Gemini 2.5 models are budget-only (`compat.thinkingBudgetRange`): 2.5-pro `128–32768` (cannot disable), 2.5-flash/flash-lite `0–24576` (`0` disables). `none` on a budget-only model maps to the range minimum (`thinkingBudget: 0` where disabling is supported, `128` where not); non-none levels are dropped on budget-only models. Declared level sets snap via nearest-declared (ties up), so `none`/`minimal` on 3.1-pro snap up to `low`. See [Thinking and reasoning](../thinking-and-reasoning.md).
|
|
98
|
+
|
|
83
99
|
## Related APIs
|
|
84
100
|
|
|
85
101
|
- [Google Vertex AI](vertex.md): enterprise ADC/workload-identity package (separate from this consumer API-key package).
|