@arnilo/prism 0.4.0 → 0.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (181) hide show
  1. package/CHANGELOG.md +41 -1
  2. package/README.md +23 -20
  3. package/dist/agent-run-state.d.ts +1 -2
  4. package/dist/agent-run-state.js +0 -3
  5. package/dist/agent-session/session/assemble.d.ts +6 -0
  6. package/dist/agent-session/session/assemble.js +391 -0
  7. package/dist/agent-session/session/persist.d.ts +28 -0
  8. package/dist/agent-session/session/persist.js +166 -0
  9. package/dist/agent-session/session/provider-round.d.ts +6 -0
  10. package/dist/agent-session/session/provider-round.js +231 -0
  11. package/dist/agent-session/session/tool-round.d.ts +31 -0
  12. package/dist/agent-session/session/tool-round.js +473 -0
  13. package/dist/agent-session/session/types.d.ts +115 -0
  14. package/dist/agent-session/session/types.js +5 -0
  15. package/dist/agent-session/session.d.ts +49 -43
  16. package/dist/agent-session/session.js +24 -1180
  17. package/dist/capture.d.ts +63 -0
  18. package/dist/capture.js +67 -0
  19. package/dist/cli-init.d.ts +18 -2
  20. package/dist/cli-init.js +2 -7
  21. package/dist/cli-runner.d.ts +2 -2
  22. package/dist/cli-runner.js +45 -9
  23. package/dist/content.d.ts +3 -3
  24. package/dist/content.js +3 -1
  25. package/dist/contracts-core/agent.d.ts +4 -0
  26. package/dist/contracts-core/batch.d.ts +97 -0
  27. package/dist/contracts-core/batch.js +65 -0
  28. package/dist/contracts-core/content.d.ts +72 -1
  29. package/dist/contracts-core/embeddings.d.ts +30 -0
  30. package/dist/contracts-core/embeddings.js +17 -0
  31. package/dist/contracts-core/images.d.ts +60 -0
  32. package/dist/contracts-core/images.js +17 -0
  33. package/dist/contracts-core/moderation.d.ts +46 -0
  34. package/dist/contracts-core/moderation.js +34 -0
  35. package/dist/contracts-core/speech.d.ts +39 -0
  36. package/dist/contracts-core/speech.js +17 -0
  37. package/dist/contracts-core/transcription.d.ts +48 -0
  38. package/dist/contracts-core/transcription.js +17 -0
  39. package/dist/contracts-core/video.d.ts +61 -0
  40. package/dist/contracts-core/video.js +17 -0
  41. package/dist/contracts-core.d.ts +7 -0
  42. package/dist/contracts-core.js +7 -0
  43. package/dist/contracts-protocol.d.ts +2 -0
  44. package/dist/index.d.ts +7 -5
  45. package/dist/index.js +5 -4
  46. package/dist/input.js +3 -2
  47. package/dist/node/agent-definitions.d.ts +1 -8
  48. package/dist/node/agent-definitions.js +0 -34
  49. package/dist/node/settings.d.ts +0 -1
  50. package/dist/node/settings.js +0 -5
  51. package/dist/pinned-fetch.js +29 -3
  52. package/dist/provider-events.js +3 -4
  53. package/dist/provider-request-policy.d.ts +15 -0
  54. package/dist/provider-request-policy.js +52 -0
  55. package/dist/providers/media.d.ts +1 -2
  56. package/dist/providers/media.js +1 -4
  57. package/dist/rpc.d.ts +1 -1
  58. package/dist/rpc.js +4 -4
  59. package/dist/testing/provider-conformance.d.ts +114 -5
  60. package/dist/testing/provider-conformance.js +342 -0
  61. package/dist/testing/tool-effect-store-conformance.d.ts +0 -1
  62. package/dist/testing/tool-effect-store-conformance.js +0 -3
  63. package/dist/thinking.d.ts +48 -9
  64. package/dist/thinking.js +134 -8
  65. package/docs/0.1.0-readiness.md +3 -3
  66. package/docs/a2a.md +2 -2
  67. package/docs/acp.md +3 -3
  68. package/docs/ag-ui-adoption.md +1 -1
  69. package/docs/ag-ui.md +1 -2
  70. package/docs/agent-definitions.md +1 -1
  71. package/docs/agent-events.md +5 -5
  72. package/docs/agent-identity.md +13 -2
  73. package/docs/agent-session-runtime.md +2 -1
  74. package/docs/audit-export.md +3 -3
  75. package/docs/batch-jobs.md +120 -0
  76. package/docs/cli-rpc.md +20 -9
  77. package/docs/coding-agent-tools.md +19 -19
  78. package/docs/coding-review-and-diagnostics.md +2 -2
  79. package/docs/coding-security.md +4 -4
  80. package/docs/coding-workspaces.md +2 -2
  81. package/docs/compaction-llm.md +2 -0
  82. package/docs/compaction-observational-memory.md +3 -0
  83. package/docs/computer-use-linux.md +13 -2
  84. package/docs/context-and-skills.md +1 -1
  85. package/docs/conversations.md +4 -4
  86. package/docs/credential-storage.md +11 -7
  87. package/docs/credentials-and-redaction.md +1 -1
  88. package/docs/data-classification.md +1 -1
  89. package/docs/database-persistence.md +4 -4
  90. package/docs/dev-inspector.md +6 -6
  91. package/docs/device-adapters.md +2 -2
  92. package/docs/diagrams.md +1 -1
  93. package/docs/document-reader.md +6 -6
  94. package/docs/documents.md +5 -4
  95. package/docs/embeddings.md +112 -0
  96. package/docs/enterprise-postgres-state.md +7 -7
  97. package/docs/evaluations.md +8 -8
  98. package/docs/extensions.md +3 -3
  99. package/docs/forge-integration.md +3 -3
  100. package/docs/graft.md +2 -2
  101. package/docs/guardrails.md +1 -1
  102. package/docs/host-security.md +15 -15
  103. package/docs/image-generation.md +129 -0
  104. package/docs/impeccable.md +5 -3
  105. package/docs/index.md +64 -36
  106. package/docs/indexed-code-search.md +2 -2
  107. package/docs/input-and-prompt-assembly.md +1 -1
  108. package/docs/language-intelligence.md +4 -4
  109. package/docs/live-testing.md +126 -0
  110. package/docs/mcp-tools.md +43 -12
  111. package/docs/middleware-hooks.md +1 -1
  112. package/docs/migrate-to-0.4.md +3 -3
  113. package/docs/migrate-to-0.5.md +144 -0
  114. package/docs/migration.md +33 -1
  115. package/docs/model-registry.md +38 -0
  116. package/docs/model-routing.md +5 -5
  117. package/docs/moderation.md +117 -0
  118. package/docs/multi-agent-patterns.md +4 -4
  119. package/docs/multimodal-content.md +26 -2
  120. package/docs/obscura.md +2 -2
  121. package/docs/observability.md +32 -7
  122. package/docs/openapi-tools.md +13 -3
  123. package/docs/operations.md +11 -0
  124. package/docs/performance.md +7 -7
  125. package/docs/persistence-credentials-multimodality-primitives.md +6 -6
  126. package/docs/policy-and-audit.md +17 -7
  127. package/docs/ponytail.md +1 -1
  128. package/docs/postgres-persistence.md +5 -5
  129. package/docs/process-sessions.md +2 -2
  130. package/docs/prompt-registry.md +7 -7
  131. package/docs/provider-caching.md +8 -2
  132. package/docs/provider-conformance.md +23 -1
  133. package/docs/provider-packages.md +49 -17
  134. package/docs/provider-primitives.md +1 -1
  135. package/docs/provider-request-policies.md +19 -6
  136. package/docs/providers/ai-sdk.md +27 -3
  137. package/docs/providers/alibaba.md +17 -1
  138. package/docs/providers/anthropic.md +16 -0
  139. package/docs/providers/azure.md +29 -1
  140. package/docs/providers/bedrock.md +27 -0
  141. package/docs/providers/clinepass.md +16 -0
  142. package/docs/providers/commandcode.md +265 -0
  143. package/docs/providers/deepseek.md +16 -0
  144. package/docs/providers/google.md +16 -0
  145. package/docs/providers/hyper.md +296 -0
  146. package/docs/providers/kimi.md +16 -0
  147. package/docs/providers/neuralwatt.md +16 -0
  148. package/docs/providers/ollama.md +27 -0
  149. package/docs/providers/openai-compatible.md +16 -0
  150. package/docs/providers/openai.md +16 -0
  151. package/docs/providers/opencode-go.md +16 -0
  152. package/docs/providers/openrouter.md +17 -1
  153. package/docs/providers/vertex.md +28 -0
  154. package/docs/providers/xai.md +16 -0
  155. package/docs/providers/zai.md +16 -0
  156. package/docs/public-contracts.md +1 -1
  157. package/docs/rag.md +26 -4
  158. package/docs/release-and-install.md +103 -46
  159. package/docs/resource-loading.md +1 -1
  160. package/docs/runs-and-usage.md +14 -2
  161. package/docs/server.md +5 -5
  162. package/docs/settings-auth-trust-security.md +7 -5
  163. package/docs/sheets.md +2 -2
  164. package/docs/speech.md +126 -0
  165. package/docs/sqlite-persistence.md +4 -4
  166. package/docs/supervisors.md +3 -3
  167. package/docs/thinking-and-reasoning.md +99 -61
  168. package/docs/tool-conformance.md +1 -1
  169. package/docs/tool-execution-primitives.md +8 -8
  170. package/docs/tools.md +4 -4
  171. package/docs/use-case-model-selection.md +1 -1
  172. package/docs/web-tools.md +1 -1
  173. package/docs/wiki.md +1 -1
  174. package/docs/work-artifacts-and-review.md +17 -6
  175. package/docs/work-connectors.md +4 -4
  176. package/docs/work-tools.md +5 -5
  177. package/docs/workflow-orchestration-primitives.md +11 -11
  178. package/docs/workflows.md +5 -5
  179. package/package.json +11 -8
  180. package/templates/init/providers.json +24 -8
  181. package/docs/antigravity-agent.md +0 -207
@@ -57,6 +57,18 @@ Live canaries stay opt-in behind host credentials; default tests are network-fre
57
57
 
58
58
  Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hosts needing Converse-only models should supply a custom provider or AI SDK bridge.
59
59
 
60
+ ## Request construction (0.5.1)
61
+
62
+ Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
63
+
64
+ | | |
65
+ | --- | --- |
66
+ | P1 session wire | none |
67
+ | Mandatory | no |
68
+ | P2 default cache | host-owned, no Prism cache fields |
69
+
70
+ See [Provider request policies](../provider-request-policies.md).
71
+
60
72
  ## Security and performance notes
61
73
 
62
74
  - No AWS SDK; package-local SigV4 only for `bedrock` service.
@@ -66,6 +78,21 @@ Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hos
66
78
  - Credential secrets are redacted from provider errors.
67
79
  - No credential prefetch at import.
68
80
 
81
+ ## Live probe
82
+
83
+ Opt-in smoke over real AWS Bedrock (package-local SigV4, static keys or session token):
84
+
85
+ ```bash
86
+ PRISM_LIVE_PROVIDER_TESTS=1 AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=us-east-1 \
87
+ node --test packages/prism-providers/dist/bedrock/__tests__/live.test.js
88
+ ```
89
+
90
+ `PRISM_LIVE_BEDROCK_MODEL` overrides the probed model (default `us.anthropic.claude-haiku-4-5-20251001-v1:0`). Without credentials the suite skips.
91
+
92
+ ## Thinking and reasoning
93
+
94
+ Bedrock OpenAI-compat chat expects snake_case `reasoning_effort` (with `effort`/`reasoningEffort` aliases) or a sanitized `reasoning` object. OpenAI-family models on Bedrock snap effort to their declared levels (gpt-5.1 → `none/low/medium/high`); non-OpenAI models pass through untouched. See [Thinking and reasoning](../thinking-and-reasoning.md).
95
+
69
96
  ## Related APIs
70
97
 
71
98
  - [OpenAI-compatible provider](openai-compatible.md)
@@ -98,6 +98,18 @@ await session.prompt("Plan the refactor", {
98
98
  - Multi-backend gateway: key compat off `api.cline.bot`, not the upstream vendor.
99
99
  - Reference USD-per-million costs are catalog metadata; ClinePass itself is a subscription.
100
100
 
101
+ ## Request construction (0.5.1)
102
+
103
+ Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
104
+
105
+ | | |
106
+ | --- | --- |
107
+ | P1 session wire | none; strips `cache_control` |
108
+ | Mandatory | no |
109
+ | P2 default cache | implicit, no markers |
110
+
111
+ See [Provider request policies](../provider-request-policies.md).
112
+
101
113
  ## Security and performance notes
102
114
 
103
115
  - No network on import, setup, build, or default tests.
@@ -106,6 +118,10 @@ await session.prompt("Plan the refactor", {
106
118
  - Provider-owned headers win. One POST per generate. Bounded error bodies.
107
119
  - Live tests: `PRISM_LIVE_PROVIDER_TESTS=1` plus `CLINE_API_KEY`.
108
120
 
121
+ ## Thinking and reasoning
122
+
123
+ ClinePass routes through `reasoning_effort` with per-model slot maps (`compat.thinkingLevelMap`) as wire authority; declared `capabilities.thinkingLevels` mirror each map's portable slots (e.g. GLM: `none/low/medium/high/xhigh`). Portable `max` never reaches the wire (upstream 500s) — the map sends `high`; unsupported slots omit the field. See [Thinking and reasoning](../thinking-and-reasoning.md).
124
+
109
125
  ## Related APIs
110
126
 
111
127
  - [Provider packages](../provider-packages.md)
@@ -0,0 +1,265 @@
1
+ # Command Code provider package
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-providers/commandcode` provides explicit, side-effect-free setup
6
+ for the [Command Code Provider API](https://commandcode.ai/docs/provider) — an
7
+ aggregator exposing every top commercial and open model through OpenAI- and
8
+ Anthropic-compatible endpoints, billed at cost with deals auto-applied. The
9
+ package dual-routes by `ModelConfig.compat.route`:
10
+
11
+ | Route | Endpoint | Official model families |
12
+ | --- | --- | --- |
13
+ | `"openai"` (default) | `POST {baseUrl}/chat/completions` | everything except `claude-*` (GPT-5.6, DeepSeek, Kimi, GLM, MiniMax, Qwen, MiMo, Gemini flash, Grok) |
14
+ | `"anthropic"` | `POST {baseUrl}/messages` | `claude-*` tiers (Opus/Sonnet/Fable/Haiku) |
15
+
16
+ Default base URL is the official Provider API root:
17
+
18
+ ```txt
19
+ https://api.commandcode.ai/provider/v1
20
+ ```
21
+
22
+ Authentication: `Authorization: Bearer <key>` on the chat route,
23
+ `x-api-key` + `anthropic-version: 2023-06-01` on the messages route (Claude
24
+ Code compatibility). The same key authenticates the CLI and the API.
25
+
26
+ ## When to use it
27
+
28
+ Use it when a host app wants Command Code models through Prism's `AgentSession`
29
+ runtime: dual-route serialization, `cache_control` breakpoints on Claude
30
+ models, implicit caching elsewhere, reasoning replay, caller-gated model
31
+ discovery, and optional zero-data-retention (`zdr: true` → provider-owned
32
+ `x-cmd-zdr: 1`, which routes only through ZDR-capable upstreams).
33
+
34
+ Do not use it for automatic credential discovery, setup-time catalog fetches,
35
+ or real-network tests (live probes are operator-gated, see below).
36
+
37
+ ## Inputs / request
38
+
39
+ ```ts
40
+ import {
41
+ createCommandCodeProviderPackage,
42
+ listCommandCodeModels,
43
+ } from "@arnilo/prism-providers/commandcode";
44
+
45
+ createCommandCodeProviderPackage(options: CommandCodeProviderPackageOptions): ProviderPackage
46
+ ```
47
+
48
+ | Field | Type | Purpose |
49
+ | --- | --- | --- |
50
+ | `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
51
+ | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
52
+ | `baseUrl` | `string` | Overrides official `https://api.commandcode.ai/provider/v1`. |
53
+ | `models` | `readonly ModelConfig[]` | Overrides featured `commandCodeModels` defaults. |
54
+ | `zdr` | `boolean` | Enforce zero data retention (`x-cmd-zdr: 1`). May route to costlier upstreams or fail `422 cmd_zdr_no_providers`. |
55
+
56
+ `ProviderRequest.options.cache.breakpoints` select messages-route
57
+ `cache_control` markers for Claude models (max 4, no `ttl`).
58
+
59
+ ## Outputs / response / events
60
+
61
+ | Surface | Behavior |
62
+ | --- | --- |
63
+ | Provider stream | Prism text, thinking, tool-call delta/final, `usage`, `done`, redacted `error`. |
64
+ | Stream completion | `done` only on completion evidence — chat route: `[DONE]` marker plus terminal `finish_reason`; messages route: `message_stop`. Truncated streams end with terminal `error`. |
65
+ | OpenAI thinking | `delta.reasoning_content` → thinking deltas; replay via top-level `reasoning_content` when `preserveThinking` (default), never folded into text. |
66
+ | Anthropic thinking | `thinking_delta` → thinking deltas; replay via Anthropic thinking blocks when `preserveThinking`. |
67
+ | Usage | Standard tokens + cache read/write per route; mapped through the shared OpenAI/Anthropic usage mappings. |
68
+ | Auth method | `api_key` for `commandcode`, credential name `apiKey`. |
69
+
70
+ ## Request/response example
71
+
72
+ ```json
73
+ {
74
+ "authorization": "Bearer cmd_…",
75
+ "content-type": "application/json"
76
+ }
77
+ ```
78
+
79
+ Messages route instead sends provider-owned `x-api-key: <key>` and
80
+ `anthropic-version: 2023-06-01`. All provider-owned headers are applied after
81
+ caller headers and cannot be overridden.
82
+
83
+ Chat-route body (thinking passthrough + preserved reasoning):
84
+
85
+ ```json
86
+ {
87
+ "model": "Qwen/Qwen3.8-Flash",
88
+ "stream": true,
89
+ "stream_options": { "include_usage": true },
90
+ "max_tokens": 512,
91
+ "messages": [
92
+ {
93
+ "role": "assistant",
94
+ "tool_calls": [{ "id": "call_1", "type": "function", "function": { "name": "lookup", "arguments": "{}" } }],
95
+ "reasoning_content": "plan the lookup"
96
+ }
97
+ ]
98
+ }
99
+ ```
100
+
101
+ ## Implementation example
102
+
103
+ ```ts
104
+ import { createExtensionKernel } from "@arnilo/prism";
105
+ import { createCommandCodeProviderPackage } from "@arnilo/prism-providers/commandcode";
106
+
107
+ const kernel = createExtensionKernel();
108
+ await kernel.load([createCommandCodeProviderPackage({ apiKey: process.env.COMMAND_CODE_API_KEY })]);
109
+ ```
110
+
111
+ Caller-gated live catalog (never runs during package setup):
112
+
113
+ ```ts
114
+ const models = await listCommandCodeModels({ fetch }); // public endpoint, no auth needed
115
+ await kernel.load([createCommandCodeProviderPackage({ apiKey: process.env.COMMAND_CODE_API_KEY, models })]);
116
+ ```
117
+
118
+ ## Featured models and routes
119
+
120
+ Featured `commandCodeModels` is a curated 38-model bootstrap catalog: ids and
121
+ context windows from the live `GET /provider/v1/models` snapshot (2026-09, 67
122
+ ids), USD-per-million-token pricing from the docs table
123
+ (<https://commandcode.ai/docs/resources/pricing-limits>). `compat.pricing_source`
124
+ records caveats: open-source models bill at the **mean per-provider price**;
125
+ DeepSeek rates are **off-peak** (17h/day; peak 2× during 01–04 & 06–10 UTC);
126
+ deals (MiniMax M3 −50%, MiMo −98/99%) are already applied upstream. Custom
127
+ pricing metadata is always stripped before the wire.
128
+
129
+ | Model family | Route | Cache kind |
130
+ | --- | --- | --- |
131
+ | `claude-opus-5/4-8/4-7`, `claude-sonnet-5/4-6`, `claude-fable-5-1/5`, `claude-haiku-4-5` | `anthropic` | `cache_control` (max 4 breakpoints, no `ttl` — undocumented) |
132
+ | `gpt-5.6-sol/terra/luna` | `openai` | `implicit` (docs cache-write price recorded in `cost.cacheWrite`; explicit-key upgrade gated on live probe — Task 9) |
133
+ | `deepseek/*`, Kimi, GLM, MiniMax, Qwen, MiMo, Gemini flash, Grok | `openai` | `implicit` |
134
+
135
+ ## Model discovery
136
+
137
+ ```txt
138
+ GET https://api.commandcode.ai/provider/v1/models
139
+ ```
140
+
141
+ Public endpoint — works without authentication and emits no auth header when no
142
+ key resolves. `listCommandCodeModels({ fetch?, baseUrl?, apiKey?, signal?,
143
+ headers? })` maps each `{ id, name, context_length }` entry to `ModelConfig`:
144
+ route from id (`claude-*` → anthropic), context window from the endpoint, and
145
+ featured docs metadata (cost/cache kind) applied when the id matches a curated
146
+ entry. The endpoint carries no pricing or capabilities — unknown ids get
147
+ route-derived cache kind and no cost. Discovery is **caller-gated** — setup
148
+ performs zero fetches.
149
+
150
+ ## Thinking / reasoning
151
+
152
+ | Surface | Behavior |
153
+ | --- | --- |
154
+ | OpenAI route stream | `reasoning_content` → thinking deltas |
155
+ | OpenAI route replay | thinking blocks → top-level `reasoning_content` when `preserveThinking`; never folded into text |
156
+ | Anthropic route stream | `thinking_delta` → thinking deltas |
157
+ | Anthropic route replay | thinking blocks when `preserveThinking` |
158
+
159
+ Owned compat keys (`route`, `preserveThinking`, `pricing_source`) are stripped
160
+ before opaque compat spread so resolved values win.
161
+
162
+ ## Extension and configuration notes
163
+
164
+ - Hosts choose base URL, model list, credential source, `fetch` impl, and ZDR.
165
+ - Route selection is explicit via `compat.route` (`"anthropic"` for `claude-*`
166
+ ids, default `"openai"`). Sending a model to the wrong endpoint 400s.
167
+ - Package contributes models via the extension `api` and an `api_key` auth method.
168
+
169
+ ### Cache and session behavior
170
+
171
+ - The chat route sends **no** `cache_control` fields; it relies on OpenAI-style
172
+ implicit caching (upstream provider behavior, passed through). Read tokens
173
+ map from `prompt_tokens_details.cached_tokens` / `cache_write_tokens` /
174
+ `prompt_cache_hit_tokens`.
175
+ - The messages route applies `cache_control: { type: "ephemeral" }` markers
176
+ only to caller-selected `cache.breakpoints` (shared `applyCacheControl()`
177
+ helper) on the last content block of each selected message. A
178
+ `system_prompt` breakpoint serializes `system` as marked text blocks (plain
179
+ string otherwise). Caching is enabled unless disabled
180
+ (`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
181
+ `ModelConfig.cache.kind: "cache_control"`.
182
+ - **No `ttl` is ever emitted**: the upstream TTL window is undocumented on the
183
+ Provider API; `cacheRetention: "long"` must not produce a marker the gateway
184
+ may reject.
185
+ - Usage accounting per route: chat route maps
186
+ `prompt_tokens_details.cached_tokens`/`cache_write_tokens` (and
187
+ `prompt_cache_hit_tokens`); messages route maps
188
+ `cache_read_input_tokens`/`cache_creation_input_tokens`.
189
+ - Session identity is simple: no session header is emitted (undocumented).
190
+
191
+ ### Live-verified mapping (findings ledger)
192
+
193
+ The following claims are encoded as operator-gated probes in
194
+ `packages/prism-providers/src/commandcode/__tests__/live.test.ts`. Each probe's
195
+ assertion encodes the documented claim, so a probe failure **is** the finding;
196
+ record the outcome here and adjust the mapping. Status: **pending operator
197
+ run** (no key in CI):
198
+
199
+ | # | Claim (documented) | Probe | Status |
200
+ | --- | --- | --- | --- |
201
+ | 1 | Warm chat-route replay reports cached tokens (implicit caching passes through) | `live_chat_route_reports_cached_tokens_on_warm_prefix_replay` | pending |
202
+ | 2 | `cache_control` on messages reports `cache_creation_input_tokens` on the creating call | `live_messages_route_cache_control_reports_creation_and_read_tokens` | pending |
203
+ | 3 | Same-prefix warm replay reads the created cache entry (TTL ≥ one request) | same probe (warm leg) | pending |
204
+ | 4 | OpenAI `prompt_cache_key` is accepted and honored for GPT-5.6 (explicit caching) — decides Task 9 | `live_gpt56_prompt_cache_key_passthrough_probe` | pending — Task 9 closed gated: **pass** → upgrade `gpt-5.6-*` to the OpenAI explicit mapping (`promptCacheKey`/`promptCacheOptions`/`applyPromptCacheBreakpoints` via `@arnilo/prism-providers/openai`, verified exported); **fail** (400/ignored/warm replay shows no cached tokens) → verified-negative, keep `implicit`, record here |
205
+ | 5 | OpenAI `reasoning_effort` is accepted (200) on the chat route | `live_reasoning_effort_is_accepted_on_chat_route` | pending |
206
+ | 6 | ZDR requests route (done) or fail `422 cmd_zdr_no_providers` when no ZDR-capable upstream exists | `live_zdr_route_probe_is_opt_in_and_routable` | pending |
207
+
208
+ Run the gate:
209
+
210
+ ```sh
211
+ PRISM_LIVE_PROVIDER_TESTS=1 COMMAND_CODE_API_KEY=cmd_... \
212
+ npm run test --workspace=@arnilo/prism-providers/commandcode
213
+ ```
214
+
215
+ ## Request construction (0.5.1)
216
+
217
+ Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
218
+
219
+ | | |
220
+ | --- | --- |
221
+ | P1 session wire | none |
222
+ | Mandatory | no |
223
+ | P2 default cache | Anthropic: `cache_control`; OpenAI: implicit (no `prompt_cache_key`) |
224
+
225
+ See [Provider request policies](../provider-request-policies.md).
226
+
227
+ ## Security and performance notes
228
+
229
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers.
230
+ - No network calls during import, setup, build, or default tests.
231
+ - No automatic environment, file, keychain, or shell credential lookup.
232
+ - API keys are resolved per request from caller-supplied values or resolvers
233
+ and redacted from errors (upstream error bodies may carry the upstream
234
+ provider's message — always redacted).
235
+ - `403 upgrade_required` (Go plan — no API access) and
236
+ `422 cmd_zdr_no_providers` are non-retryable; `429` and `5xx` are retryable
237
+ with `retry-after` surfaced as `retry_after_ms`; `400/401/422` are
238
+ non-retryable.
239
+ - Caller headers cannot override provider-owned headers (`content-type`,
240
+ `authorization` on chat, `x-api-key`/`anthropic-version` on messages,
241
+ `x-cmd-zdr` when ZDR is opted in).
242
+ - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus
243
+ `COMMAND_CODE_API_KEY`; default tests are network-free.
244
+
245
+ ## Official evidence
246
+
247
+ - Command Code Provider API docs: `https://commandcode.ai/docs/provider`
248
+ - Pricing & limits (per-model USD, deals, off-peak): `https://commandcode.ai/docs/resources/pricing-limits`
249
+ - Live `GET https://api.commandcode.ai/provider/v1/models` snapshot (2026-09) — 67 ids, context windows
250
+ - Probe ledger above pending operator-gated live run; row 4 decides plan 055 Task 9
251
+ (explicit GPT-5.6 caching upgrade) — see `plans/055-First-Class-Hyper-And-Command-Code-Providers.md`
252
+
253
+ ## Thinking and reasoning
254
+
255
+ Command Code models carry provenance-commented level tables: `claude-*` → `output_config_effort` (Anthropic-route effort sets mirroring the native Anthropic package), `gpt-5.6*` → `openai_reasoning` (`none`–`xhigh`), `deepseek-v4`/`kimi-k3`/`glm-5.3` → `reasoning_effort` (`low/high/max`), `glm-5.2` → `reasoning_effort` (`low`–`max`), Kimi-K2.x/MiniMax/Qwen → `thinking_type`, gemini-3.x → `noop` (the gateway chat route has no `thinking_level` wire, so no levels are declared), mimo/unknown → passthrough. Effort snaps to declared sets on both routes. See [Thinking and reasoning](../thinking-and-reasoning.md).
256
+
257
+ ## Related APIs
258
+
259
+ - [Provider packages](../provider-packages.md): `defineProviderPackage`,
260
+ `ModelConfig`, discovery contract, request/cache policies.
261
+ - [Thinking and reasoning](../thinking-and-reasoning.md): per-turn `ThinkingLevel` → compat families.
262
+ - [Credentials and redaction](../credentials-and-redaction.md):
263
+ `resolveCredentialValue`, `redactSecrets`.
264
+ - [Provider caching](../provider-caching.md): per-provider cache behavior matrix.
265
+ - [Provider conformance](../provider-conformance.md): network-free adapter tests.
@@ -119,6 +119,18 @@ await session.prompt("Plan the refactor", {
119
119
  - Tool-turn assistants must replay `reasoning_content` or the API returns 400.
120
120
  Non-tool multi-turn may omit it (the API ignores it).
121
121
 
122
+ ## Request construction (0.5.1)
123
+
124
+ Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
125
+
126
+ | | |
127
+ | --- | --- |
128
+ | P1 session wire | none |
129
+ | Mandatory | no |
130
+ | P2 default cache | implicit, no markers |
131
+
132
+ See [Provider request policies](../provider-request-policies.md).
133
+
122
134
  ## Security and performance notes
123
135
 
124
136
  - SSE streams and HTTP error bodies use bounded transport helpers.
@@ -129,6 +141,10 @@ await session.prompt("Plan the refactor", {
129
141
  - One POST per generate. No provider retry loop.
130
142
  - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus `DEEPSEEK_API_KEY`.
131
143
 
144
+ ## Thinking and reasoning
145
+
146
+ DeepSeek models declare `low/high/max` and stamp `reasoning_effort`; the wire table maps `medium`/`xhigh`→`high`, `none`/`minimal` stop thinking. `thinking.type: "enabled"/"disabled"` stays available (thinking on by default, `high`); a request-level `reasoning_effort: none` stops thinking only when no explicit `thinking` switch was sent. Tool turns must replay `reasoning_content` or the API returns 400. See [Thinking and reasoning](../thinking-and-reasoning.md).
147
+
132
148
  ## Related APIs
133
149
 
134
150
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
@@ -73,6 +73,18 @@ api.registerProviderPackage(createGoogleProviderPackage({ apiKey: hostKey, model
73
73
  - Vertex / enterprise identity stays out of 0.0.11.
74
74
  - Gemini CLI says third-party software accessing its backend through Gemini CLI OAuth violates applicable terms, and its FAQ directs third-party coding agents to Vertex AI or Google AI Studio API keys ([terms](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/tos-privacy.md), [FAQ](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/faq.md)). Prism therefore has no Gemini CLI OAuth API or token-import shortcut.
75
75
 
76
+ ## Request construction (0.5.1)
77
+
78
+ Agent sessions stamp `sessionId`/`cacheKey` without a host policy. Session/cache keys are correlation ids, never secrets.
79
+
80
+ | | |
81
+ | --- | --- |
82
+ | P1 session wire | `x-client-request-id` from `sessionId` |
83
+ | Mandatory | no |
84
+ | P2 default cache | none (no Prism cache markers) |
85
+
86
+ See [Provider request policies](../provider-request-policies.md).
87
+
76
88
  ## Security and performance notes
77
89
 
78
90
  - No network during import/setup/default tests; credentials host-owned and late-bound.
@@ -80,6 +92,10 @@ api.registerProviderPackage(createGoogleProviderPackage({ apiKey: hostKey, model
80
92
  - Media bounds reuse shared provider media helpers; tool args arrive complete per chunk (no partial JSON reconstruction required).
81
93
  - Offline conformance: `@arnilo/prism/testing/provider-conformance`.
82
94
 
95
+ ## Thinking and reasoning
96
+
97
+ Google models route through the `google` family: the adapter merges `compat.thinkingLevel` and the provider emits `generationConfig.thinkingConfig`. Gemini 3.x models use `thinkingLevel` with declared per-model sets: 3.6/3.5-flash and 3-flash-preview accept `minimal`–`high`; 3.1-pro accepts `low/medium/high` (default `high`); 3-pro accepts `low/high`. Gemini 2.5 models are budget-only (`compat.thinkingBudgetRange`): 2.5-pro `128–32768` (cannot disable), 2.5-flash/flash-lite `0–24576` (`0` disables). `none` on a budget-only model maps to the range minimum (`thinkingBudget: 0` where disabling is supported, `128` where not); non-none levels are dropped on budget-only models. Declared level sets snap via nearest-declared (ties up), so `none`/`minimal` on 3.1-pro snap up to `low`. See [Thinking and reasoning](../thinking-and-reasoning.md).
98
+
83
99
  ## Related APIs
84
100
 
85
101
  - [Google Vertex AI](vertex.md): enterprise ADC/workload-identity package (separate from this consumer API-key package).