@arnilo/prism 0.0.96 → 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (203) hide show
  1. package/CHANGELOG.md +285 -2
  2. package/README.md +17 -3
  3. package/dist/agent-definitions.js +2 -3
  4. package/dist/agent-event-source.d.ts +11 -0
  5. package/dist/agent-event-source.js +512 -0
  6. package/dist/agent-loops.d.ts +5 -0
  7. package/dist/agent-loops.js +99 -14
  8. package/dist/agent-run-lifecycle.d.ts +5 -2
  9. package/dist/agent-run-lifecycle.js +18 -2
  10. package/dist/agent-run-state.d.ts +27 -1
  11. package/dist/agent-run-state.js +113 -7
  12. package/dist/agents.d.ts +3 -1
  13. package/dist/agents.js +1255 -129
  14. package/dist/artifacts.d.ts +132 -0
  15. package/dist/artifacts.js +44 -0
  16. package/dist/cache-helpers.js +18 -9
  17. package/dist/checkpoints.d.ts +4 -0
  18. package/dist/checkpoints.js +17 -9
  19. package/dist/cli-init.js +3 -7
  20. package/dist/cli-runner.d.ts +2 -6
  21. package/dist/cli-runner.js +71 -33
  22. package/dist/compaction.js +5 -4
  23. package/dist/config.js +7 -4
  24. package/dist/content.js +26 -24
  25. package/dist/context-budget.d.ts +67 -0
  26. package/dist/context-budget.js +288 -0
  27. package/dist/contracts.d.ts +590 -8
  28. package/dist/contracts.js +142 -1
  29. package/dist/contribution-parsing.js +6 -2
  30. package/dist/contributions.d.ts +2 -0
  31. package/dist/contributions.js +3 -0
  32. package/dist/conversations.d.ts +50 -0
  33. package/dist/conversations.js +98 -0
  34. package/dist/credentials.d.ts +22 -2
  35. package/dist/credentials.js +18 -3
  36. package/dist/devices.d.ts +94 -0
  37. package/dist/devices.js +138 -0
  38. package/dist/event-multiplexer.js +18 -4
  39. package/dist/extensions.d.ts +18 -1
  40. package/dist/extensions.js +79 -6
  41. package/dist/feedback.js +12 -10
  42. package/dist/guardrails.d.ts +1 -1
  43. package/dist/guardrails.js +26 -17
  44. package/dist/identity.d.ts +92 -0
  45. package/dist/identity.js +265 -0
  46. package/dist/index.d.ts +94 -72
  47. package/dist/index.js +48 -36
  48. package/dist/input.d.ts +10 -1
  49. package/dist/input.js +152 -52
  50. package/dist/instruction-injection.d.ts +1 -1
  51. package/dist/middleware.js +9 -1
  52. package/dist/models.d.ts +2 -0
  53. package/dist/models.js +3 -0
  54. package/dist/node/agent-definitions.js +16 -8
  55. package/dist/node/contribution-discovery.d.ts +1 -2
  56. package/dist/node/contribution-discovery.js +3 -3
  57. package/dist/node/session-store-jsonl.js +13 -7
  58. package/dist/node/settings.d.ts +1 -1
  59. package/dist/node/settings.js +1 -1
  60. package/dist/node/system-project-prompts.js +2 -4
  61. package/dist/node/trust.js +1 -1
  62. package/dist/persistence-lifecycle.d.ts +103 -0
  63. package/dist/persistence-lifecycle.js +202 -0
  64. package/dist/provider-events.d.ts +1 -0
  65. package/dist/provider-events.js +6 -1
  66. package/dist/provider-request-policy.js +3 -4
  67. package/dist/providers/media.d.ts +1 -1
  68. package/dist/providers/openai-compatible.d.ts +46 -1
  69. package/dist/providers/openai-compatible.js +123 -53
  70. package/dist/providers/openai-primitives.js +10 -7
  71. package/dist/providers/transport.d.ts +6 -0
  72. package/dist/providers/transport.js +21 -0
  73. package/dist/providers.d.ts +2 -0
  74. package/dist/providers.js +3 -0
  75. package/dist/redaction.d.ts +1 -0
  76. package/dist/redaction.js +26 -9
  77. package/dist/resources.d.ts +2 -2
  78. package/dist/resources.js +2 -2
  79. package/dist/retry.d.ts +5 -0
  80. package/dist/retry.js +8 -1
  81. package/dist/rpc.js +55 -11
  82. package/dist/run-ledger.d.ts +6 -0
  83. package/dist/run-ledger.js +16 -13
  84. package/dist/run-limits.js +49 -10
  85. package/dist/secure-agent.js +8 -2
  86. package/dist/security.js +7 -2
  87. package/dist/session-stores.d.ts +7 -2
  88. package/dist/session-stores.js +195 -21
  89. package/dist/skill-disclosure.d.ts +35 -0
  90. package/dist/skill-disclosure.js +101 -0
  91. package/dist/skill-load.d.ts +25 -0
  92. package/dist/skill-load.js +112 -0
  93. package/dist/structured-output.d.ts +5 -1
  94. package/dist/structured-output.js +20 -2
  95. package/dist/system-prompts.js +7 -2
  96. package/dist/testing/agent-event-source-conformance.d.ts +4 -0
  97. package/dist/testing/agent-event-source-conformance.js +54 -0
  98. package/dist/testing/compaction-conformance.js +5 -1
  99. package/dist/testing/extension-conformance.js +15 -3
  100. package/dist/testing/feedback.d.ts +1 -3
  101. package/dist/testing/feedback.js +1 -1
  102. package/dist/testing/persistence-schema.d.ts +2 -2
  103. package/dist/testing/persistence-schema.js +280 -35
  104. package/dist/testing/provider-conformance.js +3 -3
  105. package/dist/testing/run-ledger-conformance.js +1 -1
  106. package/dist/testing/session-store-conformance.d.ts +6 -0
  107. package/dist/testing/session-store-conformance.js +37 -2
  108. package/dist/testing/tool-conformance.js +30 -5
  109. package/dist/testing/tool-effect-store-conformance.d.ts +9 -0
  110. package/dist/testing/tool-effect-store-conformance.js +85 -0
  111. package/dist/thinking.js +4 -1
  112. package/dist/tool-effects.d.ts +15 -0
  113. package/dist/tool-effects.js +352 -0
  114. package/dist/tool-result-fold.d.ts +40 -0
  115. package/dist/tool-result-fold.js +176 -0
  116. package/dist/tools.d.ts +8 -3
  117. package/dist/tools.js +248 -13
  118. package/docs/0.1.0-readiness.md +202 -0
  119. package/docs/a2a.md +33 -2
  120. package/docs/acp.md +126 -0
  121. package/docs/ag-ui-adoption.md +77 -0
  122. package/docs/ag-ui.md +225 -0
  123. package/docs/agent-events.md +34 -3
  124. package/docs/agent-identity.md +144 -0
  125. package/docs/agent-loops.md +17 -2
  126. package/docs/agent-session-runtime.md +21 -4
  127. package/docs/browser-automation.md +5 -0
  128. package/docs/caveman.md +129 -0
  129. package/docs/cli-rpc.md +3 -6
  130. package/docs/coding-agent-tools.md +229 -25
  131. package/docs/coding-security.md +77 -11
  132. package/docs/compaction-and-retry.md +5 -2
  133. package/docs/compaction-llm.md +20 -1
  134. package/docs/compaction-observational-memory.md +52 -8
  135. package/docs/context-and-skills.md +94 -7
  136. package/docs/contribution-registries.md +1 -0
  137. package/docs/conversations.md +135 -0
  138. package/docs/credential-storage.md +34 -1
  139. package/docs/credentials-and-redaction.md +11 -1
  140. package/docs/database-persistence.md +27 -7
  141. package/docs/device-adapters.md +97 -0
  142. package/docs/enterprise-postgres-state.md +178 -0
  143. package/docs/evaluations.md +14 -1
  144. package/docs/extensions.md +4 -1
  145. package/docs/forge-integration.md +113 -0
  146. package/docs/guardrails.md +16 -2
  147. package/docs/host-security.md +35 -4
  148. package/docs/index.md +69 -37
  149. package/docs/input-and-prompt-assembly.md +8 -7
  150. package/docs/language-intelligence.md +162 -0
  151. package/docs/mcp-tools.md +62 -5
  152. package/docs/middleware-hooks.md +2 -2
  153. package/docs/migration.md +423 -2
  154. package/docs/model-routing.md +111 -0
  155. package/docs/multimodal-content.md +8 -5
  156. package/docs/node-jsonl-session-store.md +1 -1
  157. package/docs/observability.md +2 -0
  158. package/docs/openapi-tools.md +56 -0
  159. package/docs/performance.md +282 -0
  160. package/docs/policy-and-audit.md +171 -0
  161. package/docs/ponytail.md +127 -0
  162. package/docs/postgres-persistence.md +8 -4
  163. package/docs/process-sessions.md +147 -0
  164. package/docs/provider-caching.md +13 -1
  165. package/docs/provider-conformance.md +29 -5
  166. package/docs/provider-packages.md +43 -2
  167. package/docs/provider-request-policies.md +2 -0
  168. package/docs/providers/ai-sdk.md +24 -7
  169. package/docs/providers/alibaba.md +179 -0
  170. package/docs/providers/anthropic.md +93 -0
  171. package/docs/providers/azure.md +74 -0
  172. package/docs/providers/bedrock.md +72 -0
  173. package/docs/providers/google.md +89 -0
  174. package/docs/providers/ollama.md +166 -0
  175. package/docs/providers/openai-compatible.md +31 -2
  176. package/docs/providers/openai.md +24 -5
  177. package/docs/providers/openrouter.md +2 -0
  178. package/docs/providers/vertex.md +71 -0
  179. package/docs/public-contracts.md +61 -4
  180. package/docs/rag.md +41 -12
  181. package/docs/release-and-install.md +323 -206
  182. package/docs/resource-loading.md +3 -0
  183. package/docs/runs-and-usage.md +3 -0
  184. package/docs/server.md +44 -6
  185. package/docs/session-store-conformance.md +2 -0
  186. package/docs/session-stores.md +41 -2
  187. package/docs/sqlite-persistence.md +11 -3
  188. package/docs/structured-output.md +7 -1
  189. package/docs/supervisors.md +8 -0
  190. package/docs/tool-effects.md +95 -0
  191. package/docs/tools.md +5 -0
  192. package/docs/work-artifacts-and-review.md +102 -0
  193. package/docs/work-connectors.md +32 -0
  194. package/docs/work-tools.md +137 -0
  195. package/docs/workflows.md +6 -0
  196. package/docs/working-and-semantic-memory.md +40 -7
  197. package/package.json +30 -8
  198. package/templates/init/providers.json +22 -0
  199. package/docs/review-coverage-2026-07-14.md +0 -260
  200. package/docs/review-coverage-2026-07-15.md +0 -193
  201. package/docs/review-coverage-2026-07-17-provider-validation.md +0 -192
  202. package/docs/review-coverage-2026-07-19-phase-3.md +0 -174
  203. package/docs/review-coverage-2026-07-20-phase-4.md +0 -175
@@ -1,192 +0,0 @@
1
- # Review coverage — 2026-07-17 provider validation
2
-
3
- Working evidence page for Plan 067. Freezes the 2026-07-14 P0–P2 re-verification map, first-party provider validation owners, official-doc priority URLs, Pi secondary references, cache/thinking/discovery surfaces, credential notes, and use-case model-binding inventory.
4
-
5
- **Evidence frozen:** 2026-07-17 (offline inventory; no live provider calls).
6
- **Priority rule:** official provider documentation wins; Pi (`badlogic/pi-mono`) is secondary when official docs are silent or ambiguous.
7
-
8
- Related: [2026-07-14 coverage](review-coverage-2026-07-14.md) (0.0.4), [2026-07-15 coverage](review-coverage-2026-07-15.md) (0.0.5), [provider caching](provider-caching.md), [provider packages](provider-packages.md).
9
-
10
- ## Status legend
11
-
12
- | Status | Meaning |
13
- | --- | --- |
14
- | `verify` | Prior plan marked fixed; Plan 067 must re-verify with regression tests. |
15
- | `gap` | Known missing or incorrect behavior/docs relative to official sources. |
16
- | `fixed` | Closed in this plan (updated as tasks complete). |
17
- | `by-design` | Intentional absence (document; do not “fix” into a catalog). |
18
-
19
- ## P0–P2 finding → Plan 067 owner matrix
20
-
21
- Source: `code-reviews/2026-07-14.md`. Prior Plans 053 / 054 / 058 marked these implemented for 0.0.4; this plan re-verifies rather than skipping.
22
-
23
- | Review ID | Priority | Finding | 0.0.4 owner | Plan 067 task | Current status | Credential / redaction notes |
24
- | --- | --- | --- | --- | --- | --- | --- |
25
- | R-001 | P0 | Revision request duplicated + corrupted by redaction | 053-1 | 1 | fixed | Secret canaries must not appear in repair requests, events, or redacted graphs. Re-verified 2026-07-17: `pendingHistory` + active-path redaction; revision+redactor suite asserts one repair and no `[Circular]`. |
26
- | R-002 | P1 | Multi-round tool transcript chronologically invalid | 053-2 | 1 | fixed | N/A. Re-verified 2026-07-17: two rounds × two calls keep `user → assistant → tool → tool → …` order in history and assembled request. |
27
- | R-003 | P1 | Redactor leaks secrets in object/Map keys | 053-1 | 1 | fixed | Object/Map string keys redact; collisions use deterministic `__N` suffixes. |
28
- | R-004 | P1 | Event-ledger writes lack backpressure | 053-3 | 1 | fixed | Ledger appends serialized (concurrency 1), order preserved, append failures reject run completion. |
29
- | R-008 | P1 | Unbounded SSE / error bodies; multiline `data:` | 054-1/2 | 2 | fixed | Bounded readers; error text redacted. |
30
- | R-009 | P1 | OpenAI device-code OAuth does not poll | 054-3 | 2 | fixed | Redact device/user/access/refresh codes from OAuth errors. |
31
- | R-010 | P2 | Duplicated provider protocol utilities | 054-1/2 | 2 | fixed | Shared transport/primitives remain authoritative. |
32
- | R-005 | P2 | JSONL append silent on corrupt lines | 053-4 | 2 | fixed | Dev-only store; fail closed on corrupt lines. |
33
- | R-011 | P2 | Coding-agent image read unbounded / resize no-op | 055-5 | 2 | fixed | Enforce `maxImageBytes`; deprecate `autoResizeImages` honestly. |
34
- | R-006 | P2 | Optional config ENOENT detected by message text | 053-5 | 2 | fixed | Typed `code === "ENOENT"`. |
35
- | R-012 | P2 | Release tag/version mismatch; no provenance | 058-7/9 | 2 | fixed | Tag/version gate + `--provenance`; no secrets in artifacts. |
36
-
37
- Bug-report fixes A–D remain covered by R-001 / R-003 / R-007 (malformed message shape was Plan 053; re-check under Task 1 if touched).
38
-
39
- ## Provider package validation matrix
40
-
41
- | Package | Plan 067 task | Official cache | Prism `cache.kind` | Thinking / reasoning (official → Prism) | Discovery endpoint | Static catalog (bootstrap only) | Pi secondary ref | Credential surface | Status |
42
- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
43
- | `@arnilo/prism-provider-openai` | 6 | `prompt_cache_key`; older models `prompt_cache_retention` (`24h` / `in_memory`); GPT-5.6+ `prompt_cache_options` / breakpoints tracked (host may pass via compat/extra; discovery sets `longRetention: false`) | `openai_key` (+ `longRetention`, `maxKeyLength`) | Official: Responses top-level `reasoning.effort` (+ `summary`). Prism: merges `model.compat.reasoning` + `options.compat.reasoning` (request wins); `applyThinkingLevel(..., "openai_reasoning")` | Official: `GET /models` | Featured `openAIModels` / `openAICodexModels` + **`listOpenAIModels`**; factory accepts `models?` / `codexModels?` | `packages/ai/src/api/openai-responses.ts`, `openai-responses-shared.ts` | API key; Codex OAuth device-code (`oauth.ts`) — poll + redact codes | **fixed** 2026-07-17: Responses P0s + discovery + reasoning + models override |
44
- | `@arnilo/prism-provider-kimi` | 7 | Anthropic-style `cache_control` on `/messages` when supported; OpenAI Moonshot route none | default implicit; opt-in `cache_control` | Official K2.x: `thinking.type` (`enabled`/`disabled`/`keep`); K2.7-code always enabled; K3: top-level `reasoning_effort`. Prism: `kimiThinking`/`kimiReasoningEffort`/`kimiPreserveThinking` (request wins); Coding Anthropic + Moonshot Chat Completions both callable | Official: `GET https://api.moonshot.ai/v1/models` (also `.cn`); Coding has no public list API | Featured Coding (`kimi-for-coding`, `kimi-for-coding-highspeed`, `k3`) + Moonshot (`kimi-k2.7-code`, `kimi-k3`) + **`listKimiModels`** | `kimi-coding*.ts`, `moonshotai*.ts`, `api/anthropic-messages.ts` | API key (`KIMI` / Moonshot, not interchangeable); redacted from errors | **fixed** 2026-07-17: discovery + Moonshot provider + thinking + official Coding ids |
45
- | `@arnilo/prism-provider-zai` | 8 | Implicit (no client `cache_control` / `prompt_cache_key`); official `usage.prompt_tokens_details.cached_tokens` | `implicit` | Official: `thinking` (`{type, clear_thinking?}`), `reasoning_effort` (GLM-5.2+), `tool_stream` (GLM-4.6+). Prism: `zaiThinking`/`zaiReasoningEffort`/`zaiToolStream`/`zaiClearThinking`/`zaiPreserveThinking` (request wins); Preserved Thinking replays `reasoning_content`. Obsolete `thinkingFormat`/`developerRoleFallback` docs removed | No first-class docs.z.ai list page; OpenAI-compatible `GET {baseUrl}/models` best-effort + curated featured set from Chat Completions enum / overview | Featured `zaiModels` (`glm-5.2`…`glm-4.5`) + **`listZaiModels`**; default base `https://api.z.ai/api/paas/v4` | `zai.ts`, `zai.models.ts` (Pi secondary ids only) | API key; redacted from errors / discovery | **fixed** 2026-07-17: docs drift closed; catalog + discovery + clear_thinking/preserve |
46
- | `@arnilo/prism-provider-openrouter` | 9 | Official prompt caching via `cache_control` + sticky `session_id` routing; top-level automatic when no breakpoints | `cache_control` (legacy `compat.openRouterCache`) | Official: `reasoning: { effort | max_tokens }`. Prism: `resolveOpenRouterReasoning` merge + `preserveThinking` replay as body `reasoning` | Official: `GET https://openrouter.ai/api/v1/models` | **App-controlled** `models:` + optional **`listOpenRouterModels`** (no bundled mega-catalog) | `openrouter.models.ts` (do not vendor), `api/openai-completions.ts` | API key; sanitize session/cache ids | **fixed** 2026-07-17: discovery + reasoning merge/preserve + automatic top-level cache_control |
47
- | `@arnilo/prism-provider-opencode-go` | 10 | Anthropic route: selected `cache_control`; OpenAI route: none; `x-opencode-session` from cache/session key | route-specific (`cache_control` on Anthropic; `implicit` on OpenAI) | Dual-route: Anthropic thinking blocks + OpenAI `reasoning_content`; upstream `thinking`/`reasoning_effort`/`reasoning` passthrough (request wins); `preserveThinking` default for reasoning models | Official `GET https://opencode.ai/zen/go/v1/models` (sparse) | Featured official Go ids (Grok/GLM/Kimi/MiMo/MiniMax/Qwen/DeepSeek) + **`listOpenCodeGoModels`**; default base `https://opencode.ai/zen/go/v1` | `opencode-go.ts`, `opencode-go.models.ts` (Pi secondary ids/limits only) | API key; redacted from errors / discovery | **fixed** 2026-07-18: catalog + discovery + base URL + thinking preserve |
48
- | `@arnilo/prism-provider-neuralwatt` | 11 | Implicit vLLM prefix caching; `prompt_tokens_details.cached_tokens` | `implicit` | Official: `reasoning_effort` when `capabilities.reasoning_effort` (GLM-5.2 default `max`); `thinking_token_budget`; `chat_template_kwargs` (`preserve_thinking`/`clear_thinking`/`enable_thinking`). Prism: `thinking.ts` resolves owned fields + `stripNeuralWattOwnedCompat`; `applyThinkingLevel(..., "reasoning_effort")`; Preserved Thinking replays `reasoning_content` | Official: `GET https://api.neuralwatt.com/v1/models` (auth optional for public models) | Featured `neuralWattModels` (official aliases incl. `gemma-4-31b`; no guessed pricing) + **`listNeuralWattModels()`** | **No Pi NeuralWatt provider** — official docs only | Optional API key on discovery; quota helper; redact secrets in error bodies | **fixed** 2026-07-18: catalog refresh + kwargs routing + owned-compat strip |
49
- | `@arnilo/prism-provider-ai-sdk` | 12 | Host model owns request caching; adapter maps usage only | host-owned / N/A | Host `LanguageModelV4` owns reasoning; Prism maps stream `reasoning` parts | **None by design** (host supplies model) | None | N/A (AI SDK official spec > Pi) | No package credentials; host model may hold secrets | **fixed** 2026-07-18: host-owned catalog/cache/reasoning validated; usage mapping + docs |
50
-
51
- ## Frozen official evidence sources (priority)
52
-
53
- | Provider | Frozen URLs (official) | Notes frozen 2026-07-17 |
54
- | --- | --- | --- |
55
- | OpenAI | [List models](https://developers.openai.com/api/reference/resources/models/methods/list); [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching); [Reasoning](https://developers.openai.com/api/docs/guides/reasoning); [Responses create](https://developers.openai.com/api/reference/resources/responses/methods/create/); [Models guide](https://developers.openai.com/api/docs/models) | `GET /models`. Caching: `prompt_cache_key`; pre-5.6 `prompt_cache_retention`; 5.6+ `prompt_cache_options` / breakpoints. Reasoning: `reasoning.effort`. |
56
- | Kimi / Moonshot | [List models](https://platform.kimi.ai/docs/api/list-models); [API overview](https://platform.kimi.ai/docs/api/overview); [Model parameter reference](https://platform.kimi.ai/docs/api/models-overview); [Thinking mode](https://platform.kimi.ai/docs/guide/use-kimi-k2-thinking-model); [Thinking effort](https://platform.kimi.ai/docs/guide/use-thinking-effort) | `GET /v1/models` on `api.moonshot.ai` / `.cn`. K2.x `thinking`; K3 `reasoning_effort: "max"`. Anthropic `/messages` compat remains under-documented — record empirical gaps in Task 7. |
57
- | Z.AI | [Deep thinking](https://docs.z.ai/guides/capabilities/thinking); [Thinking mode](https://docs.z.ai/guides/capabilities/thinking-mode); [Tool streaming](https://docs.z.ai/guides/capabilities/stream-tool); [Context caching](https://docs.z.ai/guides/capabilities/cache); [Chat completion](https://docs.z.ai/api-reference/llm/chat-completion); [Migrate to GLM-5.2](https://docs.z.ai/guides/overview/migrate-to-glm-new); [Overview](https://docs.z.ai/guides/overview/overview) | Code + `docs/providers/zai.md` match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking`. Historical mismatch (`thinkingFormat`) closed in Task 8. |
58
- | OpenRouter | [Get models](https://openrouter.ai/docs/api/api-reference/models/get-models); [Prompt caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching); [Reasoning tokens](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens) | `GET /api/v1/models`. Cache via `cache_control` + sticky routing. Reasoning via `reasoning` object. |
59
- | OpenCode Go | [Go](https://opencode.ai/docs/go/); [Providers](https://opencode.ai/docs/providers/) | Dual OpenAI + Anthropic compatible APIs. Official `GET /zen/go/v1/models`. Featured: Grok 4.5, GLM-5.2/5.1, Kimi K3/K2.7 Code/K2.6, MiMo, MiniMax, Qwen3.7/3.6, DeepSeek V4. Default base `https://opencode.ai/zen/go/v1`. |
60
- | NeuralWatt | [Models](https://portal.neuralwatt.com/docs/api/models); [Chat completions](https://portal.neuralwatt.com/docs/api/chat-completions); [API overview](https://portal.neuralwatt.com/docs/api/overview); [Quickstart](https://portal.neuralwatt.com/docs/quickstart) | `GET /v1/models` returns pricing/capabilities/limits metadata. Prefix caching automatic; `reasoning_effort` when capability flagged. |
61
- | AI SDK | [Custom provider / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); AI SDK usage types (`inputTokens` / cache details) | Adapter validates `specificationVersion`; maps cache read/write from usage. No Prism catalog. |
62
-
63
- ## Frozen Pi secondary references
64
-
65
- Repo: `https://github.com/badlogic/pi-mono` (`packages/ai`).
66
-
67
- | Area | Path / note |
68
- | --- | --- |
69
- | OpenAI Responses | `packages/ai/src/providers/openai-responses.ts` (also Codex / Azure variants) |
70
- | OpenAI Completions | `packages/ai/src/providers/openai-completions.ts` / `api/openai-completions.ts` |
71
- | Anthropic Messages | `packages/ai/src/providers/anthropic-messages.ts` / `api/anthropic-messages.ts` |
72
- | Generated catalogs | `packages/ai/src/providers/*.models.ts` via `scripts/generate-models.ts` — **not** copied as Prism’s sole strategy |
73
- | Provider registry docs | Pi mintlify “LLM Providers” / README provider list |
74
-
75
- Use Pi only to fill gaps or cross-check wire shapes after official docs.
76
-
77
- ## Shared discovery pattern (Task 3 decision — frozen)
78
-
79
- **Status (2026-07-17 Task 3):** pattern documented; no new core list-models primitive. Shared reuse is limited to existing transport/credential helpers.
80
-
81
- ### Inventory
82
-
83
- | Primitive | Location | Role |
84
- | --- | --- | --- |
85
- | `listNeuralWattModels({ apiKey?, fetch?, baseUrl?, signal?, headers? })` | `packages/provider-neuralwatt/src/models.ts` | **Template** — only shipped `list*Models` today |
86
- | `mapNeuralWattModel(entry)` | same | Provider-specific entry → `ModelConfig` |
87
- | `neuralWattModels` featured aliases | same | Offline bootstrap; no guessed pricing |
88
- | `readBoundedResponseText` | `@arnilo/prism/providers/transport` | Bounded error-body read + optional secret redaction |
89
- | `resolveCredentialValue` / `redactSecrets` | `@arnilo/prism` | Auth resolution + error redaction |
90
- | Package `models?: readonly ModelConfig[]` | Kimi/Z.AI/OpenRouter/OpenCode Go/NeuralWatt/**OpenAI** factories | Host override of registered catalog |
91
- | OpenAI factory `models?` / `listOpenAIModels` | **done** Task 6 | Fixed |
92
- | OpenRouter / OpenCode Go `list*Models` | OpenRouter **`listOpenRouterModels` done** Task 9; OpenCode Go **`listOpenCodeGoModels` done** Task 10; Z.AI **`listZaiModels` done** Task 8; Kimi **`listKimiModels` done** Task 7 | Per-package work |
93
- | AI SDK catalog | N/A | Host-owned `LanguageModelV4` — no discovery export |
94
-
95
- ### Decisions
96
-
97
- - Template options + return: NeuralWatt `listNeuralWattModels(...) → ModelConfig[]`.
98
- - **Never** call discovery from `create*ProviderPackage()` / extension setup.
99
- - Static catalogs = offline bootstrap / featured aliases only.
100
- - OpenRouter stays app-registration-first; **`listOpenRouterModels`** (done Task 9) feeds `models:` only.
101
- - AI SDK: no discovery export.
102
- - Prefer **package-local** helpers. Do **not** add a core model-discovery registry or OpenAI-compatible mega-mapper in Task 3. Extract a shared HTTP/list helper later only if ≥2 packages share identical parsing (unlikely: OpenRouter/NeuralWatt/OpenAI response shapes and cache/cost mapping diverge).
103
- - Discovery may populate `ModelConfig.cache` / `ModelConfig.cost` from live metadata when officially documented; static catalogs must not invent those fields.
104
- - Docs: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery), [Provider caching — Discovery and live cache/cost metadata](provider-caching.md#discovery-and-live-cache-cost-metadata), [Provider conformance — Model discovery checklist](provider-conformance.md#model-discovery-checklist).
105
-
106
- ## Shared thinking / per-turn override (Task 4 decision — frozen)
107
-
108
- | Layer | Contract |
109
- | --- | --- |
110
- | Model default | `ModelConfig.compat` (+ `capabilities.reasoning` where declared) |
111
- | Per-turn override | `ProviderRequestOptions.compat` via existing `mergeProviderRequestOptions` (request wins) |
112
- | Shared helpers | Core `applyThinkingLevel` / `thinkingCompatFor` / `thinkingFamilyForModel` → official compat fields; **not** a second options tree |
113
- | Families | `openai_reasoning` (`reasoning.effort`), `reasoning_effort`, `thinking_type` (`thinking.type`), `noop` (host-owned) — only shapes shared by ≥2 packages (or explicit no-op) |
114
- | Use-case wiring | LLM compaction + OM workers map `thinkingLevel` into `compat` via helpers (no longer inert `extra.thinkingLevel`) |
115
- | Docs | [Thinking and reasoning](thinking-and-reasoning.md) |
116
-
117
- **Decision (2026-07-17):** Implement thin core helpers; keep unique knobs (NeuralWatt budgets/kwargs, Kimi keep/all, Z.AI `tool_stream`) package-local. Core must not name forbidden provider literals (`openrouter`/`zai`/`kimi`/…). Hosts pick a family explicitly when inference is ambiguous. Per-provider first-class field hardening remains Tasks 6–12.
118
-
119
- ## Use-case model binding inventory (Task 5 — done 2026-07-17)
120
-
121
- | Site | Current binding | Session fallback today? | Thinking path today |
122
- | --- | --- | --- | --- |
123
- | `AgentSession.run` / `RunOptions.model` | Per-run override; writes `model_change` | Explicit override of session model | Host `providerOptions` / `applyThinkingLevel` |
124
- | Observational memory workers | `resolveUseCaseModel({ configured: workerModel, sessionModel })` | **Yes** — host passes `sessionModel`; `requireExplicitModel` restores skip | `thinkingLevel` → `compat` via `applyThinkingLevel` |
125
- | LLM compaction | `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` | **Yes** — `model` is the fallback slot | `thinkingLevel` → `compat` via `applyThinkingLevel` |
126
- | Supervisor children | Child `Agent` owns its own `model` | Independent | Child config |
127
- | Declarative agents | `resolveAgentDefinition` model string / `ModelConfig` | Definition-scoped | Definition / run options |
128
- | Evaluations | `runOptions` including model | Eval-owned | Via run options |
129
- | Workflows / RPC / CLI | Pass-through `runOptions.model` | Caller-owned | Via run options |
130
- | Structured output | Reuses session/run model | Yes | Same as run |
131
- | Memory / RAG embedders | Host `Embedder` — separate from chat LLM | N/A (not chat) | N/A |
132
-
133
- **Decision (2026-07-17):** Core exports `UseCaseModelBinding`, `resolveUseCaseModel`, `resolveUseCaseModelBinding`, `useCaseCredentialProviderId`. Docs: [use-case-model-selection.md](use-case-model-selection.md). OM behavior change: session fallback when `sessionModel` supplied and worker model omitted; `requireExplicitModel` preserves historical `missing_model` skip.
134
-
135
- Desired Plan 067 outcome: every use-case accepts `{ provider, model, thinking }` with **explicit session-model fallback** when no use-case default is set — **done** for OM + LLM compaction; other sites documented as already-separate.
136
-
137
- ## Known doc / code mismatches (frozen)
138
-
139
- | Item | Docs claim | Code / official | Owner task |
140
- | --- | --- | --- | --- |
141
- | Z.AI thinking | Was: `thinkingFormat: "zai"`, `developerRoleFallback` in `docs/providers/zai.md` | **fixed** Task 8 — docs+code match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking` | 8 (done) |
142
- | OpenAI reasoning | Was capabilities-only + opaque compat spread | Task 6 merges `model`/`options` `compat.reasoning` into body `reasoning`; Task 4 helper writes `compat.reasoning.effort` | 4 (done) / 6 (done) |
143
- | Kimi Moonshot | Was metadata-only (`provider: "moonshot"` without provider) | **fixed** Task 7: `createMoonshotProvider` registered when `includeMoonshotModels`; official Coding ids + `listKimiModels` | 7 (done) |
144
- | OpenCode Go catalog | Featured official Go open models + `listOpenCodeGoModels` | **done** Task 10 (removed stale `gpt-5.1-go` / `claude-sonnet-4.5-go`) | 10 (done) |
145
- | Z.AI catalog | Was: `glm-4.7` / `glm-4.5` only | **fixed** Task 8 — featured GLM-5.2…4.5 + `listZaiModels` | 8 (done) |
146
- | OM/LLM `thinkingLevel` + session fallback | Was inert `extra.thinkingLevel`; OM skipped without workerModel | Task 4 maps into `compat`; Task 5 session fallback + `requireExplicitModel` | 4 (done) / 5 (done) |
147
-
148
- ## Credential and redaction canaries (per package)
149
-
150
- | Package | Secrets in flight | Redaction canary expectation |
151
- | --- | --- | --- |
152
- | openai | API key; OAuth device/user/access/refresh | Absent from errors, events, requests, discovery failures |
153
- | kimi | API key | Absent from errors / cache keys |
154
- | zai | API key | Absent from errors / model metadata |
155
- | openrouter | API key; session/cache ids | Sanitize length; never treat cache key as secret storage |
156
- | opencode-go | API key | Absent from errors |
157
- | neuralwatt | Optional API key; quota responses | Discovery/quota error bodies redacted |
158
- | ai-sdk | Host-owned | Adapter must not echo host secrets in Prism errors |
159
-
160
- No secrets are committed in this matrix.
161
-
162
- ## Plan 067 task map (evidence owners)
163
-
164
- | Task | Owns |
165
- | --- | --- |
166
- | 0 | This page + evidence freeze — **done** |
167
- | 1 | R-001–R-004 re-verify — **fixed** 2026-07-17 |
168
- | 2 | R-005–R-006, R-008–R-012 re-verify — **fixed** 2026-07-17 |
169
- | 3 | Shared discovery pattern — **done** 2026-07-17 (package-local NeuralWatt template; no core list helper) |
170
- | 4 | Shared thinking / per-turn surface — **done** 2026-07-17 (core helpers + OM/LLM compat wiring; see thinking-and-reasoning.md) |
171
- | 5 | Use-case model selection + session fallback — **done** 2026-07-17 (`resolveUseCaseModel`, OM session fallback, use-case-model-selection.md) |
172
- | 6 | OpenAI validate/harden — **done** 2026-07-17 (Responses P0s, `listOpenAIModels`, `models?`, reasoning merge) |
173
- | 7 | Kimi validate/harden — **done** 2026-07-17 (`listKimiModels`, Moonshot provider, thinking, official Coding ids) |
174
- | 8 | Z.AI validate/harden — **done** 2026-07-17 (`listZaiModels`, GLM-5.x catalog, docs drift closed, clear_thinking/preserve) |
175
- | 9 | OpenRouter validate/harden — **done** 2026-07-17 (`listOpenRouterModels`, reasoning merge/preserve, automatic top-level `cache_control`) |
176
- | 10 | OpenCode Go validate/harden — **done** 2026-07-18 (`listOpenCodeGoModels`, official Go catalog, `zen/go/v1` base, thinking preserve) |
177
- | 11 | NeuralWatt validate/harden — **done** 2026-07-18 (featured catalog refresh, kwargs routing, owned-compat strip, applyThinkingLevel) |
178
- | 12 | AI SDK validate — **done** 2026-07-18 (host-owned catalog/cache/reasoning; usage mapping + docs) |
179
- | 13 | Cross-provider conformance + final verification — **done** 2026-07-18 (`sdk:ready`; 1,089 core tests + workspace suites + pack dry-runs) |
180
-
181
- Exact task titles live in `plans/067-provider-doc-validation-caching-discovery-and-review-hardening.md`.
182
-
183
- ## Verification for this page
184
-
185
- - Final verification completed 2026-07-18 without live provider HTTP.
186
- - Lists all seven first-party provider packages; every provider row is `fixed` or intentional by-design behavior.
187
- - Lists every 2026-07-14 P0–P2 id (R-001–R-006, R-008–R-012); all are `fixed`.
188
- - Distinguishes official-doc priority vs Pi secondary.
189
- - `phase12-boundaries.test.ts` verifies all six HTTP package discovery exports and setup zero-fetch; AI SDK no-catalog behavior has its own adapter contract test.
190
- - `provider_validation_final_contract_covers_all_adapters_and_binding_sites` verifies provider docs, cache kinds, thinking rows, use-case sites, matrix statuses, and index navigation.
191
- - `npm run sdk:ready` passes: typecheck/build, 1,089 core tests, all workspace suites, packaging/provenance guards, and publish-graph pack dry-runs.
192
- - Linked from `docs/index.md` under Release and install.
@@ -1,174 +0,0 @@
1
- # Review coverage — 2026-07-19 Phase 3
2
-
3
- Working evidence for Plan 070 Task 0. This page freezes Phase 3’s source revision, supported capability boundary, shared-primitive decision, finite-limit targets, and release evidence before public implementation begins.
4
-
5
- **Evidence frozen:** 2026-07-20. **Prism source:** `6048e82db212303f4f072ff70539830b779f35cf` (`Phase 0.0.7 completed`). **Default test rule:** all default tests use local fakes; a credential-gated live suite is separate.
6
-
7
- ## Status legend
8
-
9
- | Status | Meaning |
10
- | --- | --- |
11
- | `existing` | Current contract already covers this part. |
12
- | `extend` | Owning task extends an existing optional package/primitive. |
13
- | `new-package` | New optional package only; core stays unchanged unless a later two-consumer review proves a generic gap. |
14
- | `unsupported` | Deliberate 0.0.8 absence. Return a stable explicit error when a declared protocol operation is unavailable. |
15
-
16
- ## Frozen external references
17
-
18
- | Surface | Pinned reference | Compatibility decision |
19
- | --- | --- | --- |
20
- | OTel GenAI | [`semantic-conventions-genai@c26a2c21d1ee70d5231bd440c7b48d3c94ee506a`](https://github.com/open-telemetry/semantic-conventions-genai/tree/c26a2c21d1ee70d5231bd440c7b48d3c94ee506a/docs/gen-ai) | Development-status spans/metrics/events are adopted only where Prism has source data. Content attributes and evaluation explanations stay off by default. |
21
- | MCP specification | [2025-11-25](https://modelcontextprotocol.io/specification/2025-11-25) | JSON-RPC over stdio and Streamable HTTP only. Capability is declared only after SDK compatibility tests. |
22
- | MCP TypeScript SDK | [`@modelcontextprotocol/sdk@1.29.0`](https://github.com/modelcontextprotocol/typescript-sdk/tree/v1.29.0), lock integrity `sha512-zo37mZA9hJWpULgkRpowewez1y6ML5GsXJPY8FI0tBBCd77HEvza4jDqRKOXgHNn867PVGCyTdzqpz0izu5ZjQ==` | Current bridge is 1.29.0. Context7 verified client capability declaration, roots handlers, server capability inspection, and Streamable HTTP session construction. Task 4 pins the manifest to the tested SDK version; no untested upgrade. |
23
- | A2A | [A2A v1.0.0 `173695755607e884aa9acf8ce4feed90e32727a1`](https://github.com/a2aproject/A2A/tree/173695755607e884aa9acf8ce4feed90e32727a1) and [specification](https://a2a-protocol.org/v1.0.0/specification) | JSON-RPC/HTTPS only. Tasks, subscription, rich parts, and optional push hooks are added behind host lifecycle/auth adapters. |
24
- | Exa | [Search](https://exa.ai/docs/reference/search), [contents retrieval](https://exa.ai/docs/reference/contents-retrieval), retrieved 2026-07-20 | Direct API only, `POST /search` with requested contents. Docs have no immutable revision, so URL plus retrieval date and adapter request/response fixtures are release evidence. |
25
- | Brave | [Web Search `GET /res/v1/web/search`](https://api-dashboard.search.brave.com/api-reference/web/search/get), [versioning](https://api-dashboard.search.brave.com/documentation/guides/versioning), retrieved 2026-07-20 | Direct API only. Prism limits results to 20, matching documented maximum; token is late-bound `X-Subscription-Token`. |
26
- | Firecrawl | [v2 introduction](https://docs.firecrawl.dev/api-reference/v2-introduction), [search](https://docs.firecrawl.dev/api-reference/endpoint/search), [scrape](https://docs.firecrawl.dev/api-reference/endpoint/scrape), [extract](https://docs.firecrawl.dev/api-reference/endpoint/extract), retrieved 2026-07-20 | Direct v2 API only: `/v2/search`, `/v2/scrape`, `/v2/extract`. Markdown and returned schema data are untrusted. |
27
- | Release security | [Dependency review](https://docs.github.com/en/code-security/supply-chain-security/understanding-your-software-supply-chain/about-dependency-review), [SBOM](https://docs.github.com/en/code-security/supply-chain-security/understanding-your-software-supply-chain/about-software-bill-of-materials), [secret scanning](https://docs.github.com/en/code-security/secret-scanning/introduction/about-secret-scanning), [artifact attestations](https://docs.github.com/en/actions/security-for-github-actions/security-guides/using-artifact-attestations-to-establish-provenance-for-builds) | Prefer GitHub-native gates, immutable action revisions, minimal tokens, and protected-environment canaries. |
28
-
29
- ### External behavior relied on
30
-
31
- - GenAI inference and tool spans are `CLIENT` and `INTERNAL` respectively; `gen_ai.evaluation.result` is an event. Inputs, outputs, system instructions, and tool definitions are opt-in content fields, not default telemetry.
32
- - MCP 1.29.0 clients declare capabilities at construction; undeclared use is rejected by SDK. Roots use `roots/list` plus optional `notifications/roots/list_changed`; server capabilities are inspected after connect.
33
- - A2A v1.0 defines task get/list/cancel/subscribe, `text`/`raw`/`url`/`data` parts, ordered task updates, authenticated push configuration, and `A2A-Version` negotiation.
34
- - Brave documents a 400-character/50-word query maximum, `count <= 20`, and `offset <= 9`; Prism never expands those provider maxima.
35
- - Firecrawl documents `search.limit <= 100`, `query <= 500` characters and bearer authentication. Prism uses tighter portable defaults and validates returned Markdown/JSON itself.
36
-
37
- ## Capability traceability matrix
38
-
39
- | Roadmap criterion | Current surface | Minimum 0.0.8 gap | Owner | Required proof | Docs | Release gate |
40
- | --- | --- | --- | --- | --- | --- | --- |
41
- | Parent OTel agent/provider/tool hierarchy and context | Event-derived independent `prism.*` spans in `packages/observability-opentelemetry` | Active parent/context lifecycle; semantic map; trace reference | 1 | parent/error/detach/exporter-failure matrix | `observability.md`, `agent-events.md` | offline OTel conformance |
42
- | Metadata-safe GenAI/MCP telemetry and low-cardinality metrics | `AgentEvent`, redactor, safe counters | Convention mapping and label allow-list | 1 | secret/content/ID label-negative tests | `observability.md`, `host-security.md` | tarball/secret scan |
43
- | Final-result/trace grading, judges, pairwise, thresholds | `@arnilo/prism-evals` result scorers/datasets/experiments | Bounded owner-scoped trace resolver, optional judge, report/gate | 2 | fake judge, trace cap, deterministic threshold failure | `evaluations.md`, `runs-and-usage.md` | offline eval gate |
44
- | Batched ledger, snapshot cache, benchmarks | Serialized write-through `RunLedger`; repeated session rebuild | Optional flush/ack batch wrapper and leaf cache | 3 | order/crash/terminal flush/cache-read count benchmarks | `runs-and-usage.md`, `performance.md`, persistence pages | offline persistence + published bench |
45
- | MCP tools/resources/prompts/roots/sampling/elicitation/notifications | Tool bridge, tool-list notification, server tools, stateless Streamable HTTP | SDK-pinned capability facades/session handling | 4 | fake SDK/server capability, session, origin, overflow matrix | `mcp-tools.md`, `resource-loading.md` | offline MCP conformance |
46
- | Host-owned MCP OAuth/auth | `resolveAuthInfo`, authorizer, exact-origin pinned client transport | Exact request/session/ownership binding and host OAuth callback | 4 | wrong origin/session/owner/token-redaction cases | `mcp-tools.md`, `credential-storage.md`, `host-security.md` | secret scan + MCP conformance |
47
- | A2A durable task/cancel/reconnect/rich parts/push | Text-only immediate session execution, bounded SSE, card verification | Host task lifecycle adapter, part mapping, bounded replay/push | 5 | fake lifecycle/card/task/reconnect/push matrix | `a2a.md`, `agent-session-runtime.md`, `workflows.md` | offline A2A conformance |
48
- | Narrow web search/fetch/extract package | Tool dispatch, credentials, media SSRF bounds, bounded response reader | Optional direct-adapter package and normalized public result types | 6 | fake Brave/Exa/Firecrawl normalization/overflow/security matrix | `web-tools.md`, `tools.md` | pack/install + offline conformance |
49
- | Citation identity and schema-validated untrusted external results | `ToolResult`, `ContentBlock`, resource/media bounds | Package-local citation/document/extraction contracts | 6 | citation/schema/prompt-injection fixtures | `web-tools.md`, `host-security.md` | offline web conformance |
50
- | Official Exa/Firecrawl MCP prototype guidance only | Existing hardened MCP bridge | Documentation-only prototype boundary | 4, 6 | docs assertion: no generic remote passthrough | `mcp-tools.md`, `web-tools.md` | docs review |
51
- | SAST/dependency/secret/SBOM/license/attestation/updates/live canaries | `release.yml` readiness, pack, provenance publish | Native GitHub security workflows and protected live jobs | 7 | workflow policy/negative secret/license/SBOM fixtures | `release-and-install.md`, `host-security.md` | required PR/release jobs |
52
- | Reusable conformance and bounded performance evidence | Provider/session/tool/ledger helpers and 0.0.7 benchmark tables | Package-local protocol/web helpers plus Phase 3 benchmark script | 1–7 | each fake suite and dated benchmark output | `performance.md` | `sdk:ready` + documented live prerequisites |
53
-
54
- ## Primitive and caller inventory
55
-
56
- | Primitive | Existing contract and callers | Phase 3 disposition |
57
- | --- | --- | --- |
58
- | `AgentEvent` | `src/agents.ts` emits provider/tool/guardrail/limit events; subscribers, `RunLedger`, OTel adapter consume them | Reuse as telemetry source. Task 1 adds parent/context only; no second event bus. |
59
- | `RunLedger` / `RunLedgerRecord` | Four ordered append methods; `RuntimeAgentSession.ledgerChain` serializes appends; SQLite/PostgreSQL implement it | Task 3 may add one optional wrapper with flush/ack. Default remains write-through; no persistence schema. |
60
- | `ProductionPersistenceStore` | Cursor/ownership-scoped `queryRuns`, `queryEvents`, `queryToolCalls`, `queryUsage`; optional checkpoints/leases/feedback | Task 2 reads a bounded host-selected trace. Task 5 adapts host lifecycle/checkpoints. No generic task table. |
61
- | `RunLimits` / `RunLimitTracker` | Runtime/provider/tool limits charge before provider/tool work | Reuse for run work. Package-specific external operations get package-local finite limits until a second equal consumer exists. |
62
- | Guardrails/redaction/permission/trust | `runGuardrails`, `dispatchToolCall`, `redact*`, `assertPermission`, `assertTrusted` | All tool/protocol/web outputs route through existing dispatch/redaction. No protocol-specific security bypass. |
63
- | Durable agent/workflow state | `CheckpointStore`, `AgentRunLifecycle`, durable agent run state, workflow status/cancel | Task 5 projects host-selected lifecycle into A2A. No new worker/queue/runtime. |
64
- | `ResourceLoader` | `loadTextResource`, `loadJsonResource`, `loadBinaryResource`; permission/trust context | Task 4 maps resources only through a narrow adapter; Task 6 does not pretend remote web fetch is a generic resource loader. |
65
- | `CredentialResolver` | `resolveCredentialValue`, explicit/chained/env resolvers, OAuth store | Tasks 4–6 resolve at adapter request edge. Credentials never enter tool schemas/results/telemetry. |
66
- | Media URL/SSRF primitives | `assertSsrfAllowedUrl`, address-pinned `requestUrl`, media MIME/byte/time bounds | Reuse policy/validation semantics. Generalize pinned transport only if Task 6 proves identical requirements with media and MCP; otherwise package-local provider fetch. |
67
- | Provider transport | `readSseEvents`, `readBoundedResponseText`, bounded JSON argument parsing | Reuse bounded response/redaction semantics where shape fits. Do not force non-SSE vendor APIs through provider types. |
68
- | Tool dispatch | `ToolDefinition`, registry, schema validation, permissions, guardrails, ledger, limit tracker | Web tools and mapped MCP tools use this path; models never choose adapter/provider credentials. |
69
- | OTel adapter | Optional event projection with in-memory fixture and exporter-error isolation | Extend in place in Task 1; OTel API remains optional-package only. |
70
- | Evals package | Immutable dataset, function scorer, bounded experiment pool, package-local store | Extend in place in Task 2; judge stays host callback, not core/provider dependency. |
71
- | MCP package | Official SDK bridge, bounded tool discovery/result conversion, DNS-pinned Streamable HTTP, authorized server helper | Extend in place in Task 4; capability facades remain package-local. |
72
- | Supervisor A2A package | Card/JWS, bounded JSON-RPC/SSE, exact-origin client | Extend in place in Task 5; task/push adapters remain package-local. |
73
-
74
- ### Core primitive decision
75
-
76
- No new core primitive is authorized by Task 0. The only conditional candidates are:
77
-
78
- 1. **Trace-context carrier:** Task 1 may add one only if OTel, supervisor delegation, and another core consumer require the same dependency-free propagation shape.
79
- 2. **Flushable ledger capability:** Task 3 may add one only if one adapter wraps memory, SQLite, PostgreSQL, and host ledgers without changing `RunLedger` ordering.
80
- 3. **Pinned bounded HTTP request:** Task 6 may extract one only if media and web adapters use identical DNS-pinning, redirect, byte, abort, and error-redaction semantics.
81
-
82
- Otherwise code stays in its optional package. One-consumer interfaces, vendor types, protocol vocabularies, and task queues are rejected.
83
-
84
- ## Frozen capability boundary
85
-
86
- | Surface | Supported in 0.0.8 | Explicitly unsupported / deferred |
87
- | --- | --- | --- |
88
- | Telemetry | Optional OTel API adapter; agent/inference/tool/guardrail/delegation hierarchy; safe traces/metrics/evaluation events | Exporter registration, hosted backend, default content capture, per-delta spans, ID/content metric labels |
89
- | Evaluations | Function scorers, bounded trace target, host-supplied judge, pairwise comparison, datasets, reports, threshold assertion | Mandatory LLM/provider, evaluation service/database, automatic production grading |
90
- | MCP | Pinned-SDK tools/resources/prompts/roots/sampling/elicitation/notifications and Streamable HTTP only when SDK test proves each | Server discovery, arbitrary remote capability proxy, raw SDK callback/tool exposure, automatic OAuth/token forwarding |
91
- | A2A | v1.0 JSON-RPC/HTTPS card/task/status/list/cancel/subscribe, rich parts, bounded replay, optional host push hook | gRPC, REST/HTTP+JSON binding, endpoint discovery, JWK fetching, second durable engine |
92
- | Web tools | Host-selected Brave or Exa discovery; Firecrawl Markdown/schema extraction; direct native fetch adapters | Browser automation, vendor SDK dependency, model-selected provider, generic web/MCP passthrough |
93
- | Release | Native CI security gates and scheduled/manual protected canaries | Secrets in PR/default jobs, public-network default tests, hosted security service |
94
-
95
- ## Frozen limits and charging points
96
-
97
- Values in this table are target defaults/hard caps for later tasks, not active 0.0.7 APIs. Every count/byte/time/concurrency check happens before retaining data or starting the next request. Existing stricter package limits remain authoritative until changed with tests.
98
-
99
- | Surface | Default / hard cap | Charge before | Owner |
100
- | --- | --- | --- | --- |
101
- | OTel active spans | one agent span per run; provider/tool/guardrail/delegation only from existing bounded run work | `startSpan`; no span for deltas | 1 |
102
- | OTel content buffering | `0 / 0` by default | copying any prompt/content | 1 |
103
- | Eval trace pages | `20 / 100` per record kind | next persistence page | 2 |
104
- | Eval trace records | `1,000 / 5,000` events, tool calls, and usage records each | append snapshot item | 2 |
105
- | Eval judge request/response | `256 KiB / 1 MiB` each | serialization/body retention | 2 |
106
- | Eval judge attempts/time | `1 / 3`; `60 s / 30 min` | request attempt/timer | 2 |
107
- | Eval worker concurrency | `1 / 32` (existing) | start worker | 2 |
108
- | Ledger batch entries/bytes | `128 / 4,096`; `512 KiB / 16 MiB` | enqueue | 3 |
109
- | Ledger batch delay/in-flight flushes | `25 ms / 1 s`; `1 / 8` | timer/flush dispatch | 3 |
110
- | Snapshot cache | one current leaf snapshot per active session/run; no cross-session cache | store/reuse snapshot | 3 |
111
- | MCP existing tool bridge | existing `20/100` pages, `500/5,000` tools, `4 MiB/16 MiB` aggregate schemas, `10 MB/16 MiB` result | page/schema/result retention | 4 |
112
- | MCP resources/prompts | `20 / 100` pages; `500 / 5,000` items; `1 MiB / 8 MiB` item content | page/item conversion | 4 |
113
- | MCP roots/sampling/elicitation | `32 / 128` roots; `32 / 128` messages; `64 KiB / 1 MiB` arguments/schema | callback/request dispatch | 4 |
114
- | MCP HTTP sessions/replay | `128 / 1,024` active sessions; `1,024 / 10,000` replay events; `64 KiB / 1 MiB` event | session/replay allocation | 4 |
115
- | A2A existing request/response/stream | `64 KiB/1 MiB` request/response; `64 KiB/1 MiB` event; `10 MiB/64 MiB` stream; `10k/100k` events | body/frame/event retention | 5 |
116
- | A2A task pages/parts/artifacts | `100 / 1,000` tasks; `32 / 256` parts per message/artifact; `1 MiB / 8 MiB` raw/data part; `8 MiB / 64 MiB` aggregate artifacts | parse/decode/store/replay | 5 |
117
- | A2A push | `32 / 256` registrations; `3 / 10` attempts; `30 s / 5 min` delivery | persist/deliver/retry | 5 |
118
- | Web query/results/URLs | `4 KiB / 16 KiB`; `10 / 20` results; `5 / 20` URLs | request construction | 6 |
119
- | Web provider request/output | `256 KiB / 1 MiB` request; `2 MiB / 16 MiB` response; `1 MiB / 8 MiB` Markdown; `256 KiB / 1 MiB` extracted JSON | request/body/output retention | 6 |
120
- | Web schema/concurrency/time | `64 KiB / 256 KiB` schema; `4 / 16` active calls; `60 s / 30 min`; `2 / 4` retries | schema compile/call/retry | 6 |
121
- | CI/live canaries | default suite `0` remote calls; protected live job one bounded scenario/provider | credential resolution/network call | 7 |
122
-
123
- ### Network and credential rules
124
-
125
- 1. Provider API origins are exact allow-lists and provider API redirects are rejected. Native `fetch` uses `redirect: "error"` or equivalent pinned transport.
126
- 2. A fetched target URL is absolute HTTP(S), has no userinfo, and passes the host’s public/private policy before sending it to Firecrawl. Firecrawl-side redirects are provider behavior; Prism does not claim DNS pinning after handing a URL to Firecrawl. Hosts that need that guarantee use a controlled fetch adapter instead.
127
- 3. MCP client requests pin a validated DNS address and reject redirects today; Task 4 extends session/auth behavior without weakening that boundary.
128
- 4. A2A and MCP authorization run on every request and bind exact origin, session, and ownership. Missing/foreign resources/tasks resolve as authorized-not-found, not disclosure.
129
- 5. Credentials are resolved immediately before adapter I/O, redacted in errors before ledger/export, excluded from tool inputs/outputs/telemetry, and never forwarded automatically between MCP/A2A/provider/web surfaces.
130
- 6. Search snippets, Markdown, HTML, extracted JSON, A2A artifacts, MCP resources/prompts, and remote protocol errors are untrusted data. They cannot modify system instructions, tools, permissions, credential selection, or provider routing.
131
-
132
- ## Web normalization and external-operation ownership
133
-
134
- | Tool / adapter | Host-selected request | Normalized public output | Credential owner and redaction point |
135
- | --- | --- | --- | --- |
136
- | `web_search` / Exa | `POST https://api.exa.ai/search`; query plus explicitly requested bounded `contents` only | `title`, canonical `url`, bounded `snippet`/`highlights`, provider result ID, publication/retrieval time, provider/cost/rate metadata, citation identity | Adapter resolves `{ provider: "exa", name: "api_key" }` at request edge; redact key from headers, response/error, telemetry, and `ToolResult`. |
137
- | `web_search` / Brave | `GET https://api.search.brave.com/res/v1/web/search`; query/count/offset only | Same normalized search result; source fields missing from Brave stay absent, never guessed | Adapter resolves `{ provider: "brave", name: "subscription_token" }` at request edge; redact `X-Subscription-Token` and all error echoes. |
138
- | `web_fetch` / Firecrawl | `POST https://api.firecrawl.dev/v2/scrape`; one prevalidated public URL and Markdown format | `url`, canonical/source URL, bounded Markdown, selected metadata, retrieval time, provider/cost/rate metadata, `untrusted: true`, citation identity | Adapter resolves `{ provider: "firecrawl", name: "api_key" }` at request edge; redact bearer header/body/error echoes. |
139
- | `web_extract` / Firecrawl | `POST https://api.firecrawl.dev/v2/extract`; bounded URL list plus host-supplied JSON Schema | Same document attribution plus schema-validated bounded JSON value and `untrusted: true` | Same Firecrawl resolver/redaction rule; schema/value never influence tool permissions or system instructions. |
140
-
141
- `citationId` is deterministic: `web:<provider>:<providerResultId>` when provider returns a stable result ID, otherwise `web:<provider>:sha256(<canonicalUrl>)`. `canonicalUrl` removes fragment and normalizes only URL syntax; Prism does not follow target redirects to manufacture identity. Provider-specific fields remain under bounded metadata and never replace normalized fields.
142
-
143
- | External operation | Authorization/credential owner | Required boundary |
144
- | --- | --- | --- |
145
- | OTel export | Host exporter SDK; Prism adapter receives tracer/meter only | Exporter failure isolated; Prism receives no exporter credential. |
146
- | Model judge | Host-supplied judge callback | Judge receives bounded redacted evaluation target; no resolver/tools/workspace. |
147
- | MCP HTTP/OAuth | Host auth resolver; Task 4 session binding | Every request binds exact origin/session/ownership; no token forwarding to model/tool content. |
148
- | A2A invoke/card/push | Host authorizer/client auth/push delivery callback | Every task/push action is owner-scoped; card keys are explicitly pinned. |
149
- | Web adapters | Host `CredentialResolver` or explicit callback | Exact provider API origin, no provider API redirects, late credential resolution. |
150
- | Live canary | Protected CI environment only | Least-privilege key, redacted aggregate report, no PR/default-job secret. |
151
-
152
- ## Required test evidence by task
153
-
154
- | Task | Network-free evidence | Restricted live evidence |
155
- | --- | --- | --- |
156
- | 1 | Parent graph, context, cleanup, semantic attributes, no-content/no-ID-label, exporter isolation | Host OTel exporter smoke only if configured |
157
- | 2 | Fake trace reader/judge, redaction/cap/threshold/pairwise deterministic reports | Explicit host judge/model smoke |
158
- | 3 | Fake ledger order/flush/crash/cache invalidation; reproducible local benchmark | PostgreSQL benchmark when service exists |
159
- | 4 | Fake SDK/server capabilities, roots/resources/prompts/sampling/elicitation, sessions, auth/origin/SSRF/overflow | Configured MCP endpoint smoke |
160
- | 5 | Fake lifecycle/card/push, rich parts, cancel/reconnect/replay/owner isolation | Configured A2A endpoint smoke |
161
- | 6 | Fake Brave/Exa/Firecrawl normalization, schema, redirect/SSRF, credential/prompt-injection fixtures | Protected least-privilege Exa/Brave/Firecrawl smoke |
162
- | 7 | Workflow policy, SBOM/license/attestation, secret-negative and skipped-canary tests | Scheduled/manual protected credentials only |
163
-
164
- ## Release evidence checklist
165
-
166
- - `npm run sdk:ready`, Node 20/current import check, every workspace pack, packed offline consumer, `npm audit --audit-level=high`, tarball deny-list/secret check, and `git diff --check`.
167
- - Focused OTel/eval/ledger/MCP/A2A/web fake-server conformance suites pass without public network.
168
- - SAST, dependency review, secret scanning, SPDX SBOM/license policy, and artifact attestation jobs pass with pinned actions/minimal permissions.
169
- - Protected live canaries either pass with least-privilege credentials or are recorded as an explicit release-host prerequisite; skipped canaries never make the default suite appear live-tested.
170
- - Benchmarks publish machine/runtime, workload, p50/p95, throughput, memory, disk, cost metadata when supplied, backpressure observations, and no portability claim.
171
-
172
- ## Current exclusions
173
-
174
- No browser, coding, Office, SaaS connector, hosted observability, generic proxy, provider SDK, remote discovery, automatic token forwarding, background task worker, or new core persistence schema is authorized by this phase. Those belong to later roadmap phases or require a new evidence review.