@arnilo/prism 0.0.5 → 0.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/CHANGELOG.md +39 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +27 -16
  4. package/dist/agent-run-lifecycle.d.ts +28 -0
  5. package/dist/agent-run-lifecycle.js +33 -0
  6. package/dist/agent-run-state.d.ts +53 -0
  7. package/dist/agent-run-state.js +127 -0
  8. package/dist/agents.d.ts +3 -1
  9. package/dist/agents.js +337 -46
  10. package/dist/contracts.d.ts +205 -3
  11. package/dist/contracts.js +4 -0
  12. package/dist/guardrails.d.ts +25 -0
  13. package/dist/guardrails.js +133 -0
  14. package/dist/ids.d.ts +2 -0
  15. package/dist/ids.js +6 -0
  16. package/dist/index.d.ts +17 -3
  17. package/dist/index.js +10 -3
  18. package/dist/input.js +2 -0
  19. package/dist/resources.js +2 -1
  20. package/dist/run-limits.d.ts +34 -0
  21. package/dist/run-limits.js +163 -0
  22. package/dist/secure-agent.d.ts +3 -0
  23. package/dist/secure-agent.js +63 -0
  24. package/dist/session-stores.js +2 -3
  25. package/dist/testing/persistence-schema.d.ts +45 -7
  26. package/dist/testing/persistence-schema.js +138 -24
  27. package/dist/thinking.d.ts +42 -0
  28. package/dist/thinking.js +92 -0
  29. package/dist/tools.d.ts +10 -2
  30. package/dist/tools.js +56 -7
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +4 -2
  34. package/docs/agent-events.md +23 -16
  35. package/docs/agent-loops.md +19 -8
  36. package/docs/agent-session-runtime.md +33 -1
  37. package/docs/coding-agent-tools.md +33 -12
  38. package/docs/coding-security.md +2 -2
  39. package/docs/compaction-llm.md +17 -7
  40. package/docs/compaction-observational-memory.md +28 -4
  41. package/docs/credential-storage.md +58 -9
  42. package/docs/credentials-and-redaction.md +1 -1
  43. package/docs/database-persistence.md +8 -3
  44. package/docs/guardrails.md +75 -0
  45. package/docs/host-security.md +16 -8
  46. package/docs/index.md +26 -22
  47. package/docs/mcp-tools.md +32 -12
  48. package/docs/migration.md +164 -2
  49. package/docs/node-filesystem-config.md +1 -0
  50. package/docs/node-jsonl-session-store.md +5 -4
  51. package/docs/postgres-persistence.md +3 -3
  52. package/docs/provider-caching.md +16 -4
  53. package/docs/provider-conformance.md +39 -1
  54. package/docs/provider-packages.md +60 -3
  55. package/docs/providers/ai-sdk.md +36 -0
  56. package/docs/providers/kimi.md +124 -61
  57. package/docs/providers/neuralwatt.md +19 -13
  58. package/docs/providers/openai.md +56 -13
  59. package/docs/providers/opencode-go.md +118 -30
  60. package/docs/providers/openrouter.md +105 -35
  61. package/docs/providers/zai.md +94 -45
  62. package/docs/release-and-install.md +47 -49
  63. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  64. package/docs/runs-and-usage.md +30 -3
  65. package/docs/server.md +5 -2
  66. package/docs/sqlite-persistence.md +2 -2
  67. package/docs/structured-output.md +1 -1
  68. package/docs/thinking-and-reasoning.md +98 -0
  69. package/docs/tool-execution-primitives.md +3 -3
  70. package/docs/tools.md +21 -1
  71. package/docs/use-case-model-selection.md +109 -0
  72. package/docs/workflow-orchestration-primitives.md +1 -0
  73. package/docs/workflows.md +18 -10
  74. package/docs/working-and-semantic-memory.md +1 -0
  75. package/package.json +2 -2
@@ -0,0 +1,192 @@
1
+ # Review coverage — 2026-07-17 provider validation
2
+
3
+ Working evidence page for Plan 067. Freezes the 2026-07-14 P0–P2 re-verification map, first-party provider validation owners, official-doc priority URLs, Pi secondary references, cache/thinking/discovery surfaces, credential notes, and use-case model-binding inventory.
4
+
5
+ **Evidence frozen:** 2026-07-17 (offline inventory; no live provider calls).
6
+ **Priority rule:** official provider documentation wins; Pi (`badlogic/pi-mono`) is secondary when official docs are silent or ambiguous.
7
+
8
+ Related: [2026-07-14 coverage](review-coverage-2026-07-14.md) (0.0.4), [2026-07-15 coverage](review-coverage-2026-07-15.md) (0.0.5), [provider caching](provider-caching.md), [provider packages](provider-packages.md).
9
+
10
+ ## Status legend
11
+
12
+ | Status | Meaning |
13
+ | --- | --- |
14
+ | `verify` | Prior plan marked fixed; Plan 067 must re-verify with regression tests. |
15
+ | `gap` | Known missing or incorrect behavior/docs relative to official sources. |
16
+ | `fixed` | Closed in this plan (updated as tasks complete). |
17
+ | `by-design` | Intentional absence (document; do not “fix” into a catalog). |
18
+
19
+ ## P0–P2 finding → Plan 067 owner matrix
20
+
21
+ Source: `code-reviews/2026-07-14.md`. Prior Plans 053 / 054 / 058 marked these implemented for 0.0.4; this plan re-verifies rather than skipping.
22
+
23
+ | Review ID | Priority | Finding | 0.0.4 owner | Plan 067 task | Current status | Credential / redaction notes |
24
+ | --- | --- | --- | --- | --- | --- | --- |
25
+ | R-001 | P0 | Revision request duplicated + corrupted by redaction | 053-1 | 1 | fixed | Secret canaries must not appear in repair requests, events, or redacted graphs. Re-verified 2026-07-17: `pendingHistory` + active-path redaction; revision+redactor suite asserts one repair and no `[Circular]`. |
26
+ | R-002 | P1 | Multi-round tool transcript chronologically invalid | 053-2 | 1 | fixed | N/A. Re-verified 2026-07-17: two rounds × two calls keep `user → assistant → tool → tool → …` order in history and assembled request. |
27
+ | R-003 | P1 | Redactor leaks secrets in object/Map keys | 053-1 | 1 | fixed | Object/Map string keys redact; collisions use deterministic `__N` suffixes. |
28
+ | R-004 | P1 | Event-ledger writes lack backpressure | 053-3 | 1 | fixed | Ledger appends serialized (concurrency 1), order preserved, append failures reject run completion. |
29
+ | R-008 | P1 | Unbounded SSE / error bodies; multiline `data:` | 054-1/2 | 2 | fixed | Bounded readers; error text redacted. |
30
+ | R-009 | P1 | OpenAI device-code OAuth does not poll | 054-3 | 2 | fixed | Redact device/user/access/refresh codes from OAuth errors. |
31
+ | R-010 | P2 | Duplicated provider protocol utilities | 054-1/2 | 2 | fixed | Shared transport/primitives remain authoritative. |
32
+ | R-005 | P2 | JSONL append silent on corrupt lines | 053-4 | 2 | fixed | Dev-only store; fail closed on corrupt lines. |
33
+ | R-011 | P2 | Coding-agent image read unbounded / resize no-op | 055-5 | 2 | fixed | Enforce `maxImageBytes`; deprecate `autoResizeImages` honestly. |
34
+ | R-006 | P2 | Optional config ENOENT detected by message text | 053-5 | 2 | fixed | Typed `code === "ENOENT"`. |
35
+ | R-012 | P2 | Release tag/version mismatch; no provenance | 058-7/9 | 2 | fixed | Tag/version gate + `--provenance`; no secrets in artifacts. |
36
+
37
+ Bug-report fixes A–D remain covered by R-001 / R-003 / R-007 (malformed message shape was Plan 053; re-check under Task 1 if touched).
38
+
39
+ ## Provider package validation matrix
40
+
41
+ | Package | Plan 067 task | Official cache | Prism `cache.kind` | Thinking / reasoning (official → Prism) | Discovery endpoint | Static catalog (bootstrap only) | Pi secondary ref | Credential surface | Status |
42
+ | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
43
+ | `@arnilo/prism-provider-openai` | 6 | `prompt_cache_key`; older models `prompt_cache_retention` (`24h` / `in_memory`); GPT-5.6+ `prompt_cache_options` / breakpoints tracked (host may pass via compat/extra; discovery sets `longRetention: false`) | `openai_key` (+ `longRetention`, `maxKeyLength`) | Official: Responses top-level `reasoning.effort` (+ `summary`). Prism: merges `model.compat.reasoning` + `options.compat.reasoning` (request wins); `applyThinkingLevel(..., "openai_reasoning")` | Official: `GET /models` | Featured `openAIModels` / `openAICodexModels` + **`listOpenAIModels`**; factory accepts `models?` / `codexModels?` | `packages/ai/src/api/openai-responses.ts`, `openai-responses-shared.ts` | API key; Codex OAuth device-code (`oauth.ts`) — poll + redact codes | **fixed** 2026-07-17: Responses P0s + discovery + reasoning + models override |
44
+ | `@arnilo/prism-provider-kimi` | 7 | Anthropic-style `cache_control` on `/messages` when supported; OpenAI Moonshot route none | default implicit; opt-in `cache_control` | Official K2.x: `thinking.type` (`enabled`/`disabled`/`keep`); K2.7-code always enabled; K3: top-level `reasoning_effort`. Prism: `kimiThinking`/`kimiReasoningEffort`/`kimiPreserveThinking` (request wins); Coding Anthropic + Moonshot Chat Completions both callable | Official: `GET https://api.moonshot.ai/v1/models` (also `.cn`); Coding has no public list API | Featured Coding (`kimi-for-coding`, `kimi-for-coding-highspeed`, `k3`) + Moonshot (`kimi-k2.7-code`, `kimi-k3`) + **`listKimiModels`** | `kimi-coding*.ts`, `moonshotai*.ts`, `api/anthropic-messages.ts` | API key (`KIMI` / Moonshot, not interchangeable); redacted from errors | **fixed** 2026-07-17: discovery + Moonshot provider + thinking + official Coding ids |
45
+ | `@arnilo/prism-provider-zai` | 8 | Implicit (no client `cache_control` / `prompt_cache_key`); official `usage.prompt_tokens_details.cached_tokens` | `implicit` | Official: `thinking` (`{type, clear_thinking?}`), `reasoning_effort` (GLM-5.2+), `tool_stream` (GLM-4.6+). Prism: `zaiThinking`/`zaiReasoningEffort`/`zaiToolStream`/`zaiClearThinking`/`zaiPreserveThinking` (request wins); Preserved Thinking replays `reasoning_content`. Obsolete `thinkingFormat`/`developerRoleFallback` docs removed | No first-class docs.z.ai list page; OpenAI-compatible `GET {baseUrl}/models` best-effort + curated featured set from Chat Completions enum / overview | Featured `zaiModels` (`glm-5.2`…`glm-4.5`) + **`listZaiModels`**; default base `https://api.z.ai/api/paas/v4` | `zai.ts`, `zai.models.ts` (Pi secondary ids only) | API key; redacted from errors / discovery | **fixed** 2026-07-17: docs drift closed; catalog + discovery + clear_thinking/preserve |
46
+ | `@arnilo/prism-provider-openrouter` | 9 | Official prompt caching via `cache_control` + sticky `session_id` routing; top-level automatic when no breakpoints | `cache_control` (legacy `compat.openRouterCache`) | Official: `reasoning: { effort | max_tokens }`. Prism: `resolveOpenRouterReasoning` merge + `preserveThinking` replay as body `reasoning` | Official: `GET https://openrouter.ai/api/v1/models` | **App-controlled** `models:` + optional **`listOpenRouterModels`** (no bundled mega-catalog) | `openrouter.models.ts` (do not vendor), `api/openai-completions.ts` | API key; sanitize session/cache ids | **fixed** 2026-07-17: discovery + reasoning merge/preserve + automatic top-level cache_control |
47
+ | `@arnilo/prism-provider-opencode-go` | 10 | Anthropic route: selected `cache_control`; OpenAI route: none; `x-opencode-session` from cache/session key | route-specific (`cache_control` on Anthropic; `implicit` on OpenAI) | Dual-route: Anthropic thinking blocks + OpenAI `reasoning_content`; upstream `thinking`/`reasoning_effort`/`reasoning` passthrough (request wins); `preserveThinking` default for reasoning models | Official `GET https://opencode.ai/zen/go/v1/models` (sparse) | Featured official Go ids (Grok/GLM/Kimi/MiMo/MiniMax/Qwen/DeepSeek) + **`listOpenCodeGoModels`**; default base `https://opencode.ai/zen/go/v1` | `opencode-go.ts`, `opencode-go.models.ts` (Pi secondary ids/limits only) | API key; redacted from errors / discovery | **fixed** 2026-07-18: catalog + discovery + base URL + thinking preserve |
48
+ | `@arnilo/prism-provider-neuralwatt` | 11 | Implicit vLLM prefix caching; `prompt_tokens_details.cached_tokens` | `implicit` | Official: `reasoning_effort` when `capabilities.reasoning_effort` (GLM-5.2 default `max`); `thinking_token_budget`; `chat_template_kwargs` (`preserve_thinking`/`clear_thinking`/`enable_thinking`). Prism: `thinking.ts` resolves owned fields + `stripNeuralWattOwnedCompat`; `applyThinkingLevel(..., "reasoning_effort")`; Preserved Thinking replays `reasoning_content` | Official: `GET https://api.neuralwatt.com/v1/models` (auth optional for public models) | Featured `neuralWattModels` (official aliases incl. `gemma-4-31b`; no guessed pricing) + **`listNeuralWattModels()`** | **No Pi NeuralWatt provider** — official docs only | Optional API key on discovery; quota helper; redact secrets in error bodies | **fixed** 2026-07-18: catalog refresh + kwargs routing + owned-compat strip |
49
+ | `@arnilo/prism-provider-ai-sdk` | 12 | Host model owns request caching; adapter maps usage only | host-owned / N/A | Host `LanguageModelV4` owns reasoning; Prism maps stream `reasoning` parts | **None by design** (host supplies model) | None | N/A (AI SDK official spec > Pi) | No package credentials; host model may hold secrets | **fixed** 2026-07-18: host-owned catalog/cache/reasoning validated; usage mapping + docs |
50
+
51
+ ## Frozen official evidence sources (priority)
52
+
53
+ | Provider | Frozen URLs (official) | Notes frozen 2026-07-17 |
54
+ | --- | --- | --- |
55
+ | OpenAI | [List models](https://developers.openai.com/api/reference/resources/models/methods/list); [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching); [Reasoning](https://developers.openai.com/api/docs/guides/reasoning); [Responses create](https://developers.openai.com/api/reference/resources/responses/methods/create/); [Models guide](https://developers.openai.com/api/docs/models) | `GET /models`. Caching: `prompt_cache_key`; pre-5.6 `prompt_cache_retention`; 5.6+ `prompt_cache_options` / breakpoints. Reasoning: `reasoning.effort`. |
56
+ | Kimi / Moonshot | [List models](https://platform.kimi.ai/docs/api/list-models); [API overview](https://platform.kimi.ai/docs/api/overview); [Model parameter reference](https://platform.kimi.ai/docs/api/models-overview); [Thinking mode](https://platform.kimi.ai/docs/guide/use-kimi-k2-thinking-model); [Thinking effort](https://platform.kimi.ai/docs/guide/use-thinking-effort) | `GET /v1/models` on `api.moonshot.ai` / `.cn`. K2.x `thinking`; K3 `reasoning_effort: "max"`. Anthropic `/messages` compat remains under-documented — record empirical gaps in Task 7. |
57
+ | Z.AI | [Deep thinking](https://docs.z.ai/guides/capabilities/thinking); [Thinking mode](https://docs.z.ai/guides/capabilities/thinking-mode); [Tool streaming](https://docs.z.ai/guides/capabilities/stream-tool); [Context caching](https://docs.z.ai/guides/capabilities/cache); [Chat completion](https://docs.z.ai/api-reference/llm/chat-completion); [Migrate to GLM-5.2](https://docs.z.ai/guides/overview/migrate-to-glm-new); [Overview](https://docs.z.ai/guides/overview/overview) | Code + `docs/providers/zai.md` match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking`. Historical mismatch (`thinkingFormat`) closed in Task 8. |
58
+ | OpenRouter | [Get models](https://openrouter.ai/docs/api/api-reference/models/get-models); [Prompt caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching); [Reasoning tokens](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens) | `GET /api/v1/models`. Cache via `cache_control` + sticky routing. Reasoning via `reasoning` object. |
59
+ | OpenCode Go | [Go](https://opencode.ai/docs/go/); [Providers](https://opencode.ai/docs/providers/) | Dual OpenAI + Anthropic compatible APIs. Official `GET /zen/go/v1/models`. Featured: Grok 4.5, GLM-5.2/5.1, Kimi K3/K2.7 Code/K2.6, MiMo, MiniMax, Qwen3.7/3.6, DeepSeek V4. Default base `https://opencode.ai/zen/go/v1`. |
60
+ | NeuralWatt | [Models](https://portal.neuralwatt.com/docs/api/models); [Chat completions](https://portal.neuralwatt.com/docs/api/chat-completions); [API overview](https://portal.neuralwatt.com/docs/api/overview); [Quickstart](https://portal.neuralwatt.com/docs/quickstart) | `GET /v1/models` returns pricing/capabilities/limits metadata. Prefix caching automatic; `reasoning_effort` when capability flagged. |
61
+ | AI SDK | [Custom provider / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); AI SDK usage types (`inputTokens` / cache details) | Adapter validates `specificationVersion`; maps cache read/write from usage. No Prism catalog. |
62
+
63
+ ## Frozen Pi secondary references
64
+
65
+ Repo: `https://github.com/badlogic/pi-mono` (`packages/ai`).
66
+
67
+ | Area | Path / note |
68
+ | --- | --- |
69
+ | OpenAI Responses | `packages/ai/src/providers/openai-responses.ts` (also Codex / Azure variants) |
70
+ | OpenAI Completions | `packages/ai/src/providers/openai-completions.ts` / `api/openai-completions.ts` |
71
+ | Anthropic Messages | `packages/ai/src/providers/anthropic-messages.ts` / `api/anthropic-messages.ts` |
72
+ | Generated catalogs | `packages/ai/src/providers/*.models.ts` via `scripts/generate-models.ts` — **not** copied as Prism’s sole strategy |
73
+ | Provider registry docs | Pi mintlify “LLM Providers” / README provider list |
74
+
75
+ Use Pi only to fill gaps or cross-check wire shapes after official docs.
76
+
77
+ ## Shared discovery pattern (Task 3 decision — frozen)
78
+
79
+ **Status (2026-07-17 Task 3):** pattern documented; no new core list-models primitive. Shared reuse is limited to existing transport/credential helpers.
80
+
81
+ ### Inventory
82
+
83
+ | Primitive | Location | Role |
84
+ | --- | --- | --- |
85
+ | `listNeuralWattModels({ apiKey?, fetch?, baseUrl?, signal?, headers? })` | `packages/provider-neuralwatt/src/models.ts` | **Template** — only shipped `list*Models` today |
86
+ | `mapNeuralWattModel(entry)` | same | Provider-specific entry → `ModelConfig` |
87
+ | `neuralWattModels` featured aliases | same | Offline bootstrap; no guessed pricing |
88
+ | `readBoundedResponseText` | `@arnilo/prism/providers/transport` | Bounded error-body read + optional secret redaction |
89
+ | `resolveCredentialValue` / `redactSecrets` | `@arnilo/prism` | Auth resolution + error redaction |
90
+ | Package `models?: readonly ModelConfig[]` | Kimi/Z.AI/OpenRouter/OpenCode Go/NeuralWatt/**OpenAI** factories | Host override of registered catalog |
91
+ | OpenAI factory `models?` / `listOpenAIModels` | **done** Task 6 | Fixed |
92
+ | OpenRouter / OpenCode Go `list*Models` | OpenRouter **`listOpenRouterModels` done** Task 9; OpenCode Go **`listOpenCodeGoModels` done** Task 10; Z.AI **`listZaiModels` done** Task 8; Kimi **`listKimiModels` done** Task 7 | Per-package work |
93
+ | AI SDK catalog | N/A | Host-owned `LanguageModelV4` — no discovery export |
94
+
95
+ ### Decisions
96
+
97
+ - Template options + return: NeuralWatt `listNeuralWattModels(...) → ModelConfig[]`.
98
+ - **Never** call discovery from `create*ProviderPackage()` / extension setup.
99
+ - Static catalogs = offline bootstrap / featured aliases only.
100
+ - OpenRouter stays app-registration-first; **`listOpenRouterModels`** (done Task 9) feeds `models:` only.
101
+ - AI SDK: no discovery export.
102
+ - Prefer **package-local** helpers. Do **not** add a core model-discovery registry or OpenAI-compatible mega-mapper in Task 3. Extract a shared HTTP/list helper later only if ≥2 packages share identical parsing (unlikely: OpenRouter/NeuralWatt/OpenAI response shapes and cache/cost mapping diverge).
103
+ - Discovery may populate `ModelConfig.cache` / `ModelConfig.cost` from live metadata when officially documented; static catalogs must not invent those fields.
104
+ - Docs: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery), [Provider caching — Discovery and live cache/cost metadata](provider-caching.md#discovery-and-live-cache-cost-metadata), [Provider conformance — Model discovery checklist](provider-conformance.md#model-discovery-checklist).
105
+
106
+ ## Shared thinking / per-turn override (Task 4 decision — frozen)
107
+
108
+ | Layer | Contract |
109
+ | --- | --- |
110
+ | Model default | `ModelConfig.compat` (+ `capabilities.reasoning` where declared) |
111
+ | Per-turn override | `ProviderRequestOptions.compat` via existing `mergeProviderRequestOptions` (request wins) |
112
+ | Shared helpers | Core `applyThinkingLevel` / `thinkingCompatFor` / `thinkingFamilyForModel` → official compat fields; **not** a second options tree |
113
+ | Families | `openai_reasoning` (`reasoning.effort`), `reasoning_effort`, `thinking_type` (`thinking.type`), `noop` (host-owned) — only shapes shared by ≥2 packages (or explicit no-op) |
114
+ | Use-case wiring | LLM compaction + OM workers map `thinkingLevel` into `compat` via helpers (no longer inert `extra.thinkingLevel`) |
115
+ | Docs | [Thinking and reasoning](thinking-and-reasoning.md) |
116
+
117
+ **Decision (2026-07-17):** Implement thin core helpers; keep unique knobs (NeuralWatt budgets/kwargs, Kimi keep/all, Z.AI `tool_stream`) package-local. Core must not name forbidden provider literals (`openrouter`/`zai`/`kimi`/…). Hosts pick a family explicitly when inference is ambiguous. Per-provider first-class field hardening remains Tasks 6–12.
118
+
119
+ ## Use-case model binding inventory (Task 5 — done 2026-07-17)
120
+
121
+ | Site | Current binding | Session fallback today? | Thinking path today |
122
+ | --- | --- | --- | --- |
123
+ | `AgentSession.run` / `RunOptions.model` | Per-run override; writes `model_change` | Explicit override of session model | Host `providerOptions` / `applyThinkingLevel` |
124
+ | Observational memory workers | `resolveUseCaseModel({ configured: workerModel, sessionModel })` | **Yes** — host passes `sessionModel`; `requireExplicitModel` restores skip | `thinkingLevel` → `compat` via `applyThinkingLevel` |
125
+ | LLM compaction | `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` | **Yes** — `model` is the fallback slot | `thinkingLevel` → `compat` via `applyThinkingLevel` |
126
+ | Supervisor children | Child `Agent` owns its own `model` | Independent | Child config |
127
+ | Declarative agents | `resolveAgentDefinition` model string / `ModelConfig` | Definition-scoped | Definition / run options |
128
+ | Evaluations | `runOptions` including model | Eval-owned | Via run options |
129
+ | Workflows / RPC / CLI | Pass-through `runOptions.model` | Caller-owned | Via run options |
130
+ | Structured output | Reuses session/run model | Yes | Same as run |
131
+ | Memory / RAG embedders | Host `Embedder` — separate from chat LLM | N/A (not chat) | N/A |
132
+
133
+ **Decision (2026-07-17):** Core exports `UseCaseModelBinding`, `resolveUseCaseModel`, `resolveUseCaseModelBinding`, `useCaseCredentialProviderId`. Docs: [use-case-model-selection.md](use-case-model-selection.md). OM behavior change: session fallback when `sessionModel` supplied and worker model omitted; `requireExplicitModel` preserves historical `missing_model` skip.
134
+
135
+ Desired Plan 067 outcome: every use-case accepts `{ provider, model, thinking }` with **explicit session-model fallback** when no use-case default is set — **done** for OM + LLM compaction; other sites documented as already-separate.
136
+
137
+ ## Known doc / code mismatches (frozen)
138
+
139
+ | Item | Docs claim | Code / official | Owner task |
140
+ | --- | --- | --- | --- |
141
+ | Z.AI thinking | Was: `thinkingFormat: "zai"`, `developerRoleFallback` in `docs/providers/zai.md` | **fixed** Task 8 — docs+code match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking` | 8 (done) |
142
+ | OpenAI reasoning | Was capabilities-only + opaque compat spread | Task 6 merges `model`/`options` `compat.reasoning` into body `reasoning`; Task 4 helper writes `compat.reasoning.effort` | 4 (done) / 6 (done) |
143
+ | Kimi Moonshot | Was metadata-only (`provider: "moonshot"` without provider) | **fixed** Task 7: `createMoonshotProvider` registered when `includeMoonshotModels`; official Coding ids + `listKimiModels` | 7 (done) |
144
+ | OpenCode Go catalog | Featured official Go open models + `listOpenCodeGoModels` | **done** Task 10 (removed stale `gpt-5.1-go` / `claude-sonnet-4.5-go`) | 10 (done) |
145
+ | Z.AI catalog | Was: `glm-4.7` / `glm-4.5` only | **fixed** Task 8 — featured GLM-5.2…4.5 + `listZaiModels` | 8 (done) |
146
+ | OM/LLM `thinkingLevel` + session fallback | Was inert `extra.thinkingLevel`; OM skipped without workerModel | Task 4 maps into `compat`; Task 5 session fallback + `requireExplicitModel` | 4 (done) / 5 (done) |
147
+
148
+ ## Credential and redaction canaries (per package)
149
+
150
+ | Package | Secrets in flight | Redaction canary expectation |
151
+ | --- | --- | --- |
152
+ | openai | API key; OAuth device/user/access/refresh | Absent from errors, events, requests, discovery failures |
153
+ | kimi | API key | Absent from errors / cache keys |
154
+ | zai | API key | Absent from errors / model metadata |
155
+ | openrouter | API key; session/cache ids | Sanitize length; never treat cache key as secret storage |
156
+ | opencode-go | API key | Absent from errors |
157
+ | neuralwatt | Optional API key; quota responses | Discovery/quota error bodies redacted |
158
+ | ai-sdk | Host-owned | Adapter must not echo host secrets in Prism errors |
159
+
160
+ No secrets are committed in this matrix.
161
+
162
+ ## Plan 067 task map (evidence owners)
163
+
164
+ | Task | Owns |
165
+ | --- | --- |
166
+ | 0 | This page + evidence freeze — **done** |
167
+ | 1 | R-001–R-004 re-verify — **fixed** 2026-07-17 |
168
+ | 2 | R-005–R-006, R-008–R-012 re-verify — **fixed** 2026-07-17 |
169
+ | 3 | Shared discovery pattern — **done** 2026-07-17 (package-local NeuralWatt template; no core list helper) |
170
+ | 4 | Shared thinking / per-turn surface — **done** 2026-07-17 (core helpers + OM/LLM compat wiring; see thinking-and-reasoning.md) |
171
+ | 5 | Use-case model selection + session fallback — **done** 2026-07-17 (`resolveUseCaseModel`, OM session fallback, use-case-model-selection.md) |
172
+ | 6 | OpenAI validate/harden — **done** 2026-07-17 (Responses P0s, `listOpenAIModels`, `models?`, reasoning merge) |
173
+ | 7 | Kimi validate/harden — **done** 2026-07-17 (`listKimiModels`, Moonshot provider, thinking, official Coding ids) |
174
+ | 8 | Z.AI validate/harden — **done** 2026-07-17 (`listZaiModels`, GLM-5.x catalog, docs drift closed, clear_thinking/preserve) |
175
+ | 9 | OpenRouter validate/harden — **done** 2026-07-17 (`listOpenRouterModels`, reasoning merge/preserve, automatic top-level `cache_control`) |
176
+ | 10 | OpenCode Go validate/harden — **done** 2026-07-18 (`listOpenCodeGoModels`, official Go catalog, `zen/go/v1` base, thinking preserve) |
177
+ | 11 | NeuralWatt validate/harden — **done** 2026-07-18 (featured catalog refresh, kwargs routing, owned-compat strip, applyThinkingLevel) |
178
+ | 12 | AI SDK validate — **done** 2026-07-18 (host-owned catalog/cache/reasoning; usage mapping + docs) |
179
+ | 13 | Cross-provider conformance + final verification — **done** 2026-07-18 (`sdk:ready`; 1,089 core tests + workspace suites + pack dry-runs) |
180
+
181
+ Exact task titles live in `plans/067-provider-doc-validation-caching-discovery-and-review-hardening.md`.
182
+
183
+ ## Verification for this page
184
+
185
+ - Final verification completed 2026-07-18 without live provider HTTP.
186
+ - Lists all seven first-party provider packages; every provider row is `fixed` or intentional by-design behavior.
187
+ - Lists every 2026-07-14 P0–P2 id (R-001–R-006, R-008–R-012); all are `fixed`.
188
+ - Distinguishes official-doc priority vs Pi secondary.
189
+ - `phase12-boundaries.test.ts` verifies all six HTTP package discovery exports and setup zero-fetch; AI SDK no-catalog behavior has its own adapter contract test.
190
+ - `provider_validation_final_contract_covers_all_adapters_and_binding_sites` verifies provider docs, cache kinds, thinking rows, use-case sites, matrix statuses, and index navigation.
191
+ - `npm run sdk:ready` passes: typecheck/build, 1,089 core tests, all workspace suites, packaging/provenance guards, and publish-graph pack dry-runs.
192
+ - Linked from `docs/index.md` under Release and install.
@@ -35,13 +35,40 @@ Set the ledger and optional ownership scope/idempotency key on the agent or the
35
35
 
36
36
  | Method | Record | When called |
37
37
  | --- | --- | --- |
38
- | `appendRun` | `RunRecord` | After run starts (`running`) and again at finish (`succeeded`/`failed`/`aborted`). |
38
+ | `appendRun` | `RunRecord` | After run starts (`running`) and again at finish (`suspended`/`denied`/`succeeded`/`failed`/`aborted`). |
39
39
  | `appendEvent` | `AgentEventRecord` | After every emitted `AgentEvent`, after redaction. |
40
40
  | `appendToolCall` | `ToolCallRecord` | For each tool-call `started`, `progress`, `finished`, `error`, and `blocked` transition. |
41
41
  | `appendUsage` | `UsageRecord` | Once per terminal provider turn (`scope: "provider_turn"`) and once for the O(turns) aggregate (`scope: "run_total"`). |
42
42
 
43
43
  All methods may be sync or async (`void | Promise<void>`). The runtime awaits them at safe boundaries, so a slow adapter blocks the run.
44
44
 
45
+ ## Run limits
46
+
47
+ `RunLimits` bounds one `session.run()` across turns, provider attempts, tool rounds/calls, elapsed wall time, request/response bytes, token usage, and optional cost. Configure defaults on `AgentConfig.limits`; `RunOptions.limits` can only narrow an agent-configured value.
48
+
49
+ ```ts
50
+ await session.run("Summarize", {
51
+ limits: {
52
+ maxTurns: 4,
53
+ maxProviderAttempts: 6,
54
+ maxToolCalls: 8,
55
+ maxWallTimeMs: 30_000,
56
+ maxTotalTokens: 12_000,
57
+ maxCost: { amount: 0.25, currency: "USD" },
58
+ },
59
+ });
60
+ ```
61
+
62
+ Defaults/hard caps are respectively: turns 16/64, provider attempts 24/256, tool rounds 8/64, tool calls 32/256, wall time 120 seconds/30 minutes, request and response bytes 8/64 MiB, input tokens 40,000/1,000,000, output tokens 10,000/250,000, total tokens 50,000/1,000,000, and cost 10,000 currency units. Integer values must be positive safe integers. Cost needs a finite non-negative amount plus one currency; when cost is limited, absent, non-finite, or mixed-currency provider cost fails closed.
63
+
64
+ Prism charges turns before assembly, provider attempts and request bytes before generation, response bytes per provider event, tool rounds before a batch, tool calls before dispatch, and usage before another turn. A breach stops new work, aborts active work through the run signal, emits exactly one redacted `run_limit_exceeded` event/ledger row, and throws `AgentRunError` with `result.limit` (`limit`, `maximum`, `observed`, optional `currency`). Provider-reported token/cost totals arrive after generation, so that completed provider turn can be the unavoidable overshoot boundary.
65
+
66
+ `createRunLimitTracker()` and `resolveRunLimits()` are public for adapters that need the same validation and accounting semantics. Workflow agent nodes forward `RunWorkflowOptions.limits`; supervisor delegation narrows its step/tool/token/timeout budget into core limits; MCP tool calls use a per-call tracker.
67
+
68
+ ## Durable run state
69
+
70
+ `RunOptions.runState` writes a bounded, versioned checkpoint only at a safe interruption boundary. Its counters and absolute deadline resume with the run, while transcript history stays in `SessionStore` by session/leaf reference. `AgentRunResult.runState` exposes only redacted identity/status/version data; `interruption` excludes tool arguments. See [Agent/session runtime](agent-session-runtime.md#durable-interruption).
71
+
45
72
  ## Outputs / response / events
46
73
 
47
74
  The adapter receives these record shapes:
@@ -56,7 +83,7 @@ The adapter receives these record shapes:
56
83
  | `model` | Resolved model config for the run. |
57
84
  | `provider` | Resolved provider id for the run. |
58
85
  | `idempotencyKey` | Optional host key. |
59
- | `status` | `queued` \| `running` \| `succeeded` \| `failed` \| `aborted`. |
86
+ | `status` | `queued` \| `running` \| `suspended` \| `denied` \| `succeeded` \| `failed` \| `aborted`. |
60
87
  | `startedAt` / `finishedAt` | ISO timestamps. |
61
88
  | `abortReason` | Set when status is `aborted`. |
62
89
  | `error` | `ErrorInfo` when status is `failed`. |
@@ -241,7 +268,7 @@ console.log(cacheUsageReport(aggregate?.usage));
241
268
  - `AgentConfig.ownership` is the default ownership scope; `RunOptions.ownership` overrides it per run.
242
269
  - `AgentConfig.idempotencyKey` is the default idempotency key; `RunOptions.idempotencyKey` overrides it per run.
243
270
  - The runtime resolves `model` and `provider` from `AgentConfig`/`RunOptions`/`AgentDefinition` before writing the start `RunRecord`.
244
- - Adapters should treat appends as ordered within a `runId`: event and tool-call rows preserve emission order because the runtime drains pending appends before writing the final `RunRecord`.
271
+ - Adapters should treat appends as ordered within a `runId`: event and tool-call rows preserve emission order because the runtime serializes event ledger appends through one promise chain (concurrency 1), drains pending appends before writing the final `RunRecord`, and propagates append failures by rejecting run completion.
245
272
  - Billing queries must filter `scope = "provider_turn"`; presentation queries normally read the single `run_total`. `UsageQuery.scope`, `turn`, and `attempt` are explicit filters.
246
273
  - Adapters that need upsert semantics can use `RunRecord.id` (== `runId`) as the stable key.
247
274
  - Use `cacheUsageReport(record.usage, model)` for cache diagnostics from normalized usage. It works when a provider reports `cacheReadTokens` without `cacheWriteTokens`; missing write tokens are reported as `0`, and unavailable hit rate/savings stay `undefined`.
package/docs/server.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-server` exposes explicitly selected agents and workflows through one framework-free `(Request) => Promise<Response>` handler. It supports direct agent results, bounded agent/workflow SSE, durable workflow start/enqueue/status/cancel/resume/replay, ownership-scoped schedules, host authorization, ownership propagation, redaction, and resource ceilings.
5
+ `@arnilo/prism-server` exposes explicitly selected agents and workflows through one framework-free `(Request) => Promise<Response>` handler. It supports direct agent results, bounded agent/workflow SSE, opt-in durable agent status/resume, durable workflow start/enqueue/status/cancel/resume/replay, ownership-scoped schedules, host authorization, ownership propagation, redaction, and resource ceilings.
6
6
 
7
7
  No listener starts on import. Empty `agents`/`workflows` maps expose nothing. Authentication, authorization, route selection, durable stores, TLS, rate limiting, and framework/serverless adaptation remain host-owned.
8
8
 
@@ -17,6 +17,7 @@ Use `AgentSession` or workflow APIs directly for in-process applications. Do not
17
17
  ```ts
18
18
  const handler = createPrismHandler({
19
19
  agents?: Record<string, Agent | PrismAgentExposure>,
20
+ agentRuns?: Record<string, PrismAgentRunExposure>, // explicit durable status/resume only
20
21
  workflows?: Record<string, PrismWorkflowExposure>,
21
22
  schedules?: WorkflowSchedules | ((authorization, signal) => WorkflowSchedules),
22
23
  authorize: async ({ request, operation, capabilityId }) => false | {
@@ -38,6 +39,8 @@ At least one non-empty ownership field must come from `authorize()`. Request JSO
38
39
  | --- | --- | --- |
39
40
  | `POST /prism/agents/:id/runs` | `agent.run` | `{ "input": string | Message | Message[] }` |
40
41
  | `POST /prism/agents/:id/stream` | `agent.stream` | same; SSE response |
42
+ | `GET /prism/agents/:id/runs/:runId` | `agent.status` | none; redacted public state/version only |
43
+ | `POST /prism/agents/:id/runs/:runId/resume` | `agent.resume` | `{ "decision": "approve" | "deny", "expectedVersion": number }` |
41
44
  | `POST /prism/workflows/:id/runs` | `workflow.run` | `{ "input": unknown, "runId"?: string }` |
42
45
  | `POST /prism/workflows/:id/stream` | `workflow.stream` | same; SSE response |
43
46
  | `POST /prism/workflows/:id/enqueue` | `workflow.enqueue` | `{ "input": unknown, "runId"?: string }`; returns `202` queued handle |
@@ -125,7 +128,7 @@ Default/hard ceilings:
125
128
  - SSE uses bounded upstream subscriber queues. Consumer cancellation aborts owned work by default and releases concurrency; set `disconnectAborts: false` only when the host deliberately owns background completion.
126
129
  - Source inputs/resource URLs remain host responsibilities and use existing resource/media SSRF policies. Server package does not fetch URLs.
127
130
  - Schedule routes never accept ownership from JSON. Services carry mandatory ownership and explicit workflow/calculator registries; route authorization cannot broaden either. Replay applies workflow ownership/hash/approval checks.
128
- - No agent status/reconnect store is invented. Durable reconnect/status/resume is the workflow path; persistent agent run querying remains a host persistence API.
131
+ - Agent status/resume routes exist only for keys in `agentRuns`. Supply one core `createAgentRunLifecycle({ checkpoints, resolveAgent })` capability per selected agent; its resolver returns current `{ agent, definitionRevision }`. It reuses core checkpoint parsing/CAS/fingerprint checks, returns only public state/version, and needs a durable `SessionStore` as well as checkpoints for restart-safe resume. Empty/default configuration adds no agent lifecycle route, polling, or server cache.
129
132
 
130
133
  A2A routes are not added to `createPrismHandler()`. Install `@arnilo/prism-supervisor` and explicitly mount `createA2AHandler()` when protocol interoperability is required; this keeps cards and remote invoke absent from ordinary Prism servers.
131
134
 
@@ -56,7 +56,7 @@ import { createSqlitePersistence } from "@arnilo/prism-session-store-sqlite";
56
56
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
57
57
  | `close()` | Closes the underlying database when the adapter opened it. |
58
58
 
59
- Migrations run automatically on open and are idempotent across reopen.
59
+ Migrations run automatically on open and are idempotent across reopen. Under the SQLite migration transaction, startup checks ordered contract name/version/SHA-256 rows plus full schema-v3 PRAGMA/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
60
60
 
61
61
  ## Request/response example
62
62
 
@@ -109,7 +109,7 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
109
109
  - **No path interpolation.** The adapter opens exactly the caller-supplied `filename`; it does not expand environment variables or discover paths.
110
110
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
111
111
  - **WAL + busy timeout.** WAL is enabled by default; busy timeout defaults to 5 seconds. This meets the Plan 056 local workload target but SQLite still serializes writers — prefer PostgreSQL for high write concurrency.
112
- - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans.
112
+ - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans. Startup validation reads SQLite catalog/PRAGMA metadata only, never application rows.
113
113
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns on run and ownership tables participate in query filters; hosts must still scope writes correctly.
114
114
 
115
115
  ## Related APIs
@@ -234,7 +234,7 @@ Key cross-seam points:
234
234
  - The default parser treats assistant text as the value (`{ ok: true, value: text }`); supply a host parser whenever `T` is not `string`.
235
235
  - The default repairer builds a user message from `validation.errors[].message`; supply a host repairer for schema-specific guidance.
236
236
  - `maxRevisions` (default 3) bounds revision turns; budget exhaustion ends the loop and emits `artifact_failed` (it does not throw).
237
- - Tools are not dispatched in revision turns. Hosts needing tools in artifact turns use `singleShotLoop` or a custom loop.
237
+ - Tools are inert in artifact turns unless `loop.toolCalls: "bounded"` is explicit. Bounded mode uses run-global `maxToolRounds`, dispatches calls sequentially through normal runtime guards, skips parser/validator for tool-calling responses, and permits at most `1 + maxRevisions + maxToolRounds` provider turns. An extra tool response yields terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"` and executes nothing.
238
238
 
239
239
  ## Security and performance notes
240
240
 
@@ -0,0 +1,98 @@
1
+ # Thinking and reasoning
2
+
3
+ ## What it does
4
+
5
+ Prism keeps thinking/reasoning **provider-owned on the wire** while giving hosts one portable way to set effort per turn. Model defaults live on `ModelConfig.compat` (and `capabilities.reasoning` where declared). Per-turn overrides live on `ProviderRequestOptions.compat` and win through existing `mergeProviderRequestOptions`. Shared helpers map a portable `ThinkingLevel` into the official compat fields each family already reads — they do **not** invent a second options tree.
6
+
7
+ ## When to use it
8
+
9
+ - Session runs: pass `providerOptions.compat` (or `applyThinkingLevel`) on `RunOptions`.
10
+ - Use-case workers (LLM compaction, observational memory): pass `thinkingLevel`; packages map it into `compat` via the shared helpers.
11
+ - Provider authors: keep reading official fields from `options.compat` / `model.compat`; add package-local escape hatches only when the official API has unique knobs.
12
+
13
+ ## Contract
14
+
15
+ | Layer | Surface |
16
+ | --- | --- |
17
+ | Model default | `ModelConfig.compat` (+ `capabilities.reasoning` when the model can reason) |
18
+ | Per-turn override | `ProviderRequestOptions.compat` (request wins over model via merge) |
19
+ | Portable level | `ThinkingLevel`: `none` \| `minimal` \| `low` \| `medium` \| `high` \| `xhigh` \| `max` |
20
+ | Helpers | `thinkingCompatFor`, `applyThinkingLevel`, `thinkingFamilyForModel`, `isThinkingLevel`, `normalizeThinkingLevel`, `THINKING_LEVELS` |
21
+ | Not used | Inert `options.extra.thinkingLevel` — providers do not read `extra` for effort |
22
+
23
+ ```ts
24
+ import { applyThinkingLevel, thinkingCompatFor, thinkingFamilyForModel } from "@arnilo/prism";
25
+
26
+ // Per-turn override on a session run (OpenAI / OpenRouter family)
27
+ await session.run(input, {
28
+ providerOptions: applyThinkingLevel(undefined, "low", "openai_reasoning"),
29
+ });
30
+
31
+ // Equivalent explicit compat
32
+ await session.run(input, {
33
+ providerOptions: { compat: thinkingCompatFor("openai_reasoning", "low") },
34
+ // → { reasoning: { effort: "low" } }
35
+ });
36
+
37
+ // Use-case worker: family from model metadata (or pass an explicit family)
38
+ const family = thinkingFamilyForModel(model);
39
+ await runObserver({
40
+ ...,
41
+ providerOptions: applyThinkingLevel(base, "low", family === "noop" ? "reasoning_effort" : family),
42
+ });
43
+ ```
44
+
45
+ ## Compat families
46
+
47
+ Core maps only shapes shared by ≥2 packages (or an explicit no-op). Unique knobs stay package-local.
48
+
49
+ | Family | Compat patch | Used by (official fields) |
50
+ | --- | --- | --- |
51
+ | `openai_reasoning` | `{ reasoning: { effort } }` | OpenAI Responses `reasoning.effort`; OpenRouter `reasoning.effort` |
52
+ | `reasoning_effort` | `{ reasoning_effort }` | Z.AI `reasoning_effort`; NeuralWatt `reasoning_effort`; Kimi K3 `reasoning_effort` |
53
+ | `thinking_type` | `{ thinking: { type: "enabled" \| "disabled" } }` | Z.AI `thinking.type`; Kimi K2.x `thinking.type` (`none` → `disabled`) |
54
+ | `noop` | `{}` | AI SDK / host-owned adapters — effort is host-model settings |
55
+
56
+ `applyThinkingLevel` defaults `family` to `reasoning_effort` when omitted. For `openai_reasoning`, an existing `compat.reasoning.summary` (or other reasoning keys) is preserved when merging `effort`.
57
+
58
+ ### Recommended family by first-party package
59
+
60
+ | Package | Recommended family | Notes |
61
+ | --- | --- | --- |
62
+ | `@arnilo/prism-provider-openai` | `openai_reasoning` | First-class body `reasoning` from model + per-turn compat merge; `summary`/`mode`/`context` via compat |
63
+ | `@arnilo/prism-provider-openrouter` | `openai_reasoning` | First-class `resolveOpenRouterReasoning` merge; prefer `reasoning` object over legacy `reasoning_effort` shorthand; `preserveThinking` replays as body `reasoning` |
64
+ | `@arnilo/prism-provider-zai` | `reasoning_effort` (+ optional `thinking_type`) | Official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking`; Preserved Thinking via `reasoning_content` |
65
+ | `@arnilo/prism-provider-neuralwatt` | `reasoning_effort` | Budgets / `preserve_thinking` / `clear_thinking` / `chat_template_kwargs` stay package-local on `compat` |
66
+ | `@arnilo/prism-provider-kimi` | K3: `reasoning_effort`; K2.x: `thinking_type` | K2.7-code thinking is always on; do not send conflicting `thinking` + `reasoning_effort` |
67
+ | `@arnilo/prism-provider-opencode-go` | Anthropic route: thinking blocks (`thinking_type` family); OpenAI route: `reasoning_content` preserve + optional `thinking`/`reasoning_effort`/`reasoning` passthrough | Official dual endpoints; MiniMax/Qwen → Anthropic, others → OpenAI |
68
+ | `@arnilo/prism-provider-ai-sdk` | `noop` | Host `LanguageModelV4` owns reasoning settings |
69
+
70
+ `thinkingFamilyForModel` infers family from existing `compat` shape, then safe provider heuristics (`openai*` → `openai_reasoning`, `neuralwatt` → `reasoning_effort`), then `capabilities.reasoning` → `reasoning_effort`, else `noop`. Docs and packages may map other provider ids explicitly; core avoids provider-specific literals beyond those heuristics.
71
+
72
+ ## Merge order
73
+
74
+ 1. `ModelConfig.compat` / model defaults inside the provider
75
+ 2. `ProviderRequestOptions.compat` from agent / session policies
76
+ 3. Per-turn `RunOptions.providerOptions` or use-case `applyThinkingLevel` patch (wins)
77
+
78
+ Providers already prefer `request.options.compat.*` over `request.model.compat.*`.
79
+
80
+ ## Use-case workers
81
+
82
+ LLM compaction and observational memory accept `thinkingLevel?: string`. They call `applyThinkingLevel` into `compat` (not `extra.thinkingLevel`). When model inference returns `noop`, an explicit `thinkingLevel` still falls back to `reasoning_effort` so the host setting is never inert. Model selection for those workers (including session-model fallback) is documented in [Use-case model selection](use-case-model-selection.md).
83
+
84
+ ## Non-reasoning models
85
+
86
+ - Helper with `noop`: returns options unchanged — no invented body fields.
87
+ - Helper with a real family on a model that rejects the field: provider/API error — hosts should gate on `capabilities.reasoning` or package docs.
88
+ - `thinking_type` + `none` sets `{ type: "disabled" }`; other levels set `{ type: "enabled" }` without encoding effort (compose with `reasoning_effort` when the API supports both).
89
+
90
+ ## Related pages
91
+
92
+ - [Use-case model selection](use-case-model-selection.md) — session vs worker/summary model binding
93
+ - [Provider packages](provider-packages.md) — package boundaries and discovery
94
+ - [Provider caching](provider-caching.md) — cache retention can disable thinking on some providers (e.g. Z.AI when `cacheRetention: "none"`)
95
+ - [Provider request policies](provider-request-policies.md) — `mergeProviderRequestOptions`
96
+ - [Agent/session runtime](agent-session-runtime.md) — prior-reasoning preservation across turns
97
+ - Per-provider pages under [docs/providers](providers/)
98
+ - Evidence matrix: [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md)
@@ -123,8 +123,8 @@ Package performs **no** `PermissionPolicy`, `ToolValidator`, or trust checks of
123
123
  | --- | --- |
124
124
  | `ToolDefinition.parameters` | Stored and forwarded to providers; **not validated** by core |
125
125
  | `ToolValidator` | Host function hook; Phase 25 threads through agent runtime |
126
- | Standards-based schema validation | **Not shipped** — capability gap C-001 |
127
- | Schema compile cache | **None** — every dispatch would re-validate if host validator is naive |
126
+ | Standards-based schema validation | Optional `@arnilo/prism-tool-validator-json-schema`; host wires it through `ToolValidator` |
127
+ | Schema compile cache | Adapter-owned finite LRU; core never compiles schemas |
128
128
 
129
129
  ### MCP mapping (shipped — Task 3)
130
130
 
@@ -186,7 +186,7 @@ import { createJsonSchemaToolArgumentValidator } from "@arnilo/prism-tool-valida
186
186
  createAgent({ model, validator: createJsonSchemaToolArgumentValidator() });
187
187
  ```
188
188
 
189
- **Cache key:** stable `JSON.stringify(schema)` in adapter-owned `Map`. **Bounds:** configurable depth/properties/string/array limits before Ajv validation. **Security:** remote `$ref` rejected; prototype-pollution keys rejected in schemas and instances.
189
+ **Cache key:** stable `JSON.stringify(schema)` in an adapter-owned 256-entry LRU (hard cap 1,024); eviction removes the matching Ajv schema. **Bounds:** schemas default to 256 KiB, depth 64, 10,000 properties/keywords, and 128 refs (hard 1 MiB/128/100,000/1,024); instance depth/properties/string/array limits remain configurable. Every limit rejects non-finite, unsafe, zero/negative, and above-hard values. **Security:** only fragment-local `$ref` is accepted; prototype-pollution keys, cycles, and non-finite schema numbers reject before Ajv compilation.
190
190
 
191
191
  ### Task 2 — Parallel tool execution — **shipped**
192
192
 
package/docs/tools.md CHANGED
@@ -53,6 +53,7 @@ When multiple filters are provided, each non-empty allow list must include the t
53
53
  | `filter` | Optional exact allow/deny filter or ordered filters. |
54
54
  | `middleware` | Optional `MiddlewareRegistry`; `tool_call` runs before validation/execution and `tool_result` runs after execution. |
55
55
  | `validate` | Optional host validator returning `void`, a message string, or `ErrorInfo`. A non-`void` return blocks dispatch with reason `validation_failed` (redacted). Runs after the permission assertion and before `tool.execute()`. |
56
+ | `beforeExecute` | Optional adapter policy check immediately before the side effect; a rejection blocks dispatch with `execution_denied`. Workflows use it to retain `ExecutionPolicy` attribution after core guardrails. |
56
57
  | `emit` | Optional `AgentEvent` callback for lifecycle events. |
57
58
  | `secrets` | Known secret values to redact from thrown tool errors. |
58
59
  | `redactor` | Optional `SecretRedactor` used to redact tool-call ledger records. |
@@ -154,6 +155,10 @@ const agent = createAgent({ model, provider, tools: activeTools, permission, val
154
155
 
155
156
  Need different tools for one request? Build a short-lived agent/session with a narrower registry, or block extra calls with `PermissionPolicy` / `RunOptions.validate`. No extra per-run tool API exists yet; add one only when host apps need it.
156
157
 
158
+ ### Artifact-loop tools
159
+
160
+ `generate-validate-revise` treats provider tools as inert by default. Set `loop.toolCalls: "bounded"` and `RunOptions.maxToolRounds` only when an artifact needs a host-owned lookup before its next candidate. Each response with one-or-more calls consumes one shared round, dispatches calls sequentially through this exact `dispatchToolCall()` path, persists assistant-call then result transcript rows, and skips artifact parsing/validation for that response. A post-limit call executes nothing; the loop emits `artifact_failed` with `metadata.reason: "tool_round_limit"`. Tools do not consume `maxRevisions`, and tool schemas/context never grant authority.
161
+
157
162
  ### Runtime-supplied validators
158
163
 
159
164
  `AgentConfig.validator?` and `RunOptions.validate?` expose the same `ToolValidator` seam that `DispatchToolCallOptions.validate` already uses. The runtime threads `validate: RunOptions.validate ?? AgentConfig.validator` into every `dispatchToolCall` it issues during the tool loop, so an app can supply argument validation without taking ownership of dispatch itself. `RunOptions.validate` overrides `AgentConfig.validator` on a per-run basis (RunOptions wins). When neither is set, dispatch runs unmodified.
@@ -229,6 +234,21 @@ await session.run(input, {
229
234
  - Contribution registration and registry/filter calls do not perform provider calls, credential resolution, resource loading, network, filesystem discovery, or tool execution.
230
235
  - Dispatch performs explicit in-memory checks and executes only the selected host-active tool; it adds no retries, queues, timers, or new dependencies.
231
236
 
237
+ ## JSON Schema validator limits
238
+
239
+ Core stores `ToolDefinition.parameters` but does not compile schemas. Hosts that install `@arnilo/prism-tool-validator-json-schema` receive pre-Ajv schema limits: 256 KiB bytes, depth 64, 10,000 properties/keywords, 128 refs, and a 256-entry LRU compiled cache by default. All reject invalid values and have finite hard ceilings. Only fragment-local `$ref` values are accepted; non-local refs, cycles, forbidden keys, and non-finite schema numbers fail before tool execution.
240
+
241
+ ```ts
242
+ createJsonSchemaToolArgumentValidator({
243
+ maxSchemaBytes: 256 * 1024,
244
+ maxCompiledSchemas: 256,
245
+ });
246
+ ```
247
+
248
+ ## Guardrails
249
+
250
+ `DispatchToolCallOptions.guardrails` evaluates `tool_input` after `tool_call` middleware normalization and before lookup, permission, validation, execution policy, or side effect. `tool_output` evaluates raw completed results before redaction, event emission, ledger rows, and transcript append. A block returns a blocked result; tripwire fails the enclosing run. See [Guardrails](guardrails.md).
251
+
232
252
  ## Related APIs
233
253
 
234
254
  - [Agent/session runtime](agent-session-runtime.md): dispatches complete provider tool calls through the host-active tool harness and returns tool results on the next provider turn.
@@ -243,4 +263,4 @@ await session.run(input, {
243
263
  - [MCP client bridge](mcp-tools.md): optional `@arnilo/prism-mcp` remote tool mapping.
244
264
  - [Coding agent tools](coding-agent-tools.md): optional first-party `@arnilo/prism-coding-agent` `shell`/`read`/`write`/`edit` tools a host registers into this harness.
245
265
 
246
- `DispatchToolCallOptions.permission` can provide a `PermissionPolicy`; denial emits `tool_execution_blocked` before validation or `execute()`. Middleware cannot bypass this guard. `AgentConfig.validator`/`RunOptions.validate` run after this guard; their output is redacted through the active `SecretRedactor`. Prism does not sandbox tools. See [Security/auth/trust](settings-auth-trust-security.md).
266
+ `DispatchToolCallOptions.trust` and `.permission` run before validation or `execute()`; denial emits `tool_execution_blocked`. Middleware cannot bypass either guard. `AgentConfig.validator`/`RunOptions.validate` run after these guards; their output is redacted through the active `SecretRedactor`. `createSecureAgent()` requires all three seams plus non-empty schemas and durable pre-tool approval. Prism does not sandbox tools. See [Security/auth/trust](settings-auth-trust-security.md).