@arnilo/prism 0.0.5 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/CHANGELOG.md +28 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +26 -16
  4. package/dist/agents.js +2 -3
  5. package/dist/contracts.d.ts +2 -0
  6. package/dist/ids.d.ts +2 -0
  7. package/dist/ids.js +6 -0
  8. package/dist/index.d.ts +5 -1
  9. package/dist/index.js +3 -1
  10. package/dist/session-stores.js +2 -3
  11. package/dist/testing/persistence-schema.d.ts +45 -7
  12. package/dist/testing/persistence-schema.js +138 -24
  13. package/dist/thinking.d.ts +42 -0
  14. package/dist/thinking.js +92 -0
  15. package/dist/tools.js +2 -3
  16. package/dist/use-case-model.d.ts +63 -0
  17. package/dist/use-case-model.js +52 -0
  18. package/docs/a2a.md +4 -2
  19. package/docs/agent-events.md +10 -15
  20. package/docs/agent-loops.md +11 -8
  21. package/docs/coding-agent-tools.md +33 -12
  22. package/docs/coding-security.md +2 -2
  23. package/docs/compaction-llm.md +17 -7
  24. package/docs/compaction-observational-memory.md +28 -4
  25. package/docs/credential-storage.md +58 -9
  26. package/docs/credentials-and-redaction.md +1 -1
  27. package/docs/database-persistence.md +8 -3
  28. package/docs/host-security.md +10 -6
  29. package/docs/index.md +23 -20
  30. package/docs/mcp-tools.md +26 -10
  31. package/docs/migration.md +146 -2
  32. package/docs/node-filesystem-config.md +1 -0
  33. package/docs/node-jsonl-session-store.md +5 -4
  34. package/docs/postgres-persistence.md +3 -3
  35. package/docs/provider-caching.md +16 -4
  36. package/docs/provider-conformance.md +39 -1
  37. package/docs/provider-packages.md +60 -3
  38. package/docs/providers/ai-sdk.md +36 -0
  39. package/docs/providers/kimi.md +124 -61
  40. package/docs/providers/neuralwatt.md +19 -13
  41. package/docs/providers/openai.md +56 -13
  42. package/docs/providers/opencode-go.md +118 -30
  43. package/docs/providers/openrouter.md +105 -35
  44. package/docs/providers/zai.md +94 -45
  45. package/docs/release-and-install.md +47 -49
  46. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  47. package/docs/runs-and-usage.md +1 -1
  48. package/docs/sqlite-persistence.md +2 -2
  49. package/docs/structured-output.md +1 -1
  50. package/docs/thinking-and-reasoning.md +98 -0
  51. package/docs/tool-execution-primitives.md +3 -3
  52. package/docs/tools.md +15 -0
  53. package/docs/use-case-model-selection.md +109 -0
  54. package/docs/workflow-orchestration-primitives.md +1 -0
  55. package/docs/workflows.md +17 -10
  56. package/docs/working-and-semantic-memory.md +1 -0
  57. package/package.json +2 -2
@@ -0,0 +1,192 @@
1
+ # Review coverage — 2026-07-17 provider validation
2
+
3
+ Working evidence page for Plan 067. Freezes the 2026-07-14 P0–P2 re-verification map, first-party provider validation owners, official-doc priority URLs, Pi secondary references, cache/thinking/discovery surfaces, credential notes, and use-case model-binding inventory.
4
+
5
+ **Evidence frozen:** 2026-07-17 (offline inventory; no live provider calls).
6
+ **Priority rule:** official provider documentation wins; Pi (`badlogic/pi-mono`) is secondary when official docs are silent or ambiguous.
7
+
8
+ Related: [2026-07-14 coverage](review-coverage-2026-07-14.md) (0.0.4), [2026-07-15 coverage](review-coverage-2026-07-15.md) (0.0.5), [provider caching](provider-caching.md), [provider packages](provider-packages.md).
9
+
10
+ ## Status legend
11
+
12
+ | Status | Meaning |
13
+ | --- | --- |
14
+ | `verify` | Prior plan marked fixed; Plan 067 must re-verify with regression tests. |
15
+ | `gap` | Known missing or incorrect behavior/docs relative to official sources. |
16
+ | `fixed` | Closed in this plan (updated as tasks complete). |
17
+ | `by-design` | Intentional absence (document; do not “fix” into a catalog). |
18
+
19
+ ## P0–P2 finding → Plan 067 owner matrix
20
+
21
+ Source: `code-reviews/2026-07-14.md`. Prior Plans 053 / 054 / 058 marked these implemented for 0.0.4; this plan re-verifies rather than skipping.
22
+
23
+ | Review ID | Priority | Finding | 0.0.4 owner | Plan 067 task | Current status | Credential / redaction notes |
24
+ | --- | --- | --- | --- | --- | --- | --- |
25
+ | R-001 | P0 | Revision request duplicated + corrupted by redaction | 053-1 | 1 | fixed | Secret canaries must not appear in repair requests, events, or redacted graphs. Re-verified 2026-07-17: `pendingHistory` + active-path redaction; revision+redactor suite asserts one repair and no `[Circular]`. |
26
+ | R-002 | P1 | Multi-round tool transcript chronologically invalid | 053-2 | 1 | fixed | N/A. Re-verified 2026-07-17: two rounds × two calls keep `user → assistant → tool → tool → …` order in history and assembled request. |
27
+ | R-003 | P1 | Redactor leaks secrets in object/Map keys | 053-1 | 1 | fixed | Object/Map string keys redact; collisions use deterministic `__N` suffixes. |
28
+ | R-004 | P1 | Event-ledger writes lack backpressure | 053-3 | 1 | fixed | Ledger appends serialized (concurrency 1), order preserved, append failures reject run completion. |
29
+ | R-008 | P1 | Unbounded SSE / error bodies; multiline `data:` | 054-1/2 | 2 | fixed | Bounded readers; error text redacted. |
30
+ | R-009 | P1 | OpenAI device-code OAuth does not poll | 054-3 | 2 | fixed | Redact device/user/access/refresh codes from OAuth errors. |
31
+ | R-010 | P2 | Duplicated provider protocol utilities | 054-1/2 | 2 | fixed | Shared transport/primitives remain authoritative. |
32
+ | R-005 | P2 | JSONL append silent on corrupt lines | 053-4 | 2 | fixed | Dev-only store; fail closed on corrupt lines. |
33
+ | R-011 | P2 | Coding-agent image read unbounded / resize no-op | 055-5 | 2 | fixed | Enforce `maxImageBytes`; deprecate `autoResizeImages` honestly. |
34
+ | R-006 | P2 | Optional config ENOENT detected by message text | 053-5 | 2 | fixed | Typed `code === "ENOENT"`. |
35
+ | R-012 | P2 | Release tag/version mismatch; no provenance | 058-7/9 | 2 | fixed | Tag/version gate + `--provenance`; no secrets in artifacts. |
36
+
37
+ Bug-report fixes A–D remain covered by R-001 / R-003 / R-007 (malformed message shape was Plan 053; re-check under Task 1 if touched).
38
+
39
+ ## Provider package validation matrix
40
+
41
+ | Package | Plan 067 task | Official cache | Prism `cache.kind` | Thinking / reasoning (official → Prism) | Discovery endpoint | Static catalog (bootstrap only) | Pi secondary ref | Credential surface | Status |
42
+ | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
43
+ | `@arnilo/prism-provider-openai` | 6 | `prompt_cache_key`; older models `prompt_cache_retention` (`24h` / `in_memory`); GPT-5.6+ `prompt_cache_options` / breakpoints tracked (host may pass via compat/extra; discovery sets `longRetention: false`) | `openai_key` (+ `longRetention`, `maxKeyLength`) | Official: Responses top-level `reasoning.effort` (+ `summary`). Prism: merges `model.compat.reasoning` + `options.compat.reasoning` (request wins); `applyThinkingLevel(..., "openai_reasoning")` | Official: `GET /models` | Featured `openAIModels` / `openAICodexModels` + **`listOpenAIModels`**; factory accepts `models?` / `codexModels?` | `packages/ai/src/api/openai-responses.ts`, `openai-responses-shared.ts` | API key; Codex OAuth device-code (`oauth.ts`) — poll + redact codes | **fixed** 2026-07-17: Responses P0s + discovery + reasoning + models override |
44
+ | `@arnilo/prism-provider-kimi` | 7 | Anthropic-style `cache_control` on `/messages` when supported; OpenAI Moonshot route none | default implicit; opt-in `cache_control` | Official K2.x: `thinking.type` (`enabled`/`disabled`/`keep`); K2.7-code always enabled; K3: top-level `reasoning_effort`. Prism: `kimiThinking`/`kimiReasoningEffort`/`kimiPreserveThinking` (request wins); Coding Anthropic + Moonshot Chat Completions both callable | Official: `GET https://api.moonshot.ai/v1/models` (also `.cn`); Coding has no public list API | Featured Coding (`kimi-for-coding`, `kimi-for-coding-highspeed`, `k3`) + Moonshot (`kimi-k2.7-code`, `kimi-k3`) + **`listKimiModels`** | `kimi-coding*.ts`, `moonshotai*.ts`, `api/anthropic-messages.ts` | API key (`KIMI` / Moonshot, not interchangeable); redacted from errors | **fixed** 2026-07-17: discovery + Moonshot provider + thinking + official Coding ids |
45
+ | `@arnilo/prism-provider-zai` | 8 | Implicit (no client `cache_control` / `prompt_cache_key`); official `usage.prompt_tokens_details.cached_tokens` | `implicit` | Official: `thinking` (`{type, clear_thinking?}`), `reasoning_effort` (GLM-5.2+), `tool_stream` (GLM-4.6+). Prism: `zaiThinking`/`zaiReasoningEffort`/`zaiToolStream`/`zaiClearThinking`/`zaiPreserveThinking` (request wins); Preserved Thinking replays `reasoning_content`. Obsolete `thinkingFormat`/`developerRoleFallback` docs removed | No first-class docs.z.ai list page; OpenAI-compatible `GET {baseUrl}/models` best-effort + curated featured set from Chat Completions enum / overview | Featured `zaiModels` (`glm-5.2`…`glm-4.5`) + **`listZaiModels`**; default base `https://api.z.ai/api/paas/v4` | `zai.ts`, `zai.models.ts` (Pi secondary ids only) | API key; redacted from errors / discovery | **fixed** 2026-07-17: docs drift closed; catalog + discovery + clear_thinking/preserve |
46
+ | `@arnilo/prism-provider-openrouter` | 9 | Official prompt caching via `cache_control` + sticky `session_id` routing; top-level automatic when no breakpoints | `cache_control` (legacy `compat.openRouterCache`) | Official: `reasoning: { effort | max_tokens }`. Prism: `resolveOpenRouterReasoning` merge + `preserveThinking` replay as body `reasoning` | Official: `GET https://openrouter.ai/api/v1/models` | **App-controlled** `models:` + optional **`listOpenRouterModels`** (no bundled mega-catalog) | `openrouter.models.ts` (do not vendor), `api/openai-completions.ts` | API key; sanitize session/cache ids | **fixed** 2026-07-17: discovery + reasoning merge/preserve + automatic top-level cache_control |
47
+ | `@arnilo/prism-provider-opencode-go` | 10 | Anthropic route: selected `cache_control`; OpenAI route: none; `x-opencode-session` from cache/session key | route-specific (`cache_control` on Anthropic; `implicit` on OpenAI) | Dual-route: Anthropic thinking blocks + OpenAI `reasoning_content`; upstream `thinking`/`reasoning_effort`/`reasoning` passthrough (request wins); `preserveThinking` default for reasoning models | Official `GET https://opencode.ai/zen/go/v1/models` (sparse) | Featured official Go ids (Grok/GLM/Kimi/MiMo/MiniMax/Qwen/DeepSeek) + **`listOpenCodeGoModels`**; default base `https://opencode.ai/zen/go/v1` | `opencode-go.ts`, `opencode-go.models.ts` (Pi secondary ids/limits only) | API key; redacted from errors / discovery | **fixed** 2026-07-18: catalog + discovery + base URL + thinking preserve |
48
+ | `@arnilo/prism-provider-neuralwatt` | 11 | Implicit vLLM prefix caching; `prompt_tokens_details.cached_tokens` | `implicit` | Official: `reasoning_effort` when `capabilities.reasoning_effort` (GLM-5.2 default `max`); `thinking_token_budget`; `chat_template_kwargs` (`preserve_thinking`/`clear_thinking`/`enable_thinking`). Prism: `thinking.ts` resolves owned fields + `stripNeuralWattOwnedCompat`; `applyThinkingLevel(..., "reasoning_effort")`; Preserved Thinking replays `reasoning_content` | Official: `GET https://api.neuralwatt.com/v1/models` (auth optional for public models) | Featured `neuralWattModels` (official aliases incl. `gemma-4-31b`; no guessed pricing) + **`listNeuralWattModels()`** | **No Pi NeuralWatt provider** — official docs only | Optional API key on discovery; quota helper; redact secrets in error bodies | **fixed** 2026-07-18: catalog refresh + kwargs routing + owned-compat strip |
49
+ | `@arnilo/prism-provider-ai-sdk` | 12 | Host model owns request caching; adapter maps usage only | host-owned / N/A | Host `LanguageModelV4` owns reasoning; Prism maps stream `reasoning` parts | **None by design** (host supplies model) | None | N/A (AI SDK official spec > Pi) | No package credentials; host model may hold secrets | **fixed** 2026-07-18: host-owned catalog/cache/reasoning validated; usage mapping + docs |
50
+
51
+ ## Frozen official evidence sources (priority)
52
+
53
+ | Provider | Frozen URLs (official) | Notes frozen 2026-07-17 |
54
+ | --- | --- | --- |
55
+ | OpenAI | [List models](https://developers.openai.com/api/reference/resources/models/methods/list); [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching); [Reasoning](https://developers.openai.com/api/docs/guides/reasoning); [Responses create](https://developers.openai.com/api/reference/resources/responses/methods/create/); [Models guide](https://developers.openai.com/api/docs/models) | `GET /models`. Caching: `prompt_cache_key`; pre-5.6 `prompt_cache_retention`; 5.6+ `prompt_cache_options` / breakpoints. Reasoning: `reasoning.effort`. |
56
+ | Kimi / Moonshot | [List models](https://platform.kimi.ai/docs/api/list-models); [API overview](https://platform.kimi.ai/docs/api/overview); [Model parameter reference](https://platform.kimi.ai/docs/api/models-overview); [Thinking mode](https://platform.kimi.ai/docs/guide/use-kimi-k2-thinking-model); [Thinking effort](https://platform.kimi.ai/docs/guide/use-thinking-effort) | `GET /v1/models` on `api.moonshot.ai` / `.cn`. K2.x `thinking`; K3 `reasoning_effort: "max"`. Anthropic `/messages` compat remains under-documented — record empirical gaps in Task 7. |
57
+ | Z.AI | [Deep thinking](https://docs.z.ai/guides/capabilities/thinking); [Thinking mode](https://docs.z.ai/guides/capabilities/thinking-mode); [Tool streaming](https://docs.z.ai/guides/capabilities/stream-tool); [Context caching](https://docs.z.ai/guides/capabilities/cache); [Chat completion](https://docs.z.ai/api-reference/llm/chat-completion); [Migrate to GLM-5.2](https://docs.z.ai/guides/overview/migrate-to-glm-new); [Overview](https://docs.z.ai/guides/overview/overview) | Code + `docs/providers/zai.md` match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking`. Historical mismatch (`thinkingFormat`) closed in Task 8. |
58
+ | OpenRouter | [Get models](https://openrouter.ai/docs/api/api-reference/models/get-models); [Prompt caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching); [Reasoning tokens](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens) | `GET /api/v1/models`. Cache via `cache_control` + sticky routing. Reasoning via `reasoning` object. |
59
+ | OpenCode Go | [Go](https://opencode.ai/docs/go/); [Providers](https://opencode.ai/docs/providers/) | Dual OpenAI + Anthropic compatible APIs. Official `GET /zen/go/v1/models`. Featured: Grok 4.5, GLM-5.2/5.1, Kimi K3/K2.7 Code/K2.6, MiMo, MiniMax, Qwen3.7/3.6, DeepSeek V4. Default base `https://opencode.ai/zen/go/v1`. |
60
+ | NeuralWatt | [Models](https://portal.neuralwatt.com/docs/api/models); [Chat completions](https://portal.neuralwatt.com/docs/api/chat-completions); [API overview](https://portal.neuralwatt.com/docs/api/overview); [Quickstart](https://portal.neuralwatt.com/docs/quickstart) | `GET /v1/models` returns pricing/capabilities/limits metadata. Prefix caching automatic; `reasoning_effort` when capability flagged. |
61
+ | AI SDK | [Custom provider / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); AI SDK usage types (`inputTokens` / cache details) | Adapter validates `specificationVersion`; maps cache read/write from usage. No Prism catalog. |
62
+
63
+ ## Frozen Pi secondary references
64
+
65
+ Repo: `https://github.com/badlogic/pi-mono` (`packages/ai`).
66
+
67
+ | Area | Path / note |
68
+ | --- | --- |
69
+ | OpenAI Responses | `packages/ai/src/providers/openai-responses.ts` (also Codex / Azure variants) |
70
+ | OpenAI Completions | `packages/ai/src/providers/openai-completions.ts` / `api/openai-completions.ts` |
71
+ | Anthropic Messages | `packages/ai/src/providers/anthropic-messages.ts` / `api/anthropic-messages.ts` |
72
+ | Generated catalogs | `packages/ai/src/providers/*.models.ts` via `scripts/generate-models.ts` — **not** copied as Prism’s sole strategy |
73
+ | Provider registry docs | Pi mintlify “LLM Providers” / README provider list |
74
+
75
+ Use Pi only to fill gaps or cross-check wire shapes after official docs.
76
+
77
+ ## Shared discovery pattern (Task 3 decision — frozen)
78
+
79
+ **Status (2026-07-17 Task 3):** pattern documented; no new core list-models primitive. Shared reuse is limited to existing transport/credential helpers.
80
+
81
+ ### Inventory
82
+
83
+ | Primitive | Location | Role |
84
+ | --- | --- | --- |
85
+ | `listNeuralWattModels({ apiKey?, fetch?, baseUrl?, signal?, headers? })` | `packages/provider-neuralwatt/src/models.ts` | **Template** — only shipped `list*Models` today |
86
+ | `mapNeuralWattModel(entry)` | same | Provider-specific entry → `ModelConfig` |
87
+ | `neuralWattModels` featured aliases | same | Offline bootstrap; no guessed pricing |
88
+ | `readBoundedResponseText` | `@arnilo/prism/providers/transport` | Bounded error-body read + optional secret redaction |
89
+ | `resolveCredentialValue` / `redactSecrets` | `@arnilo/prism` | Auth resolution + error redaction |
90
+ | Package `models?: readonly ModelConfig[]` | Kimi/Z.AI/OpenRouter/OpenCode Go/NeuralWatt/**OpenAI** factories | Host override of registered catalog |
91
+ | OpenAI factory `models?` / `listOpenAIModels` | **done** Task 6 | Fixed |
92
+ | OpenRouter / OpenCode Go `list*Models` | OpenRouter **`listOpenRouterModels` done** Task 9; OpenCode Go **`listOpenCodeGoModels` done** Task 10; Z.AI **`listZaiModels` done** Task 8; Kimi **`listKimiModels` done** Task 7 | Per-package work |
93
+ | AI SDK catalog | N/A | Host-owned `LanguageModelV4` — no discovery export |
94
+
95
+ ### Decisions
96
+
97
+ - Template options + return: NeuralWatt `listNeuralWattModels(...) → ModelConfig[]`.
98
+ - **Never** call discovery from `create*ProviderPackage()` / extension setup.
99
+ - Static catalogs = offline bootstrap / featured aliases only.
100
+ - OpenRouter stays app-registration-first; **`listOpenRouterModels`** (done Task 9) feeds `models:` only.
101
+ - AI SDK: no discovery export.
102
+ - Prefer **package-local** helpers. Do **not** add a core model-discovery registry or OpenAI-compatible mega-mapper in Task 3. Extract a shared HTTP/list helper later only if ≥2 packages share identical parsing (unlikely: OpenRouter/NeuralWatt/OpenAI response shapes and cache/cost mapping diverge).
103
+ - Discovery may populate `ModelConfig.cache` / `ModelConfig.cost` from live metadata when officially documented; static catalogs must not invent those fields.
104
+ - Docs: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery), [Provider caching — Discovery and live cache/cost metadata](provider-caching.md#discovery-and-live-cache-cost-metadata), [Provider conformance — Model discovery checklist](provider-conformance.md#model-discovery-checklist).
105
+
106
+ ## Shared thinking / per-turn override (Task 4 decision — frozen)
107
+
108
+ | Layer | Contract |
109
+ | --- | --- |
110
+ | Model default | `ModelConfig.compat` (+ `capabilities.reasoning` where declared) |
111
+ | Per-turn override | `ProviderRequestOptions.compat` via existing `mergeProviderRequestOptions` (request wins) |
112
+ | Shared helpers | Core `applyThinkingLevel` / `thinkingCompatFor` / `thinkingFamilyForModel` → official compat fields; **not** a second options tree |
113
+ | Families | `openai_reasoning` (`reasoning.effort`), `reasoning_effort`, `thinking_type` (`thinking.type`), `noop` (host-owned) — only shapes shared by ≥2 packages (or explicit no-op) |
114
+ | Use-case wiring | LLM compaction + OM workers map `thinkingLevel` into `compat` via helpers (no longer inert `extra.thinkingLevel`) |
115
+ | Docs | [Thinking and reasoning](thinking-and-reasoning.md) |
116
+
117
+ **Decision (2026-07-17):** Implement thin core helpers; keep unique knobs (NeuralWatt budgets/kwargs, Kimi keep/all, Z.AI `tool_stream`) package-local. Core must not name forbidden provider literals (`openrouter`/`zai`/`kimi`/…). Hosts pick a family explicitly when inference is ambiguous. Per-provider first-class field hardening remains Tasks 6–12.
118
+
119
+ ## Use-case model binding inventory (Task 5 — done 2026-07-17)
120
+
121
+ | Site | Current binding | Session fallback today? | Thinking path today |
122
+ | --- | --- | --- | --- |
123
+ | `AgentSession.run` / `RunOptions.model` | Per-run override; writes `model_change` | Explicit override of session model | Host `providerOptions` / `applyThinkingLevel` |
124
+ | Observational memory workers | `resolveUseCaseModel({ configured: workerModel, sessionModel })` | **Yes** — host passes `sessionModel`; `requireExplicitModel` restores skip | `thinkingLevel` → `compat` via `applyThinkingLevel` |
125
+ | LLM compaction | `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` | **Yes** — `model` is the fallback slot | `thinkingLevel` → `compat` via `applyThinkingLevel` |
126
+ | Supervisor children | Child `Agent` owns its own `model` | Independent | Child config |
127
+ | Declarative agents | `resolveAgentDefinition` model string / `ModelConfig` | Definition-scoped | Definition / run options |
128
+ | Evaluations | `runOptions` including model | Eval-owned | Via run options |
129
+ | Workflows / RPC / CLI | Pass-through `runOptions.model` | Caller-owned | Via run options |
130
+ | Structured output | Reuses session/run model | Yes | Same as run |
131
+ | Memory / RAG embedders | Host `Embedder` — separate from chat LLM | N/A (not chat) | N/A |
132
+
133
+ **Decision (2026-07-17):** Core exports `UseCaseModelBinding`, `resolveUseCaseModel`, `resolveUseCaseModelBinding`, `useCaseCredentialProviderId`. Docs: [use-case-model-selection.md](use-case-model-selection.md). OM behavior change: session fallback when `sessionModel` supplied and worker model omitted; `requireExplicitModel` preserves historical `missing_model` skip.
134
+
135
+ Desired Plan 067 outcome: every use-case accepts `{ provider, model, thinking }` with **explicit session-model fallback** when no use-case default is set — **done** for OM + LLM compaction; other sites documented as already-separate.
136
+
137
+ ## Known doc / code mismatches (frozen)
138
+
139
+ | Item | Docs claim | Code / official | Owner task |
140
+ | --- | --- | --- | --- |
141
+ | Z.AI thinking | Was: `thinkingFormat: "zai"`, `developerRoleFallback` in `docs/providers/zai.md` | **fixed** Task 8 — docs+code match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking` | 8 (done) |
142
+ | OpenAI reasoning | Was capabilities-only + opaque compat spread | Task 6 merges `model`/`options` `compat.reasoning` into body `reasoning`; Task 4 helper writes `compat.reasoning.effort` | 4 (done) / 6 (done) |
143
+ | Kimi Moonshot | Was metadata-only (`provider: "moonshot"` without provider) | **fixed** Task 7: `createMoonshotProvider` registered when `includeMoonshotModels`; official Coding ids + `listKimiModels` | 7 (done) |
144
+ | OpenCode Go catalog | Featured official Go open models + `listOpenCodeGoModels` | **done** Task 10 (removed stale `gpt-5.1-go` / `claude-sonnet-4.5-go`) | 10 (done) |
145
+ | Z.AI catalog | Was: `glm-4.7` / `glm-4.5` only | **fixed** Task 8 — featured GLM-5.2…4.5 + `listZaiModels` | 8 (done) |
146
+ | OM/LLM `thinkingLevel` + session fallback | Was inert `extra.thinkingLevel`; OM skipped without workerModel | Task 4 maps into `compat`; Task 5 session fallback + `requireExplicitModel` | 4 (done) / 5 (done) |
147
+
148
+ ## Credential and redaction canaries (per package)
149
+
150
+ | Package | Secrets in flight | Redaction canary expectation |
151
+ | --- | --- | --- |
152
+ | openai | API key; OAuth device/user/access/refresh | Absent from errors, events, requests, discovery failures |
153
+ | kimi | API key | Absent from errors / cache keys |
154
+ | zai | API key | Absent from errors / model metadata |
155
+ | openrouter | API key; session/cache ids | Sanitize length; never treat cache key as secret storage |
156
+ | opencode-go | API key | Absent from errors |
157
+ | neuralwatt | Optional API key; quota responses | Discovery/quota error bodies redacted |
158
+ | ai-sdk | Host-owned | Adapter must not echo host secrets in Prism errors |
159
+
160
+ No secrets are committed in this matrix.
161
+
162
+ ## Plan 067 task map (evidence owners)
163
+
164
+ | Task | Owns |
165
+ | --- | --- |
166
+ | 0 | This page + evidence freeze — **done** |
167
+ | 1 | R-001–R-004 re-verify — **fixed** 2026-07-17 |
168
+ | 2 | R-005–R-006, R-008–R-012 re-verify — **fixed** 2026-07-17 |
169
+ | 3 | Shared discovery pattern — **done** 2026-07-17 (package-local NeuralWatt template; no core list helper) |
170
+ | 4 | Shared thinking / per-turn surface — **done** 2026-07-17 (core helpers + OM/LLM compat wiring; see thinking-and-reasoning.md) |
171
+ | 5 | Use-case model selection + session fallback — **done** 2026-07-17 (`resolveUseCaseModel`, OM session fallback, use-case-model-selection.md) |
172
+ | 6 | OpenAI validate/harden — **done** 2026-07-17 (Responses P0s, `listOpenAIModels`, `models?`, reasoning merge) |
173
+ | 7 | Kimi validate/harden — **done** 2026-07-17 (`listKimiModels`, Moonshot provider, thinking, official Coding ids) |
174
+ | 8 | Z.AI validate/harden — **done** 2026-07-17 (`listZaiModels`, GLM-5.x catalog, docs drift closed, clear_thinking/preserve) |
175
+ | 9 | OpenRouter validate/harden — **done** 2026-07-17 (`listOpenRouterModels`, reasoning merge/preserve, automatic top-level `cache_control`) |
176
+ | 10 | OpenCode Go validate/harden — **done** 2026-07-18 (`listOpenCodeGoModels`, official Go catalog, `zen/go/v1` base, thinking preserve) |
177
+ | 11 | NeuralWatt validate/harden — **done** 2026-07-18 (featured catalog refresh, kwargs routing, owned-compat strip, applyThinkingLevel) |
178
+ | 12 | AI SDK validate — **done** 2026-07-18 (host-owned catalog/cache/reasoning; usage mapping + docs) |
179
+ | 13 | Cross-provider conformance + final verification — **done** 2026-07-18 (`sdk:ready`; 1,089 core tests + workspace suites + pack dry-runs) |
180
+
181
+ Exact task titles live in `plans/067-provider-doc-validation-caching-discovery-and-review-hardening.md`.
182
+
183
+ ## Verification for this page
184
+
185
+ - Final verification completed 2026-07-18 without live provider HTTP.
186
+ - Lists all seven first-party provider packages; every provider row is `fixed` or intentional by-design behavior.
187
+ - Lists every 2026-07-14 P0–P2 id (R-001–R-006, R-008–R-012); all are `fixed`.
188
+ - Distinguishes official-doc priority vs Pi secondary.
189
+ - `phase12-boundaries.test.ts` verifies all six HTTP package discovery exports and setup zero-fetch; AI SDK no-catalog behavior has its own adapter contract test.
190
+ - `provider_validation_final_contract_covers_all_adapters_and_binding_sites` verifies provider docs, cache kinds, thinking rows, use-case sites, matrix statuses, and index navigation.
191
+ - `npm run sdk:ready` passes: typecheck/build, 1,089 core tests, all workspace suites, packaging/provenance guards, and publish-graph pack dry-runs.
192
+ - Linked from `docs/index.md` under Release and install.
@@ -241,7 +241,7 @@ console.log(cacheUsageReport(aggregate?.usage));
241
241
  - `AgentConfig.ownership` is the default ownership scope; `RunOptions.ownership` overrides it per run.
242
242
  - `AgentConfig.idempotencyKey` is the default idempotency key; `RunOptions.idempotencyKey` overrides it per run.
243
243
  - The runtime resolves `model` and `provider` from `AgentConfig`/`RunOptions`/`AgentDefinition` before writing the start `RunRecord`.
244
- - Adapters should treat appends as ordered within a `runId`: event and tool-call rows preserve emission order because the runtime drains pending appends before writing the final `RunRecord`.
244
+ - Adapters should treat appends as ordered within a `runId`: event and tool-call rows preserve emission order because the runtime serializes event ledger appends through one promise chain (concurrency 1), drains pending appends before writing the final `RunRecord`, and propagates append failures by rejecting run completion.
245
245
  - Billing queries must filter `scope = "provider_turn"`; presentation queries normally read the single `run_total`. `UsageQuery.scope`, `turn`, and `attempt` are explicit filters.
246
246
  - Adapters that need upsert semantics can use `RunRecord.id` (== `runId`) as the stable key.
247
247
  - Use `cacheUsageReport(record.usage, model)` for cache diagnostics from normalized usage. It works when a provider reports `cacheReadTokens` without `cacheWriteTokens`; missing write tokens are reported as `0`, and unavailable hit rate/savings stay `undefined`.
@@ -56,7 +56,7 @@ import { createSqlitePersistence } from "@arnilo/prism-session-store-sqlite";
56
56
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
57
57
  | `close()` | Closes the underlying database when the adapter opened it. |
58
58
 
59
- Migrations run automatically on open and are idempotent across reopen.
59
+ Migrations run automatically on open and are idempotent across reopen. Under the SQLite migration transaction, startup checks ordered contract name/version/SHA-256 rows plus full schema-v3 PRAGMA/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
60
60
 
61
61
  ## Request/response example
62
62
 
@@ -109,7 +109,7 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
109
109
  - **No path interpolation.** The adapter opens exactly the caller-supplied `filename`; it does not expand environment variables or discover paths.
110
110
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
111
111
  - **WAL + busy timeout.** WAL is enabled by default; busy timeout defaults to 5 seconds. This meets the Plan 056 local workload target but SQLite still serializes writers — prefer PostgreSQL for high write concurrency.
112
- - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans.
112
+ - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans. Startup validation reads SQLite catalog/PRAGMA metadata only, never application rows.
113
113
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns on run and ownership tables participate in query filters; hosts must still scope writes correctly.
114
114
 
115
115
  ## Related APIs
@@ -234,7 +234,7 @@ Key cross-seam points:
234
234
  - The default parser treats assistant text as the value (`{ ok: true, value: text }`); supply a host parser whenever `T` is not `string`.
235
235
  - The default repairer builds a user message from `validation.errors[].message`; supply a host repairer for schema-specific guidance.
236
236
  - `maxRevisions` (default 3) bounds revision turns; budget exhaustion ends the loop and emits `artifact_failed` (it does not throw).
237
- - Tools are not dispatched in revision turns. Hosts needing tools in artifact turns use `singleShotLoop` or a custom loop.
237
+ - Tools are inert in artifact turns unless `loop.toolCalls: "bounded"` is explicit. Bounded mode uses run-global `maxToolRounds`, dispatches calls sequentially through normal runtime guards, skips parser/validator for tool-calling responses, and permits at most `1 + maxRevisions + maxToolRounds` provider turns. An extra tool response yields terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"` and executes nothing.
238
238
 
239
239
  ## Security and performance notes
240
240
 
@@ -0,0 +1,98 @@
1
+ # Thinking and reasoning
2
+
3
+ ## What it does
4
+
5
+ Prism keeps thinking/reasoning **provider-owned on the wire** while giving hosts one portable way to set effort per turn. Model defaults live on `ModelConfig.compat` (and `capabilities.reasoning` where declared). Per-turn overrides live on `ProviderRequestOptions.compat` and win through existing `mergeProviderRequestOptions`. Shared helpers map a portable `ThinkingLevel` into the official compat fields each family already reads — they do **not** invent a second options tree.
6
+
7
+ ## When to use it
8
+
9
+ - Session runs: pass `providerOptions.compat` (or `applyThinkingLevel`) on `RunOptions`.
10
+ - Use-case workers (LLM compaction, observational memory): pass `thinkingLevel`; packages map it into `compat` via the shared helpers.
11
+ - Provider authors: keep reading official fields from `options.compat` / `model.compat`; add package-local escape hatches only when the official API has unique knobs.
12
+
13
+ ## Contract
14
+
15
+ | Layer | Surface |
16
+ | --- | --- |
17
+ | Model default | `ModelConfig.compat` (+ `capabilities.reasoning` when the model can reason) |
18
+ | Per-turn override | `ProviderRequestOptions.compat` (request wins over model via merge) |
19
+ | Portable level | `ThinkingLevel`: `none` \| `minimal` \| `low` \| `medium` \| `high` \| `xhigh` \| `max` |
20
+ | Helpers | `thinkingCompatFor`, `applyThinkingLevel`, `thinkingFamilyForModel`, `isThinkingLevel`, `normalizeThinkingLevel`, `THINKING_LEVELS` |
21
+ | Not used | Inert `options.extra.thinkingLevel` — providers do not read `extra` for effort |
22
+
23
+ ```ts
24
+ import { applyThinkingLevel, thinkingCompatFor, thinkingFamilyForModel } from "@arnilo/prism";
25
+
26
+ // Per-turn override on a session run (OpenAI / OpenRouter family)
27
+ await session.run(input, {
28
+ providerOptions: applyThinkingLevel(undefined, "low", "openai_reasoning"),
29
+ });
30
+
31
+ // Equivalent explicit compat
32
+ await session.run(input, {
33
+ providerOptions: { compat: thinkingCompatFor("openai_reasoning", "low") },
34
+ // → { reasoning: { effort: "low" } }
35
+ });
36
+
37
+ // Use-case worker: family from model metadata (or pass an explicit family)
38
+ const family = thinkingFamilyForModel(model);
39
+ await runObserver({
40
+ ...,
41
+ providerOptions: applyThinkingLevel(base, "low", family === "noop" ? "reasoning_effort" : family),
42
+ });
43
+ ```
44
+
45
+ ## Compat families
46
+
47
+ Core maps only shapes shared by ≥2 packages (or an explicit no-op). Unique knobs stay package-local.
48
+
49
+ | Family | Compat patch | Used by (official fields) |
50
+ | --- | --- | --- |
51
+ | `openai_reasoning` | `{ reasoning: { effort } }` | OpenAI Responses `reasoning.effort`; OpenRouter `reasoning.effort` |
52
+ | `reasoning_effort` | `{ reasoning_effort }` | Z.AI `reasoning_effort`; NeuralWatt `reasoning_effort`; Kimi K3 `reasoning_effort` |
53
+ | `thinking_type` | `{ thinking: { type: "enabled" \| "disabled" } }` | Z.AI `thinking.type`; Kimi K2.x `thinking.type` (`none` → `disabled`) |
54
+ | `noop` | `{}` | AI SDK / host-owned adapters — effort is host-model settings |
55
+
56
+ `applyThinkingLevel` defaults `family` to `reasoning_effort` when omitted. For `openai_reasoning`, an existing `compat.reasoning.summary` (or other reasoning keys) is preserved when merging `effort`.
57
+
58
+ ### Recommended family by first-party package
59
+
60
+ | Package | Recommended family | Notes |
61
+ | --- | --- | --- |
62
+ | `@arnilo/prism-provider-openai` | `openai_reasoning` | First-class body `reasoning` from model + per-turn compat merge; `summary`/`mode`/`context` via compat |
63
+ | `@arnilo/prism-provider-openrouter` | `openai_reasoning` | First-class `resolveOpenRouterReasoning` merge; prefer `reasoning` object over legacy `reasoning_effort` shorthand; `preserveThinking` replays as body `reasoning` |
64
+ | `@arnilo/prism-provider-zai` | `reasoning_effort` (+ optional `thinking_type`) | Official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking`; Preserved Thinking via `reasoning_content` |
65
+ | `@arnilo/prism-provider-neuralwatt` | `reasoning_effort` | Budgets / `preserve_thinking` / `clear_thinking` / `chat_template_kwargs` stay package-local on `compat` |
66
+ | `@arnilo/prism-provider-kimi` | K3: `reasoning_effort`; K2.x: `thinking_type` | K2.7-code thinking is always on; do not send conflicting `thinking` + `reasoning_effort` |
67
+ | `@arnilo/prism-provider-opencode-go` | Anthropic route: thinking blocks (`thinking_type` family); OpenAI route: `reasoning_content` preserve + optional `thinking`/`reasoning_effort`/`reasoning` passthrough | Official dual endpoints; MiniMax/Qwen → Anthropic, others → OpenAI |
68
+ | `@arnilo/prism-provider-ai-sdk` | `noop` | Host `LanguageModelV4` owns reasoning settings |
69
+
70
+ `thinkingFamilyForModel` infers family from existing `compat` shape, then safe provider heuristics (`openai*` → `openai_reasoning`, `neuralwatt` → `reasoning_effort`), then `capabilities.reasoning` → `reasoning_effort`, else `noop`. Docs and packages may map other provider ids explicitly; core avoids provider-specific literals beyond those heuristics.
71
+
72
+ ## Merge order
73
+
74
+ 1. `ModelConfig.compat` / model defaults inside the provider
75
+ 2. `ProviderRequestOptions.compat` from agent / session policies
76
+ 3. Per-turn `RunOptions.providerOptions` or use-case `applyThinkingLevel` patch (wins)
77
+
78
+ Providers already prefer `request.options.compat.*` over `request.model.compat.*`.
79
+
80
+ ## Use-case workers
81
+
82
+ LLM compaction and observational memory accept `thinkingLevel?: string`. They call `applyThinkingLevel` into `compat` (not `extra.thinkingLevel`). When model inference returns `noop`, an explicit `thinkingLevel` still falls back to `reasoning_effort` so the host setting is never inert. Model selection for those workers (including session-model fallback) is documented in [Use-case model selection](use-case-model-selection.md).
83
+
84
+ ## Non-reasoning models
85
+
86
+ - Helper with `noop`: returns options unchanged — no invented body fields.
87
+ - Helper with a real family on a model that rejects the field: provider/API error — hosts should gate on `capabilities.reasoning` or package docs.
88
+ - `thinking_type` + `none` sets `{ type: "disabled" }`; other levels set `{ type: "enabled" }` without encoding effort (compose with `reasoning_effort` when the API supports both).
89
+
90
+ ## Related pages
91
+
92
+ - [Use-case model selection](use-case-model-selection.md) — session vs worker/summary model binding
93
+ - [Provider packages](provider-packages.md) — package boundaries and discovery
94
+ - [Provider caching](provider-caching.md) — cache retention can disable thinking on some providers (e.g. Z.AI when `cacheRetention: "none"`)
95
+ - [Provider request policies](provider-request-policies.md) — `mergeProviderRequestOptions`
96
+ - [Agent/session runtime](agent-session-runtime.md) — prior-reasoning preservation across turns
97
+ - Per-provider pages under [docs/providers](providers/)
98
+ - Evidence matrix: [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md)
@@ -123,8 +123,8 @@ Package performs **no** `PermissionPolicy`, `ToolValidator`, or trust checks of
123
123
  | --- | --- |
124
124
  | `ToolDefinition.parameters` | Stored and forwarded to providers; **not validated** by core |
125
125
  | `ToolValidator` | Host function hook; Phase 25 threads through agent runtime |
126
- | Standards-based schema validation | **Not shipped** — capability gap C-001 |
127
- | Schema compile cache | **None** — every dispatch would re-validate if host validator is naive |
126
+ | Standards-based schema validation | Optional `@arnilo/prism-tool-validator-json-schema`; host wires it through `ToolValidator` |
127
+ | Schema compile cache | Adapter-owned finite LRU; core never compiles schemas |
128
128
 
129
129
  ### MCP mapping (shipped — Task 3)
130
130
 
@@ -186,7 +186,7 @@ import { createJsonSchemaToolArgumentValidator } from "@arnilo/prism-tool-valida
186
186
  createAgent({ model, validator: createJsonSchemaToolArgumentValidator() });
187
187
  ```
188
188
 
189
- **Cache key:** stable `JSON.stringify(schema)` in adapter-owned `Map`. **Bounds:** configurable depth/properties/string/array limits before Ajv validation. **Security:** remote `$ref` rejected; prototype-pollution keys rejected in schemas and instances.
189
+ **Cache key:** stable `JSON.stringify(schema)` in an adapter-owned 256-entry LRU (hard cap 1,024); eviction removes the matching Ajv schema. **Bounds:** schemas default to 256 KiB, depth 64, 10,000 properties/keywords, and 128 refs (hard 1 MiB/128/100,000/1,024); instance depth/properties/string/array limits remain configurable. Every limit rejects non-finite, unsafe, zero/negative, and above-hard values. **Security:** only fragment-local `$ref` is accepted; prototype-pollution keys, cycles, and non-finite schema numbers reject before Ajv compilation.
190
190
 
191
191
  ### Task 2 — Parallel tool execution — **shipped**
192
192
 
package/docs/tools.md CHANGED
@@ -154,6 +154,10 @@ const agent = createAgent({ model, provider, tools: activeTools, permission, val
154
154
 
155
155
  Need different tools for one request? Build a short-lived agent/session with a narrower registry, or block extra calls with `PermissionPolicy` / `RunOptions.validate`. No extra per-run tool API exists yet; add one only when host apps need it.
156
156
 
157
+ ### Artifact-loop tools
158
+
159
+ `generate-validate-revise` treats provider tools as inert by default. Set `loop.toolCalls: "bounded"` and `RunOptions.maxToolRounds` only when an artifact needs a host-owned lookup before its next candidate. Each response with one-or-more calls consumes one shared round, dispatches calls sequentially through this exact `dispatchToolCall()` path, persists assistant-call then result transcript rows, and skips artifact parsing/validation for that response. A post-limit call executes nothing; the loop emits `artifact_failed` with `metadata.reason: "tool_round_limit"`. Tools do not consume `maxRevisions`, and tool schemas/context never grant authority.
160
+
157
161
  ### Runtime-supplied validators
158
162
 
159
163
  `AgentConfig.validator?` and `RunOptions.validate?` expose the same `ToolValidator` seam that `DispatchToolCallOptions.validate` already uses. The runtime threads `validate: RunOptions.validate ?? AgentConfig.validator` into every `dispatchToolCall` it issues during the tool loop, so an app can supply argument validation without taking ownership of dispatch itself. `RunOptions.validate` overrides `AgentConfig.validator` on a per-run basis (RunOptions wins). When neither is set, dispatch runs unmodified.
@@ -229,6 +233,17 @@ await session.run(input, {
229
233
  - Contribution registration and registry/filter calls do not perform provider calls, credential resolution, resource loading, network, filesystem discovery, or tool execution.
230
234
  - Dispatch performs explicit in-memory checks and executes only the selected host-active tool; it adds no retries, queues, timers, or new dependencies.
231
235
 
236
+ ## JSON Schema validator limits
237
+
238
+ Core stores `ToolDefinition.parameters` but does not compile schemas. Hosts that install `@arnilo/prism-tool-validator-json-schema` receive pre-Ajv schema limits: 256 KiB bytes, depth 64, 10,000 properties/keywords, 128 refs, and a 256-entry LRU compiled cache by default. All reject invalid values and have finite hard ceilings. Only fragment-local `$ref` values are accepted; non-local refs, cycles, forbidden keys, and non-finite schema numbers fail before tool execution.
239
+
240
+ ```ts
241
+ createJsonSchemaToolArgumentValidator({
242
+ maxSchemaBytes: 256 * 1024,
243
+ maxCompiledSchemas: 256,
244
+ });
245
+ ```
246
+
232
247
  ## Related APIs
233
248
 
234
249
  - [Agent/session runtime](agent-session-runtime.md): dispatches complete provider tool calls through the host-active tool harness and returns tool results on the next provider turn.
@@ -0,0 +1,109 @@
1
+ # Use-case model selection
2
+
3
+ ## What it does
4
+
5
+ Prism separates the **session chat model** (`AgentConfig.model` / `RunOptions.model`) from **use-case models** used by background or adjacent LLM jobs (observational memory workers, LLM compaction summarizers, declarative agents, supervisor children, evals). Hosts bind `{ model?, provider?, providerOptions?, thinkingLevel? }` per use case. When the use-case omits `model`, resolution falls back to the active session model. Workers never write `model_change` session entries for their own jobs.
6
+
7
+ ## When to use it
8
+
9
+ - Observational memory should run a cheaper/faster model than the chat session (or inherit the session model when unset).
10
+ - LLM compaction should summarize with an explicit `summaryModel`, falling back to a host-supplied session `model`.
11
+ - Declarative agents, supervisor children, evals, and RPC/CLI runs already own their models — document them as use-case sites that stay separate from a parent session.
12
+ - Memory/RAG `Embedder` selection is related but **not** a chat `ModelConfig` binding.
13
+
14
+ ## Contract
15
+
16
+ | Layer | Surface |
17
+ | --- | --- |
18
+ | Binding | `UseCaseModelBinding`: `{ model?, provider?, providerOptions?, thinkingLevel?, requireExplicitModel? }` |
19
+ | Resolver | `resolveUseCaseModel({ configured, sessionModel, requireExplicitModel?, … })` → `{ model, source }` or `undefined` |
20
+ | Binding helper | `resolveUseCaseModelBinding(binding, sessionModel)` |
21
+ | Credential id | `useCaseCredentialProviderId(resolved, binding?)` — always the **resolved** `model.provider` |
22
+ | Escape hatch | `requireExplicitModel: true` skips session fallback (OM historical `missing_model`) |
23
+
24
+ ```ts
25
+ import { resolveUseCaseModel, applyThinkingLevel, thinkingFamilyForModel } from "@arnilo/prism";
26
+
27
+ // Prefer an explicit worker; otherwise inherit the session/agent model.
28
+ const resolved = resolveUseCaseModel({
29
+ configured: settings.workerModel, // optional use-case ModelConfig
30
+ sessionModel: agent.config.model, // host-supplied; AgentSession does not expose agent
31
+ thinkingLevel: settings.thinkingLevel,
32
+ });
33
+ if (!resolved) {
34
+ // skip — neither configured nor session model (or requireExplicitModel)
35
+ }
36
+
37
+ const family = thinkingFamilyForModel(resolved.model);
38
+ const providerOptions = resolved.thinkingLevel
39
+ ? applyThinkingLevel(resolved.providerOptions, resolved.thinkingLevel, family === "noop" ? "reasoning_effort" : family)
40
+ : resolved.providerOptions;
41
+ ```
42
+
43
+ ### Precedence
44
+
45
+ 1. `configured` / `binding.model` → `source: "configured"`
46
+ 2. Else `sessionModel` when `requireExplicitModel` is not set → `source: "session"`
47
+ 3. Else `undefined` (package skips or throws)
48
+
49
+ Resolution is O(1) and network-free. It does not mutate session history.
50
+
51
+ ## Binding sites
52
+
53
+ | Site | How hosts bind | Session fallback |
54
+ | --- | --- | --- |
55
+ | Observational memory | `workerModel` / settings `workerModel` + runtime `sessionModel` | Yes — pass `sessionModel: agent.config.model`; `requireExplicitModel` restores skip |
56
+ | LLM compaction | `summaryModel` with `model` as fallback slot | Yes — `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` |
57
+ | `RunOptions.model` | Per-run override on the **session** | N/A — this *is* the session/run model (writes `model_change`) |
58
+ | Declarative `AgentDefinition` | Definition `model` / registry resolve | Definition-scoped (independent agent) |
59
+ | Supervisor children | Child `createSession` / child `AgentConfig.model` | Independent child session |
60
+ | Evals / workflows / RPC / CLI | Caller `runOptions.model` | Caller-owned |
61
+ | Structured output | Reuses session/run model | Same as run |
62
+ | Memory / RAG | Host `Embedder` | Not a chat model — see [Working and semantic memory](working-and-semantic-memory.md) |
63
+
64
+ ## Observational memory
65
+
66
+ ```ts
67
+ createObservationalMemoryRuntime({
68
+ session,
69
+ appendEntry: (entry) => store.append(entry),
70
+ workerProvider,
71
+ sessionModel: agent.config.model, // enables fallback when workerModel unset
72
+ // workerModel: { provider: "neuralwatt", model: "glm-5.2-fast" }, // optional override
73
+ overrides: { thinkingLevel: "low", observeAfterTokens: 1 },
74
+ });
75
+ ```
76
+
77
+ - Default: no `workerModel` + `sessionModel` set → workers use the session model.
78
+ - Explicit `workerModel` (or settings `workerModel`) always wins.
79
+ - `requireExplicitModel: true` (runtime or settings) → `skipped: "missing_model"` when no worker model, even if `sessionModel` is set.
80
+ - Neither worker nor session model → `skipped: "missing_model"`.
81
+ - Default credential request uses the **resolved** model's `provider` id.
82
+
83
+ ## LLM compaction
84
+
85
+ ```ts
86
+ createLlmCompactionStrategy({
87
+ provider: summaryProvider,
88
+ summaryModel: { provider: "example", model: "cheap-summary" }, // optional
89
+ model: agent.config.model, // session fallback when summaryModel omitted
90
+ thinkingLevel: "low",
91
+ });
92
+ ```
93
+
94
+ `summaryModel` wins; otherwise `model` is required. Thinking maps into `compat` via `applyThinkingLevel` ([Thinking and reasoning](thinking-and-reasoning.md)).
95
+
96
+ ## Security
97
+
98
+ - Credential requests for worker calls must target the **resolved** model’s provider — not ambient session credentials for a different provider unless the host wires that explicitly.
99
+ - Pass known secrets into worker/compaction options so prompts, ledger custom entries, and errors stay redacted.
100
+ - Background workers must not append `model_change` entries or otherwise rewrite the chat session’s model timeline.
101
+
102
+ ## Related pages
103
+
104
+ - [Thinking and reasoning](thinking-and-reasoning.md) — per-turn `thinkingLevel` → `compat`
105
+ - [Observational memory compaction package](compaction-observational-memory.md)
106
+ - [LLM compaction package](compaction-llm.md)
107
+ - [Agent/session runtime](agent-session-runtime.md) — `RunOptions.model` / `model_change`
108
+ - [Working and semantic memory](working-and-semantic-memory.md) — `Embedder` (non-chat)
109
+ - Evidence matrix: [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md)
@@ -387,6 +387,7 @@ const review = functionNode({ execute: async (ctx) => lint(ctx.upstream.draft) }
387
387
 
388
388
  const workflow = defineWorkflow({
389
389
  id: "research-draft-review",
390
+ revision: "2026-07-19.1",
390
391
  nodes: { research, draft, review },
391
392
  edges: [
392
393
  ["research", "draft"],
package/docs/workflows.md CHANGED
@@ -29,29 +29,32 @@ Use `createWorkflowCoordinator()` when multiple processes share SQLite/PostgreSQ
29
29
 
30
30
  ## Inputs / request
31
31
 
32
- `defineWorkflow({ id, nodes, edges, limits? })`:
32
+ `defineWorkflow({ id, revision, nodes, edges, limits? })`:
33
33
 
34
34
  | Field | Notes |
35
35
  | --- | --- |
36
36
  | `id` | Stable workflow id (required) |
37
+ | `revision` | Non-empty host-authored definition revision (required); parent and nested revisions enter `definitionHash` |
37
38
  | `nodes` | Record of node definitions (`kind` + typed fields) |
38
39
  | `edges` | `[from, to]` pairs; must be acyclic; unknown ids rejected |
39
- | `limits.maxNodes` | Default 1000 |
40
- | `limits.maxFanOut` | Default 64 |
41
- | `limits.maxConcurrency` | Default 8 |
42
- | `limits.maxNodeOutputBytes` | Default 4 MiB |
43
- | `limits.maxCheckpointBytes` | Default 1 MiB |
40
+ | `limits.maxNodes` | Default 1,000 / hard cap 10,000 |
41
+ | `limits.maxFanOut` | Default 64 / hard cap 1,024 |
42
+ | `limits.maxConcurrency` | Default 8 / hard cap 256 |
43
+ | `limits.maxNodeOutputBytes` | Default 4 MiB / hard cap 16 MiB |
44
+ | `limits.maxCheckpointBytes` | Default 1 MiB / hard cap 8 MiB |
44
45
  | `limits.maxNestedDepth` / hard cap | 8 / 32; inherited by child workflows |
45
46
  | `limits.maxStateBytes` / hard cap | 64 KiB / 512 KiB |
46
47
  | `limits.maxStateHistory` / hard cap | 32 / 128 state snapshots; updates stop before evidence would be discarded |
47
48
  | `limits.maxReplayDepth` / hard cap | 8 / 32 lineage generations |
48
49
  | `state.initial` / `state.schema` | Initial shared JSON object and optional host-validated schema |
49
50
 
51
+ All workflow limits and runtime `concurrency` reject non-safe integers, zero, negatives, NaN, `Infinity`, and values above the named hard cap. Node retries allow 0–100; an explicit node timeout allows 1–86,400,000 ms. Omitting `timeoutMs` remains an explicit host choice.
52
+
50
53
  `runWorkflow(workflow, input, options?)`:
51
54
 
52
55
  | Option | Notes |
53
56
  | --- | --- |
54
- | `concurrency` | Worker pool size (capped by workflow/global limits) |
57
+ | `concurrency` | Worker pool size; positive safe integer, hard cap 256, and capped by the workflow limit |
55
58
  | `checkpoints` | `WorkflowCheckpointAdapter` for save/load/list |
56
59
  | `agentFactory` | `(agentName) => AgentSession` for agent nodes |
57
60
  | `tools` | Tool registry/lookup for tool nodes |
@@ -100,6 +103,7 @@ Package-local `WorkflowEvent` types: `workflow_started`, `workflow_suspended`, `
100
103
  ```json
101
104
  {
102
105
  "id": "research-draft",
106
+ "revision": "2026-07-19.1",
103
107
  "nodes": ["research", "draft"],
104
108
  "edges": [["research", "draft"]],
105
109
  "limits": { "maxNodes": 256, "maxFanOut": 32, "maxConcurrency": 4 }
@@ -159,6 +163,7 @@ const publish = functionNode({
159
163
 
160
164
  const workflow = defineWorkflow({
161
165
  id: "research-draft",
166
+ revision: "2026-07-19.1",
162
167
  nodes: { research, draft, publish },
163
168
  edges: [["research", "draft"], ["draft", "publish"]],
164
169
  limits: { maxNodes: 256, maxFanOut: 32, maxConcurrency: 4 },
@@ -209,6 +214,7 @@ if (result.status === "suspended") {
209
214
  await cancelWorkflowRun({
210
215
  workflowId: workflow.id,
211
216
  runId: result.runId,
217
+ workflow,
212
218
  checkpoints,
213
219
  ownership: { tenantId: "t1" },
214
220
  });
@@ -258,15 +264,16 @@ runRpcServer({
258
264
 
259
265
  ## Security and performance notes
260
266
 
261
- - Definitions fail closed on cycles, unknown edges, self-edges, and `maxNodes` overflow.
262
- - Fan-out length is bounded by `maxFanOut`; concurrency by `maxConcurrency`.
267
+ - Definitions require a non-empty host-authored `revision` and fail closed on cycles, unknown edges, self-edges, invalid limits, and `maxNodes` overflow. Revision and every nested revision enter the deterministic definition hash; hosts must bump revision when function/tool behavior changes.
268
+ - Fan-out length is bounded by `maxFanOut`; concurrency by `maxConcurrency`; every count/byte/runtime option has a finite hard cap.
263
269
  - Node outputs, shared state/history, schedule input/records, and checkpoints are byte/count/depth bounded. Checkpoint size remains the final aggregate ceiling.
264
270
  - Event buses use a bounded buffer (default 2048) with `close` / `drop_oldest` / `drop_newest` overflow.
265
271
  - Checkpoints redact suspension/resume payloads via `SecretRedactor` / `secrets` before save; resume rejects tenant, schema, definition-hash, and expected-version mismatch.
266
272
  - Suspension requires a checkpoint adapter, consumes no worker/polling slot, and is ignored by distributed coordinators until explicit resume.
267
273
  - Concurrent resumes race on checkpoint CAS before node execution; one wins and stale/duplicate reviewers fail closed. Approved tool nodes then re-run current `ExecutionPolicy`, so durable approval cannot grant stale permissions.
268
274
  - `toolNode({ approval: { reason, data?, resumeSchema? } })` suspends before tool execution. Denial is terminal `denied`; no tool side effect occurs.
269
- - `cancelWorkflowRun` aborts local runs immediately and writes a durable cancellation request for a remotely leased run; workers check it during lease renewal.
275
+ - `cancelWorkflowRun` requires the current workflow definition and exact tenant/account/user ownership. It verifies recursive definition hash before abort/mutation, then aborts local runs or writes a durable cancellation request for remotely leased work. Tenant-only or missing ownership cannot cancel a more-specific owned run.
276
+ - Active registry identity includes workflow ID, run ID, and exact ownership. Exact duplicates fail instead of overwriting; distinct owners remain isolated in lookup/list/cancel/unregister.
270
277
  - Tool nodes attach `workflowId` / `nodeId` on `ExecutionAction.metadata` for approval/audit context.
271
278
  - Nested workflows inherit host registries/policies and cannot inject broader tools, agents, ownership, or credentials. Nested depth is inherited; child suspension bubbles to the parent review cursor.
272
279
  - Replay source ownership/hash/status/node eligibility are checked before a new checkpoint is created. Source records are immutable, lineage is bounded, and copied approval-bearing paths are rejected.
@@ -152,6 +152,7 @@ await runMemoryConformance(() => ({
152
152
  - Configure `secrets` / `redactor` so memory text and metadata cannot persist or inject raw canaries.
153
153
  - Injected context is inert text — it cannot grant tools or permissions.
154
154
  - Hard caps: top-K ≤ 32, messageRange ≤ 4, embed batch ≤ 128, injected tokens ≤ 8000, payload/working-memory byte limits enforced.
155
+ - Every embedding is a non-empty finite number vector. `embedBatched()`, in-memory `VectorStore` upserts/queries, and PostgreSQL/pgvector parameters reject NaN, ±Infinity, non-numbers, and wrong configured dimensions before similarity scoring or SQL. Custom adapters can call `assertFiniteVector(vector, label, expectedLength?)` at their trust boundary.
155
156
  - Default `remember()` does not block agent completion; pass `{ wait: true }` when indexing must finish first.
156
157
  - PostgreSQL live suite is gated by `PRISM_TEST_POSTGRES_URL` and requires the `vector` extension.
157
158