@arnilo/prism 0.0.4 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/CHANGELOG.md +46 -1
  2. package/README.md +34 -10
  3. package/dist/agent-loops.d.ts +1 -0
  4. package/dist/agent-loops.js +26 -16
  5. package/dist/agents.js +147 -21
  6. package/dist/cli-init.d.ts +41 -0
  7. package/dist/cli-init.js +390 -0
  8. package/dist/cli-runner.d.ts +7 -1
  9. package/dist/cli-runner.js +13 -1
  10. package/dist/content.d.ts +19 -0
  11. package/dist/content.js +197 -69
  12. package/dist/contracts.d.ts +96 -9
  13. package/dist/contracts.js +8 -0
  14. package/dist/feedback.d.ts +48 -0
  15. package/dist/feedback.js +230 -0
  16. package/dist/ids.d.ts +2 -0
  17. package/dist/ids.js +6 -0
  18. package/dist/index.d.ts +10 -4
  19. package/dist/index.js +6 -3
  20. package/dist/providers/media.d.ts +3 -1
  21. package/dist/providers/media.js +11 -1
  22. package/dist/session-stores.js +2 -3
  23. package/dist/testing/feedback.d.ts +6 -0
  24. package/dist/testing/feedback.js +37 -0
  25. package/dist/testing/persistence-schema.d.ts +48 -10
  26. package/dist/testing/persistence-schema.js +166 -22
  27. package/dist/testing/run-ledger-conformance.js +7 -1
  28. package/dist/thinking.d.ts +42 -0
  29. package/dist/thinking.js +92 -0
  30. package/dist/tools.js +2 -3
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +75 -0
  34. package/docs/agent-events.md +14 -21
  35. package/docs/agent-loops.md +12 -9
  36. package/docs/agent-session-runtime.md +14 -16
  37. package/docs/cli-rpc.md +35 -7
  38. package/docs/coding-agent-tools.md +35 -14
  39. package/docs/coding-security.md +7 -3
  40. package/docs/compaction-llm.md +17 -7
  41. package/docs/compaction-observational-memory.md +30 -4
  42. package/docs/context-and-skills.md +1 -0
  43. package/docs/credential-storage.md +58 -9
  44. package/docs/credentials-and-redaction.md +3 -3
  45. package/docs/database-persistence.md +17 -9
  46. package/docs/evaluations.md +122 -0
  47. package/docs/extensions.md +2 -2
  48. package/docs/host-security.md +26 -5
  49. package/docs/index.md +43 -28
  50. package/docs/mcp-tools.md +74 -13
  51. package/docs/migration.md +177 -3
  52. package/docs/multimodal-content.md +14 -6
  53. package/docs/node-filesystem-config.md +1 -0
  54. package/docs/node-jsonl-session-store.md +5 -4
  55. package/docs/observability.md +14 -6
  56. package/docs/performance.md +209 -0
  57. package/docs/postgres-persistence.md +8 -6
  58. package/docs/provider-caching.md +16 -4
  59. package/docs/provider-conformance.md +40 -1
  60. package/docs/provider-packages.md +62 -3
  61. package/docs/providers/ai-sdk.md +149 -0
  62. package/docs/providers/kimi.md +124 -61
  63. package/docs/providers/neuralwatt.md +19 -13
  64. package/docs/providers/openai.md +56 -13
  65. package/docs/providers/opencode-go.md +118 -30
  66. package/docs/providers/openrouter.md +105 -35
  67. package/docs/providers/zai.md +94 -45
  68. package/docs/public-contracts.md +6 -5
  69. package/docs/rag.md +113 -0
  70. package/docs/release-and-install.md +100 -79
  71. package/docs/review-coverage-2026-07-15.md +193 -0
  72. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  73. package/docs/runs-and-usage.md +42 -5
  74. package/docs/server.md +139 -0
  75. package/docs/settings-auth-trust-security.md +5 -5
  76. package/docs/sqlite-persistence.md +6 -5
  77. package/docs/structured-output.md +1 -1
  78. package/docs/supervisors.md +71 -0
  79. package/docs/thinking-and-reasoning.md +98 -0
  80. package/docs/tool-execution-primitives.md +3 -3
  81. package/docs/tools.md +15 -0
  82. package/docs/use-case-model-selection.md +109 -0
  83. package/docs/workflow-orchestration-primitives.md +20 -3
  84. package/docs/workflows.md +114 -33
  85. package/docs/working-and-semantic-memory.md +170 -0
  86. package/package.json +13 -3
  87. package/templates/init/README.md.tmpl +28 -0
  88. package/templates/init/env.example.tmpl +1 -0
  89. package/templates/init/gitignore.tmpl +11 -0
  90. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  91. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  92. package/templates/init/package.json.tmpl +22 -0
  93. package/templates/init/providers.json +76 -0
  94. package/templates/init/src/agent.ts.tmpl +10 -0
  95. package/templates/init/src/index.ts.tmpl +12 -0
  96. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  97. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -0,0 +1,193 @@
1
+ # Review coverage — 2026-07-15
2
+
3
+ This page freezes Prism 0.0.5 scope at commit `f5128a816ae204c52f3e2f089de71c99bd5de6d4` after the Prism/Mastra review. It maps every confirmed finding and accepted capability to one owning roadmap phase, package/public surface, focused test owner, and documentation owner.
4
+
5
+ Status: **scope frozen; Phases 1-14 implementation complete; publication handoff pending**. Phase 0 changes no public runtime behavior.
6
+
7
+ ## Release decision
8
+
9
+ - Target release is **0.0.5**. Do not publish the currently versioned but unpublished 0.0.4 package graph.
10
+ - npm currently reports `@arnilo/prism@0.0.3`; Phase 14 has retargeted every repository package and lock entry to 0.0.5. Publication remains pending the clean signed-tag workflow.
11
+ - Core remains a Node.js 20-compatible, zero-runtime-dependency harness.
12
+ - New capabilities are opt-in and package-owned. No server, database, credential store, telemetry exporter, memory worker, schedule, remote agent, or privileged tool activates on install.
13
+ - Any row below that changes a trust boundary remains a release blocker until its owning phase and regression checks pass.
14
+
15
+ ## Frozen confirmed findings
16
+
17
+ Each ID appears once. “Planned” means owned, not fixed.
18
+
19
+ | ID | Confirmed finding | Disposition and single owner | Package / public surface | Focused test owner | Documentation owner | Status |
20
+ | --- | --- | --- | --- | --- | --- | --- |
21
+ | S-001 | Bracket-normalization lets IPv6 loopback, link-local, ULA, and IPv4-mapped private literals bypass media SSRF checks; DNS resolution/rebinding behavior is not sufficient for a private-network denial claim. | Phase 1 | Core `src/content.ts`; `SsrfPolicy` and media fetch path | `src/__tests__/content.test.ts` with injected resolver/requester matrix | `multimodal-content.md`, `host-security.md` | completed 2026-07-15 |
22
+ | S-002 | `createReadOnlyTools()` does not apply the aggregate `executionPolicy`. | Phase 1 | `@arnilo/prism-coding-agent` read-only aggregator | coding-agent execution-policy tests | `coding-agent-tools.md` | completed 2026-07-15 |
23
+ | S-003 | Approval caching defaults to a process-lifetime map instead of `none`; `run` and `session` scopes have no identity in the cache key. | Phase 1 | `@arnilo/prism-coding-security`; approval policy and execution identity metadata | coding-security approval tests | `coding-security.md`, `host-security.md` | completed 2026-07-15 |
24
+ | C-001 | Sandbox adapter returns only an exit code and cannot forward stdout/stderr to coding-agent `onData`. | Phase 2 | coding-security sandbox adapter / coding-agent bash operations bridge | sandbox and shell integration tests | `coding-security.md`, `coding-agent-tools.md` | completed 2026-07-15 |
25
+ | C-002 | OpenTelemetry agent spans end only on `agent_finished`; provider failure can leave `prism.agent.run` active. | Phase 2 | `@arnilo/prism-observability-opentelemetry` instrumentation lifecycle | in-memory telemetry failed/aborted-run tests | `observability.md` | completed 2026-07-15 |
26
+ | C-003 | `prism.provider.tokens` is incremented for provider-turn and agent-total events, double-counting one usage source. | Phase 2 | OTel token metric names/scopes | in-memory metric assertions | `observability.md`, `runs-and-usage.md` | completed 2026-07-15 |
27
+ | C-004 | Usage rows do not distinguish provider-turn values from run totals; the loop returns the latest turn rather than an explicit aggregate. | Phase 2 | core `Usage`, `UsageRecord`, runtime accumulator, persistence adapters | multi-turn run-ledger conformance and runtime tests | `runs-and-usage.md`, persistence docs | completed 2026-07-15 |
28
+ | C-005 | Providers resolve media one block at a time, so request-wide 32-item/32-MiB bounds are not accumulated before provider/upload side effects. | Phase 2 | core media request resolver and provider media serializers | content/provider-media/provider package tests | `multimodal-content.md` | completed 2026-07-15 |
29
+ | A-001 | Workflow coordinator concurrency test was timing-sensitive under load. | Phase 0 | test only; production workflow API unchanged | `packages/workflows/src/__tests__/coordinator.test.ts` repeated five times plus full suite | this page, `performance.md` | completed at frozen HEAD `f5128a8` |
30
+ | A-002 | `AgentConfig.extensions`, `settings`, and `credentials` are accepted but explicitly ignored by agent/session runtime. | Phase 3 | core `AgentConfig` and host composition docs | compile fixtures and runtime behavior tests | `agent-session-runtime.md`, `customization.md`, `migration.md` | completed 2026-07-15 |
31
+ | A-003 | `AgentSession.run()`/`prompt()` return `Promise<void>` and integrated streaming requires a separately started subscriber. | Phase 3 | core `AgentSession`, new run-result/stream contract | agent/session, CLI/RPC, workflow tests | agent/session, events, CLI/RPC, workflow docs | completed 2026-07-15 |
32
+ | A-004 | Direct source/phase-text boundary tests make refactors fail without behavioral changes. | Phase 14 | test architecture; replace implementation assertions when touched, retain manifest/docs/absence boundary checks | docs/export/behavior/pack tests replacing implementation-text assertions | testing/release coverage docs | completed for touched tests; broad historical conversion deferred |
33
+ | A-005 | `src/contracts.ts`, `src/agents.ts`, and `packages/workflows/src/run.ts` are conflict hotspots; docs/plans are larger than production source. | Phase 14 | maintainability gate; split only where completed work proves a cohesive domain | typecheck, public exports, behavior suites, docs links | review coverage and release docs | reviewed; no safe release-time cohesive split, deferred with measurements |
34
+ | R-001 | 0.0.4 must not be published after the review; all package/version/tag/provenance checks must target 0.0.5. | Phase 14 | 30-package graph, lockfile, release workflow/script | release, packaging, install, Node 20/24, registry checks | `release-and-install.md`, changelogs, migration docs | completed 2026-07-16; publication handoff pending |
35
+
36
+ ## Accepted capability scope
37
+
38
+ These are the twelve accepted Mastra-comparison recommendations. Each has one implementation owner; adjacent phases may consume the result but do not own a duplicate implementation.
39
+
40
+ | ID | Capability | Single owner | Reused Prism primitives | Minimum missing surface | Test owner | Documentation owner |
41
+ | --- | --- | --- | --- | --- | --- | --- |
42
+ | F-001 | Direct run result and integrated stream | Phase 3 / core | `AgentSession`, loop, `AgentEvent`, bounded subscriber | `AgentRunResult`; race-free `session.stream()` wrapper over one execution | core agent/session suite | agent/session and events docs |
43
+ | F-002 | Minimal scorers, sampling, datasets, batch experiments | Phase 4 / `@arnilo/prism-evals` (completed 2026-07-15) | F-001 result, run/trace IDs, package-local evaluation store | package-local scorer/dataset/experiment records and runner | eval package tests | `evaluations.md` |
44
+ | F-003 | Minimal project scaffold | Phase 5 / existing root CLI (completed 2026-07-15) | CLI parser, examples, provider packages, pack smoke | deterministic `prism init`; tiny templates | CLI temp-project pack/install/typecheck/test | CLI, README, release/install docs |
45
+ | F-004 | AI SDK model interoperability | Phase 6 / `@arnilo/prism-provider-ai-sdk` (completed 2026-07-15) | `AIProvider`, provider events, transport/content/structured-output conformance | one supported AI SDK language-model adapter | provider conformance with fake AI SDK model | provider package and AI SDK page |
46
+ | F-005 | Working memory and semantic recall | Phase 7 / `@arnilo/prism-memory` (completed 2026-07-15) | context providers, middleware, ownership, Postgres package | package-local `Embedder`, `VectorStore`, working-memory store; one pgvector adapter | memory conformance and PostgreSQL opt-in suite | `working-and-semantic-memory.md` |
47
+ | F-006 | Durable human approval and suspend/resume | Phase 8 / workflows plus coding-security bridge (completed 2026-07-15) | checkpoints, leases, fencing, workflow resume/cancel, execution policy | suspended state, validated payload, exact-once resume cursor | workflow restart/race/authorization tests | workflow and coding-security docs |
48
+ | F-007 | Small text/Markdown RAG | Phase 9 / `@arnilo/prism-rag` (completed 2026-07-16) | Phase 7 embed/vector contracts, resource/content bounds, context provider | deterministic chunk/index/retrieve/citation helpers | RAG package tests | `rag.md` |
49
+ | F-008 | Web-standard agent/workflow handler and MCP server | Phase 10 / `@arnilo/prism-server` plus existing MCP package (completed 2026-07-16) | F-001 result/stream, workflow commands, MCP SDK, abort/redaction | `Request -> Response` handler; explicit MCP server registration | server and MCP in-memory tests | `server.md`, `mcp-tools.md` |
50
+ | F-009 | Durable schedules and reconnectable background runs | Phase 11 / workflows (completed 2026-07-16) | coordinator, checkpoints, leases, active-run registry | schedule records/claims and one-time/interval/host-calculated next run | workflow multi-coordinator tests | workflow/persistence/server docs |
51
+ | F-010 | Workflow composition, state, and replay | Phase 11 / workflows (completed 2026-07-16) | DAG runner, node adapters, checkpoints, lineage/events | workflow node, bounded typed state, replay lineage/cursor | workflow nested/replay tests | workflow docs |
52
+ | F-011 | Trace/run feedback linked to evaluations | Phase 12 / persistence plus OTel/evals (completed 2026-07-16) | run/trace IDs, cursor queries, redaction, OTel | bounded feedback record/query and safe OTel projection | persistence conformance and OTel tests | observability/evaluation docs |
53
+ | F-012 | Supervisor delegation and A2A | Phase 13 / `@arnilo/prism-supervisor` (completed 2026-07-16) | F-001, memory scope, permission/abort/budget, web handler | bounded child delegation and current A2A protocol mapping/signing | supervisor isolation and A2A protocol tests | `supervisors.md`, `a2a.md` |
54
+
55
+ ## Existing primitive inventory
56
+
57
+ | Domain | Existing reusable primitive | What it already covers | Frozen gap/decision |
58
+ | --- | --- | --- | --- |
59
+ | Agent execution | `Agent`, `AgentSession`, `AgentLoopStrategy`, `singleShotLoop`, generate/validate/revise loop | provider/tool turns, retry, compaction, abort, stores, middleware, ledger | Add only F-001 result/stream in Phase 3; do not add a second engine or central application object. |
60
+ | Events and fan-in | `AgentEvent`, `createEventMultiplexer<T>()`, bounded `subscribe()` queues | normalized live events, ordered source fan-in, overflow policies | Reuse for server, workflow, and supervisor streams. Durable replay remains storage-owned. |
61
+ | Provider boundary | `AIProvider.generate()`, `ProviderEvent`, request policies, bounded SSE/body/argument transport, provider/media/openai helpers | first-party model streams, tool fragments, structured output, abort, metadata | AI SDK is an optional adapter only; no new provider abstraction. |
62
+ | Input/context | `InputBuilder`, `PromptBuilder`, `ContextProvider`, instruction injectors, resource loader, ten middleware hook names | explicit context/prompt assembly and host contributions | Memory and RAG inject through context/middleware; no memory-specific core hook. |
63
+ | Tools | `ToolDefinition`, validator, registry/filter, permission policy, `ExecutionPolicy`, bounded parallel dispatch | schema validation, allow/deny, ownership metadata, ordered tool transcript | Fix policy propagation/identity; durable approval extends workflow checkpoints rather than tool loop internals. |
64
+ | Redaction | `SecretRedactor`, message/event/request/session/ledger redactors | active-path cycle handling, object/Map key redaction, metadata-safe errors | Reuse at every new storage/remote boundary; no second redaction system. |
65
+ | Session/run persistence | `SessionStore`, `RunLedger`, `ProductionPersistenceStore`, cursor pages, ownership scope | sessions, branches, entries, runs, events, tool calls, usage, definitions, retention, migrations | Add narrowly typed records only when package-local stores cannot preserve cross-package run linkage. Vector search stays outside `SessionStore`. |
66
+ | Durable coordination | `CheckpointStore`, `LeaseStore`, fencing tokens; memory/SQLite/PostgreSQL implementations | versioned CAS state, atomic claims, expiry/takeover | Reuse for suspend, schedules, background runs, and replay. No new queue/lock engine. |
67
+ | Workflows | bounded DAG with agent/function/tool/conditional/fan-out/join nodes; checkpoint/resume/cancel/coordinator/RPC | deterministic local and multi-process orchestration | Extended for suspension, schedules/background runs, composition/state/replay without a durable-agent engine. |
68
+ | MCP | client bridge over SDK stdio and Streamable HTTP; list cache, timeout, abort, bounded results | consuming remote MCP tools | Add explicit server exposure in existing package; expose nothing by default. |
69
+ | CLI/RPC | `prism` CLI, LF-delimited RPC, `CommandDefinition`, workflow commands | host-controlled run and durable workflow operations | Add stdlib-only `init`; server remains optional package-owned. |
70
+ | Testing | mock provider plus provider/session/ledger/compaction/tool/extension/persistence conformance | network-free adapter and behavior checks | Add package-specific conformance only for genuinely reusable memory/eval/server seams. Avoid new source-text boundary tests. |
71
+ | Packaging | root plus 29 workspaces; 23 capability packages and six profile bundles | independent installation, dependency-ordered release, pack/import smoke | `prism-all` reaches all 30 packages; provider umbrella reaches all seven adapters; focused profiles stay unchanged. Core keeps zero runtime dependencies. |
72
+
73
+ ## Current package and export inventory
74
+
75
+ The publishable graph is root plus 28 workspaces (29 packages). Every package remains explicit; no package discovery exists.
76
+
77
+ | Export group | Current public surface |
78
+ | --- | --- |
79
+ | Root runtime | `@arnilo/prism` |
80
+ | Root provider helpers | `@arnilo/prism/providers/openai-compatible`, `/transport`, `/openai`, `/media` |
81
+ | Root testing helpers | `/testing/provider-conformance`, `/session-store-conformance`, `/compaction-conformance`, `/tool-conformance`, `/extension-conformance`, `/persistence-schema`, `/run-ledger-conformance` |
82
+ | Root Node helpers | `/node/config`, `/settings`, `/trust`, `/session-store-jsonl`, `/contribution-discovery`, `/instruction-injectors`, `/system-prompts`, `/agent-definitions` |
83
+ | Provider packages | `prism-provider-openai`, `-opencode-go`, `-openrouter`, `-zai`, `-kimi`, `-neuralwatt` |
84
+ | Compaction packages | `prism-compaction-llm`, `prism-compaction-observational-memory` |
85
+ | Optional feature packages | `prism-observability-opentelemetry`, `prism-tool-validator-json-schema`, `prism-mcp`, `prism-coding-agent`, `prism-coding-security`, `prism-session-store-sqlite`, `prism-session-store-postgres`, `prism-credentials-node`, `prism-workflows`, `prism-evals`, `prism-provider-ai-sdk`, `prism-memory`, `prism-rag`, `prism-server`, `prism-supervisor` |
86
+ | Profile bundles | `prism-all`, `prism-base`, `prism-code`, `prism-compaction`, `prism-providers`, `prism-sdk` |
87
+
88
+ Core middleware exposes exactly ten built-in hook names: `provider_request`, `input_assembly`, `prompt_build`, `context`, `tool_call`, `tool_result`, `retry`, `compaction`, `session_start`, and `session_shutdown`. Custom string hooks remain extension-owned. New phases reuse these hooks or package APIs unless the owning phase documents a generic cross-package gap.
89
+
90
+ ## Package and primitive decisions by phase
91
+
92
+ | Phase | Owner | Existing primitive consumed | Why it is insufficient | Minimum generic addition allowed |
93
+ | --- | --- | --- | --- | --- |
94
+ | 1 | core, coding-agent, coding-security | content policy, execution policy, approval policy | hostname normalization/DNS pinning and cache identity are incomplete | normalized/pinned network seam and execution run/session identity only |
95
+ | 2 | core, coding-security, OTel, providers | bash callback, runtime events, usage rows, media resolver | callbacks/terminal states/accounting/request aggregation are incomplete | sandbox output callback; usage scope/aggregate; request-level media resolver |
96
+ | 3 | core | session run, loop, events | no direct terminal result or integrated stream | `AgentRunResult` and one stream wrapper |
97
+ | 4 | optional eval package | Phase 3 result, IDs, persistence | no scorer/dataset record or bounded runner | package-local contracts; generic persistence addition only for durable cross-package linkage |
98
+ | 5 | root CLI | CLI and compile-checked examples | `prism init` + `templates/init` | templates and command only; no new library primitive |
99
+ | 6 | optional AI SDK provider | `AIProvider` and provider conformance | AI SDK models do not implement Prism's interface | adapter only; no core change |
100
+ | 7 | optional memory plus PostgreSQL adapter | context, ownership, middleware | no embedding/vector/working-memory contract | package-owned contracts and one production adapter |
101
+ | 8 | workflows/coding-security | checkpoint/resume/lease/policy | completed: suspended/denied state, validated/redacted resume, expected-version CAS, tool policy recheck | workflow status/checkpoint extension; no generic scheduler |
102
+ | 9 | optional RAG | Phase 7 and context/resource limits | completed: bounded text/Markdown chunk/index/filter/retrieve/citation/context helpers | package-only helpers |
103
+ | 10 | optional server and MCP | Web APIs, result/stream, workflow commands, MCP SDK | completed: bounded authorized agent/workflow Web handler plus explicit MCP tool/command registration and SDK Web transport | package-owned handler/router and MCP registrations |
104
+ | 11 | workflows/persistence | coordinator/checkpoint/lease/DAG | completed: ownership-scoped one-time/interval/calculated schedules, deterministic background enqueue, nested runner composition, bounded validated state/history, immutable replay lineage and fresh approval checks | package records over existing generic checkpoint/lease primitives; no migration or second worker engine |
105
+ | 12 | persistence/evals/OTel | run/trace IDs and cursor queries | completed: immutable exact-owned feedback append/query/delete, ID-only evaluation linkage, migration-003 SQLite/PostgreSQL stores, safe span projection and fixed-label counters | one bounded typed record/store plus optional OTel handlers; no vendor exporter |
106
+ | 13 | optional supervisor/A2A | session result, memory scope, policy, event multiplexer, Request/Response, WebCrypto | completed: explicit bounded child delegation, narrowing permissions/hooks, derived scopes, A2A 1.0 cards/ES256/JSON-RPC/SSE/exact-origin client | one zero-dependency optional package; text-only protocol subset |
107
+ | 14 | release/docs/tests | deterministic release and pack checks | completed: graph/docs/changelogs retargeted to 0.0.5; packed cross-capability and release matrix evidence recorded | no release-time source split; all umbrella inclusion preserves independent direct installs and inert activation |
108
+
109
+ ## Threat-boundary matrix
110
+
111
+ | Boundary | Untrusted input / asset | Primary risks | Required 0.0.5 controls | Owner |
112
+ | --- | --- | --- | --- | --- |
113
+ | Media network fetch | URL, hostname, DNS answer, response bytes/MIME | private-network access, rebinding, redirects, decompression/size abuse | normalize literals, classify all private ranges, resolve and pin public address or require allow-list/injected hardened loader, reject redirects, timeout/abort, per-item and whole-request bounds before provider side effects | Phases 1-2 |
114
+ | Tool execution | model-produced name/arguments/path/command | shell execution, traversal/symlink escape, policy bypass, output exhaustion | schema/filter/permission/execution policy on every aggregator, realpath containment, approval, sandbox, abort, bounded/redacted output | Phases 1-2 |
115
+ | Approval cache/resume | action and human decision | cross-run/session/tenant decision reuse, stale approval, duplicate side effect | default no cache, identity-bound keys, durable owner check, version/fencing/idempotency, policy recheck on resume | Phases 1 and 8 |
116
+ | Remote HTTP/MCP | request body, identity headers, selected capability, stream consumer | unauthorized capability use, DoS, data leak, orphaned runs | expose nothing by default, host auth callback, owner checks, body/event/concurrency/time bounds, disconnect abort, redacted errors | Phase 10 |
117
+ | Working/semantic memory and RAG | stored user text, embeddings, retrieved documents/metadata | cross-tenant recall, prompt injection, retention leak, unbounded context | mandatory tenant/resource/thread scope, bounded top-K/context, redaction, inert context, host retention/deletion | Phases 7 and 9 |
118
+ | Workflow suspend/schedule/replay | resume payload, checkpoint, schedule, source run | forged resume, stale worker, duplicate fire/side effect, approval replay | schema validation, authorization, CAS/fencing, idempotency, immutable lineage, bounded state/history | Phases 8 and 11 |
119
+ | Feedback/evaluation | comments, tags, scorer output, dataset items | PII/secret persistence, tenant leak, metric-cardinality explosion | ownership, redaction, limits, retention, no free text in metric labels, scorer isolation | Phases 4 and 12 |
120
+ | Supervisor delegation | child prompt/result, selected child/tool, delegated memory | privilege amplification, recursive/cyclic delegation, budget exhaustion, memory/credential leak | implemented: explicit allow-list, AND-only permission narrowing, unique resources, depth/concurrency/token/time/tool limits, cancellation propagation | Phase 13 complete |
121
+ | A2A | remote card, signature, task/message stream | impersonation, replay, malicious/oversized remote content, SSRF/credential forwarding | implemented: protocol-1.0 validation, host auth, ES256 signature/expiry verification, exact HTTPS origin allow-list/redirect rejection, bounds/timeouts/abort, untrusted output mapping | Phase 13 complete |
122
+
123
+ ## Capabilities already present — do not reimplement
124
+
125
+ | Existing capability | Evidence | 0.0.5 rule |
126
+ | --- | --- | --- |
127
+ | Revision ownership/redaction fix and valid multi-round tool transcript | `src/agent-loops.ts`, `src/redaction.ts`, regression tests | Preserve; Phase 3 reuses loop output. |
128
+ | Object and `Map` key redaction, active-path cycle handling | redaction suite | Use the same redactor everywhere. |
129
+ | Bounded provider SSE, error body, argument parsing, retry, native structured output | provider primitives and conformance | Adapters reuse these helpers/limits. |
130
+ | Polling OpenAI device OAuth | provider-openai OAuth tests | No new credential flow in core. |
131
+ | JSON Schema tool validation, bounded parallel calls, exclusive tools | core/tool-validator and loop tests | Preserve shared dispatch. |
132
+ | MCP client bridge | `@arnilo/prism-mcp` | Phase 10 adds server direction only. |
133
+ | SQLite/PostgreSQL production persistence, checkpoints, leases | adapter conformance/live CI | Extend only required records/migrations. |
134
+ | Encrypted file/keychain credentials | `@arnilo/prism-credentials-node` | Hosts keep explicit credential ownership. |
135
+ | Audio/file/document blocks and MIME/size checks | core/provider media tests | Fix SSRF/request aggregation; do not add a second media model. |
136
+ | OpenTelemetry agent/provider/tool instrumentation | optional OTel package | Correct lifecycle/accounting; no vendor matrix. |
137
+ | Bounded DAG workflows and distributed coordinator | workflows package | Extend the same engine. |
138
+ | Deterministic resumable package release | `scripts/release.mjs`, release workflow/tests | Retarget to 0.0.5 only after all gates. |
139
+
140
+ ## Explicit exclusions
141
+
142
+ 0.0.5 does not include Studio/editor/cloud services, browser automation, voice providers, chat-channel adapters, framework-specific HTTP adapters, auth-provider packages, deployment-provider packages, vendor observability exporters, a vector-store matrix, advanced document parsers/GraphRAG, an interactive TUI, automatic discovery/activation, or a mandatory central application object.
143
+
144
+ A cron-expression parser is also excluded: Phase 11 provides one-time timestamps, fixed intervals, and a host-supplied next-run calculator. Add a cron adapter only from a concrete requirement.
145
+
146
+ ## Frozen baseline summary
147
+
148
+ Detailed commands and measurements are in [Performance limits](performance.md#005-phase-0-baseline-2026-07-15).
149
+
150
+ | Area | Frozen value |
151
+ | --- | --- |
152
+ | Runtime/toolchain | Node 24.18.0 measurement host; npm 11.16.0; supported runtime remains Node >=20 |
153
+ | Full network-free tests | 1,475 total; 1,450 pass; 25 explicit live skips; 0 fail; 25.750 s |
154
+ | `sdk:ready` | pass; 54.341 s |
155
+ | Publishable graph | 24 packages at Phase 0 freeze; 25 after Phase 4 (evals); 26 after Phase 6 (AI SDK); 27 after Phase 7 (memory); 28 after Phase 9 (RAG); 29 after Phase 10 (server); 30 after Phase 13 (supervisor/A2A). Phase 14 follow-up puts all 30 behind `prism-all` and all seven adapters behind `prism-providers`. |
156
+ | Tarballs | 542,993 packed bytes / 2,084,900 unpacked bytes aggregate; root 346.0 kB / 1.3 MB |
157
+ | Installed workspace | 72 MiB `node_modules` |
158
+ | Source/tests | 189 production TypeScript files / 26,828 lines; 144 test files / 23,535 lines |
159
+ | Docs/plans/examples | 70 docs / 12,662 lines; 58 plans / 24,270 lines; 39 examples / 3,134 lines |
160
+ | Current generator | `prism init`; default sources ~3.3 KB / 8 files; default clean install ~27.5 MB vs Mastra 439 MB |
161
+ | Mastra comparator | default scaffold measured 439 MB `node_modules`, 300 MB build output, 427 installed packages |
162
+
163
+ ## Phase 0 verification
164
+
165
+ Executed from repository root at frozen HEAD:
166
+
167
+ ```bash
168
+ npm test
169
+ npm run sdk:ready
170
+ for i in 1 2 3 4 5; do node --test packages/workflows/dist/__tests__/coordinator.test.js; done
171
+ node /tmp/prism-phase0-bench.mjs
172
+ ```
173
+
174
+ Results:
175
+
176
+ - Pre-change baseline and post-documentation full checks passed with zero failures. Final `sdk:ready` completed in 57.764 s with all 24 dry-run packs; the frozen pre-change value remains 54.341 s.
177
+ - Coordinator suite passed five consecutive runs after the pre-freeze test correction.
178
+ - Traceability validation found 14 unique confirmed-finding IDs, 12 unique capability IDs, all roadmap phases 0-14, and no missing or duplicate owner.
179
+ - The benchmark script was temporary because these measurements are dated evidence, not CI wall-clock assertions.
180
+
181
+ ## Documentation reviewed
182
+
183
+ - Local: `docs/index.md`, `public-contracts.md`, `agent-session-runtime.md`, `agent-loops.md`, `runs-and-usage.md`, `workflows.md`, `workflow-orchestration-primitives.md`, `host-security.md`, `release-and-install.md`, `performance.md`, and `review-coverage-2026-07-14.md`.
184
+ - Source: core contracts/runtime/middleware/input/tools/content/redaction/persistence/checkpoint/lease/provider primitives, workflow package, optional package manifests/exports, and all seven testing conformance helpers.
185
+ - Comparison: Mastra repository commit `2745031d1d4a4978f037092da371428c32e2842a` and current docs/source reviewed on 2026-07-15 for agents, memory, RAG, workflows, evals, observability, server, MCP, scheduling, supervisors, and A2A.
186
+
187
+ ## Related pages
188
+
189
+ - [Roadmap](../roadmap.md): executable Phase 0-14 plan and release gate.
190
+ - [Review coverage — 2026-07-14](review-coverage-2026-07-14.md): evidence for capabilities already shipped before this review.
191
+ - [Performance limits](performance.md): dated baseline details and existing runtime limits.
192
+ - [Host security](host-security.md): current host responsibilities pending the Phase 1-2 corrections.
193
+ - [Release and install](release-and-install.md): current package graph and deterministic publication process.
@@ -0,0 +1,192 @@
1
+ # Review coverage — 2026-07-17 provider validation
2
+
3
+ Working evidence page for Plan 067. Freezes the 2026-07-14 P0–P2 re-verification map, first-party provider validation owners, official-doc priority URLs, Pi secondary references, cache/thinking/discovery surfaces, credential notes, and use-case model-binding inventory.
4
+
5
+ **Evidence frozen:** 2026-07-17 (offline inventory; no live provider calls).
6
+ **Priority rule:** official provider documentation wins; Pi (`badlogic/pi-mono`) is secondary when official docs are silent or ambiguous.
7
+
8
+ Related: [2026-07-14 coverage](review-coverage-2026-07-14.md) (0.0.4), [2026-07-15 coverage](review-coverage-2026-07-15.md) (0.0.5), [provider caching](provider-caching.md), [provider packages](provider-packages.md).
9
+
10
+ ## Status legend
11
+
12
+ | Status | Meaning |
13
+ | --- | --- |
14
+ | `verify` | Prior plan marked fixed; Plan 067 must re-verify with regression tests. |
15
+ | `gap` | Known missing or incorrect behavior/docs relative to official sources. |
16
+ | `fixed` | Closed in this plan (updated as tasks complete). |
17
+ | `by-design` | Intentional absence (document; do not “fix” into a catalog). |
18
+
19
+ ## P0–P2 finding → Plan 067 owner matrix
20
+
21
+ Source: `code-reviews/2026-07-14.md`. Prior Plans 053 / 054 / 058 marked these implemented for 0.0.4; this plan re-verifies rather than skipping.
22
+
23
+ | Review ID | Priority | Finding | 0.0.4 owner | Plan 067 task | Current status | Credential / redaction notes |
24
+ | --- | --- | --- | --- | --- | --- | --- |
25
+ | R-001 | P0 | Revision request duplicated + corrupted by redaction | 053-1 | 1 | fixed | Secret canaries must not appear in repair requests, events, or redacted graphs. Re-verified 2026-07-17: `pendingHistory` + active-path redaction; revision+redactor suite asserts one repair and no `[Circular]`. |
26
+ | R-002 | P1 | Multi-round tool transcript chronologically invalid | 053-2 | 1 | fixed | N/A. Re-verified 2026-07-17: two rounds × two calls keep `user → assistant → tool → tool → …` order in history and assembled request. |
27
+ | R-003 | P1 | Redactor leaks secrets in object/Map keys | 053-1 | 1 | fixed | Object/Map string keys redact; collisions use deterministic `__N` suffixes. |
28
+ | R-004 | P1 | Event-ledger writes lack backpressure | 053-3 | 1 | fixed | Ledger appends serialized (concurrency 1), order preserved, append failures reject run completion. |
29
+ | R-008 | P1 | Unbounded SSE / error bodies; multiline `data:` | 054-1/2 | 2 | fixed | Bounded readers; error text redacted. |
30
+ | R-009 | P1 | OpenAI device-code OAuth does not poll | 054-3 | 2 | fixed | Redact device/user/access/refresh codes from OAuth errors. |
31
+ | R-010 | P2 | Duplicated provider protocol utilities | 054-1/2 | 2 | fixed | Shared transport/primitives remain authoritative. |
32
+ | R-005 | P2 | JSONL append silent on corrupt lines | 053-4 | 2 | fixed | Dev-only store; fail closed on corrupt lines. |
33
+ | R-011 | P2 | Coding-agent image read unbounded / resize no-op | 055-5 | 2 | fixed | Enforce `maxImageBytes`; deprecate `autoResizeImages` honestly. |
34
+ | R-006 | P2 | Optional config ENOENT detected by message text | 053-5 | 2 | fixed | Typed `code === "ENOENT"`. |
35
+ | R-012 | P2 | Release tag/version mismatch; no provenance | 058-7/9 | 2 | fixed | Tag/version gate + `--provenance`; no secrets in artifacts. |
36
+
37
+ Bug-report fixes A–D remain covered by R-001 / R-003 / R-007 (malformed message shape was Plan 053; re-check under Task 1 if touched).
38
+
39
+ ## Provider package validation matrix
40
+
41
+ | Package | Plan 067 task | Official cache | Prism `cache.kind` | Thinking / reasoning (official → Prism) | Discovery endpoint | Static catalog (bootstrap only) | Pi secondary ref | Credential surface | Status |
42
+ | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
43
+ | `@arnilo/prism-provider-openai` | 6 | `prompt_cache_key`; older models `prompt_cache_retention` (`24h` / `in_memory`); GPT-5.6+ `prompt_cache_options` / breakpoints tracked (host may pass via compat/extra; discovery sets `longRetention: false`) | `openai_key` (+ `longRetention`, `maxKeyLength`) | Official: Responses top-level `reasoning.effort` (+ `summary`). Prism: merges `model.compat.reasoning` + `options.compat.reasoning` (request wins); `applyThinkingLevel(..., "openai_reasoning")` | Official: `GET /models` | Featured `openAIModels` / `openAICodexModels` + **`listOpenAIModels`**; factory accepts `models?` / `codexModels?` | `packages/ai/src/api/openai-responses.ts`, `openai-responses-shared.ts` | API key; Codex OAuth device-code (`oauth.ts`) — poll + redact codes | **fixed** 2026-07-17: Responses P0s + discovery + reasoning + models override |
44
+ | `@arnilo/prism-provider-kimi` | 7 | Anthropic-style `cache_control` on `/messages` when supported; OpenAI Moonshot route none | default implicit; opt-in `cache_control` | Official K2.x: `thinking.type` (`enabled`/`disabled`/`keep`); K2.7-code always enabled; K3: top-level `reasoning_effort`. Prism: `kimiThinking`/`kimiReasoningEffort`/`kimiPreserveThinking` (request wins); Coding Anthropic + Moonshot Chat Completions both callable | Official: `GET https://api.moonshot.ai/v1/models` (also `.cn`); Coding has no public list API | Featured Coding (`kimi-for-coding`, `kimi-for-coding-highspeed`, `k3`) + Moonshot (`kimi-k2.7-code`, `kimi-k3`) + **`listKimiModels`** | `kimi-coding*.ts`, `moonshotai*.ts`, `api/anthropic-messages.ts` | API key (`KIMI` / Moonshot, not interchangeable); redacted from errors | **fixed** 2026-07-17: discovery + Moonshot provider + thinking + official Coding ids |
45
+ | `@arnilo/prism-provider-zai` | 8 | Implicit (no client `cache_control` / `prompt_cache_key`); official `usage.prompt_tokens_details.cached_tokens` | `implicit` | Official: `thinking` (`{type, clear_thinking?}`), `reasoning_effort` (GLM-5.2+), `tool_stream` (GLM-4.6+). Prism: `zaiThinking`/`zaiReasoningEffort`/`zaiToolStream`/`zaiClearThinking`/`zaiPreserveThinking` (request wins); Preserved Thinking replays `reasoning_content`. Obsolete `thinkingFormat`/`developerRoleFallback` docs removed | No first-class docs.z.ai list page; OpenAI-compatible `GET {baseUrl}/models` best-effort + curated featured set from Chat Completions enum / overview | Featured `zaiModels` (`glm-5.2`…`glm-4.5`) + **`listZaiModels`**; default base `https://api.z.ai/api/paas/v4` | `zai.ts`, `zai.models.ts` (Pi secondary ids only) | API key; redacted from errors / discovery | **fixed** 2026-07-17: docs drift closed; catalog + discovery + clear_thinking/preserve |
46
+ | `@arnilo/prism-provider-openrouter` | 9 | Official prompt caching via `cache_control` + sticky `session_id` routing; top-level automatic when no breakpoints | `cache_control` (legacy `compat.openRouterCache`) | Official: `reasoning: { effort | max_tokens }`. Prism: `resolveOpenRouterReasoning` merge + `preserveThinking` replay as body `reasoning` | Official: `GET https://openrouter.ai/api/v1/models` | **App-controlled** `models:` + optional **`listOpenRouterModels`** (no bundled mega-catalog) | `openrouter.models.ts` (do not vendor), `api/openai-completions.ts` | API key; sanitize session/cache ids | **fixed** 2026-07-17: discovery + reasoning merge/preserve + automatic top-level cache_control |
47
+ | `@arnilo/prism-provider-opencode-go` | 10 | Anthropic route: selected `cache_control`; OpenAI route: none; `x-opencode-session` from cache/session key | route-specific (`cache_control` on Anthropic; `implicit` on OpenAI) | Dual-route: Anthropic thinking blocks + OpenAI `reasoning_content`; upstream `thinking`/`reasoning_effort`/`reasoning` passthrough (request wins); `preserveThinking` default for reasoning models | Official `GET https://opencode.ai/zen/go/v1/models` (sparse) | Featured official Go ids (Grok/GLM/Kimi/MiMo/MiniMax/Qwen/DeepSeek) + **`listOpenCodeGoModels`**; default base `https://opencode.ai/zen/go/v1` | `opencode-go.ts`, `opencode-go.models.ts` (Pi secondary ids/limits only) | API key; redacted from errors / discovery | **fixed** 2026-07-18: catalog + discovery + base URL + thinking preserve |
48
+ | `@arnilo/prism-provider-neuralwatt` | 11 | Implicit vLLM prefix caching; `prompt_tokens_details.cached_tokens` | `implicit` | Official: `reasoning_effort` when `capabilities.reasoning_effort` (GLM-5.2 default `max`); `thinking_token_budget`; `chat_template_kwargs` (`preserve_thinking`/`clear_thinking`/`enable_thinking`). Prism: `thinking.ts` resolves owned fields + `stripNeuralWattOwnedCompat`; `applyThinkingLevel(..., "reasoning_effort")`; Preserved Thinking replays `reasoning_content` | Official: `GET https://api.neuralwatt.com/v1/models` (auth optional for public models) | Featured `neuralWattModels` (official aliases incl. `gemma-4-31b`; no guessed pricing) + **`listNeuralWattModels()`** | **No Pi NeuralWatt provider** — official docs only | Optional API key on discovery; quota helper; redact secrets in error bodies | **fixed** 2026-07-18: catalog refresh + kwargs routing + owned-compat strip |
49
+ | `@arnilo/prism-provider-ai-sdk` | 12 | Host model owns request caching; adapter maps usage only | host-owned / N/A | Host `LanguageModelV4` owns reasoning; Prism maps stream `reasoning` parts | **None by design** (host supplies model) | None | N/A (AI SDK official spec > Pi) | No package credentials; host model may hold secrets | **fixed** 2026-07-18: host-owned catalog/cache/reasoning validated; usage mapping + docs |
50
+
51
+ ## Frozen official evidence sources (priority)
52
+
53
+ | Provider | Frozen URLs (official) | Notes frozen 2026-07-17 |
54
+ | --- | --- | --- |
55
+ | OpenAI | [List models](https://developers.openai.com/api/reference/resources/models/methods/list); [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching); [Reasoning](https://developers.openai.com/api/docs/guides/reasoning); [Responses create](https://developers.openai.com/api/reference/resources/responses/methods/create/); [Models guide](https://developers.openai.com/api/docs/models) | `GET /models`. Caching: `prompt_cache_key`; pre-5.6 `prompt_cache_retention`; 5.6+ `prompt_cache_options` / breakpoints. Reasoning: `reasoning.effort`. |
56
+ | Kimi / Moonshot | [List models](https://platform.kimi.ai/docs/api/list-models); [API overview](https://platform.kimi.ai/docs/api/overview); [Model parameter reference](https://platform.kimi.ai/docs/api/models-overview); [Thinking mode](https://platform.kimi.ai/docs/guide/use-kimi-k2-thinking-model); [Thinking effort](https://platform.kimi.ai/docs/guide/use-thinking-effort) | `GET /v1/models` on `api.moonshot.ai` / `.cn`. K2.x `thinking`; K3 `reasoning_effort: "max"`. Anthropic `/messages` compat remains under-documented — record empirical gaps in Task 7. |
57
+ | Z.AI | [Deep thinking](https://docs.z.ai/guides/capabilities/thinking); [Thinking mode](https://docs.z.ai/guides/capabilities/thinking-mode); [Tool streaming](https://docs.z.ai/guides/capabilities/stream-tool); [Context caching](https://docs.z.ai/guides/capabilities/cache); [Chat completion](https://docs.z.ai/api-reference/llm/chat-completion); [Migrate to GLM-5.2](https://docs.z.ai/guides/overview/migrate-to-glm-new); [Overview](https://docs.z.ai/guides/overview/overview) | Code + `docs/providers/zai.md` match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking`. Historical mismatch (`thinkingFormat`) closed in Task 8. |
58
+ | OpenRouter | [Get models](https://openrouter.ai/docs/api/api-reference/models/get-models); [Prompt caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching); [Reasoning tokens](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens) | `GET /api/v1/models`. Cache via `cache_control` + sticky routing. Reasoning via `reasoning` object. |
59
+ | OpenCode Go | [Go](https://opencode.ai/docs/go/); [Providers](https://opencode.ai/docs/providers/) | Dual OpenAI + Anthropic compatible APIs. Official `GET /zen/go/v1/models`. Featured: Grok 4.5, GLM-5.2/5.1, Kimi K3/K2.7 Code/K2.6, MiMo, MiniMax, Qwen3.7/3.6, DeepSeek V4. Default base `https://opencode.ai/zen/go/v1`. |
60
+ | NeuralWatt | [Models](https://portal.neuralwatt.com/docs/api/models); [Chat completions](https://portal.neuralwatt.com/docs/api/chat-completions); [API overview](https://portal.neuralwatt.com/docs/api/overview); [Quickstart](https://portal.neuralwatt.com/docs/quickstart) | `GET /v1/models` returns pricing/capabilities/limits metadata. Prefix caching automatic; `reasoning_effort` when capability flagged. |
61
+ | AI SDK | [Custom provider / LanguageModelV4](https://ai-sdk.dev/providers/community-providers/custom-providers); AI SDK usage types (`inputTokens` / cache details) | Adapter validates `specificationVersion`; maps cache read/write from usage. No Prism catalog. |
62
+
63
+ ## Frozen Pi secondary references
64
+
65
+ Repo: `https://github.com/badlogic/pi-mono` (`packages/ai`).
66
+
67
+ | Area | Path / note |
68
+ | --- | --- |
69
+ | OpenAI Responses | `packages/ai/src/providers/openai-responses.ts` (also Codex / Azure variants) |
70
+ | OpenAI Completions | `packages/ai/src/providers/openai-completions.ts` / `api/openai-completions.ts` |
71
+ | Anthropic Messages | `packages/ai/src/providers/anthropic-messages.ts` / `api/anthropic-messages.ts` |
72
+ | Generated catalogs | `packages/ai/src/providers/*.models.ts` via `scripts/generate-models.ts` — **not** copied as Prism’s sole strategy |
73
+ | Provider registry docs | Pi mintlify “LLM Providers” / README provider list |
74
+
75
+ Use Pi only to fill gaps or cross-check wire shapes after official docs.
76
+
77
+ ## Shared discovery pattern (Task 3 decision — frozen)
78
+
79
+ **Status (2026-07-17 Task 3):** pattern documented; no new core list-models primitive. Shared reuse is limited to existing transport/credential helpers.
80
+
81
+ ### Inventory
82
+
83
+ | Primitive | Location | Role |
84
+ | --- | --- | --- |
85
+ | `listNeuralWattModels({ apiKey?, fetch?, baseUrl?, signal?, headers? })` | `packages/provider-neuralwatt/src/models.ts` | **Template** — only shipped `list*Models` today |
86
+ | `mapNeuralWattModel(entry)` | same | Provider-specific entry → `ModelConfig` |
87
+ | `neuralWattModels` featured aliases | same | Offline bootstrap; no guessed pricing |
88
+ | `readBoundedResponseText` | `@arnilo/prism/providers/transport` | Bounded error-body read + optional secret redaction |
89
+ | `resolveCredentialValue` / `redactSecrets` | `@arnilo/prism` | Auth resolution + error redaction |
90
+ | Package `models?: readonly ModelConfig[]` | Kimi/Z.AI/OpenRouter/OpenCode Go/NeuralWatt/**OpenAI** factories | Host override of registered catalog |
91
+ | OpenAI factory `models?` / `listOpenAIModels` | **done** Task 6 | Fixed |
92
+ | OpenRouter / OpenCode Go `list*Models` | OpenRouter **`listOpenRouterModels` done** Task 9; OpenCode Go **`listOpenCodeGoModels` done** Task 10; Z.AI **`listZaiModels` done** Task 8; Kimi **`listKimiModels` done** Task 7 | Per-package work |
93
+ | AI SDK catalog | N/A | Host-owned `LanguageModelV4` — no discovery export |
94
+
95
+ ### Decisions
96
+
97
+ - Template options + return: NeuralWatt `listNeuralWattModels(...) → ModelConfig[]`.
98
+ - **Never** call discovery from `create*ProviderPackage()` / extension setup.
99
+ - Static catalogs = offline bootstrap / featured aliases only.
100
+ - OpenRouter stays app-registration-first; **`listOpenRouterModels`** (done Task 9) feeds `models:` only.
101
+ - AI SDK: no discovery export.
102
+ - Prefer **package-local** helpers. Do **not** add a core model-discovery registry or OpenAI-compatible mega-mapper in Task 3. Extract a shared HTTP/list helper later only if ≥2 packages share identical parsing (unlikely: OpenRouter/NeuralWatt/OpenAI response shapes and cache/cost mapping diverge).
103
+ - Discovery may populate `ModelConfig.cache` / `ModelConfig.cost` from live metadata when officially documented; static catalogs must not invent those fields.
104
+ - Docs: [Provider packages — Caller-gated model discovery](provider-packages.md#caller-gated-model-discovery), [Provider caching — Discovery and live cache/cost metadata](provider-caching.md#discovery-and-live-cache-cost-metadata), [Provider conformance — Model discovery checklist](provider-conformance.md#model-discovery-checklist).
105
+
106
+ ## Shared thinking / per-turn override (Task 4 decision — frozen)
107
+
108
+ | Layer | Contract |
109
+ | --- | --- |
110
+ | Model default | `ModelConfig.compat` (+ `capabilities.reasoning` where declared) |
111
+ | Per-turn override | `ProviderRequestOptions.compat` via existing `mergeProviderRequestOptions` (request wins) |
112
+ | Shared helpers | Core `applyThinkingLevel` / `thinkingCompatFor` / `thinkingFamilyForModel` → official compat fields; **not** a second options tree |
113
+ | Families | `openai_reasoning` (`reasoning.effort`), `reasoning_effort`, `thinking_type` (`thinking.type`), `noop` (host-owned) — only shapes shared by ≥2 packages (or explicit no-op) |
114
+ | Use-case wiring | LLM compaction + OM workers map `thinkingLevel` into `compat` via helpers (no longer inert `extra.thinkingLevel`) |
115
+ | Docs | [Thinking and reasoning](thinking-and-reasoning.md) |
116
+
117
+ **Decision (2026-07-17):** Implement thin core helpers; keep unique knobs (NeuralWatt budgets/kwargs, Kimi keep/all, Z.AI `tool_stream`) package-local. Core must not name forbidden provider literals (`openrouter`/`zai`/`kimi`/…). Hosts pick a family explicitly when inference is ambiguous. Per-provider first-class field hardening remains Tasks 6–12.
118
+
119
+ ## Use-case model binding inventory (Task 5 — done 2026-07-17)
120
+
121
+ | Site | Current binding | Session fallback today? | Thinking path today |
122
+ | --- | --- | --- | --- |
123
+ | `AgentSession.run` / `RunOptions.model` | Per-run override; writes `model_change` | Explicit override of session model | Host `providerOptions` / `applyThinkingLevel` |
124
+ | Observational memory workers | `resolveUseCaseModel({ configured: workerModel, sessionModel })` | **Yes** — host passes `sessionModel`; `requireExplicitModel` restores skip | `thinkingLevel` → `compat` via `applyThinkingLevel` |
125
+ | LLM compaction | `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` | **Yes** — `model` is the fallback slot | `thinkingLevel` → `compat` via `applyThinkingLevel` |
126
+ | Supervisor children | Child `Agent` owns its own `model` | Independent | Child config |
127
+ | Declarative agents | `resolveAgentDefinition` model string / `ModelConfig` | Definition-scoped | Definition / run options |
128
+ | Evaluations | `runOptions` including model | Eval-owned | Via run options |
129
+ | Workflows / RPC / CLI | Pass-through `runOptions.model` | Caller-owned | Via run options |
130
+ | Structured output | Reuses session/run model | Yes | Same as run |
131
+ | Memory / RAG embedders | Host `Embedder` — separate from chat LLM | N/A (not chat) | N/A |
132
+
133
+ **Decision (2026-07-17):** Core exports `UseCaseModelBinding`, `resolveUseCaseModel`, `resolveUseCaseModelBinding`, `useCaseCredentialProviderId`. Docs: [use-case-model-selection.md](use-case-model-selection.md). OM behavior change: session fallback when `sessionModel` supplied and worker model omitted; `requireExplicitModel` preserves historical `missing_model` skip.
134
+
135
+ Desired Plan 067 outcome: every use-case accepts `{ provider, model, thinking }` with **explicit session-model fallback** when no use-case default is set — **done** for OM + LLM compaction; other sites documented as already-separate.
136
+
137
+ ## Known doc / code mismatches (frozen)
138
+
139
+ | Item | Docs claim | Code / official | Owner task |
140
+ | --- | --- | --- | --- |
141
+ | Z.AI thinking | Was: `thinkingFormat: "zai"`, `developerRoleFallback` in `docs/providers/zai.md` | **fixed** Task 8 — docs+code match official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking` | 8 (done) |
142
+ | OpenAI reasoning | Was capabilities-only + opaque compat spread | Task 6 merges `model`/`options` `compat.reasoning` into body `reasoning`; Task 4 helper writes `compat.reasoning.effort` | 4 (done) / 6 (done) |
143
+ | Kimi Moonshot | Was metadata-only (`provider: "moonshot"` without provider) | **fixed** Task 7: `createMoonshotProvider` registered when `includeMoonshotModels`; official Coding ids + `listKimiModels` | 7 (done) |
144
+ | OpenCode Go catalog | Featured official Go open models + `listOpenCodeGoModels` | **done** Task 10 (removed stale `gpt-5.1-go` / `claude-sonnet-4.5-go`) | 10 (done) |
145
+ | Z.AI catalog | Was: `glm-4.7` / `glm-4.5` only | **fixed** Task 8 — featured GLM-5.2…4.5 + `listZaiModels` | 8 (done) |
146
+ | OM/LLM `thinkingLevel` + session fallback | Was inert `extra.thinkingLevel`; OM skipped without workerModel | Task 4 maps into `compat`; Task 5 session fallback + `requireExplicitModel` | 4 (done) / 5 (done) |
147
+
148
+ ## Credential and redaction canaries (per package)
149
+
150
+ | Package | Secrets in flight | Redaction canary expectation |
151
+ | --- | --- | --- |
152
+ | openai | API key; OAuth device/user/access/refresh | Absent from errors, events, requests, discovery failures |
153
+ | kimi | API key | Absent from errors / cache keys |
154
+ | zai | API key | Absent from errors / model metadata |
155
+ | openrouter | API key; session/cache ids | Sanitize length; never treat cache key as secret storage |
156
+ | opencode-go | API key | Absent from errors |
157
+ | neuralwatt | Optional API key; quota responses | Discovery/quota error bodies redacted |
158
+ | ai-sdk | Host-owned | Adapter must not echo host secrets in Prism errors |
159
+
160
+ No secrets are committed in this matrix.
161
+
162
+ ## Plan 067 task map (evidence owners)
163
+
164
+ | Task | Owns |
165
+ | --- | --- |
166
+ | 0 | This page + evidence freeze — **done** |
167
+ | 1 | R-001–R-004 re-verify — **fixed** 2026-07-17 |
168
+ | 2 | R-005–R-006, R-008–R-012 re-verify — **fixed** 2026-07-17 |
169
+ | 3 | Shared discovery pattern — **done** 2026-07-17 (package-local NeuralWatt template; no core list helper) |
170
+ | 4 | Shared thinking / per-turn surface — **done** 2026-07-17 (core helpers + OM/LLM compat wiring; see thinking-and-reasoning.md) |
171
+ | 5 | Use-case model selection + session fallback — **done** 2026-07-17 (`resolveUseCaseModel`, OM session fallback, use-case-model-selection.md) |
172
+ | 6 | OpenAI validate/harden — **done** 2026-07-17 (Responses P0s, `listOpenAIModels`, `models?`, reasoning merge) |
173
+ | 7 | Kimi validate/harden — **done** 2026-07-17 (`listKimiModels`, Moonshot provider, thinking, official Coding ids) |
174
+ | 8 | Z.AI validate/harden — **done** 2026-07-17 (`listZaiModels`, GLM-5.x catalog, docs drift closed, clear_thinking/preserve) |
175
+ | 9 | OpenRouter validate/harden — **done** 2026-07-17 (`listOpenRouterModels`, reasoning merge/preserve, automatic top-level `cache_control`) |
176
+ | 10 | OpenCode Go validate/harden — **done** 2026-07-18 (`listOpenCodeGoModels`, official Go catalog, `zen/go/v1` base, thinking preserve) |
177
+ | 11 | NeuralWatt validate/harden — **done** 2026-07-18 (featured catalog refresh, kwargs routing, owned-compat strip, applyThinkingLevel) |
178
+ | 12 | AI SDK validate — **done** 2026-07-18 (host-owned catalog/cache/reasoning; usage mapping + docs) |
179
+ | 13 | Cross-provider conformance + final verification — **done** 2026-07-18 (`sdk:ready`; 1,089 core tests + workspace suites + pack dry-runs) |
180
+
181
+ Exact task titles live in `plans/067-provider-doc-validation-caching-discovery-and-review-hardening.md`.
182
+
183
+ ## Verification for this page
184
+
185
+ - Final verification completed 2026-07-18 without live provider HTTP.
186
+ - Lists all seven first-party provider packages; every provider row is `fixed` or intentional by-design behavior.
187
+ - Lists every 2026-07-14 P0–P2 id (R-001–R-006, R-008–R-012); all are `fixed`.
188
+ - Distinguishes official-doc priority vs Pi secondary.
189
+ - `phase12-boundaries.test.ts` verifies all six HTTP package discovery exports and setup zero-fetch; AI SDK no-catalog behavior has its own adapter contract test.
190
+ - `provider_validation_final_contract_covers_all_adapters_and_binding_sites` verifies provider docs, cache kinds, thinking rows, use-case sites, matrix statuses, and index navigation.
191
+ - `npm run sdk:ready` passes: typecheck/build, 1,089 core tests, all workspace suites, packaging/provenance guards, and publish-graph pack dry-runs.
192
+ - Linked from `docs/index.md` under Release and install.
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
5
+ `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. `RunFeedbackStore` is the separate post-run seam for immutable ratings, comments, tags, and evaluation links. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
6
6
 
7
7
  APIs:
8
8
 
@@ -13,6 +13,7 @@ APIs:
13
13
  - `ToolCallRecord` / `ToolCallStatus`
14
14
  - `UsageRecord`
15
15
  - `redactRunLedgerRecord()`
16
+ - `RunFeedbackRecord` / `RunFeedbackStore` / `createMemoryRunFeedbackStore()`
16
17
 
17
18
  ## When to use it
18
19
 
@@ -37,7 +38,7 @@ Set the ledger and optional ownership scope/idempotency key on the agent or the
37
38
  | `appendRun` | `RunRecord` | After run starts (`running`) and again at finish (`succeeded`/`failed`/`aborted`). |
38
39
  | `appendEvent` | `AgentEventRecord` | After every emitted `AgentEvent`, after redaction. |
39
40
  | `appendToolCall` | `ToolCallRecord` | For each tool-call `started`, `progress`, `finished`, `error`, and `blocked` transition. |
40
- | `appendUsage` | `UsageRecord` | For each provider `usage` event and for the final loop usage. |
41
+ | `appendUsage` | `UsageRecord` | Once per terminal provider turn (`scope: "provider_turn"`) and once for the O(turns) aggregate (`scope: "run_total"`). |
41
42
 
42
43
  All methods may be sync or async (`void | Promise<void>`). The runtime awaits them at safe boundaries, so a slow adapter blocks the run.
43
44
 
@@ -92,9 +93,40 @@ The adapter receives these record shapes:
92
93
  | --- | --- |
93
94
  | `id` | Unique ledger row id. |
94
95
  | `runId` / `sessionId` / `entryId` | Correlation ids. |
96
+ | `scope` | `provider_turn` for billable source rows; `run_total` for the aggregate. Never sum both scopes. |
97
+ | `turn` / `attempt` | Provider-turn attribution; absent on `run_total`. |
95
98
  | `usage` | `Usage` shape: input/output/total/cache tokens, cost, currency. |
96
99
  | `recordedAt` | ISO timestamp. |
97
100
 
101
+ ## Run/trace feedback
102
+
103
+ `RunFeedbackStore.append()` accepts an immutable record only when `resolveRun` finds the same `runId` under the exact `{ tenantId, accountId?, userId? }` scope. A tenant plus account or user is mandatory. Records contain `sessionId`, optional `traceId`, finite `rating` in `[-1, 1]`, comment, tags, scorer IDs, evaluation IDs, timestamp, creator, and metadata. Correction appends a new ID; records are never updated in place. `delete()` is the explicit privacy/retention operation.
104
+
105
+ ```ts
106
+ import { createMemoryRunFeedbackStore } from "@arnilo/prism";
107
+
108
+ const feedback = createMemoryRunFeedbackStore({
109
+ resolveRun: ({ runId }) => runId === result.runId
110
+ ? { runId, sessionId: result.sessionId, tenantId: "t1", userId: "u1" }
111
+ : false,
112
+ redactor,
113
+ });
114
+ await feedback.append({
115
+ id: "fb_1",
116
+ runId: result.runId,
117
+ rating: 1,
118
+ comment: "Useful and cited",
119
+ tags: ["reviewed"],
120
+ evaluationIds: ["eval_1"],
121
+ tenantId: "t1",
122
+ userId: "u1",
123
+ });
124
+ const page = await feedback.query({ runId: result.runId, tenantId: "t1", userId: "u1", limit: 50 });
125
+ await feedback.delete({ id: "fb_1", tenantId: "t1", userId: "u1" });
126
+ ```
127
+
128
+ Default/hard bounds: comment 4/16 KiB, tags 16/64, scorer/evaluation IDs 16/64 each, metadata 16/64 KiB, query page 100/500; tags are 64 characters and identifiers 128. The store redacts comment/tags/metadata after run ownership validation and before persistence. IDs are linked, not scorer payloads. `ProductionPersistenceStore.feedback?` exposes this capability; first-party SQLite/PostgreSQL adapters implement it in schema migration `003_run_feedback` and reject missing/cross-owned runs.
129
+
98
130
  ## Status transitions
99
131
 
100
132
  ```
@@ -189,7 +221,9 @@ await session.run("Hello", {
189
221
  });
190
222
 
191
223
  console.log(runs.at(-1)?.status); // succeeded
192
- console.log(cacheUsageReport(usageRows.at(-1)?.usage));
224
+ const billable = usageRows.filter((row) => row.scope === "provider_turn");
225
+ const aggregate = usageRows.find((row) => row.scope === "run_total");
226
+ console.log(cacheUsageReport(aggregate?.usage));
193
227
  // { cacheReadTokens: 0, cacheWriteTokens: 0, ... } when provider usage is present
194
228
  ```
195
229
 
@@ -207,7 +241,8 @@ console.log(cacheUsageReport(usageRows.at(-1)?.usage));
207
241
  - `AgentConfig.ownership` is the default ownership scope; `RunOptions.ownership` overrides it per run.
208
242
  - `AgentConfig.idempotencyKey` is the default idempotency key; `RunOptions.idempotencyKey` overrides it per run.
209
243
  - The runtime resolves `model` and `provider` from `AgentConfig`/`RunOptions`/`AgentDefinition` before writing the start `RunRecord`.
210
- - Adapters should treat appends as ordered within a `runId`: event and tool-call rows preserve emission order because the runtime drains pending appends before writing the final `RunRecord`.
244
+ - Adapters should treat appends as ordered within a `runId`: event and tool-call rows preserve emission order because the runtime serializes event ledger appends through one promise chain (concurrency 1), drains pending appends before writing the final `RunRecord`, and propagates append failures by rejecting run completion.
245
+ - Billing queries must filter `scope = "provider_turn"`; presentation queries normally read the single `run_total`. `UsageQuery.scope`, `turn`, and `attempt` are explicit filters.
211
246
  - Adapters that need upsert semantics can use `RunRecord.id` (== `runId`) as the stable key.
212
247
  - Use `cacheUsageReport(record.usage, model)` for cache diagnostics from normalized usage. It works when a provider reports `cacheReadTokens` without `cacheWriteTokens`; missing write tokens are reported as `0`, and unavailable hit rate/savings stay `undefined`.
213
248
  - **Provider-specific telemetry is package-owned.** Core `Usage` carries token counts and `cost`/`currency`; it has no energy or detailed cost-breakdown fields. Providers that surface extra telemetry (e.g. `@arnilo/prism-provider-neuralwatt` exposes `neuralWattEventsWithTelemetry()`, `parseNeuralWattComment()`, and `mapNeuralWattTelemetry()` for `: energy`/`: cost` SSE comments and non-streaming top-level fields) keep that data in package-specific helpers/types. Telemetry never enters `RunLedger` usage rows unless the host explicitly copies it in; it carries usage/cost numbers only — never prompts, API keys, or headers. Account-level quota is likewise package-owned: `@arnilo/prism-provider-neuralwatt` exports an explicit `getNeuralWattQuota()` helper that the host calls on demand (never during generation); NeuralWatt rate-limits that endpoint to 1 request per second per customer, so the caller owns throttling.
@@ -219,9 +254,11 @@ console.log(cacheUsageReport(usageRows.at(-1)?.usage));
219
254
  - **Redaction.** The runtime calls `redactRunLedgerRecord()` and `redactAgentEvent()` with the active `SecretRedactor` before handing records to the adapter. `AgentEventRecord.redacted` and `ToolCallRecord.redacted` are set to `true` when a redactor is configured. Hosts should still redact before writing to durable storage if they perform additional transformations.
220
255
  - **Message content stays in `SessionStore`.** `AgentEventRecord.event` may contain `message_delta` / `message_finished` payloads; these are redacted but still belong conceptually to the session store. Do not use the ledger as the source of truth for messages.
221
256
  - **Cache diagnostics stay numeric.** `cacheUsageReport()` derives reports from `Usage` numbers and optional `ModelConfig.cost`; do not add prompt text, cache keys, headers, credentials, or provider payloads to usage rows.
257
+ - **No double billing.** Sum `provider_turn` rows or read `run_total`; never sum both. Run totals add every turn/attempt in O(turns), derive missing per-turn totals from input/output tokens, and omit aggregate cost when reported currencies conflict.
222
258
  - **Synchronous adapters block the run.** An adapter that performs network or heavy DB writes inline will slow down the agent loop. For high-throughput hosts, buffer or batch inside the adapter and return quickly; the runtime awaits the returned promise. If batching, preserve per-run order before acknowledging a batch: `appendEvent` rows should be pageable by `(runId, sequence)`, run rows by `(sessionId, startedAt, id)`, and usage rows by `(runId, recordedAt, id)`.
223
259
  - **Idempotency is host-owned.** The runtime writes the key into `RunRecord.idempotencyKey`; enforcing unique keys and deduplicating retries is the host adapter's responsibility.
224
- - **Tenant isolation.** `OwnershipScope` fields are copied from the active ownership scope, but the runtime does not enforce tenant isolation. Host adapters must apply their own access controls when querying persisted ledger rows.
260
+ - **Tenant isolation.** `OwnershipScope` fields are copied from the active ownership scope, but the runtime does not enforce tenant isolation for ledger rows. Feedback is stricter: append/query/delete require tenant plus account/user, and first-party stores compare the exact scope to the linked run.
261
+ - **Feedback privacy.** Comments/tags/metadata can contain PII. Configure a feedback redactor, apply retention, and call owned `delete()` for erasure. Never copy comments or tag values into metric labels.
225
262
 
226
263
  ## Related APIs
227
264