@arnilo/prism 0.0.4 → 0.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (70) hide show
  1. package/CHANGELOG.md +18 -0
  2. package/README.md +34 -10
  3. package/dist/agents.js +146 -19
  4. package/dist/cli-init.d.ts +41 -0
  5. package/dist/cli-init.js +390 -0
  6. package/dist/cli-runner.d.ts +7 -1
  7. package/dist/cli-runner.js +13 -1
  8. package/dist/content.d.ts +19 -0
  9. package/dist/content.js +197 -69
  10. package/dist/contracts.d.ts +94 -9
  11. package/dist/contracts.js +8 -0
  12. package/dist/feedback.d.ts +48 -0
  13. package/dist/feedback.js +230 -0
  14. package/dist/index.d.ts +6 -4
  15. package/dist/index.js +4 -3
  16. package/dist/providers/media.d.ts +3 -1
  17. package/dist/providers/media.js +11 -1
  18. package/dist/testing/feedback.d.ts +6 -0
  19. package/dist/testing/feedback.js +37 -0
  20. package/dist/testing/persistence-schema.d.ts +3 -3
  21. package/dist/testing/persistence-schema.js +32 -2
  22. package/dist/testing/run-ledger-conformance.js +7 -1
  23. package/docs/a2a.md +73 -0
  24. package/docs/agent-events.md +4 -6
  25. package/docs/agent-loops.md +1 -1
  26. package/docs/agent-session-runtime.md +14 -16
  27. package/docs/cli-rpc.md +35 -7
  28. package/docs/coding-agent-tools.md +2 -2
  29. package/docs/coding-security.md +7 -3
  30. package/docs/compaction-observational-memory.md +2 -0
  31. package/docs/context-and-skills.md +1 -0
  32. package/docs/credentials-and-redaction.md +2 -2
  33. package/docs/database-persistence.md +9 -6
  34. package/docs/evaluations.md +122 -0
  35. package/docs/extensions.md +2 -2
  36. package/docs/host-security.md +20 -3
  37. package/docs/index.md +29 -17
  38. package/docs/mcp-tools.md +49 -4
  39. package/docs/migration.md +33 -3
  40. package/docs/multimodal-content.md +14 -6
  41. package/docs/observability.md +14 -6
  42. package/docs/performance.md +209 -0
  43. package/docs/postgres-persistence.md +6 -4
  44. package/docs/provider-conformance.md +1 -0
  45. package/docs/provider-packages.md +2 -0
  46. package/docs/providers/ai-sdk.md +113 -0
  47. package/docs/public-contracts.md +6 -5
  48. package/docs/rag.md +113 -0
  49. package/docs/release-and-install.md +100 -77
  50. package/docs/review-coverage-2026-07-15.md +193 -0
  51. package/docs/runs-and-usage.md +41 -4
  52. package/docs/server.md +139 -0
  53. package/docs/settings-auth-trust-security.md +5 -5
  54. package/docs/sqlite-persistence.md +4 -3
  55. package/docs/supervisors.md +71 -0
  56. package/docs/workflow-orchestration-primitives.md +19 -3
  57. package/docs/workflows.md +97 -23
  58. package/docs/working-and-semantic-memory.md +169 -0
  59. package/package.json +12 -2
  60. package/templates/init/README.md.tmpl +28 -0
  61. package/templates/init/env.example.tmpl +1 -0
  62. package/templates/init/gitignore.tmpl +11 -0
  63. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  64. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  65. package/templates/init/package.json.tmpl +22 -0
  66. package/templates/init/providers.json +76 -0
  67. package/templates/init/src/agent.ts.tmpl +10 -0
  68. package/templates/init/src/index.ts.tmpl +12 -0
  69. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  70. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -0,0 +1,193 @@
1
+ # Review coverage — 2026-07-15
2
+
3
+ This page freezes Prism 0.0.5 scope at commit `f5128a816ae204c52f3e2f089de71c99bd5de6d4` after the Prism/Mastra review. It maps every confirmed finding and accepted capability to one owning roadmap phase, package/public surface, focused test owner, and documentation owner.
4
+
5
+ Status: **scope frozen; Phases 1-14 implementation complete; publication handoff pending**. Phase 0 changes no public runtime behavior.
6
+
7
+ ## Release decision
8
+
9
+ - Target release is **0.0.5**. Do not publish the currently versioned but unpublished 0.0.4 package graph.
10
+ - npm currently reports `@arnilo/prism@0.0.3`; Phase 14 has retargeted every repository package and lock entry to 0.0.5. Publication remains pending the clean signed-tag workflow.
11
+ - Core remains a Node.js 20-compatible, zero-runtime-dependency harness.
12
+ - New capabilities are opt-in and package-owned. No server, database, credential store, telemetry exporter, memory worker, schedule, remote agent, or privileged tool activates on install.
13
+ - Any row below that changes a trust boundary remains a release blocker until its owning phase and regression checks pass.
14
+
15
+ ## Frozen confirmed findings
16
+
17
+ Each ID appears once. “Planned” means owned, not fixed.
18
+
19
+ | ID | Confirmed finding | Disposition and single owner | Package / public surface | Focused test owner | Documentation owner | Status |
20
+ | --- | --- | --- | --- | --- | --- | --- |
21
+ | S-001 | Bracket-normalization lets IPv6 loopback, link-local, ULA, and IPv4-mapped private literals bypass media SSRF checks; DNS resolution/rebinding behavior is not sufficient for a private-network denial claim. | Phase 1 | Core `src/content.ts`; `SsrfPolicy` and media fetch path | `src/__tests__/content.test.ts` with injected resolver/requester matrix | `multimodal-content.md`, `host-security.md` | completed 2026-07-15 |
22
+ | S-002 | `createReadOnlyTools()` does not apply the aggregate `executionPolicy`. | Phase 1 | `@arnilo/prism-coding-agent` read-only aggregator | coding-agent execution-policy tests | `coding-agent-tools.md` | completed 2026-07-15 |
23
+ | S-003 | Approval caching defaults to a process-lifetime map instead of `none`; `run` and `session` scopes have no identity in the cache key. | Phase 1 | `@arnilo/prism-coding-security`; approval policy and execution identity metadata | coding-security approval tests | `coding-security.md`, `host-security.md` | completed 2026-07-15 |
24
+ | C-001 | Sandbox adapter returns only an exit code and cannot forward stdout/stderr to coding-agent `onData`. | Phase 2 | coding-security sandbox adapter / coding-agent bash operations bridge | sandbox and shell integration tests | `coding-security.md`, `coding-agent-tools.md` | completed 2026-07-15 |
25
+ | C-002 | OpenTelemetry agent spans end only on `agent_finished`; provider failure can leave `prism.agent.run` active. | Phase 2 | `@arnilo/prism-observability-opentelemetry` instrumentation lifecycle | in-memory telemetry failed/aborted-run tests | `observability.md` | completed 2026-07-15 |
26
+ | C-003 | `prism.provider.tokens` is incremented for provider-turn and agent-total events, double-counting one usage source. | Phase 2 | OTel token metric names/scopes | in-memory metric assertions | `observability.md`, `runs-and-usage.md` | completed 2026-07-15 |
27
+ | C-004 | Usage rows do not distinguish provider-turn values from run totals; the loop returns the latest turn rather than an explicit aggregate. | Phase 2 | core `Usage`, `UsageRecord`, runtime accumulator, persistence adapters | multi-turn run-ledger conformance and runtime tests | `runs-and-usage.md`, persistence docs | completed 2026-07-15 |
28
+ | C-005 | Providers resolve media one block at a time, so request-wide 32-item/32-MiB bounds are not accumulated before provider/upload side effects. | Phase 2 | core media request resolver and provider media serializers | content/provider-media/provider package tests | `multimodal-content.md` | completed 2026-07-15 |
29
+ | A-001 | Workflow coordinator concurrency test was timing-sensitive under load. | Phase 0 | test only; production workflow API unchanged | `packages/workflows/src/__tests__/coordinator.test.ts` repeated five times plus full suite | this page, `performance.md` | completed at frozen HEAD `f5128a8` |
30
+ | A-002 | `AgentConfig.extensions`, `settings`, and `credentials` are accepted but explicitly ignored by agent/session runtime. | Phase 3 | core `AgentConfig` and host composition docs | compile fixtures and runtime behavior tests | `agent-session-runtime.md`, `customization.md`, `migration.md` | completed 2026-07-15 |
31
+ | A-003 | `AgentSession.run()`/`prompt()` return `Promise<void>` and integrated streaming requires a separately started subscriber. | Phase 3 | core `AgentSession`, new run-result/stream contract | agent/session, CLI/RPC, workflow tests | agent/session, events, CLI/RPC, workflow docs | completed 2026-07-15 |
32
+ | A-004 | Direct source/phase-text boundary tests make refactors fail without behavioral changes. | Phase 14 | test architecture; replace implementation assertions when touched, retain manifest/docs/absence boundary checks | docs/export/behavior/pack tests replacing implementation-text assertions | testing/release coverage docs | completed for touched tests; broad historical conversion deferred |
33
+ | A-005 | `src/contracts.ts`, `src/agents.ts`, and `packages/workflows/src/run.ts` are conflict hotspots; docs/plans are larger than production source. | Phase 14 | maintainability gate; split only where completed work proves a cohesive domain | typecheck, public exports, behavior suites, docs links | review coverage and release docs | reviewed; no safe release-time cohesive split, deferred with measurements |
34
+ | R-001 | 0.0.4 must not be published after the review; all package/version/tag/provenance checks must target 0.0.5. | Phase 14 | 30-package graph, lockfile, release workflow/script | release, packaging, install, Node 20/24, registry checks | `release-and-install.md`, changelogs, migration docs | completed 2026-07-16; publication handoff pending |
35
+
36
+ ## Accepted capability scope
37
+
38
+ These are the twelve accepted Mastra-comparison recommendations. Each has one implementation owner; adjacent phases may consume the result but do not own a duplicate implementation.
39
+
40
+ | ID | Capability | Single owner | Reused Prism primitives | Minimum missing surface | Test owner | Documentation owner |
41
+ | --- | --- | --- | --- | --- | --- | --- |
42
+ | F-001 | Direct run result and integrated stream | Phase 3 / core | `AgentSession`, loop, `AgentEvent`, bounded subscriber | `AgentRunResult`; race-free `session.stream()` wrapper over one execution | core agent/session suite | agent/session and events docs |
43
+ | F-002 | Minimal scorers, sampling, datasets, batch experiments | Phase 4 / `@arnilo/prism-evals` (completed 2026-07-15) | F-001 result, run/trace IDs, package-local evaluation store | package-local scorer/dataset/experiment records and runner | eval package tests | `evaluations.md` |
44
+ | F-003 | Minimal project scaffold | Phase 5 / existing root CLI (completed 2026-07-15) | CLI parser, examples, provider packages, pack smoke | deterministic `prism init`; tiny templates | CLI temp-project pack/install/typecheck/test | CLI, README, release/install docs |
45
+ | F-004 | AI SDK model interoperability | Phase 6 / `@arnilo/prism-provider-ai-sdk` (completed 2026-07-15) | `AIProvider`, provider events, transport/content/structured-output conformance | one supported AI SDK language-model adapter | provider conformance with fake AI SDK model | provider package and AI SDK page |
46
+ | F-005 | Working memory and semantic recall | Phase 7 / `@arnilo/prism-memory` (completed 2026-07-15) | context providers, middleware, ownership, Postgres package | package-local `Embedder`, `VectorStore`, working-memory store; one pgvector adapter | memory conformance and PostgreSQL opt-in suite | `working-and-semantic-memory.md` |
47
+ | F-006 | Durable human approval and suspend/resume | Phase 8 / workflows plus coding-security bridge (completed 2026-07-15) | checkpoints, leases, fencing, workflow resume/cancel, execution policy | suspended state, validated payload, exact-once resume cursor | workflow restart/race/authorization tests | workflow and coding-security docs |
48
+ | F-007 | Small text/Markdown RAG | Phase 9 / `@arnilo/prism-rag` (completed 2026-07-16) | Phase 7 embed/vector contracts, resource/content bounds, context provider | deterministic chunk/index/retrieve/citation helpers | RAG package tests | `rag.md` |
49
+ | F-008 | Web-standard agent/workflow handler and MCP server | Phase 10 / `@arnilo/prism-server` plus existing MCP package (completed 2026-07-16) | F-001 result/stream, workflow commands, MCP SDK, abort/redaction | `Request -> Response` handler; explicit MCP server registration | server and MCP in-memory tests | `server.md`, `mcp-tools.md` |
50
+ | F-009 | Durable schedules and reconnectable background runs | Phase 11 / workflows (completed 2026-07-16) | coordinator, checkpoints, leases, active-run registry | schedule records/claims and one-time/interval/host-calculated next run | workflow multi-coordinator tests | workflow/persistence/server docs |
51
+ | F-010 | Workflow composition, state, and replay | Phase 11 / workflows (completed 2026-07-16) | DAG runner, node adapters, checkpoints, lineage/events | workflow node, bounded typed state, replay lineage/cursor | workflow nested/replay tests | workflow docs |
52
+ | F-011 | Trace/run feedback linked to evaluations | Phase 12 / persistence plus OTel/evals (completed 2026-07-16) | run/trace IDs, cursor queries, redaction, OTel | bounded feedback record/query and safe OTel projection | persistence conformance and OTel tests | observability/evaluation docs |
53
+ | F-012 | Supervisor delegation and A2A | Phase 13 / `@arnilo/prism-supervisor` (completed 2026-07-16) | F-001, memory scope, permission/abort/budget, web handler | bounded child delegation and current A2A protocol mapping/signing | supervisor isolation and A2A protocol tests | `supervisors.md`, `a2a.md` |
54
+
55
+ ## Existing primitive inventory
56
+
57
+ | Domain | Existing reusable primitive | What it already covers | Frozen gap/decision |
58
+ | --- | --- | --- | --- |
59
+ | Agent execution | `Agent`, `AgentSession`, `AgentLoopStrategy`, `singleShotLoop`, generate/validate/revise loop | provider/tool turns, retry, compaction, abort, stores, middleware, ledger | Add only F-001 result/stream in Phase 3; do not add a second engine or central application object. |
60
+ | Events and fan-in | `AgentEvent`, `createEventMultiplexer<T>()`, bounded `subscribe()` queues | normalized live events, ordered source fan-in, overflow policies | Reuse for server, workflow, and supervisor streams. Durable replay remains storage-owned. |
61
+ | Provider boundary | `AIProvider.generate()`, `ProviderEvent`, request policies, bounded SSE/body/argument transport, provider/media/openai helpers | first-party model streams, tool fragments, structured output, abort, metadata | AI SDK is an optional adapter only; no new provider abstraction. |
62
+ | Input/context | `InputBuilder`, `PromptBuilder`, `ContextProvider`, instruction injectors, resource loader, ten middleware hook names | explicit context/prompt assembly and host contributions | Memory and RAG inject through context/middleware; no memory-specific core hook. |
63
+ | Tools | `ToolDefinition`, validator, registry/filter, permission policy, `ExecutionPolicy`, bounded parallel dispatch | schema validation, allow/deny, ownership metadata, ordered tool transcript | Fix policy propagation/identity; durable approval extends workflow checkpoints rather than tool loop internals. |
64
+ | Redaction | `SecretRedactor`, message/event/request/session/ledger redactors | active-path cycle handling, object/Map key redaction, metadata-safe errors | Reuse at every new storage/remote boundary; no second redaction system. |
65
+ | Session/run persistence | `SessionStore`, `RunLedger`, `ProductionPersistenceStore`, cursor pages, ownership scope | sessions, branches, entries, runs, events, tool calls, usage, definitions, retention, migrations | Add narrowly typed records only when package-local stores cannot preserve cross-package run linkage. Vector search stays outside `SessionStore`. |
66
+ | Durable coordination | `CheckpointStore`, `LeaseStore`, fencing tokens; memory/SQLite/PostgreSQL implementations | versioned CAS state, atomic claims, expiry/takeover | Reuse for suspend, schedules, background runs, and replay. No new queue/lock engine. |
67
+ | Workflows | bounded DAG with agent/function/tool/conditional/fan-out/join nodes; checkpoint/resume/cancel/coordinator/RPC | deterministic local and multi-process orchestration | Extended for suspension, schedules/background runs, composition/state/replay without a durable-agent engine. |
68
+ | MCP | client bridge over SDK stdio and Streamable HTTP; list cache, timeout, abort, bounded results | consuming remote MCP tools | Add explicit server exposure in existing package; expose nothing by default. |
69
+ | CLI/RPC | `prism` CLI, LF-delimited RPC, `CommandDefinition`, workflow commands | host-controlled run and durable workflow operations | Add stdlib-only `init`; server remains optional package-owned. |
70
+ | Testing | mock provider plus provider/session/ledger/compaction/tool/extension/persistence conformance | network-free adapter and behavior checks | Add package-specific conformance only for genuinely reusable memory/eval/server seams. Avoid new source-text boundary tests. |
71
+ | Packaging | root plus 29 workspaces; 23 capability packages and six profile bundles | independent installation, dependency-ordered release, pack/import smoke | `prism-all` reaches all 30 packages; provider umbrella reaches all seven adapters; focused profiles stay unchanged. Core keeps zero runtime dependencies. |
72
+
73
+ ## Current package and export inventory
74
+
75
+ The publishable graph is root plus 28 workspaces (29 packages). Every package remains explicit; no package discovery exists.
76
+
77
+ | Export group | Current public surface |
78
+ | --- | --- |
79
+ | Root runtime | `@arnilo/prism` |
80
+ | Root provider helpers | `@arnilo/prism/providers/openai-compatible`, `/transport`, `/openai`, `/media` |
81
+ | Root testing helpers | `/testing/provider-conformance`, `/session-store-conformance`, `/compaction-conformance`, `/tool-conformance`, `/extension-conformance`, `/persistence-schema`, `/run-ledger-conformance` |
82
+ | Root Node helpers | `/node/config`, `/settings`, `/trust`, `/session-store-jsonl`, `/contribution-discovery`, `/instruction-injectors`, `/system-prompts`, `/agent-definitions` |
83
+ | Provider packages | `prism-provider-openai`, `-opencode-go`, `-openrouter`, `-zai`, `-kimi`, `-neuralwatt` |
84
+ | Compaction packages | `prism-compaction-llm`, `prism-compaction-observational-memory` |
85
+ | Optional feature packages | `prism-observability-opentelemetry`, `prism-tool-validator-json-schema`, `prism-mcp`, `prism-coding-agent`, `prism-coding-security`, `prism-session-store-sqlite`, `prism-session-store-postgres`, `prism-credentials-node`, `prism-workflows`, `prism-evals`, `prism-provider-ai-sdk`, `prism-memory`, `prism-rag`, `prism-server`, `prism-supervisor` |
86
+ | Profile bundles | `prism-all`, `prism-base`, `prism-code`, `prism-compaction`, `prism-providers`, `prism-sdk` |
87
+
88
+ Core middleware exposes exactly ten built-in hook names: `provider_request`, `input_assembly`, `prompt_build`, `context`, `tool_call`, `tool_result`, `retry`, `compaction`, `session_start`, and `session_shutdown`. Custom string hooks remain extension-owned. New phases reuse these hooks or package APIs unless the owning phase documents a generic cross-package gap.
89
+
90
+ ## Package and primitive decisions by phase
91
+
92
+ | Phase | Owner | Existing primitive consumed | Why it is insufficient | Minimum generic addition allowed |
93
+ | --- | --- | --- | --- | --- |
94
+ | 1 | core, coding-agent, coding-security | content policy, execution policy, approval policy | hostname normalization/DNS pinning and cache identity are incomplete | normalized/pinned network seam and execution run/session identity only |
95
+ | 2 | core, coding-security, OTel, providers | bash callback, runtime events, usage rows, media resolver | callbacks/terminal states/accounting/request aggregation are incomplete | sandbox output callback; usage scope/aggregate; request-level media resolver |
96
+ | 3 | core | session run, loop, events | no direct terminal result or integrated stream | `AgentRunResult` and one stream wrapper |
97
+ | 4 | optional eval package | Phase 3 result, IDs, persistence | no scorer/dataset record or bounded runner | package-local contracts; generic persistence addition only for durable cross-package linkage |
98
+ | 5 | root CLI | CLI and compile-checked examples | `prism init` + `templates/init` | templates and command only; no new library primitive |
99
+ | 6 | optional AI SDK provider | `AIProvider` and provider conformance | AI SDK models do not implement Prism's interface | adapter only; no core change |
100
+ | 7 | optional memory plus PostgreSQL adapter | context, ownership, middleware | no embedding/vector/working-memory contract | package-owned contracts and one production adapter |
101
+ | 8 | workflows/coding-security | checkpoint/resume/lease/policy | completed: suspended/denied state, validated/redacted resume, expected-version CAS, tool policy recheck | workflow status/checkpoint extension; no generic scheduler |
102
+ | 9 | optional RAG | Phase 7 and context/resource limits | completed: bounded text/Markdown chunk/index/filter/retrieve/citation/context helpers | package-only helpers |
103
+ | 10 | optional server and MCP | Web APIs, result/stream, workflow commands, MCP SDK | completed: bounded authorized agent/workflow Web handler plus explicit MCP tool/command registration and SDK Web transport | package-owned handler/router and MCP registrations |
104
+ | 11 | workflows/persistence | coordinator/checkpoint/lease/DAG | completed: ownership-scoped one-time/interval/calculated schedules, deterministic background enqueue, nested runner composition, bounded validated state/history, immutable replay lineage and fresh approval checks | package records over existing generic checkpoint/lease primitives; no migration or second worker engine |
105
+ | 12 | persistence/evals/OTel | run/trace IDs and cursor queries | completed: immutable exact-owned feedback append/query/delete, ID-only evaluation linkage, migration-003 SQLite/PostgreSQL stores, safe span projection and fixed-label counters | one bounded typed record/store plus optional OTel handlers; no vendor exporter |
106
+ | 13 | optional supervisor/A2A | session result, memory scope, policy, event multiplexer, Request/Response, WebCrypto | completed: explicit bounded child delegation, narrowing permissions/hooks, derived scopes, A2A 1.0 cards/ES256/JSON-RPC/SSE/exact-origin client | one zero-dependency optional package; text-only protocol subset |
107
+ | 14 | release/docs/tests | deterministic release and pack checks | completed: graph/docs/changelogs retargeted to 0.0.5; packed cross-capability and release matrix evidence recorded | no release-time source split; all umbrella inclusion preserves independent direct installs and inert activation |
108
+
109
+ ## Threat-boundary matrix
110
+
111
+ | Boundary | Untrusted input / asset | Primary risks | Required 0.0.5 controls | Owner |
112
+ | --- | --- | --- | --- | --- |
113
+ | Media network fetch | URL, hostname, DNS answer, response bytes/MIME | private-network access, rebinding, redirects, decompression/size abuse | normalize literals, classify all private ranges, resolve and pin public address or require allow-list/injected hardened loader, reject redirects, timeout/abort, per-item and whole-request bounds before provider side effects | Phases 1-2 |
114
+ | Tool execution | model-produced name/arguments/path/command | shell execution, traversal/symlink escape, policy bypass, output exhaustion | schema/filter/permission/execution policy on every aggregator, realpath containment, approval, sandbox, abort, bounded/redacted output | Phases 1-2 |
115
+ | Approval cache/resume | action and human decision | cross-run/session/tenant decision reuse, stale approval, duplicate side effect | default no cache, identity-bound keys, durable owner check, version/fencing/idempotency, policy recheck on resume | Phases 1 and 8 |
116
+ | Remote HTTP/MCP | request body, identity headers, selected capability, stream consumer | unauthorized capability use, DoS, data leak, orphaned runs | expose nothing by default, host auth callback, owner checks, body/event/concurrency/time bounds, disconnect abort, redacted errors | Phase 10 |
117
+ | Working/semantic memory and RAG | stored user text, embeddings, retrieved documents/metadata | cross-tenant recall, prompt injection, retention leak, unbounded context | mandatory tenant/resource/thread scope, bounded top-K/context, redaction, inert context, host retention/deletion | Phases 7 and 9 |
118
+ | Workflow suspend/schedule/replay | resume payload, checkpoint, schedule, source run | forged resume, stale worker, duplicate fire/side effect, approval replay | schema validation, authorization, CAS/fencing, idempotency, immutable lineage, bounded state/history | Phases 8 and 11 |
119
+ | Feedback/evaluation | comments, tags, scorer output, dataset items | PII/secret persistence, tenant leak, metric-cardinality explosion | ownership, redaction, limits, retention, no free text in metric labels, scorer isolation | Phases 4 and 12 |
120
+ | Supervisor delegation | child prompt/result, selected child/tool, delegated memory | privilege amplification, recursive/cyclic delegation, budget exhaustion, memory/credential leak | implemented: explicit allow-list, AND-only permission narrowing, unique resources, depth/concurrency/token/time/tool limits, cancellation propagation | Phase 13 complete |
121
+ | A2A | remote card, signature, task/message stream | impersonation, replay, malicious/oversized remote content, SSRF/credential forwarding | implemented: protocol-1.0 validation, host auth, ES256 signature/expiry verification, exact HTTPS origin allow-list/redirect rejection, bounds/timeouts/abort, untrusted output mapping | Phase 13 complete |
122
+
123
+ ## Capabilities already present — do not reimplement
124
+
125
+ | Existing capability | Evidence | 0.0.5 rule |
126
+ | --- | --- | --- |
127
+ | Revision ownership/redaction fix and valid multi-round tool transcript | `src/agent-loops.ts`, `src/redaction.ts`, regression tests | Preserve; Phase 3 reuses loop output. |
128
+ | Object and `Map` key redaction, active-path cycle handling | redaction suite | Use the same redactor everywhere. |
129
+ | Bounded provider SSE, error body, argument parsing, retry, native structured output | provider primitives and conformance | Adapters reuse these helpers/limits. |
130
+ | Polling OpenAI device OAuth | provider-openai OAuth tests | No new credential flow in core. |
131
+ | JSON Schema tool validation, bounded parallel calls, exclusive tools | core/tool-validator and loop tests | Preserve shared dispatch. |
132
+ | MCP client bridge | `@arnilo/prism-mcp` | Phase 10 adds server direction only. |
133
+ | SQLite/PostgreSQL production persistence, checkpoints, leases | adapter conformance/live CI | Extend only required records/migrations. |
134
+ | Encrypted file/keychain credentials | `@arnilo/prism-credentials-node` | Hosts keep explicit credential ownership. |
135
+ | Audio/file/document blocks and MIME/size checks | core/provider media tests | Fix SSRF/request aggregation; do not add a second media model. |
136
+ | OpenTelemetry agent/provider/tool instrumentation | optional OTel package | Correct lifecycle/accounting; no vendor matrix. |
137
+ | Bounded DAG workflows and distributed coordinator | workflows package | Extend the same engine. |
138
+ | Deterministic resumable package release | `scripts/release.mjs`, release workflow/tests | Retarget to 0.0.5 only after all gates. |
139
+
140
+ ## Explicit exclusions
141
+
142
+ 0.0.5 does not include Studio/editor/cloud services, browser automation, voice providers, chat-channel adapters, framework-specific HTTP adapters, auth-provider packages, deployment-provider packages, vendor observability exporters, a vector-store matrix, advanced document parsers/GraphRAG, an interactive TUI, automatic discovery/activation, or a mandatory central application object.
143
+
144
+ A cron-expression parser is also excluded: Phase 11 provides one-time timestamps, fixed intervals, and a host-supplied next-run calculator. Add a cron adapter only from a concrete requirement.
145
+
146
+ ## Frozen baseline summary
147
+
148
+ Detailed commands and measurements are in [Performance limits](performance.md#005-phase-0-baseline-2026-07-15).
149
+
150
+ | Area | Frozen value |
151
+ | --- | --- |
152
+ | Runtime/toolchain | Node 24.18.0 measurement host; npm 11.16.0; supported runtime remains Node >=20 |
153
+ | Full network-free tests | 1,475 total; 1,450 pass; 25 explicit live skips; 0 fail; 25.750 s |
154
+ | `sdk:ready` | pass; 54.341 s |
155
+ | Publishable graph | 24 packages at Phase 0 freeze; 25 after Phase 4 (evals); 26 after Phase 6 (AI SDK); 27 after Phase 7 (memory); 28 after Phase 9 (RAG); 29 after Phase 10 (server); 30 after Phase 13 (supervisor/A2A). Phase 14 follow-up puts all 30 behind `prism-all` and all seven adapters behind `prism-providers`. |
156
+ | Tarballs | 542,993 packed bytes / 2,084,900 unpacked bytes aggregate; root 346.0 kB / 1.3 MB |
157
+ | Installed workspace | 72 MiB `node_modules` |
158
+ | Source/tests | 189 production TypeScript files / 26,828 lines; 144 test files / 23,535 lines |
159
+ | Docs/plans/examples | 70 docs / 12,662 lines; 58 plans / 24,270 lines; 39 examples / 3,134 lines |
160
+ | Current generator | `prism init`; default sources ~3.3 KB / 8 files; default clean install ~27.5 MB vs Mastra 439 MB |
161
+ | Mastra comparator | default scaffold measured 439 MB `node_modules`, 300 MB build output, 427 installed packages |
162
+
163
+ ## Phase 0 verification
164
+
165
+ Executed from repository root at frozen HEAD:
166
+
167
+ ```bash
168
+ npm test
169
+ npm run sdk:ready
170
+ for i in 1 2 3 4 5; do node --test packages/workflows/dist/__tests__/coordinator.test.js; done
171
+ node /tmp/prism-phase0-bench.mjs
172
+ ```
173
+
174
+ Results:
175
+
176
+ - Pre-change baseline and post-documentation full checks passed with zero failures. Final `sdk:ready` completed in 57.764 s with all 24 dry-run packs; the frozen pre-change value remains 54.341 s.
177
+ - Coordinator suite passed five consecutive runs after the pre-freeze test correction.
178
+ - Traceability validation found 14 unique confirmed-finding IDs, 12 unique capability IDs, all roadmap phases 0-14, and no missing or duplicate owner.
179
+ - The benchmark script was temporary because these measurements are dated evidence, not CI wall-clock assertions.
180
+
181
+ ## Documentation reviewed
182
+
183
+ - Local: `docs/index.md`, `public-contracts.md`, `agent-session-runtime.md`, `agent-loops.md`, `runs-and-usage.md`, `workflows.md`, `workflow-orchestration-primitives.md`, `host-security.md`, `release-and-install.md`, `performance.md`, and `review-coverage-2026-07-14.md`.
184
+ - Source: core contracts/runtime/middleware/input/tools/content/redaction/persistence/checkpoint/lease/provider primitives, workflow package, optional package manifests/exports, and all seven testing conformance helpers.
185
+ - Comparison: Mastra repository commit `2745031d1d4a4978f037092da371428c32e2842a` and current docs/source reviewed on 2026-07-15 for agents, memory, RAG, workflows, evals, observability, server, MCP, scheduling, supervisors, and A2A.
186
+
187
+ ## Related pages
188
+
189
+ - [Roadmap](../roadmap.md): executable Phase 0-14 plan and release gate.
190
+ - [Review coverage — 2026-07-14](review-coverage-2026-07-14.md): evidence for capabilities already shipped before this review.
191
+ - [Performance limits](performance.md): dated baseline details and existing runtime limits.
192
+ - [Host security](host-security.md): current host responsibilities pending the Phase 1-2 corrections.
193
+ - [Release and install](release-and-install.md): current package graph and deterministic publication process.
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
5
+ `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. `RunFeedbackStore` is the separate post-run seam for immutable ratings, comments, tags, and evaluation links. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
6
6
 
7
7
  APIs:
8
8
 
@@ -13,6 +13,7 @@ APIs:
13
13
  - `ToolCallRecord` / `ToolCallStatus`
14
14
  - `UsageRecord`
15
15
  - `redactRunLedgerRecord()`
16
+ - `RunFeedbackRecord` / `RunFeedbackStore` / `createMemoryRunFeedbackStore()`
16
17
 
17
18
  ## When to use it
18
19
 
@@ -37,7 +38,7 @@ Set the ledger and optional ownership scope/idempotency key on the agent or the
37
38
  | `appendRun` | `RunRecord` | After run starts (`running`) and again at finish (`succeeded`/`failed`/`aborted`). |
38
39
  | `appendEvent` | `AgentEventRecord` | After every emitted `AgentEvent`, after redaction. |
39
40
  | `appendToolCall` | `ToolCallRecord` | For each tool-call `started`, `progress`, `finished`, `error`, and `blocked` transition. |
40
- | `appendUsage` | `UsageRecord` | For each provider `usage` event and for the final loop usage. |
41
+ | `appendUsage` | `UsageRecord` | Once per terminal provider turn (`scope: "provider_turn"`) and once for the O(turns) aggregate (`scope: "run_total"`). |
41
42
 
42
43
  All methods may be sync or async (`void | Promise<void>`). The runtime awaits them at safe boundaries, so a slow adapter blocks the run.
43
44
 
@@ -92,9 +93,40 @@ The adapter receives these record shapes:
92
93
  | --- | --- |
93
94
  | `id` | Unique ledger row id. |
94
95
  | `runId` / `sessionId` / `entryId` | Correlation ids. |
96
+ | `scope` | `provider_turn` for billable source rows; `run_total` for the aggregate. Never sum both scopes. |
97
+ | `turn` / `attempt` | Provider-turn attribution; absent on `run_total`. |
95
98
  | `usage` | `Usage` shape: input/output/total/cache tokens, cost, currency. |
96
99
  | `recordedAt` | ISO timestamp. |
97
100
 
101
+ ## Run/trace feedback
102
+
103
+ `RunFeedbackStore.append()` accepts an immutable record only when `resolveRun` finds the same `runId` under the exact `{ tenantId, accountId?, userId? }` scope. A tenant plus account or user is mandatory. Records contain `sessionId`, optional `traceId`, finite `rating` in `[-1, 1]`, comment, tags, scorer IDs, evaluation IDs, timestamp, creator, and metadata. Correction appends a new ID; records are never updated in place. `delete()` is the explicit privacy/retention operation.
104
+
105
+ ```ts
106
+ import { createMemoryRunFeedbackStore } from "@arnilo/prism";
107
+
108
+ const feedback = createMemoryRunFeedbackStore({
109
+ resolveRun: ({ runId }) => runId === result.runId
110
+ ? { runId, sessionId: result.sessionId, tenantId: "t1", userId: "u1" }
111
+ : false,
112
+ redactor,
113
+ });
114
+ await feedback.append({
115
+ id: "fb_1",
116
+ runId: result.runId,
117
+ rating: 1,
118
+ comment: "Useful and cited",
119
+ tags: ["reviewed"],
120
+ evaluationIds: ["eval_1"],
121
+ tenantId: "t1",
122
+ userId: "u1",
123
+ });
124
+ const page = await feedback.query({ runId: result.runId, tenantId: "t1", userId: "u1", limit: 50 });
125
+ await feedback.delete({ id: "fb_1", tenantId: "t1", userId: "u1" });
126
+ ```
127
+
128
+ Default/hard bounds: comment 4/16 KiB, tags 16/64, scorer/evaluation IDs 16/64 each, metadata 16/64 KiB, query page 100/500; tags are 64 characters and identifiers 128. The store redacts comment/tags/metadata after run ownership validation and before persistence. IDs are linked, not scorer payloads. `ProductionPersistenceStore.feedback?` exposes this capability; first-party SQLite/PostgreSQL adapters implement it in schema migration `003_run_feedback` and reject missing/cross-owned runs.
129
+
98
130
  ## Status transitions
99
131
 
100
132
  ```
@@ -189,7 +221,9 @@ await session.run("Hello", {
189
221
  });
190
222
 
191
223
  console.log(runs.at(-1)?.status); // succeeded
192
- console.log(cacheUsageReport(usageRows.at(-1)?.usage));
224
+ const billable = usageRows.filter((row) => row.scope === "provider_turn");
225
+ const aggregate = usageRows.find((row) => row.scope === "run_total");
226
+ console.log(cacheUsageReport(aggregate?.usage));
193
227
  // { cacheReadTokens: 0, cacheWriteTokens: 0, ... } when provider usage is present
194
228
  ```
195
229
 
@@ -208,6 +242,7 @@ console.log(cacheUsageReport(usageRows.at(-1)?.usage));
208
242
  - `AgentConfig.idempotencyKey` is the default idempotency key; `RunOptions.idempotencyKey` overrides it per run.
209
243
  - The runtime resolves `model` and `provider` from `AgentConfig`/`RunOptions`/`AgentDefinition` before writing the start `RunRecord`.
210
244
  - Adapters should treat appends as ordered within a `runId`: event and tool-call rows preserve emission order because the runtime drains pending appends before writing the final `RunRecord`.
245
+ - Billing queries must filter `scope = "provider_turn"`; presentation queries normally read the single `run_total`. `UsageQuery.scope`, `turn`, and `attempt` are explicit filters.
211
246
  - Adapters that need upsert semantics can use `RunRecord.id` (== `runId`) as the stable key.
212
247
  - Use `cacheUsageReport(record.usage, model)` for cache diagnostics from normalized usage. It works when a provider reports `cacheReadTokens` without `cacheWriteTokens`; missing write tokens are reported as `0`, and unavailable hit rate/savings stay `undefined`.
213
248
  - **Provider-specific telemetry is package-owned.** Core `Usage` carries token counts and `cost`/`currency`; it has no energy or detailed cost-breakdown fields. Providers that surface extra telemetry (e.g. `@arnilo/prism-provider-neuralwatt` exposes `neuralWattEventsWithTelemetry()`, `parseNeuralWattComment()`, and `mapNeuralWattTelemetry()` for `: energy`/`: cost` SSE comments and non-streaming top-level fields) keep that data in package-specific helpers/types. Telemetry never enters `RunLedger` usage rows unless the host explicitly copies it in; it carries usage/cost numbers only — never prompts, API keys, or headers. Account-level quota is likewise package-owned: `@arnilo/prism-provider-neuralwatt` exports an explicit `getNeuralWattQuota()` helper that the host calls on demand (never during generation); NeuralWatt rate-limits that endpoint to 1 request per second per customer, so the caller owns throttling.
@@ -219,9 +254,11 @@ console.log(cacheUsageReport(usageRows.at(-1)?.usage));
219
254
  - **Redaction.** The runtime calls `redactRunLedgerRecord()` and `redactAgentEvent()` with the active `SecretRedactor` before handing records to the adapter. `AgentEventRecord.redacted` and `ToolCallRecord.redacted` are set to `true` when a redactor is configured. Hosts should still redact before writing to durable storage if they perform additional transformations.
220
255
  - **Message content stays in `SessionStore`.** `AgentEventRecord.event` may contain `message_delta` / `message_finished` payloads; these are redacted but still belong conceptually to the session store. Do not use the ledger as the source of truth for messages.
221
256
  - **Cache diagnostics stay numeric.** `cacheUsageReport()` derives reports from `Usage` numbers and optional `ModelConfig.cost`; do not add prompt text, cache keys, headers, credentials, or provider payloads to usage rows.
257
+ - **No double billing.** Sum `provider_turn` rows or read `run_total`; never sum both. Run totals add every turn/attempt in O(turns), derive missing per-turn totals from input/output tokens, and omit aggregate cost when reported currencies conflict.
222
258
  - **Synchronous adapters block the run.** An adapter that performs network or heavy DB writes inline will slow down the agent loop. For high-throughput hosts, buffer or batch inside the adapter and return quickly; the runtime awaits the returned promise. If batching, preserve per-run order before acknowledging a batch: `appendEvent` rows should be pageable by `(runId, sequence)`, run rows by `(sessionId, startedAt, id)`, and usage rows by `(runId, recordedAt, id)`.
223
259
  - **Idempotency is host-owned.** The runtime writes the key into `RunRecord.idempotencyKey`; enforcing unique keys and deduplicating retries is the host adapter's responsibility.
224
- - **Tenant isolation.** `OwnershipScope` fields are copied from the active ownership scope, but the runtime does not enforce tenant isolation. Host adapters must apply their own access controls when querying persisted ledger rows.
260
+ - **Tenant isolation.** `OwnershipScope` fields are copied from the active ownership scope, but the runtime does not enforce tenant isolation for ledger rows. Feedback is stricter: append/query/delete require tenant plus account/user, and first-party stores compare the exact scope to the linked run.
261
+ - **Feedback privacy.** Comments/tags/metadata can contain PII. Configure a feedback redactor, apply retention, and call owned `delete()` for erasure. Never copy comments or tag values into metric labels.
225
262
 
226
263
  ## Related APIs
227
264
 
package/docs/server.md ADDED
@@ -0,0 +1,139 @@
1
+ # Web-standard server handler
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-server` exposes explicitly selected agents and workflows through one framework-free `(Request) => Promise<Response>` handler. It supports direct agent results, bounded agent/workflow SSE, durable workflow start/enqueue/status/cancel/resume/replay, ownership-scoped schedules, host authorization, ownership propagation, redaction, and resource ceilings.
6
+
7
+ No listener starts on import. Empty `agents`/`workflows` maps expose nothing. Authentication, authorization, route selection, durable stores, TLS, rate limiting, and framework/serverless adaptation remain host-owned.
8
+
9
+ ## When to use it
10
+
11
+ Use it when a Node 20, serverless, worker, or framework host already speaks Web `Request`/`Response` and needs a small Prism API boundary. Wrap it in the platform's native adapter rather than adding Express, Fastify, Hono, Koa, Nest, or Next to Prism.
12
+
13
+ Use `AgentSession` or workflow APIs directly for in-process applications. Do not treat this package as an auth provider, user database, firewall, durable agent-result store, or public listener.
14
+
15
+ ## Inputs / request
16
+
17
+ ```ts
18
+ const handler = createPrismHandler({
19
+ agents?: Record<string, Agent | PrismAgentExposure>,
20
+ workflows?: Record<string, PrismWorkflowExposure>,
21
+ schedules?: WorkflowSchedules | ((authorization, signal) => WorkflowSchedules),
22
+ authorize: async ({ request, operation, capabilityId }) => false | {
23
+ ownership: { tenantId?: string; accountId?: string; userId?: string },
24
+ metadata?: Record<string, unknown>,
25
+ },
26
+ basePath?: "/prism",
27
+ allowedHosts?: string[],
28
+ allowedOrigins?: string[],
29
+ redactor?: SecretRedactor,
30
+ limits?: PrismServerLimits,
31
+ disconnectAborts?: boolean,
32
+ });
33
+ ```
34
+
35
+ At least one non-empty ownership field must come from `authorize()`. Request JSON never chooses ownership.
36
+
37
+ | Method and route | Authorization operation | Body |
38
+ | --- | --- | --- |
39
+ | `POST /prism/agents/:id/runs` | `agent.run` | `{ "input": string | Message | Message[] }` |
40
+ | `POST /prism/agents/:id/stream` | `agent.stream` | same; SSE response |
41
+ | `POST /prism/workflows/:id/runs` | `workflow.run` | `{ "input": unknown, "runId"?: string }` |
42
+ | `POST /prism/workflows/:id/stream` | `workflow.stream` | same; SSE response |
43
+ | `POST /prism/workflows/:id/enqueue` | `workflow.enqueue` | `{ "input": unknown, "runId"?: string }`; returns `202` queued handle |
44
+ | `GET /prism/workflows/:id/runs/:runId` | `workflow.status` | none |
45
+ | `DELETE /prism/workflows/:id/runs/:runId` | `workflow.cancel` | none |
46
+ | `POST /prism/workflows/:id/runs/:runId/resume` | `workflow.resume` | `{ "decision": "approve" | "deny", "input"?: unknown, "expectedVersion": number }` |
47
+ | `POST /prism/workflows/:id/runs/:runId/replay` | `workflow.replay` | `{ "fromNodeId": string, "runId"?: string }` |
48
+ | `POST /prism/schedules/:id` | `schedule.create` | `{ "workflowId", "nextRunAt", "input"?, "intervalMs"?, "calculatorId"?, "paused"?, "metadata"? }` |
49
+ | `GET /prism/schedules?status=&cursor=&limit=` | `schedule.list` | none |
50
+ | `POST /prism/schedules/:id/pause` | `schedule.pause` | `{}` |
51
+ | `POST /prism/schedules/:id/resume` | `schedule.resume` | `{ "nextRunAt"?: string }` |
52
+ | `POST /prism/schedules/:id/trigger` | `schedule.trigger` | `{ "idempotencyKey": string }` |
53
+ | `DELETE /prism/schedules/:id` | `schedule.delete` | none |
54
+
55
+ POST routes require `Content-Type: application/json`. Capability/run IDs are bounded URL-safe identifiers. A custom `PrismAgentExposure.sessionFactory` can build sessions from authorized host context; otherwise an `Agent` creates a fresh session.
56
+
57
+ ## Outputs / response / events
58
+
59
+ Direct routes return bounded JSON. Stream routes return `text/event-stream`; every event is one `data: <AgentEvent|WorkflowEvent>` frame. Status returns the ownership-scoped durable checkpoint record. Resume uses Phase 8 expected-version CAS. Cancel aborts active work or marks eligible durable checkpoints aborted.
60
+
61
+ Errors use `{ "error": { "code", "message" } }`. Unknown routes/capabilities are `404`, authorization/policy denial `403`, malformed input `400`, unsupported content type `415`, body overflow `413`, concurrency overflow `429`, and result overflow `507`. Unexpected errors are generic and never include stacks.
62
+
63
+ ## Request/response example
64
+
65
+ ```json
66
+ {
67
+ "request": { "method": "POST", "path": "/prism/agents/support/runs", "body": { "input": "Summarize this" } },
68
+ "response": { "status": "succeeded", "sessionId": "...", "runId": "...", "text": "Summary" }
69
+ }
70
+ ```
71
+
72
+ ## Implementation example
73
+
74
+ ```ts
75
+ import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
76
+ import { createPrismHandler } from "@arnilo/prism-server";
77
+
78
+ const agent = createAgent({
79
+ model: { provider: "mock", model: "offline" },
80
+ provider: createMockProvider([providerTextDelta("ready"), providerDone()]),
81
+ });
82
+
83
+ const handler = createPrismHandler({
84
+ agents: { support: agent },
85
+ authorize: async ({ request }) => request.headers.get("authorization") === "Bearer host-validated"
86
+ ? { ownership: { tenantId: "tenant-1", userId: "user-1" } }
87
+ : false,
88
+ allowedHosts: ["api.example.test"],
89
+ allowedOrigins: ["https://app.example.test"],
90
+ });
91
+
92
+ // Cloudflare/Bun/Deno-style: export { handler as fetch }.
93
+ // Node/framework hosts adapt their request to Web Request and return Web Response.
94
+ ```
95
+
96
+ ## Extension and configuration notes
97
+
98
+ - `basePath` defaults to `/prism`; URL root exposure is rejected.
99
+ - Agent maps and workflow maps are immutable host selections. No registry/package discovery runs.
100
+ - Workflow exposure requires its existing `WorkflowCheckpointAdapter`; no server-owned database exists.
101
+ - Schedule exposure is optional and may be one service or an authorization-selected resolver. Returned service ownership must exactly match authorized tenant/account/user scope; otherwise request is forbidden.
102
+ - `PrismWorkflowExposure.runOptions` can supply agent/tool/policy/resume-validator wiring. Server-owned ownership, signal, checkpoint, redactor, run ID, and event bus fields cannot be overridden.
103
+ - Host/origin checks and CORS headers activate only when their allow-lists are configured. Hosts still own reverse-proxy trust and canonical host handling.
104
+
105
+ Default/hard ceilings:
106
+
107
+ | Limit | Default | Hard cap |
108
+ | --- | ---: | ---: |
109
+ | JSON request | 64 KiB | 1 MiB |
110
+ | direct response | 1 MiB | 8 MiB |
111
+ | SSE event | 64 KiB | 1 MiB |
112
+ | SSE total | 10 MiB | 64 MiB |
113
+ | SSE event count | 10,000 | 100,000 |
114
+ | concurrent runs | 16 | 256 |
115
+ | subscriber queue | 128 | 4,096 |
116
+ | request/run timeout | 120 s | 30 min |
117
+
118
+ ## Security and performance notes
119
+
120
+ - `authorize()` is required and runs for every matched operation before capability lookup or body execution. Return `false` on missing/invalid credentials. Do not trust caller ownership fields.
121
+ - Use authorization metadata only for non-secret audit context. Never put credentials in metadata, input, route IDs, run IDs, checkpoints, events, or responses.
122
+ - Configure `SecretRedactor` before runs. Redaction matches known secrets; it is not DLP.
123
+ - Agent tools and workflow tool nodes still need their own `PermissionPolicy`, `ToolValidator`, and `ExecutionPolicy`. HTTP authorization does not replace side-effect policy.
124
+ - Host and origin allow-lists are exact string matches. Configure reverse-proxy normalization, TLS, rate limiting, IP policy, CSRF/cookie policy, and authentication outside Prism.
125
+ - SSE uses bounded upstream subscriber queues. Consumer cancellation aborts owned work by default and releases concurrency; set `disconnectAborts: false` only when the host deliberately owns background completion.
126
+ - Source inputs/resource URLs remain host responsibilities and use existing resource/media SSRF policies. Server package does not fetch URLs.
127
+ - Schedule routes never accept ownership from JSON. Services carry mandatory ownership and explicit workflow/calculator registries; route authorization cannot broaden either. Replay applies workflow ownership/hash/approval checks.
128
+ - No agent status/reconnect store is invented. Durable reconnect/status/resume is the workflow path; persistent agent run querying remains a host persistence API.
129
+
130
+ A2A routes are not added to `createPrismHandler()`. Install `@arnilo/prism-supervisor` and explicitly mount `createA2AHandler()` when protocol interoperability is required; this keeps cards and remote invoke absent from ordinary Prism servers.
131
+
132
+ ## Related APIs
133
+
134
+ - [Agent/session runtime](agent-session-runtime.md): direct result and event stream semantics.
135
+ - [Workflows](workflows.md): durable checkpoints, status, cancellation, and exact-once resume.
136
+ - [MCP client and server exposure](mcp-tools.md): selected MCP capabilities and web-standard MCP transport.
137
+ - [Host security guide](host-security.md): remote-boundary checklist.
138
+ - [A2A interoperability](a2a.md): separately mounted A2A 1.0 handler/client.
139
+ - [Release and install](release-and-install.md): optional package installation and profiles.
@@ -18,7 +18,7 @@ Use these APIs when a host wants one explicit place to compose settings, resolve
18
18
  - Node-only subpaths: `@arnilo/prism/node/settings` for caller-named JSON settings files and `@arnilo/prism/node/trust` for explicit trusted path roots with symlink-aware realpath checks.
19
19
 
20
20
  ## Outputs / response / events
21
- Settings and credential helpers return existing `SettingsProvider` and `CredentialResolver` contracts. `AgentConfig.settings` and `AgentConfig.credentials` are host-owned metadata for compatibility; `createAgent()` / `session.run()` do not call `settings.get()` or `credentials.resolve()`. Permission denial blocks tool execution, extension setup, and resource loader calls before side effects. A configured `AgentConfig.redactor` or `RunOptions.redactor` redacts provider requests, emitted `AgentEvent` payloads, stored `SessionEntry` values, and runtime `InstructionContext` input/history seen by instruction injectors.
21
+ Settings and credential helpers return existing `SettingsProvider` and `CredentialResolver` contracts. These seams are host-owned outside `AgentConfig`; `createAgent()` / `session.run()` do not call `settings.get()` or `credentials.resolve()`. Permission denial blocks tool execution, extension setup, and resource loader calls before side effects. A configured `AgentConfig.redactor` or `RunOptions.redactor` redacts provider requests, emitted `AgentEvent` payloads, stored `SessionEntry` values, and runtime `InstructionContext` input/history seen by instruction injectors.
22
22
 
23
23
  ## Request/response example
24
24
  ```ts
@@ -47,17 +47,17 @@ const apiKey = await resolveCredentialValue(credentials, { name: "api", provider
47
47
 
48
48
  const agent = createAgent({
49
49
  model: { provider: "demo", model: "model" },
50
- // host-owned metadata; runtime does not read/resolve these fields
51
- settings,
52
- credentials,
50
+ // Resolve credentials at the provider edge; register known secrets for redaction.
53
51
  redactor: apiKey ? createSecretRedactor([apiKey]) : undefined,
54
52
  });
53
+ void settings;
54
+ void credentials;
55
55
  void trust;
56
56
  void agent;
57
57
  ```
58
58
 
59
59
  ## Extension and configuration notes
60
- Root imports stay filesystem-free. Node settings files are caller-named and read once; optional missing files are skipped. Trust storage, prompts, approval UI, OAuth token storage, environment-variable selection, and persistent credentials belong in the host or an extension package. For Node.js hosts, [`@arnilo/prism-credentials-node`](credential-storage.md) provides encrypted-file and system-keychain backends. Passing `settings` / `credentials` on `AgentConfig` does not wire hidden runtime reads; hosts pass concrete values or resolvers to the provider/request edge that needs them.
60
+ Root imports stay filesystem-free. Node settings files are caller-named and read once; optional missing files are skipped. Trust storage, prompts, approval UI, OAuth token storage, environment-variable selection, and persistent credentials belong in the host or an extension package. For Node.js hosts, [`@arnilo/prism-credentials-node`](credential-storage.md) provides encrypted-file and system-keychain backends. Pass concrete settings values or credential resolvers to the provider/request edge that needs them; do not place them on `AgentConfig`.
61
61
 
62
62
  ## Security and performance notes
63
63
  Prism does not sandbox host tools or extensions. Prism does not read environment variables, keychains, user config files, package manifests, resources, settings providers, credential resolvers, or project-local extensions unless the host explicitly wires those operations. Redaction is exact known-secret replacement only; it is not secret detection. Permission and trust checks are one operation per guarded call and add no workers, watchers, retries, network, or filesystem scans.
@@ -37,6 +37,7 @@ import { createSqlitePersistence } from "@arnilo/prism-session-store-sqlite";
37
37
  | `filename` | `string` | SQLite database path. Use `:memory:` for ephemeral tests. |
38
38
  | `wal` | `boolean` | Enable WAL journal mode. Defaults to `true`. |
39
39
  | `busyTimeoutMs` | `number` | SQLite `busy_timeout` in milliseconds. Defaults to `5000`. |
40
+ | `feedbackRedactor` | `SecretRedactor` | Optional redaction for feedback comment/tags/metadata before insert. |
40
41
  | `fileMode` | `number` | Unix file mode for newly created database files. Defaults to `0o600`. |
41
42
  | `database` | `Database` | Advanced: supply an existing `better-sqlite3` handle (caller owns lifecycle). |
42
43
 
@@ -51,7 +52,7 @@ import { createSqlitePersistence } from "@arnilo/prism-session-store-sqlite";
51
52
  | `SessionStore.readBranchPath` | Recursive ancestor query from `leafId` (or latest leaf) in root→leaf order. |
52
53
  | `RunLedger.append*` | Inserts run/event/tool/usage rows; events receive monotonic per-run `sequence` values. |
53
54
  | `ProductionPersistenceStore.query*` | Parameterized cursor pagination on indexed columns. |
54
- | `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, and bounded pagination. |
55
+ | `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, bounded pagination, and workflow suspended/denied/schedule/state/replay values without a schema migration. |
55
56
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
56
57
  | `close()` | Closes the underlying database when the adapter opened it. |
57
58
 
@@ -98,7 +99,7 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
98
99
  - The package is optional and workspace-local; `@arnilo/prism` core has no SQLite dependency.
99
100
  - Hosts choose the database path and own backup, retention enforcement, and filesystem permissions.
100
101
  - `SessionAppendOptions` idempotency rows are durable in `prism_session_append_idempotency` and survive reopen.
101
- - Schema version **1** (`001_init`) matches `@arnilo/prism/testing/persistence-schema` — PostgreSQL adapters share the same model with dialect-local DDL.
102
+ - Schema version **3** applies `001_init`, additive `002_usage_scope`, and `003_run_feedback`. Migration 003 adds immutable `prism_run_feedback` rows with run FK/cascade deletion and owner/run/trace cursor indexes. `persistence.feedback` validates exact run ownership, bounds/redacts through optional `feedbackRedactor`, queries bounded pages, and deletes only exact-owned IDs. PostgreSQL shares the same model with dialect-local DDL.
102
103
  - Pass an existing `better-sqlite3` `Database` via `database` when your host already manages connections.
103
104
 
104
105
  ## Security and performance notes
@@ -118,5 +119,5 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
118
119
  - [Run ledger conformance](run-ledger-conformance.md): `assertRunLedgerConforms` / `runRunLedgerConformance`.
119
120
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): package matrix and threat model.
120
121
  - [Node JSONL session store](node-jsonl-session-store.md): dev-only single-process alternative.
121
- - [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` for durable multi-process workflow execution.
122
+ - [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` and `createWorkflowSchedules()` for durable background execution and schedules.
122
123
  - [Migration guide](migration.md): moving from JSONL/in-memory to database-backed persistence.
@@ -0,0 +1,71 @@
1
+ # Supervisor delegation
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-supervisor` adds optional runtime-selected delegation to an explicit local child allow-list. It returns normal `AgentRunResult` values and does not modify core `createAgent()` or deterministic workflows.
6
+
7
+ ## When to use it
8
+
9
+ Use a supervisor when a host or agent must choose a child dynamically. Use `@arnilo/prism-workflows` for known DAGs, durable checkpoints, schedules, replay, or human suspension.
10
+
11
+ ## Inputs / request
12
+
13
+ | API/field | Meaning |
14
+ | --- | --- |
15
+ | `createSupervisor({ ownership, children })` | Creates one ownership-scoped supervisor. |
16
+ | `SupervisorChild.createAgent(context)` | Child-owned factory; receives derived resource/thread IDs, narrowed permission, abort signal, and nested `delegate`. |
17
+ | `delegate({ childId, input, threadId?, limits?, signal? })` | Invokes one allow-listed child. Input is text and byte-bounded. |
18
+ | `hooks.before` | May reject, modify redacted input, or narrow limits/policy. |
19
+ | `hooks.after` | Observes redacted terminal summary; failures cannot alter settled result. |
20
+ | `limits` | Depth 4/16, active children 4/32, input 64 KiB/1 MiB, steps 8/64, tools 32/256, tokens 20k/1m, timeout 60s/30m, event queue 128/4096 default/hard. |
21
+
22
+ ## Outputs / response / events
23
+
24
+ `delegate()` returns the child's `AgentRunResult` or throws its `AgentRunError`/a supervisor denial or limit error. `subscribe()` emits bounded `delegation_started`, `delegation_finished`, `delegation_rejected`, and `delegation_error` metadata events.
25
+
26
+ ## Request/response example
27
+
28
+ ```json
29
+ {"childId":"research","input":"Check primary sources","limits":{"maxTokens":4000}}
30
+ ```
31
+
32
+ ## Implementation example
33
+
34
+ ```ts
35
+ import { createSupervisor } from "@arnilo/prism-supervisor";
36
+
37
+ const supervisor = createSupervisor({
38
+ ownership: { tenantId: "tenant", userId: "user" },
39
+ permission: parentPolicy,
40
+ children: {
41
+ research: {
42
+ permission: readOnlyPolicy,
43
+ createAgent: ({ resourceId, threadId, permission, delegate }) =>
44
+ createResearchAgent({ resourceId, threadId, permission, delegate }),
45
+ },
46
+ },
47
+ hooks: { before: ({ input }) => ({ input, limits: { maxTokens: 4000 } }) },
48
+ });
49
+
50
+ const result = await supervisor.delegate({ childId: "research", input: "Check sources" });
51
+ ```
52
+
53
+ ## Extension and configuration notes
54
+
55
+ Child factories resolve their own providers/credentials and construct context/memory using the supplied IDs. Parent, child, returned-agent, budget, and hook permission policies are AND-composed. Child/request/hook limits can only lower inherited limits. A nested factory can call the supplied `delegate()`; immutable path state rejects cycles and depth overflow.
56
+
57
+ ## Security and performance notes
58
+
59
+ - Child IDs are explicit; no package/provider discovery occurs.
60
+ - `resourceId` and `threadId` include supervisor/delegation/child identity. Do not replace them with parent memory IDs.
61
+ - Tool budget is checked before side effects. Token usage is enforced on terminal aggregate usage and can exceed by at most one provider turn because providers report tokens after generation.
62
+ - Abort and timeout cover hooks, child creation, nested delegation, and the run. Host child code must cooperate with `AbortSignal`.
63
+ - Redaction applies before hook input, run metadata/results, completion hooks, and events. Child credentials are never supplied in delegation context.
64
+ - Static workflows remain smaller and more reproducible for known graphs.
65
+
66
+ ## Related APIs
67
+
68
+ - [A2A interoperability](a2a.md): remote protocol boundary.
69
+ - [Workflows](workflows.md): preferred deterministic orchestration.
70
+ - [Working and semantic memory](working-and-semantic-memory.md): child scope construction.
71
+ - [Host security](host-security.md): permission and credential boundaries.