@arnilo/prism 0.3.2 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (116) hide show
  1. package/CHANGELOG.md +17 -0
  2. package/README.md +34 -57
  3. package/dist/agent-run-lifecycle.js +4 -0
  4. package/dist/agent-run-state.d.ts +4 -0
  5. package/dist/agent-run-state.js +18 -5
  6. package/dist/agent-session/session.d.ts +7 -0
  7. package/dist/agent-session/session.js +59 -2
  8. package/dist/cli-dev.d.ts +29 -0
  9. package/dist/cli-dev.js +52 -0
  10. package/dist/cli-init.d.ts +17 -2
  11. package/dist/cli-init.js +194 -21
  12. package/dist/cli-runner.d.ts +5 -1
  13. package/dist/cli-runner.js +12 -1
  14. package/dist/contracts-core/agent.d.ts +6 -0
  15. package/dist/contracts-protocol.d.ts +18 -0
  16. package/dist/contracts-run-state.d.ts +1 -2
  17. package/dist/index.d.ts +3 -1
  18. package/dist/index.js +2 -1
  19. package/dist/input.d.ts +8 -0
  20. package/dist/input.js +4 -0
  21. package/dist/testing/persistence-schema.d.ts +1 -1
  22. package/dist/testing/persistence-schema.js +32 -28
  23. package/dist/testing/tool-conformance.d.ts +25 -0
  24. package/dist/testing/tool-conformance.js +128 -1
  25. package/dist/tool-search.d.ts +76 -0
  26. package/dist/tool-search.js +199 -0
  27. package/docs/0.1.0-readiness.md +2 -2
  28. package/docs/acp-agent.md +1 -1
  29. package/docs/antigravity-agent.md +1 -1
  30. package/docs/browser-automation.md +5 -5
  31. package/docs/caveman.md +2 -2
  32. package/docs/cli-rpc.md +26 -3
  33. package/docs/coding-security.md +1 -1
  34. package/docs/coding-tools.md +82 -0
  35. package/docs/compaction-and-retry.md +2 -2
  36. package/docs/compaction-llm.md +4 -4
  37. package/docs/compaction-observational-memory.md +3 -3
  38. package/docs/context-and-skills.md +2 -0
  39. package/docs/core.md +85 -0
  40. package/docs/credential-storage.md +1 -1
  41. package/docs/database-persistence.md +4 -0
  42. package/docs/dev-inspector.md +103 -0
  43. package/docs/diagrams.md +247 -0
  44. package/docs/documents.md +213 -0
  45. package/docs/evaluations.md +35 -1
  46. package/docs/graft.md +3 -3
  47. package/docs/guardrails.md +1 -1
  48. package/docs/host-security.md +4 -3
  49. package/docs/impeccable.md +2 -2
  50. package/docs/index.md +31 -20
  51. package/docs/mcp-tools.md +1 -1
  52. package/docs/migrate-to-0.4.md +312 -0
  53. package/docs/migration.md +22 -0
  54. package/docs/model-routing.md +1 -1
  55. package/docs/multi-agent-patterns.md +177 -0
  56. package/docs/multimodal-content.md +1 -1
  57. package/docs/obscura.md +10 -10
  58. package/docs/openapi-tools.md +1 -1
  59. package/docs/performance.md +23 -3
  60. package/docs/persistence-credentials-multimodality-primitives.md +1 -1
  61. package/docs/policy-and-audit.md +1 -1
  62. package/docs/ponytail.md +2 -2
  63. package/docs/prompt-registry.md +106 -0
  64. package/docs/provider-caching.md +32 -32
  65. package/docs/provider-conformance.md +1 -1
  66. package/docs/provider-packages.md +19 -19
  67. package/docs/provider-primitives.md +4 -4
  68. package/docs/providers/ai-sdk.md +3 -3
  69. package/docs/providers/alibaba.md +5 -5
  70. package/docs/providers/anthropic.md +6 -6
  71. package/docs/providers/azure.md +3 -3
  72. package/docs/providers/bedrock.md +3 -3
  73. package/docs/providers/clinepass.md +3 -3
  74. package/docs/providers/deepseek.md +3 -3
  75. package/docs/providers/google.md +4 -4
  76. package/docs/providers/kimi.md +3 -3
  77. package/docs/providers/neuralwatt.md +8 -8
  78. package/docs/providers/ollama.md +3 -3
  79. package/docs/providers/openai-compatible.md +1 -1
  80. package/docs/providers/openai.md +5 -5
  81. package/docs/providers/opencode-go.md +4 -4
  82. package/docs/providers/openrouter.md +3 -3
  83. package/docs/providers/vertex.md +5 -5
  84. package/docs/providers/xai.md +3 -3
  85. package/docs/providers/zai.md +3 -3
  86. package/docs/rag.md +5 -5
  87. package/docs/release-and-install.md +98 -50
  88. package/docs/runs-and-usage.md +14 -1
  89. package/docs/server.md +90 -1
  90. package/docs/sheets.md +229 -0
  91. package/docs/supervisors.md +1 -0
  92. package/docs/thinking-and-reasoning.md +10 -10
  93. package/docs/tool-conformance.md +27 -2
  94. package/docs/tools.md +29 -2
  95. package/docs/web-tools.md +2 -2
  96. package/docs/wiki.md +6 -6
  97. package/docs/workflow-orchestration-primitives.md +24 -0
  98. package/docs/workflows.md +70 -9
  99. package/docs/working-and-semantic-memory.md +53 -5
  100. package/package.json +10 -30
  101. package/templates/README.md +23 -0
  102. package/templates/deep-research/README.md.tmpl +47 -0
  103. package/templates/deep-research/env.example.tmpl +12 -0
  104. package/templates/deep-research/gitignore.tmpl +7 -0
  105. package/templates/deep-research/manifest.json +12 -0
  106. package/templates/deep-research/package.json.tmpl +23 -0
  107. package/templates/deep-research/src/agent.ts.tmpl +81 -0
  108. package/templates/deep-research/src/index.ts.tmpl +53 -0
  109. package/templates/deep-research/src/tests/research.test.ts.tmpl +114 -0
  110. package/templates/deep-research/src/tools.ts.tmpl +86 -0
  111. package/templates/deep-research/src/types.ts.tmpl +45 -0
  112. package/templates/deep-research/src/workflow.ts.tmpl +156 -0
  113. package/templates/deep-research/tsconfig.json.tmpl +15 -0
  114. package/templates/init/manifest.json +5 -0
  115. package/templates/init/package.json.tmpl +2 -1
  116. package/templates/init/providers.json +16 -16
package/docs/migration.md CHANGED
@@ -1,5 +1,27 @@
1
1
  # Migration guide
2
2
 
3
+ ## 0.3.3 → 0.4.0 package reorganization (breaking)
4
+
5
+ Prism 0.4 consolidates package names into explicit family subpaths. It is a dependency and import-specifier migration, not a persisted-data migration. See the complete [legacy 0.3 → 0.4 guide](migrate-to-0.4.md) for all 54 retired package mappings, profile replacements, optional peers/host binaries, security checks, rollback, and npm legacy-warning behavior.
6
+
7
+ ## 0.3.1 → 0.3.2: bounded workflow loop durability (additive, no migration)
8
+
9
+ `@arnilo/prism-workflows@0.3.2` adds the bounded `loopNode` durable extension. It adds optional `WorkflowNodeCheckpoint.iterations` records, each carrying `schemaVersion: 1`, a zero-based `iteration`, stable `iterationId`, and bounded/redacted output. Existing `WorkflowCheckpointValue.schemaVersion` remains `1`; older checkpoints without `iterations` remain readable through the legacy `iteration`/`lastOutput` cursor, and older hosts ignore the additive field. No SQL or generic checkpoint-store migration is required. Replay creates a new run and never mutates source iteration evidence. Hosts using saga compensation keep one saga step/aggregate and register per-iteration compensation by `iterationId` in reverse order.
10
+
11
+ This independent package patch freezes budget accounting: `maxNodes` counts declared DAG nodes once, while loop body executions consume only the required hard-capped `maxIterations` budget. Rollback is package-version rollback; no persisted migration is needed.
12
+
13
+ ## 0.3.2 → 0.3.3: run-ledger prompt provenance (additive, schema version 9)
14
+
15
+ Plan 042 adds an optional typed `promptVersion` ref (`{ name, version, hash }`) to `RunOptions` and `RunRecord`. Hosts resolve a prompt from `@arnilo/prism-prompts` and stamp the run: the ref is copied onto the start/finish ledger records and persisted by the first-party SQLite/PostgreSQL stores as a nullable `prompt_version` JSON column (shared schema migration `009_run_prompt_version`, schema version 8 → 9, forward-only and applied automatically by the adapters' checksummed `prism_migrations`). Strictly additive: unset `promptVersion` produces byte-identical rows and records, legacy rows read back without the field, and no exported declaration was removed. The ref carries identity only (`sha256:` body hash) — prompt bodies stay in the separate `@arnilo/prism-prompts` tables and out of run rows, metadata, and telemetry.
16
+
17
+ ## 0.3.2 → 0.3.3: tool progressive disclosure (additive, no migration)
18
+
19
+ Plan 041 adds opt-in progressive tool loading to `@arnilo/prism`: `toolsDisclosure` (default `"all"`, byte-identical to previous releases) and `toolsSearch.topK` on `AgentConfig` / `RunOptions`, plus the generated `search_tools` tool in search mode. Strictly additive — no exported declaration removed, no persisted shape repurposed. Durable run state gains an optional `sessionState.activatedToolNames` (names only, capped at 128); stores that ignore it resume exactly as before. Set nothing and behavior is unchanged; see [Tools](tools.md#tool-disclosure-progressive-tool-loading).
20
+
21
+ ## 0.3.1 → 0.3.2 memory package: composite recall scoring (additive, no migration)
22
+
23
+ `@arnilo/prism-memory@0.3.2` adds opt-in `RecallOptions.scoring`: sum-normalized similarity/recency/importance blending, with a positive `halfLifeMs` required only when `recencyWeight > 0`. Default recall (no `scoring`) keeps its existing ordering and query count. `MemoryVectorRecord.importance?` persists through an additive nullable `importance REAL` column (`ADD COLUMN IF NOT EXISTS`); legacy NULL rows score neutral `1.0`, so no re-index or data migration is required. At write, hosts may pass a clamped `[0,1]` `entry.importance` or an `importanceFrom` hook over a redacted reflection; it runs once at write, never at recall. Rollback is package-version rollback only: old readers ignore the nullable column, and new readers treat absent values neutrally.
24
+
3
25
  ## 0.3.0 → 0.3.1 production RAG engine (independent patch)
4
26
 
5
27
  Only `@arnilo/prism-rag`, `@arnilo/prism-memory`, and `@arnilo/prism-observability-opentelemetry` move to `0.3.1`. Keep every other first-party package on `^0.3.0` — those ranges already satisfy `0.3.1`.
@@ -156,4 +156,4 @@ await router.recordOutcome({ identity, provider, model, success: true, latencyMs
156
156
  - [Policy and audit](policy-and-audit.md)
157
157
  - [Agent identity](agent-identity.md)
158
158
  - [Enterprise PostgreSQL state](enterprise-postgres-state.md): durable router state, migration, cleanup, and ownership requirements.
159
- - Package README: [`@arnilo/prism-model-router`](../packages/model-router/README.md)
159
+ - Package README: [`@arnilo/prism-core`](../packages/prism-core/README.md)
@@ -0,0 +1,177 @@
1
+ # Multi-agent patterns: handoff, hierarchical crew, supervisor delegation, A2A
2
+
3
+ ## What it does
4
+
5
+ Maps the four Prism answers for "more than one agent" onto one decision table. All four compose existing seams — none introduces a new runtime:
6
+
7
+ - **In-session handoff (swarm)** — agent A transfers control of the ongoing conversation to agent B by calling a host-built `handoff` tool; the host resolves the target `AgentDefinition` with `resolveAgentDefinition` and opens the specialist against the same session (same store + session id, previous run's `leafId`). One transcript, no new session. No helper primitive ships; the tool factory lives in [`examples/handoff-swarm.ts`](../examples/handoff-swarm.ts).
8
+ - **Hierarchical crew** — a manager agent decomposes a goal into typed tasks (`{ tasks: [{ role, instruction }] }`) via structured output ([`Artifact*`](structured-output.md)), fans out to parallel role specialists with bounded `maxFanOut` ([`fanOutNode`](workflows.md)), aggregates deliverables with host reduce ([`joinNode`](workflows.md)), and validates outputs with conditional routing to completion or revision ([`conditionalNode`](workflows.md)). The entire process is a deterministic DAG workflow with zero new runtime primitives. Live demo in [`examples/crew-hierarchy.ts`](../examples/crew-hierarchy.ts).
9
+ - **Supervisor delegation** — `@arnilo/prism-supervisor` `delegate()` invokes allow-listed child agents as bounded runs and returns their result to the parent. Separate child transcripts, hooks, budgets, narrowing.
10
+ - **A2A 1.0** — cross-service interop over the JSON-RPC/HTTPS binding; the remote peer's lifecycle is host-owned behind `A2ATaskLifecycle`.
11
+
12
+ ## When to use it
13
+
14
+ | Pattern | Use when | Conversation boundary | Ownership / identity | Telemetry |
15
+ | --- | --- | --- | --- | --- |
16
+ | In-session handoff | One host, one ongoing conversation; the model decides **when** to transfer; specialists are alternate definitions of the same app | One continuous transcript chain (same store, session id, `leafId`) | Same session scope; give the specialist its own identity via its definition (`AgentConfig.identity` / `RunOptions.identity`) | Attribution is per-run: each `session.run()`'s events/result belong to the active definition — record the swap in host bookkeeping; no `delegated_agent_step` event exists for in-process swaps |
17
+ | Hierarchical crew | A goal requires dynamic decomposition by a manager LLM, parallel execution by role specialists, host aggregation, and conditional validation/revision loop | Workflow DAG execution — each specialist executes a bounded child task session; final deliverable returns to host | Workflow tenant/ownership scopes propagate; specialists activate only their own narrowed `tools` | Workflow node events (`node_started`/`node_finished`/`agent_event`); task attribution per role in the aggregated deliverable |
18
+ | Supervisor delegation | Parent agent needs a child as a *tool call*: bounded budget, hooks that redact/narrow, nested delegation, durable child approvals | Separate runs; child result returns to the parent transcript | Parent identity/effectStore propagate; child factories receive derived resource/thread ids and AND-composed permission | Dedicated `delegation_started/finished/rejected/error` events, projectable through observability `handleDelegation()`; opt-in `delegation_child_event` passthrough |
19
+ | A2A 1.0 | The other agent is owned by a **different service/deployment**; cross-org or cross-cluster; needs durable task lifecycle, push configs, streaming | Protocol boundary (JSON-RPC/HTTPS agent card); replay/reconnect via host-owned task adapter | Exact-origin verified client, `A2AAuthorization` per operation, principal-scoped push configs | Host-owned task adapter records the remote lifecycle; Prism creates no worker/store |
20
+
21
+ Rule of thumb: same conversation → handoff; dynamic task decomposition + parallel execution → hierarchical crew; same process but a subtask → supervisor delegation; different deployment/trust boundary → A2A.
22
+
23
+ ## How in-session handoff works
24
+
25
+ The pattern is a definition swap over existing seams — triage keeps calling `handoff(target)`; the host authorizes, swaps, and continues the same session:
26
+
27
+ ```ts
28
+ // Host-built allow-list tool; untrusted target -> fail-closed tool error result.
29
+ const handoffTool: ToolDefinition = {
30
+ name: "handoff",
31
+ description: "Transfer this conversation to a named specialist agent.",
32
+ parameters: { type: "object", required: ["target"], properties: { target: { type: "string" } } },
33
+ execute(args, ctx): ToolResult {
34
+ const target = String((args as { target: string }).target);
35
+ if (!(target in handoffTargets)) {
36
+ return { toolCallId: ctx.toolCallId, name: "handoff", error: { message: `Unknown handoff target: ${target}` } };
37
+ }
38
+ return { toolCallId: ctx.toolCallId, name: "handoff", value: { transferredTo: target } };
39
+ },
40
+ };
41
+
42
+ // Host authorizes the transfer mid-run: resolve the target definition and
43
+ // continue the SAME session (same store + sessionId). The previous run's
44
+ // leafId carries the transcript pointer; without it the next append forks
45
+ // a sibling branch and the specialist loses the carried context.
46
+ const specialist = await resolveAgentDefinition(handoffTargets[target], {
47
+ tools: [refundTool], // narrowed: no handoff tool unless the host allows it
48
+ overrides: { provider: specialistProvider },
49
+ });
50
+ const specialistSession = createAgentSession({ agent: specialist, store, id: "handoff-demo", leafId: triageRun.leafId });
51
+ ```
52
+
53
+ Live demo: [`examples/handoff-swarm.ts`](../examples/handoff-swarm.ts) — triage → billing transfer with fail-closed unknown target, a specialist whose re-handoff attempt is blocked (`unknown_tool`), and zero provider calls for the swap (it is a registry-level operation).
54
+
55
+ Two non-obvious details the example encodes:
56
+
57
+ 1. **`leafId` carries the transcript pointer.** The specialist session must pass the triage run's `leafId`; creating the session without it appends to a sibling branch and the specialist loses the carried context.
58
+ 2. **Carried context is the transcript chain itself.** Handoff is not delegation: there is no input-payload boundary to sanitize; whatever was said to triage is what the specialist reads.
59
+
60
+ ## How hierarchical crew orchestration works
61
+
62
+ Hierarchical multi-agent orchestration (the CrewAI "Hierarchical Process" pattern) decomposes a high-level goal into structured tasks, assigns each task to a role specialist agent in parallel, aggregates deliverables, and validates the outcome in a deterministic workflow:
63
+
64
+ ```ts
65
+ // 1. Manager produces a typed task plan via structured output.
66
+ // Untrusted model output is validated against the schema before updating state.
67
+ const manager = agentNode({
68
+ agent: "manager",
69
+ input: (ctx) => ({ goal: ctx.workflowInput }),
70
+ output: async (ctx) => {
71
+ const plan = parseTaskPlan(await getSessionOutput(ctx.session));
72
+ if (!plan.ok) throw new Error(`Invalid task plan: ${plan.error}`);
73
+ await ctx.updateState({ plan: plan.value });
74
+ return plan.value;
75
+ },
76
+ });
77
+
78
+ // 2. fan_out maps each task item to its corresponding role specialist.
79
+ const fan = fanOutNode({
80
+ items: (ctx) => (ctx.state.plan as TaskPlan).tasks,
81
+ map: async (task, _index, _ctx) => {
82
+ const agent = await resolveAgentDefinition(definitions[task.role], { tools });
83
+ const session = createAgentSession({ agent });
84
+ const result = await session.run(task.instruction);
85
+ return { role: task.role, result: result.text, attribution: { agent: task.role } };
86
+ },
87
+ maxFanOut: 8,
88
+ });
89
+
90
+ // 3. join aggregates all specialist deliverables and computes per-role attribution.
91
+ const aggregate = joinNode({
92
+ from: "fan",
93
+ reduce: async (items, ctx) => ({
94
+ deliverables: items,
95
+ summary: items.map((d) => `[${d.role}]: ${d.result}`).join("\n"),
96
+ validationPassed: evaluateQuality(items),
97
+ }),
98
+ });
99
+
100
+ // 4. conditional validation routes to completion or revision.
101
+ const validate = conditionalNode({
102
+ when: async (ctx) => Boolean((ctx.upstream.aggregate as AggregatedDeliverable).validationPassed),
103
+ then: ["complete"],
104
+ else: ["revise"],
105
+ });
106
+
107
+ const complete = functionNode({ execute: async (ctx) => formatDeliverable(ctx.upstream.aggregate) });
108
+ const revise = functionNode({ execute: async (ctx) => formatRevision(ctx.upstream.aggregate) });
109
+
110
+ // 5. Entire flow is a single defineWorkflow DAG with fixed revision ID.
111
+ const crewWorkflow = defineWorkflow({
112
+ revision: "crew-demo-1",
113
+ id: "hierarchical-crew",
114
+ nodes: { manager, fan, aggregate, validate, complete, revise },
115
+ edges: [
116
+ ["manager", "fan"],
117
+ ["fan", "aggregate"],
118
+ ["aggregate", "validate"],
119
+ ["validate", "complete"],
120
+ ["validate", "revise"],
121
+ ],
122
+ limits: { maxFanOut: 8, maxConcurrency: 4, maxNodes: 32 },
123
+ });
124
+ ```
125
+
126
+ Live demo: [`examples/crew-hierarchy.ts`](../examples/crew-hierarchy.ts) — manager structured task decomposition, parallel role specialists (`researcher`, `writer`), host reduce aggregation with per-role attribution, and conditional validation/revision routing.
127
+
128
+ ## CrewAI to Prism mapping table
129
+
130
+ | CrewAI Concept | Prism Primitive | Notes & Documentation |
131
+ | --- | --- | --- |
132
+ | **Crew** | Workflow ([`defineWorkflow`](workflows.md)) | A deterministic DAG with explicit revision id, node concurrency, and checkpoint persistence. |
133
+ | **Manager Agent** | Agent Node ([`agentNode`](workflows.md)) + Structured Output ([`generateValidateReviseLoop`](structured-output.md)) | Manager emits a typed `{ tasks: [{ role, instruction }] }` schema via `ArtifactValidator`/`ArtifactParser`. |
134
+ | **Task** | Fan-out item ([`fanOutNode`](workflows.md)) | Bounded dynamic fan-out (`maxFanOut`), mapping each decomposed task to a role specialist session. |
135
+ | **Role Agent (Specialist)** | Agent Definition ([`resolveAgentDefinition`](agent-definitions.md)) | Declarative agent with fail-closed tool narrowing; activated per role during `fan_out.map`. |
136
+ | **Process (Sequential / Hierarchical)** | Workflow DAG ([`defineWorkflow`](workflows.md) / Edges) | Edges define data and execution dependencies; no unconstrained agent-to-agent loops. |
137
+ | **Task Output Aggregation** | Join Node ([`joinNode`](workflows.md) + `reduce`) | Host-controlled reduction aggregating specialist outputs and computing per-role attribution. |
138
+ | **Validation & Quality Review** | Conditional Node ([`conditionalNode`](workflows.md)) | Deterministic branch routing to `complete` or `revise` based on validation criteria. |
139
+ | **Process Revision Loop** | Node Retries / DAG Branching / Loop Node ([`loopNode`](workflows.md)) | Bounded retry/revision path or bounded in-graph loop iteration. |
140
+
141
+ ## Where Prism is stronger
142
+
143
+ - **Durable Human-in-the-Loop (HITL)**: Prism workflows support durable pause and resume via [`suspend()`](workflows.md#durable-suspension-and-resumption) and [`resumeWorkflow()`](workflows.md) across worker restarts or approval gates ([Agent durable approval](agent-session-runtime.md)).
144
+ - **Strict Budget & Concurrency Caps**: Workflows enforce hard limits on `maxNodes`, `maxFanOut`, `maxConcurrency`, and timeout bounds ([Workflow limits](workflows.md)).
145
+ - **Fail-Closed Capability Narrowing**: Specialists receive only their explicitly authorized `tools` via [`resolveAgentDefinition`](agent-definitions.md); managers cannot invoke specialist tools directly, preventing accidental tool leakage.
146
+ - **Untrusted Model Output Validation**: Manager task plans are treated as untrusted LLM output and validated against a typed schema before triggering fan-out ([Structured output](structured-output.md)).
147
+ - **Durable Audit & Telemetry**: Every node start/finish and agent event is emitted with deterministic sequence numbers and can be persisted to signed audit ledgers ([Policy and audit](policy-and-audit.md), [Observability](observability.md)).
148
+
149
+ ## Security and performance notes
150
+
151
+ - **Transfers are explicit model-initiated, host-authorized.** The `handoff` tool exists on the triage agent's allow-list only; the target name is validated against the host-authored targets map before any definition resolves. Unknown names fail closed as a standard tool error (`Unknown handoff target: <name>`).
152
+ - **No permission escalation through handoff or delegation.** The specialist's capabilities come solely from its own `AgentDefinition` as resolved by `resolveAgentDefinition` (fail-closed for omitted capabilities). Handoff or fan-out grants nothing: tools/identity are what the host put on that definition. The specialist cannot invoke manager tools unless its definition explicitly includes them — the standard `unknown_tool` block applies otherwise.
153
+ - **Narrowing on transfer, never widening.** If the specialist needs the caller's verified identity, project it through `narrowIdentity` / `assertIdentityPropagation` ([Agent identity](agent-identity.md)) so scopes and tenant cannot widen across the swap. For delegation the same discipline is built in (`narrowIdentity`, AND-composed policies); for A2A the exact-origin client plus per-operation authorization is the boundary.
154
+ - **Manager-generated task plans are untrusted model output.** Manager plan outputs are validated against the typed schema via `ArtifactValidator` before being persisted to workflow state or dispatched to `fan_out`. Malformed or invalid plans trigger the artifact repair loop or fail closed before any specialist is invoked.
155
+ - **Redaction of carried context.** Handoff carries the raw transcript by design — same rows a human replay would read. Apply the session egress seams on the way out: `redactSessionEntry` / `redactMessage` with a host field policy (see [Data classification](data-classification.md)) and `AgentConfig.redactor`; for durable replay across tenants reuse the redacted transcript seam discipline used by ACP `sessions.transcript` ([ACP interop](acp.md)).
156
+ - **Telemetry attribution.** Which agent produced which turn is not stored on message entries; the host knows (it performed the swap or aggregated fan-out results) and should pin it per run via `RunOptions.identity` (principal kind `agent`) so `identityTelemetryAttributes` (`prism.identity.*`) carries redacted attribution on telemetry, or via observability metadata. Supervisor runs emit dedicated `delegation_*` events; adapter-owned delegation timelines emit `delegated_agent_step` through `createDelegatedAgentStep` (see [Agent events](agent-events.md)) — an in-process definition swap has no session seam to emit it, which is why the swap records attribution host-side.
157
+ - **Performance.** The swap performs zero provider calls; it costs one registry resolution plus one session open (~sub-millisecond in the example fixture). The transferred turn costs what any tool round costs.
158
+
159
+ ## Extension and configuration notes
160
+
161
+ - Handoff targets may be code-defined `AgentDefinition` objects or `<configRoot>/agents/<name>/AGENT.md` bundles resolved via `resolveAgentBundle` — the allow-list maps names to either.
162
+ - Hosts wanting the pattern behind a UI timeline can emit their own step events from the swap; `delegated_agent_step` remains reserved for adapter-driven delegation loops (`@arnilo/prism-antigravity-agent`).
163
+ - A reusable in-session handoff helper was evaluated and **not** shipped in 0.3.x: the unavoidable boilerplate is a ~20-line allow-list tool plus one `createAgentSession` call. Revisit only if multiple hosts show materially different swap semantics.
164
+ - Hierarchical crew patterns compose entirely on existing `@arnilo/prism-workflows` and `@arnilo/prism` primitives (`agentNode`, `fanOutNode`, `joinNode`, `conditionalNode`, `ArtifactValidator`, `resolveAgentDefinition`); no separate helper package is needed.
165
+
166
+ ## Related APIs
167
+
168
+ - [Workflows](workflows.md): `defineWorkflow`, `fanOutNode`, `joinNode`, `conditionalNode`, `runWorkflow`.
169
+ - [Structured output](structured-output.md): `ArtifactParser`, `ArtifactValidator`, `generateValidateReviseLoop`.
170
+ - [Agent definitions](agent-definitions.md): `resolveAgentDefinition` and fail-closed capability activation — the swap seam itself.
171
+ - [Supervisor delegation](supervisors.md): same-process subtasks with budgets, hooks, and durable nested approvals.
172
+ - [A2A interoperability](a2a.md): the cross-service protocol boundary.
173
+ - [Agent identity](agent-identity.md): verified identity propagation and narrowing (`narrowIdentity`, `assertIdentityPropagation`).
174
+ - [Agent events](agent-events.md): `delegated_agent_step` and delegation event surfaces for timelines.
175
+ - [Policy and audit](policy-and-audit.md): decision ledger, approval gates, and signed audit export.
176
+ - [Observability](observability.md): OpenTelemetry agent/provider/tool hierarchy and metrics.
177
+ - [Middleware hooks](middleware-hooks.md): context bridging and redaction without permission grants.
@@ -136,7 +136,7 @@ try {
136
136
  - Supplying `fetch` is a trusted compatibility/custom-transport escape hatch: Prism still checks URL literals and host allow-lists, but the host-provided fetch owns DNS resolution, rebinding protection, redirects, proxies, TLS, auth, and logging.
137
137
  - `resourceUri` resolution requires a caller-provided `ResourceLoader` and optional `ResourceLoadContext.permission` check.
138
138
  - Local filesystem paths should use trust policies such as `createPathTrustPolicy()` before exposing URIs to loaders.
139
- - Provider upload/create/delete lifecycles are provider-package-local. `@arnilo/prism-provider-openai` inlines files under 4 MiB as `data:<mediaType>;base64,...` `file_data`, otherwise uses a bounded per-run upload cache and best-effort `DELETE /v1/files` cleanup after each stream.
139
+ - Provider upload/create/delete lifecycles are provider-package-local. `@arnilo/prism-providers/openai` inlines files under 4 MiB as `data:<mediaType>;base64,...` `file_data`, otherwise uses a bounded per-run upload cache and best-effort `DELETE /v1/files` cleanup after each stream.
140
140
  - Shared wire helpers live in `@arnilo/prism/providers/media` (`resolveProviderMediaMessages`, `serializeOpenAIResponsesInputFile`, `serializePdfDocumentWireBlock`, `createBoundedUploadCache`). OpenAI Responses, Kimi, and OpenCode Go Anthropic routes resolve their complete media collection once before serialization or upload.
141
141
  - OpenAI Realtime audio is a bidirectional `RealtimeSession` stream, not a `ContentBlock`: provide host-captured `Uint8Array` chunks with `sendAudio()` and consume untrusted `audio_delta` / transcript events. It has a fixed 256 events/s, 1 MiB/s, and 600 s default ceiling.
142
142
 
package/docs/obscura.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # Obscura browser engine
2
2
 
3
- Optional `@arnilo/prism-obscura` support for a host-installed
3
+ Optional `@arnilo/prism-web-tools/obscura` support for a host-installed
4
4
  [Obscura](https://github.com/h4ckf0r0day/obscura) headless browser. Obscura is never
5
5
  bundled — install the binary (or use the `h4ckf0r0day/obscura` Docker image) and point
6
6
  the package at it.
@@ -11,7 +11,7 @@ the package at it.
11
11
  ## Install
12
12
 
13
13
  ```bash
14
- npm install @arnilo/prism-obscura
14
+ npm install @arnilo/prism-web-tools @arnilo/prism-mcp
15
15
  ```
16
16
 
17
17
  ## Process lifecycle (`spawnObscuraProcess`)
@@ -23,7 +23,7 @@ environment (`PATH`, `HOME`), and insecure flags (`--allow-private-network`,
23
23
  is set explicitly.
24
24
 
25
25
  ```ts
26
- import { spawnObscuraProcess } from "@arnilo/prism-obscura";
26
+ import { spawnObscuraProcess } from "@arnilo/prism-web-tools/obscura";
27
27
 
28
28
  const obscura = spawnObscuraProcess({
29
29
  command: "/usr/local/bin/obscura",
@@ -53,7 +53,7 @@ Connects to `obscura mcp` (stdio or Streamable HTTP) through
53
53
  allow-list, so future Obscura tools keep flowing through.
54
54
 
55
55
  ```ts
56
- import { createObscuraMcpTools } from "@arnilo/prism-obscura";
56
+ import { createObscuraMcpTools } from "@arnilo/prism-web-tools/obscura";
57
57
 
58
58
  const obscura = await createObscuraMcpTools({
59
59
  transport: { type: "stdio", command: "/usr/local/bin/obscura", args: ["mcp"] },
@@ -68,7 +68,7 @@ await obscura.close();
68
68
  - Effects: read/diagnostic/waiter/capture tools are effect-free; navigation,
69
69
  interaction, evaluation, cookie/storage writes, tabs, and any unknown future tool
70
70
  are exclusive, serialized external mutations (Obscura keeps one live page).
71
- - Naming: default `obscura_` prefix coexists with `@arnilo/prism-browser`;
71
+ - Naming: default `obscura_` prefix coexists with the `browser` subpath;
72
72
  `namePrefix: ""` preserves native Obscura names.
73
73
  - Transports: stdio configs are validated with the same fail-closed command policy;
74
74
  Streamable HTTP endpoints outside loopback require explicit `allowRemoteHttp` and
@@ -82,8 +82,8 @@ host's Playwright via `chromium.connectOverCDP`. `connect()` and browser launch
82
82
  never used; Prism never launches browsers.
83
83
 
84
84
  ```ts
85
- import { connectObscuraCdp } from "@arnilo/prism-obscura";
86
- import { createBrowserTools } from "@arnilo/prism-browser";
85
+ import { connectObscuraCdp } from "@arnilo/prism-web-tools/obscura";
86
+ import { createBrowserTools } from "@arnilo/prism-web-tools/browser";
87
87
 
88
88
  const session = await connectObscuraCdp({
89
89
  command: "/usr/local/bin/obscura",
@@ -107,7 +107,7 @@ await session.close(); // browser first, then the owned process
107
107
  APIs (`browser.newBrowserCDPSession()`, `context.newCDPSession(page)`); the package
108
108
  adds no CDP command allow-list.
109
109
  - Concurrency limit: pages served by one Obscura worker share one V8 isolate —
110
- CPU-bound page JavaScript can delay sibling pages. Keep `@arnilo/prism-browser`
110
+ CPU-bound page JavaScript can delay sibling pages. Keep the `browser` subpath
111
111
  limits authoritative; size Obscura's `--workers` for the host.
112
112
  - Screenshots/PDF require a render-enabled Obscura build and still obey the browser
113
113
  package's artifact/byte policy.
@@ -119,7 +119,7 @@ child processes. Returns standard Prism `web_search`/`web_fetch` tools plus expl
119
119
  `obscura_fetch`/`obscura_scrape` (disable with `nativeTools: false`).
120
120
 
121
121
  ```ts
122
- import { createObscuraWebTools } from "@arnilo/prism-obscura";
122
+ import { createObscuraWebTools } from "@arnilo/prism-web-tools/obscura";
123
123
 
124
124
  const web = createObscuraWebTools({ command: "/usr/local/bin/obscura" });
125
125
  agent.tools = [...agent.tools, ...web.tools];
@@ -149,7 +149,7 @@ agent.tools = [...agent.tools, ...web.tools];
149
149
  - Docker-style invocations work through `argsBefore` (e.g.
150
150
  `["run", "--rm", "-i", "h4ckf0r0day/obscura"]`).
151
151
  - An opt-in live smoke test runs against a real installed binary with
152
- `npm run test:live -w @arnilo/prism-obscura` plus `PRISM_LIVE_OBSCURA=1` and
152
+ `npm run test:live -w @arnilo/prism-web-tools` plus `PRISM_LIVE_OBSCURA=1` and
153
153
  `PRISM_OBSCURA_BIN=/path/to/obscura`.
154
154
 
155
155
  ## Host conformance (one generic integration)
@@ -53,4 +53,4 @@ Defaults and hard caps (frozen in `scripts/phase11-freeze-manifest.json`): `maxD
53
53
  - [Tools](tools.md): registry, dispatch, validation
54
54
  - [Recoverable tool effects](tool-effects.md): approval + idempotency contracts
55
55
  - [Host security guide](host-security.md): permission, trust, validation checklist
56
- - Package README: [`@arnilo/prism-openapi-tools`](../packages/prism-openapi-tools/README.md)
56
+ - Package README: [`@arnilo/prism-coding-tools`](../packages/prism-coding-tools/README.md)
@@ -65,6 +65,26 @@ serialized provider event at the exact response-byte cap succeeds; one byte belo
65
65
  fails closed, including multibyte Unicode deltas. Context-budget omission order and
66
66
  newest-history preservation remain covered by the root context-budget tests.
67
67
 
68
+ ## Tool progressive disclosure (plan 041)
69
+
70
+ `node scripts/benchmark.mjs --scenario tool-search` is network-free (mock assembly, in-memory, no credentials). It builds a 128-tool fixture registry and assembles the provider input once per mode through `assembleProviderInput`: `toolsDisclosure "all"` (default, full tool set) vs `"search"` (top-k 16 plus the generated `search_tools` tool), then asserts provider-request tool-definition bytes shrink ≥ 60% and the index+score pass stays well under a turn. Frozen caps live in `scripts/budgets.json#toolSearch` (reduction floor 0.6, index+score ceiling 50 ms, disclosed-count ceiling 33 — sanity bounds, machine-dependent). Schema/caps/network-free gating in `npm test`: `scripts/benchmark-tool-search.test.mjs`.
71
+
72
+ ```bash
73
+ node scripts/benchmark.mjs --scenario tool-search --out /tmp/prism-tool-search.json
74
+ ```
75
+
76
+ Recorded 2026-08-30, Node v24.19.0 / Linux x64: tool bytes 31,923 → 4,329 (**86.4% reduction**, floor 60%), index+score 1.9–2.5 ms across three runs, disclosed 17 tools (top-k 16 + `search_tools`). Tool-accuracy fixtures (mock provider picking by name among 64/128 distractors, scripted scanner reading only the disclosed list) show search mode at full-exposure pick accuracy in both sizes — the conformance floor `search ≥ all` holds (`src/__tests__/tool-search.test.ts`).
77
+
78
+ ## Workflow loop refinement (plan 045)
79
+
80
+ `node scripts/benchmark.mjs --scenario workflow-loop` is network-free: five serial `loopNode` iterations each run one refinement through a mock provider and an in-memory checkpoint adapter. The frozen budget in `scripts/budgets.json#workflowLoop` allows 50 ms p95 per node execution, or 250 ms across all five iterations. The scenario also checks five provider calls, five finished iteration records, peak provider concurrency of one, and zero active work after completion.
81
+
82
+ Recorded 2026-08-31 on Node v24.19.0 / Linux x64: 5 warmups + 20 measured runs, p50 **2.356 ms**, p95 **6.443 ms** (**1.289 ms/iteration**), 352.88 runs/s. `maxNodes` remains the declared-node count; `maxIterations` is the independent runtime budget and stays hard-capped at 64. These timings are local evidence, not portable SLOs.
83
+
84
+ ```bash
85
+ node scripts/benchmark.mjs --scenario workflow-loop --out /tmp/prism-workflow-loop.json
86
+ ```
87
+
68
88
  ## Current-line root artifact diet
69
89
 
70
90
  `npm pack --dry-run --json` on `@arnilo/prism` is gated by `scripts/budget-gate.test.mjs` against `scripts/budgets.json#root` (±5%). Repository-only history stays out of the tarball: `docs/_evidence/**`, `docs/release-*-evidence.md`, `docs/api-page-template.md`, `dist/__tests__`, and `*.map`. Every page linked from shipped `docs/index.md` must be in the pack. Recorded 2026-08-27: **923,045 packed / 3,149,665 unpacked / 375 files** (226 `dist` js+d.ts, 124 index-linked docs, 25 other). 0.1.0 freeze 713,454 / 293 stays historical.
@@ -365,7 +385,7 @@ Structured Git/check/handoff defaults/hard caps: paths 1,000/10,000; refs 1 KiB/
365
385
 
366
386
  Durable coding plan/checkpoint defaults/hard caps: plan Markdown 256 KiB/1 MiB; todos 1,000/10,000 with 512 B/4 KiB text; checkpoint metadata 64 KiB/512 KiB; artifact references 16/64 at 256 MiB/2 GiB each; check summaries 1 KiB/8 KiB. Checkpoints store URI/hash/summaries/fingerprints only; resume revalidates workspace root, base branch, plan hash, and tool/policy/image fingerprints before import.
367
387
 
368
- Browser automation defaults/hard caps from `@arnilo/prism-browser`: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2,000/10,000; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30 s/120 s; action 10 s/60 s; wait 30 s/120 s; run wall 20 min/30 min; popups 4/16; dialogs 16/64; listeners 64/256; action input 64 KiB/256 KiB; close grace 5 s/30 s; network requests 1,000/10,000 with 10/32 redirects per request and 8/32 WebSockets; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate. Caps charge before context/page/action/queue/snapshot/network/artifact retention. Host supplies Playwright and egress proxy attestation; package import launches nothing.
388
+ Browser automation defaults/hard caps from the `browser` subpath: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2,000/10,000; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30 s/120 s; action 10 s/60 s; wait 30 s/120 s; run wall 20 min/30 min; popups 4/16; dialogs 16/64; listeners 64/256; action input 64 KiB/256 KiB; close grace 5 s/30 s; network requests 1,000/10,000 with 10/32 redirects per request and 8/32 WebSockets; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate. Caps charge before context/page/action/queue/snapshot/network/artifact retention. Host supplies Playwright and egress proxy attestation; package import launches nothing.
369
389
 
370
390
  0.0.14 co-work defaults/hard caps (frozen in [Phase 9 evidence](_evidence/review-coverage-2026-07-25-phase-9.md)): conversation thread list pages 50/200, active branches per thread 16/64, replay/export page 100/500 events; artifact revisions per artifact 32/128, artifacts per thread 64/256, metadata record 8/64 KiB, preview 16/64 KiB, citations 32/128 (2/8 KiB each), delivery-link TTL 5 min/24 h, delivery token 4/16 KiB, compare exactly 2 revisions; memory retention batch 500/5000; proactive capability TTL 24 h/31 d, capability token record 16 KiB; browser checkpoint URL 8 KiB/16 KiB, domain-state hash 256 B/1 KiB, host-data ref 2 KiB/8 KiB, 16/64 checkpoints per run; device stream chunk 1 MiB/8 MiB, concurrent device sessions per identity 1/4 (device wall/turns/tool calls consume shared `RunLimits`). All caps charge before persist/emit and fail closed on overflow. Benchmark placeholder: `node scripts/benchmark-0.0.14.mjs` (release Task 12) reports conversation replay, memory injection/consent, artifact revision/delivery, AG-UI co-work mapping, and connector refresh overhead against these budgets.
371
391
 
@@ -595,7 +615,7 @@ Experiment concurrency is capped at 32 workers and defaults to 1. Scorers operat
595
615
 
596
616
  ### 0.0.5 Phase 6 verification (2026-07-15)
597
617
 
598
- Optional `@arnilo/prism-provider-ai-sdk` adapts AI SDK `LanguageModelV4` streams to Prism without adding an AI SDK dependency to core.
618
+ Optional `@arnilo/prism-providers/ai-sdk` adapts AI SDK `LanguageModelV4` streams to Prism without adding an AI SDK dependency to core.
599
619
 
600
620
  | Surface | Result |
601
621
  | --- | --- |
@@ -751,7 +771,7 @@ Offline behavior tests (identity propagation, policy export, router deny paths,
751
771
 
752
772
  ### 0.3.x Phase 39 Obscura browser-engine envelopes (2026-08-29)
753
773
 
754
- `@arnilo/prism-obscura` binary-backed legs, network-free, driven by a deterministic fake CLI: `node scripts/benchmark-obscura.mjs` (3 runs, medians vs reviewed ceilings; artifact `scripts/benchmark-obscura.json`). Startup leg probes SIG-0 liveness after spawn — a real host waits on its readiness endpoint inside the same bound.
774
+ `obscura` binary-backed legs, network-free, driven by a deterministic fake CLI: `node scripts/benchmark-obscura.mjs` (3 runs, medians vs reviewed ceilings; artifact `scripts/benchmark-obscura.json`). Startup leg probes SIG-0 liveness after spawn — a real host waits on its readiness endpoint inside the same bound.
755
775
 
756
776
  | Leg | Median (3 runs) | Ceiling | Notes |
757
777
  | --- | --- | --- | --- |
@@ -70,7 +70,7 @@ Static review of `src/contracts.ts`, `src/session-stores.ts`, `src/credentials.t
70
70
  | `refreshOAuthCredential` | `src/credentials.ts` | Calls `OAuthProvider.refresh`; optional `OAuthCredentialStore.set` |
71
71
  | `OAuthCredentialStore` | `src/contracts.ts` | `set(provider, credentials)` only — no `get`/`delete` in core contract |
72
72
  | `OAuthProvider` / `OAuthCredentials` | `src/contracts.ts` | `login`, optional `refresh`, optional `getCredential` |
73
- | Device-code OAuth | `packages/provider-openai` | Bounded polling; abort via `OAuthLoginCallbacks.signal` |
73
+ | Device-code OAuth | `packages/prism-providers/src/openai` | Bounded polling; abort via `OAuthLoginCallbacks.signal` |
74
74
  | Redaction | `src/redaction.ts` | Exact known-secret replacement; not secret detection |
75
75
 
76
76
  **Gaps (C-011):** No encrypted file store, no system keychain adapter, no versioned credential envelope, no `OAuthCredentialStore` `get`/`delete`/`list` in core (Task 4 package may extend store interface locally while integrating `refreshOAuthCredential`).
@@ -203,4 +203,4 @@ await evaluateAndAppend(request, { store: state.policy, evaluator, id: crypto.ra
203
203
  - [Workflows](workflows.md): proactive schedule capability enable/revoke events bridge here via `onCapability`.
204
204
  - [Host security](host-security.md)
205
205
  - [Enterprise PostgreSQL state](enterprise-postgres-state.md): durable policy/evaluation/work/router composition.
206
- - Package README: [`@arnilo/prism-policy`](../packages/policy/README.md)
206
+ - Package README: [`@arnilo/prism-core`](../packages/prism-core/README.md)
package/docs/ponytail.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-ponytail` is an optional package that wires [DietrichGebert/ponytail](https://github.com/DietrichGebert/ponytail) into Prism contribution contracts.
5
+ `@arnilo/prism-coding-tools/ponytail` is an optional package that wires [DietrichGebert/ponytail](https://github.com/DietrichGebert/ponytail) into Prism contribution contracts.
6
6
 
7
7
  It registers upstream skills and commands, injects active mode instructions via upstream `getPonytailInstructions` / `filterSkillBodyForMode`, and persists mode as session custom `ponytail-mode` entries. Import is inert; missing upstream fails closed at `setup` with a bounded redacted error.
8
8
 
@@ -66,7 +66,7 @@ Deactivation: exact phrases `stop ponytail` and `normal mode`.
66
66
  ## Implementation example
67
67
 
68
68
  ```ts
69
- import { createPonytailExtension } from "@arnilo/prism-ponytail";
69
+ import { createPonytailExtension } from "@arnilo/prism-coding-tools/ponytail";
70
70
  import {
71
71
  createExtensionKernel,
72
72
  createLoadSkillTool,
@@ -0,0 +1,106 @@
1
+ # Versioned prompt registry
2
+
3
+ ## What it does
4
+
5
+ The optional `@arnilo/prism-prompts` package stores prompt assets as immutable, content-hashed versions. It provides a memory store plus SQLite and PostgreSQL adapters. The registry returns prompt data; it does not compose system-prompt layers, evaluate prompt quality, discover files, or activate text.
6
+
7
+ ## Inputs / request
8
+
9
+ ```ts
10
+ import { createMemoryPromptStore } from "@arnilo/prism-prompts";
11
+
12
+ const store = createMemoryPromptStore();
13
+ const version = await store.put({
14
+ tenantId: "tenant-1",
15
+ name: "support-agent",
16
+ body: "Answer support questions briefly.",
17
+ labels: ["production"],
18
+ metadata: { owner: "support" },
19
+ });
20
+ ```
21
+
22
+ `put` always appends the next version for one ownership/name scope. It computes `hash` as `sha256:<64 lowercase hex>` over exact UTF-8 body bytes. Records, labels, and JSON metadata are frozen before return.
23
+
24
+ Ownership fields (`tenantId`, optional `accountId` and `userId`) are direct fields on every operation. Omitted ownership is a separate local scope, never a wildcard over tenant-owned rows.
25
+
26
+ ## Outputs / response / events
27
+
28
+ ```ts
29
+ const latest = await store.resolve({ tenantId: "tenant-1", name: "support-agent" });
30
+ const production = await store.resolve({ tenantId: "tenant-1", name: "support-agent", label: "production" });
31
+ const exact = await store.resolve({ tenantId: "tenant-1", name: "support-agent", version: 1 });
32
+
33
+ for (let page = await store.list({ tenantId: "tenant-1", name: "support-agent", limit: 50 });; ) {
34
+ consume(page.items);
35
+ if (!page.nextCursor) break;
36
+ page = await store.list({ tenantId: "tenant-1", name: "support-agent", cursor: page.nextCursor, limit: 50 });
37
+ }
38
+
39
+ const diff = await store.diff({ tenantId: "tenant-1", name: "support-agent", fromVersion: 1, toVersion: 2 });
40
+ ```
41
+
42
+ `resolve` returns the latest version by default, or the latest version carrying `label`; exact `version` can be combined with a label. `list` uses bounded keyset cursors ordered by name/version. `diff` returns bounded `context`/`add`/`remove` lines plus `added`, `removed`, and `truncated` counts. The host decides how a resolved body enters the existing `composeSystemPrompt` layers.
43
+
44
+ ## Run provenance
45
+
46
+ Pass the resolved version's identity to a run so every ledger record answers "which prompt version produced this output":
47
+
48
+ ```ts
49
+ const resolved = await store.resolve({ tenantId, name: "support-agent" });
50
+ await session.run(input, {
51
+ promptVersion: { name: resolved.name, version: resolved.version, hash: resolved.hash },
52
+ });
53
+ ```
54
+
55
+ The ref is opaque identity — name, version number, and the store's SHA-256 body hash — never prompt content. It rides on the start/finish `RunRecord`s, round-trips through first-party SQLite/PostgreSQL run rows (`prompt_version` column, schema migration `009_run_prompt_version`), and stays subject to the existing ledger redaction and field-policy boundaries. See [Runs and usage](runs-and-usage.md#prompt-provenance).
56
+
57
+ ## Durable adapters
58
+
59
+ ```ts
60
+ import { createSqlitePromptStore } from "@arnilo/prism-prompts";
61
+ const sqlite = createSqlitePromptStore({ filename: "./prompts.db" });
62
+
63
+ import { createPostgresPromptStore } from "@arnilo/prism-prompts";
64
+ const postgres = await createPostgresPromptStore({
65
+ connectionString: process.env.DATABASE_URL,
66
+ schema: "prism",
67
+ });
68
+ ```
69
+
70
+ SQLite uses `better-sqlite3`; PostgreSQL uses a caller-supplied or adapter-owned `pg` pool. Both adapters use the package-owned `prism_prompts` and `prism_prompt_labels` tables, exact ownership predicates, bound values, and an indexed label lookup. Startup applies checked `001_init` migration history and refuses checksum drift. SQLite exposes `applySqlitePromptMigrations` for managed setup tests; PostgreSQL migration setup is guarded by `pg_advisory_xact_lock`.
71
+
72
+ ## Eval-gated promotion
73
+
74
+ `assertPromptPromotion` composes [evaluations](evaluations.md) with the store to answer one question — should this candidate version replace the baseline? It resolves both versions (read-only), runs them head-to-head over a dataset through `runComparison`, and returns a typed verdict. It never promotes anything, writes nothing, and never touches a live agent:
75
+
76
+ ```ts
77
+ import { assertPromptPromotion } from "@arnilo/prism-prompts";
78
+
79
+ const v = await assertPromptPromotion({
80
+ store,
81
+ name: "support-agent",
82
+ candidate: { label: "candidate" }, // or an exact version
83
+ baseline: { label: "production" }, // must resolve to a different version
84
+ dataset,
85
+ scorers,
86
+ run: (prompt) => hostRunnerFactory(prompt.body), // host bridge: body → candidate
87
+ minimumWinRate: 0.8, // optional; default gate is a strict win majority
88
+ thresholds: { maximumFailures: 0 }, // optional; forwarded to assertEvaluationThreshold
89
+ });
90
+ if (v.verdict === "promote") await store.put({ ...hostInput, body: v.candidate.body, labels: ["production"] });
91
+ ```
92
+
93
+ The verdict carries `promote`/`hold`, per-scorer `wins/losses/ties/failures`, `winRate`, the raw `ComparisonReport`, a redacted bounded `reportJson` (`serializeEvaluationReport`), and `reasons` on hold. The default gate holds unless the candidate wins strictly more scored comparisons than the baseline; `minimumWinRate` and `thresholds` add stricter gates, and threshold equality passes. Requires the optional peer `@arnilo/prism-evals` (install it or the helper fails closed with `ERR_PRISM_PROMPT_EVALS_PEER`). Promotion itself stays a host decision: applying the verdict means `put`-ing a new version with labels — the helper never does.
94
+
95
+ ## Limits and security
96
+
97
+ Names, bodies, labels, metadata, cursors, pages, and diffs have finite defaults and hard caps. Prompt bodies are data: no evaluation, template execution, file discovery, or implicit layer injection occurs. Store body hashes are integrity checks; a durable row whose hash no longer matches its body fails closed. Never put credentials or provider clients in prompt metadata.
98
+
99
+ Threat model: the registry is **host-trusted data**. Anyone who can write versions into the store is inside the trust boundary — `put`, label management, and `assertPromptPromotion` verdicts are host operations, never agent-reachable surfaces. Untrusted prompt-injection defense stays at Prism's existing untrusted-content boundaries (tool results, attachments, and provider output), which the store neither bypasses nor weakens: a resolved body enters the system-prompt layer exactly like a host-authored constant. The optional `@arnilo/prism-evals` peer is only loaded by `assertPromptPromotion` and never makes the store itself depend on evaluation infrastructure.
100
+
101
+ ## Related APIs
102
+
103
+ - [System prompts](system-prompts.md): existing explicit layering and file adapters.
104
+ - [Input and prompt assembly](input-and-prompt-assembly.md): host-controlled message/context assembly.
105
+ - [Evaluations](evaluations.md): bounded evaluation primitives; `assertPromptPromotion` composes `runComparison` + `assertEvaluationThreshold`.
106
+ - [Database persistence](database-persistence.md): persistence and ownership conventions.