@arnilo/prism 0.0.7 → 0.0.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,174 @@
1
+ # Review coverage — 2026-07-19 Phase 3
2
+
3
+ Working evidence for Plan 070 Task 0. This page freezes Phase 3’s source revision, supported capability boundary, shared-primitive decision, finite-limit targets, and release evidence before public implementation begins.
4
+
5
+ **Evidence frozen:** 2026-07-20. **Prism source:** `6048e82db212303f4f072ff70539830b779f35cf` (`Phase 0.0.7 completed`). **Default test rule:** all default tests use local fakes; a credential-gated live suite is separate.
6
+
7
+ ## Status legend
8
+
9
+ | Status | Meaning |
10
+ | --- | --- |
11
+ | `existing` | Current contract already covers this part. |
12
+ | `extend` | Owning task extends an existing optional package/primitive. |
13
+ | `new-package` | New optional package only; core stays unchanged unless a later two-consumer review proves a generic gap. |
14
+ | `unsupported` | Deliberate 0.0.8 absence. Return a stable explicit error when a declared protocol operation is unavailable. |
15
+
16
+ ## Frozen external references
17
+
18
+ | Surface | Pinned reference | Compatibility decision |
19
+ | --- | --- | --- |
20
+ | OTel GenAI | [`semantic-conventions-genai@c26a2c21d1ee70d5231bd440c7b48d3c94ee506a`](https://github.com/open-telemetry/semantic-conventions-genai/tree/c26a2c21d1ee70d5231bd440c7b48d3c94ee506a/docs/gen-ai) | Development-status spans/metrics/events are adopted only where Prism has source data. Content attributes and evaluation explanations stay off by default. |
21
+ | MCP specification | [2025-11-25](https://modelcontextprotocol.io/specification/2025-11-25) | JSON-RPC over stdio and Streamable HTTP only. Capability is declared only after SDK compatibility tests. |
22
+ | MCP TypeScript SDK | [`@modelcontextprotocol/sdk@1.29.0`](https://github.com/modelcontextprotocol/typescript-sdk/tree/v1.29.0), lock integrity `sha512-zo37mZA9hJWpULgkRpowewez1y6ML5GsXJPY8FI0tBBCd77HEvza4jDqRKOXgHNn867PVGCyTdzqpz0izu5ZjQ==` | Current bridge is 1.29.0. Context7 verified client capability declaration, roots handlers, server capability inspection, and Streamable HTTP session construction. Task 4 pins the manifest to the tested SDK version; no untested upgrade. |
23
+ | A2A | [A2A v1.0.0 `173695755607e884aa9acf8ce4feed90e32727a1`](https://github.com/a2aproject/A2A/tree/173695755607e884aa9acf8ce4feed90e32727a1) and [specification](https://a2a-protocol.org/v1.0.0/specification) | JSON-RPC/HTTPS only. Tasks, subscription, rich parts, and optional push hooks are added behind host lifecycle/auth adapters. |
24
+ | Exa | [Search](https://exa.ai/docs/reference/search), [contents retrieval](https://exa.ai/docs/reference/contents-retrieval), retrieved 2026-07-20 | Direct API only, `POST /search` with requested contents. Docs have no immutable revision, so URL plus retrieval date and adapter request/response fixtures are release evidence. |
25
+ | Brave | [Web Search `GET /res/v1/web/search`](https://api-dashboard.search.brave.com/api-reference/web/search/get), [versioning](https://api-dashboard.search.brave.com/documentation/guides/versioning), retrieved 2026-07-20 | Direct API only. Prism limits results to 20, matching documented maximum; token is late-bound `X-Subscription-Token`. |
26
+ | Firecrawl | [v2 introduction](https://docs.firecrawl.dev/api-reference/v2-introduction), [search](https://docs.firecrawl.dev/api-reference/endpoint/search), [scrape](https://docs.firecrawl.dev/api-reference/endpoint/scrape), [extract](https://docs.firecrawl.dev/api-reference/endpoint/extract), retrieved 2026-07-20 | Direct v2 API only: `/v2/search`, `/v2/scrape`, `/v2/extract`. Markdown and returned schema data are untrusted. |
27
+ | Release security | [Dependency review](https://docs.github.com/en/code-security/supply-chain-security/understanding-your-software-supply-chain/about-dependency-review), [SBOM](https://docs.github.com/en/code-security/supply-chain-security/understanding-your-software-supply-chain/about-software-bill-of-materials), [secret scanning](https://docs.github.com/en/code-security/secret-scanning/introduction/about-secret-scanning), [artifact attestations](https://docs.github.com/en/actions/security-for-github-actions/security-guides/using-artifact-attestations-to-establish-provenance-for-builds) | Prefer GitHub-native gates, immutable action revisions, minimal tokens, and protected-environment canaries. |
28
+
29
+ ### External behavior relied on
30
+
31
+ - GenAI inference and tool spans are `CLIENT` and `INTERNAL` respectively; `gen_ai.evaluation.result` is an event. Inputs, outputs, system instructions, and tool definitions are opt-in content fields, not default telemetry.
32
+ - MCP 1.29.0 clients declare capabilities at construction; undeclared use is rejected by SDK. Roots use `roots/list` plus optional `notifications/roots/list_changed`; server capabilities are inspected after connect.
33
+ - A2A v1.0 defines task get/list/cancel/subscribe, `text`/`raw`/`url`/`data` parts, ordered task updates, authenticated push configuration, and `A2A-Version` negotiation.
34
+ - Brave documents a 400-character/50-word query maximum, `count <= 20`, and `offset <= 9`; Prism never expands those provider maxima.
35
+ - Firecrawl documents `search.limit <= 100`, `query <= 500` characters and bearer authentication. Prism uses tighter portable defaults and validates returned Markdown/JSON itself.
36
+
37
+ ## Capability traceability matrix
38
+
39
+ | Roadmap criterion | Current surface | Minimum 0.0.8 gap | Owner | Required proof | Docs | Release gate |
40
+ | --- | --- | --- | --- | --- | --- | --- |
41
+ | Parent OTel agent/provider/tool hierarchy and context | Event-derived independent `prism.*` spans in `packages/observability-opentelemetry` | Active parent/context lifecycle; semantic map; trace reference | 1 | parent/error/detach/exporter-failure matrix | `observability.md`, `agent-events.md` | offline OTel conformance |
42
+ | Metadata-safe GenAI/MCP telemetry and low-cardinality metrics | `AgentEvent`, redactor, safe counters | Convention mapping and label allow-list | 1 | secret/content/ID label-negative tests | `observability.md`, `host-security.md` | tarball/secret scan |
43
+ | Final-result/trace grading, judges, pairwise, thresholds | `@arnilo/prism-evals` result scorers/datasets/experiments | Bounded owner-scoped trace resolver, optional judge, report/gate | 2 | fake judge, trace cap, deterministic threshold failure | `evaluations.md`, `runs-and-usage.md` | offline eval gate |
44
+ | Batched ledger, snapshot cache, benchmarks | Serialized write-through `RunLedger`; repeated session rebuild | Optional flush/ack batch wrapper and leaf cache | 3 | order/crash/terminal flush/cache-read count benchmarks | `runs-and-usage.md`, `performance.md`, persistence pages | offline persistence + published bench |
45
+ | MCP tools/resources/prompts/roots/sampling/elicitation/notifications | Tool bridge, tool-list notification, server tools, stateless Streamable HTTP | SDK-pinned capability facades/session handling | 4 | fake SDK/server capability, session, origin, overflow matrix | `mcp-tools.md`, `resource-loading.md` | offline MCP conformance |
46
+ | Host-owned MCP OAuth/auth | `resolveAuthInfo`, authorizer, exact-origin pinned client transport | Exact request/session/ownership binding and host OAuth callback | 4 | wrong origin/session/owner/token-redaction cases | `mcp-tools.md`, `credential-storage.md`, `host-security.md` | secret scan + MCP conformance |
47
+ | A2A durable task/cancel/reconnect/rich parts/push | Text-only immediate session execution, bounded SSE, card verification | Host task lifecycle adapter, part mapping, bounded replay/push | 5 | fake lifecycle/card/task/reconnect/push matrix | `a2a.md`, `agent-session-runtime.md`, `workflows.md` | offline A2A conformance |
48
+ | Narrow web search/fetch/extract package | Tool dispatch, credentials, media SSRF bounds, bounded response reader | Optional direct-adapter package and normalized public result types | 6 | fake Brave/Exa/Firecrawl normalization/overflow/security matrix | `web-tools.md`, `tools.md` | pack/install + offline conformance |
49
+ | Citation identity and schema-validated untrusted external results | `ToolResult`, `ContentBlock`, resource/media bounds | Package-local citation/document/extraction contracts | 6 | citation/schema/prompt-injection fixtures | `web-tools.md`, `host-security.md` | offline web conformance |
50
+ | Official Exa/Firecrawl MCP prototype guidance only | Existing hardened MCP bridge | Documentation-only prototype boundary | 4, 6 | docs assertion: no generic remote passthrough | `mcp-tools.md`, `web-tools.md` | docs review |
51
+ | SAST/dependency/secret/SBOM/license/attestation/updates/live canaries | `release.yml` readiness, pack, provenance publish | Native GitHub security workflows and protected live jobs | 7 | workflow policy/negative secret/license/SBOM fixtures | `release-and-install.md`, `host-security.md` | required PR/release jobs |
52
+ | Reusable conformance and bounded performance evidence | Provider/session/tool/ledger helpers and 0.0.7 benchmark tables | Package-local protocol/web helpers plus Phase 3 benchmark script | 1–7 | each fake suite and dated benchmark output | `performance.md` | `sdk:ready` + documented live prerequisites |
53
+
54
+ ## Primitive and caller inventory
55
+
56
+ | Primitive | Existing contract and callers | Phase 3 disposition |
57
+ | --- | --- | --- |
58
+ | `AgentEvent` | `src/agents.ts` emits provider/tool/guardrail/limit events; subscribers, `RunLedger`, OTel adapter consume them | Reuse as telemetry source. Task 1 adds parent/context only; no second event bus. |
59
+ | `RunLedger` / `RunLedgerRecord` | Four ordered append methods; `RuntimeAgentSession.ledgerChain` serializes appends; SQLite/PostgreSQL implement it | Task 3 may add one optional wrapper with flush/ack. Default remains write-through; no persistence schema. |
60
+ | `ProductionPersistenceStore` | Cursor/ownership-scoped `queryRuns`, `queryEvents`, `queryToolCalls`, `queryUsage`; optional checkpoints/leases/feedback | Task 2 reads a bounded host-selected trace. Task 5 adapts host lifecycle/checkpoints. No generic task table. |
61
+ | `RunLimits` / `RunLimitTracker` | Runtime/provider/tool limits charge before provider/tool work | Reuse for run work. Package-specific external operations get package-local finite limits until a second equal consumer exists. |
62
+ | Guardrails/redaction/permission/trust | `runGuardrails`, `dispatchToolCall`, `redact*`, `assertPermission`, `assertTrusted` | All tool/protocol/web outputs route through existing dispatch/redaction. No protocol-specific security bypass. |
63
+ | Durable agent/workflow state | `CheckpointStore`, `AgentRunLifecycle`, durable agent run state, workflow status/cancel | Task 5 projects host-selected lifecycle into A2A. No new worker/queue/runtime. |
64
+ | `ResourceLoader` | `loadTextResource`, `loadJsonResource`, `loadBinaryResource`; permission/trust context | Task 4 maps resources only through a narrow adapter; Task 6 does not pretend remote web fetch is a generic resource loader. |
65
+ | `CredentialResolver` | `resolveCredentialValue`, explicit/chained/env resolvers, OAuth store | Tasks 4–6 resolve at adapter request edge. Credentials never enter tool schemas/results/telemetry. |
66
+ | Media URL/SSRF primitives | `assertSsrfAllowedUrl`, address-pinned `requestUrl`, media MIME/byte/time bounds | Reuse policy/validation semantics. Generalize pinned transport only if Task 6 proves identical requirements with media and MCP; otherwise package-local provider fetch. |
67
+ | Provider transport | `readSseEvents`, `readBoundedResponseText`, bounded JSON argument parsing | Reuse bounded response/redaction semantics where shape fits. Do not force non-SSE vendor APIs through provider types. |
68
+ | Tool dispatch | `ToolDefinition`, registry, schema validation, permissions, guardrails, ledger, limit tracker | Web tools and mapped MCP tools use this path; models never choose adapter/provider credentials. |
69
+ | OTel adapter | Optional event projection with in-memory fixture and exporter-error isolation | Extend in place in Task 1; OTel API remains optional-package only. |
70
+ | Evals package | Immutable dataset, function scorer, bounded experiment pool, package-local store | Extend in place in Task 2; judge stays host callback, not core/provider dependency. |
71
+ | MCP package | Official SDK bridge, bounded tool discovery/result conversion, DNS-pinned Streamable HTTP, authorized server helper | Extend in place in Task 4; capability facades remain package-local. |
72
+ | Supervisor A2A package | Card/JWS, bounded JSON-RPC/SSE, exact-origin client | Extend in place in Task 5; task/push adapters remain package-local. |
73
+
74
+ ### Core primitive decision
75
+
76
+ No new core primitive is authorized by Task 0. The only conditional candidates are:
77
+
78
+ 1. **Trace-context carrier:** Task 1 may add one only if OTel, supervisor delegation, and another core consumer require the same dependency-free propagation shape.
79
+ 2. **Flushable ledger capability:** Task 3 may add one only if one adapter wraps memory, SQLite, PostgreSQL, and host ledgers without changing `RunLedger` ordering.
80
+ 3. **Pinned bounded HTTP request:** Task 6 may extract one only if media and web adapters use identical DNS-pinning, redirect, byte, abort, and error-redaction semantics.
81
+
82
+ Otherwise code stays in its optional package. One-consumer interfaces, vendor types, protocol vocabularies, and task queues are rejected.
83
+
84
+ ## Frozen capability boundary
85
+
86
+ | Surface | Supported in 0.0.8 | Explicitly unsupported / deferred |
87
+ | --- | --- | --- |
88
+ | Telemetry | Optional OTel API adapter; agent/inference/tool/guardrail/delegation hierarchy; safe traces/metrics/evaluation events | Exporter registration, hosted backend, default content capture, per-delta spans, ID/content metric labels |
89
+ | Evaluations | Function scorers, bounded trace target, host-supplied judge, pairwise comparison, datasets, reports, threshold assertion | Mandatory LLM/provider, evaluation service/database, automatic production grading |
90
+ | MCP | Pinned-SDK tools/resources/prompts/roots/sampling/elicitation/notifications and Streamable HTTP only when SDK test proves each | Server discovery, arbitrary remote capability proxy, raw SDK callback/tool exposure, automatic OAuth/token forwarding |
91
+ | A2A | v1.0 JSON-RPC/HTTPS card/task/status/list/cancel/subscribe, rich parts, bounded replay, optional host push hook | gRPC, REST/HTTP+JSON binding, endpoint discovery, JWK fetching, second durable engine |
92
+ | Web tools | Host-selected Brave or Exa discovery; Firecrawl Markdown/schema extraction; direct native fetch adapters | Browser automation, vendor SDK dependency, model-selected provider, generic web/MCP passthrough |
93
+ | Release | Native CI security gates and scheduled/manual protected canaries | Secrets in PR/default jobs, public-network default tests, hosted security service |
94
+
95
+ ## Frozen limits and charging points
96
+
97
+ Values in this table are target defaults/hard caps for later tasks, not active 0.0.7 APIs. Every count/byte/time/concurrency check happens before retaining data or starting the next request. Existing stricter package limits remain authoritative until changed with tests.
98
+
99
+ | Surface | Default / hard cap | Charge before | Owner |
100
+ | --- | --- | --- | --- |
101
+ | OTel active spans | one agent span per run; provider/tool/guardrail/delegation only from existing bounded run work | `startSpan`; no span for deltas | 1 |
102
+ | OTel content buffering | `0 / 0` by default | copying any prompt/content | 1 |
103
+ | Eval trace pages | `20 / 100` per record kind | next persistence page | 2 |
104
+ | Eval trace records | `1,000 / 5,000` events, tool calls, and usage records each | append snapshot item | 2 |
105
+ | Eval judge request/response | `256 KiB / 1 MiB` each | serialization/body retention | 2 |
106
+ | Eval judge attempts/time | `1 / 3`; `60 s / 30 min` | request attempt/timer | 2 |
107
+ | Eval worker concurrency | `1 / 32` (existing) | start worker | 2 |
108
+ | Ledger batch entries/bytes | `128 / 4,096`; `512 KiB / 16 MiB` | enqueue | 3 |
109
+ | Ledger batch delay/in-flight flushes | `25 ms / 1 s`; `1 / 8` | timer/flush dispatch | 3 |
110
+ | Snapshot cache | one current leaf snapshot per active session/run; no cross-session cache | store/reuse snapshot | 3 |
111
+ | MCP existing tool bridge | existing `20/100` pages, `500/5,000` tools, `4 MiB/16 MiB` aggregate schemas, `10 MB/16 MiB` result | page/schema/result retention | 4 |
112
+ | MCP resources/prompts | `20 / 100` pages; `500 / 5,000` items; `1 MiB / 8 MiB` item content | page/item conversion | 4 |
113
+ | MCP roots/sampling/elicitation | `32 / 128` roots; `32 / 128` messages; `64 KiB / 1 MiB` arguments/schema | callback/request dispatch | 4 |
114
+ | MCP HTTP sessions/replay | `128 / 1,024` active sessions; `1,024 / 10,000` replay events; `64 KiB / 1 MiB` event | session/replay allocation | 4 |
115
+ | A2A existing request/response/stream | `64 KiB/1 MiB` request/response; `64 KiB/1 MiB` event; `10 MiB/64 MiB` stream; `10k/100k` events | body/frame/event retention | 5 |
116
+ | A2A task pages/parts/artifacts | `100 / 1,000` tasks; `32 / 256` parts per message/artifact; `1 MiB / 8 MiB` raw/data part; `8 MiB / 64 MiB` aggregate artifacts | parse/decode/store/replay | 5 |
117
+ | A2A push | `32 / 256` registrations; `3 / 10` attempts; `30 s / 5 min` delivery | persist/deliver/retry | 5 |
118
+ | Web query/results/URLs | `4 KiB / 16 KiB`; `10 / 20` results; `5 / 20` URLs | request construction | 6 |
119
+ | Web provider request/output | `256 KiB / 1 MiB` request; `2 MiB / 16 MiB` response; `1 MiB / 8 MiB` Markdown; `256 KiB / 1 MiB` extracted JSON | request/body/output retention | 6 |
120
+ | Web schema/concurrency/time | `64 KiB / 256 KiB` schema; `4 / 16` active calls; `60 s / 30 min`; `2 / 4` retries | schema compile/call/retry | 6 |
121
+ | CI/live canaries | default suite `0` remote calls; protected live job one bounded scenario/provider | credential resolution/network call | 7 |
122
+
123
+ ### Network and credential rules
124
+
125
+ 1. Provider API origins are exact allow-lists and provider API redirects are rejected. Native `fetch` uses `redirect: "error"` or equivalent pinned transport.
126
+ 2. A fetched target URL is absolute HTTP(S), has no userinfo, and passes the host’s public/private policy before sending it to Firecrawl. Firecrawl-side redirects are provider behavior; Prism does not claim DNS pinning after handing a URL to Firecrawl. Hosts that need that guarantee use a controlled fetch adapter instead.
127
+ 3. MCP client requests pin a validated DNS address and reject redirects today; Task 4 extends session/auth behavior without weakening that boundary.
128
+ 4. A2A and MCP authorization run on every request and bind exact origin, session, and ownership. Missing/foreign resources/tasks resolve as authorized-not-found, not disclosure.
129
+ 5. Credentials are resolved immediately before adapter I/O, redacted in errors before ledger/export, excluded from tool inputs/outputs/telemetry, and never forwarded automatically between MCP/A2A/provider/web surfaces.
130
+ 6. Search snippets, Markdown, HTML, extracted JSON, A2A artifacts, MCP resources/prompts, and remote protocol errors are untrusted data. They cannot modify system instructions, tools, permissions, credential selection, or provider routing.
131
+
132
+ ## Web normalization and external-operation ownership
133
+
134
+ | Tool / adapter | Host-selected request | Normalized public output | Credential owner and redaction point |
135
+ | --- | --- | --- | --- |
136
+ | `web_search` / Exa | `POST https://api.exa.ai/search`; query plus explicitly requested bounded `contents` only | `title`, canonical `url`, bounded `snippet`/`highlights`, provider result ID, publication/retrieval time, provider/cost/rate metadata, citation identity | Adapter resolves `{ provider: "exa", name: "api_key" }` at request edge; redact key from headers, response/error, telemetry, and `ToolResult`. |
137
+ | `web_search` / Brave | `GET https://api.search.brave.com/res/v1/web/search`; query/count/offset only | Same normalized search result; source fields missing from Brave stay absent, never guessed | Adapter resolves `{ provider: "brave", name: "subscription_token" }` at request edge; redact `X-Subscription-Token` and all error echoes. |
138
+ | `web_fetch` / Firecrawl | `POST https://api.firecrawl.dev/v2/scrape`; one prevalidated public URL and Markdown format | `url`, canonical/source URL, bounded Markdown, selected metadata, retrieval time, provider/cost/rate metadata, `untrusted: true`, citation identity | Adapter resolves `{ provider: "firecrawl", name: "api_key" }` at request edge; redact bearer header/body/error echoes. |
139
+ | `web_extract` / Firecrawl | `POST https://api.firecrawl.dev/v2/extract`; bounded URL list plus host-supplied JSON Schema | Same document attribution plus schema-validated bounded JSON value and `untrusted: true` | Same Firecrawl resolver/redaction rule; schema/value never influence tool permissions or system instructions. |
140
+
141
+ `citationId` is deterministic: `web:<provider>:<providerResultId>` when provider returns a stable result ID, otherwise `web:<provider>:sha256(<canonicalUrl>)`. `canonicalUrl` removes fragment and normalizes only URL syntax; Prism does not follow target redirects to manufacture identity. Provider-specific fields remain under bounded metadata and never replace normalized fields.
142
+
143
+ | External operation | Authorization/credential owner | Required boundary |
144
+ | --- | --- | --- |
145
+ | OTel export | Host exporter SDK; Prism adapter receives tracer/meter only | Exporter failure isolated; Prism receives no exporter credential. |
146
+ | Model judge | Host-supplied judge callback | Judge receives bounded redacted evaluation target; no resolver/tools/workspace. |
147
+ | MCP HTTP/OAuth | Host auth resolver; Task 4 session binding | Every request binds exact origin/session/ownership; no token forwarding to model/tool content. |
148
+ | A2A invoke/card/push | Host authorizer/client auth/push delivery callback | Every task/push action is owner-scoped; card keys are explicitly pinned. |
149
+ | Web adapters | Host `CredentialResolver` or explicit callback | Exact provider API origin, no provider API redirects, late credential resolution. |
150
+ | Live canary | Protected CI environment only | Least-privilege key, redacted aggregate report, no PR/default-job secret. |
151
+
152
+ ## Required test evidence by task
153
+
154
+ | Task | Network-free evidence | Restricted live evidence |
155
+ | --- | --- | --- |
156
+ | 1 | Parent graph, context, cleanup, semantic attributes, no-content/no-ID-label, exporter isolation | Host OTel exporter smoke only if configured |
157
+ | 2 | Fake trace reader/judge, redaction/cap/threshold/pairwise deterministic reports | Explicit host judge/model smoke |
158
+ | 3 | Fake ledger order/flush/crash/cache invalidation; reproducible local benchmark | PostgreSQL benchmark when service exists |
159
+ | 4 | Fake SDK/server capabilities, roots/resources/prompts/sampling/elicitation, sessions, auth/origin/SSRF/overflow | Configured MCP endpoint smoke |
160
+ | 5 | Fake lifecycle/card/push, rich parts, cancel/reconnect/replay/owner isolation | Configured A2A endpoint smoke |
161
+ | 6 | Fake Brave/Exa/Firecrawl normalization, schema, redirect/SSRF, credential/prompt-injection fixtures | Protected least-privilege Exa/Brave/Firecrawl smoke |
162
+ | 7 | Workflow policy, SBOM/license/attestation, secret-negative and skipped-canary tests | Scheduled/manual protected credentials only |
163
+
164
+ ## Release evidence checklist
165
+
166
+ - `npm run sdk:ready`, Node 20/current import check, every workspace pack, packed offline consumer, `npm audit --audit-level=high`, tarball deny-list/secret check, and `git diff --check`.
167
+ - Focused OTel/eval/ledger/MCP/A2A/web fake-server conformance suites pass without public network.
168
+ - SAST, dependency review, secret scanning, SPDX SBOM/license policy, and artifact attestation jobs pass with pinned actions/minimal permissions.
169
+ - Protected live canaries either pass with least-privilege credentials or are recorded as an explicit release-host prerequisite; skipped canaries never make the default suite appear live-tested.
170
+ - Benchmarks publish machine/runtime, workload, p50/p95, throughput, memory, disk, cost metadata when supplied, backpressure observations, and no portability claim.
171
+
172
+ ## Current exclusions
173
+
174
+ No browser, coding, Office, SaaS connector, hosted observability, generic proxy, provider SDK, remote discovery, automatic token forwarding, background task worker, or new core persistence schema is authorized by this phase. Those belong to later roadmap phases or require a new evidence review.
@@ -80,6 +80,7 @@ await assertRunLedgerConforms({ ledger, readRuns: () => runs, readEvents: () =>
80
80
 
81
81
  - `read*` callbacks are optional for smoke tests but required for ordering, tenant, and reopen probes.
82
82
  - `RunLedger` is write-only from Prism's perspective; replay/query APIs live on `ProductionPersistenceStore` or host-owned reads.
83
+ - Run conformance against the underlying write-through adapter. Test `createBatchedRunLedger()` separately for FIFO, bounds, terminal/manual flush, retained failure, and documented buffered crash loss; batching does not weaken adapter conformance.
83
84
  - Redaction is not asserted here — the runtime calls `redactRunLedgerRecord()` before writes when a `SecretRedactor` is active.
84
85
 
85
86
  ## Security and performance notes
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. `RunFeedbackStore` is the separate post-run seam for immutable ratings, comments, tags, and evaluation links. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
5
+ `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. `createBatchedRunLedger()` is an explicit optional wrapper; direct ledger writes remain default. `RunFeedbackStore` is the separate post-run seam for immutable ratings, comments, tags, and evaluation links. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
6
6
 
7
7
  APIs:
8
8
 
@@ -152,7 +152,7 @@ const page = await feedback.query({ runId: result.runId, tenantId: "t1", userId:
152
152
  await feedback.delete({ id: "fb_1", tenantId: "t1", userId: "u1" });
153
153
  ```
154
154
 
155
- Default/hard bounds: comment 4/16 KiB, tags 16/64, scorer/evaluation IDs 16/64 each, metadata 16/64 KiB, query page 100/500; tags are 64 characters and identifiers 128. The store redacts comment/tags/metadata after run ownership validation and before persistence. IDs are linked, not scorer payloads. `ProductionPersistenceStore.feedback?` exposes this capability; first-party SQLite/PostgreSQL adapters implement it in schema migration `003_run_feedback` and reject missing/cross-owned runs.
155
+ Default/hard bounds: comment 4/16 KiB, tags 16/64, scorer/evaluation IDs 16/64 each, metadata 16/64 KiB, query page 100/500; tags are 64 characters and identifiers 128. `@arnilo/prism-evals` may read `queryRuns/queryEvents/queryToolCalls/queryUsage` only through an explicit owner/session/run-scoped trace resolver with finite cursor pages and aggregate bytes. The store redacts comment/tags/metadata after run ownership validation and before persistence. IDs are linked, not scorer payloads. `ProductionPersistenceStore.feedback?` exposes this capability; first-party SQLite/PostgreSQL adapters implement it in schema migration `003_run_feedback` and reject missing/cross-owned runs.
156
156
 
157
157
  ## Status transitions
158
158
 
@@ -287,6 +287,21 @@ console.log(cacheUsageReport(aggregate?.usage));
287
287
  - **Tenant isolation.** `OwnershipScope` fields are copied from the active ownership scope, but the runtime does not enforce tenant isolation for ledger rows. Feedback is stricter: append/query/delete require tenant plus account/user, and first-party stores compare the exact scope to the linked run.
288
288
  - **Feedback privacy.** Comments/tags/metadata can contain PII. Configure a feedback redactor, apply retention, and call owned `delete()` for erasure. Never copy comments or tag values into metric labels.
289
289
 
290
+ ## Optional batching and durability
291
+
292
+ ```ts
293
+ const ledger = createBatchedRunLedger(store, {
294
+ maxBatchEntries: 128,
295
+ maxBatchBytes: 512 * 1024,
296
+ maxDelayMs: 25,
297
+ durability: "flush_on_terminal",
298
+ });
299
+ ```
300
+
301
+ Modes: `write_through` acknowledges each target write; `flush_on_terminal` buffers but runtime awaits terminal flush; `buffered` acknowledges enqueue only and requires host `flush()` for durability. `status()` distinguishes accepted/flushed/buffered counts. Defaults/hard caps: 128/4,096 batch entries, 512 KiB/8 MiB batch bytes, 25 ms/60 s delay; buffered count/bytes apply backpressure before enqueue. FIFO spans all record kinds. Inputs are already runtime-redacted. Flush errors propagate and retain the failing record for retry. `dispose({ flush: false })` clears memory but deliberately loses unflushed records—same crash-before-flush ceiling as process failure.
302
+
303
+ Runtime session snapshots cache one leaf/generation for at most one second. Successful append, compaction append, checkout, and durable resume invalidate; failed append does not advance leaf/cache. Cache is session-local and never shared across ownership/session/branch.
304
+
290
305
  ## Related APIs
291
306
 
292
307
  - [Performance limits](performance.md): batching, cursor keys, and production sizing assumptions.
@@ -108,6 +108,7 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
108
108
  - **File ownership.** Create database files on a host-controlled path with restrictive permissions (`0600` default on Unix via `fileMode`).
109
109
  - **No path interpolation.** The adapter opens exactly the caller-supplied `filename`; it does not expand environment variables or discover paths.
110
110
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
111
+ - **Optional batching.** SQLite remains write-through by default. Hosts may wrap its ledger with core `createBatchedRunLedger()`; `flush_on_terminal` preserves terminal acknowledgement, while `buffered` explicitly risks crash-before-flush loss.
111
112
  - **WAL + busy timeout.** WAL is enabled by default; busy timeout defaults to 5 seconds. This meets the Plan 056 local workload target but SQLite still serializes writers — prefer PostgreSQL for high write concurrency.
112
113
  - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans. Startup validation reads SQLite catalog/PRAGMA metadata only, never application rows.
113
114
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns on run and ownership tables participate in query filters; hosts must still scope writes correctly.
@@ -21,7 +21,7 @@ Use a supervisor when a host or agent must choose a child dynamically. Use `@arn
21
21
 
22
22
  ## Outputs / response / events
23
23
 
24
- `delegate()` returns the child's `AgentRunResult` or throws its `AgentRunError`/a supervisor denial or limit error. `subscribe()` emits bounded `delegation_started`, `delegation_finished`, `delegation_rejected`, and `delegation_error` metadata events.
24
+ `delegate()` returns the child's `AgentRunResult` or throws its `AgentRunError`/a supervisor denial or limit error. `subscribe()` emits bounded `delegation_started`, `delegation_finished`, `delegation_rejected`, and `delegation_error` metadata events. Hosts may project those events through observability `handleDelegation()` using the parent Prism run ID; no OpenTelemetry dependency enters this package.
25
25
 
26
26
  ## Request/response example
27
27
 
@@ -65,7 +65,7 @@ Child factories resolve their own providers/credentials and construct context/me
65
65
 
66
66
  ## Related APIs
67
67
 
68
- - [A2A interoperability](a2a.md): remote protocol boundary.
68
+ - [A2A interoperability](a2a.md): separate remote protocol boundary. `A2ATaskLifecycle` adapts host durable agent/workflow state directly; it does not route A2A execution through local supervisor child planning.
69
69
  - [Workflows](workflows.md): preferred deterministic orchestration.
70
70
  - [Working and semantic memory](working-and-semantic-memory.md): child scope construction.
71
71
  - [Host security](host-security.md): permission and credential boundaries.
package/docs/tools.md CHANGED
@@ -10,6 +10,7 @@ APIs:
10
10
  - `filterTools()`
11
11
  - `dispatchToolCall()`
12
12
  - `ToolFilter`, `ToolFilterInput`, `ToolValidator`, `DispatchToolCallOptions`
13
+ - Optional [Web search, fetch, and extraction](web-tools.md): three narrow host-selected `ToolDefinition`s with untrusted bounded outputs.
13
14
 
14
15
  ## When to use it
15
16
 
@@ -260,7 +261,7 @@ createJsonSchemaToolArgumentValidator({
260
261
  - [Credentials and redaction](credentials-and-redaction.md): redaction helpers used for tool execution errors.
261
262
  - [Observational memory compaction package](compaction-observational-memory.md): optional exact-id recall tool factory.
262
263
  - [Tool execution primitives](tool-execution-primitives.md): JSON Schema adapter, parallelism, MCP bridge, and execution-policy designs.
263
- - [MCP client bridge](mcp-tools.md): optional `@arnilo/prism-mcp` remote tool mapping.
264
+ - [MCP client bridge](mcp-tools.md): optional remote tool mapping plus separate bounded resource/prompt facades; non-tool MCP capabilities never bypass tool dispatch by masquerading as `ToolDefinition`.
264
265
  - [Coding agent tools](coding-agent-tools.md): optional first-party `@arnilo/prism-coding-agent` `shell`/`read`/`write`/`edit` tools a host registers into this harness.
265
266
 
266
267
  `DispatchToolCallOptions.trust` and `.permission` run before validation or `execute()`; denial emits `tool_execution_blocked`. Middleware cannot bypass either guard. `AgentConfig.validator`/`RunOptions.validate` run after these guards; their output is redacted through the active `SecretRedactor`. `createSecureAgent()` requires all three seams plus non-empty schemas and durable pre-tool approval. Prism does not sandbox tools. See [Security/auth/trust](settings-auth-trust-security.md).
@@ -0,0 +1,78 @@
1
+ # Web search, fetch, and extraction
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-web-tools` provides three separate bounded `ToolDefinition`s: host-selected Brave or Exa `web_search`, Firecrawl Markdown `web_fetch`, and Firecrawl JSON `web_extract`. Package uses native `fetch`; no vendor SDK or browser is installed.
6
+
7
+ ## When to use it
8
+
9
+ Use when agent needs explicit public-web discovery or host-approved document retrieval/extraction. Keep search separate from fetch/extract so model cannot select provider, credential, API origin, extraction schema, or cost path.
10
+
11
+ ## Inputs / request
12
+
13
+ | Tool | Model-visible input | Host-only construction input |
14
+ | --- | --- | --- |
15
+ | `web_search` | `query`, optional `count` | exactly one `createBraveSearch()` or `createExaSearch()`, credential source, limits |
16
+ | `web_fetch` | one absolute public `url` | `createFirecrawlFetch()`, credential source, SSRF/DNS policy, limits |
17
+ | `web_extract` | bounded `urls` array | `createFirecrawlExtractor()`, fixed JSON Schema, validator, credential source, SSRF/DNS policy, limits |
18
+
19
+ Credentials may be explicit callbacks or core `CredentialResolver`s. They resolve at request edge as Brave `subscription_token`, Exa `api_key`, or Firecrawl `api_key`. `allowedOrigins` must contain fixed provider API origin; adapters reject redirects. Firecrawl target URLs reject userinfo, non-HTTP(S), private literals, and hostnames denied by `SsrfPolicy`; supply `validateUrl` for host DNS/rebinding/egress checks before handing URL to Firecrawl.
20
+
21
+ ## Outputs / response / events
22
+
23
+ Search results retain bounded `title`, canonical fragment-free `url`, `snippet`, highlights, provider result ID, publication/retrieval time, returned request/cost/rate facts, and stable citation identity. Citation is `web:<provider>:<sourceId>` when provider supplies ID, otherwise SHA-256 of canonical URL.
24
+
25
+ Fetch returns bounded Markdown and selected attribution. Extract validates host schema before I/O and validates returned JSON through host `ToolArgumentValidator`. Every result has `untrusted: true`; tool result metadata is `trust: "untrusted_external"`. Missing provider facts remain absent—Prism never guesses billing or freshness.
26
+
27
+ ## Request/response example
28
+
29
+ ```json
30
+ {
31
+ "tool": "web_search",
32
+ "arguments": { "query": "Prism TypeScript SDK", "count": 5 },
33
+ "result": {
34
+ "provider": "brave",
35
+ "untrusted": true,
36
+ "results": [{ "citationId": "web:brave:…", "url": "https://example.com/", "title": "Example" }]
37
+ }
38
+ }
39
+ ```
40
+
41
+ ## Implementation example
42
+
43
+ ```ts
44
+ import { createEnvCredentialResolver } from "@arnilo/prism";
45
+ import { createJsonSchemaArgumentValidator } from "@arnilo/prism-tool-validator-json-schema";
46
+ import { createBraveSearch, createFirecrawlExtractor, createFirecrawlFetch, createWebTools } from "@arnilo/prism-web-tools";
47
+
48
+ const credentials = createEnvCredentialResolver(process.env, {
49
+ "brave:subscription_token": "BRAVE_SEARCH_TOKEN",
50
+ "firecrawl:api_key": "FIRECRAWL_API_KEY",
51
+ });
52
+ const schema = { type: "object", properties: { title: { type: "string" } }, required: ["title"], additionalProperties: false };
53
+ const tools = createWebTools({
54
+ search: createBraveSearch({ credentials }),
55
+ fetch: createFirecrawlFetch({ credentials, validateUrl: publicDnsPolicy }),
56
+ extract: createFirecrawlExtractor({ credentials, schema, validator: createJsonSchemaArgumentValidator(), validateUrl: publicDnsPolicy }),
57
+ });
58
+ ```
59
+
60
+ ## Extension and configuration notes
61
+
62
+ Hosts substitute `createExaSearch()` for Brave; no runtime/model routing exists. Root export includes all adapters; `./brave`, `./exa`, and `./firecrawl` subpaths support atomic imports. Official vendor MCP servers are prototypes only: use hardened MCP origin/auth/capability policy, never generic remote passthrough.
63
+
64
+ Default/hard limits: query 4/16 KiB; results 10/20; URLs 5/20; request 256 KiB/1 MiB; response and aggregate 2/16 MiB; Markdown 1/8 MiB; extraction 256 KiB/1 MiB; schema 64/256 KiB; JSON depth 64/128 and properties 10k/100k; retries 2/4; rate delay 5/60 seconds; concurrency 4/16 active plus the same bounded waiting queue; polling 20/100; wall time 60 seconds/30 minutes. Hosts may only narrow or raise within hard caps.
65
+
66
+ ## Security and performance notes
67
+
68
+ Provider credentials never enter tool schemas/results, prompts, telemetry, URLs, or errors. Error text excludes remote bodies. Search snippets, Markdown, and extracted JSON are prompt-injection-capable data: never concatenate them into system instructions or use them to modify tools, permissions, credentials, trust, routing, or schemas. Firecrawl fetches target URLs remotely; Prism cannot claim target DNS pinning after handoff. Use controlled host fetch when that guarantee is required.
69
+
70
+ Default tests use injected fake fetch and make no public request. Restricted smoke: `PRISM_LIVE_WEB=1 npm run test:live -w @arnilo/prism-web-tools` plus least-privilege provider environment credential. Browser automation, arbitrary HTML execution, model-selected providers, automatic OAuth forwarding, and generic web/MCP passthrough are unsupported.
71
+
72
+ ## Related APIs
73
+
74
+ - [Tools](tools.md): registry, validation, permission, trust, guardrails, and ledger dispatch.
75
+ - [Credential storage](credential-storage.md): explicit resolver composition and environment mapping.
76
+ - [Host security](host-security.md): SSRF, untrusted-content, and secret boundaries.
77
+ - [MCP tools](mcp-tools.md): hardened prototype path for official vendor MCP servers.
78
+ - [Performance and resource limits](performance.md): operational ceilings and benchmark evidence.
package/docs/workflows.md CHANGED
@@ -290,6 +290,7 @@ Use workflows for known, durable, replayable graphs. Use optional supervisor del
290
290
  - [Agent/session runtime](agent-session-runtime.md): `AgentSession.run()`/`stream()`, abort, subscribe
291
291
  - [Guardrails](guardrails.md): `RunWorkflowOptions.guardrails` routes tool nodes through core dispatch before policy and side effects.
292
292
  - [Supervisor delegation](supervisors.md): bounded dynamic child selection.
293
+ - [A2A interoperability](a2a.md): hosts may adapt existing exact-owner workflow status/list/cancel/checkpoint/event surfaces to `A2ATaskLifecycle`; A2A package adds no workflow worker, queue, or schema.
293
294
  - [Agent events](agent-events.md): core `AgentEvent` wrapped by `agent_event`
294
295
  - [Session stores and branching](session-stores-and-branching.md): session `leafId` reuse on resume
295
296
  - [CLI/RPC](cli-rpc.md): host control seam; wire `createWorkflowCommands()` into `runRpcServer`
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@arnilo/prism",
3
- "version": "0.0.7",
3
+ "version": "0.0.8",
4
4
  "description": "Agent harness for AI providers, agents, sessions, and tools.",
5
5
  "type": "module",
6
6
  "main": "./dist/index.js",
@@ -118,6 +118,7 @@
118
118
  "packages/rag",
119
119
  "packages/server",
120
120
  "packages/supervisor",
121
+ "packages/web-tools",
121
122
  "packages/prism-*"
122
123
  ],
123
124
  "scripts": {