@arnilo/prism 0.0.7 → 0.0.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/CHANGELOG.md +50 -2
  2. package/README.md +3 -1
  3. package/dist/agent-loops.js +14 -8
  4. package/dist/agents.js +37 -3
  5. package/dist/contracts.d.ts +17 -0
  6. package/dist/index.d.ts +4 -2
  7. package/dist/index.js +3 -2
  8. package/dist/provider-events.d.ts +2 -0
  9. package/dist/provider-events.js +21 -13
  10. package/dist/providers/openai-compatible.js +8 -5
  11. package/dist/providers/transport.d.ts +10 -1
  12. package/dist/providers/transport.js +24 -8
  13. package/dist/run-ledger.d.ts +21 -0
  14. package/dist/run-ledger.js +115 -0
  15. package/dist/tools.js +2 -0
  16. package/docs/a2a.md +61 -42
  17. package/docs/agent-events.md +5 -4
  18. package/docs/agent-loops.md +2 -2
  19. package/docs/agent-session-runtime.md +1 -0
  20. package/docs/browser-automation.md +124 -0
  21. package/docs/coding-agent-tools.md +111 -14
  22. package/docs/coding-security.md +84 -11
  23. package/docs/credential-storage.md +9 -0
  24. package/docs/database-persistence.md +1 -1
  25. package/docs/evaluations.md +38 -4
  26. package/docs/guardrails.md +3 -2
  27. package/docs/host-security.md +29 -4
  28. package/docs/index.md +20 -15
  29. package/docs/mcp-tools.md +29 -5
  30. package/docs/migration.md +104 -0
  31. package/docs/observability.md +26 -14
  32. package/docs/performance.md +54 -0
  33. package/docs/postgres-persistence.md +1 -0
  34. package/docs/provider-conformance.md +1 -1
  35. package/docs/provider-primitives.md +7 -1
  36. package/docs/providers/kimi.md +16 -2
  37. package/docs/providers/opencode-go.md +43 -2
  38. package/docs/release-and-install.md +117 -62
  39. package/docs/resource-loading.md +4 -0
  40. package/docs/review-coverage-2026-07-19-phase-3.md +174 -0
  41. package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
  42. package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
  43. package/docs/run-ledger-conformance.md +1 -0
  44. package/docs/runs-and-usage.md +17 -2
  45. package/docs/sqlite-persistence.md +1 -0
  46. package/docs/structured-output.md +2 -2
  47. package/docs/supervisors.md +2 -2
  48. package/docs/tools.md +5 -1
  49. package/docs/web-tools.md +78 -0
  50. package/docs/workflows.md +2 -0
  51. package/package.json +6 -4
@@ -2,12 +2,17 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects.
5
+ `@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools and one disposable Docker/OCI sandbox reference. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects, and optionally contains untrusted coding work in a host-invoked container.
6
6
 
7
7
  | Export | Purpose |
8
8
  | --- | --- |
9
9
  | `createCodingApprovalPolicy(options)` | Returns an `ExecutionPolicy` with trusted roots, read-only mode, command allow/deny rules, approval caching, and timeout/abort-aware approval waits. |
10
10
  | `createSandboxBashOperations(adapter)` | Maps a host-owned `SandboxAdapter` to coding-agent `BashOperations` for delegated shell execution. |
11
+ | `createSandboxCodingComposition(cwd, options)` | Authoritative construction: returns `{ tools, composition }` with required `workspaceMode` (`"host"` \| `"sandbox"`), fail-closed mixed wiring, and containment metadata. |
12
+ | `createSandboxReadOnlyComposition(cwd, options)` | Same contract for read-only tools (`read`/`repo_list`/`repo_search`). |
13
+ | `createSandboxCodingTools` / `createSandboxReadOnlyTools` | Thin wrappers that return `tools` only (compat); still require `workspaceMode`. |
14
+ | `createSandboxFilesystemOperations` / `createSandboxRepositoryOperations` | Optional execFile-backed FS/list/search backends for a disposable sandbox tree. |
15
+ | `createDockerSandbox(options)` | Creates one disposable non-root Docker container with read-only root/source, bounded tmpfs workspace, typed `execFile`, import/export, and stop/kill/cleanup. |
11
16
  | `assertPathInsideRoots`, `isPathInsideReal` | Symlink-aware path containment helpers. |
12
17
  | `evaluateCommandRules`, `hasShellMetacharacters` | Command classification helpers. |
13
18
 
@@ -21,7 +26,7 @@ import type { ExecutionAction, ExecutionPolicy, ExecutionDecision } from "@arnil
21
26
 
22
27
  Use this package when coding tools need path scoping, human approval, command rules, or a pluggable sandbox backend. Wire the returned policy through `createCodingTools(cwd, { executionPolicy })` or per-tool `executionPolicy` options.
23
28
 
24
- Prism does **not** claim OS-level isolation unless the host provides a sandbox adapter. Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
29
+ Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
25
30
 
26
31
  ## Inputs / request
27
32
 
@@ -36,10 +41,39 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
36
41
 
37
42
  `run` caching keys decisions by the tool execution context's `runId`; `session` uses `sessionId`. Coding tools pass both identities to the policy. A missing/empty identity disables caching for that check rather than creating a global bucket. Identical actions in different runs/sessions never share approvals or denials.
38
43
 
44
+ ### Docker sandbox inputs
45
+
46
+ | Option | Default | Purpose |
47
+ | --- | --- | --- |
48
+ | `docker` | required | Absolute host Docker executable. |
49
+ | `image` | required | Digest-pinned image (`name@sha256:<64-hex>`). Never pulled (`--pull=never`). |
50
+ | `sourceRoot` | required | Absolute host directory imported into `/workspace`. |
51
+ | `user` | required | Non-root `uid:gid`. |
52
+ | `network` | `{ mode: "none" }` | Default no network; custom mode requires a pre-created network name and does not claim DNS containment. |
53
+ | `env` | `{}` | Exact allow-list only; host environment is never inherited. |
54
+ | `secrets` | `[]` | Canaries redacted from CLI/adapter errors. |
55
+ | `limits` | package defaults | CPU/memory/PID/FD/tmpfs/command/export/time caps validated before create. |
56
+
57
+ ### Workspace mode inputs (`createSandboxCodingComposition`)
58
+
59
+ | Option | Default | Purpose |
60
+ | --- | --- | --- |
61
+ | `workspaceMode` | **required** | `"host"` (all tools on host cwd; never claims containment) or `"sandbox"` (shell + FS/list/search share one disposable tree). |
62
+ | `sandbox` | optional in host; required for sandbox unless custom ops supplied | `SandboxAdapter` / `DisposableSandbox`. |
63
+ | `workspaceRoot` | `"/workspace"` in sandbox mode | Tree root used as tool cwd when sandbox backends are bound. |
64
+ | `allowMixedWorkspaceWiring` | `false` | Escape hatch: allow sandbox shell + host FS backends. Records `composition.warnings`; forces `containmentClaim: false`. Missing hatch throws. |
65
+ | `read`/`write`/`edit`/`repository.operations` | auto-wired from `DisposableSandbox` in sandbox mode | Host may supply custom tree backends instead of auto-wire. |
66
+
67
+ `0.0.9` silent split (sandbox shell + host FS) is **superseded**. Mixed wiring is never the default.
68
+
39
69
  ## Outputs / response / events
40
70
 
41
71
  `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations`, receive `onData(Buffer)` for ordered stdout/stderr forwarding through the shell tool's existing bounded accumulator, and never grant policy approval themselves.
42
72
 
73
+ `createSandboxCodingComposition()` returns `{ tools, composition }` where `SandboxCodingComposition` carries `workspaceMode`, `containmentClaim`, `mixedWiringAllowed`, `warnings`, `workspaceRoot`, and optional `treeIdentity` (from `importIdentity` / `lastExportIdentity`). `containmentClaim` is `true` only for sandbox mode with tree backends bound and mixed wiring denied. Host mode and escape-hatch mixed wiring always set `containmentClaim: false` — never treat host mode as contained execution.
74
+
75
+ `createDockerSandbox()` returns a `DisposableSandbox`: typed `execFile(file, args)`, shell-compatible `exec`, `status`, cooperative `stop`, forced `kill`, and idempotent `close`. Import may surface `importIdentity`; successful export updates `lastExportIdentity`. `close({ export })` can stream a bounded workspace tar plus SHA-256/entry/byte metadata through a host callback; checkpoints should retain only host artifact references/hashes, never whole workspaces.
76
+
43
77
  ## Request/response example
44
78
 
45
79
  ```json
@@ -52,8 +86,12 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
52
86
  ## Implementation example
53
87
 
54
88
  ```ts
55
- import { createCodingTools } from "@arnilo/prism-coding-agent";
56
- import { createCodingApprovalPolicy, createSandboxBashOperations } from "@arnilo/prism-coding-security";
89
+ import {
90
+ createCodingApprovalPolicy,
91
+ createDockerSandbox,
92
+ createSandboxCodingComposition,
93
+ } from "@arnilo/prism-coding-security";
94
+ import { createGitTools } from "@arnilo/prism-coding-agent";
57
95
 
58
96
  const policy = createCodingApprovalPolicy({
59
97
  roots: [workspaceRoot],
@@ -62,27 +100,62 @@ const policy = createCodingApprovalPolicy({
62
100
  approvalTimeoutMs: 60_000,
63
101
  });
64
102
 
65
- const tools = createCodingTools(workspaceRoot, {
103
+ const sandbox = await createDockerSandbox({
104
+ docker: "/usr/bin/docker",
105
+ image: "registry.example/prism-code@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
106
+ sourceRoot: "/srv/jobs/task-1/source",
107
+ user: "10001:10001",
108
+ network: { mode: "none" },
109
+ env: { CI: "1" },
110
+ limits: { cpus: 2, memoryBytes: 2 * 1024 ** 3, maxPids: 256, workspaceBytes: 1024 ** 3 },
111
+ });
112
+
113
+ // Sandbox mode: shell/read/write/edit/list/search share one disposable tree.
114
+ const { tools, composition } = createSandboxCodingComposition("/srv/jobs/task-1/source", {
115
+ workspaceMode: "sandbox",
116
+ sandbox,
66
117
  executionPolicy: policy,
67
- shell: {
68
- operations: createSandboxBashOperations(mySandboxAdapter),
69
- },
118
+ repository: { exclude: [".git", "node_modules", "dist"] },
119
+ });
120
+ // composition.containmentClaim === true when backends are bound
121
+
122
+ // Same-tree Git/check (opt-in; not folded into coding tools):
123
+ const gitTools = createGitTools(composition.workspaceRoot, {
124
+ execFile: sandbox.execFile.bind(sandbox),
125
+ commitIdentity: { name: "bot", email: "bot@example.com" },
126
+ });
127
+
128
+ // Host mode (explicit non-contained): omit sandbox; never claim containment.
129
+ const host = createSandboxCodingComposition(hostCwd, { workspaceMode: "host", executionPolicy: policy });
130
+ // host.composition.containmentClaim === false
131
+
132
+ await sandbox.execFile({ file: "npm", args: ["test"], cwd: "/workspace" });
133
+ await sandbox.close({
134
+ export: async (stream, meta) => hostArtifacts.write(stream, meta),
70
135
  });
71
136
  ```
72
137
 
73
138
  ## Extension and configuration notes
74
139
 
75
- Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers.
140
+ Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()`/`createSandboxCodingComposition()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` / `DisposableSandbox` are replaceable and host-owned; approval policy and sandboxing are separate layers. Custom remote sandboxes can implement `DisposableSandbox` without using Docker.
141
+
142
+ `createSandboxCodingComposition()` requires `workspaceMode`. Sandbox mode auto-wires FS/list/search through `DisposableSandbox.execFile` (or host-supplied custom operations) so mutations stay on the disposable tree until export. Host mode runs every coding tool against the host cwd and never sets `containmentClaim`. Sandbox shell + host FS throws unless `allowMixedWorkspaceWiring: true` (warnings + `containmentClaim: false`). Opt-in structured Git tools (`createGitTools(composition.workspaceRoot, { execFile: sandbox.execFile, commitIdentity })`) share the same tree/cwd; Prism still never pushes or opens PRs. Optional `@arnilo/prism-browser` can share the same disposable boundary: use `assertBrowserSandboxNetwork()` before browse-ready custom networks, and `createSharedSandboxBrowserOptions({ workspaceRoot, downloadsRoot, containedProxyAttestation })` so uploads/downloads align with `/workspace` and `/downloads`. Close the browser context before disposing the sandbox.
143
+
144
+ The Docker reference adapter starts by recorded container ID/label, uses argument arrays only, mounts source read-only, populates a size-bounded tmpfs `/workspace`, drops all capabilities, enables `no-new-privileges`, runs with `--init`, and never exposes the Docker socket, privileged mode, or host PID/IPC namespaces. Image pull/build/update stays outside Prism. Protected real-Docker checks are opt-in via `PRISM_TEST_DOCKER_SANDBOX=1` with host-supplied `PRISM_TEST_DOCKER_BIN` and digest-pinned `PRISM_TEST_DOCKER_IMAGE`.
76
145
 
77
146
  Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
78
147
 
79
148
  ## Security and performance notes
80
149
 
81
- Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
150
+ Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, repository list/search walks, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe.
151
+
152
+ Docker sandbox containment—not command regexes—enforces filesystem/network/process boundaries for the reference adapter. Network defaults to none; a custom Docker network still requires a host firewall/proxy for DNS/egress claims. Import rejects symlink escapes, devices, FIFOs, and sockets; export counts entries/bytes and hashes before host retention. Secrets in `secrets` are redacted from adapter errors and never exported as environment metadata. Unified workspace mode reuses existing sandbox/repo/coding hard caps and does not introduce unbounded host↔container sync loops. Host mode and `allowMixedWorkspaceWiring` never claim disposable containment. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter and Docker daemon.
82
153
 
83
154
  ## Related APIs
84
155
 
85
- - [Coding agent tools](coding-agent-tools.md)
156
+ - [Coding agent tools](coding-agent-tools.md): durable plan/todo Markdown helpers and `state.coding` checkpoint metadata for restart/resume without a second runtime
157
+ - [Workflows](workflows.md): `runWorkflow` / `resumeWorkflow` / `startWorkflowBackground` composition for coding tasks
86
158
  - [Host security guide](host-security.md)
159
+ - [Performance limits](performance.md)
87
160
  - [Tool execution primitives](tool-execution-primitives.md)
88
161
  - [Security/auth/trust](settings-auth-trust-security.md)
@@ -218,9 +218,18 @@ const providers = createOpenAIProviderPackage({ apiKey });
218
218
  - Never log passphrases, derived keys, or decrypted credential payloads.
219
219
  - Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
220
220
 
221
+ ## MCP authentication boundary
222
+
223
+ MCP credentials remain host inputs: resolve them before constructing client `requestInit` or inside server `resolveAuthInfo`. Stateful server `resolveIdentity` receives validated SDK auth metadata only to derive a stable non-secret principal ID. Never copy access/refresh tokens into MCP resource/prompt/sampling/elicitation payloads, telemetry, errors, or session bindings; Prism does not refresh or persist MCP OAuth automatically.
224
+
225
+ ## Web adapter credential boundary
226
+
227
+ `@arnilo/prism-web-tools` accepts explicit callbacks or `CredentialResolver`. Brave resolves `subscription_token`; Exa and Firecrawl resolve `api_key` immediately before each fixed-origin request. Keys never enter tool arguments/results, URLs, provider metadata, errors, telemetry, or prompts. Use separate least-privilege credentials and do not forward MCP/provider tokens between adapters.
228
+
221
229
  ## Related APIs
222
230
 
223
231
  - [Credentials and redaction](credentials-and-redaction.md): core resolver helpers and `refreshOAuthCredential()`
232
+ - [Web search, fetch, and extraction](web-tools.md): late-bound Brave/Exa/Firecrawl credentials
224
233
  - [Security/auth/trust](settings-auth-trust-security.md): host-owned settings/credentials boundaries
225
234
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 threat model and conformance matrix rows 7–10
226
235
  - `@arnilo/prism`: `CredentialResolver`, `OAuthCredentialStore`, `createMemoryCredentialStore()`
@@ -241,7 +241,7 @@ Minimum production guidance:
241
241
 
242
242
  - **Branch context:** implement `SessionStore.readBranchPath(query)` with an ancestor query / recursive CTE. Treat `SessionStore.list(sessionId)` as an O(n) development fallback only.
243
243
  - **Cursor pagination:** every `query*` method should honor `cursor`, `limit`, and `order`. Encode cursors from indexed columns such as `(timestamp, id)`, `(started_at, id)`, `(recorded_at, id)`, or `(run_id, sequence)`; never use offset pagination for long sessions.
244
- - **Batch appends:** `SessionStore.append()` is single-entry because the runtime advances one branch leaf at a time. Hosts may batch inside their DB/ledger adapters for `RunLedger` rows, but the adapter must preserve per-run event order and must not acknowledge writes before durable enqueue/commit.
244
+ - **Batch appends:** `SessionStore.append()` stays single-entry because runtime advances one branch leaf at a time. Optional `createBatchedRunLedger()` wraps any ledger with bounded FIFO/backpressure and explicit `write_through`, `flush_on_terminal`, or crash-loss-capable `buffered` acknowledgement semantics; SQLite/PostgreSQL defaults remain direct durable writes.
245
245
  - **Event sequence allocation:** allocate a monotonic `sequence` per `run_id` when inserting `prism_agent_events`. Use it with `run_id` for stable event timeline pagination when timestamps collide.
246
246
  - **Run/event/usage query shapes:** runs page by `(session_id, started_at, id)` or `(branch_id, started_at, id)`; events page by `(run_id, sequence)` or `(session_id, timestamp, id)`; usage pages by `(run_id, recorded_at, id)` or `(session_id, recorded_at, id)`.
247
247
  - **Host-owned sizing:** hosts own connection pools, transaction timeouts, page-size caps, queue/batch size, retention jobs, partitioning, and tenant/account/user isolation. Prism does not guess production limits.
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, live post-run scoring, and bounded batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
5
+ `@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, bounded persistence-trace grading, explicit host model judges, pairwise comparisons, CI thresholds, live post-run scoring, and batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
6
6
 
7
7
  ## When to use it
8
8
 
@@ -18,6 +18,10 @@ Use this package when a host needs offline quality checks or sampled live scorin
18
18
  | `runExperiment` | `agent`, dataset, scorers, bounded `concurrency`, optional store/ownership |
19
19
  | `createMemoryEvaluationStore` | optional seed records |
20
20
  | `appendEvaluationFeedback` | `RunFeedbackStore`, `EvaluationStore`, feedback fields, and 1–64 known evaluation IDs |
21
+ | `createPersistenceTraceResolver` | explicit `ProductionPersistenceStore`, exact session/run/ownership, page/byte bounds |
22
+ | `createModelJudge` | host judge callback, stable rubric/version, timeout/attempt/output bounds |
23
+ | `runComparison` | immutable dataset, 2–8 named candidates by default, pairwise scorers |
24
+ | `assertEvaluationThreshold` / `serializeEvaluationReport` | mean/failure/per-scorer gates and bounded redacted JSON |
21
25
 
22
26
  ## Outputs / response / events
23
27
 
@@ -112,11 +116,41 @@ console.log(report.aggregate.meanScore, linked.evaluationIds);
112
116
  - Scorers receive result/item data only. Credentials, tools, and workspace access are not provided unless the host deliberately closes over them.
113
117
  - Records pass through `SecretRedactor` / `secrets` before store append.
114
118
  - Queries filter by ownership scope. Feedback linkage additionally requires tenant plus account/user and the feedback store re-verifies the run.
115
- - Experiment concurrency defaults to `1` and is capped at `32`. Scoring can reference run IDs without duplicating unbounded event payloads.
119
+ - Experiment concurrency defaults to `1` and is capped at `32`. Datasets cap at 10,000 items.
120
+ - Trace reads default to 100 rows × 20 pages with a 4 MiB aggregate cap (hard: 1,000 × 100 and 32 MiB). Repeated/missing cursors, identity drift, ownership drift, and overflow fail closed before scoring.
121
+ - Model judges are host callbacks, not providers: Prism passes rubric/version plus bounded target only—never credential resolvers, tools, or workspace. Defaults are one attempt, 30 seconds, and 16 KiB output; failures become redacted evaluation records.
122
+ - Pairwise candidates are sorted by name, executed once per item, compared in stable item/pair/scorer order, and record ties/failures without choosing a winner. Candidate and scorer outputs have byte caps.
123
+ - `assertEvaluationThreshold()` throws `ERR_PRISM_EVAL_THRESHOLD`; an uncaught error gives CI a non-zero exit. Keep model-judge/live gates credential-gated and outside the network-free default suite. `serializeEvaluationReport()` bounds/redacts checked-in artifacts.
124
+
125
+ ## Trace, judge, comparison, and CI example
126
+
127
+ ```ts
128
+ const traceResolver = createPersistenceTraceResolver(persistence);
129
+ const judge = createModelJudge({
130
+ id: "quality", rubric: "Score factual quality from 0 to 1", rubricVersion: "2026-07-20",
131
+ judge: hostStructuredJudge,
132
+ });
133
+ const evaluations = await scoreRun({ result, scorers: [judge], traceResolver, ownership });
134
+ const comparison = await runComparison({ dataset, candidates: { baseline, candidate }, scorers: [preference] });
135
+ assertEvaluationThreshold(report, { minimumMean: 0.9, maximumFailures: 0 });
136
+ ```
137
+
138
+ `traceResolver` is explicit; no arbitrary run search occurs. `baseline`/`candidate` are host functions returning `AgentRunResult`. See `examples/evaluation-gate.ts` for a network-free gate and `examples/coding-browser-evaluation.ts` for coding/browser adversarial fixtures.
139
+
140
+ ## Coding and browser adversarial evaluations (0.0.9)
141
+
142
+ Release 0.0.9 ships curated network-free adversarial fixtures in package tests:
143
+
144
+ - `@arnilo/prism-coding-agent` `eval-fixtures.test.ts`: safe native list vs shell, Git path/ref injection, dirty-tree rollback, unknown named-check failure, PR-handoff artifact completeness, and prompt-injection file content under read-only tools.
145
+ - `@arnilo/prism-browser` `eval-fixtures.test.ts`: stale snapshot refs, side-effect approval, private/loopback/file deny, upload/download/screenshot policy, CSS/evaluate target rejection, and hostile accessible-name text.
146
+
147
+ Fixtures reuse `@arnilo/prism-evals` (`defineDataset` / `defineScorer` / `scoreRun` / `assertEvaluationThreshold` / `serializeEvaluationReport`). Optional SWE-bench-compatible or live-browser harnesses remain host adapters — they are not default dependencies or quality claims. Protected real Docker/Playwright gates stay env-gated (`PRISM_TEST_DOCKER_SANDBOX`, `PRISM_LIVE_PLAYWRIGHT`) and never enter `sdk:ready`.
116
148
 
117
149
  ## Related APIs
118
150
 
119
151
  - [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
120
152
  - [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
121
- - [Observability](observability.md): trace/run metadata hosts may copy into `traceId`
122
- - [Release and install](release-and-install.md): optional package install
153
+ - [Observability](observability.md): use `onTraceReference` or bounded `traceId(runId)` to supply `ScoreRunOptions.traceId`; evaluation telemetry emits no reason/explanation content
154
+ - [Coding agent tools](coding-agent-tools.md) / [Browser automation](browser-automation.md) / [Workflows](workflows.md): network-free coding-task composition at `examples/durable-coding-workflow.ts`; adversarial coding/browser eval example at `examples/coding-browser-evaluation.ts`
155
+ - [Performance limits](performance.md): `scripts/benchmark-0.0.10.mjs` workspace-mode evidence and `scripts/benchmark-0.0.9.mjs` coding/browser evidence fields
156
+ - [Release and install](release-and-install.md): optional package install and protected sandbox-browser workflow
@@ -30,7 +30,7 @@ Decisions are `allow`, `block`, `tripwire`, or `interrupt`. Evaluation defaults
30
30
 
31
31
  ## Outputs / response / events
32
32
 
33
- Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
33
+ Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. Optional OpenTelemetry instrumentation records only controlled stage/action on a short run-child span; guardrail name, reason, and metadata are excluded. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
34
34
 
35
35
  Ordering is fixed:
36
36
 
@@ -65,11 +65,12 @@ Guardrails are callbacks supplied by the host. Prism does not discover, load, re
65
65
 
66
66
  ## Security and performance notes
67
67
 
68
- Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work.
68
+ Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work. Browser snapshots and page text from `@arnilo/prism-browser` are untrusted external content: never allow them to modify tools, permissions, credentials, or policy. Browser mutations still require host `ExecutionPolicy`/approval; prompt-injection text in a page cannot grant upload/download release.
69
69
 
70
70
  ## Related APIs
71
71
 
72
72
  - [Agent/session runtime](agent-session-runtime.md)
73
73
  - [Tools](tools.md)
74
+ - [Browser automation](browser-automation.md)
74
75
  - [Agent events](agent-events.md)
75
76
  - [Host security](host-security.md)
@@ -30,6 +30,7 @@ Start from explicit host inputs. Do not let runtime code discover security state
30
30
  | Remote media policy | public/default pinned DNS or explicit trusted transport | `SsrfPolicy`, `resolveMediaContentBlock()` |
31
31
  | Durable history | host database adapter | `SessionStore`, `assertSessionStoreConforms()` |
32
32
  | Durable audit | host ledger adapter | `RunLedger`, `redactRunLedgerRecord()` |
33
+ | Telemetry | host OpenTelemetry SDK/exporter | metadata-only adapter, controlled metric labels, `onTraceReference` |
33
34
  | Durable interruption | host checkpoint + session stores, exact ownership | `RunOptions.runState`, `resumeAgentRun()`, `createAgentRunLifecycle()`, `createSecureAgent()` |
34
35
  | Extensions | explicit package imports only | `createExtensionKernel`, `ExtensionAPI` |
35
36
  | Remote agent/workflow API | host authentication + ownership mapping | `@arnilo/prism-server`, `createPrismHandler()` |
@@ -45,7 +46,7 @@ Security controls fail closed before side effects when wired at the guarded edge
45
46
  - validator failures emit `tool_execution_blocked` with `validation_failed`
46
47
  - configured guardrails fail closed; output stages buffer blocked provider/tool content before events, ledgers, session entries, or MCP responses
47
48
  - configured redactors scrub provider requests, agent events, session entries, ledger records, tool errors, extension errors, injector context, and durable run checkpoints
48
- - durable resume requires host-derived exact ownership and checkpoint version; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
49
+ - durable resume requires host-derived exact ownership and checkpoint version; coding-task resume also revalidates plan/workspace artifact hashes plus tool/policy/image fingerprints via `assertCodingResumeAllowed` before import; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
49
50
 
50
51
  These checks are explicit function calls during load, assembly, dispatch, append, or run handling. Prism adds no background watchers, filesystem scanners, network probes, credential polling, or automatic extension discovery.
51
52
 
@@ -120,6 +121,8 @@ Wire those values where they matter: provider adapters receive the resolved cred
120
121
  - Use `createContributionRegistries({ duplicate: "error" })` and prefixed names for third-party packages to prevent silent shadowing.
121
122
  - Extension contributions are inert until selected. Loading an extension package runs its `setup(api)` code, so hosts should load only trusted packages or isolate untrusted code outside Prism.
122
123
  - Skills and instruction injectors grant no tools, permissions, validators, or resource access. Host-active tools and permission policies still decide execution.
124
+ - Optional ledger batching accepts runtime-redacted records only. Prefer `flush_on_terminal`; `buffered` explicitly permits crash-before-flush loss. Flush failures propagate; hosts can call `dispose({ flush: false })` to clear queued objects when deliberately discarding an aborted buffered workload.
125
+ - Session snapshot cache holds one session-local leaf for at most one second and invalidates after committed mutation, checkout, compaction, and resume; it never crosses session/branch/ownership.
123
126
  - For production persistence, implement a database-backed `SessionStore`/`RunLedger`, run `assertSessionStoreConforms()` against the store, and follow the database schema guidance. Do not ship provider instances, credential resolvers, or secrets into durable rows.
124
127
 
125
128
  ## Security and performance notes
@@ -130,10 +133,12 @@ Wire those values where they matter: provider adapters receive the resolved cred
130
133
  - Known secrets must be passed into redactors before data is emitted or persisted. Redact again in host adapters if they transform records after Prism redaction.
131
134
  - Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects. Its untrusted-schema adapter rejects non-local refs, forbidden keys/cycles/non-finite values and bounds bytes/depth/properties/keywords/refs plus its LRU cache before Ajv compilation; do not raise caps above documented hard limits.
132
135
  - Treat embeddings as untrusted numeric input. `@arnilo/prism-memory` rejects empty, non-number, NaN, and infinite vectors before in-memory similarity or pgvector parameters; custom `Embedder`/`VectorStore` implementations must retain the same boundary.
136
+ - Evaluation trace readers require exact supplied ownership plus session/run identity, reject cursor/identity drift, and redact before bounded scorer/judge input. Model-judge callbacks receive no credential resolver, tools, or workspace; keep live judges outside default CI and redact report artifacts.
133
137
  - Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
134
- - MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands, requires per-call `authorize`, and retains core gates. Its web handler still needs host `resolveAuthInfo`, TLS, edge rate limiting, and exact host/origin policy. See [MCP client/server exposure](mcp-tools.md).
138
+ - MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands/resources/prompts, requires per-operation `authorize`, and retains core gates. Sampling, roots, model/credential selection, and elicitation consent stay host-owned; URL elicitation is never opened automatically. Stateful web mode requires host `resolveAuthInfo` plus `resolveIdentity`, exact origin policy, and binds every POST/GET/DELETE/SSE request to one non-secret principal; mismatches return 404. Handler still needs TLS and edge rate limiting. See [MCP client/server exposure](mcp-tools.md).
135
139
  - `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
136
- - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Limits are not containment: Prism provides no OS sandbox unless the host supplies one.
140
+ - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, repository list/search depth/entry/match/scan/time caps, structured Git path/ref/message/output/patch/worktree caps, named-check concurrency/output caps, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Opt-in `createGitTools()` uses argument arrays with hooks/credential prompts/external diff disabled, requires host `commitIdentity` for commits, and never pushes or opens PRs. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell/repository backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, identity-scoped approval caching, required `workspaceMode` on `createSandboxCodingComposition()` / `createSandboxCodingTools()`, and the optional `createDockerSandbox()` reference adapter. **Host mode is never contained execution** (`containmentClaim: false`). Sandbox mode claims containment only when FS backends target the disposable tree; mixed wiring requires `allowMixedWorkspaceWiring` and still does not claim containment. Limits alone are not containment: construct the Docker adapter (absolute CLI, digest-pinned image, network none by default) or an equivalent host sandbox before treating coding execution as production-safe. Docker daemon/image trust, egress firewall/proxy, and artifact retention remain host-owned.
141
+ - Optional `@arnilo/prism-browser` requires a host-supplied Playwright Browser (`playwright-core@1.61.0` peer). Import is inert. One non-persistent context belongs to one run; actions serialize; refs are snapshot-scoped; CSS/evaluate/CDP/persistent profiles are denied. Context routing + `serviceWorkers: "block"` deny file/data/blob/devtools/private/loopback by default and require contained-proxy attestation for external egress (Playwright routing is defense in depth, not DNS containment). Uploads are realpath-rooted; downloads quarantine with hash/MIME until host `approveRelease`; screenshots return bounded `ImageContent`. Observation vs mutation/high-impact actions map to `ExecutionPolicy`. Treat snapshot/page text as untrusted external content. Close contexts with `browser_close` or `manager.closeRun(runId)` on abort/terminal. Browser control endpoint, binary/image pin, and real egress firewall/proxy remain host-owned. Shared sandbox: `createSharedSandboxBrowserOptions()` + `assertBrowserSandboxNetwork()`.
137
142
  - `@arnilo/prism-credentials-node` rejects oversized/malformed envelopes and excessive scrypt work before KDF allocation, uses async scrypt, and requires restrictive existing/new Unix vault modes. Keep vault ownership and parent-directory access host-controlled; review before `chmod 600`, never auto-weaken a file policy. Keychain calls use abort-aware native async work with finite timeout/payload caps and sanitized errors. OS prompts, service availability, and whether a native backend promptly honors cancellation remain host/platform boundaries; no plaintext fallback is attempted.
138
143
  - LLM compaction always sends finite summary `maxTokens`, retains bounded deltas/events, and bounds/redacts provider/factory/policy error detail. Observational-memory workers cap turns, calls, arguments, results, transcript, and surfaced errors; unknown tools fail before execution, while invalid results can only be rejected after a host tool returns and may therefore follow side effects. Pass all known provider/credential/tool secrets into compaction/runtime options; exact replacement is not secret discovery.
139
144
  - Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
@@ -160,7 +165,27 @@ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy
160
165
  - Keep depth, active children, input, turn/tool/token, timeout, and queue ceilings finite; propagate abort through nested calls.
161
166
  - Expose A2A only behind per-request authentication/authorization, TLS, edge rate limits, and replay policy. Public card discovery grants no invoke access.
162
167
  - Remote A2A endpoints/card URLs require exact HTTPS origin allow-lists and redirect rejection. Pin ES256 card keys/expiry; never auto-fetch untrusted `jku`.
163
- - Treat cards, task status, errors, artifacts, and SSE frames as untrusted bounded input and redact before logs/hooks/events. Client streaming uses one fatal UTF-8 decoder, bounded incremental LF/CRLF/multiline SSE parsing, and rejects partial/post-terminal frames; malformed remote bytes never become replacement characters in JSON.
168
+ - Treat cards, rich parts, task status/history, errors, artifacts, and SSE replay frames as untrusted bounded input and redact before logs/hooks/events. URL parts require host public/pinned-network policy and are never auto-fetched. Durable task/push adapters repeat exact-owner checks; foreign/missing records share not-found responses.
169
+ - Push delivery remains host-owned: validate every attempt/redirect against SSRF/rebinding policy, cap retries/time/output, authenticate webhook payloads, deduplicate event IDs, and keep token/auth credentials out of configs returned over A2A. Client streaming uses fatal UTF-8 decoding and rejects partial/post-terminal frames.
170
+
171
+ ## Web research boundaries
172
+
173
+ - Construct `@arnilo/prism-web-tools` with one host-selected Brave or Exa adapter; never expose adapter/provider/credential/schema selection to model arguments.
174
+ - Provider API origins are fixed exact HTTPS origins and redirects fail. Credentials resolve immediately before I/O; remote bodies and secrets are excluded from errors/results/telemetry.
175
+ - Firecrawl targets reject userinfo, non-HTTP(S), private literals, and policy-denied hosts. Supply `validateUrl` for host DNS/rebinding/egress checks. Firecrawl performs remote retrieval, so Prism cannot pin target DNS after handoff.
176
+ - Treat every snippet, highlight, Markdown byte, metadata field, and extracted JSON value as prompt-injection-capable untrusted data. Never elevate it into system instructions or let it modify tools, permissions, trust, credentials, routing, or extraction schema.
177
+ - Keep counts/bytes/retries/rate delays/polling/concurrency/wall time finite. Live credentials belong only in explicit protected `PRISM_LIVE_WEB=1` runs.
178
+
179
+ ## Supply-chain and live-canary boundaries
180
+
181
+ - Require `security / codeql`, `security / supply-chain`, PR dependency review, release readiness, and PostgreSQL integration in protected-branch rules. Enable GitHub secret scanning and push protection as repository settings; checked-in workflows cannot enable those service controls.
182
+ - Actions are pinned to full commit revisions. Dependabot proposes weekly npm/action revision changes; review upstream release notes before merge rather than replacing pins with moving tags.
183
+ - `scripts/verify-sbom.mjs` accepts only bounded SPDX 2.3 inventory with exact checked-in permissive licenses. Any missing/new expression fails until reviewed; do not widen policy merely to unblock CI.
184
+ - `scripts/scan-secrets.mjs` checks tracked source and unpacked public tarballs for high-confidence credential/private-key forms without printing matched values. It complements GitHub secret scanning; it is not entropy scanning or DLP.
185
+ - Tag publication alone receives npm/OIDC/attestation permissions. Untrusted pull-request code receives no canary, npm, or OIDC secret and no workflow uses `pull_request_target`.
186
+ - Scheduled/manual canaries run only in protected `live-canaries` environment. Use dedicated read-only/low-quota credentials and provider account spend limits. Runner performs four probes, at most one MCP cleanup, one provider output token, one Brave result, 64-KiB responses, and finite timeouts; report excludes endpoints, headers, bodies, credentials, and MCP session IDs.
187
+ - Scheduled/manual coding/browser containment checks run in protected `sandbox-browser` environment (`.github/workflows/sandbox-browser.yml`). They receive no provider/npm/OIDC secrets; Docker/Playwright enablement is variable-gated with host-preloaded digest-pinned images/binaries; uploads are redacted aggregate status only.
188
+ - Live endpoint operators own TLS, egress allow-lists, account-dollar budget, cleanup beyond MCP session DELETE, and revocation. Failed canaries log only operation kind plus status/timeout; inspect provider-side audit logs for details.
164
189
 
165
190
  ## Related APIs
166
191
 
package/docs/index.md CHANGED
@@ -10,11 +10,11 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
10
10
  - [Agent definitions](agent-definitions.md): resolve declarative `AgentDefinition` values via `resolveAgentDefinition`, and turn app-config `<configRoot>/agents/<name>/AGENT.md` bundles into runnable agents via `discoverAgentBundles` / `resolveAgentBundle` (explicit tool/skill activation by name, fail-closed omitted capabilities, migration-only `activateAllCapabilities`, strict duplicate scope checks, configurable prompt layers, no auto-discovery).
11
11
  - [Agent loops](agent-loops.md): replaceable per-run control loops — `singleShotLoop` default and opt-in bounded artifact-loop tool rounds with host-supplied `validator`/`parser`/`repairer` callbacks.
12
12
  - [Guardrails](guardrails.md): typed fail-closed input/output/tool checks with buffered provider output and redacted decision records.
13
- - [Agent events](agent-events.md): the `AgentEvent` stream — agent/turn/message (including live `tool_call_delta` fragments), provider turn timing, tool execution, queue/subscriber overflow, compaction/retry, artifact validation/refinement, and error variants, redacted via `redactAgentEvent`.
14
- - [Observability](observability.md): metadata-only provider/tool and run-feedback/evaluation projection, terminal span cleanup, low-cardinality metrics, and optional `@arnilo/prism-observability-opentelemetry` adapter.
15
- - [Evaluations](evaluations.md): optional deterministic scorers/datasets/experiments plus ID-only linkage from evaluation records to immutable owned run feedback.
16
- - [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence plus bounded immutable run/trace feedback, evaluation links, ownership, redaction, query, and deletion semantics.
17
- - [Performance limits](performance.md): bounded live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
13
+ - [Agent events](agent-events.md): redacted lifecycle stream used by UIs, ledgers, and metadata-only parented telemetry; message/progress deltas never create spans.
14
+ - [Observability](observability.md): OTel GenAI agent/provider/tool hierarchy, host context parenting, bounded trace linkage, safe evaluation events, controlled metrics, and exporter isolation.
15
+ - [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, coding/browser adversarial fixtures, and ID-only linkage to immutable owned run feedback.
16
+ - [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence, optional bounded FIFO durability policies, session snapshot caching, and immutable run/trace feedback.
17
+ - [Performance limits](performance.md): bounded evaluation traces/judges/reports, 0.0.10 workspace-mode and 0.0.9 coding/browser benchmark evidence, security scan/live-canary backstops, live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
18
18
  - [Structured output](structured-output.md): the `Artifact*` seam plus provider-native `StructuredOutputOptions` / `structuredOutputMode` for capable models.
19
19
 
20
20
  ## Compaction/session memory
@@ -27,7 +27,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
27
27
  - [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention, and NoSQL mapping.
28
28
  - [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, and transactionally verified/backfilled migration-v3 metadata.
29
29
  - [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
30
- - [Migration guide](migration.md): 0.0.3 compatibility, 0.0.6 hardening, and 0.0.7 guardrails, RunLimits, durable approval/resume, and secure composition.
30
+ - [Migration guide](migration.md): 0.0.3 compatibility through 0.0.10 workspace modes (required `workspaceMode`, fail-closed mixed wiring) and 0.0.9 coding/browser sandbox, repository/Git, durable plans, and Playwright automation changes.
31
31
  - [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety.
32
32
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
33
33
 
@@ -57,9 +57,11 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
57
57
  - [Tools](tools.md): register host-owned active tools with replace-or-error duplicate policy, apply exact allow/deny filtering, dispatch normal or opt-in bounded artifact-loop calls, and optionally bound untrusted JSON Schema compilation.
58
58
  - [Tool execution primitives](tool-execution-primitives.md): finite JSON Schema LRU validation, exclusive-aware bounded parallel dispatch, MCP bridge mapping, coding execution policy, and image-read bounds.
59
59
  - [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
60
- - [MCP client bridge and server exposure](mcp-tools.md): optional bounded atomic tool discovery/results, exact-origin DNS-pinned HTTPS/loopback-only HTTP client transport, and explicitly authorized Prism tools/commands/durable agent lifecycle on SDK `McpServer`.
61
- - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, and `edit` definitions with streamed text pages, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
62
- - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, and abort-aware streaming sandbox adapters for coding tools.
60
+ - [MCP client bridge and server exposure](mcp-tools.md): SDK-1.29.0 bounded tools/resources/prompts, host-owned roots/sampling/elicitation, exact-origin DNS-pinned client transport, and principal-bound opt-in Streamable HTTP sessions.
61
+ - [Web search, fetch, and extraction](web-tools.md): optional host-selected Brave/Exa discovery and Firecrawl Markdown/schema tools with native fetch, stable citations, late credentials, finite limits, and explicit untrusted-content boundaries.
62
+ - [Browser automation](browser-automation.md): optional `@arnilo/prism-browser` with host-supplied Playwright contexts, AI-mode snapshots/refs, ordered `browser_open`/`browser_snapshot`/`browser_act`/`browser_close`, egress/side-effect/upload/download/screenshot policy, and finite page/action/snapshot/network/artifact caps.
63
+ - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, `edit`, `repo_list`, and `repo_search` definitions plus opt-in `createGitTools()` / `coding_check` for structured Git status/diff/branch/worktree/apply/commit/PR-handoff and named checks; durable plan/todo Markdown helpers with workflow `state.coding` checkpoint metadata; streamed text pages, bounded repository list/search, finite Git/check/plan caps, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
64
+ - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, required `workspaceMode` (`host`/`sandbox`) with fail-closed mixed wiring, `createSandboxCodingComposition()` containment metadata, and the disposable Docker/OCI sandbox reference with bounded workspace import/export.
63
65
 
64
66
  ## Extensions/plugins
65
67
  - [Contribution discovery (workspace)](contribution-discovery.md): opt-in, realpath-contained directory scanner turning `SKILL.md`/`manifest.json` into inert `DiscoveredContribution` envelopes the host registers — no `import()`, no auto-activate, no provider scanning. Per-agent bundles remain app-controlled and are documented under Agent/session runtime.
@@ -77,17 +79,17 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
77
79
  - [Web-standard server handler](server.md): optional framework-free authorized direct/SSE agent, explicitly selected durable agent lifecycle, and durable workflow routes with explicit bounds and zero default exposure.
78
80
 
79
81
  ## Multi-agent and interoperability
80
- - [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, and finite budgets.
81
- - [A2A interoperability](a2a.md): optional A2A 1.0 cards, ES256 signatures, authorized JSON-RPC/SSE handler, and exact-origin client with fatal streaming UTF-8 plus bounded LF/CRLF/multiline SSE parsing.
82
+ - [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, finite budgets, host-projected delegation telemetry, and separate A2A durable adapter boundary.
83
+ - [A2A interoperability](a2a.md): A2A 1.0 JSON-RPC/HTTPS cards plus host-owned durable task get/list/cancel/subscribe, bounded rich parts/replay, principal-scoped push configs, and exact-origin verified client.
82
84
 
83
85
  ## CLI/RPC
84
86
  - [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
85
- - [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Interactive TUI (C-012) deferred.
87
+ - [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Compose coding plans/checkpoints via workspace Markdown + `state.coding` without a second runtime. Interactive TUI (C-012) deferred.
86
88
  - [Workflow orchestration primitives](workflow-orchestration-primitives.md): architecture inventory — workflow adapters consume core `CheckpointStore`, `LeaseStore`, and bounded `EventMultiplexer`; run control and optional RPC commands stay package-local.
87
89
  - [Workflow/TUI scope](workflow-tui-primitives.md): records why 0.0.5 ships workflow APIs/RPC control but no interactive terminal UI.
88
90
 
89
91
  ## Security and credentials
90
- - [Host security guide](host-security.md): fail-closed checklist for bounded encrypted vault/KDF/keychain, JSON Schema/vector/cryptographic-ID, MCP discovery/result/transport operations, settings, redaction, trust roots, remote media, exact workflow ownership/revision checks, finite coding I/O and spill ownership, permission/approval policies, persistence, extension loading, and tool validation.
92
+ - [Host security guide](host-security.md): fail-closed checklist for supply-chain/attestation/canary isolation, bounded credentials, JSON/schema/vector/crypto, MCP/A2A/web remote boundaries, untrusted external content, settings, redaction, trust roots, workflow ownership, coding I/O, permissions, persistence, extensions, and tool validation.
91
93
  - [Security/auth/trust](settings-auth-trust-security.md): settings providers, credential helpers, trust/permission policies, redaction controls, host-owned settings/credentials wiring outside `AgentConfig`, and security-boundary hardening summary.
92
94
  - [Credentials and redaction](credentials-and-redaction.md): compose explicit credential resolver order, use caller-supplied env objects/OAuth refresh helpers, resolve credentials only at the provider edge, and redact known secret values.
93
95
  - [Credential storage](credential-storage.md): optional `@arnilo/prism-credentials-node` adapter with strict bounded AES-GCM envelopes, async finite scrypt, restrictive Unix files, and abort-aware bounded system-keychain calls.
@@ -96,14 +98,17 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
96
98
  - Provider test doubles: `createMockProvider()` and provider event helpers are documented on the canonical Provider layer page above.
97
99
  - [Provider conformance](provider-conformance.md): run network-free provider adapter assertions (stream order, abort, tool-call reconstruction, cache usage, content coverage, protected header ownership, secret leak) from `@arnilo/prism/testing/provider-conformance`.
98
100
  - [Session store conformance](session-store-conformance.md): assert any `SessionStore` adapter satisfies append/idempotency/conflict/branch invariants from `@arnilo/prism/testing/session-store-conformance`.
99
- - [Run ledger conformance](run-ledger-conformance.md): assert durable run/event/tool/usage writes and reopen survival. Run-feedback stores use `@arnilo/prism/testing/feedback` for append/query/delete/ownership linkage conformance.
101
+ - [Run ledger conformance](run-ledger-conformance.md): assert durable run/event/tool/usage writes and reopen survival; batch-wrapper FIFO/bounds/flush checks remain separate. Run-feedback stores use `@arnilo/prism/testing/feedback` for append/query/delete/ownership linkage conformance.
100
102
  - [Compaction conformance](compaction-conformance.md): assert any `CompactionStrategy` returns a non-empty redacted summary and observes abort from `@arnilo/prism/testing/compaction-conformance`.
101
103
  - [Tool conformance](tool-conformance.md): assert the tool-dispatch blocked-reason matrix (unknown/denied/invalid/permission/validator) and success path from `@arnilo/prism/testing/tool-conformance`.
102
104
  - [Extension conformance](extension-conformance.md): assert an `Extension` setup runs, contributions stay inert, and setup errors are redacted or rethrown from `@arnilo/prism/testing/extension-conformance`.
103
105
  - `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, cache-aware prompt assembly, NeuralWatt agent run, stores/branching, compaction, observational-memory recall, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
104
106
 
105
107
  ## Release and install
106
- - [Release and install](release-and-install.md): 30-package graph and profiles, install/tarball rules, deterministic resumable provenance publication, and offline test budget.
108
+ - [Release and install](release-and-install.md): 32-package graph (including optional browser), install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, protected live canaries, and sandbox-browser Docker/Playwright gates.
109
+ - [Review coverage (2026-07-21 Phase 5)](review-coverage-2026-07-21-phase-5.md): Plan 073 evidence freeze — unified workspace modes, primitive ownership, reused finite limits, threats, and 0.0.10 release gates.
110
+ - [Review coverage (2026-07-20 Phase 4)](review-coverage-2026-07-20-phase-4.md): Plan 072 evidence freeze — revised coding/browser-only scope, external revisions, primitive ownership, finite limits, threats, and 0.0.9 release gates.
111
+ - [Review coverage (2026-07-19 Phase 3)](review-coverage-2026-07-19-phase-3.md): Plan 070 evidence freeze — exact protocol/vendor references, capability/primitive/limit matrices, supported boundaries, and 0.0.8 release evidence.
107
112
  - [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md): Plan 067 evidence freeze — P0–P2 re-verification owners, seven first-party provider packages mapped to official-doc URLs, Pi secondary refs, cache/thinking/discovery surfaces, credential canaries, and use-case model-binding inventory.
108
113
  - [Review coverage (2026-07-15)](review-coverage-2026-07-15.md): frozen 0.0.5 finding/feature ownership, existing-primitive inventory, package decisions, threat boundaries, exclusions, and measured Phase 0 baseline.
109
114
  - [Review coverage (2026-07-14)](review-coverage-2026-07-14.md): traceability matrix linking review findings and bug-report fixes to plan tasks, tests, and documentation for release 0.0.4.
package/docs/mcp-tools.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-mcp` has two explicit directions. Its client bridge connects hosts to remote [Model Context Protocol](https://modelcontextprotocol.io) servers and maps discovered tools to ordinary `ToolDefinition`s. Its server API registers selected Prism `ToolDefinition` and `CommandDefinition` values on the official SDK `McpServer`, with required authorization and a bounded optional Web-standard Streamable HTTP handler. The package wraps `@modelcontextprotocol/sdk` v1.29+ and adds no MCP branch to core Prism.
5
+ `@arnilo/prism-mcp` has two explicit directions. Its client bridge connects hosts to remote [Model Context Protocol](https://modelcontextprotocol.io) servers and maps discovered tools to ordinary `ToolDefinition`s. Its server API registers selected Prism `ToolDefinition` and `CommandDefinition` values on the official SDK `McpServer`, with required authorization and a bounded optional Web-standard Streamable HTTP handler. The package pins `@modelcontextprotocol/sdk` **1.29.0** (MCP protocol negotiation remains SDK-owned) and adds no MCP branch to core Prism.
6
6
 
7
7
  Primary API:
8
8
 
@@ -19,7 +19,21 @@ await bridge.refresh(); // re-list after notifications or TTL expiry
19
19
  await bridge.close(); // close client + transport
20
20
  ```
21
21
 
22
- Advanced hosts that manage their own `Client` + `Transport` can call `attachMcpToolBridge(client, transport, options)` after `client.connect(transport)`.
22
+ Advanced hosts that manage their own `Client` + `Transport` can call `attachMcpToolBridge()` or `attachMcpCapabilities()` after connect. `connectMcpCapabilities()` keeps resources/prompts as host-facing facades rather than converting them into model tools, and declares roots/sampling/elicitation only when callbacks are supplied.
23
+
24
+ ```ts
25
+ const bridge = await connectMcpCapabilities({
26
+ serverId: "research",
27
+ transport: { type: "streamable-http", url, allowedOrigins: [origin] },
28
+ roots: () => [{ uri: "file:///workspace", name: "workspace" }],
29
+ sampling: hostSampling, // host selects model/provider/credentials
30
+ elicitation: hostElicitation, // URL mode returns approval; Prism never opens/fetches URL
31
+ });
32
+ await bridge.listResources();
33
+ await bridge.getPrompt("review", { topic: "security" });
34
+ ```
35
+
36
+ Server capability matrix for SDK 1.29.0: tools/resources/prompts and their list-change notifications are supported through official registrations; roots/sampling/form+URL elicitation are supported as explicit client callbacks. Missing server resources/prompts throw `McpUnsupportedCapabilityError` with `ERR_PRISM_MCP_UNSUPPORTED_CAPABILITY`. Resource/prompt results and sampling/elicitation inputs/results are bounded JSON. Accepted form/URL elicitation requires host-only `humanInteraction: true`; bridge strips marker before protocol output and fails closed when absent. Automatic root discovery/consent, model selection, credential resolution, URL navigation, generic command proxying, and custom JSON-RPC are unsupported.
23
37
 
24
38
  Server direction:
25
39
 
@@ -41,10 +55,13 @@ const handleMcp = await createPrismMcpWebHandler(server, {
41
55
  resolveAuthInfo: authenticateRequest,
42
56
  allowedHosts: ["api.example.test"],
43
57
  allowedOrigins: ["https://app.example.test"],
58
+ // Omit these two for bounded stateless JSON mode.
59
+ sessionIdGenerator: crypto.randomUUID,
60
+ resolveIdentity: (_request, auth) => auth ? { id: validatedPrincipalId(auth) } : false,
44
61
  });
45
62
  ```
46
63
 
47
- `McpServer.connect(transport)` remains available for SDK stdio or in-memory transports. The helper uses SDK `WebStandardStreamableHTTPServerTransport` in bounded stateless JSON-response mode; it does not start a listener.
64
+ `McpServer.connect(transport)` remains available for SDK stdio or in-memory transports. The helper uses SDK `WebStandardStreamableHTTPServerTransport`; it does not start a listener. Default remains bounded stateless JSON-response mode. Supplying `sessionIdGenerator` enables SDK `MCP-Session-Id` POST/GET/DELETE/SSE lifecycle and requires exact `allowedOrigins` plus host `resolveIdentity`. Every request re-authenticates, and a different principal receives non-disclosing 404. SDK owns protocol-version/session headers and SSE semantics. SDK 1.29.0's in-memory event store is not enabled, so `Last-Event-ID` replay is explicitly unsupported; reconnect starts only through SDK-supported active session GET.
48
65
 
49
66
  ## When to use it
50
67
 
@@ -160,6 +177,7 @@ Plaintext is accepted only when `allowLoopbackHttp: true`, the URL hostname is l
160
177
  | Option | Default | Purpose |
161
178
  | --- | --- | --- |
162
179
  | `tools` / `commands` | empty | Explicit allow-list; zero default exposure |
180
+ | `resources` / `prompts` | empty | Static URI/name registrations with bounded host callbacks and per-read/get authorization |
163
181
  | `agentRuns` | empty | Explicit `{ [agentId]: { lifecycle } }` map; registers `agent.<id>.status` and `agent.<id>.resume` only |
164
182
  | `authorize` | required | Per-call host authz using SDK auth/session metadata |
165
183
  | `permission` / `validate` / `redactor` | none | Core tool-dispatch gates and known-secret redaction |
@@ -167,7 +185,7 @@ Plaintext is accepted only when `allowLoopbackHttp: true`, the URL hostname is l
167
185
  | `maxConcurrentCalls` | 16 (256 hard) | Bound active tool/command execution |
168
186
  | `callTimeoutMs` | 60 s (30 min hard) | Abort and return timed-out calls |
169
187
 
170
- Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard), 32 concurrent requests (512 hard), and 60 s timeout (30 min hard). It parses bounded JSON before passing `parsedBody` to the SDK transport. `allowedHosts`/`allowedOrigins` activate SDK DNS-rebinding checks only when explicitly configured. Authentication data comes only from host `resolveAuthInfo()`.
188
+ Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard), 32 concurrent requests (512 hard), 60 s timeout (30 min hard), and 32 sessions (512 hard). Stateful mode is intentionally one official SDK transport/session lineage per handler; use one handler/server instance per independently hosted endpoint when multi-tenant transport isolation is required. It parses bounded JSON before passing `parsedBody` to the SDK transport. `allowedHosts`/`allowedOrigins` activate SDK DNS-rebinding checks only when explicitly configured. Authentication data comes only from host `resolveAuthInfo()`.
171
189
 
172
190
  ## Security and performance notes
173
191
 
@@ -183,7 +201,8 @@ Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard),
183
201
  | Accidental server exposure | Empty default arrays/maps, duplicate-name rejection, explicit tools/commands/lifecycle only |
184
202
  | Agent lifecycle data leak or cross-tenant resume | `agentRuns` requires exact tenant plus account/user ownership; core lifecycle returns public redacted state only and CAS-resumes with current agent/revision |
185
203
  | Unbounded MCP HTTP | Bounded pre-parsed JSON, response bytes, concurrent requests, call timeout, SDK web-standard transport |
186
- | Cross-tenant operation | Authorizer derives ownership from validated auth and passes it to tool dispatch/selected workflow commands; never trust arguments as identity |
204
+ | Cross-tenant operation | Authorizer derives ownership from validated auth and passes it to tool/resource/prompt dispatch; stateful handler binds session to stable validated principal on every request; never trust arguments as identity |
205
+ | Sampling / elicitation authority | Host callbacks alone choose model/provider/credentials or obtain consent; bounded URL elicitation is returned to host UI and never fetched/opened automatically; auth tokens never enter callback params/results |
187
206
 
188
207
  For durable lifecycle exposure, construct `createAgentRunLifecycle({ checkpoints, resolveAgent })` in core, then pass selected entries as `agentRuns: { support: { lifecycle } }`. MCP registers two tools: `agent.support.status` accepts `{ runId, sessionId? }`; `agent.support.resume` accepts `{ runId, sessionId?, decision, expectedVersion }`. Do not expose an agent without durable checkpoints and a restart-safe `SessionStore`; no lifecycle tool appears by default.
189
208
 
@@ -191,9 +210,14 @@ MCP output is untrusted. Register bridge tools through core dispatch with a `Sec
191
210
 
192
211
  Discovery validation is atomic: cursor/page/tool/name/description/schema failures reject `refresh()` and preserve the previous immutable tool-array reference. The bridge intentionally uses raw SDK `request()` for `tools/list` and `tools/call`; this avoids eager Ajv compilation/validation of untrusted remote output schemas. Host `ToolValidator` remains the argument-validation owner.
193
212
 
213
+ ## Vendor web MCP prototype boundary
214
+
215
+ Official Exa/Firecrawl MCP servers may be tested only as explicit hardened prototypes: pin endpoint/origin/auth, inspect declared capabilities, allow-list individual tools/resources, retain all MCP bounds, and never expose generic remote passthrough. Production web research uses direct host-selected `@arnilo/prism-web-tools` adapters so provider choice, credentials, schema, and costs remain outside model control.
216
+
194
217
  ## Related APIs
195
218
 
196
219
  - [Tools](tools.md): registry, dispatch, validation
220
+ - [Web search, fetch, and extraction](web-tools.md): preferred direct bounded Brave/Exa/Firecrawl production path
197
221
  - [Tool execution primitives](tool-execution-primitives.md): Plan 055 design and conformance matrix
198
222
  - [Host security guide](host-security.md): permission, trust, validation checklist
199
223
  - [Web-standard server handler](server.md): agent/workflow HTTP routes and shared remote-boundary rules