@arnilo/prism 0.0.7 → 0.0.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +50 -2
- package/README.md +3 -1
- package/dist/agent-loops.js +14 -8
- package/dist/agents.js +37 -3
- package/dist/contracts.d.ts +17 -0
- package/dist/index.d.ts +4 -2
- package/dist/index.js +3 -2
- package/dist/provider-events.d.ts +2 -0
- package/dist/provider-events.js +21 -13
- package/dist/providers/openai-compatible.js +8 -5
- package/dist/providers/transport.d.ts +10 -1
- package/dist/providers/transport.js +24 -8
- package/dist/run-ledger.d.ts +21 -0
- package/dist/run-ledger.js +115 -0
- package/dist/tools.js +2 -0
- package/docs/a2a.md +61 -42
- package/docs/agent-events.md +5 -4
- package/docs/agent-loops.md +2 -2
- package/docs/agent-session-runtime.md +1 -0
- package/docs/browser-automation.md +124 -0
- package/docs/coding-agent-tools.md +111 -14
- package/docs/coding-security.md +84 -11
- package/docs/credential-storage.md +9 -0
- package/docs/database-persistence.md +1 -1
- package/docs/evaluations.md +38 -4
- package/docs/guardrails.md +3 -2
- package/docs/host-security.md +29 -4
- package/docs/index.md +20 -15
- package/docs/mcp-tools.md +29 -5
- package/docs/migration.md +104 -0
- package/docs/observability.md +26 -14
- package/docs/performance.md +54 -0
- package/docs/postgres-persistence.md +1 -0
- package/docs/provider-conformance.md +1 -1
- package/docs/provider-primitives.md +7 -1
- package/docs/providers/kimi.md +16 -2
- package/docs/providers/opencode-go.md +43 -2
- package/docs/release-and-install.md +117 -62
- package/docs/resource-loading.md +4 -0
- package/docs/review-coverage-2026-07-19-phase-3.md +174 -0
- package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
- package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
- package/docs/run-ledger-conformance.md +1 -0
- package/docs/runs-and-usage.md +17 -2
- package/docs/sqlite-persistence.md +1 -0
- package/docs/structured-output.md +2 -2
- package/docs/supervisors.md +2 -2
- package/docs/tools.md +5 -1
- package/docs/web-tools.md +78 -0
- package/docs/workflows.md +2 -0
- package/package.json +6 -4
package/docs/coding-security.md
CHANGED
|
@@ -2,12 +2,17 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects.
|
|
5
|
+
`@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools and one disposable Docker/OCI sandbox reference. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects, and optionally contains untrusted coding work in a host-invoked container.
|
|
6
6
|
|
|
7
7
|
| Export | Purpose |
|
|
8
8
|
| --- | --- |
|
|
9
9
|
| `createCodingApprovalPolicy(options)` | Returns an `ExecutionPolicy` with trusted roots, read-only mode, command allow/deny rules, approval caching, and timeout/abort-aware approval waits. |
|
|
10
10
|
| `createSandboxBashOperations(adapter)` | Maps a host-owned `SandboxAdapter` to coding-agent `BashOperations` for delegated shell execution. |
|
|
11
|
+
| `createSandboxCodingComposition(cwd, options)` | Authoritative construction: returns `{ tools, composition }` with required `workspaceMode` (`"host"` \| `"sandbox"`), fail-closed mixed wiring, and containment metadata. |
|
|
12
|
+
| `createSandboxReadOnlyComposition(cwd, options)` | Same contract for read-only tools (`read`/`repo_list`/`repo_search`). |
|
|
13
|
+
| `createSandboxCodingTools` / `createSandboxReadOnlyTools` | Thin wrappers that return `tools` only (compat); still require `workspaceMode`. |
|
|
14
|
+
| `createSandboxFilesystemOperations` / `createSandboxRepositoryOperations` | Optional execFile-backed FS/list/search backends for a disposable sandbox tree. |
|
|
15
|
+
| `createDockerSandbox(options)` | Creates one disposable non-root Docker container with read-only root/source, bounded tmpfs workspace, typed `execFile`, import/export, and stop/kill/cleanup. |
|
|
11
16
|
| `assertPathInsideRoots`, `isPathInsideReal` | Symlink-aware path containment helpers. |
|
|
12
17
|
| `evaluateCommandRules`, `hasShellMetacharacters` | Command classification helpers. |
|
|
13
18
|
|
|
@@ -21,7 +26,7 @@ import type { ExecutionAction, ExecutionPolicy, ExecutionDecision } from "@arnil
|
|
|
21
26
|
|
|
22
27
|
Use this package when coding tools need path scoping, human approval, command rules, or a pluggable sandbox backend. Wire the returned policy through `createCodingTools(cwd, { executionPolicy })` or per-tool `executionPolicy` options.
|
|
23
28
|
|
|
24
|
-
Prism does **not** claim OS-level isolation unless the host
|
|
29
|
+
Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
|
|
25
30
|
|
|
26
31
|
## Inputs / request
|
|
27
32
|
|
|
@@ -36,10 +41,39 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
|
|
|
36
41
|
|
|
37
42
|
`run` caching keys decisions by the tool execution context's `runId`; `session` uses `sessionId`. Coding tools pass both identities to the policy. A missing/empty identity disables caching for that check rather than creating a global bucket. Identical actions in different runs/sessions never share approvals or denials.
|
|
38
43
|
|
|
44
|
+
### Docker sandbox inputs
|
|
45
|
+
|
|
46
|
+
| Option | Default | Purpose |
|
|
47
|
+
| --- | --- | --- |
|
|
48
|
+
| `docker` | required | Absolute host Docker executable. |
|
|
49
|
+
| `image` | required | Digest-pinned image (`name@sha256:<64-hex>`). Never pulled (`--pull=never`). |
|
|
50
|
+
| `sourceRoot` | required | Absolute host directory imported into `/workspace`. |
|
|
51
|
+
| `user` | required | Non-root `uid:gid`. |
|
|
52
|
+
| `network` | `{ mode: "none" }` | Default no network; custom mode requires a pre-created network name and does not claim DNS containment. |
|
|
53
|
+
| `env` | `{}` | Exact allow-list only; host environment is never inherited. |
|
|
54
|
+
| `secrets` | `[]` | Canaries redacted from CLI/adapter errors. |
|
|
55
|
+
| `limits` | package defaults | CPU/memory/PID/FD/tmpfs/command/export/time caps validated before create. |
|
|
56
|
+
|
|
57
|
+
### Workspace mode inputs (`createSandboxCodingComposition`)
|
|
58
|
+
|
|
59
|
+
| Option | Default | Purpose |
|
|
60
|
+
| --- | --- | --- |
|
|
61
|
+
| `workspaceMode` | **required** | `"host"` (all tools on host cwd; never claims containment) or `"sandbox"` (shell + FS/list/search share one disposable tree). |
|
|
62
|
+
| `sandbox` | optional in host; required for sandbox unless custom ops supplied | `SandboxAdapter` / `DisposableSandbox`. |
|
|
63
|
+
| `workspaceRoot` | `"/workspace"` in sandbox mode | Tree root used as tool cwd when sandbox backends are bound. |
|
|
64
|
+
| `allowMixedWorkspaceWiring` | `false` | Escape hatch: allow sandbox shell + host FS backends. Records `composition.warnings`; forces `containmentClaim: false`. Missing hatch throws. |
|
|
65
|
+
| `read`/`write`/`edit`/`repository.operations` | auto-wired from `DisposableSandbox` in sandbox mode | Host may supply custom tree backends instead of auto-wire. |
|
|
66
|
+
|
|
67
|
+
`0.0.9` silent split (sandbox shell + host FS) is **superseded**. Mixed wiring is never the default.
|
|
68
|
+
|
|
39
69
|
## Outputs / response / events
|
|
40
70
|
|
|
41
71
|
`createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations`, receive `onData(Buffer)` for ordered stdout/stderr forwarding through the shell tool's existing bounded accumulator, and never grant policy approval themselves.
|
|
42
72
|
|
|
73
|
+
`createSandboxCodingComposition()` returns `{ tools, composition }` where `SandboxCodingComposition` carries `workspaceMode`, `containmentClaim`, `mixedWiringAllowed`, `warnings`, `workspaceRoot`, and optional `treeIdentity` (from `importIdentity` / `lastExportIdentity`). `containmentClaim` is `true` only for sandbox mode with tree backends bound and mixed wiring denied. Host mode and escape-hatch mixed wiring always set `containmentClaim: false` — never treat host mode as contained execution.
|
|
74
|
+
|
|
75
|
+
`createDockerSandbox()` returns a `DisposableSandbox`: typed `execFile(file, args)`, shell-compatible `exec`, `status`, cooperative `stop`, forced `kill`, and idempotent `close`. Import may surface `importIdentity`; successful export updates `lastExportIdentity`. `close({ export })` can stream a bounded workspace tar plus SHA-256/entry/byte metadata through a host callback; checkpoints should retain only host artifact references/hashes, never whole workspaces.
|
|
76
|
+
|
|
43
77
|
## Request/response example
|
|
44
78
|
|
|
45
79
|
```json
|
|
@@ -52,8 +86,12 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
|
|
|
52
86
|
## Implementation example
|
|
53
87
|
|
|
54
88
|
```ts
|
|
55
|
-
import {
|
|
56
|
-
|
|
89
|
+
import {
|
|
90
|
+
createCodingApprovalPolicy,
|
|
91
|
+
createDockerSandbox,
|
|
92
|
+
createSandboxCodingComposition,
|
|
93
|
+
} from "@arnilo/prism-coding-security";
|
|
94
|
+
import { createGitTools } from "@arnilo/prism-coding-agent";
|
|
57
95
|
|
|
58
96
|
const policy = createCodingApprovalPolicy({
|
|
59
97
|
roots: [workspaceRoot],
|
|
@@ -62,27 +100,62 @@ const policy = createCodingApprovalPolicy({
|
|
|
62
100
|
approvalTimeoutMs: 60_000,
|
|
63
101
|
});
|
|
64
102
|
|
|
65
|
-
const
|
|
103
|
+
const sandbox = await createDockerSandbox({
|
|
104
|
+
docker: "/usr/bin/docker",
|
|
105
|
+
image: "registry.example/prism-code@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
|
106
|
+
sourceRoot: "/srv/jobs/task-1/source",
|
|
107
|
+
user: "10001:10001",
|
|
108
|
+
network: { mode: "none" },
|
|
109
|
+
env: { CI: "1" },
|
|
110
|
+
limits: { cpus: 2, memoryBytes: 2 * 1024 ** 3, maxPids: 256, workspaceBytes: 1024 ** 3 },
|
|
111
|
+
});
|
|
112
|
+
|
|
113
|
+
// Sandbox mode: shell/read/write/edit/list/search share one disposable tree.
|
|
114
|
+
const { tools, composition } = createSandboxCodingComposition("/srv/jobs/task-1/source", {
|
|
115
|
+
workspaceMode: "sandbox",
|
|
116
|
+
sandbox,
|
|
66
117
|
executionPolicy: policy,
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
118
|
+
repository: { exclude: [".git", "node_modules", "dist"] },
|
|
119
|
+
});
|
|
120
|
+
// composition.containmentClaim === true when backends are bound
|
|
121
|
+
|
|
122
|
+
// Same-tree Git/check (opt-in; not folded into coding tools):
|
|
123
|
+
const gitTools = createGitTools(composition.workspaceRoot, {
|
|
124
|
+
execFile: sandbox.execFile.bind(sandbox),
|
|
125
|
+
commitIdentity: { name: "bot", email: "bot@example.com" },
|
|
126
|
+
});
|
|
127
|
+
|
|
128
|
+
// Host mode (explicit non-contained): omit sandbox; never claim containment.
|
|
129
|
+
const host = createSandboxCodingComposition(hostCwd, { workspaceMode: "host", executionPolicy: policy });
|
|
130
|
+
// host.composition.containmentClaim === false
|
|
131
|
+
|
|
132
|
+
await sandbox.execFile({ file: "npm", args: ["test"], cwd: "/workspace" });
|
|
133
|
+
await sandbox.close({
|
|
134
|
+
export: async (stream, meta) => hostArtifacts.write(stream, meta),
|
|
70
135
|
});
|
|
71
136
|
```
|
|
72
137
|
|
|
73
138
|
## Extension and configuration notes
|
|
74
139
|
|
|
75
|
-
Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter`
|
|
140
|
+
Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()`/`createSandboxCodingComposition()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` / `DisposableSandbox` are replaceable and host-owned; approval policy and sandboxing are separate layers. Custom remote sandboxes can implement `DisposableSandbox` without using Docker.
|
|
141
|
+
|
|
142
|
+
`createSandboxCodingComposition()` requires `workspaceMode`. Sandbox mode auto-wires FS/list/search through `DisposableSandbox.execFile` (or host-supplied custom operations) so mutations stay on the disposable tree until export. Host mode runs every coding tool against the host cwd and never sets `containmentClaim`. Sandbox shell + host FS throws unless `allowMixedWorkspaceWiring: true` (warnings + `containmentClaim: false`). Opt-in structured Git tools (`createGitTools(composition.workspaceRoot, { execFile: sandbox.execFile, commitIdentity })`) share the same tree/cwd; Prism still never pushes or opens PRs. Optional `@arnilo/prism-browser` can share the same disposable boundary: use `assertBrowserSandboxNetwork()` before browse-ready custom networks, and `createSharedSandboxBrowserOptions({ workspaceRoot, downloadsRoot, containedProxyAttestation })` so uploads/downloads align with `/workspace` and `/downloads`. Close the browser context before disposing the sandbox.
|
|
143
|
+
|
|
144
|
+
The Docker reference adapter starts by recorded container ID/label, uses argument arrays only, mounts source read-only, populates a size-bounded tmpfs `/workspace`, drops all capabilities, enables `no-new-privileges`, runs with `--init`, and never exposes the Docker socket, privileged mode, or host PID/IPC namespaces. Image pull/build/update stays outside Prism. Protected real-Docker checks are opt-in via `PRISM_TEST_DOCKER_SANDBOX=1` with host-supplied `PRISM_TEST_DOCKER_BIN` and digest-pinned `PRISM_TEST_DOCKER_IMAGE`.
|
|
76
145
|
|
|
77
146
|
Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
|
|
78
147
|
|
|
79
148
|
## Security and performance notes
|
|
80
149
|
|
|
81
|
-
Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe.
|
|
150
|
+
Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, repository list/search walks, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe.
|
|
151
|
+
|
|
152
|
+
Docker sandbox containment—not command regexes—enforces filesystem/network/process boundaries for the reference adapter. Network defaults to none; a custom Docker network still requires a host firewall/proxy for DNS/egress claims. Import rejects symlink escapes, devices, FIFOs, and sockets; export counts entries/bytes and hashes before host retention. Secrets in `secrets` are redacted from adapter errors and never exported as environment metadata. Unified workspace mode reuses existing sandbox/repo/coding hard caps and does not introduce unbounded host↔container sync loops. Host mode and `allowMixedWorkspaceWiring` never claim disposable containment. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter and Docker daemon.
|
|
82
153
|
|
|
83
154
|
## Related APIs
|
|
84
155
|
|
|
85
|
-
- [Coding agent tools](coding-agent-tools.md)
|
|
156
|
+
- [Coding agent tools](coding-agent-tools.md): durable plan/todo Markdown helpers and `state.coding` checkpoint metadata for restart/resume without a second runtime
|
|
157
|
+
- [Workflows](workflows.md): `runWorkflow` / `resumeWorkflow` / `startWorkflowBackground` composition for coding tasks
|
|
86
158
|
- [Host security guide](host-security.md)
|
|
159
|
+
- [Performance limits](performance.md)
|
|
87
160
|
- [Tool execution primitives](tool-execution-primitives.md)
|
|
88
161
|
- [Security/auth/trust](settings-auth-trust-security.md)
|
|
@@ -218,9 +218,18 @@ const providers = createOpenAIProviderPackage({ apiKey });
|
|
|
218
218
|
- Never log passphrases, derived keys, or decrypted credential payloads.
|
|
219
219
|
- Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
|
|
220
220
|
|
|
221
|
+
## MCP authentication boundary
|
|
222
|
+
|
|
223
|
+
MCP credentials remain host inputs: resolve them before constructing client `requestInit` or inside server `resolveAuthInfo`. Stateful server `resolveIdentity` receives validated SDK auth metadata only to derive a stable non-secret principal ID. Never copy access/refresh tokens into MCP resource/prompt/sampling/elicitation payloads, telemetry, errors, or session bindings; Prism does not refresh or persist MCP OAuth automatically.
|
|
224
|
+
|
|
225
|
+
## Web adapter credential boundary
|
|
226
|
+
|
|
227
|
+
`@arnilo/prism-web-tools` accepts explicit callbacks or `CredentialResolver`. Brave resolves `subscription_token`; Exa and Firecrawl resolve `api_key` immediately before each fixed-origin request. Keys never enter tool arguments/results, URLs, provider metadata, errors, telemetry, or prompts. Use separate least-privilege credentials and do not forward MCP/provider tokens between adapters.
|
|
228
|
+
|
|
221
229
|
## Related APIs
|
|
222
230
|
|
|
223
231
|
- [Credentials and redaction](credentials-and-redaction.md): core resolver helpers and `refreshOAuthCredential()`
|
|
232
|
+
- [Web search, fetch, and extraction](web-tools.md): late-bound Brave/Exa/Firecrawl credentials
|
|
224
233
|
- [Security/auth/trust](settings-auth-trust-security.md): host-owned settings/credentials boundaries
|
|
225
234
|
- [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 threat model and conformance matrix rows 7–10
|
|
226
235
|
- `@arnilo/prism`: `CredentialResolver`, `OAuthCredentialStore`, `createMemoryCredentialStore()`
|
|
@@ -241,7 +241,7 @@ Minimum production guidance:
|
|
|
241
241
|
|
|
242
242
|
- **Branch context:** implement `SessionStore.readBranchPath(query)` with an ancestor query / recursive CTE. Treat `SessionStore.list(sessionId)` as an O(n) development fallback only.
|
|
243
243
|
- **Cursor pagination:** every `query*` method should honor `cursor`, `limit`, and `order`. Encode cursors from indexed columns such as `(timestamp, id)`, `(started_at, id)`, `(recorded_at, id)`, or `(run_id, sequence)`; never use offset pagination for long sessions.
|
|
244
|
-
- **Batch appends:** `SessionStore.append()`
|
|
244
|
+
- **Batch appends:** `SessionStore.append()` stays single-entry because runtime advances one branch leaf at a time. Optional `createBatchedRunLedger()` wraps any ledger with bounded FIFO/backpressure and explicit `write_through`, `flush_on_terminal`, or crash-loss-capable `buffered` acknowledgement semantics; SQLite/PostgreSQL defaults remain direct durable writes.
|
|
245
245
|
- **Event sequence allocation:** allocate a monotonic `sequence` per `run_id` when inserting `prism_agent_events`. Use it with `run_id` for stable event timeline pagination when timestamps collide.
|
|
246
246
|
- **Run/event/usage query shapes:** runs page by `(session_id, started_at, id)` or `(branch_id, started_at, id)`; events page by `(run_id, sequence)` or `(session_id, timestamp, id)`; usage pages by `(run_id, recorded_at, id)` or `(session_id, recorded_at, id)`.
|
|
247
247
|
- **Host-owned sizing:** hosts own connection pools, transaction timeouts, page-size caps, queue/batch size, retention jobs, partitioning, and tenant/account/user isolation. Prism does not guess production limits.
|
package/docs/evaluations.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, live post-run scoring, and
|
|
5
|
+
`@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, bounded persistence-trace grading, explicit host model judges, pairwise comparisons, CI thresholds, live post-run scoring, and batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
|
|
6
6
|
|
|
7
7
|
## When to use it
|
|
8
8
|
|
|
@@ -18,6 +18,10 @@ Use this package when a host needs offline quality checks or sampled live scorin
|
|
|
18
18
|
| `runExperiment` | `agent`, dataset, scorers, bounded `concurrency`, optional store/ownership |
|
|
19
19
|
| `createMemoryEvaluationStore` | optional seed records |
|
|
20
20
|
| `appendEvaluationFeedback` | `RunFeedbackStore`, `EvaluationStore`, feedback fields, and 1–64 known evaluation IDs |
|
|
21
|
+
| `createPersistenceTraceResolver` | explicit `ProductionPersistenceStore`, exact session/run/ownership, page/byte bounds |
|
|
22
|
+
| `createModelJudge` | host judge callback, stable rubric/version, timeout/attempt/output bounds |
|
|
23
|
+
| `runComparison` | immutable dataset, 2–8 named candidates by default, pairwise scorers |
|
|
24
|
+
| `assertEvaluationThreshold` / `serializeEvaluationReport` | mean/failure/per-scorer gates and bounded redacted JSON |
|
|
21
25
|
|
|
22
26
|
## Outputs / response / events
|
|
23
27
|
|
|
@@ -112,11 +116,41 @@ console.log(report.aggregate.meanScore, linked.evaluationIds);
|
|
|
112
116
|
- Scorers receive result/item data only. Credentials, tools, and workspace access are not provided unless the host deliberately closes over them.
|
|
113
117
|
- Records pass through `SecretRedactor` / `secrets` before store append.
|
|
114
118
|
- Queries filter by ownership scope. Feedback linkage additionally requires tenant plus account/user and the feedback store re-verifies the run.
|
|
115
|
-
- Experiment concurrency defaults to `1` and is capped at `32`.
|
|
119
|
+
- Experiment concurrency defaults to `1` and is capped at `32`. Datasets cap at 10,000 items.
|
|
120
|
+
- Trace reads default to 100 rows × 20 pages with a 4 MiB aggregate cap (hard: 1,000 × 100 and 32 MiB). Repeated/missing cursors, identity drift, ownership drift, and overflow fail closed before scoring.
|
|
121
|
+
- Model judges are host callbacks, not providers: Prism passes rubric/version plus bounded target only—never credential resolvers, tools, or workspace. Defaults are one attempt, 30 seconds, and 16 KiB output; failures become redacted evaluation records.
|
|
122
|
+
- Pairwise candidates are sorted by name, executed once per item, compared in stable item/pair/scorer order, and record ties/failures without choosing a winner. Candidate and scorer outputs have byte caps.
|
|
123
|
+
- `assertEvaluationThreshold()` throws `ERR_PRISM_EVAL_THRESHOLD`; an uncaught error gives CI a non-zero exit. Keep model-judge/live gates credential-gated and outside the network-free default suite. `serializeEvaluationReport()` bounds/redacts checked-in artifacts.
|
|
124
|
+
|
|
125
|
+
## Trace, judge, comparison, and CI example
|
|
126
|
+
|
|
127
|
+
```ts
|
|
128
|
+
const traceResolver = createPersistenceTraceResolver(persistence);
|
|
129
|
+
const judge = createModelJudge({
|
|
130
|
+
id: "quality", rubric: "Score factual quality from 0 to 1", rubricVersion: "2026-07-20",
|
|
131
|
+
judge: hostStructuredJudge,
|
|
132
|
+
});
|
|
133
|
+
const evaluations = await scoreRun({ result, scorers: [judge], traceResolver, ownership });
|
|
134
|
+
const comparison = await runComparison({ dataset, candidates: { baseline, candidate }, scorers: [preference] });
|
|
135
|
+
assertEvaluationThreshold(report, { minimumMean: 0.9, maximumFailures: 0 });
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
`traceResolver` is explicit; no arbitrary run search occurs. `baseline`/`candidate` are host functions returning `AgentRunResult`. See `examples/evaluation-gate.ts` for a network-free gate and `examples/coding-browser-evaluation.ts` for coding/browser adversarial fixtures.
|
|
139
|
+
|
|
140
|
+
## Coding and browser adversarial evaluations (0.0.9)
|
|
141
|
+
|
|
142
|
+
Release 0.0.9 ships curated network-free adversarial fixtures in package tests:
|
|
143
|
+
|
|
144
|
+
- `@arnilo/prism-coding-agent` `eval-fixtures.test.ts`: safe native list vs shell, Git path/ref injection, dirty-tree rollback, unknown named-check failure, PR-handoff artifact completeness, and prompt-injection file content under read-only tools.
|
|
145
|
+
- `@arnilo/prism-browser` `eval-fixtures.test.ts`: stale snapshot refs, side-effect approval, private/loopback/file deny, upload/download/screenshot policy, CSS/evaluate target rejection, and hostile accessible-name text.
|
|
146
|
+
|
|
147
|
+
Fixtures reuse `@arnilo/prism-evals` (`defineDataset` / `defineScorer` / `scoreRun` / `assertEvaluationThreshold` / `serializeEvaluationReport`). Optional SWE-bench-compatible or live-browser harnesses remain host adapters — they are not default dependencies or quality claims. Protected real Docker/Playwright gates stay env-gated (`PRISM_TEST_DOCKER_SANDBOX`, `PRISM_LIVE_PLAYWRIGHT`) and never enter `sdk:ready`.
|
|
116
148
|
|
|
117
149
|
## Related APIs
|
|
118
150
|
|
|
119
151
|
- [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
|
|
120
152
|
- [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
|
|
121
|
-
- [Observability](observability.md):
|
|
122
|
-
- [
|
|
153
|
+
- [Observability](observability.md): use `onTraceReference` or bounded `traceId(runId)` to supply `ScoreRunOptions.traceId`; evaluation telemetry emits no reason/explanation content
|
|
154
|
+
- [Coding agent tools](coding-agent-tools.md) / [Browser automation](browser-automation.md) / [Workflows](workflows.md): network-free coding-task composition at `examples/durable-coding-workflow.ts`; adversarial coding/browser eval example at `examples/coding-browser-evaluation.ts`
|
|
155
|
+
- [Performance limits](performance.md): `scripts/benchmark-0.0.10.mjs` workspace-mode evidence and `scripts/benchmark-0.0.9.mjs` coding/browser evidence fields
|
|
156
|
+
- [Release and install](release-and-install.md): optional package install and protected sandbox-browser workflow
|
package/docs/guardrails.md
CHANGED
|
@@ -30,7 +30,7 @@ Decisions are `allow`, `block`, `tripwire`, or `interrupt`. Evaluation defaults
|
|
|
30
30
|
|
|
31
31
|
## Outputs / response / events
|
|
32
32
|
|
|
33
|
-
Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
|
|
33
|
+
Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. Optional OpenTelemetry instrumentation records only controlled stage/action on a short run-child span; guardrail name, reason, and metadata are excluded. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
|
|
34
34
|
|
|
35
35
|
Ordering is fixed:
|
|
36
36
|
|
|
@@ -65,11 +65,12 @@ Guardrails are callbacks supplied by the host. Prism does not discover, load, re
|
|
|
65
65
|
|
|
66
66
|
## Security and performance notes
|
|
67
67
|
|
|
68
|
-
Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work.
|
|
68
|
+
Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work. Browser snapshots and page text from `@arnilo/prism-browser` are untrusted external content: never allow them to modify tools, permissions, credentials, or policy. Browser mutations still require host `ExecutionPolicy`/approval; prompt-injection text in a page cannot grant upload/download release.
|
|
69
69
|
|
|
70
70
|
## Related APIs
|
|
71
71
|
|
|
72
72
|
- [Agent/session runtime](agent-session-runtime.md)
|
|
73
73
|
- [Tools](tools.md)
|
|
74
|
+
- [Browser automation](browser-automation.md)
|
|
74
75
|
- [Agent events](agent-events.md)
|
|
75
76
|
- [Host security](host-security.md)
|
package/docs/host-security.md
CHANGED
|
@@ -30,6 +30,7 @@ Start from explicit host inputs. Do not let runtime code discover security state
|
|
|
30
30
|
| Remote media policy | public/default pinned DNS or explicit trusted transport | `SsrfPolicy`, `resolveMediaContentBlock()` |
|
|
31
31
|
| Durable history | host database adapter | `SessionStore`, `assertSessionStoreConforms()` |
|
|
32
32
|
| Durable audit | host ledger adapter | `RunLedger`, `redactRunLedgerRecord()` |
|
|
33
|
+
| Telemetry | host OpenTelemetry SDK/exporter | metadata-only adapter, controlled metric labels, `onTraceReference` |
|
|
33
34
|
| Durable interruption | host checkpoint + session stores, exact ownership | `RunOptions.runState`, `resumeAgentRun()`, `createAgentRunLifecycle()`, `createSecureAgent()` |
|
|
34
35
|
| Extensions | explicit package imports only | `createExtensionKernel`, `ExtensionAPI` |
|
|
35
36
|
| Remote agent/workflow API | host authentication + ownership mapping | `@arnilo/prism-server`, `createPrismHandler()` |
|
|
@@ -45,7 +46,7 @@ Security controls fail closed before side effects when wired at the guarded edge
|
|
|
45
46
|
- validator failures emit `tool_execution_blocked` with `validation_failed`
|
|
46
47
|
- configured guardrails fail closed; output stages buffer blocked provider/tool content before events, ledgers, session entries, or MCP responses
|
|
47
48
|
- configured redactors scrub provider requests, agent events, session entries, ledger records, tool errors, extension errors, injector context, and durable run checkpoints
|
|
48
|
-
- durable resume requires host-derived exact ownership and checkpoint version; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
|
|
49
|
+
- durable resume requires host-derived exact ownership and checkpoint version; coding-task resume also revalidates plan/workspace artifact hashes plus tool/policy/image fingerprints via `assertCodingResumeAllowed` before import; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
|
|
49
50
|
|
|
50
51
|
These checks are explicit function calls during load, assembly, dispatch, append, or run handling. Prism adds no background watchers, filesystem scanners, network probes, credential polling, or automatic extension discovery.
|
|
51
52
|
|
|
@@ -120,6 +121,8 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
120
121
|
- Use `createContributionRegistries({ duplicate: "error" })` and prefixed names for third-party packages to prevent silent shadowing.
|
|
121
122
|
- Extension contributions are inert until selected. Loading an extension package runs its `setup(api)` code, so hosts should load only trusted packages or isolate untrusted code outside Prism.
|
|
122
123
|
- Skills and instruction injectors grant no tools, permissions, validators, or resource access. Host-active tools and permission policies still decide execution.
|
|
124
|
+
- Optional ledger batching accepts runtime-redacted records only. Prefer `flush_on_terminal`; `buffered` explicitly permits crash-before-flush loss. Flush failures propagate; hosts can call `dispose({ flush: false })` to clear queued objects when deliberately discarding an aborted buffered workload.
|
|
125
|
+
- Session snapshot cache holds one session-local leaf for at most one second and invalidates after committed mutation, checkout, compaction, and resume; it never crosses session/branch/ownership.
|
|
123
126
|
- For production persistence, implement a database-backed `SessionStore`/`RunLedger`, run `assertSessionStoreConforms()` against the store, and follow the database schema guidance. Do not ship provider instances, credential resolvers, or secrets into durable rows.
|
|
124
127
|
|
|
125
128
|
## Security and performance notes
|
|
@@ -130,10 +133,12 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
130
133
|
- Known secrets must be passed into redactors before data is emitted or persisted. Redact again in host adapters if they transform records after Prism redaction.
|
|
131
134
|
- Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects. Its untrusted-schema adapter rejects non-local refs, forbidden keys/cycles/non-finite values and bounds bytes/depth/properties/keywords/refs plus its LRU cache before Ajv compilation; do not raise caps above documented hard limits.
|
|
132
135
|
- Treat embeddings as untrusted numeric input. `@arnilo/prism-memory` rejects empty, non-number, NaN, and infinite vectors before in-memory similarity or pgvector parameters; custom `Embedder`/`VectorStore` implementations must retain the same boundary.
|
|
136
|
+
- Evaluation trace readers require exact supplied ownership plus session/run identity, reject cursor/identity drift, and redact before bounded scorer/judge input. Model-judge callbacks receive no credential resolver, tools, or workspace; keep live judges outside default CI and redact report artifacts.
|
|
133
137
|
- Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
|
|
134
|
-
- MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands, requires per-
|
|
138
|
+
- MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands/resources/prompts, requires per-operation `authorize`, and retains core gates. Sampling, roots, model/credential selection, and elicitation consent stay host-owned; URL elicitation is never opened automatically. Stateful web mode requires host `resolveAuthInfo` plus `resolveIdentity`, exact origin policy, and binds every POST/GET/DELETE/SSE request to one non-secret principal; mismatches return 404. Handler still needs TLS and edge rate limiting. See [MCP client/server exposure](mcp-tools.md).
|
|
135
139
|
- `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
|
|
136
|
-
- Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules,
|
|
140
|
+
- Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, repository list/search depth/entry/match/scan/time caps, structured Git path/ref/message/output/patch/worktree caps, named-check concurrency/output caps, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Opt-in `createGitTools()` uses argument arrays with hooks/credential prompts/external diff disabled, requires host `commitIdentity` for commits, and never pushes or opens PRs. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell/repository backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, identity-scoped approval caching, required `workspaceMode` on `createSandboxCodingComposition()` / `createSandboxCodingTools()`, and the optional `createDockerSandbox()` reference adapter. **Host mode is never contained execution** (`containmentClaim: false`). Sandbox mode claims containment only when FS backends target the disposable tree; mixed wiring requires `allowMixedWorkspaceWiring` and still does not claim containment. Limits alone are not containment: construct the Docker adapter (absolute CLI, digest-pinned image, network none by default) or an equivalent host sandbox before treating coding execution as production-safe. Docker daemon/image trust, egress firewall/proxy, and artifact retention remain host-owned.
|
|
141
|
+
- Optional `@arnilo/prism-browser` requires a host-supplied Playwright Browser (`playwright-core@1.61.0` peer). Import is inert. One non-persistent context belongs to one run; actions serialize; refs are snapshot-scoped; CSS/evaluate/CDP/persistent profiles are denied. Context routing + `serviceWorkers: "block"` deny file/data/blob/devtools/private/loopback by default and require contained-proxy attestation for external egress (Playwright routing is defense in depth, not DNS containment). Uploads are realpath-rooted; downloads quarantine with hash/MIME until host `approveRelease`; screenshots return bounded `ImageContent`. Observation vs mutation/high-impact actions map to `ExecutionPolicy`. Treat snapshot/page text as untrusted external content. Close contexts with `browser_close` or `manager.closeRun(runId)` on abort/terminal. Browser control endpoint, binary/image pin, and real egress firewall/proxy remain host-owned. Shared sandbox: `createSharedSandboxBrowserOptions()` + `assertBrowserSandboxNetwork()`.
|
|
137
142
|
- `@arnilo/prism-credentials-node` rejects oversized/malformed envelopes and excessive scrypt work before KDF allocation, uses async scrypt, and requires restrictive existing/new Unix vault modes. Keep vault ownership and parent-directory access host-controlled; review before `chmod 600`, never auto-weaken a file policy. Keychain calls use abort-aware native async work with finite timeout/payload caps and sanitized errors. OS prompts, service availability, and whether a native backend promptly honors cancellation remain host/platform boundaries; no plaintext fallback is attempted.
|
|
138
143
|
- LLM compaction always sends finite summary `maxTokens`, retains bounded deltas/events, and bounds/redacts provider/factory/policy error detail. Observational-memory workers cap turns, calls, arguments, results, transcript, and surfaced errors; unknown tools fail before execution, while invalid results can only be rejected after a host tool returns and may therefore follow side effects. Pass all known provider/credential/tool secrets into compaction/runtime options; exact replacement is not secret discovery.
|
|
139
144
|
- Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
|
|
@@ -160,7 +165,27 @@ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy
|
|
|
160
165
|
- Keep depth, active children, input, turn/tool/token, timeout, and queue ceilings finite; propagate abort through nested calls.
|
|
161
166
|
- Expose A2A only behind per-request authentication/authorization, TLS, edge rate limits, and replay policy. Public card discovery grants no invoke access.
|
|
162
167
|
- Remote A2A endpoints/card URLs require exact HTTPS origin allow-lists and redirect rejection. Pin ES256 card keys/expiry; never auto-fetch untrusted `jku`.
|
|
163
|
-
- Treat cards, task status, errors, artifacts, and SSE frames as untrusted bounded input and redact before logs/hooks/events.
|
|
168
|
+
- Treat cards, rich parts, task status/history, errors, artifacts, and SSE replay frames as untrusted bounded input and redact before logs/hooks/events. URL parts require host public/pinned-network policy and are never auto-fetched. Durable task/push adapters repeat exact-owner checks; foreign/missing records share not-found responses.
|
|
169
|
+
- Push delivery remains host-owned: validate every attempt/redirect against SSRF/rebinding policy, cap retries/time/output, authenticate webhook payloads, deduplicate event IDs, and keep token/auth credentials out of configs returned over A2A. Client streaming uses fatal UTF-8 decoding and rejects partial/post-terminal frames.
|
|
170
|
+
|
|
171
|
+
## Web research boundaries
|
|
172
|
+
|
|
173
|
+
- Construct `@arnilo/prism-web-tools` with one host-selected Brave or Exa adapter; never expose adapter/provider/credential/schema selection to model arguments.
|
|
174
|
+
- Provider API origins are fixed exact HTTPS origins and redirects fail. Credentials resolve immediately before I/O; remote bodies and secrets are excluded from errors/results/telemetry.
|
|
175
|
+
- Firecrawl targets reject userinfo, non-HTTP(S), private literals, and policy-denied hosts. Supply `validateUrl` for host DNS/rebinding/egress checks. Firecrawl performs remote retrieval, so Prism cannot pin target DNS after handoff.
|
|
176
|
+
- Treat every snippet, highlight, Markdown byte, metadata field, and extracted JSON value as prompt-injection-capable untrusted data. Never elevate it into system instructions or let it modify tools, permissions, trust, credentials, routing, or extraction schema.
|
|
177
|
+
- Keep counts/bytes/retries/rate delays/polling/concurrency/wall time finite. Live credentials belong only in explicit protected `PRISM_LIVE_WEB=1` runs.
|
|
178
|
+
|
|
179
|
+
## Supply-chain and live-canary boundaries
|
|
180
|
+
|
|
181
|
+
- Require `security / codeql`, `security / supply-chain`, PR dependency review, release readiness, and PostgreSQL integration in protected-branch rules. Enable GitHub secret scanning and push protection as repository settings; checked-in workflows cannot enable those service controls.
|
|
182
|
+
- Actions are pinned to full commit revisions. Dependabot proposes weekly npm/action revision changes; review upstream release notes before merge rather than replacing pins with moving tags.
|
|
183
|
+
- `scripts/verify-sbom.mjs` accepts only bounded SPDX 2.3 inventory with exact checked-in permissive licenses. Any missing/new expression fails until reviewed; do not widen policy merely to unblock CI.
|
|
184
|
+
- `scripts/scan-secrets.mjs` checks tracked source and unpacked public tarballs for high-confidence credential/private-key forms without printing matched values. It complements GitHub secret scanning; it is not entropy scanning or DLP.
|
|
185
|
+
- Tag publication alone receives npm/OIDC/attestation permissions. Untrusted pull-request code receives no canary, npm, or OIDC secret and no workflow uses `pull_request_target`.
|
|
186
|
+
- Scheduled/manual canaries run only in protected `live-canaries` environment. Use dedicated read-only/low-quota credentials and provider account spend limits. Runner performs four probes, at most one MCP cleanup, one provider output token, one Brave result, 64-KiB responses, and finite timeouts; report excludes endpoints, headers, bodies, credentials, and MCP session IDs.
|
|
187
|
+
- Scheduled/manual coding/browser containment checks run in protected `sandbox-browser` environment (`.github/workflows/sandbox-browser.yml`). They receive no provider/npm/OIDC secrets; Docker/Playwright enablement is variable-gated with host-preloaded digest-pinned images/binaries; uploads are redacted aggregate status only.
|
|
188
|
+
- Live endpoint operators own TLS, egress allow-lists, account-dollar budget, cleanup beyond MCP session DELETE, and revocation. Failed canaries log only operation kind plus status/timeout; inspect provider-side audit logs for details.
|
|
164
189
|
|
|
165
190
|
## Related APIs
|
|
166
191
|
|
package/docs/index.md
CHANGED
|
@@ -10,11 +10,11 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
10
10
|
- [Agent definitions](agent-definitions.md): resolve declarative `AgentDefinition` values via `resolveAgentDefinition`, and turn app-config `<configRoot>/agents/<name>/AGENT.md` bundles into runnable agents via `discoverAgentBundles` / `resolveAgentBundle` (explicit tool/skill activation by name, fail-closed omitted capabilities, migration-only `activateAllCapabilities`, strict duplicate scope checks, configurable prompt layers, no auto-discovery).
|
|
11
11
|
- [Agent loops](agent-loops.md): replaceable per-run control loops — `singleShotLoop` default and opt-in bounded artifact-loop tool rounds with host-supplied `validator`/`parser`/`repairer` callbacks.
|
|
12
12
|
- [Guardrails](guardrails.md): typed fail-closed input/output/tool checks with buffered provider output and redacted decision records.
|
|
13
|
-
- [Agent events](agent-events.md):
|
|
14
|
-
- [Observability](observability.md):
|
|
15
|
-
- [Evaluations](evaluations.md):
|
|
16
|
-
- [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence
|
|
17
|
-
- [Performance limits](performance.md): bounded live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
|
|
13
|
+
- [Agent events](agent-events.md): redacted lifecycle stream used by UIs, ledgers, and metadata-only parented telemetry; message/progress deltas never create spans.
|
|
14
|
+
- [Observability](observability.md): OTel GenAI agent/provider/tool hierarchy, host context parenting, bounded trace linkage, safe evaluation events, controlled metrics, and exporter isolation.
|
|
15
|
+
- [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, coding/browser adversarial fixtures, and ID-only linkage to immutable owned run feedback.
|
|
16
|
+
- [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence, optional bounded FIFO durability policies, session snapshot caching, and immutable run/trace feedback.
|
|
17
|
+
- [Performance limits](performance.md): bounded evaluation traces/judges/reports, 0.0.10 workspace-mode and 0.0.9 coding/browser benchmark evidence, security scan/live-canary backstops, live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
|
|
18
18
|
- [Structured output](structured-output.md): the `Artifact*` seam plus provider-native `StructuredOutputOptions` / `structuredOutputMode` for capable models.
|
|
19
19
|
|
|
20
20
|
## Compaction/session memory
|
|
@@ -27,7 +27,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
27
27
|
- [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention, and NoSQL mapping.
|
|
28
28
|
- [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, and transactionally verified/backfilled migration-v3 metadata.
|
|
29
29
|
- [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
|
|
30
|
-
- [Migration guide](migration.md): 0.0.3 compatibility
|
|
30
|
+
- [Migration guide](migration.md): 0.0.3 compatibility through 0.0.10 workspace modes (required `workspaceMode`, fail-closed mixed wiring) and 0.0.9 coding/browser sandbox, repository/Git, durable plans, and Playwright automation changes.
|
|
31
31
|
- [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety.
|
|
32
32
|
- [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
|
|
33
33
|
|
|
@@ -57,9 +57,11 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
57
57
|
- [Tools](tools.md): register host-owned active tools with replace-or-error duplicate policy, apply exact allow/deny filtering, dispatch normal or opt-in bounded artifact-loop calls, and optionally bound untrusted JSON Schema compilation.
|
|
58
58
|
- [Tool execution primitives](tool-execution-primitives.md): finite JSON Schema LRU validation, exclusive-aware bounded parallel dispatch, MCP bridge mapping, coding execution policy, and image-read bounds.
|
|
59
59
|
- [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
|
|
60
|
-
- [MCP client bridge and server exposure](mcp-tools.md):
|
|
61
|
-
- [
|
|
62
|
-
- [
|
|
60
|
+
- [MCP client bridge and server exposure](mcp-tools.md): SDK-1.29.0 bounded tools/resources/prompts, host-owned roots/sampling/elicitation, exact-origin DNS-pinned client transport, and principal-bound opt-in Streamable HTTP sessions.
|
|
61
|
+
- [Web search, fetch, and extraction](web-tools.md): optional host-selected Brave/Exa discovery and Firecrawl Markdown/schema tools with native fetch, stable citations, late credentials, finite limits, and explicit untrusted-content boundaries.
|
|
62
|
+
- [Browser automation](browser-automation.md): optional `@arnilo/prism-browser` with host-supplied Playwright contexts, AI-mode snapshots/refs, ordered `browser_open`/`browser_snapshot`/`browser_act`/`browser_close`, egress/side-effect/upload/download/screenshot policy, and finite page/action/snapshot/network/artifact caps.
|
|
63
|
+
- [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, `edit`, `repo_list`, and `repo_search` definitions plus opt-in `createGitTools()` / `coding_check` for structured Git status/diff/branch/worktree/apply/commit/PR-handoff and named checks; durable plan/todo Markdown helpers with workflow `state.coding` checkpoint metadata; streamed text pages, bounded repository list/search, finite Git/check/plan caps, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
|
|
64
|
+
- [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, required `workspaceMode` (`host`/`sandbox`) with fail-closed mixed wiring, `createSandboxCodingComposition()` containment metadata, and the disposable Docker/OCI sandbox reference with bounded workspace import/export.
|
|
63
65
|
|
|
64
66
|
## Extensions/plugins
|
|
65
67
|
- [Contribution discovery (workspace)](contribution-discovery.md): opt-in, realpath-contained directory scanner turning `SKILL.md`/`manifest.json` into inert `DiscoveredContribution` envelopes the host registers — no `import()`, no auto-activate, no provider scanning. Per-agent bundles remain app-controlled and are documented under Agent/session runtime.
|
|
@@ -77,17 +79,17 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
77
79
|
- [Web-standard server handler](server.md): optional framework-free authorized direct/SSE agent, explicitly selected durable agent lifecycle, and durable workflow routes with explicit bounds and zero default exposure.
|
|
78
80
|
|
|
79
81
|
## Multi-agent and interoperability
|
|
80
|
-
- [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, and
|
|
81
|
-
- [A2A interoperability](a2a.md):
|
|
82
|
+
- [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, finite budgets, host-projected delegation telemetry, and separate A2A durable adapter boundary.
|
|
83
|
+
- [A2A interoperability](a2a.md): A2A 1.0 JSON-RPC/HTTPS cards plus host-owned durable task get/list/cancel/subscribe, bounded rich parts/replay, principal-scoped push configs, and exact-origin verified client.
|
|
82
84
|
|
|
83
85
|
## CLI/RPC
|
|
84
86
|
- [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
|
|
85
|
-
- [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Interactive TUI (C-012) deferred.
|
|
87
|
+
- [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Compose coding plans/checkpoints via workspace Markdown + `state.coding` without a second runtime. Interactive TUI (C-012) deferred.
|
|
86
88
|
- [Workflow orchestration primitives](workflow-orchestration-primitives.md): architecture inventory — workflow adapters consume core `CheckpointStore`, `LeaseStore`, and bounded `EventMultiplexer`; run control and optional RPC commands stay package-local.
|
|
87
89
|
- [Workflow/TUI scope](workflow-tui-primitives.md): records why 0.0.5 ships workflow APIs/RPC control but no interactive terminal UI.
|
|
88
90
|
|
|
89
91
|
## Security and credentials
|
|
90
|
-
- [Host security guide](host-security.md): fail-closed checklist for
|
|
92
|
+
- [Host security guide](host-security.md): fail-closed checklist for supply-chain/attestation/canary isolation, bounded credentials, JSON/schema/vector/crypto, MCP/A2A/web remote boundaries, untrusted external content, settings, redaction, trust roots, workflow ownership, coding I/O, permissions, persistence, extensions, and tool validation.
|
|
91
93
|
- [Security/auth/trust](settings-auth-trust-security.md): settings providers, credential helpers, trust/permission policies, redaction controls, host-owned settings/credentials wiring outside `AgentConfig`, and security-boundary hardening summary.
|
|
92
94
|
- [Credentials and redaction](credentials-and-redaction.md): compose explicit credential resolver order, use caller-supplied env objects/OAuth refresh helpers, resolve credentials only at the provider edge, and redact known secret values.
|
|
93
95
|
- [Credential storage](credential-storage.md): optional `@arnilo/prism-credentials-node` adapter with strict bounded AES-GCM envelopes, async finite scrypt, restrictive Unix files, and abort-aware bounded system-keychain calls.
|
|
@@ -96,14 +98,17 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
96
98
|
- Provider test doubles: `createMockProvider()` and provider event helpers are documented on the canonical Provider layer page above.
|
|
97
99
|
- [Provider conformance](provider-conformance.md): run network-free provider adapter assertions (stream order, abort, tool-call reconstruction, cache usage, content coverage, protected header ownership, secret leak) from `@arnilo/prism/testing/provider-conformance`.
|
|
98
100
|
- [Session store conformance](session-store-conformance.md): assert any `SessionStore` adapter satisfies append/idempotency/conflict/branch invariants from `@arnilo/prism/testing/session-store-conformance`.
|
|
99
|
-
- [Run ledger conformance](run-ledger-conformance.md): assert durable run/event/tool/usage writes and reopen survival. Run-feedback stores use `@arnilo/prism/testing/feedback` for append/query/delete/ownership linkage conformance.
|
|
101
|
+
- [Run ledger conformance](run-ledger-conformance.md): assert durable run/event/tool/usage writes and reopen survival; batch-wrapper FIFO/bounds/flush checks remain separate. Run-feedback stores use `@arnilo/prism/testing/feedback` for append/query/delete/ownership linkage conformance.
|
|
100
102
|
- [Compaction conformance](compaction-conformance.md): assert any `CompactionStrategy` returns a non-empty redacted summary and observes abort from `@arnilo/prism/testing/compaction-conformance`.
|
|
101
103
|
- [Tool conformance](tool-conformance.md): assert the tool-dispatch blocked-reason matrix (unknown/denied/invalid/permission/validator) and success path from `@arnilo/prism/testing/tool-conformance`.
|
|
102
104
|
- [Extension conformance](extension-conformance.md): assert an `Extension` setup runs, contributions stay inert, and setup errors are redacted or rethrown from `@arnilo/prism/testing/extension-conformance`.
|
|
103
105
|
- `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, cache-aware prompt assembly, NeuralWatt agent run, stores/branching, compaction, observational-memory recall, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
|
|
104
106
|
|
|
105
107
|
## Release and install
|
|
106
|
-
- [Release and install](release-and-install.md):
|
|
108
|
+
- [Release and install](release-and-install.md): 32-package graph (including optional browser), install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, protected live canaries, and sandbox-browser Docker/Playwright gates.
|
|
109
|
+
- [Review coverage (2026-07-21 Phase 5)](review-coverage-2026-07-21-phase-5.md): Plan 073 evidence freeze — unified workspace modes, primitive ownership, reused finite limits, threats, and 0.0.10 release gates.
|
|
110
|
+
- [Review coverage (2026-07-20 Phase 4)](review-coverage-2026-07-20-phase-4.md): Plan 072 evidence freeze — revised coding/browser-only scope, external revisions, primitive ownership, finite limits, threats, and 0.0.9 release gates.
|
|
111
|
+
- [Review coverage (2026-07-19 Phase 3)](review-coverage-2026-07-19-phase-3.md): Plan 070 evidence freeze — exact protocol/vendor references, capability/primitive/limit matrices, supported boundaries, and 0.0.8 release evidence.
|
|
107
112
|
- [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md): Plan 067 evidence freeze — P0–P2 re-verification owners, seven first-party provider packages mapped to official-doc URLs, Pi secondary refs, cache/thinking/discovery surfaces, credential canaries, and use-case model-binding inventory.
|
|
108
113
|
- [Review coverage (2026-07-15)](review-coverage-2026-07-15.md): frozen 0.0.5 finding/feature ownership, existing-primitive inventory, package decisions, threat boundaries, exclusions, and measured Phase 0 baseline.
|
|
109
114
|
- [Review coverage (2026-07-14)](review-coverage-2026-07-14.md): traceability matrix linking review findings and bug-report fixes to plan tasks, tests, and documentation for release 0.0.4.
|
package/docs/mcp-tools.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-mcp` has two explicit directions. Its client bridge connects hosts to remote [Model Context Protocol](https://modelcontextprotocol.io) servers and maps discovered tools to ordinary `ToolDefinition`s. Its server API registers selected Prism `ToolDefinition` and `CommandDefinition` values on the official SDK `McpServer`, with required authorization and a bounded optional Web-standard Streamable HTTP handler. The package
|
|
5
|
+
`@arnilo/prism-mcp` has two explicit directions. Its client bridge connects hosts to remote [Model Context Protocol](https://modelcontextprotocol.io) servers and maps discovered tools to ordinary `ToolDefinition`s. Its server API registers selected Prism `ToolDefinition` and `CommandDefinition` values on the official SDK `McpServer`, with required authorization and a bounded optional Web-standard Streamable HTTP handler. The package pins `@modelcontextprotocol/sdk` **1.29.0** (MCP protocol negotiation remains SDK-owned) and adds no MCP branch to core Prism.
|
|
6
6
|
|
|
7
7
|
Primary API:
|
|
8
8
|
|
|
@@ -19,7 +19,21 @@ await bridge.refresh(); // re-list after notifications or TTL expiry
|
|
|
19
19
|
await bridge.close(); // close client + transport
|
|
20
20
|
```
|
|
21
21
|
|
|
22
|
-
Advanced hosts that manage their own `Client` + `Transport` can call `attachMcpToolBridge(
|
|
22
|
+
Advanced hosts that manage their own `Client` + `Transport` can call `attachMcpToolBridge()` or `attachMcpCapabilities()` after connect. `connectMcpCapabilities()` keeps resources/prompts as host-facing facades rather than converting them into model tools, and declares roots/sampling/elicitation only when callbacks are supplied.
|
|
23
|
+
|
|
24
|
+
```ts
|
|
25
|
+
const bridge = await connectMcpCapabilities({
|
|
26
|
+
serverId: "research",
|
|
27
|
+
transport: { type: "streamable-http", url, allowedOrigins: [origin] },
|
|
28
|
+
roots: () => [{ uri: "file:///workspace", name: "workspace" }],
|
|
29
|
+
sampling: hostSampling, // host selects model/provider/credentials
|
|
30
|
+
elicitation: hostElicitation, // URL mode returns approval; Prism never opens/fetches URL
|
|
31
|
+
});
|
|
32
|
+
await bridge.listResources();
|
|
33
|
+
await bridge.getPrompt("review", { topic: "security" });
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
Server capability matrix for SDK 1.29.0: tools/resources/prompts and their list-change notifications are supported through official registrations; roots/sampling/form+URL elicitation are supported as explicit client callbacks. Missing server resources/prompts throw `McpUnsupportedCapabilityError` with `ERR_PRISM_MCP_UNSUPPORTED_CAPABILITY`. Resource/prompt results and sampling/elicitation inputs/results are bounded JSON. Accepted form/URL elicitation requires host-only `humanInteraction: true`; bridge strips marker before protocol output and fails closed when absent. Automatic root discovery/consent, model selection, credential resolution, URL navigation, generic command proxying, and custom JSON-RPC are unsupported.
|
|
23
37
|
|
|
24
38
|
Server direction:
|
|
25
39
|
|
|
@@ -41,10 +55,13 @@ const handleMcp = await createPrismMcpWebHandler(server, {
|
|
|
41
55
|
resolveAuthInfo: authenticateRequest,
|
|
42
56
|
allowedHosts: ["api.example.test"],
|
|
43
57
|
allowedOrigins: ["https://app.example.test"],
|
|
58
|
+
// Omit these two for bounded stateless JSON mode.
|
|
59
|
+
sessionIdGenerator: crypto.randomUUID,
|
|
60
|
+
resolveIdentity: (_request, auth) => auth ? { id: validatedPrincipalId(auth) } : false,
|
|
44
61
|
});
|
|
45
62
|
```
|
|
46
63
|
|
|
47
|
-
`McpServer.connect(transport)` remains available for SDK stdio or in-memory transports. The helper uses SDK `WebStandardStreamableHTTPServerTransport
|
|
64
|
+
`McpServer.connect(transport)` remains available for SDK stdio or in-memory transports. The helper uses SDK `WebStandardStreamableHTTPServerTransport`; it does not start a listener. Default remains bounded stateless JSON-response mode. Supplying `sessionIdGenerator` enables SDK `MCP-Session-Id` POST/GET/DELETE/SSE lifecycle and requires exact `allowedOrigins` plus host `resolveIdentity`. Every request re-authenticates, and a different principal receives non-disclosing 404. SDK owns protocol-version/session headers and SSE semantics. SDK 1.29.0's in-memory event store is not enabled, so `Last-Event-ID` replay is explicitly unsupported; reconnect starts only through SDK-supported active session GET.
|
|
48
65
|
|
|
49
66
|
## When to use it
|
|
50
67
|
|
|
@@ -160,6 +177,7 @@ Plaintext is accepted only when `allowLoopbackHttp: true`, the URL hostname is l
|
|
|
160
177
|
| Option | Default | Purpose |
|
|
161
178
|
| --- | --- | --- |
|
|
162
179
|
| `tools` / `commands` | empty | Explicit allow-list; zero default exposure |
|
|
180
|
+
| `resources` / `prompts` | empty | Static URI/name registrations with bounded host callbacks and per-read/get authorization |
|
|
163
181
|
| `agentRuns` | empty | Explicit `{ [agentId]: { lifecycle } }` map; registers `agent.<id>.status` and `agent.<id>.resume` only |
|
|
164
182
|
| `authorize` | required | Per-call host authz using SDK auth/session metadata |
|
|
165
183
|
| `permission` / `validate` / `redactor` | none | Core tool-dispatch gates and known-secret redaction |
|
|
@@ -167,7 +185,7 @@ Plaintext is accepted only when `allowLoopbackHttp: true`, the URL hostname is l
|
|
|
167
185
|
| `maxConcurrentCalls` | 16 (256 hard) | Bound active tool/command execution |
|
|
168
186
|
| `callTimeoutMs` | 60 s (30 min hard) | Abort and return timed-out calls |
|
|
169
187
|
|
|
170
|
-
Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard), 32 concurrent requests (512 hard),
|
|
188
|
+
Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard), 32 concurrent requests (512 hard), 60 s timeout (30 min hard), and 32 sessions (512 hard). Stateful mode is intentionally one official SDK transport/session lineage per handler; use one handler/server instance per independently hosted endpoint when multi-tenant transport isolation is required. It parses bounded JSON before passing `parsedBody` to the SDK transport. `allowedHosts`/`allowedOrigins` activate SDK DNS-rebinding checks only when explicitly configured. Authentication data comes only from host `resolveAuthInfo()`.
|
|
171
189
|
|
|
172
190
|
## Security and performance notes
|
|
173
191
|
|
|
@@ -183,7 +201,8 @@ Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard),
|
|
|
183
201
|
| Accidental server exposure | Empty default arrays/maps, duplicate-name rejection, explicit tools/commands/lifecycle only |
|
|
184
202
|
| Agent lifecycle data leak or cross-tenant resume | `agentRuns` requires exact tenant plus account/user ownership; core lifecycle returns public redacted state only and CAS-resumes with current agent/revision |
|
|
185
203
|
| Unbounded MCP HTTP | Bounded pre-parsed JSON, response bytes, concurrent requests, call timeout, SDK web-standard transport |
|
|
186
|
-
| Cross-tenant operation | Authorizer derives ownership from validated auth and passes it to tool dispatch
|
|
204
|
+
| Cross-tenant operation | Authorizer derives ownership from validated auth and passes it to tool/resource/prompt dispatch; stateful handler binds session to stable validated principal on every request; never trust arguments as identity |
|
|
205
|
+
| Sampling / elicitation authority | Host callbacks alone choose model/provider/credentials or obtain consent; bounded URL elicitation is returned to host UI and never fetched/opened automatically; auth tokens never enter callback params/results |
|
|
187
206
|
|
|
188
207
|
For durable lifecycle exposure, construct `createAgentRunLifecycle({ checkpoints, resolveAgent })` in core, then pass selected entries as `agentRuns: { support: { lifecycle } }`. MCP registers two tools: `agent.support.status` accepts `{ runId, sessionId? }`; `agent.support.resume` accepts `{ runId, sessionId?, decision, expectedVersion }`. Do not expose an agent without durable checkpoints and a restart-safe `SessionStore`; no lifecycle tool appears by default.
|
|
189
208
|
|
|
@@ -191,9 +210,14 @@ MCP output is untrusted. Register bridge tools through core dispatch with a `Sec
|
|
|
191
210
|
|
|
192
211
|
Discovery validation is atomic: cursor/page/tool/name/description/schema failures reject `refresh()` and preserve the previous immutable tool-array reference. The bridge intentionally uses raw SDK `request()` for `tools/list` and `tools/call`; this avoids eager Ajv compilation/validation of untrusted remote output schemas. Host `ToolValidator` remains the argument-validation owner.
|
|
193
212
|
|
|
213
|
+
## Vendor web MCP prototype boundary
|
|
214
|
+
|
|
215
|
+
Official Exa/Firecrawl MCP servers may be tested only as explicit hardened prototypes: pin endpoint/origin/auth, inspect declared capabilities, allow-list individual tools/resources, retain all MCP bounds, and never expose generic remote passthrough. Production web research uses direct host-selected `@arnilo/prism-web-tools` adapters so provider choice, credentials, schema, and costs remain outside model control.
|
|
216
|
+
|
|
194
217
|
## Related APIs
|
|
195
218
|
|
|
196
219
|
- [Tools](tools.md): registry, dispatch, validation
|
|
220
|
+
- [Web search, fetch, and extraction](web-tools.md): preferred direct bounded Brave/Exa/Firecrawl production path
|
|
197
221
|
- [Tool execution primitives](tool-execution-primitives.md): Plan 055 design and conformance matrix
|
|
198
222
|
- [Host security guide](host-security.md): permission, trust, validation checklist
|
|
199
223
|
- [Web-standard server handler](server.md): agent/workflow HTTP routes and shared remote-boundary rules
|