@arnilo/prism 0.0.8 → 0.0.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (62) hide show
  1. package/CHANGELOG.md +44 -3
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +44 -5
  4. package/dist/agents.js +141 -7
  5. package/dist/context-budget.d.ts +63 -0
  6. package/dist/context-budget.js +235 -0
  7. package/dist/contracts.d.ts +107 -0
  8. package/dist/contracts.js +77 -0
  9. package/dist/index.d.ts +7 -5
  10. package/dist/index.js +5 -4
  11. package/dist/input.d.ts +3 -0
  12. package/dist/input.js +71 -28
  13. package/dist/node/session-store-jsonl.js +4 -1
  14. package/dist/provider-events.d.ts +2 -0
  15. package/dist/provider-events.js +21 -13
  16. package/dist/providers/openai-compatible.js +8 -5
  17. package/dist/providers/transport.d.ts +10 -1
  18. package/dist/providers/transport.js +24 -8
  19. package/dist/rpc.js +13 -2
  20. package/dist/session-stores.d.ts +7 -2
  21. package/dist/session-stores.js +174 -4
  22. package/dist/structured-output.d.ts +5 -1
  23. package/dist/structured-output.js +18 -0
  24. package/dist/testing/persistence-schema.d.ts +1 -1
  25. package/dist/testing/persistence-schema.js +8 -2
  26. package/dist/testing/session-store-conformance.d.ts +6 -0
  27. package/dist/testing/session-store-conformance.js +36 -1
  28. package/dist/tools.js +2 -0
  29. package/docs/agent-events.md +3 -2
  30. package/docs/agent-loops.md +10 -3
  31. package/docs/agent-session-runtime.md +5 -1
  32. package/docs/browser-automation.md +124 -0
  33. package/docs/cli-rpc.md +2 -1
  34. package/docs/coding-agent-tools.md +178 -14
  35. package/docs/coding-security.md +84 -11
  36. package/docs/evaluations.md +13 -2
  37. package/docs/guardrails.md +2 -1
  38. package/docs/host-security.md +4 -2
  39. package/docs/index.md +19 -15
  40. package/docs/input-and-prompt-assembly.md +4 -1
  41. package/docs/migration.md +84 -0
  42. package/docs/node-jsonl-session-store.md +1 -1
  43. package/docs/performance.md +46 -0
  44. package/docs/postgres-persistence.md +3 -3
  45. package/docs/provider-conformance.md +1 -1
  46. package/docs/provider-packages.md +1 -1
  47. package/docs/provider-primitives.md +7 -1
  48. package/docs/providers/anthropic.md +92 -0
  49. package/docs/providers/google.md +87 -0
  50. package/docs/public-contracts.md +4 -0
  51. package/docs/release-and-install.md +197 -64
  52. package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
  53. package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
  54. package/docs/review-coverage-2026-07-22-phase-6.md +209 -0
  55. package/docs/session-store-conformance.md +2 -0
  56. package/docs/session-stores.md +40 -1
  57. package/docs/sqlite-persistence.md +3 -3
  58. package/docs/structured-output.md +9 -3
  59. package/docs/tools.md +3 -0
  60. package/docs/web-tools.md +1 -1
  61. package/docs/workflows.md +3 -0
  62. package/package.json +6 -5
@@ -2,12 +2,17 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects.
5
+ `@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools and one disposable Docker/OCI sandbox reference. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects, and optionally contains untrusted coding work in a host-invoked container.
6
6
 
7
7
  | Export | Purpose |
8
8
  | --- | --- |
9
9
  | `createCodingApprovalPolicy(options)` | Returns an `ExecutionPolicy` with trusted roots, read-only mode, command allow/deny rules, approval caching, and timeout/abort-aware approval waits. |
10
10
  | `createSandboxBashOperations(adapter)` | Maps a host-owned `SandboxAdapter` to coding-agent `BashOperations` for delegated shell execution. |
11
+ | `createSandboxCodingComposition(cwd, options)` | Authoritative construction: returns `{ tools, composition }` with required `workspaceMode` (`"host"` \| `"sandbox"`), fail-closed mixed wiring, and containment metadata. |
12
+ | `createSandboxReadOnlyComposition(cwd, options)` | Same contract for read-only tools (`read`/`repo_list`/`repo_search`). |
13
+ | `createSandboxCodingTools` / `createSandboxReadOnlyTools` | Thin wrappers that return `tools` only (compat); still require `workspaceMode`. |
14
+ | `createSandboxFilesystemOperations` / `createSandboxRepositoryOperations` | Optional execFile-backed FS/list/search backends for a disposable sandbox tree. |
15
+ | `createDockerSandbox(options)` | Creates one disposable non-root Docker container with read-only root/source, bounded tmpfs workspace, typed `execFile`, import/export, and stop/kill/cleanup. |
11
16
  | `assertPathInsideRoots`, `isPathInsideReal` | Symlink-aware path containment helpers. |
12
17
  | `evaluateCommandRules`, `hasShellMetacharacters` | Command classification helpers. |
13
18
 
@@ -21,7 +26,7 @@ import type { ExecutionAction, ExecutionPolicy, ExecutionDecision } from "@arnil
21
26
 
22
27
  Use this package when coding tools need path scoping, human approval, command rules, or a pluggable sandbox backend. Wire the returned policy through `createCodingTools(cwd, { executionPolicy })` or per-tool `executionPolicy` options.
23
28
 
24
- Prism does **not** claim OS-level isolation unless the host provides a sandbox adapter. Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
29
+ Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
25
30
 
26
31
  ## Inputs / request
27
32
 
@@ -36,10 +41,39 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
36
41
 
37
42
  `run` caching keys decisions by the tool execution context's `runId`; `session` uses `sessionId`. Coding tools pass both identities to the policy. A missing/empty identity disables caching for that check rather than creating a global bucket. Identical actions in different runs/sessions never share approvals or denials.
38
43
 
44
+ ### Docker sandbox inputs
45
+
46
+ | Option | Default | Purpose |
47
+ | --- | --- | --- |
48
+ | `docker` | required | Absolute host Docker executable. |
49
+ | `image` | required | Digest-pinned image (`name@sha256:<64-hex>`). Never pulled (`--pull=never`). |
50
+ | `sourceRoot` | required | Absolute host directory imported into `/workspace`. |
51
+ | `user` | required | Non-root `uid:gid`. |
52
+ | `network` | `{ mode: "none" }` | Default no network; custom mode requires a pre-created network name and does not claim DNS containment. |
53
+ | `env` | `{}` | Exact allow-list only; host environment is never inherited. |
54
+ | `secrets` | `[]` | Canaries redacted from CLI/adapter errors. |
55
+ | `limits` | package defaults | CPU/memory/PID/FD/tmpfs/command/export/time caps validated before create. |
56
+
57
+ ### Workspace mode inputs (`createSandboxCodingComposition`)
58
+
59
+ | Option | Default | Purpose |
60
+ | --- | --- | --- |
61
+ | `workspaceMode` | **required** | `"host"` (all tools on host cwd; never claims containment) or `"sandbox"` (shell + FS/list/search share one disposable tree). |
62
+ | `sandbox` | optional in host; required for sandbox unless custom ops supplied | `SandboxAdapter` / `DisposableSandbox`. |
63
+ | `workspaceRoot` | `"/workspace"` in sandbox mode | Tree root used as tool cwd when sandbox backends are bound. |
64
+ | `allowMixedWorkspaceWiring` | `false` | Escape hatch: allow sandbox shell + host FS backends. Records `composition.warnings`; forces `containmentClaim: false`. Missing hatch throws. |
65
+ | `read`/`write`/`edit`/`repository.operations` | auto-wired from `DisposableSandbox` in sandbox mode | Host may supply custom tree backends instead of auto-wire. |
66
+
67
+ `0.0.9` silent split (sandbox shell + host FS) is **superseded**. Mixed wiring is never the default.
68
+
39
69
  ## Outputs / response / events
40
70
 
41
71
  `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations`, receive `onData(Buffer)` for ordered stdout/stderr forwarding through the shell tool's existing bounded accumulator, and never grant policy approval themselves.
42
72
 
73
+ `createSandboxCodingComposition()` returns `{ tools, composition }` where `SandboxCodingComposition` carries `workspaceMode`, `containmentClaim`, `mixedWiringAllowed`, `warnings`, `workspaceRoot`, and optional `treeIdentity` (from `importIdentity` / `lastExportIdentity`). `containmentClaim` is `true` only for sandbox mode with tree backends bound and mixed wiring denied. Host mode and escape-hatch mixed wiring always set `containmentClaim: false` — never treat host mode as contained execution.
74
+
75
+ `createDockerSandbox()` returns a `DisposableSandbox`: typed `execFile(file, args)`, shell-compatible `exec`, `status`, cooperative `stop`, forced `kill`, and idempotent `close`. Import may surface `importIdentity`; successful export updates `lastExportIdentity`. `close({ export })` can stream a bounded workspace tar plus SHA-256/entry/byte metadata through a host callback; checkpoints should retain only host artifact references/hashes, never whole workspaces.
76
+
43
77
  ## Request/response example
44
78
 
45
79
  ```json
@@ -52,8 +86,12 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
52
86
  ## Implementation example
53
87
 
54
88
  ```ts
55
- import { createCodingTools } from "@arnilo/prism-coding-agent";
56
- import { createCodingApprovalPolicy, createSandboxBashOperations } from "@arnilo/prism-coding-security";
89
+ import {
90
+ createCodingApprovalPolicy,
91
+ createDockerSandbox,
92
+ createSandboxCodingComposition,
93
+ } from "@arnilo/prism-coding-security";
94
+ import { createGitTools } from "@arnilo/prism-coding-agent";
57
95
 
58
96
  const policy = createCodingApprovalPolicy({
59
97
  roots: [workspaceRoot],
@@ -62,27 +100,62 @@ const policy = createCodingApprovalPolicy({
62
100
  approvalTimeoutMs: 60_000,
63
101
  });
64
102
 
65
- const tools = createCodingTools(workspaceRoot, {
103
+ const sandbox = await createDockerSandbox({
104
+ docker: "/usr/bin/docker",
105
+ image: "registry.example/prism-code@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
106
+ sourceRoot: "/srv/jobs/task-1/source",
107
+ user: "10001:10001",
108
+ network: { mode: "none" },
109
+ env: { CI: "1" },
110
+ limits: { cpus: 2, memoryBytes: 2 * 1024 ** 3, maxPids: 256, workspaceBytes: 1024 ** 3 },
111
+ });
112
+
113
+ // Sandbox mode: shell/read/write/edit/list/search share one disposable tree.
114
+ const { tools, composition } = createSandboxCodingComposition("/srv/jobs/task-1/source", {
115
+ workspaceMode: "sandbox",
116
+ sandbox,
66
117
  executionPolicy: policy,
67
- shell: {
68
- operations: createSandboxBashOperations(mySandboxAdapter),
69
- },
118
+ repository: { exclude: [".git", "node_modules", "dist"] },
119
+ });
120
+ // composition.containmentClaim === true when backends are bound
121
+
122
+ // Same-tree Git/check (opt-in; not folded into coding tools):
123
+ const gitTools = createGitTools(composition.workspaceRoot, {
124
+ execFile: sandbox.execFile.bind(sandbox),
125
+ commitIdentity: { name: "bot", email: "bot@example.com" },
126
+ });
127
+
128
+ // Host mode (explicit non-contained): omit sandbox; never claim containment.
129
+ const host = createSandboxCodingComposition(hostCwd, { workspaceMode: "host", executionPolicy: policy });
130
+ // host.composition.containmentClaim === false
131
+
132
+ await sandbox.execFile({ file: "npm", args: ["test"], cwd: "/workspace" });
133
+ await sandbox.close({
134
+ export: async (stream, meta) => hostArtifacts.write(stream, meta),
70
135
  });
71
136
  ```
72
137
 
73
138
  ## Extension and configuration notes
74
139
 
75
- Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers.
140
+ Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()`/`createSandboxCodingComposition()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` / `DisposableSandbox` are replaceable and host-owned; approval policy and sandboxing are separate layers. Custom remote sandboxes can implement `DisposableSandbox` without using Docker.
141
+
142
+ `createSandboxCodingComposition()` requires `workspaceMode`. Sandbox mode auto-wires FS/list/search through `DisposableSandbox.execFile` (or host-supplied custom operations) so mutations stay on the disposable tree until export. Host mode runs every coding tool against the host cwd and never sets `containmentClaim`. Sandbox shell + host FS throws unless `allowMixedWorkspaceWiring: true` (warnings + `containmentClaim: false`). Opt-in structured Git tools (`createGitTools(composition.workspaceRoot, { execFile: sandbox.execFile, commitIdentity })`) share the same tree/cwd; Prism still never pushes or opens PRs. Optional `@arnilo/prism-browser` can share the same disposable boundary: use `assertBrowserSandboxNetwork()` before browse-ready custom networks, and `createSharedSandboxBrowserOptions({ workspaceRoot, downloadsRoot, containedProxyAttestation })` so uploads/downloads align with `/workspace` and `/downloads`. Close the browser context before disposing the sandbox.
143
+
144
+ The Docker reference adapter starts by recorded container ID/label, uses argument arrays only, mounts source read-only, populates a size-bounded tmpfs `/workspace`, drops all capabilities, enables `no-new-privileges`, runs with `--init`, and never exposes the Docker socket, privileged mode, or host PID/IPC namespaces. Image pull/build/update stays outside Prism. Protected real-Docker checks are opt-in via `PRISM_TEST_DOCKER_SANDBOX=1` with host-supplied `PRISM_TEST_DOCKER_BIN` and digest-pinned `PRISM_TEST_DOCKER_IMAGE`.
76
145
 
77
146
  Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
78
147
 
79
148
  ## Security and performance notes
80
149
 
81
- Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
150
+ Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, repository list/search walks, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe.
151
+
152
+ Docker sandbox containment—not command regexes—enforces filesystem/network/process boundaries for the reference adapter. Network defaults to none; a custom Docker network still requires a host firewall/proxy for DNS/egress claims. Import rejects symlink escapes, devices, FIFOs, and sockets; export counts entries/bytes and hashes before host retention. Secrets in `secrets` are redacted from adapter errors and never exported as environment metadata. Unified workspace mode reuses existing sandbox/repo/coding hard caps and does not introduce unbounded host↔container sync loops. Host mode and `allowMixedWorkspaceWiring` never claim disposable containment. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter and Docker daemon.
82
153
 
83
154
  ## Related APIs
84
155
 
85
- - [Coding agent tools](coding-agent-tools.md)
156
+ - [Coding agent tools](coding-agent-tools.md): durable plan/todo Markdown helpers and `state.coding` checkpoint metadata for restart/resume without a second runtime
157
+ - [Workflows](workflows.md): `runWorkflow` / `resumeWorkflow` / `startWorkflowBackground` composition for coding tasks
86
158
  - [Host security guide](host-security.md)
159
+ - [Performance limits](performance.md)
87
160
  - [Tool execution primitives](tool-execution-primitives.md)
88
161
  - [Security/auth/trust](settings-auth-trust-security.md)
@@ -135,11 +135,22 @@ const comparison = await runComparison({ dataset, candidates: { baseline, candid
135
135
  assertEvaluationThreshold(report, { minimumMean: 0.9, maximumFailures: 0 });
136
136
  ```
137
137
 
138
- `traceResolver` is explicit; no arbitrary run search occurs. `baseline`/`candidate` are host functions returning `AgentRunResult`. See `examples/evaluation-gate.ts` for a network-free gate.
138
+ `traceResolver` is explicit; no arbitrary run search occurs. `baseline`/`candidate` are host functions returning `AgentRunResult`. See `examples/evaluation-gate.ts` for a network-free gate and `examples/coding-browser-evaluation.ts` for coding/browser adversarial fixtures.
139
+
140
+ ## Coding and browser adversarial evaluations (0.0.9)
141
+
142
+ Release 0.0.9 ships curated network-free adversarial fixtures in package tests:
143
+
144
+ - `@arnilo/prism-coding-agent` `eval-fixtures.test.ts`: safe native list vs shell, Git path/ref injection, dirty-tree rollback, unknown named-check failure, PR-handoff artifact completeness, and prompt-injection file content under read-only tools.
145
+ - `@arnilo/prism-browser` `eval-fixtures.test.ts`: stale snapshot refs, side-effect approval, private/loopback/file deny, upload/download/screenshot policy, CSS/evaluate target rejection, and hostile accessible-name text.
146
+
147
+ Fixtures reuse `@arnilo/prism-evals` (`defineDataset` / `defineScorer` / `scoreRun` / `assertEvaluationThreshold` / `serializeEvaluationReport`). Optional SWE-bench-compatible or live-browser harnesses remain host adapters — they are not default dependencies or quality claims. Protected real Docker/Playwright gates stay env-gated (`PRISM_TEST_DOCKER_SANDBOX`, `PRISM_LIVE_PLAYWRIGHT`) and never enter `sdk:ready`.
139
148
 
140
149
  ## Related APIs
141
150
 
142
151
  - [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
143
152
  - [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
144
153
  - [Observability](observability.md): use `onTraceReference` or bounded `traceId(runId)` to supply `ScoreRunOptions.traceId`; evaluation telemetry emits no reason/explanation content
145
- - [Release and install](release-and-install.md): optional package install
154
+ - [Coding agent tools](coding-agent-tools.md) / [Browser automation](browser-automation.md) / [Workflows](workflows.md): network-free coding-task composition at `examples/durable-coding-workflow.ts`; adversarial coding/browser eval example at `examples/coding-browser-evaluation.ts`
155
+ - [Performance limits](performance.md): `scripts/benchmark-0.0.11.mjs` search/budget evidence, `scripts/benchmark-0.0.10.mjs` workspace-mode evidence, and `scripts/benchmark-0.0.9.mjs` coding/browser evidence fields
156
+ - [Release and install](release-and-install.md): optional package install and protected sandbox-browser workflow
@@ -65,11 +65,12 @@ Guardrails are callbacks supplied by the host. Prism does not discover, load, re
65
65
 
66
66
  ## Security and performance notes
67
67
 
68
- Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work.
68
+ Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work. Browser snapshots and page text from `@arnilo/prism-browser` are untrusted external content: never allow them to modify tools, permissions, credentials, or policy. Browser mutations still require host `ExecutionPolicy`/approval; prompt-injection text in a page cannot grant upload/download release.
69
69
 
70
70
  ## Related APIs
71
71
 
72
72
  - [Agent/session runtime](agent-session-runtime.md)
73
73
  - [Tools](tools.md)
74
+ - [Browser automation](browser-automation.md)
74
75
  - [Agent events](agent-events.md)
75
76
  - [Host security](host-security.md)
@@ -46,7 +46,7 @@ Security controls fail closed before side effects when wired at the guarded edge
46
46
  - validator failures emit `tool_execution_blocked` with `validation_failed`
47
47
  - configured guardrails fail closed; output stages buffer blocked provider/tool content before events, ledgers, session entries, or MCP responses
48
48
  - configured redactors scrub provider requests, agent events, session entries, ledger records, tool errors, extension errors, injector context, and durable run checkpoints
49
- - durable resume requires host-derived exact ownership and checkpoint version; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
49
+ - durable resume requires host-derived exact ownership and checkpoint version; coding-task resume also revalidates plan/workspace artifact hashes plus tool/policy/image fingerprints via `assertCodingResumeAllowed` before import; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
50
50
 
51
51
  These checks are explicit function calls during load, assembly, dispatch, append, or run handling. Prism adds no background watchers, filesystem scanners, network probes, credential polling, or automatic extension discovery.
52
52
 
@@ -137,7 +137,8 @@ Wire those values where they matter: provider adapters receive the resolved cred
137
137
  - Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
138
138
  - MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands/resources/prompts, requires per-operation `authorize`, and retains core gates. Sampling, roots, model/credential selection, and elicitation consent stay host-owned; URL elicitation is never opened automatically. Stateful web mode requires host `resolveAuthInfo` plus `resolveIdentity`, exact origin policy, and binds every POST/GET/DELETE/SSE request to one non-secret principal; mismatches return 404. Handler still needs TLS and edge rate limiting. See [MCP client/server exposure](mcp-tools.md).
139
139
  - `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
140
- - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Limits are not containment: Prism provides no OS sandbox unless the host supplies one.
140
+ - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, repository list/search depth/entry/match/scan/time caps, structured Git path/ref/message/output/patch/worktree caps, named-check concurrency/output caps, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Opt-in `createGitTools()` uses argument arrays with hooks/credential prompts/external diff disabled, requires host `commitIdentity` for commits, and never pushes or opens PRs. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell/repository backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, identity-scoped approval caching, required `workspaceMode` on `createSandboxCodingComposition()` / `createSandboxCodingTools()`, and the optional `createDockerSandbox()` reference adapter. **Host mode is never contained execution** (`containmentClaim: false`). Sandbox mode claims containment only when FS backends target the disposable tree; mixed wiring requires `allowMixedWorkspaceWiring` and still does not claim containment. Limits alone are not containment: construct the Docker adapter (absolute CLI, digest-pinned image, network none by default) or an equivalent host sandbox before treating coding execution as production-safe. Docker daemon/image trust, egress firewall/proxy, and artifact retention remain host-owned.
141
+ - Optional `@arnilo/prism-browser` requires a host-supplied Playwright Browser (`playwright-core@1.61.0` peer). Import is inert. One non-persistent context belongs to one run; actions serialize; refs are snapshot-scoped; CSS/evaluate/CDP/persistent profiles are denied. Context routing + `serviceWorkers: "block"` deny file/data/blob/devtools/private/loopback by default and require contained-proxy attestation for external egress (Playwright routing is defense in depth, not DNS containment). Uploads are realpath-rooted; downloads quarantine with hash/MIME until host `approveRelease`; screenshots return bounded `ImageContent`. Observation vs mutation/high-impact actions map to `ExecutionPolicy`. Treat snapshot/page text as untrusted external content. Close contexts with `browser_close` or `manager.closeRun(runId)` on abort/terminal. Browser control endpoint, binary/image pin, and real egress firewall/proxy remain host-owned. Shared sandbox: `createSharedSandboxBrowserOptions()` + `assertBrowserSandboxNetwork()`.
141
142
  - `@arnilo/prism-credentials-node` rejects oversized/malformed envelopes and excessive scrypt work before KDF allocation, uses async scrypt, and requires restrictive existing/new Unix vault modes. Keep vault ownership and parent-directory access host-controlled; review before `chmod 600`, never auto-weaken a file policy. Keychain calls use abort-aware native async work with finite timeout/payload caps and sanitized errors. OS prompts, service availability, and whether a native backend promptly honors cancellation remain host/platform boundaries; no plaintext fallback is attempted.
142
143
  - LLM compaction always sends finite summary `maxTokens`, retains bounded deltas/events, and bounds/redacts provider/factory/policy error detail. Observational-memory workers cap turns, calls, arguments, results, transcript, and surfaced errors; unknown tools fail before execution, while invalid results can only be rejected after a host tool returns and may therefore follow side effects. Pass all known provider/credential/tool secrets into compaction/runtime options; exact replacement is not secret discovery.
143
144
  - Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
@@ -183,6 +184,7 @@ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy
183
184
  - `scripts/scan-secrets.mjs` checks tracked source and unpacked public tarballs for high-confidence credential/private-key forms without printing matched values. It complements GitHub secret scanning; it is not entropy scanning or DLP.
184
185
  - Tag publication alone receives npm/OIDC/attestation permissions. Untrusted pull-request code receives no canary, npm, or OIDC secret and no workflow uses `pull_request_target`.
185
186
  - Scheduled/manual canaries run only in protected `live-canaries` environment. Use dedicated read-only/low-quota credentials and provider account spend limits. Runner performs four probes, at most one MCP cleanup, one provider output token, one Brave result, 64-KiB responses, and finite timeouts; report excludes endpoints, headers, bodies, credentials, and MCP session IDs.
187
+ - Scheduled/manual coding/browser containment checks run in protected `sandbox-browser` environment (`.github/workflows/sandbox-browser.yml`). They receive no provider/npm/OIDC secrets; Docker/Playwright enablement is variable-gated with host-preloaded digest-pinned images/binaries; uploads are redacted aggregate status only.
186
188
  - Live endpoint operators own TLS, egress allow-lists, account-dollar budget, cleanup beyond MCP session DELETE, and revocation. Failed canaries log only operation kind plus status/timeout; inspect provider-side audit logs for details.
187
189
 
188
190
  ## Related APIs
package/docs/index.md CHANGED
@@ -6,15 +6,15 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
6
6
  - [Public contracts](public-contracts.md): type shapes for messages, agents, tools, stores, generic `CheckpointStore`, atomic `LeaseStore`, bounded `EventMultiplexer`, resources, credentials, and events.
7
7
 
8
8
  ## Agent/session runtime
9
- - [Agent/session runtime](agent-session-runtime.md): create explicit or opt-in secure agents/sessions, get direct `AgentRunResult` values from `run`/`prompt`, use integrated `stream()`, subscribe to normalized events, and expose opted-in durable lifecycle capabilities.
9
+ - [Agent/session runtime](agent-session-runtime.md): create explicit or opt-in secure agents/sessions, get direct `AgentRunResult` values from `run`/`prompt`, mid-run `steer` (turn-boundary or softInterrupt), use integrated `stream()`, subscribe to normalized events, and expose opted-in durable lifecycle capabilities.
10
10
  - [Agent definitions](agent-definitions.md): resolve declarative `AgentDefinition` values via `resolveAgentDefinition`, and turn app-config `<configRoot>/agents/<name>/AGENT.md` bundles into runnable agents via `discoverAgentBundles` / `resolveAgentBundle` (explicit tool/skill activation by name, fail-closed omitted capabilities, migration-only `activateAllCapabilities`, strict duplicate scope checks, configurable prompt layers, no auto-discovery).
11
11
  - [Agent loops](agent-loops.md): replaceable per-run control loops — `singleShotLoop` default and opt-in bounded artifact-loop tool rounds with host-supplied `validator`/`parser`/`repairer` callbacks.
12
12
  - [Guardrails](guardrails.md): typed fail-closed input/output/tool checks with buffered provider output and redacted decision records.
13
13
  - [Agent events](agent-events.md): redacted lifecycle stream used by UIs, ledgers, and metadata-only parented telemetry; message/progress deltas never create spans.
14
14
  - [Observability](observability.md): OTel GenAI agent/provider/tool hierarchy, host context parenting, bounded trace linkage, safe evaluation events, controlled metrics, and exporter isolation.
15
- - [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, and ID-only linkage to immutable owned run feedback.
15
+ - [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, coding/browser adversarial fixtures, and ID-only linkage to immutable owned run feedback.
16
16
  - [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence, optional bounded FIFO durability policies, session snapshot caching, and immutable run/trace feedback.
17
- - [Performance limits](performance.md): bounded evaluation traces/judges/reports, security scan/live-canary backstops, live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
17
+ - [Performance limits](performance.md): bounded evaluation traces/judges/reports, 0.0.11 search/budget, 0.0.10 workspace-mode, and 0.0.9 coding/browser benchmark evidence, security scan/live-canary backstops, live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
18
18
  - [Structured output](structured-output.md): the `Artifact*` seam plus provider-native `StructuredOutputOptions` / `structuredOutputMode` for capable models.
19
19
 
20
20
  ## Compaction/session memory
@@ -22,13 +22,13 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
22
22
  - [LLM compaction package](compaction-llm.md): optional provider-backed strategy with finite summary/reserve/error caps, bounded redacted streaming retention, and mandatory finite post-policy `model.parameters.maxTokens`.
23
23
  - [Observational memory compaction package](compaction-observational-memory.md): optional source-backed memory with owned append callback, finite turn/call/argument/result/transcript/error worker limits, redacted provider-valid transcripts, fast compaction, recall, and status/view commands; worker model falls back to host-supplied `sessionModel`.
24
24
  - [Working and semantic memory](working-and-semantic-memory.md): optional `@arnilo/prism-memory` working-memory store, semantic recall, finite Embedder/VectorStore contracts, in-memory adapters, and PostgreSQL/pgvector path.
25
- - [Session stores](session-stores.md): `SessionStore` contract, `SessionAppendOptions`, `SessionAppendConflictError`, branch handles, `readBranchPath`, and dev-vs-production branch reads — start here for session persistence.
25
+ - [Session stores](session-stores.md): `SessionStore` contract, `SessionAppendOptions`, `SessionAppendConflictError`, branch handles, `readBranchPath`, optional bounded `searchSessions` / `SessionIndex` (memory linear|unsupported), and dev-vs-production branch reads — start here for session persistence.
26
26
  - [Session stores and branching](session-stores-and-branching.md): detailed branch semantics and helper reference (kept for compatibility; links back to the canonical atomic append / branch-handle sections).
27
27
  - [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention, and NoSQL mapping.
28
- - [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, and transactionally verified/backfilled migration-v3 metadata.
29
- - [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
30
- - [Migration guide](migration.md): 0.0.3 compatibility through 0.0.8 telemetry/evaluation, MCP/A2A, web research, ledger batching, and release-security changes.
31
- - [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety.
28
+ - [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, FTS `searchSessions` (migration-v4), and transactionally verified/backfilled migration metadata.
29
+ - [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, FTS `searchSessions` (migration-v4), advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
30
+ - [Migration guide](migration.md): 0.0.3 compatibility through **0.0.11** coding-harness fundamentals (SessionIndex, contextBudget, Anthropic/Google providers, steer, ask_user_decision, goal→verify), 0.0.10 workspace modes, and 0.0.9 coding/browser surfaces.
31
+ - [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety; `searchSessions` throws `SessionSearchUnsupportedError`.
32
32
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
33
33
 
34
34
  ## Provider and model connection
@@ -40,13 +40,13 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
40
40
  - [Use-case model selection](use-case-model-selection.md): bind `{ model?, provider?, thinkingLevel? }` for observational memory, LLM compaction, and other non-session LLM jobs with explicit session-model fallback via `resolveUseCaseModel`.
41
41
  - [Provider request policies](provider-request-policies.md): chain `ProviderRequestPolicy` hooks, use `createSessionCachePolicy`, and merge legacy/structured cache options safely.
42
42
  - [Provider packages](provider-packages.md): define explicit provider packages, model metadata, auth descriptors, request/cache policies, and provider-owned header precedence without package discovery or provider-specific core behavior; includes a first-party cache behavior summary and the **caller-gated on-demand model discovery** contract (`list*Models`, setup zero-fetch).
43
- - Phase 12 package workspaces: [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md) (official Go open models, dual-route Anthropic/OpenAI, caller-gated `listOpenCodeGoModels`, `reasoning_content`/thinking preserve), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md) (app-controlled catalog, caller-gated `listOpenRouterModels`, `reasoning` merge/preserve, `cache_control` + sticky `session_id`), [`@arnilo/prism-provider-zai`](providers/zai.md) (official `thinking`/`reasoning_effort`/`tool_stream`, implicit cache, caller-gated `listZaiModels`), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md) with implicit vLLM prefix caching, reasoning controls (`reasoning_effort`/`thinking_token_budget`/`enable_thinking`/`preserve_thinking`/`clear_thinking`), reasoning preservation, OpenAI-style tool-call loop, quota, telemetry, and retry classification helpers.
43
+ - Phase 12 package workspaces: [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-anthropic`](providers/anthropic.md) (native Messages, `cache_control`, thinking, caller-gated `listAnthropicModels`), [`@arnilo/prism-provider-google`](providers/google.md) (native Gemini `generateContent` SSE, caller-gated `listGoogleModels`; Vertex deferred), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md) (official Go open models, dual-route Anthropic/OpenAI, caller-gated `listOpenCodeGoModels`, `reasoning_content`/thinking preserve), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md) (app-controlled catalog, caller-gated `listOpenRouterModels`, `reasoning` merge/preserve, `cache_control` + sticky `session_id`), [`@arnilo/prism-provider-zai`](providers/zai.md) (official `thinking`/`reasoning_effort`/`tool_stream`, implicit cache, caller-gated `listZaiModels`), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md) with implicit vLLM prefix caching, reasoning controls (`reasoning_effort`/`thinking_token_budget`/`enable_thinking`/`preserve_thinking`/`clear_thinking`), reasoning preservation, OpenAI-style tool-call loop, quota, telemetry, and retry classification helpers.
44
44
  - Optional AI SDK adapter: [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md) maps host-owned `LanguageModelV4` models onto Prism `AIProvider` streams (specification v4; no Prism catalog; maps `finish.usage` cache read/write tokens; reasoning is host-model-owned).
45
45
  - [OpenAI-compatible provider](providers/openai-compatible.md): optional provider subpath using native or injected `fetch` for Chat Completions streaming.
46
46
 
47
47
  ## Input, prompt, and context assembly
48
48
  - [SDK customization guide](customization.md): map provider resolution, middleware, context, builders, injectors, loops, compaction, retry, stores, and skills to explicit host-wired APIs.
49
- - [Input and prompt assembly](input-and-prompt-assembly.md): render tiny prompt templates and turn common host input, history, attachments, explicit resources, summaries, and tool results into messages with replaceable builders, provider-input assembly, legacy default order, and opt-in cache-aware ordering. Audio/file/document `ContentBlock` types and capability checks are documented there.
49
+ - [Input and prompt assembly](input-and-prompt-assembly.md): render tiny prompt templates and turn common host input, history, attachments, explicit resources, summaries, and tool results into messages with replaceable builders, provider-input assembly, legacy default order, opt-in cache-aware ordering, and optional `contextBudget` eviction + omission reports. Audio/file/document `ContentBlock` types and capability checks are documented there.
50
50
  - [Multimodal content](multimodal-content.md): complete-request media resolution and aggregate bounds, DNS-classified/address-pinned URLs, SSRF/MIME policy, and `ModelCapabilities.input` tags.
51
51
  - [System prompts](system-prompts.md): compose explicit user/package/app/run system prompt layers, auto-load the standard `AGENTS.md` (workspace) / `SYSTEM.md` prompt files via the Node `loadSystemPromptFiles` loader (trust-gated for `AGENTS.md`), and append `SYSTEM.md` → per-agent `AGENT.md` body → repo `AGENTS.md` layers from a discovered agent bundle via `resolveAgentBundle`.
52
52
  - [Instruction injection](instruction-injection.md): register package injectors that layer redacted instructions/context blocks without granting tools, permissions, or resource escapes.
@@ -59,8 +59,9 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
59
59
  - [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
60
60
  - [MCP client bridge and server exposure](mcp-tools.md): SDK-1.29.0 bounded tools/resources/prompts, host-owned roots/sampling/elicitation, exact-origin DNS-pinned client transport, and principal-bound opt-in Streamable HTTP sessions.
61
61
  - [Web search, fetch, and extraction](web-tools.md): optional host-selected Brave/Exa discovery and Firecrawl Markdown/schema tools with native fetch, stable citations, late credentials, finite limits, and explicit untrusted-content boundaries.
62
- - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, and `edit` definitions with streamed text pages, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
63
- - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, and abort-aware streaming sandbox adapters for coding tools.
62
+ - [Browser automation](browser-automation.md): optional `@arnilo/prism-browser` with host-supplied Playwright contexts, AI-mode snapshots/refs, ordered `browser_open`/`browser_snapshot`/`browser_act`/`browser_close`, egress/side-effect/upload/download/screenshot policy, and finite page/action/snapshot/network/artifact caps.
63
+ - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, `edit`, `repo_list`, and `repo_search` definitions plus opt-in `createGitTools()` / `coding_check`, opt-in `createAskUserDecisionTool` (single/multi/free-text + durable suspend glue), and `runCodingGoalVerify`; durable plan/todo Markdown helpers with workflow `state.coding` checkpoint metadata; streamed text pages, bounded repository list/search, finite Git/check/plan/ask caps, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
64
+ - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, required `workspaceMode` (`host`/`sandbox`) with fail-closed mixed wiring, `createSandboxCodingComposition()` containment metadata, and the disposable Docker/OCI sandbox reference with bounded workspace import/export.
64
65
 
65
66
  ## Extensions/plugins
66
67
  - [Contribution discovery (workspace)](contribution-discovery.md): opt-in, realpath-contained directory scanner turning `SKILL.md`/`manifest.json` into inert `DiscoveredContribution` envelopes the host registers — no `import()`, no auto-activate, no provider scanning. Per-agent bundles remain app-controlled and are documented under Agent/session runtime.
@@ -82,8 +83,8 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
82
83
  - [A2A interoperability](a2a.md): A2A 1.0 JSON-RPC/HTTPS cards plus host-owned durable task get/list/cancel/subscribe, bounded rich parts/replay, principal-scoped push configs, and exact-origin verified client.
83
84
 
84
85
  ## CLI/RPC
85
- - [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
86
- - [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Interactive TUI (C-012) deferred.
86
+ - [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including mid-run `steer`, branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
87
+ - [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Compose coding plans/checkpoints via workspace Markdown + `state.coding` without a second runtime. Interactive TUI (C-012) deferred.
87
88
  - [Workflow orchestration primitives](workflow-orchestration-primitives.md): architecture inventory — workflow adapters consume core `CheckpointStore`, `LeaseStore`, and bounded `EventMultiplexer`; run control and optional RPC commands stay package-local.
88
89
  - [Workflow/TUI scope](workflow-tui-primitives.md): records why 0.0.5 ships workflow APIs/RPC control but no interactive terminal UI.
89
90
 
@@ -104,7 +105,10 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
104
105
  - `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, cache-aware prompt assembly, NeuralWatt agent run, stores/branching, compaction, observational-memory recall, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
105
106
 
106
107
  ## Release and install
107
- - [Release and install](release-and-install.md): 31-package graph, install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, and protected live canaries.
108
+ - [Release and install](release-and-install.md): 32-package graph (including optional browser), install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, protected live canaries, and sandbox-browser Docker/Playwright gates.
109
+ - [Review coverage (2026-07-22 Phase 6)](review-coverage-2026-07-22-phase-6.md): Plan 074 evidence freeze — SessionIndex/search, contextBudget, native Anthropic/Google packages, goal→verify, steer, ask_user_decision (multi/free-text/suspend), finite limits, threats, and 0.0.11 release gates.
110
+ - [Review coverage (2026-07-21 Phase 5)](review-coverage-2026-07-21-phase-5.md): Plan 073 evidence freeze — unified workspace modes, primitive ownership, reused finite limits, threats, and 0.0.10 release gates.
111
+ - [Review coverage (2026-07-20 Phase 4)](review-coverage-2026-07-20-phase-4.md): Plan 072 evidence freeze — revised coding/browser-only scope, external revisions, primitive ownership, finite limits, threats, and 0.0.9 release gates.
108
112
  - [Review coverage (2026-07-19 Phase 3)](review-coverage-2026-07-19-phase-3.md): Plan 070 evidence freeze — exact protocol/vendor references, capability/primitive/limit matrices, supported boundaries, and 0.0.8 release evidence.
109
113
  - [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md): Plan 067 evidence freeze — P0–P2 re-verification owners, seven first-party provider packages mapped to official-doc URLs, Pi secondary refs, cache/thinking/discovery surfaces, credential canaries, and use-case model-binding inventory.
110
114
  - [Review coverage (2026-07-15)](review-coverage-2026-07-15.md): frozen 0.0.5 finding/feature ownership, existing-primitive inventory, package decisions, threat boundaries, exclusions, and measured Phase 0 baseline.
@@ -63,7 +63,8 @@ Useful exported types:
63
63
  - `InputAttachment`: already-loaded text/content blocks (including `audio`, `file`, and `document`) or an explicit URI loaded through a caller-provided `ResourceLoader`.
64
64
  - `PromptInstruction`: labeled system instruction text.
65
65
  - `DefaultPromptBuilder`: the default `PromptBuilder`.
66
- - `AssembleProviderInputOptions`: model, input, optional builders, context providers, selected skills, active tools, generic provider options, metadata, and signal.
66
+ - `AssembleProviderInputOptions`: model, input, optional builders, context providers, selected skills, active tools, generic provider options, metadata, signal, and optional `contextBudget` (`maxInputTokens` / `maxInputBytes` / `reportOmissions`).
67
+ - `applyContextBudget` / `getContextBudgetReport` / `resolveContextBudget`: deterministic eviction + omission report helpers (estimate = UTF-16 code units ÷ 4).
67
68
  - `PromptTemplateOptions`: missing-variable behavior for `renderPromptTemplate()`.
68
69
 
69
70
  ## Outputs / response / events
@@ -86,6 +87,8 @@ The default prompt builder still prepends context, selected skills, and tool dec
86
87
  - Tool results are tool messages containing `tool_result` content; the agent/session runtime uses this to feed dispatched tool results into the next provider turn, placing the assistant `tool_call` and the matching role `tool` `tool_result` before any final assistant content. Cache-aware layout keeps tool results before the current user suffix so it does not split tool transcripts.
87
88
  - Middleware runs only when `middleware` is supplied in the context.
88
89
  - `assembleProviderInput()` returns a `ProviderRequest` with the caller's model/tools/provider options/metadata/signal and composed messages/context. It also calls `assertMessagesSupportModelCapabilities()` so unsupported `audio`/`file`/`document`/`image` blocks fail with `UnsupportedModalityError` when the model declares `capabilities.input`.
90
+ - Optional `contextBudget` (at least one of `maxInputTokens` / `maxInputBytes`) runs after default message groups are built and before final flatten. Eviction drops droppable sections first (toolResults → history → summaries → context → skills → attachments; layout-aware). Protected instructions + current user `input` (+ tools catalog) fail closed with `ContextBudgetError` if they alone exceed the budget. When `reportOmissions: true`, attach `ProviderRequest.metadata[CONTEXT_BUDGET_REPORT_METADATA_KEY]` and read via `getContextBudgetReport(request)` (kinds/ids/sizes only — no secrets). Raw session store entries are never deleted.
91
+ - Optional `contextBudget` (at least one of `maxInputTokens` / `maxInputBytes`) runs after default message groups are built and before final flatten. Eviction drops droppable sections first (toolResults → history → summaries → context → skills → attachments; layout-aware). Protected instructions + current user `input` (+ tools catalog) fail closed with `ContextBudgetError` if they alone exceed the budget. When `reportOmissions: true`, attach `ProviderRequest.metadata[CONTEXT_BUDGET_REPORT_METADATA_KEY]` and read via `getContextBudgetReport(request)` (kinds/ids/sizes only — no secrets). Raw session store entries are never deleted.
89
92
  - `renderPromptTemplate()` replaces top-level `{{name}}` variables with caller-supplied JSON-compatible values. Strings are inserted directly; numbers, booleans, `null`, arrays, and objects are stringified deterministically with sorted object keys. Missing variables throw by default or stay unchanged with `{ missing: "preserve" }`.
90
93
 
91
94
  ## Request/response example
package/docs/migration.md CHANGED
@@ -7,6 +7,90 @@ Prism 0.0.6 preserves documented 0.0.3 agent construction except for two intenti
7
7
  1. **`session.run()` / `session.prompt()` return `AgentRunResult`** and `session.stream()` starts one owned run after subscribing. Callers that ignored the previous `Promise<void>` keep working; failed/aborted runs reject with `AgentRunError` (`.result` attached).
8
8
  2. **`AgentConfig.extensions` / `settings` / `credentials` are removed.** Wire extensions through `createExtensionKernel()`, read settings in the host, and pass credential resolvers to the provider edge.
9
9
 
10
+ ## 0.0.10 → 0.0.11 coding harness fundamentals (additive)
11
+
12
+ Release **0.0.11** adds SessionIndex/search, assembler `contextBudget`, native Anthropic + Google provider packages, mid-run `steer`, coding-agent goal→verify + `ask_user_decision` (multi/free-text/suspend glue). Package count: **32 → 34** (adds `@arnilo/prism-provider-anthropic`, `@arnilo/prism-provider-google`). Version bump itself is Task 13 / release gate — treat this section as the behavioral migration map.
13
+
14
+ | Surface | Before (0.0.10) | After (0.0.11) |
15
+ | --- | --- | --- |
16
+ | Session search | No `searchSessions` / `SessionIndex` | Optional store search; SQLite/Postgres FTS migration `004_session_search` (schema **v4**); memory `sessionSearchMode: "linear" | "unsupported"` (default linear); JSONL throws `SessionSearchUnsupportedError` |
17
+ | Context budget | Assembler has no token/byte eviction | Opt-in `contextBudget` on `assembleProviderInput`; omission report via metadata helper |
18
+ | Providers | OpenCode Go Anthropic *route*; no first-party Google | `@arnilo/prism-provider-anthropic` (`createAnthropicProviderPackage`) + `@arnilo/prism-provider-google` (`createGoogleProviderPackage`); AI SDK remains escape hatch |
19
+ | Mid-run input | RPC `steer` unsupported / no queue | `AgentSession.steer` + RPC `steer` (queue 8 / 64 KiB; optional softInterrupt) |
20
+ | Coding helper | Compose manually from plan/checks/workflows | `runCodingGoalVerify` + `examples/coding-goal-verify.ts` |
21
+ | Ask user | n/a | Opt-in `createAskUserDecisionTool`; durable `suspendAskUserDecision` (no new agent interruption kinds) |
22
+ | Structured output + tools | Native schema attached every GVR provider turn | Opt-in `structuredOutputTiming: "final-turn-only"` (default `"every-turn"`): tool-eligible turns omit schema; artifact/revision turns schema-on / tools-off |
23
+
24
+ **Host actions:** reopen SQLite/Postgres stores so migration 004 applies; set `metadata.workspaceRoot` when filtering by workspace; wire Anthropic/Google packages explicitly; do not expect JSONL search. Benchmarks: `scripts/benchmark-0.0.11.mjs` (lands with release Task 13). See [Phase 6 evidence](review-coverage-2026-07-22-phase-6.md).
25
+
26
+ ## 0.0.9 / 0.0.96 → 0.0.10 coding workspace modes (breaking composition)
27
+
28
+ `@arnilo/prism-coding-security` composition now requires explicit `workspaceMode: "host" | "sandbox"`. Missing mode throws at construction. The `0.0.9` default that wired sandbox shell while keeping read/write/edit/list/search on the host cwd is **superseded** and fail-closed.
29
+
30
+ | Before (0.0.9) | After (0.0.10) |
31
+ | --- | --- |
32
+ | `createSandboxCodingTools(cwd, { sandbox })` — shell in sandbox, FS on host | Must pass `workspaceMode`. Prefer `createSandboxCodingComposition(...)`. |
33
+ | Silent split-brain treated as normal | Throws unless `allowMixedWorkspaceWiring: true` (warnings; `containmentClaim: false`). |
34
+ | No containment metadata | `composition.containmentClaim` / `warnings` / optional `treeIdentity`. Host mode never claims containment. |
35
+
36
+ ```ts
37
+ // Contained: one disposable tree
38
+ const { tools, composition } = createSandboxCodingComposition(sourceRoot, {
39
+ workspaceMode: "sandbox",
40
+ sandbox, // DisposableSandbox auto-wires FS backends
41
+ });
42
+
43
+ // Explicit host (non-contained)
44
+ createSandboxCodingTools(cwd, { workspaceMode: "host" });
45
+
46
+ // Escape hatch (documented split; no containment claim)
47
+ createSandboxCodingTools(cwd, {
48
+ workspaceMode: "sandbox",
49
+ sandbox,
50
+ allowMixedWorkspaceWiring: true,
51
+ });
52
+
53
+ // Same-tree Git
54
+ createGitTools(composition.workspaceRoot, {
55
+ execFile: sandbox.execFile.bind(sandbox),
56
+ commitIdentity: { name: "bot", email: "bot@example.com" },
57
+ });
58
+ ```
59
+
60
+ Docker defaults unchanged: digest-pinned image, non-root user, network none, absolute Docker CLI, no host-env inheritance. Unified mode adds no unbounded sync; caps stay in `sandbox-limits.ts` / coding-agent limits. Benchmark evidence: `scripts/benchmark-0.0.10.mjs`.
61
+
62
+ ## 0.0.8 → 0.0.9 release overview
63
+
64
+ All 32 first-party manifests and exact internal ranges move together to `0.0.9`; mixed first-party versions are unsupported. Core remains dependency-free at runtime and existing low-level agent/session APIs remain compatible. New coding sandbox, repository/Git, durable coding-plan, and browser surfaces are opt-in. `@arnilo/prism-browser` is included by `@arnilo/prism-all` but not by `@arnilo/prism-code` — install it explicitly when interactive browser automation is required. Office execution remains outside Prism packaging (host-selected skills/instructions only). No tag or publication is automatic from this migration.
65
+
66
+ ### Malformed streamed tool-call arguments (recoverable)
67
+
68
+ Malformed streamed tool-call JSON (id+name present) no longer terminates the run as `ProviderTransportError("invalid_json_arguments")`. First-party providers emit a tool call carrying `argumentsError`; dispatch blocks with `tool_execution_blocked` / `invalid_arguments` (`error.code: "invalid_json_arguments"`), never calls `execute()`, and the model can self-correct within existing turn/tool-round budgets. Prefer `toolCallFromArgumentsText` / `tryParseJsonObjectArguments` in custom providers.
69
+
70
+ ### Incomplete tool-call deltas (typed failure)
71
+
72
+ Tool-call deltas missing `id` and/or `name` at stream end no longer throw a bare `Error("Incomplete tool call delta...")`. Core reconstruction and the openai-compatible finalizer surface `ProviderTransportError` / `ErrorInfo.code: "incomplete_delta"`, fail the provider turn (no tool execution), and keep OpenCode Go / Kimi dangling fail-closed behavior. Distinguish from Defect 1a: missing identity fails the turn; present identity with bad JSON recovers via failed tool results.
73
+
74
+ ### Empty call-free artifact candidates (parse_error)
75
+
76
+ `generateValidateReviseLoop` treats empty/whitespace-only call-free assistant text (including thinking-only/reasoning-only turns) as `parse_error` before the host parser/identity default. Session runs succeed only after `artifact_finished`; terminal `artifact_failed` fails the run (`AgentRunError`, typically `error.code: "parse_error"`).
77
+
78
+ ## 0.0.9 coding-security Docker sandbox (additive)
79
+
80
+ `@arnilo/prism-coding-security` adds `createDockerSandbox()` / `DisposableSandbox` while preserving `SandboxAdapter.exec` and `createSandboxBashOperations()`. Hosts opt in with an absolute Docker executable and digest-pinned image; default network is none, host env is never inherited, and workspace export is an explicit bounded host callback. Existing approval-policy callers need no changes.
81
+
82
+ ## 0.0.9 coding-agent repository list/search (additive behavior change)
83
+
84
+ `@arnilo/prism-coding-agent` adds native `repo_list` / `repo_search` tools. `createCodingTools()` / `createAllTools()` now return six tools. **`createReadOnlyTools()` deliberately expands from `[read]` to `[read, repo_list, repo_search]`** — update hosts that asserted the previous read-only membership. Prefer `createSandboxCodingComposition(cwd, { workspaceMode, sandbox, repository })` (or the tools-only wrappers) from `@arnilo/prism-coding-security`. Pass required `workspaceMode`; sandbox mode keeps shell and FS/list/search on one disposable tree. The 0.0.9 split (sandbox shell + host FS) is superseded — see **0.0.9 / 0.0.96 → 0.0.10 coding workspace modes** above.
85
+
86
+ Opt-in structured Git/check tools are available via `createGitTools(cwd, { commitIdentity, checks? })` and are **not** added to `createCodingTools()`/`createAllTools()`. Commits require an explicit host `commitIdentity`; PR handoff returns bounded metadata/artifacts only and never pushes.
87
+
88
+ Durable coding-task composition uses existing workflows plus coding-agent helpers (`writeCodingPlanFile`, `buildCodingCheckpointMetadata`, `assertCodingResumeAllowed`). Plan/todos remain workspace Markdown; checkpoint state keeps only references/hashes/summaries/fingerprints under `state.coding`. No `CodingRun` or todo database is introduced. See `examples/durable-coding-workflow.ts`.
89
+
90
+ ## 0.0.9 browser automation (additive)
91
+
92
+ Install `@arnilo/prism-browser` explicitly (or through `@arnilo/prism-all`) for interactive browser tools. Hosts supply a pinned Playwright `Browser` (`playwright-core@1.61.0` optional peer); package import launches and downloads nothing. `createBrowserTools()` returns exactly `browser_open`, `browser_snapshot`, `browser_act`, and `browser_close` (all `exclusive: true`). Network policy defaults to require contained-proxy attestation; configure `uploads`/`downloads` for file transfer; `browser_act` adds `upload`/`screenshot`/`download_release`. Use `createBrowserManager().closeRun(runId)` / `close()` on terminal/abort. Align with a disposable sandbox via `createSharedSandboxBrowserOptions()` and `assertBrowserSandboxNetwork()`. CSS/XPath/evaluate/CDP/persistent profiles remain unsupported.
93
+
10
94
  ## 0.0.7 → 0.0.8 release overview
11
95
 
12
96
  All 31 first-party manifests and exact internal ranges move together to `0.0.8`; mixed first-party versions are unsupported. Core remains dependency-free at runtime and existing low-level agent/session APIs remain compatible. New telemetry, evaluation, MCP, A2A, ledger batching, and web research surfaces are opt-in. Release CI now requires CodeQL, dependency/license/SBOM/secret checks, packed-artifact attestations, PostgreSQL integration, and protected live-canary prerequisites; no tag or publication is automatic from this migration.
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- The optional `@arnilo/prism/node/session-store-jsonl` subpath stores `SessionEntry` records in a caller-named JSONL file: one JSON object per line.
5
+ The optional `@arnilo/prism/node/session-store-jsonl` subpath stores `SessionEntry` records in a caller-named JSONL file: one JSON object per line. `searchSessions` is unsupported and throws `SessionSearchUnsupportedError` (use memory linear mode or a DB adapter for search).
6
6
 
7
7
  APIs:
8
8
 
@@ -6,6 +6,42 @@ Evaluation defaults are finite: 100 trace rows × 20 pages and 4 MiB aggregate t
6
6
 
7
7
  This page states Prism runtime limits that keep slow consumers and long sessions from becoming unbounded memory or latency problems.
8
8
 
9
+ ## Release 0.0.11 session search / context budget / steer caps
10
+
11
+ Finite caps (defaults / hard) — full matrix in [Phase 6 evidence](review-coverage-2026-07-22-phase-6.md):
12
+
13
+ | Resource | Default / hard |
14
+ | --- | --- |
15
+ | Session search page | 20 / 100 |
16
+ | Search query string | 4 KiB / 16 KiB |
17
+ | Search snippet | 512 B / 4 KiB |
18
+ | Memory linear sessions / entries / bytes | 1000/5000 · 10000/50000 · 8 MiB/64 MiB |
19
+ | FTS candidates | 1000 / 5000 |
20
+ | Context budget tokens / bytes | caller-set / hard 2_000_000 tokens · 32 MiB |
21
+ | Context omission rows | 256 / 1024 |
22
+ | Pending steers | 8 messages / 64 KiB |
23
+
24
+ Run `node scripts/benchmark-0.0.11.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.11.test.mjs`. Default mode is network-free: memory-linear `searchSessions` (label + query) plus assembler `contextBudget` eviction/fit. Emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals. Search never default-scans an unbounded store; budget fails closed on mandatory prefix overflow; steer overflow fails closed. These are evidence fields, not CI timing gates.
25
+
26
+ ## Release 0.0.10 reproducible workspace-mode evidence
27
+
28
+ Run `node scripts/benchmark-0.0.10.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.10.test.mjs`. Default mode is network-free: host-composition write/read/list plus sandbox-fake composition write/read/list/search (in-memory `DisposableSandbox`). Emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals. Optional `PRISM_BENCH_DOCKER=1` (with `PRISM_TEST_DOCKER_*`) appends real local Docker composition rows. Unified workspace mode reuses existing sandbox/repo hard caps and adds no unbounded host↔container sync. These are evidence fields, not CI timing gates.
29
+
30
+ ## Release 0.0.9 reproducible coding/browser evidence
31
+
32
+ Run `node scripts/benchmark-0.0.9.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.9.test.mjs`. Default mode is network-free fake/in-process only and emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals for repository list/search, Git status, and browser open/snapshot/action/close. Optional `PRISM_BENCH_DOCKER=1` (with `PRISM_TEST_DOCKER_*`) and `PRISM_BENCH_PLAYWRIGHT=1` append real local Docker / protected Playwright rows. These are evidence fields, not CI timing gates.
33
+
34
+ 2026-07-21 baseline: Node v24.18.0, Linux x64, 100 iterations/scenario, network=false, credentials=false, docker=false, playwright=false.
35
+
36
+ | Scenario | mode | ops/s | p95 ms | heap bytes | disk bytes | processes | cost USD | backpressure | resource limits |
37
+ | --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
38
+ | repo-list | fake-in-process | 1,343 | 1.11 | 14,747,560 | 0 | 1 | 0 | 0 | 0 |
39
+ | repo-search | fake-in-process | 380 | 3.72 | 16,093,176 | 0 | 1 | 0 | 0 | 0 |
40
+ | git-status | fake-in-process | 479 | 2.62 | 13,821,688 | 0 | 1 | 0 | 0 | 0 |
41
+ | browser-open-snapshot-action-close | fake-in-process | 17,141 | 0.11 | 18,590,488 | 0 | 1 | 0 | 0 | 0 |
42
+
43
+ Rows exercise shipped repository/Git helpers and fake Playwright APIs only. Real Docker sandbox and Playwright browser timings remain explicit protected-gate evidence (`PRISM_TEST_DOCKER_SANDBOX=1`, `PRISM_LIVE_PLAYWRIGHT=1` / `PRISM_BENCH_DOCKER=1` / `PRISM_BENCH_PLAYWRIGHT=1`) because this release-candidate host did not enable those gates for the dated baseline. No live claim is inferred from skipped gates.
44
+
9
45
  ## Release 0.0.8 reproducible synthetic evidence
10
46
 
11
47
  Run `node scripts/benchmark-0.0.8.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000. Script uses no network/credentials and emits environment, throughput, p50/p95 latency, heap, synthetic disk bytes, zero external cost, and backpressure signals. These are evidence fields, not CI timing gates.
@@ -29,6 +65,16 @@ Security automation is isolated from `npm test`: CodeQL/supply-chain jobs have 1
29
65
 
30
66
  Web tools default/hard ceilings are query 4/16 KiB, results 10/20, URLs 5/20, request 256 KiB/1 MiB, response/aggregate 2/16 MiB, Markdown 1/8 MiB, extraction 256 KiB/1 MiB, schema 64/256 KiB, concurrency 4/16, retries 2/4, polling 20/100, and wall time 60 seconds/30 minutes. Bounds charge before request, retention, retry, or polling; overflow fails rather than truncating citation/extraction evidence.
31
67
 
68
+ Docker sandbox defaults/hard caps from `@arnilo/prism-coding-security`: startup 30 s/120 s; wall 20 min/30 min; idle 5 min/15 min; CPUs 2/8; memory 2 GiB/16 GiB (swap equal to memory); PIDs 256/1,024; FDs 1,024/8,192; workspace/tmp/download tmpfs 1 GiB/8 GiB, 256 MiB/2 GiB, 64 MiB/512 MiB; commands 100/256 with concurrent execs 1/8; env 64/256 names and 64 KiB/256 KiB values; export 50,000/250,000 entries and 256 MiB/2 GiB bytes with 16/64 retained artifacts; stop grace 5 s/30 s and cleanup 30 s/120 s. Caps validate before `docker create`/exec/export; overflow aborts and cleans the recorded container. Output still streams into the coding-agent `OutputAccumulator` ceilings (64 MiB/1 GiB).
69
+
70
+ Repository list/search defaults/hard caps from `@arnilo/prism-coding-agent`: depth 32/128; entries/files 10,000/100,000; page/results 1,000/10,000; search scan 64 MiB/1 GiB aggregate and 8 MiB/64 MiB per file; matches 1,000/10,000; pattern 512 B/4 KiB; line 50 KiB/1 MiB; context 5/20; wall 30 s/300 s; concurrency config 8/32. Walks stream via `opendir`/`lstat`, never follow symlink escapes, and stop immediately on aggregate limits or abort.
71
+
72
+ Structured Git/check/handoff defaults/hard caps: paths 1,000/10,000; refs 1 KiB/4 KiB; commit message 64 KiB/256 KiB; inline Git output 4 MiB/64 MiB; diff lines 10,000/100,000; changed files 1,000/10,000; patch input 16 MiB/64 MiB; worktrees 4/16; named checks 8/32 names, concurrency 1/4, timeout 10 min/60 min, diagnostic lines 2,000/100,000, output 4 MiB/64 MiB; PR handoff JSON 256 KiB/1 MiB with 100/1,000 commits. Git tools use typed argument arrays (never shell), disable hooks/credential prompts/external diff by default, and emit host-owned PR handoff data only — no push/network/PR client.
73
+
74
+ Durable coding plan/checkpoint defaults/hard caps: plan Markdown 256 KiB/1 MiB; todos 1,000/10,000 with 512 B/4 KiB text; checkpoint metadata 64 KiB/512 KiB; artifact references 16/64 at 256 MiB/2 GiB each; check summaries 1 KiB/8 KiB. Checkpoints store URI/hash/summaries/fingerprints only; resume revalidates workspace root, base branch, plan hash, and tool/policy/image fingerprints before import.
75
+
76
+ Browser automation defaults/hard caps from `@arnilo/prism-browser`: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2,000/10,000; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30 s/120 s; action 10 s/60 s; wait 30 s/120 s; run wall 20 min/30 min; popups 4/16; dialogs 16/64; listeners 64/256; action input 64 KiB/256 KiB; close grace 5 s/30 s; network requests 1,000/10,000 with 10/32 redirects per request and 8/32 WebSockets; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate. Caps charge before context/page/action/queue/snapshot/network/artifact retention. Host supplies Playwright and egress proxy attestation; package import launches nothing.
77
+
32
78
  Current surfaces:
33
79
 
34
80
  - `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
@@ -4,7 +4,7 @@
4
4
 
5
5
  The optional `@arnilo/prism-session-store-postgres` package ships a production-oriented PostgreSQL adapter that implements:
6
6
 
7
- - `SessionStore` — atomic `append` / `list` / `get` / `readBranchPath`
7
+ - `SessionStore` — atomic `append` / `list` / `get` / `readBranchPath` / bounded `searchSessions` / bounded `searchSessions`
8
8
  - `RunLedger` — durable run, event, tool-call, and usage rows
9
9
  - `ProductionPersistenceStore` — cursor-paginated `query*` reads plus generic `checkpoints` and atomic `leases` capabilities
10
10
 
@@ -59,7 +59,7 @@ Hosts own TLS (`ssl` in `poolConfig`), credentials, connection limits, and backu
59
59
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
60
60
  | `close()` | Ends the pool when the adapter created it from `connectionString`. |
61
61
 
62
- Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks. While holding that lock, startup verifies ordered contract name/version/SHA-256 rows and full schema-v3 `information_schema`/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
62
+ Migrations run automatically on open and are idempotent across reopen. Concurrent setup uses per-schema advisory transaction locks. While holding that lock, startup verifies ordered contract name/version/SHA-256 rows and full schema-v4 `information_schema`/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
63
63
 
64
64
  ## Request/response example
65
65
 
@@ -116,7 +116,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
116
116
  - The package is optional and workspace-local; `@arnilo/prism` core has no PostgreSQL dependency.
117
117
  - Schema names must match `^[a-zA-Z_][a-zA-Z0-9_]*$`; the adapter quotes them and never interpolates user values into identifier positions.
118
118
  - `SessionAppendOptions` idempotency rows are durable in `prism_session_append_idempotency` and survive reopen.
119
- - Schema version **3** applies `001_init`, additive `002_usage_scope`, and `003_run_feedback`. Migration 003 adds immutable `prism_run_feedback` rows with run FK/cascade deletion and owner/run/trace cursor indexes. `persistence.feedback` validates exact run ownership, bounds/redacts through optional `feedbackRedactor`, queries bounded pages, and deletes only exact-owned IDs. SQLite shares the same model with dialect-local DDL.
119
+ - Schema version **4** applies `001_init`, `002_usage_scope`, `003_run_feedback`, and `004_session_search`. Migration 003 adds immutable `prism_run_feedback` rows with run FK/cascade deletion and owner/run/trace cursor indexes. Migration 004 adds session search FTS (Postgres `tsvector` FTS table dual-written on append) plus `prism_sessions(updated_at, id)` cursor index; existing entries are backfilled once. `persistence.feedback` validates exact run ownership, bounds/redacts through optional `feedbackRedactor`, queries bounded pages, and deletes only exact-owned IDs. Search hits never include credentials; ownership filters apply when present. SQLite shares the same model with dialect-local DDL.
120
120
  - Pass an existing `pg` `Pool` when your host already manages pooling, TLS, and credential rotation.
121
121
 
122
122
  ## Security and performance notes