@arnilo/prism 0.0.8 → 0.0.96

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem tools as Prism `ToolDefinition` objects. It ships four tools — `shell`, `read`, `write`, `edit` — plus aggregator factories. The tools are **inert** until a host imports them and registers them into a `ToolRegistry`. Behavior is a behavioral port of the pi coding agent's `bash`/`read`/`write`/`edit` tools, adapted to Prism's `ToolDefinition` / `ToolResult` contracts (no `@earendil-works/*` or `typebox` dependencies; only `diff` plus the Node standard library).
5
+ `@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem/repository tools as Prism `ToolDefinition` objects. It ships six default coding tools — `shell`, `read`, `write`, `edit`, `repo_list`, `repo_search` — plus an opt-in structured Git/check set (`createGitTools`) for status/diff/branch/worktree/apply/commit/PR-handoff and named checks. Bounded coding-plan/checkpoint helpers compose ordinary workspace Markdown with workflow checkpoint state (references/hashes/summaries/fingerprints only). The tools are **inert** until a host imports them and registers them into a `ToolRegistry`. Behavior for shell/read/write/edit is a behavioral port of the pi coding agent's tools, adapted to Prism's `ToolDefinition` / `ToolResult` contracts (no `@earendil-works/*` or `typebox` dependencies; only `diff` plus the Node standard library). List/search/Git are native Prism tools with no glob/ripgrep/Git-library dependency.
6
6
 
7
7
  | Export | Purpose |
8
8
  | --- | --- |
@@ -10,13 +10,23 @@
10
10
  | `createReadTool(cwd, options?)` | `read` tool: read a text or image file into `TextContent` / `ImageContent`. |
11
11
  | `createWriteTool(cwd, options?)` | `write` tool: create or overwrite a file, creating parent directories. |
12
12
  | `createEditTool(cwd, options?)` | `edit` tool: precise exact-then-fuzzy text replacement in an existing file. |
13
- | `createCodingTools(cwd, options?)` | All four tools (`shell`, `read`, `write`, `edit`). |
14
- | `createReadOnlyTools(cwd, options?)` | Read-only subset: `read` only. |
15
- | `createAllTools(cwd, options?)` | Every tool the package provides (currently identical to `createCodingTools`). |
13
+ | `createRepoListTool(cwd, options?)` | `repo_list` tool: bounded deterministic repository listing. |
14
+ | `createRepoSearchTool(cwd, options?)` | `repo_search` tool: bounded literal/regex text search. |
15
+ | `createCodingTools(cwd, options?)` | Default six tools (`shell`, `read`, `write`, `edit`, `repo_list`, `repo_search`). |
16
+ | `createReadOnlyTools(cwd, options?)` | Read-only subset: `read`, `repo_list`, `repo_search`. |
17
+ | `createAllTools(cwd, options?)` | Identical to `createCodingTools` (Git tools remain opt-in via `createGitTools`). |
18
+ | `createGitTools(cwd, options?)` | Opt-in Git tools (`git_status`/`git_diff`/`git_branch`/`git_worktree`/`git_apply`/`git_commit`/`git_pr_handoff`) plus optional `coding_check`. |
19
+ | `createCodingCheckTool(cwd, options)` | Named host-declared checks; model selects only a name. |
20
+ | `createLocalRepositoryOperations(limits?)` | Default streaming Node filesystem backend for list/search. |
21
+ | `createGitOperations(options)` | Typed Git operations backend (argument arrays, safe config, finite output). |
22
+ | `buildCodingCheckpointMetadata` / `validateCodingCheckpointMetadata` / `assertCodingResumeAllowed` | Bounded durable coding-task metadata for workflow `state.coding` (no second runtime). |
23
+ | `writeCodingPlanFile` / `readCodingPlanFile` / `createCodingPlanMarkdown` / `parseCodingPlanTodos` | Workspace plan/todo Markdown helpers with finite byte/todo caps and hash verification. |
24
+ | `fingerprintJson` / `CODING_STATE_KEY` | Stable tool/policy fingerprints and the shared-state key for coding metadata. |
16
25
  | `detectSupportedImageMimeType(buf)` / `detectSupportedImageMimeTypeFromFile(path)` | Magic-byte image MIME detection (PNG/JPEG/GIF/WebP/BMP) used by `read`. |
17
26
  | `DEFAULT_MAX_IMAGE_BYTES` | Default `read` image size ceiling (10 MB). |
18
- | `DEFAULT_*` / `HARD_*` coding limit constants | Published text-scan, image, write/edit, shell timeout, display, and total-output ceilings. |
27
+ | `DEFAULT_*` / `HARD_*` coding limit constants | Published text-scan, image, write/edit, shell, repository, Git, check, handoff, and plan/checkpoint ceilings. |
19
28
  | `ReadTextOptions` / `ReadTextResult` | Bounded text-page contract required by custom `ReadOperations`. |
29
+ | `RepositoryOperations` / `RepositoryLimitOptions` | Pluggable list/search backend and finite caps. |
20
30
  | `TransformImage` / `TransformImageInput` | Types for the optional `read` `transformImage` callback. |
21
31
  | `withFileMutationQueue(path, fn)` | Per-path serialization primitive re-exported for hosts. |
22
32
 
@@ -55,6 +65,7 @@ const tools = createCodingTools(workspaceRoot, {
55
65
  | `read` | `read` |
56
66
  | `write` | `write` |
57
67
  | `edit` | `edit` |
68
+ | `repo_list` / `repo_search` | _(native; no pi equivalent)_ |
58
69
 
59
70
  ## Inputs / request
60
71
 
@@ -159,6 +170,80 @@ Each `edits[].oldText` must match a unique, non-overlapping region of the origin
159
170
 
160
171
  `edit` result `metadata`: `{ diff, patch, firstChangedLine }` — a display-oriented diff, a standard unified patch, and the first changed line in the new file. These are host-readable; the model only sees the short confirmation (keeps model context small).
161
172
 
173
+ ### `repo_list`
174
+
175
+ List repository entries with deterministic relative paths. Uses Node `opendir`/`lstat` only — no glob dependency. Does not follow symlinks; rejects path escapes outside the workspace root. Hidden names and excluded basenames (default `.git`, `node_modules`, `dist`) are skipped unless `includeHidden` is set / host `exclude` is overridden.
176
+
177
+ **Inputs:**
178
+
179
+ | Field | Type | Purpose |
180
+ | --- | --- | --- |
181
+ | `path` | `string` | Workspace-relative directory or file to list (default root). |
182
+ | `includeHidden` | `boolean` | Include dot names (default false). |
183
+ | `maxDepth` | `number` | Directory depth cap (default 32, hard 128). |
184
+ | `maxResults` | `number` | Page size (default 1,000, hard 10,000). |
185
+ | `offset` | `number` | Entries to skip before retaining (default 0). |
186
+
187
+ **Outputs:** text lines `kind\trelative/path[\tsize]` plus metadata (`truncated`, `truncatedBy`, `nextOffset`, `entries`, scan counts). Continue with `offset=nextOffset` when truncated by results.
188
+
189
+ ### `repo_search`
190
+
191
+ Search text files under the workspace. Default mode is literal substring match; `mode: "regex"` enables length-bounded regular expressions. Binary files (NUL in a bounded prefix) and oversize files are skipped. Aggregate scanned bytes, matches, line bytes, pattern bytes, and wall time are finite.
192
+
193
+ **Inputs:**
194
+
195
+ | Field | Type | Purpose |
196
+ | --- | --- | --- |
197
+ | `query` | `string` | Literal or regex pattern (required). |
198
+ | `path` | `string` | Workspace-relative start path. |
199
+ | `mode` | `"literal" \| "regex"` | Default `literal`. |
200
+ | `caseSensitive` | `boolean` | Default false. |
201
+ | `includeHidden` | `boolean` | Default false. |
202
+ | `context` | `number` | Context lines before/after each match (default 5, hard 20). |
203
+ | `maxMatches` | `number` | Match cap (default 1,000, hard 10,000). |
204
+
205
+ **Outputs:** ripgrep-like lines `path:line:column:text` with optional `path-` / `path+` context, plus metadata (`matches`, `truncated`, scan/skip counts).
206
+
207
+ ### Structured Git tools (`createGitTools`)
208
+
209
+ Opt-in tools over a host-pinned Git executable (`gitPath`, default `/usr/bin/git`) or sandbox `execFile`. Every invocation uses argument arrays with safe config (`core.hooksPath=/dev/null`, empty credential helper, pager disabled, `GIT_TERMINAL_PROMPT=0`). Shell is never used internally. Git tools are **not** included in `createCodingTools()` / `createAllTools()`.
210
+
211
+ | Tool | Purpose |
212
+ | --- | --- |
213
+ | `git_status` | `status --porcelain=v2 -z --branch` → structured branch + entries + `dirty`. |
214
+ | `git_diff` | Bounded `--no-ext-diff --no-textconv` diff; oversized output may spill via `artifactWriter`. |
215
+ | `git_branch` | `validate` / `list` / `create` / `switch` with `git check-ref-format --branch`. Switch refuses unrelated dirty trees unless `createCheckpoint=true`. |
216
+ | `git_worktree` | `list` / `add` / `remove` within finite worktree caps. |
217
+ | `git_apply` | `check` / `apply` / `reverse`; always `--check` before mutating apply. Apply requires clean/checkpoint; failures restore. |
218
+ | `git_commit` | Explicit-path `add` + `commit --no-verify -F <tempfile>`; requires host `commitIdentity`. Allows dirty entries that are exactly the requested paths; unrelated dirt requires checkpoint. Never pushes. |
219
+ | `git_pr_handoff` | Bounded `{ base, head, commits, changedPaths, diffstat, checks, artifact? }` for host PR creation. Never authenticates or opens a PR. |
220
+ | `coding_check` | Included when `checks` are declared: model selects only a name; executable/args/env are host-fixed. |
221
+
222
+ ```ts
223
+ import { createGitTools } from "@arnilo/prism-coding-agent";
224
+
225
+ const gitTools = createGitTools(workspaceRoot, {
226
+ gitPath: "/usr/bin/git",
227
+ commitIdentity: { name: "Prism Bot", email: "bot@example.com" },
228
+ checks: {
229
+ test: { file: "/usr/bin/npm", args: ["test"] },
230
+ },
231
+ });
232
+ ```
233
+
234
+ ### Durable coding plans and checkpoints
235
+
236
+ There is no `CodingRun`, todo database, or second approval engine. Persist executable plan/todos as ordinary workspace Markdown (for example `plans/<task>.md`) and store only bounded metadata under workflow `state.coding`:
237
+
238
+ | Field group | Stored in checkpoint | Not stored |
239
+ | --- | --- | --- |
240
+ | Plan / workspace export / patch artifacts | URI + SHA-256 + byte count | File contents, credentials, raw command output |
241
+ | Branch / worktree / base | Paths and ref names | Full diffs |
242
+ | Named checks | Name + exit code + short summary | Full stdout/stderr |
243
+ | Fingerprints | Workflow revision, definition hash, tool/policy fingerprints, optional image digest | Browser storage state, secrets, env |
244
+
245
+ Use `writeCodingPlanFile` / `readCodingPlanFile` for the workspace artifact, `buildCodingCheckpointMetadata` before `ctx.updateState({ coding })`, and `assertCodingResumeAllowed` before import/resume. Wrong owner/revision/hash/fingerprint fails closed. See `examples/durable-coding-workflow.ts` for a network-free plan → branch → edit → check → approval → handoff composition over `runWorkflow` / `resumeWorkflow` / `startWorkflowBackground`.
246
+
162
247
  ## Outputs / response / events
163
248
 
164
249
  Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`. `write` and `edit` serialize per realpath through `withFileMutationQueue` so concurrent calls targeting one file do not interleave. `shell` is marked `exclusive`; tool dispatch serializes it at the turn level. The package emits no events of its own; hosts observe tool execution through the normal Prism `AgentEvent` stream via `dispatchToolCall`.
@@ -197,10 +282,10 @@ Minimal drop-in for any Prism app:
197
282
  import { createToolRegistry } from "@arnilo/prism";
198
283
  import { createCodingTools, createReadOnlyTools } from "@arnilo/prism-coding-agent";
199
284
 
200
- // Full coding set (shell + read + write + edit) against the project root:
285
+ // Full coding set (shell + read + write + edit + repo_list + repo_search) against the project root:
201
286
  const tools = createToolRegistry(createCodingTools(process.cwd()));
202
287
 
203
- // Or a read-only set for inspection-only agents:
288
+ // Or a read-only set for inspection-only agents (read + repo_list + repo_search):
204
289
  const ro = createToolRegistry(createReadOnlyTools(process.cwd()));
205
290
  ```
206
291
 
@@ -227,17 +312,18 @@ const remoteWrite = createWriteTool("/repo", {
227
312
 
228
313
  ## Extension and configuration notes
229
314
 
230
- - **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
231
- - **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`. Invalid/non-finite/unsafe/above-hard-cap values throw during tool construction; request `timeout` errors before spawn.
232
- - **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override.
315
+ - **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. Custom `RepositoryOperations` must honor depth/entry/file/match/scan/time caps and abort. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
316
+ - **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`; list/search accept `repository` limits and shared aggregator `ToolsOptions.repository`.
317
+ - **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit?, list?, search?, repository? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override. Read-only membership is deliberately `read` + `repo_list` + `repo_search` (0.0.9 behavior change).
318
+ - **Sandbox composition.** Prefer `@arnilo/prism-coding-security` `createSandboxCodingTools(cwd, { sandbox, ... })` to wire shell through a `SandboxAdapter` while sharing repository options. Filesystem tools still use the host `cwd` unless custom operations are supplied; Docker tmpfs workspace mutations stay inside the container until export.
233
319
  - **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
234
320
  - No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
235
321
 
236
322
  ## Security and performance notes
237
323
 
238
- - **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
324
+ - **Host shell/filesystem access.** These tools run real commands and read/write/list/search real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
239
325
  - **Non-zero exit is not an error.** A failing command is a normal `shell` result (exit code in metadata); only timeout/abort/spawn failures are error results. Do not assume `error == undefined` means the command succeeded.
240
- - **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
326
+ - **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `repo_list`/`repo_search` stream walks and charge depth/entry/file/match/scan/time before retention. Structured Git tools use argument arrays with finite output/path/ref/message/patch caps, disable hooks/credential prompts/external diff by default, and never push or open PRs. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
241
327
  - **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
242
328
  - **Bounded image reads.** `read` rejects images over `maxImageBytes` (default 10 MB) by `stat` before read when possible; MIME is detected from magic bytes only. Optional `transformImage` is host-owned — the base package has no image-processing dependency.
243
329
 
@@ -252,8 +338,19 @@ const remoteWrite = createWriteTool("/repo", {
252
338
  | Edit target / input / count | 8 MiB / 2 MiB / 100 | 64 MiB / 16 MiB / 1,000 | before target read/matching/write |
253
339
  | Shell wall time | 600 seconds | 3,600 seconds | process-tree kill |
254
340
  | Shell total stdout+stderr | 64 MiB | 1 GiB | process-tree kill; spill removal |
255
-
256
- Every configurable value is a positive safe integer; Prism rejects rather than clamps invalid values. Limits control resources, not authority: they do not replace root containment, approval, validation, or a sandbox.
341
+ | Repo depth / entries / files / page | 32 / 10,000 / 10,000 / 1,000 | 128 / 100,000 / 100,000 / 10,000 | before descending/retaining next entry |
342
+ | Search scan / file / matches | 64 MiB / 8 MiB / 1,000 | 1 GiB / 64 MiB / 10,000 | before next file/match retention |
343
+ | Search pattern / line / context / time | 512 B / 50 KiB / 5 / 30 s | 4 KiB / 1 MiB / 20 / 300 s | before regex compile / line retain / deadline |
344
+ | Git paths / refs / message | 1,000 / 1 KiB / 64 KiB | 10,000 / 4 KiB / 256 KiB | before process/temp-file creation |
345
+ | Git output / diff lines / changed files / patch | 4 MiB / 10,000 / 1,000 / 16 MiB | 64 MiB / 100,000 / 10,000 / 64 MiB | stream before retain; artifact spill optional |
346
+ | Worktrees | 4 | 16 | before add |
347
+ | Named checks (names / concurrency / time / lines / output) | 8 / 1 / 10 min / 2,000 / 4 MiB | 32 / 4 / 60 min / 100,000 / 64 MiB | construction / before start / line retention |
348
+ | PR handoff JSON / commits | 256 KiB / 100 | 1 MiB / 1,000 | before result exposure |
349
+ | Plan markdown / todos / todo text | 256 KiB / 1,000 / 512 B | 1 MiB / 10,000 / 4 KiB | before write/parse/checkpoint |
350
+ | Coding checkpoint metadata / artifact refs / artifact bytes | 64 KiB / 16 / 256 MiB | 512 KiB / 64 / 2 GiB | before state save / resume verify |
351
+ | Check summary text | 1 KiB | 8 KiB | before checkpoint retention |
352
+
353
+ Every configurable value is a positive safe integer (context may be zero); Prism rejects rather than clamps invalid values. Limits control resources, not authority: they do not replace root containment, approval, validation, or a sandbox.
257
354
 
258
355
  ## Related APIs
259
356
 
@@ -2,12 +2,15 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects.
5
+ `@arnilo/prism-coding-security` is an optional package that supplies structured execution policy for `@arnilo/prism-coding-agent` tools and one disposable Docker/OCI sandbox reference. It complements name-based `PermissionPolicy` at dispatch time with path/command context checked **inside** each tool before side effects, and optionally contains untrusted coding work in a host-invoked container.
6
6
 
7
7
  | Export | Purpose |
8
8
  | --- | --- |
9
9
  | `createCodingApprovalPolicy(options)` | Returns an `ExecutionPolicy` with trusted roots, read-only mode, command allow/deny rules, approval caching, and timeout/abort-aware approval waits. |
10
10
  | `createSandboxBashOperations(adapter)` | Maps a host-owned `SandboxAdapter` to coding-agent `BashOperations` for delegated shell execution. |
11
+ | `createSandboxCodingTools(cwd, options)` | One construction path: full coding tools with shell wired to `options.sandbox` and shared repository options. |
12
+ | `createSandboxReadOnlyTools(cwd, options)` | Read-only coding tools (`read`/`repo_list`/`repo_search`) with shared repository options. |
13
+ | `createDockerSandbox(options)` | Creates one disposable non-root Docker container with read-only root/source, bounded tmpfs workspace, typed `execFile`, import/export, and stop/kill/cleanup. |
11
14
  | `assertPathInsideRoots`, `isPathInsideReal` | Symlink-aware path containment helpers. |
12
15
  | `evaluateCommandRules`, `hasShellMetacharacters` | Command classification helpers. |
13
16
 
@@ -21,7 +24,7 @@ import type { ExecutionAction, ExecutionPolicy, ExecutionDecision } from "@arnil
21
24
 
22
25
  Use this package when coding tools need path scoping, human approval, command rules, or a pluggable sandbox backend. Wire the returned policy through `createCodingTools(cwd, { executionPolicy })` or per-tool `executionPolicy` options.
23
26
 
24
- Prism does **not** claim OS-level isolation unless the host provides a sandbox adapter. Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
27
+ Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
25
28
 
26
29
  ## Inputs / request
27
30
 
@@ -36,10 +39,25 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
36
39
 
37
40
  `run` caching keys decisions by the tool execution context's `runId`; `session` uses `sessionId`. Coding tools pass both identities to the policy. A missing/empty identity disables caching for that check rather than creating a global bucket. Identical actions in different runs/sessions never share approvals or denials.
38
41
 
42
+ ### Docker sandbox inputs
43
+
44
+ | Option | Default | Purpose |
45
+ | --- | --- | --- |
46
+ | `docker` | required | Absolute host Docker executable. |
47
+ | `image` | required | Digest-pinned image (`name@sha256:<64-hex>`). Never pulled (`--pull=never`). |
48
+ | `sourceRoot` | required | Absolute host directory imported into `/workspace`. |
49
+ | `user` | required | Non-root `uid:gid`. |
50
+ | `network` | `{ mode: "none" }` | Default no network; custom mode requires a pre-created network name and does not claim DNS containment. |
51
+ | `env` | `{}` | Exact allow-list only; host environment is never inherited. |
52
+ | `secrets` | `[]` | Canaries redacted from CLI/adapter errors. |
53
+ | `limits` | package defaults | CPU/memory/PID/FD/tmpfs/command/export/time caps validated before create. |
54
+
39
55
  ## Outputs / response / events
40
56
 
41
57
  `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations`, receive `onData(Buffer)` for ordered stdout/stderr forwarding through the shell tool's existing bounded accumulator, and never grant policy approval themselves.
42
58
 
59
+ `createDockerSandbox()` returns a `DisposableSandbox`: typed `execFile(file, args)`, shell-compatible `exec`, `status`, cooperative `stop`, forced `kill`, and idempotent `close`. `close({ export })` can stream a bounded workspace tar plus SHA-256/entry/byte metadata through a host callback; checkpoints should retain only host artifact references/hashes, never whole workspaces.
60
+
43
61
  ## Request/response example
44
62
 
45
63
  ```json
@@ -52,8 +70,11 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
52
70
  ## Implementation example
53
71
 
54
72
  ```ts
55
- import { createCodingTools } from "@arnilo/prism-coding-agent";
56
- import { createCodingApprovalPolicy, createSandboxBashOperations } from "@arnilo/prism-coding-security";
73
+ import {
74
+ createCodingApprovalPolicy,
75
+ createDockerSandbox,
76
+ createSandboxCodingTools,
77
+ } from "@arnilo/prism-coding-security";
57
78
 
58
79
  const policy = createCodingApprovalPolicy({
59
80
  roots: [workspaceRoot],
@@ -62,27 +83,48 @@ const policy = createCodingApprovalPolicy({
62
83
  approvalTimeoutMs: 60_000,
63
84
  });
64
85
 
65
- const tools = createCodingTools(workspaceRoot, {
86
+ const sandbox = await createDockerSandbox({
87
+ docker: "/usr/bin/docker",
88
+ image: "registry.example/prism-code@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
89
+ sourceRoot: "/srv/jobs/task-1/source",
90
+ user: "10001:10001",
91
+ network: { mode: "none" },
92
+ env: { CI: "1" },
93
+ limits: { cpus: 2, memoryBytes: 2 * 1024 ** 3, maxPids: 256, workspaceBytes: 1024 ** 3 },
94
+ });
95
+
96
+ // Host cwd is the inspected workspace; shell runs inside the sandbox.
97
+ const tools = createSandboxCodingTools("/srv/jobs/task-1/source", {
98
+ sandbox,
66
99
  executionPolicy: policy,
67
- shell: {
68
- operations: createSandboxBashOperations(mySandboxAdapter),
69
- },
100
+ repository: { exclude: [".git", "node_modules", "dist"] },
101
+ });
102
+
103
+ await sandbox.execFile({ file: "npm", args: ["test"], cwd: "/workspace" });
104
+ await sandbox.close({
105
+ export: async (stream, meta) => hostArtifacts.write(stream, meta),
70
106
  });
71
107
  ```
72
108
 
73
109
  ## Extension and configuration notes
74
110
 
75
- Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers.
111
+ Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()`/`createSandboxCodingTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` / `DisposableSandbox` are replaceable and host-owned; approval policy and sandboxing are separate layers. Custom remote sandboxes can implement `DisposableSandbox` without using Docker. `createSandboxCodingTools()` wires shell through the adapter while list/search/read/write/edit keep the host `cwd` unless custom operations are supplied — Docker tmpfs mutations remain inside the container until export. Opt-in structured Git tools from `@arnilo/prism-coding-agent` (`createGitTools`) can target the same disposable sandbox by passing `execFile: sandbox.execFile` and a host `commitIdentity`; Prism still never pushes or opens PRs. Optional `@arnilo/prism-browser` can share the same disposable boundary: use `assertBrowserSandboxNetwork()` before browse-ready custom networks, and `createSharedSandboxBrowserOptions({ workspaceRoot, downloadsRoot, containedProxyAttestation })` so uploads/downloads align with `/workspace` and `/downloads`. Close the browser context before disposing the sandbox.
112
+
113
+ The Docker reference adapter starts by recorded container ID/label, uses argument arrays only, mounts source read-only, populates a size-bounded tmpfs `/workspace`, drops all capabilities, enables `no-new-privileges`, runs with `--init`, and never exposes the Docker socket, privileged mode, or host PID/IPC namespaces. Image pull/build/update stays outside Prism. Protected real-Docker checks are opt-in via `PRISM_TEST_DOCKER_SANDBOX=1` with host-supplied `PRISM_TEST_DOCKER_BIN` and digest-pinned `PRISM_TEST_DOCKER_IMAGE`.
76
114
 
77
115
  Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
78
116
 
79
117
  ## Security and performance notes
80
118
 
81
- Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
119
+ Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, repository list/search walks, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe.
120
+
121
+ Docker sandbox containment—not command regexes—enforces filesystem/network/process boundaries for the reference adapter. Network defaults to none; a custom Docker network still requires a host firewall/proxy for DNS/egress claims. Import rejects symlink escapes, devices, FIFOs, and sockets; export counts entries/bytes and hashes before host retention. Secrets in `secrets` are redacted from adapter errors and never exported as environment metadata. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter and Docker daemon.
82
122
 
83
123
  ## Related APIs
84
124
 
85
- - [Coding agent tools](coding-agent-tools.md)
125
+ - [Coding agent tools](coding-agent-tools.md): durable plan/todo Markdown helpers and `state.coding` checkpoint metadata for restart/resume without a second runtime
126
+ - [Workflows](workflows.md): `runWorkflow` / `resumeWorkflow` / `startWorkflowBackground` composition for coding tasks
86
127
  - [Host security guide](host-security.md)
128
+ - [Performance limits](performance.md)
87
129
  - [Tool execution primitives](tool-execution-primitives.md)
88
130
  - [Security/auth/trust](settings-auth-trust-security.md)
@@ -135,11 +135,22 @@ const comparison = await runComparison({ dataset, candidates: { baseline, candid
135
135
  assertEvaluationThreshold(report, { minimumMean: 0.9, maximumFailures: 0 });
136
136
  ```
137
137
 
138
- `traceResolver` is explicit; no arbitrary run search occurs. `baseline`/`candidate` are host functions returning `AgentRunResult`. See `examples/evaluation-gate.ts` for a network-free gate.
138
+ `traceResolver` is explicit; no arbitrary run search occurs. `baseline`/`candidate` are host functions returning `AgentRunResult`. See `examples/evaluation-gate.ts` for a network-free gate and `examples/coding-browser-evaluation.ts` for coding/browser adversarial fixtures.
139
+
140
+ ## Coding and browser adversarial evaluations (0.0.9)
141
+
142
+ Release 0.0.9 ships curated network-free adversarial fixtures in package tests:
143
+
144
+ - `@arnilo/prism-coding-agent` `eval-fixtures.test.ts`: safe native list vs shell, Git path/ref injection, dirty-tree rollback, unknown named-check failure, PR-handoff artifact completeness, and prompt-injection file content under read-only tools.
145
+ - `@arnilo/prism-browser` `eval-fixtures.test.ts`: stale snapshot refs, side-effect approval, private/loopback/file deny, upload/download/screenshot policy, CSS/evaluate target rejection, and hostile accessible-name text.
146
+
147
+ Fixtures reuse `@arnilo/prism-evals` (`defineDataset` / `defineScorer` / `scoreRun` / `assertEvaluationThreshold` / `serializeEvaluationReport`). Optional SWE-bench-compatible or live-browser harnesses remain host adapters — they are not default dependencies or quality claims. Protected real Docker/Playwright gates stay env-gated (`PRISM_TEST_DOCKER_SANDBOX`, `PRISM_LIVE_PLAYWRIGHT`) and never enter `sdk:ready`.
139
148
 
140
149
  ## Related APIs
141
150
 
142
151
  - [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
143
152
  - [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
144
153
  - [Observability](observability.md): use `onTraceReference` or bounded `traceId(runId)` to supply `ScoreRunOptions.traceId`; evaluation telemetry emits no reason/explanation content
145
- - [Release and install](release-and-install.md): optional package install
154
+ - [Coding agent tools](coding-agent-tools.md) / [Browser automation](browser-automation.md) / [Workflows](workflows.md): network-free coding-task composition at `examples/durable-coding-workflow.ts`; adversarial coding/browser eval example at `examples/coding-browser-evaluation.ts`
155
+ - [Performance limits](performance.md): `scripts/benchmark-0.0.9.mjs` coding/browser evidence fields
156
+ - [Release and install](release-and-install.md): optional package install and protected sandbox-browser workflow
@@ -65,11 +65,12 @@ Guardrails are callbacks supplied by the host. Prism does not discover, load, re
65
65
 
66
66
  ## Security and performance notes
67
67
 
68
- Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work.
68
+ Output buffering prevents blocked provider content from reaching subscribers, session entries, ledgers, parsers, delegation, or tools. Tool-output checks receive raw results but Prism discards blocked raw output before event, ledger, transcript, or MCP exposure. Redaction replaces exact known values only; it is not general secret detection. Parallel checks receive an abort signal, but callback code must honor it to stop in-flight work. Browser snapshots and page text from `@arnilo/prism-browser` are untrusted external content: never allow them to modify tools, permissions, credentials, or policy. Browser mutations still require host `ExecutionPolicy`/approval; prompt-injection text in a page cannot grant upload/download release.
69
69
 
70
70
  ## Related APIs
71
71
 
72
72
  - [Agent/session runtime](agent-session-runtime.md)
73
73
  - [Tools](tools.md)
74
+ - [Browser automation](browser-automation.md)
74
75
  - [Agent events](agent-events.md)
75
76
  - [Host security](host-security.md)
@@ -46,7 +46,7 @@ Security controls fail closed before side effects when wired at the guarded edge
46
46
  - validator failures emit `tool_execution_blocked` with `validation_failed`
47
47
  - configured guardrails fail closed; output stages buffer blocked provider/tool content before events, ledgers, session entries, or MCP responses
48
48
  - configured redactors scrub provider requests, agent events, session entries, ledger records, tool errors, extension errors, injector context, and durable run checkpoints
49
- - durable resume requires host-derived exact ownership and checkpoint version; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
49
+ - durable resume requires host-derived exact ownership and checkpoint version; coding-task resume also revalidates plan/workspace artifact hashes plus tool/policy/image fingerprints via `assertCodingResumeAllowed` before import; `createAgentRunLifecycle()` exposes only public state through explicitly selected server/MCP capabilities; never accept ownership or resume input from an approval body
50
50
 
51
51
  These checks are explicit function calls during load, assembly, dispatch, append, or run handling. Prism adds no background watchers, filesystem scanners, network probes, credential polling, or automatic extension discovery.
52
52
 
@@ -137,7 +137,8 @@ Wire those values where they matter: provider adapters receive the resolved cred
137
137
  - Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
138
138
  - MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands/resources/prompts, requires per-operation `authorize`, and retains core gates. Sampling, roots, model/credential selection, and elicitation consent stay host-owned; URL elicitation is never opened automatically. Stateful web mode requires host `resolveAuthInfo` plus `resolveIdentity`, exact origin policy, and binds every POST/GET/DELETE/SSE request to one non-secret principal; mismatches return 404. Handler still needs TLS and edge rate limiting. See [MCP client/server exposure](mcp-tools.md).
139
139
  - `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
140
- - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Limits are not containment: Prism provides no OS sandbox unless the host supplies one.
140
+ - Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, repository list/search depth/entry/match/scan/time caps, structured Git path/ref/message/output/patch/worktree caps, named-check concurrency/output caps, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Opt-in `createGitTools()` uses argument arrays with hooks/credential prompts/external diff disabled, requires host `commitIdentity` for commits, and never pushes or opens PRs. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell/repository backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, identity-scoped approval caching, `createSandboxCodingTools()`, and the optional `createDockerSandbox()` reference adapter. Limits alone are not containment: construct the Docker adapter (absolute CLI, digest-pinned image, network none by default) or an equivalent host sandbox before treating coding execution as production-safe. Docker daemon/image trust, egress firewall/proxy, and artifact retention remain host-owned.
141
+ - Optional `@arnilo/prism-browser` requires a host-supplied Playwright Browser (`playwright-core@1.61.0` peer). Import is inert. One non-persistent context belongs to one run; actions serialize; refs are snapshot-scoped; CSS/evaluate/CDP/persistent profiles are denied. Context routing + `serviceWorkers: "block"` deny file/data/blob/devtools/private/loopback by default and require contained-proxy attestation for external egress (Playwright routing is defense in depth, not DNS containment). Uploads are realpath-rooted; downloads quarantine with hash/MIME until host `approveRelease`; screenshots return bounded `ImageContent`. Observation vs mutation/high-impact actions map to `ExecutionPolicy`. Treat snapshot/page text as untrusted external content. Close contexts with `browser_close` or `manager.closeRun(runId)` on abort/terminal. Browser control endpoint, binary/image pin, and real egress firewall/proxy remain host-owned. Shared sandbox: `createSharedSandboxBrowserOptions()` + `assertBrowserSandboxNetwork()`.
141
142
  - `@arnilo/prism-credentials-node` rejects oversized/malformed envelopes and excessive scrypt work before KDF allocation, uses async scrypt, and requires restrictive existing/new Unix vault modes. Keep vault ownership and parent-directory access host-controlled; review before `chmod 600`, never auto-weaken a file policy. Keychain calls use abort-aware native async work with finite timeout/payload caps and sanitized errors. OS prompts, service availability, and whether a native backend promptly honors cancellation remain host/platform boundaries; no plaintext fallback is attempted.
142
143
  - LLM compaction always sends finite summary `maxTokens`, retains bounded deltas/events, and bounds/redacts provider/factory/policy error detail. Observational-memory workers cap turns, calls, arguments, results, transcript, and surfaced errors; unknown tools fail before execution, while invalid results can only be rejected after a host tool returns and may therefore follow side effects. Pass all known provider/credential/tool secrets into compaction/runtime options; exact replacement is not secret discovery.
143
144
  - Default remote-media loading resolves every DNS answer, rejects the hostname if any address is non-public, and pins one validated address through the request. Explicit `allowedHostnames` can trust private destinations. A host-supplied `fetch` owns DNS/rebinding/proxy/redirect safety; a custom `requestUrl` must connect to its supplied validated address.
@@ -183,6 +184,7 @@ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy
183
184
  - `scripts/scan-secrets.mjs` checks tracked source and unpacked public tarballs for high-confidence credential/private-key forms without printing matched values. It complements GitHub secret scanning; it is not entropy scanning or DLP.
184
185
  - Tag publication alone receives npm/OIDC/attestation permissions. Untrusted pull-request code receives no canary, npm, or OIDC secret and no workflow uses `pull_request_target`.
185
186
  - Scheduled/manual canaries run only in protected `live-canaries` environment. Use dedicated read-only/low-quota credentials and provider account spend limits. Runner performs four probes, at most one MCP cleanup, one provider output token, one Brave result, 64-KiB responses, and finite timeouts; report excludes endpoints, headers, bodies, credentials, and MCP session IDs.
187
+ - Scheduled/manual coding/browser containment checks run in protected `sandbox-browser` environment (`.github/workflows/sandbox-browser.yml`). They receive no provider/npm/OIDC secrets; Docker/Playwright enablement is variable-gated with host-preloaded digest-pinned images/binaries; uploads are redacted aggregate status only.
186
188
  - Live endpoint operators own TLS, egress allow-lists, account-dollar budget, cleanup beyond MCP session DELETE, and revocation. Failed canaries log only operation kind plus status/timeout; inspect provider-side audit logs for details.
187
189
 
188
190
  ## Related APIs
package/docs/index.md CHANGED
@@ -12,9 +12,9 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
12
12
  - [Guardrails](guardrails.md): typed fail-closed input/output/tool checks with buffered provider output and redacted decision records.
13
13
  - [Agent events](agent-events.md): redacted lifecycle stream used by UIs, ledgers, and metadata-only parented telemetry; message/progress deltas never create spans.
14
14
  - [Observability](observability.md): OTel GenAI agent/provider/tool hierarchy, host context parenting, bounded trace linkage, safe evaluation events, controlled metrics, and exporter isolation.
15
- - [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, and ID-only linkage to immutable owned run feedback.
15
+ - [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, coding/browser adversarial fixtures, and ID-only linkage to immutable owned run feedback.
16
16
  - [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence, optional bounded FIFO durability policies, session snapshot caching, and immutable run/trace feedback.
17
- - [Performance limits](performance.md): bounded evaluation traces/judges/reports, security scan/live-canary backstops, live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
17
+ - [Performance limits](performance.md): bounded evaluation traces/judges/reports, 0.0.9 coding/browser benchmark evidence, security scan/live-canary backstops, live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
18
18
  - [Structured output](structured-output.md): the `Artifact*` seam plus provider-native `StructuredOutputOptions` / `structuredOutputMode` for capable models.
19
19
 
20
20
  ## Compaction/session memory
@@ -27,7 +27,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
27
27
  - [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention, and NoSQL mapping.
28
28
  - [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, and transactionally verified/backfilled migration-v3 metadata.
29
29
  - [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
30
- - [Migration guide](migration.md): 0.0.3 compatibility through 0.0.8 telemetry/evaluation, MCP/A2A, web research, ledger batching, and release-security changes.
30
+ - [Migration guide](migration.md): 0.0.3 compatibility through 0.0.9 coding/browser sandbox, repository/Git, durable plans, and Playwright automation changes.
31
31
  - [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety.
32
32
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
33
33
 
@@ -59,8 +59,9 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
59
59
  - [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
60
60
  - [MCP client bridge and server exposure](mcp-tools.md): SDK-1.29.0 bounded tools/resources/prompts, host-owned roots/sampling/elicitation, exact-origin DNS-pinned client transport, and principal-bound opt-in Streamable HTTP sessions.
61
61
  - [Web search, fetch, and extraction](web-tools.md): optional host-selected Brave/Exa discovery and Firecrawl Markdown/schema tools with native fetch, stable citations, late credentials, finite limits, and explicit untrusted-content boundaries.
62
- - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, and `edit` definitions with streamed text pages, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
63
- - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, and abort-aware streaming sandbox adapters for coding tools.
62
+ - [Browser automation](browser-automation.md): optional `@arnilo/prism-browser` with host-supplied Playwright contexts, AI-mode snapshots/refs, ordered `browser_open`/`browser_snapshot`/`browser_act`/`browser_close`, egress/side-effect/upload/download/screenshot policy, and finite page/action/snapshot/network/artifact caps.
63
+ - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, `edit`, `repo_list`, and `repo_search` definitions plus opt-in `createGitTools()` / `coding_check` for structured Git status/diff/branch/worktree/apply/commit/PR-handoff and named checks; durable plan/todo Markdown helpers with workflow `state.coding` checkpoint metadata; streamed text pages, bounded repository list/search, finite Git/check/plan caps, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
64
+ - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, abort-aware streaming sandbox adapters, `createSandboxCodingTools()` composition, and the disposable Docker/OCI sandbox reference with bounded workspace import/export.
64
65
 
65
66
  ## Extensions/plugins
66
67
  - [Contribution discovery (workspace)](contribution-discovery.md): opt-in, realpath-contained directory scanner turning `SKILL.md`/`manifest.json` into inert `DiscoveredContribution` envelopes the host registers — no `import()`, no auto-activate, no provider scanning. Per-agent bundles remain app-controlled and are documented under Agent/session runtime.
@@ -83,7 +84,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
83
84
 
84
85
  ## CLI/RPC
85
86
  - [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
86
- - [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Interactive TUI (C-012) deferred.
87
+ - [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Compose coding plans/checkpoints via workspace Markdown + `state.coding` without a second runtime. Interactive TUI (C-012) deferred.
87
88
  - [Workflow orchestration primitives](workflow-orchestration-primitives.md): architecture inventory — workflow adapters consume core `CheckpointStore`, `LeaseStore`, and bounded `EventMultiplexer`; run control and optional RPC commands stay package-local.
88
89
  - [Workflow/TUI scope](workflow-tui-primitives.md): records why 0.0.5 ships workflow APIs/RPC control but no interactive terminal UI.
89
90
 
@@ -104,7 +105,8 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
104
105
  - `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, cache-aware prompt assembly, NeuralWatt agent run, stores/branching, compaction, observational-memory recall, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
105
106
 
106
107
  ## Release and install
107
- - [Release and install](release-and-install.md): 31-package graph, install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, and protected live canaries.
108
+ - [Release and install](release-and-install.md): 32-package graph (including optional browser), install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, protected live canaries, and sandbox-browser Docker/Playwright gates.
109
+ - [Review coverage (2026-07-20 Phase 4)](review-coverage-2026-07-20-phase-4.md): Plan 072 evidence freeze — revised coding/browser-only scope, external revisions, primitive ownership, finite limits, threats, and 0.0.9 release gates.
108
110
  - [Review coverage (2026-07-19 Phase 3)](review-coverage-2026-07-19-phase-3.md): Plan 070 evidence freeze — exact protocol/vendor references, capability/primitive/limit matrices, supported boundaries, and 0.0.8 release evidence.
109
111
  - [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md): Plan 067 evidence freeze — P0–P2 re-verification owners, seven first-party provider packages mapped to official-doc URLs, Pi secondary refs, cache/thinking/discovery surfaces, credential canaries, and use-case model-binding inventory.
110
112
  - [Review coverage (2026-07-15)](review-coverage-2026-07-15.md): frozen 0.0.5 finding/feature ownership, existing-primitive inventory, package decisions, threat boundaries, exclusions, and measured Phase 0 baseline.
package/docs/migration.md CHANGED
@@ -7,6 +7,38 @@ Prism 0.0.6 preserves documented 0.0.3 agent construction except for two intenti
7
7
  1. **`session.run()` / `session.prompt()` return `AgentRunResult`** and `session.stream()` starts one owned run after subscribing. Callers that ignored the previous `Promise<void>` keep working; failed/aborted runs reject with `AgentRunError` (`.result` attached).
8
8
  2. **`AgentConfig.extensions` / `settings` / `credentials` are removed.** Wire extensions through `createExtensionKernel()`, read settings in the host, and pass credential resolvers to the provider edge.
9
9
 
10
+ ## 0.0.8 → 0.0.9 release overview
11
+
12
+ All 32 first-party manifests and exact internal ranges move together to `0.0.9`; mixed first-party versions are unsupported. Core remains dependency-free at runtime and existing low-level agent/session APIs remain compatible. New coding sandbox, repository/Git, durable coding-plan, and browser surfaces are opt-in. `@arnilo/prism-browser` is included by `@arnilo/prism-all` but not by `@arnilo/prism-code` — install it explicitly when interactive browser automation is required. Office execution remains outside Prism packaging (host-selected skills/instructions only). No tag or publication is automatic from this migration.
13
+
14
+ ### Malformed streamed tool-call arguments (recoverable)
15
+
16
+ Malformed streamed tool-call JSON (id+name present) no longer terminates the run as `ProviderTransportError("invalid_json_arguments")`. First-party providers emit a tool call carrying `argumentsError`; dispatch blocks with `tool_execution_blocked` / `invalid_arguments` (`error.code: "invalid_json_arguments"`), never calls `execute()`, and the model can self-correct within existing turn/tool-round budgets. Prefer `toolCallFromArgumentsText` / `tryParseJsonObjectArguments` in custom providers.
17
+
18
+ ### Incomplete tool-call deltas (typed failure)
19
+
20
+ Tool-call deltas missing `id` and/or `name` at stream end no longer throw a bare `Error("Incomplete tool call delta...")`. Core reconstruction and the openai-compatible finalizer surface `ProviderTransportError` / `ErrorInfo.code: "incomplete_delta"`, fail the provider turn (no tool execution), and keep OpenCode Go / Kimi dangling fail-closed behavior. Distinguish from Defect 1a: missing identity fails the turn; present identity with bad JSON recovers via failed tool results.
21
+
22
+ ### Empty call-free artifact candidates (parse_error)
23
+
24
+ `generateValidateReviseLoop` treats empty/whitespace-only call-free assistant text (including thinking-only/reasoning-only turns) as `parse_error` before the host parser/identity default. Session runs succeed only after `artifact_finished`; terminal `artifact_failed` fails the run (`AgentRunError`, typically `error.code: "parse_error"`).
25
+
26
+ ## 0.0.9 coding-security Docker sandbox (additive)
27
+
28
+ `@arnilo/prism-coding-security` adds `createDockerSandbox()` / `DisposableSandbox` while preserving `SandboxAdapter.exec` and `createSandboxBashOperations()`. Hosts opt in with an absolute Docker executable and digest-pinned image; default network is none, host env is never inherited, and workspace export is an explicit bounded host callback. Existing approval-policy callers need no changes.
29
+
30
+ ## 0.0.9 coding-agent repository list/search (additive behavior change)
31
+
32
+ `@arnilo/prism-coding-agent` adds native `repo_list` / `repo_search` tools. `createCodingTools()` / `createAllTools()` now return six tools. **`createReadOnlyTools()` deliberately expands from `[read]` to `[read, repo_list, repo_search]`** — update hosts that asserted the previous read-only membership. Prefer `createSandboxCodingTools(cwd, { sandbox, repository })` from `@arnilo/prism-coding-security` when shell must run inside a sandbox adapter while list/search inspect a host workspace path.
33
+
34
+ Opt-in structured Git/check tools are available via `createGitTools(cwd, { commitIdentity, checks? })` and are **not** added to `createCodingTools()`/`createAllTools()`. Commits require an explicit host `commitIdentity`; PR handoff returns bounded metadata/artifacts only and never pushes.
35
+
36
+ Durable coding-task composition uses existing workflows plus coding-agent helpers (`writeCodingPlanFile`, `buildCodingCheckpointMetadata`, `assertCodingResumeAllowed`). Plan/todos remain workspace Markdown; checkpoint state keeps only references/hashes/summaries/fingerprints under `state.coding`. No `CodingRun` or todo database is introduced. See `examples/durable-coding-workflow.ts`.
37
+
38
+ ## 0.0.9 browser automation (additive)
39
+
40
+ Install `@arnilo/prism-browser` explicitly (or through `@arnilo/prism-all`) for interactive browser tools. Hosts supply a pinned Playwright `Browser` (`playwright-core@1.61.0` optional peer); package import launches and downloads nothing. `createBrowserTools()` returns exactly `browser_open`, `browser_snapshot`, `browser_act`, and `browser_close` (all `exclusive: true`). Network policy defaults to require contained-proxy attestation; configure `uploads`/`downloads` for file transfer; `browser_act` adds `upload`/`screenshot`/`download_release`. Use `createBrowserManager().closeRun(runId)` / `close()` on terminal/abort. Align with a disposable sandbox via `createSharedSandboxBrowserOptions()` and `assertBrowserSandboxNetwork()`. CSS/XPath/evaluate/CDP/persistent profiles remain unsupported.
41
+
10
42
  ## 0.0.7 → 0.0.8 release overview
11
43
 
12
44
  All 31 first-party manifests and exact internal ranges move together to `0.0.8`; mixed first-party versions are unsupported. Core remains dependency-free at runtime and existing low-level agent/session APIs remain compatible. New telemetry, evaluation, MCP, A2A, ledger batching, and web research surfaces are opt-in. Release CI now requires CodeQL, dependency/license/SBOM/secret checks, packed-artifact attestations, PostgreSQL integration, and protected live-canary prerequisites; no tag or publication is automatic from this migration.
@@ -6,6 +6,21 @@ Evaluation defaults are finite: 100 trace rows × 20 pages and 4 MiB aggregate t
6
6
 
7
7
  This page states Prism runtime limits that keep slow consumers and long sessions from becoming unbounded memory or latency problems.
8
8
 
9
+ ## Release 0.0.9 reproducible coding/browser evidence
10
+
11
+ Run `node scripts/benchmark-0.0.9.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.9.test.mjs`. Default mode is network-free fake/in-process only and emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals for repository list/search, Git status, and browser open/snapshot/action/close. Optional `PRISM_BENCH_DOCKER=1` (with `PRISM_TEST_DOCKER_*`) and `PRISM_BENCH_PLAYWRIGHT=1` append real local Docker / protected Playwright rows. These are evidence fields, not CI timing gates.
12
+
13
+ 2026-07-21 baseline: Node v24.18.0, Linux x64, 100 iterations/scenario, network=false, credentials=false, docker=false, playwright=false.
14
+
15
+ | Scenario | mode | ops/s | p95 ms | heap bytes | disk bytes | processes | cost USD | backpressure | resource limits |
16
+ | --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
17
+ | repo-list | fake-in-process | 1,343 | 1.11 | 14,747,560 | 0 | 1 | 0 | 0 | 0 |
18
+ | repo-search | fake-in-process | 380 | 3.72 | 16,093,176 | 0 | 1 | 0 | 0 | 0 |
19
+ | git-status | fake-in-process | 479 | 2.62 | 13,821,688 | 0 | 1 | 0 | 0 | 0 |
20
+ | browser-open-snapshot-action-close | fake-in-process | 17,141 | 0.11 | 18,590,488 | 0 | 1 | 0 | 0 | 0 |
21
+
22
+ Rows exercise shipped repository/Git helpers and fake Playwright APIs only. Real Docker sandbox and Playwright browser timings remain explicit protected-gate evidence (`PRISM_TEST_DOCKER_SANDBOX=1`, `PRISM_LIVE_PLAYWRIGHT=1` / `PRISM_BENCH_DOCKER=1` / `PRISM_BENCH_PLAYWRIGHT=1`) because this release-candidate host did not enable those gates for the dated baseline. No live claim is inferred from skipped gates.
23
+
9
24
  ## Release 0.0.8 reproducible synthetic evidence
10
25
 
11
26
  Run `node scripts/benchmark-0.0.8.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000. Script uses no network/credentials and emits environment, throughput, p50/p95 latency, heap, synthetic disk bytes, zero external cost, and backpressure signals. These are evidence fields, not CI timing gates.
@@ -29,6 +44,16 @@ Security automation is isolated from `npm test`: CodeQL/supply-chain jobs have 1
29
44
 
30
45
  Web tools default/hard ceilings are query 4/16 KiB, results 10/20, URLs 5/20, request 256 KiB/1 MiB, response/aggregate 2/16 MiB, Markdown 1/8 MiB, extraction 256 KiB/1 MiB, schema 64/256 KiB, concurrency 4/16, retries 2/4, polling 20/100, and wall time 60 seconds/30 minutes. Bounds charge before request, retention, retry, or polling; overflow fails rather than truncating citation/extraction evidence.
31
46
 
47
+ Docker sandbox defaults/hard caps from `@arnilo/prism-coding-security`: startup 30 s/120 s; wall 20 min/30 min; idle 5 min/15 min; CPUs 2/8; memory 2 GiB/16 GiB (swap equal to memory); PIDs 256/1,024; FDs 1,024/8,192; workspace/tmp/download tmpfs 1 GiB/8 GiB, 256 MiB/2 GiB, 64 MiB/512 MiB; commands 100/256 with concurrent execs 1/8; env 64/256 names and 64 KiB/256 KiB values; export 50,000/250,000 entries and 256 MiB/2 GiB bytes with 16/64 retained artifacts; stop grace 5 s/30 s and cleanup 30 s/120 s. Caps validate before `docker create`/exec/export; overflow aborts and cleans the recorded container. Output still streams into the coding-agent `OutputAccumulator` ceilings (64 MiB/1 GiB).
48
+
49
+ Repository list/search defaults/hard caps from `@arnilo/prism-coding-agent`: depth 32/128; entries/files 10,000/100,000; page/results 1,000/10,000; search scan 64 MiB/1 GiB aggregate and 8 MiB/64 MiB per file; matches 1,000/10,000; pattern 512 B/4 KiB; line 50 KiB/1 MiB; context 5/20; wall 30 s/300 s; concurrency config 8/32. Walks stream via `opendir`/`lstat`, never follow symlink escapes, and stop immediately on aggregate limits or abort.
50
+
51
+ Structured Git/check/handoff defaults/hard caps: paths 1,000/10,000; refs 1 KiB/4 KiB; commit message 64 KiB/256 KiB; inline Git output 4 MiB/64 MiB; diff lines 10,000/100,000; changed files 1,000/10,000; patch input 16 MiB/64 MiB; worktrees 4/16; named checks 8/32 names, concurrency 1/4, timeout 10 min/60 min, diagnostic lines 2,000/100,000, output 4 MiB/64 MiB; PR handoff JSON 256 KiB/1 MiB with 100/1,000 commits. Git tools use typed argument arrays (never shell), disable hooks/credential prompts/external diff by default, and emit host-owned PR handoff data only — no push/network/PR client.
52
+
53
+ Durable coding plan/checkpoint defaults/hard caps: plan Markdown 256 KiB/1 MiB; todos 1,000/10,000 with 512 B/4 KiB text; checkpoint metadata 64 KiB/512 KiB; artifact references 16/64 at 256 MiB/2 GiB each; check summaries 1 KiB/8 KiB. Checkpoints store URI/hash/summaries/fingerprints only; resume revalidates workspace root, base branch, plan hash, and tool/policy/image fingerprints before import.
54
+
55
+ Browser automation defaults/hard caps from `@arnilo/prism-browser`: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2,000/10,000; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30 s/120 s; action 10 s/60 s; wait 30 s/120 s; run wall 20 min/30 min; popups 4/16; dialogs 16/64; listeners 64/256; action input 64 KiB/256 KiB; close grace 5 s/30 s; network requests 1,000/10,000 with 10/32 redirects per request and 8/32 WebSockets; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate. Caps charge before context/page/action/queue/snapshot/network/artifact retention. Host supplies Playwright and egress proxy attestation; package import launches nothing.
56
+
32
57
  Current surfaces:
33
58
 
34
59
  - `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
@@ -45,7 +45,7 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
45
45
  - `collectProviderEvents()` returns provider events in stream order.
46
46
  - `assertProviderStreamConforms()` returns collected events after verifying the stream ends with `done` or `error`, terminal events are last, and optional text/usage expectations match.
47
47
  - `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject. This is the supported timeout primitive; use a host abort controller or `RunOptions.signal` rather than deprecated provider-level `timeoutMs`.
48
- - `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. The runtime uses the same reconstruction behavior before tool execution when a provider streams deltas.
48
+ - `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. Malformed JSON with id+name present yields `argumentsError` (no throw); missing id/name throws typed `incomplete_delta`. The runtime uses the same reconstruction before tool execution when a provider streams deltas.
49
49
  - `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`. This is the provider-neutral check for normalized cache read/write token extraction; every first-party provider package exercises it against server-specific fields (`cached_tokens`, `cache_read_input_tokens`, etc.).
50
50
  - `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
51
51
  - `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.