@arnilo/prism 0.0.19 → 0.0.22

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem/repository tools as Prism `ToolDefinition` objects. It ships six default coding tools — `shell`, `read`, `write`, `edit`, `repo_list`, `repo_search` — plus opt-in structured Git/check set (`createGitTools`), opt-in `createAskUserDecisionTool({ ask })`, and bounded coding-plan/checkpoint helpers. The tools are **inert** until a host imports them and registers them into a `ToolRegistry`. Hosts may register any subset, omit aggregators entirely, or mix first-party tools with host-owned `ToolDefinition`s. Behavior for shell/read/write/edit is a behavioral port of the pi coding agent's tools, adapted to Prism's `ToolDefinition` / `ToolResult` contracts (no `@earendil-works/*` or `typebox` dependencies; only `diff` plus the Node standard library). List/search/Git are native Prism tools with no glob/ripgrep/Git-library dependency.
5
+ `@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem/repository tools as Prism `ToolDefinition` objects. It ships nine default coding tools — `shell`, `read`, `write`, `edit`, `repo_list`, `repo_search`, `glob`, `delete`, `move` — plus opt-in structured Git/check set (`createGitTools`), opt-in `createAskUserDecisionTool({ ask })`, and bounded coding-plan/checkpoint helpers. The tools are **inert** until a host imports them and registers them into a `ToolRegistry`. Hosts may register any subset, omit aggregators entirely, or mix first-party tools with host-owned `ToolDefinition`s. Behavior for shell/read/write/edit is a behavioral port of the pi coding agent's tools, adapted to Prism's `ToolDefinition` / `ToolResult` contracts (no `@earendil-works/*` or `typebox` dependencies; only `diff` plus the Node standard library). List/search/glob/Git are native Prism tools with no picomatch/ripgrep/Git-library dependency (hand-rolled `*`/`?`/`**` glob matcher).
6
6
 
7
7
  | Export | Purpose |
8
8
  | --- | --- |
@@ -11,14 +11,18 @@
11
11
  | `createWriteTool(cwd, options?)` | `write` tool: create or overwrite a file, creating parent directories. |
12
12
  | `createEditTool(cwd, options?)` | `edit` tool: precise exact-then-fuzzy text replacement in an existing file. |
13
13
  | `createRepoListTool(cwd, options?)` | `repo_list` tool: bounded deterministic repository listing. |
14
- | `createRepoSearchTool(cwd, options?)` | `repo_search` tool: bounded literal text search. |
15
- | `createCodingTools(cwd, options?)` | Default six tools (`shell`, `read`, `write`, `edit`, `repo_list`, `repo_search`). |
16
- | `createReadOnlyTools(cwd, options?)` | Read-only subset: `read`, `repo_list`, `repo_search`. |
14
+ | `createRepoSearchTool(cwd, options?)` | `repo_search` tool: bounded literal text search (`outputMode`: content / files_with_matches / count). |
15
+ | `createGlobTool(cwd, options?)` | `glob` tool: bounded filename-pattern match (`*` / `?` / `**`; no brace expansion). |
16
+ | `createDeleteTool(cwd, options?)` | `delete` tool: high-risk delete of a file or empty directory (no recursive delete, no trash). |
17
+ | `createMoveTool(cwd, options?)` | `move` tool: high-risk rename/move within the workspace (`overwrite` default false). |
18
+ | `createReadPathSet()` | Session-scoped path set for optional `requireReadBeforeWrite` soft guard. |
19
+ | `createCodingTools(cwd, options?)` | Default nine tools (`shell`, `read`, `write`, `edit`, `repo_list`, `repo_search`, `glob`, `delete`, `move`). |
20
+ | `createReadOnlyTools(cwd, options?)` | Read-only subset: `read`, `repo_list`, `repo_search`, `glob`. |
17
21
  | `createAllTools(cwd, options?)` | Identical to `createCodingTools` (Git tools remain opt-in via `createGitTools`). |
18
22
  | `createGitTools(cwd, options?)` | Opt-in Git tools (`git_status`/`git_diff`/`git_branch`/`git_worktree`/`git_apply`/`git_commit`/`git_pr_handoff`) plus optional `coding_check`. |
19
23
  | `createCodingCheckTool(cwd, options)` | Named host-declared checks; model selects only a name. |
20
24
  | `createAskUserDecisionTool(options)` | Opt-in user decision tool (`ask_user_decision`); host supplies `ask` callback. Not in default aggregators. |
21
- | `createLocalRepositoryOperations(limits?)` | Default streaming Node filesystem backend for list/search. |
25
+ | `createLocalRepositoryOperations(limits?)` | Default streaming Node filesystem backend for list/search/glob. |
22
26
  | `createGitOperations(options)` | Typed Git operations backend (argument arrays, safe config, finite output). |
23
27
  | `buildCodingCheckpointMetadata` / `validateCodingCheckpointMetadata` / `assertCodingResumeAllowed` | Bounded durable coding-task metadata for workflow `state.coding` (no second runtime). |
24
28
  | `writeCodingPlanFile` / `readCodingPlanFile` / `createCodingPlanMarkdown` / `parseCodingPlanTodos` | Workspace plan/todo Markdown helpers with finite byte/todo caps and hash verification. |
@@ -66,13 +70,39 @@ const tools = createCodingTools(workspaceRoot, {
66
70
  | `read` | `read` |
67
71
  | `write` | `write` |
68
72
  | `edit` | `edit` |
69
- | `repo_list` / `repo_search` | _(native; no pi equivalent)_ |
73
+ | `repo_list` / `repo_search` / `glob` | _(native; no pi equivalent)_ |
74
+ | `delete` / `move` | _(native; no pi equivalent)_ |
75
+
76
+ ### Tool selection guide
77
+
78
+ | Need | Prefer | Avoid |
79
+ | --- | --- | --- |
80
+ | Enumerate directories | `repo_list` | `shell` `find`/`ls` |
81
+ | Match filename patterns | `glob` | `shell` `find` |
82
+ | Find text in files | `repo_search` | `shell` `grep`/`rg` |
83
+ | Read one file (paged) | `read` | `shell` `cat` |
84
+ | Create / full overwrite | `write` | — |
85
+ | Targeted replace | `edit` | full `write` rewrite when a small edit works |
86
+ | Remove file / empty dir | `delete` | `shell` `rm` |
87
+ | Rename / relocate | `move` | `shell` `mv` |
88
+ | Arbitrary process | `shell` | dedicated tools above |
89
+
90
+ ### Phase 4 non-goals (0.0.21)
91
+
92
+ These are **out of scope** for this package release (see roadmap Phase 9 / later):
93
+
94
+ - **No PDF / document reader** — text and supported images only via `read`.
95
+ - **No trash / recycle daemon** — `delete` / `move` are permanent; host undo is not automatic.
96
+ - **No PTY / interactive process control** — `shell` is one-shot exec with bounded capture.
97
+ - **No LSP / language-server tools** — use host-owned tools if needed.
98
+ - **No recursive directory delete** — `delete` refuses non-empty directories.
99
+ - **No brace-expansion globs** — `glob` supports only `*`, `?`, and `**`.
70
100
 
71
101
  ## Inputs / request
72
102
 
73
103
  ### `shell`
74
104
 
75
- Run a shell command and return combined stdout+stderr.
105
+ Run a shell command and return combined stdout+stderr. Prefer dedicated coding tools (table above) when they fit.
76
106
 
77
107
  **Inputs:**
78
108
 
@@ -141,7 +171,7 @@ const read = createReadTool(cwd, {
141
171
 
142
172
  ### `write`
143
173
 
144
- Create or overwrite a file, creating parent directories as needed.
174
+ Create or **overwrite** a file (full replace), creating parent directories as needed. Prefer `edit` for targeted changes.
145
175
 
146
176
  **Inputs:**
147
177
 
@@ -149,6 +179,7 @@ Create or overwrite a file, creating parent directories as needed.
149
179
  | --- | --- | --- |
150
180
  | `path` | `string` | Path to the file to write (relative or absolute). Required. |
151
181
  | `content` | `string` | Content to write (empty string creates an empty file). Required. |
182
+ | `force` | `boolean` | Bypass optional read-before-write guard when the host enabled `requireReadBeforeWrite`. |
152
183
 
153
184
  **Outputs:** a `TextContent` confirmation naming the **absolute path** with UTF-8 byte and line counts (e.g. `Successfully wrote 42 bytes (3 lines) to /abs/path.txt`). `maxInputBytes` defaults to 8 MiB (64 MiB hard cap); oversized UTF-8 input fails before policy evaluation, directory creation, or write. Write failures and abort are error results. Empty `content` is valid.
154
185
 
@@ -156,6 +187,19 @@ Default local `writeFile` uses same-directory temp + `rename` so a crash mid-wri
156
187
 
157
188
  `write` result `metadata`: `{ bytes, lines, path }` (absolute path). Concurrent writes to the same path serialize through `withFileMutationQueue`; writes to different paths run in parallel.
158
189
 
190
+ ### Optional read-before-write guard
191
+
192
+ Hosts may opt in to a session-scoped soft guard: share one `createReadPathSet()` across `read` / `write` / `edit` and set `requireReadBeforeWrite: true` on write/edit options. Successful `read` marks the path; unread existing-file writes/edits fail with a clear error unless `force: true`. Default is **off** (no behavior change for hosts that ignore it).
193
+
194
+ ```ts
195
+ import { createReadPathSet, createReadTool, createWriteTool, createEditTool } from "@arnilo/prism-coding-agent";
196
+
197
+ const readPaths = createReadPathSet();
198
+ const read = createReadTool(cwd, { readPathSet: readPaths });
199
+ const write = createWriteTool(cwd, { requireReadBeforeWrite: true, readPathSet: readPaths });
200
+ const edit = createEditTool(cwd, { requireReadBeforeWrite: true, readPathSet: readPaths });
201
+ ```
202
+
159
203
  ### `edit`
160
204
 
161
205
  Precise text replacement in an existing file via exact-then-fuzzy matching.
@@ -166,8 +210,13 @@ Precise text replacement in an existing file via exact-then-fuzzy matching.
166
210
  | --- | --- | --- |
167
211
  | `path` | `string` | Path to the file to edit. Required. |
168
212
  | `edits` | `Array<{ oldText: string, newText: string }>` | Targeted replacements, each matched against the **original** file (not incrementally). No overlapping/nested edits. Required, non-empty. |
213
+ | `force` | `boolean` | Bypass optional read-before-write guard when enabled. |
214
+
215
+ Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse).
169
216
 
170
- Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse). **Fuzzy matching can apply a wrong region when `oldText` is slightly off** — tradeoff for imprecise models; prefer exact `oldText` when possible. A BOM is stripped before matching and re-prepended on write; original line endings are restored. Defaults reject targets over 8 MiB, aggregate old/new UTF-8 input over 2 MiB, or more than 100 edits (hard caps: 64 MiB, 16 MiB, and 1,000). Stat and bounded read checks run before matching or mutation. Default local `writeFile` uses same-directory temp + `rename` (crash-safe replace).
217
+ **Fuzzy silent-success tradeoff (loud):** when exact match fails, fuzzy may still apply a replacement **without warning the model**. That can edit the wrong region if `oldText` is slightly off (extra/missing whitespace, unicode lookalikes). Prefer exact `oldText` copied from a fresh `read`. Duplicate / non-unique matches already **fail closed** and leave the file unchanged — ambiguity is not silently resolved by picking the first hit.
218
+
219
+ A BOM is stripped before matching and re-prepended on write; original line endings are restored. Defaults reject targets over 8 MiB, aggregate old/new UTF-8 input over 2 MiB, or more than 100 edits (hard caps: 64 MiB, 16 MiB, and 1,000). Stat and bounded read checks run before matching or mutation. Default local `writeFile` uses same-directory temp + `rename` (crash-safe replace).
171
220
 
172
221
  **Outputs:** a `TextContent` confirmation (`Successfully replaced N block(s) in {path}.`) plus `metadata`. Any failure — missing/unreadable file, no match, duplicate (non-unique) match, overlap, empty `oldText`, no-op edit, or abort — is an error result, and the file is left **unchanged** (the match runs before the write).
173
222
 
@@ -175,7 +224,7 @@ Each `edits[].oldText` must match a unique, non-overlapping region of the origin
175
224
 
176
225
  ### `repo_list`
177
226
 
178
- List repository entries with deterministic relative paths. Uses Node `opendir`/`lstat` only — no glob dependency. Does not follow symlinks; rejects path escapes outside the workspace root. Hidden names and excluded basenames (default `.git`, `node_modules`, `dist`) are skipped unless `includeHidden` is set / host `exclude` is overridden.
227
+ List repository entries with deterministic relative paths. Uses Node `opendir`/`lstat` only — no glob dependency. Prefer `glob` when you already know a filename pattern. Prefer `repo_search` to find text inside files. Does not follow symlinks; rejects path escapes outside the workspace root. Hidden names and excluded basenames (default `.git`, `node_modules`, `dist`) are skipped unless `includeHidden` is set / host `exclude` is overridden.
179
228
 
180
229
  **Inputs:**
181
230
 
@@ -202,10 +251,59 @@ Search text files under the workspace using literal substring match. Binary file
202
251
  | `mode` | `"literal"` | Literal only (default). `regex` removed in 0.0.18. |
203
252
  | `caseSensitive` | `boolean` | Default false. |
204
253
  | `includeHidden` | `boolean` | Default false. |
205
- | `context` | `number` | Context lines before/after each match (default 5, hard 20). |
254
+ | `context` | `number` | Context lines before/after each match (default 5, hard 20). Ignored for non-content `outputMode`. |
206
255
  | `maxMatches` | `number` | Match cap (default 1,000, hard 10,000). |
256
+ | `outputMode` | `"content"` \| `"files_with_matches"` \| `"count"` | Result shape (default `content`). |
257
+
258
+ **Outputs:**
259
+ - `content` (default): ripgrep-like lines `path:line:column:text` with optional `path-` / `path+` context.
260
+ - `files_with_matches`: unique matching paths only.
261
+ - `count`: totals (`N matches in M files`) without line bodies.
262
+
263
+ Metadata includes `matches`, `truncated`, scan/skip counts; non-content modes also expose `fileCount`.
264
+
265
+ ### `glob`
266
+
267
+ Find workspace files by filename pattern without shell `find`. Hand-rolled matcher: `*` (one path segment), `?` (one char), `**` (directories). Brace expansion (`{a,b}`) is **rejected**. Patterns match workspace-relative full paths (e.g. `src/util/a.ts`). Returns **files only** (directories traversed but not listed). Same exclude/hidden/depth/page/time caps as `repo_list`.
207
268
 
208
- **Outputs:** ripgrep-like lines `path:line:column:text` with optional `path-` / `path+` context, plus metadata (`matches`, `truncated`, scan/skip counts).
269
+ **Inputs:**
270
+
271
+ | Field | Type | Purpose |
272
+ | --- | --- | --- |
273
+ | `pattern` | `string` | Glob pattern (required). |
274
+ | `path` | `string` | Workspace-relative start directory (default root). |
275
+ | `includeHidden` | `boolean` | Default false. |
276
+ | `maxDepth` | `number` | Depth cap (default 32, hard 128). |
277
+ | `maxResults` | `number` | Page size (default 1,000, hard 10,000). |
278
+ | `offset` | `number` | Matches to skip (default 0). |
279
+
280
+ **Outputs:** one relative path per line plus metadata (`truncated`, `truncatedBy`, `nextOffset`, scan counts). Continue with `offset=nextOffset` when truncated.
281
+
282
+ ### `delete`
283
+
284
+ High-risk: permanently delete a **single file or empty directory**. Non-empty directories fail closed (no recursive delete). Symlinks are unlinked as links (targets not followed for containment). **No trash daemon** — host undo is not automatic; gate with approval policy.
285
+
286
+ **Inputs:**
287
+
288
+ | Field | Type | Purpose |
289
+ | --- | --- | --- |
290
+ | `path` | `string` | File or empty directory to delete. Required. |
291
+
292
+ **Outputs:** confirmation with absolute path, or error (missing, non-empty dir, escape, abort).
293
+
294
+ ### `move`
295
+
296
+ High-risk: rename or move a file within the workspace. Dual-path mutation queue (lexicographic lock order). `overwrite` defaults **false**; when true, replaces an existing destination **file** only. Does not create parent directories. **No trash** — host undo is not automatic.
297
+
298
+ **Inputs:**
299
+
300
+ | Field | Type | Purpose |
301
+ | --- | --- | --- |
302
+ | `from` | `string` | Source path. Required. |
303
+ | `to` | `string` | Destination path. Required. |
304
+ | `overwrite` | `boolean` | Replace existing destination file (default false). |
305
+
306
+ **Outputs:** confirmation with absolute from/to, or error (missing source, dest exists without overwrite, escape, abort).
209
307
 
210
308
  ### Structured Git tools (`createGitTools`)
211
309
 
@@ -351,10 +449,10 @@ Minimal drop-in for any Prism app:
351
449
  import { createToolRegistry } from "@arnilo/prism";
352
450
  import { createCodingTools, createReadOnlyTools } from "@arnilo/prism-coding-agent";
353
451
 
354
- // Full coding set (shell + read + write + edit + repo_list + repo_search) against the project root:
452
+ // Full coding set (shell + read + write + edit + repo_list + repo_search + glob + delete + move):
355
453
  const tools = createToolRegistry(createCodingTools(process.cwd()));
356
454
 
357
- // Or a read-only set for inspection-only agents (read + repo_list + repo_search):
455
+ // Or a read-only set for inspection-only agents (read + repo_list + repo_search + glob):
358
456
  const ro = createToolRegistry(createReadOnlyTools(process.cwd()));
359
457
  ```
360
458
 
@@ -381,23 +479,27 @@ const remoteWrite = createWriteTool("/repo", {
381
479
  });
382
480
  ```
383
481
 
482
+ Packed capability demo: `examples/coding-tools-capability-gaps.ts` (search modes, glob, read-before-write, delete/move).
483
+
384
484
  ## Extension and configuration notes
385
485
 
386
486
  - **Long coding sessions.** Use `createCodingCompactionStrategy()` from optional `@arnilo/prism-compaction-llm` when history needs a bounded coding handoff. It is selected explicitly through normal `session.compact()` / agent compaction configuration, preserves raw session entries, and prioritizes file paths, patch intent, checks, plan/todo state, blockers, and verification steps. It does not read files, retain full diffs, or create a second coding runtime.
387
- - **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. Custom `RepositoryOperations` must honor depth/entry/file/match/scan/time caps and abort. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
388
- - **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`; list/search accept `repository` limits and shared aggregator `ToolsOptions.repository`.
389
- - **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit?, list?, search?, repository? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override. Read-only membership is deliberately `read` + `repo_list` + `repo_search` (0.0.9 behavior change).
390
- - **Sandbox composition.** Prefer `@arnilo/prism-coding-security` `createSandboxCodingComposition(cwd, { workspaceMode, sandbox, ... })` (or tools-only wrappers). `workspaceMode` is required: `"sandbox"` keeps shell/read/write/edit/list/search on one disposable tree; `"host"` runs against host cwd and never claims containment. Mixed sandbox-shell + host-FS wiring throws unless `allowMixedWorkspaceWiring: true`. Same-tree Git: `createGitTools(composition.workspaceRoot, { execFile: sandbox.execFile, commitIdentity })`.
487
+ - **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. Custom `RepositoryOperations` must honor depth/entry/file/match/scan/time caps and abort (including `glob`). Custom `DeleteOperations` / `MoveOperations` must honor containment and abort. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
488
+ - **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes` and optional `readPathSet`; `WriteToolOptions` / `EditToolOptions` add input caps plus optional `requireReadBeforeWrite` / `readPathSet` / `force`; list/search/glob accept `repository` limits and shared aggregator `ToolsOptions.repository`.
489
+ - **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit?, delete?, move?, list?, search?, glob?, repository? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override. Full membership is nine tools; read-only is `read` + `repo_list` + `repo_search` + `glob`.
490
+ - **Sandbox composition.** Prefer `@arnilo/prism-coding-security` `createSandboxCodingComposition(cwd, { workspaceMode, sandbox, ... })` (or tools-only wrappers). `workspaceMode` is required: `"sandbox"` keeps shell/read/write/edit/list/search/glob/delete/move on one disposable tree; `"host"` runs against host cwd and never claims containment. Mixed sandbox-shell + host-FS wiring throws unless `allowMixedWorkspaceWiring: true`. Same-tree Git: `createGitTools(composition.workspaceRoot, { execFile: sandbox.execFile, commitIdentity })`.
391
491
  - **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
392
492
  - No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
393
493
 
394
494
  ## Security and performance notes
395
495
 
396
- - **Host shell/filesystem access.** These tools run real commands and read/write/list/search real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
496
+ - **Host shell/filesystem access.** These tools run real commands and read/write/list/search/glob/delete/move real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
497
+ - **High-risk mutations.** `delete` and `move` are permanent (no trash). Prefer host confirmation via `ExecutionPolicy` before allowing them. Do not instruct models to bypass policy/sandbox.
397
498
  - **Non-zero exit is not an error.** A failing command is a normal `shell` result (exit code in metadata); only timeout/abort/spawn failures are error results. Do not assume `error == undefined` means the command succeeded.
398
- - **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `repo_list`/`repo_search` stream walks and charge depth/entry/file/match/scan/time before retention. Structured Git tools use argument arrays with finite output/path/ref/message/patch caps, disable hooks/credential prompts/external diff by default, and never push or open PRs. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
399
- - **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
499
+ - **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `repo_list`/`repo_search`/`glob` stream walks and charge depth/entry/file/match/scan/time before retention. Structured Git tools use argument arrays with finite output/path/ref/message/patch caps, disable hooks/credential prompts/external diff by default, and never push or open PRs. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
500
+ - **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. `move` locks both paths in lexicographic order. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
400
501
  - **Bounded image reads.** `read` rejects images over `maxImageBytes` (default 10 MB) by `stat` before read when possible; MIME is detected from magic bytes only. Optional `transformImage` is host-owned — the base package has no image-processing dependency.
502
+ - **Fuzzy edit risk.** Silent fuzzy success can mis-apply edits; duplicate matches fail closed. See the `edit` section above.
401
503
 
402
504
  ### Resource-limit defaults and hard caps
403
505
 
@@ -26,14 +26,14 @@ import type { ExecutionAction, ExecutionPolicy, ExecutionDecision } from "@arnil
26
26
 
27
27
  Use this package when coding tools need path scoping, human approval, command rules, or a pluggable sandbox backend. Wire the returned policy through `createCodingTools(cwd, { executionPolicy })` or per-tool `executionPolicy` options.
28
28
 
29
- Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
29
+ Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit/delete/move without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
30
30
 
31
31
  ## Inputs / request
32
32
 
33
33
  | Option | Default | Purpose |
34
34
  | --- | --- | --- |
35
35
  | `roots` | required | Realpath-contained filesystem roots. |
36
- | `readOnly` | `false` | Deny shell/write/edit actions. |
36
+ | `readOnly` | `false` | Deny non-`read` actions (including shell/write/edit/delete/move). |
37
37
  | `commandRules` | `[]` | Ordered allow/deny/approval command classification. |
38
38
  | `approve` | none | Host callback for actions not statically allowed; omission fails closed. |
39
39
  | `approvalCacheScope` | `"none"` | Optional `run` or `session` decision cache scope. |
@@ -110,7 +110,7 @@ const sandbox = await createDockerSandbox({
110
110
  limits: { cpus: 2, memoryBytes: 2 * 1024 ** 3, maxPids: 256, workspaceBytes: 1024 ** 3 },
111
111
  });
112
112
 
113
- // Sandbox mode: shell/read/write/edit/list/search share one disposable tree.
113
+ // Sandbox mode: shell/read/write/edit/list/search/glob/delete/move share one disposable tree.
114
114
  const { tools, composition } = createSandboxCodingComposition("/srv/jobs/task-1/source", {
115
115
  workspaceMode: "sandbox",
116
116
  sandbox,
@@ -6,7 +6,7 @@
6
6
 
7
7
  ## When to use it
8
8
 
9
- Use context resolution when a host wants project/session/context blocks resolved before prompt composition. Use the skill registry when a host wants explicit progressive skill disclosure. Declarative `AgentDefinition.skills` are inactive unless listed; omitted skills means none unless the host uses the migration-only `activateAllCapabilities: true` option.
9
+ Use context resolution when a host wants project/session/context blocks resolved before prompt composition. Use the skill registry when a host wants explicit progressive skill disclosure: catalog `name` + `description` every turn by default, full `instructions` only after `load_skill` or when the host opts into eager mode. Declarative `AgentDefinition.skills` are inactive unless listed; omitted skills means none unless the host uses the migration-only `activateAllCapabilities: true` option.
10
10
 
11
11
  Do not use these helpers as an agent loop, package discovery mechanism, context cache, token budgeter, retrier, credential resolver, semantic skill ranker, tool activator, or permission system.
12
12
 
@@ -107,7 +107,8 @@ The agent/session runtime resolves skills per run and wires each active skill's
107
107
  | Surface | Config shape | Run override | Active skills |
108
108
  | --- | --- | --- | --- |
109
109
  | Runtime agent | `AgentConfig.skills: SkillRegistry` | `RunOptions.activeSkills: ["brief"]` | Named skills only, resolved with `resolveActiveSkills({ registry, names, tools })`. |
110
- | Runtime agent | `AgentConfig.skills: SkillRegistry` | no `activeSkills` / no `skills` | All registry skills (`SkillRegistry.list()`). |
110
+ | Runtime agent | `AgentConfig.skills: SkillRegistry` | no `activeSkills` / no `skills` / no `activateAllSkills` | No skills active (fail-closed default). |
111
+ | Runtime agent | `AgentConfig.skills: SkillRegistry` | `activateAllSkills: true` (run or agent) | All registry skills (`SkillRegistry.list()`), migration opt-in. |
111
112
  | Runtime agent | `AgentConfig.skills: Skill[]` | `RunOptions.skills: [...]` | Override array only. |
112
113
  | Runtime agent | `AgentConfig.skills: Skill[]` | no `RunOptions.skills` | All configured array skills. |
113
114
  | Declarative definition | `AgentDefinition.skills: ["brief"]` | later runtime `activeSkills` optional | Listed names only. |
@@ -118,13 +119,13 @@ Runtime selection precedence mirrors the other `RunOptions` overrides (`redactor
118
119
 
119
120
  1. `AgentConfig.skills` is a `SkillRegistry` and `RunOptions.activeSkills: readonly string[]` (names) is set → the runtime calls `resolveActiveSkills({ registry, names, tools })`.
120
121
  2. `RunOptions.skills: readonly Skill[]` is set → that array replaces `AgentConfig.skills` for the run. This override exists for the case where `AgentConfig.skills` is a plain `Skill[]` (no registry), so name resolution is impossible.
121
- 3. Neither set → all configured runtime skills are active (current behavior; `SkillRegistry.list()` or the plain array as-is). This is not the declarative default.
122
+ 3. Neither set → no skills active when `AgentConfig.skills` is a `SkillRegistry` (fail-closed). Use `activateAllSkills: true` on the run or agent to restore prior list-all behavior (`SkillRegistry.list()`). Plain `Skill[]` configs still activate every configured array skill. This is not the declarative default.
122
123
 
123
124
  names win when a registry exists. `RunOptions.activeSkills` cannot be used against a plain-array `AgentConfig.skills` — use `RunOptions.skills` instead. Use `RunOptions.skills: []` for an explicit no-skills runtime run.
124
125
 
125
- Each active skill contributes two things the runtime now wires together:
126
+ Each active skill contributes two things the runtime wires together:
126
127
 
127
- - `Skill.instructions` → rendered as system messages by `skillMessages()` (active set only).
128
+ - `Skill` prompt text → rendered as system messages by `skillMessages()` / `skillPromptText()` (active set only). Default `skillsDisclosure: "progressive"` sends `Skill <name>: <description>`; full `instructions` appear only when the skill is in the session `LoadedSkillSet` or disclosure is `"eager"`.
128
129
  - `Skill.context: ContextProvider[]` → collected across active skills (`activeSkills.flatMap(s => s.context ?? [])`), resolved through the existing `resolveContextProviders(...)`, and merged into the request's `context` **after** host `AgentConfig.context` blocks. Inactive skills contribute neither instructions nor context.
129
130
 
130
131
  `toolNames` enforcement is live: because selection routes through `resolveActiveSkills()`, a skill demanding a host-inactive tool throws with `Skill ${name} requires inactive tool: ${missing}` **before the first provider turn** — no provider call, no store write, no partial side effect. This is the fail-fast contract the docs already claimed; the runtime now honors it.
@@ -153,7 +154,80 @@ await session.run(input, { skills: [{ name: "verbose", instructions: "Be verbose
153
154
  await session.run(input, { skills: [] }); // explicit no skills for this run
154
155
  ```
155
156
 
156
- Skill selection grants no tool access and cannot bypass permissions — a skill's `toolNames` can only *require* host-active tools, never activate or grant them. Declarative skills also do not activate themselves by presence in a registry; list names on `AgentDefinition.skills` (or pass runtime `activeSkills`) when wanted. Per-skill token budgeting is deferred; the merge order (host context, then skill context) is the only priority knob today.
157
+ Skill selection grants no tool access and cannot bypass permissions — a skill's `toolNames` can only *require* host-active tools, never activate or grant them. Declarative skills also do not activate themselves by presence in a registry; list names on `AgentDefinition.skills` (or pass runtime `activeSkills`) when wanted.
158
+
159
+ ### Progressive skill disclosure
160
+
161
+ `skillsDisclosure` on `AgentConfig` / `RunOptions` (`"progressive"` default, `"eager"` opt-in; run wins) controls how active skills render in provider input:
162
+
163
+ | Mode | Provider view per active skill |
164
+ | --- | --- |
165
+ | `"progressive"` (default) | `Skill <name>: <description>` (or `(no description)` when empty) |
166
+ | `"eager"` | `Skill <name>:\n<instructions>` every turn (pre-0.0.20 behavior) |
167
+
168
+ Catalog caps: **64** entries default / **256** hard; descriptions **512 B** default / **4 KiB** hard; instruction bodies **32 KiB** default / **256 KiB** hard on load/eager render. Oversize catalog/description/instruction payloads fail closed (`SkillDisclosureError` / `SkillLoadError`).
169
+
170
+ ```ts
171
+ import { assembleProviderInput, createLoadedSkillSet } from "@arnilo/prism";
172
+
173
+ const loaded = createLoadedSkillSet(); // session-owned; not checkpoint-persisted in 0.0.20
174
+ const request = await assembleProviderInput({
175
+ model,
176
+ input: "Hi",
177
+ skills: active,
178
+ skillsDisclosure: "progressive",
179
+ loadedSkills: loaded,
180
+ });
181
+ ```
182
+
183
+ ### On-demand skill load (`load_skill`)
184
+
185
+ Hosts opt in by registering `createLoadSkillTool({ registry, loaded })` on the active tool set. The model calls `load_skill { name }` with an exact registry name; success adds the name to the session `LoadedSkillSet` so later turns include `instructions` under progressive mode. The tool does **not** activate tools, widen permissions, or load skills that were not active for the run.
186
+
187
+ Fail-closed cases: unknown name, inactive skill for the run, inactive required `toolNames`, oversize body, duplicate load, missing session loaded-set wiring. Tool output and errors are size-capped; skill text is untrusted host/extension data.
188
+
189
+ ```ts
190
+ import { createAgent, createLoadSkillTool, createSkillRegistry } from "@arnilo/prism";
191
+
192
+ const registry = createSkillRegistry([ponytail, brief]);
193
+ const loadSkill = createLoadSkillTool({ registry }); // session injects loadedSkills at dispatch
194
+ const agent = createAgent({ model, provider, skills: registry, tools: [loadSkill, /* host */] });
195
+ await agent.createSession().run("…", { activeSkills: ["ponytail"] });
196
+ // Turn 1: catalog only. After load_skill({ name: "ponytail" }), later turns include instructions.
197
+ ```
198
+
199
+ ### Third-party behavior packages (Caveman, Ponytail)
200
+
201
+ `@arnilo/prism-caveman` and `@arnilo/prism-ponytail` register upstream skills into the extension kernel skill registry. Hosts should:
202
+
203
+ 1. `kernel.load([createCavemanExtension(...), createPonytailExtension(...)])` with session `appendEntry` / `getEntries` callbacks.
204
+ 2. Build `createSkillRegistry(kernel.registries.skills.list())` and pass `activeSkills` / `resolveActiveSkills` names.
205
+ 3. Keep `skillsDisclosure: "progressive"` and register `createLoadSkillTool` — full `SKILL.md` bodies stay catalog-only until `load_skill`.
206
+ 4. Select `instructionInjectors: ["caveman-mode", "ponytail-mode"]` (or subset) for mode/level slices **without** forcing `skillsDisclosure: "eager"`.
207
+
208
+ Mode slices and skill bodies are independent: the injector can add `PONYTAIL MODE ACTIVE` while `ponytail-audit` remains catalog-only until loaded. See [Caveman](caveman.md), [Ponytail](ponytail.md), and `examples/caveman-ponytail.ts`.
209
+
210
+ Pure validation without the tool: `resolveSkillLoad({ registry, name, tools, loaded, activeSkillNames })`.
211
+
212
+ ### Context budget priority and skill demotion
213
+
214
+ When `assembleProviderInput` runs with `contextBudget`, `applyContextBudget` evicts droppable sections in layout order. Within `context` blocks and skills, victims sort by ascending `ContextBlock.priority` (missing = **0**), then LIFO within the same priority.
215
+
216
+ Under pressure on a skill with a loaded body, eviction may demote to catalog-only first (`ContextBudgetOmissionKind: "skill_body"`), then remove the skill entirely (`"skills"`). Demoted bodies render as description-only even when the name remains in `LoadedSkillSet`. See [Input and prompt assembly](input-and-prompt-assembly.md).
217
+
218
+ ### Optional tool-result fold
219
+
220
+ `toolResultFold` on `AgentConfig` / `RunOptions` (run wins) is **off** unless the host supplies a `summarize` callback. When enabled, aged large tool-result messages in the **provider view** become a one-line header plus bounded summary text; session store entries stay raw. Defaults: `minAgeTurns` **2**, `minBytes` **4096**, `maxSummaryBytes` **512** (hard **4096**). Summarizer failure keeps the raw tool result (fail closed). Not a second memory system — use observational memory / compaction for durable recall.
221
+
222
+ ```ts
223
+ await session.run("…", {
224
+ toolResultFold: {
225
+ minAgeTurns: 2,
226
+ minBytes: 4_096,
227
+ summarize: async ({ toolCallId, text }) => `ref:${toolCallId} ${text.slice(0, 80)}`,
228
+ },
229
+ });
230
+ ```
157
231
 
158
232
  ### Migration note
159
233
 
@@ -164,12 +238,25 @@ For declarative agents, old configs that omitted `skills` should now add explici
164
238
  resolveAgentDefinition({ name: "doc", model, skills: ["brief"] }, context);
165
239
  ```
166
240
 
167
- Use `activateAllCapabilities: true` only as a temporary all-skills/all-tools compatibility opt-in during migration. Runtime `RunOptions.activeSkills` remains the per-run narrowing tool after an agent has a skill registry configured.
241
+ Runtime hosts that relied on `SkillRegistry.list()` when `activeSkills` was omitted must opt in explicitly:
242
+
243
+ ```ts
244
+ // Restore pre-0.0.20 list-all activation (still subject to progressive disclosure):
245
+ await session.run("Hi", { activateAllSkills: true });
246
+
247
+ // Or restore full instruction bodies every turn:
248
+ const agent = createAgent({ model, provider, skills: registry, skillsDisclosure: "eager" });
249
+ ```
250
+
251
+ Use `activateAllCapabilities: true` only as a temporary all-skills/all-tools compatibility opt-in during migration for **declarative** definitions. Runtime `RunOptions.activeSkills` remains the per-run narrowing tool after an agent has a skill registry configured.
168
252
 
169
253
  ## Security and performance notes
170
254
 
171
255
  - Context providers run sequentially and deterministically in caller order.
172
256
  - Skill registry lookup is `Map`-backed, and selection is linear in requested skills plus active tools. Strict duplicate mode adds one O(1) `Map.has()` check during registration only.
257
+ - Progressive catalog render is O(active skills) with byte/count caps; `load_skill` lookup is O(1). Budget eviction over context/skills is O(n log n) worst case.
258
+ - `load_skill` cannot grant tools; loaded instructions are untrusted text bounded by hard caps. `toolResultFold` summarizer output is untrusted and capped; failures keep raw tool results.
259
+ - Loaded-skill names are session-scoped in memory only in 0.0.20 — not checkpoint-persisted; new sessions start catalog-only until reload.
173
260
  - These helpers perform no provider calls, tool execution, resource loading, package discovery, filesystem/network access, retries, timers, or watchers by themselves.
174
261
  - Context and skill output is host/extension data. Do not include secrets unless the host explicitly accepts that prompt exposure.
175
262
  - Active tools remain host-supplied; skills and middleware do not activate tools or grant permissions. Use `duplicate: "error"` when loading third-party skills to prevent silent name shadowing.
@@ -140,6 +140,8 @@ await kernel.middleware.run("provider_request", { metadata: {} });
140
140
  - [Compaction and retry policies](compaction-and-retry.md): compaction strategy/retry policy contributions and `compaction`/`retry` middleware runtime behavior.
141
141
  - [LLM compaction package](compaction-llm.md): optional extension helper that registers a provider-backed compaction strategy.
142
142
  - [Observational memory compaction package](compaction-observational-memory.md): optional extension helper that registers an inert fast memory compaction strategy.
143
+ - [Caveman behavior integration](caveman.md): optional `@arnilo/prism-caveman` upstream Caveman skills, commands, level injector, and session `caveman-level` persistence.
144
+ - [Ponytail behavior integration](ponytail.md): optional `@arnilo/prism-ponytail` upstream Ponytail skills, commands, mode injector, and session `ponytail-mode` persistence.
143
145
  - [Public contracts](public-contracts.md): `Extension`, `ExtensionAPI`, and contribution contract types.
144
146
  - [Credentials and redaction](credentials-and-redaction.md): secret-redaction behavior used for extension errors.
145
147
 
package/docs/index.md CHANGED
@@ -34,7 +34,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
34
34
  - [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention/legal-hold/quota lifecycle (`lifecycle`), and NoSQL mapping.
35
35
  - [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, FTS `searchSessions` (migration-v4), and transactionally verified/backfilled migration metadata.
36
36
  - [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, FTS `searchSessions` (migration-v4), advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
37
- - [Migration guide](migration.md): **0.0.15** OpenAI hosted tools/continuation/Realtime, exact AI SDK v4 matrix, RAG lifecycle/reranking/trust/status, and memory export/rebuild; **0.0.14** conversations, memory consent/lifecycle, artifact co-work review, AG-UI co-work events, scoped M365/GWS OAuth connectors, browser checkpoints, device contracts, and Alibaba/Ollama providers; plus prior release migrations.
37
+ - [Migration guide](migration.md): **0.0.22** third-party behavior integrations (Caveman, Ponytail); **0.0.21** coding-tool capability gaps (`outputMode`, glob, read-before-write, delete/move, aggregator 9/4); **0.0.20** progressive skill disclosure, empty registry default, `load_skill`, priority budget demotion, optional tool-result fold; **0.0.19** observational memory lifecycle; **0.0.15** OpenAI hosted tools/continuation/Realtime, exact AI SDK v4 matrix, RAG lifecycle/reranking/trust/status, and memory export/rebuild; **0.0.14** conversations, memory consent/lifecycle, artifact co-work review, AG-UI co-work events, scoped M365/GWS OAuth connectors, browser checkpoints, device contracts, and Alibaba/Ollama providers; plus prior release migrations.
38
38
  - [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety; `searchSessions` throws `SessionSearchUnsupportedError`.
39
39
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
40
40
 
@@ -58,7 +58,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
58
58
  - [Multimodal content](multimodal-content.md): complete-request media resolution and aggregate bounds, DNS-classified/address-pinned URLs, SSRF/MIME policy, `ModelCapabilities.input` tags, and first-party content-type mapping.
59
59
  - [System prompts](system-prompts.md): compose explicit user/package/app/run system prompt layers, auto-load the standard `AGENTS.md` (workspace) / `SYSTEM.md` prompt files via the Node `loadSystemPromptFiles` loader (trust-gated for `AGENTS.md`), and append `SYSTEM.md` → per-agent `AGENT.md` body → repo `AGENTS.md` layers from a discovered agent bundle via `resolveAgentBundle`.
60
60
  - [Instruction injection](instruction-injection.md): register package injectors that layer redacted instructions/context blocks without granting tools, permissions, or resource escapes.
61
- - [Context and skills](context-and-skills.md): resolve ordered context providers and keep context/skill selection host-owned; omitted declarative skills stay inactive by default, `toolNames` fail closed before provider turns, and strict skill registries prevent silent shadowing.
61
+ - [Context and skills](context-and-skills.md): resolve ordered context providers; progressive skill catalog (`skillsDisclosure`, default catalog-only), `load_skill` on-demand bodies, fail-closed registry activation (`activateAllSkills` migration opt-in), `toolNames` fail closed before provider turns, priority-aware budget demotion, and optional `toolResultFold`.
62
62
  - [Retrieval-augmented generation](rag.md): optional bounded source lifecycle, document adapters, host reranking, ingestion status, attributable citations, and inert context injection.
63
63
 
64
64
  ## Tools
@@ -71,7 +71,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
71
71
  - [Work connectors](work-connectors.md): connector principles, capability gates, scoped OAuth establishment (0.0.14), and out-of-scope boundaries (Slack/Teams channels not shipped) for Microsoft 365 / Google Workspace.
72
72
  - [Browser automation](browser-automation.md): optional `@arnilo/prism-browser` with host-supplied Playwright contexts, AI-mode snapshots/refs, ordered `browser_open`/`browser_snapshot`/`browser_act`/`browser_close`, egress/side-effect/upload/download/screenshot policy, finite page/action/snapshot/network/artifact caps, and 0.0.14 verified-state checkpoints with reload/verify-before-side-effect.
73
73
  - [Device adapters](device-adapters.md): deny-by-default realtime voice / desktop-control contract + conformance (0.0.14); no vendor package — admission fails closed without explicit consent+sandbox+approval, stream bounds, shared `RunLimits`, redacted telemetry.
74
- - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, `edit`, `repo_list`, and `repo_search` definitions plus opt-in `createGitTools()` / `coding_check`, opt-in `createAskUserDecisionTool` (single/multi/free-text + durable suspend glue), and `runCodingGoalVerify`; durable plan/todo Markdown helpers with workflow `state.coding` checkpoint metadata; streamed text pages, bounded repository list/search, finite Git/check/plan/ask caps, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
74
+ - [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, `edit`, `repo_list`, `repo_search`, `glob`, `delete`, and `move` definitions plus opt-in `createGitTools()` / `coding_check`, opt-in `createAskUserDecisionTool` (single/multi/free-text + durable suspend glue), and `runCodingGoalVerify`; durable plan/todo Markdown helpers with workflow `state.coding` checkpoint metadata; streamed text pages, `repo_search` `outputMode`, bounded glob, optional read-before-write, finite Git/check/plan/ask caps, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. No PDF/trash/PTY/LSP in 0.0.21. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
75
75
  - [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, required `workspaceMode` (`host`/`sandbox`) with fail-closed mixed wiring, `createSandboxCodingComposition()` containment metadata, and the disposable Docker/OCI sandbox reference with bounded workspace import/export.
76
76
 
77
77
  ## Extensions/plugins
@@ -114,11 +114,15 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
114
114
  - [Compaction conformance](compaction-conformance.md): assert any `CompactionStrategy` returns a non-empty redacted summary and observes abort from `@arnilo/prism/testing/compaction-conformance`.
115
115
  - [Tool conformance](tool-conformance.md): assert the tool-dispatch blocked-reason matrix (unknown/denied/invalid/permission/validator) and success path from `@arnilo/prism/testing/tool-conformance`.
116
116
  - [Extension conformance](extension-conformance.md): assert an `Extension` setup runs, contributions stay inert, and setup errors are redacted or rethrown from `@arnilo/prism/testing/extension-conformance`.
117
- - `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts), [`examples/enterprise-identity.ts`](../examples/enterprise-identity.ts), [`examples/enterprise-policy-audit.ts`](../examples/enterprise-policy-audit.ts), [`examples/enterprise-work-connectors.ts`](../examples/enterprise-work-connectors.ts), [`examples/conversation-durable-replay.ts`](../examples/conversation-durable-replay.ts), [`examples/artifact-review-delivery.ts`](../examples/artifact-review-delivery.ts), [`examples/server-deployment-seams.ts`](../examples/server-deployment-seams.ts), cache-aware prompt assembly, NeuralWatt agent run ([`examples/neuralwatt-agent-run.ts`](../examples/neuralwatt-agent-run.ts)), [`examples/coding-compaction.ts`](../examples/coding-compaction.ts), stores/branching, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
117
+ - `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts), [`examples/enterprise-identity.ts`](../examples/enterprise-identity.ts), [`examples/enterprise-policy-audit.ts`](../examples/enterprise-policy-audit.ts), [`examples/enterprise-work-connectors.ts`](../examples/enterprise-work-connectors.ts), [`examples/conversation-durable-replay.ts`](../examples/conversation-durable-replay.ts), [`examples/artifact-review-delivery.ts`](../examples/artifact-review-delivery.ts), [`examples/server-deployment-seams.ts`](../examples/server-deployment-seams.ts), cache-aware prompt assembly, NeuralWatt agent run ([`examples/neuralwatt-agent-run.ts`](../examples/neuralwatt-agent-run.ts)), [`examples/coding-compaction.ts`](../examples/coding-compaction.ts), [`examples/caveman-ponytail.ts`](../examples/caveman-ponytail.ts), stores/branching, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
118
+
119
+ ## Third-party integrations
120
+ - [Caveman behavior integration](caveman.md): optional `@arnilo/prism-caveman` — upstream Caveman skills/commands, `caveman-mode` injector, session `caveman-level` persistence, progressive catalog + `load_skill`; requires host `upstreamPath` and session attach callbacks; inert until `kernel.load`.
121
+ - [Ponytail behavior integration](ponytail.md): optional `@arnilo/prism-ponytail` — upstream Ponytail skills/commands, `ponytail-mode` injector, session `ponytail-mode` persistence; resolves peer `@dietrichgebert/ponytail` or `upstreamPath`; opt-in (not in code/sdk profiles).
118
122
 
119
123
  ## Release and install
120
- - [Release and install](release-and-install.md): current **0.0.19** 44-package graph (Phase 2 observational memory lifecycle; plan 002), exact-peer/install/tarball rules, deterministic resumable publication and publish dry-run, pinned supply-chain gates, offline tests, the 0.0.15 provider/AI-SDK/RAG/memory protected live-canary matrix, and sandbox-browser Docker/Playwright gates.
121
- - [0.1.0 / 1.0 readiness gates](0.1.0-readiness.md): command-per-gate 1.0 readiness table — frozen API surface + compat gate, migration/docs tripwires, budget table, live-suite matrix, security matrix, current-line status (**0.0.19** published target), signed-publication/live-canary prerequisites for 1.0, and Phase 12 demand-evidence entry criteria.
124
+ - [Release and install](release-and-install.md): current **0.0.22** 46-package graph (Phase 5 Caveman/Ponytail behavior integrations; plan 005), exact-peer/install/tarball rules, deterministic resumable publication and publish dry-run, pinned supply-chain gates, offline tests, the 0.0.15 provider/AI-SDK/RAG/memory protected live-canary matrix, and sandbox-browser Docker/Playwright gates.
125
+ - [0.1.0 / 1.0 readiness gates](0.1.0-readiness.md): command-per-gate 1.0 readiness table — frozen API surface + compat gate, migration/docs tripwires, budget table, live-suite matrix, security matrix, current-line status (**0.0.22** published target), signed-publication/live-canary prerequisites for 1.0, and Phase 12 demand-evidence entry criteria.
122
126
  - [Review coverage (2026-07-26 Phase 11)](review-coverage-2026-07-26-phase-11.md): Plan 079 evidence freeze — baseline size/startup/benchmark budgets, hotspot domain extraction table, confirmed duplication survivors (redactor/cleanJson/row-codecs/checkpoints/exec-runner/approval/ownership), profile adoption recommendations, and tarball artifact-diet findings for 0.0.16.
123
127
  - [Review coverage (2026-07-26 Phase 10)](review-coverage-2026-07-26-phase-10.md): Plan 078 evidence freeze — OpenAI hosted tools/continuation/realtime, AI SDK version matrix, remaining provider metadata parity, RAG replaceSource/loaders/parsers/reranker/provenance/ingestion-status, memory export/rebuild/conformance, and 0.0.15 (43 → 43 manifests; no new package) release gates.
124
128
  - [Review coverage (2026-07-25 Phase 9)](review-coverage-2026-07-25-phase-9.md): Plan 077 evidence freeze — conversation service, memory consent/lifecycle, artifact co-work review, AG-UI co-work events, scoped M365/GWS OAuth, browser checkpoint composition, and deny-by-default device contracts for 0.0.14 (41 → 43 manifests; only the two provider packages are new).
package/docs/migration.md CHANGED
@@ -1,5 +1,76 @@
1
1
  # Migration guide
2
2
 
3
+ ## 0.0.21 → 0.0.22 third-party behavior integrations (additive)
4
+
5
+ Release **0.0.22** adds two optional behavior packages; core `@arnilo/prism` runtime behavior is unchanged.
6
+
7
+ 1. **New packages (opt-in).** `@arnilo/prism-caveman` and `@arnilo/prism-ponytail` wire upstream Caveman and Ponytail into Prism extension contracts. They are **not** included in `@arnilo/prism-code`, `@arnilo/prism-sdk`, or `@arnilo/prism-all` by default — install explicitly when needed.
8
+ 2. **Inert until loaded.** Import registers nothing. Host calls `createExtensionKernel().load([createCavemanExtension(...)])` / `createPonytailExtension(...)`.
9
+ 3. **Session attach required.** Both factories require host `appendEntry` and `getEntries` callbacks (same pattern as observational memory `attach`) for mode/level persistence (`caveman-level`, `ponytail-mode` custom entries).
10
+ 4. **Progressive disclosure.** Keep `skillsDisclosure: "progressive"` and register `createLoadSkillTool`; mode/level slices come from `caveman-mode` / `ponytail-mode` instruction injectors, not eager full `SKILL.md` bodies.
11
+ 5. **Upstream resolution.** Caveman requires `upstreamPath` to a [juliusbrussee/caveman](https://github.com/juliusbrussee/caveman) checkout (`skills/` marker). Ponytail resolves optional peer `@dietrichgebert/ponytail@^4.8.4` or `upstreamPath`. Missing upstream → `setup` throws; zero contributions registered.
12
+ 6. **Publish graph.** Publishable manifest count is **46** (was 44).
13
+
14
+ Example: `node examples/caveman-ponytail.ts` (network-free fixture upstream trees).
15
+
16
+ ```ts
17
+ import { createCavemanExtension } from "@arnilo/prism-caveman";
18
+ import { createPonytailExtension } from "@arnilo/prism-ponytail";
19
+
20
+ await kernel.load([
21
+ createCavemanExtension({ upstreamPath: "/path/to/caveman", appendEntry, getEntries }),
22
+ createPonytailExtension({ defaultMode: "full", appendEntry, getEntries }),
23
+ ]);
24
+ ```
25
+
26
+ No breaking changes for hosts that do not install the new packages.
27
+
28
+ ## 0.0.20 → 0.0.21 coding-tool capability gaps (small intentional breaks)
29
+
30
+ Release **0.0.21** completes Phase 4 coding-tool capability gaps in `@arnilo/prism-coding-agent` / `@arnilo/prism-coding-security`:
31
+
32
+ 1. **`repo_search` gains `outputMode`.** Optional `outputMode?: "content" | "files_with_matches" | "count"` (default `"content"`). Files-only and count modes omit match body text from model content; invalid values fail closed.
33
+ 2. **Bounded `glob` tool.** `createGlobTool` / aggregator membership; `*` / `?` / `**` only (no brace expansion); reuses repository walk limits; files only.
34
+ 3. **Optional read-before-write.** Host sets `requireReadBeforeWrite: true` with a shared `ReadPathSet` on read/write/edit; unread paths fail unless `force: true`. In-memory / session-scoped only — not checkpoint-persisted.
35
+ 4. **Bounded `delete` and `move`.** File or empty directory delete (no recursive); move with `overwrite` default `false`; high-risk `ExecutionPolicy` kinds; host undo is not automatic.
36
+ 5. **Aggregator membership.** `createCodingTools` → **9** tools (adds `glob`, `delete`, `move`); `createReadOnlyTools` → **4** (adds `glob`). Hosts asserting exact `.length` must update.
37
+ 6. **Approval / sandbox.** `isMutatingKind` includes `delete` and `move` (not `glob`). Full sandbox custom ops must supply delete/move backends; `RepositoryOperations` requires `glob`.
38
+
39
+ Example: `node examples/coding-tools-capability-gaps.ts` (network-free).
40
+
41
+ ```ts
42
+ const tools = createCodingTools(cwd); // length 9
43
+ const search = createRepoSearchTool(cwd);
44
+ await search.execute({ query: "TODO", outputMode: "files_with_matches" }, ctx);
45
+
46
+ const readPathSet = createReadPathSet();
47
+ const write = createWriteTool(cwd, { requireReadBeforeWrite: true, readPathSet });
48
+ ```
49
+
50
+ Fuzzy edit may still succeed silently on a normalized whitespace/unicode match — docs state that tradeoff; ambiguous multi-match already fails closed. No PDF/trash/PTY/LSP in 0.0.21.
51
+
52
+ ## 0.0.19 → 0.0.20 skills and context progressive disclosure (small intentional breaks)
53
+
54
+ Release **0.0.20** completes Phase 3 progressive skill disclosure in core `@arnilo/prism`:
55
+
56
+ 1. **Default skill prompt is catalog-only.** Active skills render `Skill <name>: <description>` every turn (`skillsDisclosure: "progressive"` default). Full `instructions` appear only after a successful `load_skill` for that session or when the host sets `skillsDisclosure: "eager"`.
57
+ 2. **Runtime `SkillRegistry` without activation is empty.** When `AgentConfig.skills` is a `SkillRegistry` and neither `RunOptions.activeSkills` nor `RunOptions.skills` is set, **zero** skills activate (was `SkillRegistry.list()`). Migration: `activateAllSkills: true` on the run or agent restores list-all activation (still subject to disclosure rules). Plain `Skill[]` configs are unchanged.
58
+ 3. **`load_skill` is host-opt-in.** Export `createLoadSkillTool({ registry, loaded })` from `@arnilo/prism`; register on the active tool set. Unknown names, inactive required tools, oversize bodies, and duplicate loads fail closed; load cannot widen tools or permissions.
59
+ 4. **Context budget honors `ContextBlock.priority`.** Within `context` and `skills` victims, lower priority drops first (missing = 0), then LIFO. Skills with loaded bodies may demote to description-only (`skill_body` omission) before full removal.
60
+ 5. **Optional `toolResultFold`.** Off by default; host `summarize` + thresholds fold aged large tool results in provider view only (session store untouched). Summarizer failure keeps raw results.
61
+
62
+ Example: `node examples/skills-progressive-disclosure.ts` (network-free).
63
+
64
+ ```ts
65
+ // Migration for hosts that relied on activate-all registry behavior:
66
+ await session.run("Hi", { activateAllSkills: true });
67
+
68
+ // Migration for hosts that want full bodies every turn without load_skill:
69
+ const agent = createAgent({ model, provider, skills: registry, skillsDisclosure: "eager" });
70
+ ```
71
+
72
+ Declarative `activateAllCapabilities: true` is unchanged and does **not** set runtime `activateAllSkills`.
73
+
3
74
  ## 0.0.18 → 0.0.19 observational memory lifecycle (small intentional breaks)
4
75
 
5
76
  Release **0.0.19** completes Phase 2 observational memory in `@arnilo/prism-compaction-observational-memory` only; core `@arnilo/prism` runtime behavior is unchanged.