@arnilo/prism 0.0.96 → 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (203) hide show
  1. package/CHANGELOG.md +290 -2
  2. package/README.md +17 -3
  3. package/dist/agent-definitions.js +2 -3
  4. package/dist/agent-event-source.d.ts +11 -0
  5. package/dist/agent-event-source.js +512 -0
  6. package/dist/agent-loops.d.ts +5 -0
  7. package/dist/agent-loops.js +99 -14
  8. package/dist/agent-run-lifecycle.d.ts +5 -2
  9. package/dist/agent-run-lifecycle.js +18 -2
  10. package/dist/agent-run-state.d.ts +27 -1
  11. package/dist/agent-run-state.js +113 -7
  12. package/dist/agents.d.ts +3 -1
  13. package/dist/agents.js +1255 -129
  14. package/dist/artifacts.d.ts +132 -0
  15. package/dist/artifacts.js +44 -0
  16. package/dist/cache-helpers.js +18 -9
  17. package/dist/checkpoints.d.ts +4 -0
  18. package/dist/checkpoints.js +17 -9
  19. package/dist/cli-init.js +3 -7
  20. package/dist/cli-runner.d.ts +2 -6
  21. package/dist/cli-runner.js +71 -33
  22. package/dist/compaction.js +5 -4
  23. package/dist/config.js +7 -4
  24. package/dist/content.js +26 -24
  25. package/dist/context-budget.d.ts +67 -0
  26. package/dist/context-budget.js +288 -0
  27. package/dist/contracts.d.ts +590 -8
  28. package/dist/contracts.js +142 -1
  29. package/dist/contribution-parsing.js +6 -2
  30. package/dist/contributions.d.ts +2 -0
  31. package/dist/contributions.js +3 -0
  32. package/dist/conversations.d.ts +50 -0
  33. package/dist/conversations.js +98 -0
  34. package/dist/credentials.d.ts +22 -2
  35. package/dist/credentials.js +18 -3
  36. package/dist/devices.d.ts +94 -0
  37. package/dist/devices.js +138 -0
  38. package/dist/event-multiplexer.js +18 -4
  39. package/dist/extensions.d.ts +18 -1
  40. package/dist/extensions.js +79 -6
  41. package/dist/feedback.js +12 -10
  42. package/dist/guardrails.d.ts +1 -1
  43. package/dist/guardrails.js +26 -17
  44. package/dist/identity.d.ts +92 -0
  45. package/dist/identity.js +265 -0
  46. package/dist/index.d.ts +94 -72
  47. package/dist/index.js +48 -36
  48. package/dist/input.d.ts +10 -1
  49. package/dist/input.js +152 -52
  50. package/dist/instruction-injection.d.ts +1 -1
  51. package/dist/middleware.js +9 -1
  52. package/dist/models.d.ts +2 -0
  53. package/dist/models.js +3 -0
  54. package/dist/node/agent-definitions.js +16 -8
  55. package/dist/node/contribution-discovery.d.ts +1 -2
  56. package/dist/node/contribution-discovery.js +3 -3
  57. package/dist/node/session-store-jsonl.js +13 -7
  58. package/dist/node/settings.d.ts +1 -1
  59. package/dist/node/settings.js +1 -1
  60. package/dist/node/system-project-prompts.js +2 -4
  61. package/dist/node/trust.js +1 -1
  62. package/dist/persistence-lifecycle.d.ts +103 -0
  63. package/dist/persistence-lifecycle.js +202 -0
  64. package/dist/provider-events.d.ts +1 -0
  65. package/dist/provider-events.js +6 -1
  66. package/dist/provider-request-policy.js +3 -4
  67. package/dist/providers/media.d.ts +1 -1
  68. package/dist/providers/openai-compatible.d.ts +46 -1
  69. package/dist/providers/openai-compatible.js +123 -53
  70. package/dist/providers/openai-primitives.js +10 -7
  71. package/dist/providers/transport.d.ts +6 -0
  72. package/dist/providers/transport.js +21 -0
  73. package/dist/providers.d.ts +2 -0
  74. package/dist/providers.js +3 -0
  75. package/dist/redaction.d.ts +1 -0
  76. package/dist/redaction.js +26 -9
  77. package/dist/resources.d.ts +2 -2
  78. package/dist/resources.js +2 -2
  79. package/dist/retry.d.ts +5 -0
  80. package/dist/retry.js +8 -1
  81. package/dist/rpc.js +55 -11
  82. package/dist/run-ledger.d.ts +6 -0
  83. package/dist/run-ledger.js +16 -13
  84. package/dist/run-limits.js +49 -10
  85. package/dist/secure-agent.js +8 -2
  86. package/dist/security.js +7 -2
  87. package/dist/session-stores.d.ts +7 -2
  88. package/dist/session-stores.js +195 -21
  89. package/dist/skill-disclosure.d.ts +35 -0
  90. package/dist/skill-disclosure.js +101 -0
  91. package/dist/skill-load.d.ts +25 -0
  92. package/dist/skill-load.js +112 -0
  93. package/dist/structured-output.d.ts +5 -1
  94. package/dist/structured-output.js +20 -2
  95. package/dist/system-prompts.js +7 -2
  96. package/dist/testing/agent-event-source-conformance.d.ts +4 -0
  97. package/dist/testing/agent-event-source-conformance.js +54 -0
  98. package/dist/testing/compaction-conformance.js +5 -1
  99. package/dist/testing/extension-conformance.js +15 -3
  100. package/dist/testing/feedback.d.ts +1 -3
  101. package/dist/testing/feedback.js +1 -1
  102. package/dist/testing/persistence-schema.d.ts +2 -2
  103. package/dist/testing/persistence-schema.js +280 -35
  104. package/dist/testing/provider-conformance.js +3 -3
  105. package/dist/testing/run-ledger-conformance.js +1 -1
  106. package/dist/testing/session-store-conformance.d.ts +6 -0
  107. package/dist/testing/session-store-conformance.js +37 -2
  108. package/dist/testing/tool-conformance.js +30 -5
  109. package/dist/testing/tool-effect-store-conformance.d.ts +9 -0
  110. package/dist/testing/tool-effect-store-conformance.js +85 -0
  111. package/dist/thinking.js +4 -1
  112. package/dist/tool-effects.d.ts +15 -0
  113. package/dist/tool-effects.js +352 -0
  114. package/dist/tool-result-fold.d.ts +40 -0
  115. package/dist/tool-result-fold.js +176 -0
  116. package/dist/tools.d.ts +8 -3
  117. package/dist/tools.js +248 -13
  118. package/docs/0.1.0-readiness.md +215 -0
  119. package/docs/a2a.md +33 -2
  120. package/docs/acp.md +152 -0
  121. package/docs/ag-ui-adoption.md +77 -0
  122. package/docs/ag-ui.md +225 -0
  123. package/docs/agent-events.md +34 -3
  124. package/docs/agent-identity.md +144 -0
  125. package/docs/agent-loops.md +17 -2
  126. package/docs/agent-session-runtime.md +21 -4
  127. package/docs/browser-automation.md +5 -0
  128. package/docs/caveman.md +129 -0
  129. package/docs/cli-rpc.md +3 -6
  130. package/docs/coding-agent-tools.md +229 -25
  131. package/docs/coding-security.md +77 -11
  132. package/docs/compaction-and-retry.md +5 -2
  133. package/docs/compaction-llm.md +20 -1
  134. package/docs/compaction-observational-memory.md +52 -8
  135. package/docs/context-and-skills.md +94 -7
  136. package/docs/contribution-registries.md +1 -0
  137. package/docs/conversations.md +135 -0
  138. package/docs/credential-storage.md +34 -1
  139. package/docs/credentials-and-redaction.md +11 -1
  140. package/docs/database-persistence.md +27 -7
  141. package/docs/device-adapters.md +97 -0
  142. package/docs/enterprise-postgres-state.md +178 -0
  143. package/docs/evaluations.md +14 -1
  144. package/docs/extensions.md +4 -1
  145. package/docs/forge-integration.md +113 -0
  146. package/docs/guardrails.md +16 -2
  147. package/docs/host-security.md +35 -4
  148. package/docs/index.md +69 -37
  149. package/docs/input-and-prompt-assembly.md +8 -7
  150. package/docs/language-intelligence.md +162 -0
  151. package/docs/mcp-tools.md +62 -5
  152. package/docs/middleware-hooks.md +2 -2
  153. package/docs/migration.md +427 -2
  154. package/docs/model-routing.md +111 -0
  155. package/docs/multimodal-content.md +8 -5
  156. package/docs/node-jsonl-session-store.md +1 -1
  157. package/docs/observability.md +2 -0
  158. package/docs/openapi-tools.md +56 -0
  159. package/docs/performance.md +282 -0
  160. package/docs/policy-and-audit.md +171 -0
  161. package/docs/ponytail.md +127 -0
  162. package/docs/postgres-persistence.md +8 -4
  163. package/docs/process-sessions.md +147 -0
  164. package/docs/provider-caching.md +13 -1
  165. package/docs/provider-conformance.md +29 -5
  166. package/docs/provider-packages.md +43 -2
  167. package/docs/provider-request-policies.md +2 -0
  168. package/docs/providers/ai-sdk.md +24 -7
  169. package/docs/providers/alibaba.md +179 -0
  170. package/docs/providers/anthropic.md +93 -0
  171. package/docs/providers/azure.md +74 -0
  172. package/docs/providers/bedrock.md +72 -0
  173. package/docs/providers/google.md +89 -0
  174. package/docs/providers/ollama.md +166 -0
  175. package/docs/providers/openai-compatible.md +31 -2
  176. package/docs/providers/openai.md +24 -5
  177. package/docs/providers/openrouter.md +2 -0
  178. package/docs/providers/vertex.md +71 -0
  179. package/docs/public-contracts.md +68 -4
  180. package/docs/rag.md +41 -12
  181. package/docs/release-and-install.md +362 -208
  182. package/docs/resource-loading.md +3 -0
  183. package/docs/runs-and-usage.md +3 -0
  184. package/docs/server.md +44 -6
  185. package/docs/session-store-conformance.md +2 -0
  186. package/docs/session-stores.md +41 -2
  187. package/docs/sqlite-persistence.md +11 -3
  188. package/docs/structured-output.md +7 -1
  189. package/docs/supervisors.md +8 -0
  190. package/docs/tool-effects.md +95 -0
  191. package/docs/tools.md +5 -0
  192. package/docs/work-artifacts-and-review.md +102 -0
  193. package/docs/work-connectors.md +32 -0
  194. package/docs/work-tools.md +137 -0
  195. package/docs/workflows.md +6 -0
  196. package/docs/working-and-semantic-memory.md +40 -7
  197. package/package.json +30 -7
  198. package/templates/init/providers.json +22 -0
  199. package/docs/review-coverage-2026-07-14.md +0 -260
  200. package/docs/review-coverage-2026-07-15.md +0 -193
  201. package/docs/review-coverage-2026-07-17-provider-validation.md +0 -192
  202. package/docs/review-coverage-2026-07-19-phase-3.md +0 -174
  203. package/docs/review-coverage-2026-07-20-phase-4.md +0 -175
@@ -8,9 +8,16 @@
8
8
  | --- | --- |
9
9
  | `createCodingApprovalPolicy(options)` | Returns an `ExecutionPolicy` with trusted roots, read-only mode, command allow/deny rules, approval caching, and timeout/abort-aware approval waits. |
10
10
  | `createSandboxBashOperations(adapter)` | Maps a host-owned `SandboxAdapter` to coding-agent `BashOperations` for delegated shell execution. |
11
- | `createSandboxCodingTools(cwd, options)` | One construction path: full coding tools with shell wired to `options.sandbox` and shared repository options. |
12
- | `createSandboxReadOnlyTools(cwd, options)` | Read-only coding tools (`read`/`repo_list`/`repo_search`) with shared repository options. |
11
+ | `createSandboxCodingComposition(cwd, options)` | Authoritative construction: returns `{ tools, composition }` with required `workspaceMode` (`"host"` \| `"sandbox"`), fail-closed mixed wiring, and containment metadata. |
12
+ | `createSandboxReadOnlyComposition(cwd, options)` | Same contract for read-only tools (`read`/`repo_list`/`repo_search`). |
13
+ | `createSandboxCodingTools` / `createSandboxReadOnlyTools` | Thin wrappers that return `tools` only (compat); still require `workspaceMode`. |
14
+ | `createSandboxFilesystemOperations` / `createSandboxRepositoryOperations` | Optional execFile-backed FS/list/search backends for a disposable sandbox tree. |
13
15
  | `createDockerSandbox(options)` | Creates one disposable non-root Docker container with read-only root/source, bounded tmpfs workspace, typed `execFile`, import/export, and stop/kill/cleanup. |
16
+ | `SandboxProcessHandle` | Optional long-running process handle (`write`/`signal`/`kill`/`release`/`wait`) returned by `DisposableSandbox.startProcess?`. |
17
+ | `createEgressPolicy(options)` | Deny-all allow-list policy: exact host/port/protocol rules plus frozen `npm-registry` / `github` presets; SHA-256 fingerprint. |
18
+ | `createAllowListEgressProxy(options)` | HTTP forward proxy + CONNECT tunnel enforcing the policy: pinned DNS (rebinding defense), private/metadata IP denial, redirect re-validation + hop cap, byte/time caps, per-decision audit, attestation for sandbox composition. |
19
+ | `composeEgressSandboxNetwork(attestation, name)` | Validated custom Docker network carrying proxy attestation; recorded as `prism.egress.*` container labels. |
20
+ | `assertEgressAttestation(attestation)` | Fail-closed validation of proxy attestation evidence. |
14
21
  | `assertPathInsideRoots`, `isPathInsideReal` | Symlink-aware path containment helpers. |
15
22
  | `evaluateCommandRules`, `hasShellMetacharacters` | Command classification helpers. |
16
23
 
@@ -24,14 +31,40 @@ import type { ExecutionAction, ExecutionPolicy, ExecutionDecision } from "@arnil
24
31
 
25
32
  Use this package when coding tools need path scoping, human approval, command rules, or a pluggable sandbox backend. Wire the returned policy through `createCodingTools(cwd, { executionPolicy })` or per-tool `executionPolicy` options.
26
33
 
27
- Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
34
+ Use `createDockerSandbox()` when the host wants a production-reference containment boundary. Prism does **not** claim OS-level isolation unless the host constructs this adapter (or supplies an equivalent custom `DisposableSandbox`). Default policy denies shell/write/edit/delete/move without an `approve` callback and rejects paths outside configured roots. Coding shell definitions are marked `exclusive: true`, matching the approval policy's shell decision, so a single-shot turn containing shell work runs sequentially even when `toolConcurrency > 1`. Non-shell turns retain configured parallelism.
35
+
36
+ Use `createEgressPolicy()` + `createAllowListEgressProxy()` when a coding agent needs outbound network access under an explicit allow list: package installs, forge API calls, or source fetches — never unrestricted egress. The proxy is inert until `start()`; nothing binds or resolves on import or construction.
37
+
38
+ ## Allow-list egress composition
39
+
40
+ `createEgressPolicy({ allow, presets })` builds a deny-all policy. Rules are exact `{ host, port, protocol }` triples — no wildcards, no CIDR, no regex. Presets (`npm-registry`, `github`) expand to explicit rule lists at construction. The policy exposes a stable SHA-256 `fingerprint` over the canonical rule set.
41
+
42
+ ```ts
43
+ import { createEgressPolicy, createAllowListEgressProxy } from "@arnilo/prism-coding-security";
44
+
45
+ const policy = createEgressPolicy({
46
+ allow: [{ host: "api.github.com", port: 443, protocol: "https" }],
47
+ presets: ["npm-registry"],
48
+ });
49
+ const proxy = createAllowListEgressProxy({ policy, audit: (record) => host.recordEgress(record) });
50
+ const endpoint = await proxy.start(); // 127.0.0.1:0 by default; bind a reachable interface for containers
51
+ ```
52
+
53
+ Every request is checked against the policy before any DNS or connect. HTTPS goes through CONNECT tunnels with TLS passed through untouched — no interception, no MITM. DNS answers are resolved once, pinned, and the connected socket's remote address is verified against the pinned set (rebinding defense); private/link-local/metadata ranges (`10/8`, `172.16/12`, `192.168/16`, `127/8`, `169.254/16` incl. `169.254.169.254`, CGNAT, ULA, `::1`, `fe80::/10`) are denied unless the matching rule sets `allowPrivate: true`. Plain-HTTP redirects are followed up to `redirectHops` with every hop re-validated against policy; redirects to unlisted hosts or non-http targets fail closed. Request/response bytes and total transfer time are capped; oversized or slow-loris transfers are cut with `ERR_PRISM_EGRESS_LIMIT`. Every allow/deny writes an `EgressAuditRecord` (id, ts, decision, host, port, protocol, reason, bytes, duration, client address) — never headers, bodies, or tokens. `reloadPolicy()` is the only way to change rules and bumps `policyVersion`; `attestation()` returns `{ proxyEndpoint, denyDirectEgress: true, policyFingerprint, policyVersion, startedAt }` for sandbox composition.
54
+
55
+ Sandbox composition: `composeEgressSandboxNetwork(proxy.attestation(), networkName)` returns a custom `DockerNetworkConfig` whose attestation is validated and recorded as `prism.egress.endpoint` / `prism.egress.fingerprint` / `prism.egress.policyVersion` / `prism.egress.denyDirect=1` container labels. The adapter records evidence; the host must actually restrict the Docker network so the proxy is the only reachable path (e.g., a dedicated network with only the proxy container attached). A custom network without valid attestation fails closed for egress claims, mirroring `assertBrowserSandboxNetwork`.
56
+
57
+ ```ts
58
+ const network = composeEgressSandboxNetwork(proxy.attestation(), "egress-net");
59
+ const sandbox = await createDockerSandbox({ docker, image, sourceRoot, user, network, limits });
60
+ ```
28
61
 
29
62
  ## Inputs / request
30
63
 
31
64
  | Option | Default | Purpose |
32
65
  | --- | --- | --- |
33
66
  | `roots` | required | Realpath-contained filesystem roots. |
34
- | `readOnly` | `false` | Deny shell/write/edit actions. |
67
+ | `readOnly` | `false` | Deny non-`read` actions (including shell/write/edit/delete/move). |
35
68
  | `commandRules` | `[]` | Ordered allow/deny/approval command classification. |
36
69
  | `approve` | none | Host callback for actions not statically allowed; omission fails closed. |
37
70
  | `approvalCacheScope` | `"none"` | Optional `run` or `session` decision cache scope. |
@@ -52,11 +85,25 @@ Use `createDockerSandbox()` when the host wants a production-reference containme
52
85
  | `secrets` | `[]` | Canaries redacted from CLI/adapter errors. |
53
86
  | `limits` | package defaults | CPU/memory/PID/FD/tmpfs/command/export/time caps validated before create. |
54
87
 
88
+ ### Workspace mode inputs (`createSandboxCodingComposition`)
89
+
90
+ | Option | Default | Purpose |
91
+ | --- | --- | --- |
92
+ | `workspaceMode` | **required** | `"host"` (all tools on host cwd; never claims containment) or `"sandbox"` (shell + FS/list/search share one disposable tree). |
93
+ | `sandbox` | optional in host; required for sandbox unless custom ops supplied | `SandboxAdapter` / `DisposableSandbox`. |
94
+ | `workspaceRoot` | `"/workspace"` in sandbox mode | Tree root used as tool cwd when sandbox backends are bound. |
95
+ | `allowMixedWorkspaceWiring` | `false` | Escape hatch: allow sandbox shell + host FS backends. Records `composition.warnings`; forces `containmentClaim: false`. Missing hatch throws. |
96
+ | `read`/`write`/`edit`/`repository.operations` | auto-wired from `DisposableSandbox` in sandbox mode | Host may supply custom tree backends instead of auto-wire. |
97
+
98
+ `0.0.9` silent split (sandbox shell + host FS) is **superseded**. Mixed wiring is never the default.
99
+
55
100
  ## Outputs / response / events
56
101
 
57
102
  `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations`, receive `onData(Buffer)` for ordered stdout/stderr forwarding through the shell tool's existing bounded accumulator, and never grant policy approval themselves.
58
103
 
59
- `createDockerSandbox()` returns a `DisposableSandbox`: typed `execFile(file, args)`, shell-compatible `exec`, `status`, cooperative `stop`, forced `kill`, and idempotent `close`. `close({ export })` can stream a bounded workspace tar plus SHA-256/entry/byte metadata through a host callback; checkpoints should retain only host artifact references/hashes, never whole workspaces.
104
+ `createSandboxCodingComposition()` returns `{ tools, composition }` where `SandboxCodingComposition` carries `workspaceMode`, `containmentClaim`, `mixedWiringAllowed`, `warnings`, `workspaceRoot`, and optional `treeIdentity` (from `importIdentity` / `lastExportIdentity`). `containmentClaim` is `true` only for sandbox mode with tree backends bound and mixed wiring denied. Host mode and escape-hatch mixed wiring always set `containmentClaim: false` — never treat host mode as contained execution.
105
+
106
+ `createDockerSandbox()` returns a `DisposableSandbox`: typed `execFile(file, args)`, shell-compatible `exec`, `status`, cooperative `stop`, forced `kill`, and idempotent `close`. Import may surface `importIdentity`; successful export updates `lastExportIdentity`. `close({ export })` can stream a bounded workspace tar plus SHA-256/entry/byte metadata through a host callback; checkpoints should retain only host artifact references/hashes, never whole workspaces. Optional `startProcess?(SandboxExecFileRequest)` returns a `SandboxProcessHandle` for long-running work consumed by coding-agent `createProcessSessions({ sandbox })`; absence means one-shot-only — ProcessSessions fails closed with `ERR_PRISM_PROCESS_UNSUPPORTED` (no native fallback). The Docker reference adapter does not implement `startProcess` yet; capability is detected, never assumed. See [Process sessions](process-sessions.md).
60
107
 
61
108
  ## Request/response example
62
109
 
@@ -73,8 +120,9 @@ Use `createDockerSandbox()` when the host wants a production-reference containme
73
120
  import {
74
121
  createCodingApprovalPolicy,
75
122
  createDockerSandbox,
76
- createSandboxCodingTools,
123
+ createSandboxCodingComposition,
77
124
  } from "@arnilo/prism-coding-security";
125
+ import { createGitTools } from "@arnilo/prism-coding-agent";
78
126
 
79
127
  const policy = createCodingApprovalPolicy({
80
128
  roots: [workspaceRoot],
@@ -93,12 +141,24 @@ const sandbox = await createDockerSandbox({
93
141
  limits: { cpus: 2, memoryBytes: 2 * 1024 ** 3, maxPids: 256, workspaceBytes: 1024 ** 3 },
94
142
  });
95
143
 
96
- // Host cwd is the inspected workspace; shell runs inside the sandbox.
97
- const tools = createSandboxCodingTools("/srv/jobs/task-1/source", {
144
+ // Sandbox mode: shell/read/write/edit/list/search/glob/delete/move share one disposable tree.
145
+ const { tools, composition } = createSandboxCodingComposition("/srv/jobs/task-1/source", {
146
+ workspaceMode: "sandbox",
98
147
  sandbox,
99
148
  executionPolicy: policy,
100
149
  repository: { exclude: [".git", "node_modules", "dist"] },
101
150
  });
151
+ // composition.containmentClaim === true when backends are bound
152
+
153
+ // Same-tree Git/check (opt-in; not folded into coding tools):
154
+ const gitTools = createGitTools(composition.workspaceRoot, {
155
+ execFile: sandbox.execFile.bind(sandbox),
156
+ commitIdentity: { name: "bot", email: "bot@example.com" },
157
+ });
158
+
159
+ // Host mode (explicit non-contained): omit sandbox; never claim containment.
160
+ const host = createSandboxCodingComposition(hostCwd, { workspaceMode: "host", executionPolicy: policy });
161
+ // host.composition.containmentClaim === false
102
162
 
103
163
  await sandbox.execFile({ file: "npm", args: ["test"], cwd: "/workspace" });
104
164
  await sandbox.close({
@@ -108,17 +168,21 @@ await sandbox.close({
108
168
 
109
169
  ## Extension and configuration notes
110
170
 
111
- Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()`/`createSandboxCodingTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` / `DisposableSandbox` are replaceable and host-owned; approval policy and sandboxing are separate layers. Custom remote sandboxes can implement `DisposableSandbox` without using Docker. `createSandboxCodingTools()` wires shell through the adapter while list/search/read/write/edit keep the host `cwd` unless custom operations are supplied — Docker tmpfs mutations remain inside the container until export. Opt-in structured Git tools from `@arnilo/prism-coding-agent` (`createGitTools`) can target the same disposable sandbox by passing `execFile: sandbox.execFile` and a host `commitIdentity`; Prism still never pushes or opens PRs. Optional `@arnilo/prism-browser` can share the same disposable boundary: use `assertBrowserSandboxNetwork()` before browse-ready custom networks, and `createSharedSandboxBrowserOptions({ workspaceRoot, downloadsRoot, containedProxyAttestation })` so uploads/downloads align with `/workspace` and `/downloads`. Close the browser context before disposing the sandbox.
171
+ Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()`/`createSandboxCodingComposition()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` / `DisposableSandbox` are replaceable and host-owned; approval policy and sandboxing are separate layers. Custom remote sandboxes can implement `DisposableSandbox` without using Docker.
172
+
173
+ `createSandboxCodingComposition()` requires `workspaceMode`. Sandbox mode auto-wires FS/list/search through `DisposableSandbox.execFile` (or host-supplied custom operations) so mutations stay on the disposable tree until export. Host mode runs every coding tool against the host cwd and never sets `containmentClaim`. Sandbox shell + host FS throws unless `allowMixedWorkspaceWiring: true` (warnings + `containmentClaim: false`). Opt-in structured Git tools (`createGitTools(composition.workspaceRoot, { execFile: sandbox.execFile, commitIdentity })`) share the same tree/cwd; Prism still never pushes or opens PRs. Optional `@arnilo/prism-browser` can share the same disposable boundary: use `assertBrowserSandboxNetwork()` before browse-ready custom networks, and `createSharedSandboxBrowserOptions({ workspaceRoot, downloadsRoot, containedProxyAttestation })` so uploads/downloads align with `/workspace` and `/downloads`. Close the browser context before disposing the sandbox.
112
174
 
113
175
  The Docker reference adapter starts by recorded container ID/label, uses argument arrays only, mounts source read-only, populates a size-bounded tmpfs `/workspace`, drops all capabilities, enables `no-new-privileges`, runs with `--init`, and never exposes the Docker socket, privileged mode, or host PID/IPC namespaces. Image pull/build/update stays outside Prism. Protected real-Docker checks are opt-in via `PRISM_TEST_DOCKER_SANDBOX=1` with host-supplied `PRISM_TEST_DOCKER_BIN` and digest-pinned `PRISM_TEST_DOCKER_IMAGE`.
114
176
 
115
- Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
177
+ Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. `ask_user_decision` also maps onto the shared decision model: inside a durable gated agent run its call suspends as a kind-`elicitation` pending decision whose schema carries the choice contract (option-id enums) plus the full question/options UX payload on a Prism-owned schema extension property; the resume decision's `elicitation` payload (a `selectedId`/`selectedIds`/`customText` answer) is validated against the schema and the tool-level answer-shape rules, then resolves the call without invoking the blocking `ask()` callback. The process-local `ask()` path and the workflow suspend/resume path are unchanged. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
116
178
 
117
179
  ## Security and performance notes
118
180
 
119
181
  Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, repository list/search walks, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe.
120
182
 
121
- Docker sandbox containment—not command regexes—enforces filesystem/network/process boundaries for the reference adapter. Network defaults to none; a custom Docker network still requires a host firewall/proxy for DNS/egress claims. Import rejects symlink escapes, devices, FIFOs, and sockets; export counts entries/bytes and hashes before host retention. Secrets in `secrets` are redacted from adapter errors and never exported as environment metadata. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter and Docker daemon.
183
+ Docker sandbox containment—not command regexes—enforces filesystem/network/process boundaries for the reference adapter. Network defaults to none; a custom Docker network still requires a host firewall/proxy for DNS/egress claims. Import rejects symlink escapes, devices, FIFOs, and sockets; export counts entries/bytes and hashes before host retention. Secrets in `secrets` are redacted from adapter errors and never exported as environment metadata. Unified workspace mode reuses existing sandbox/repo/coding hard caps and does not introduce unbounded host↔container sync loops. Host mode and `allowMixedWorkspaceWiring` never claim disposable containment. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter and Docker daemon.
184
+
185
+ The egress proxy is a policy enforcer, not a firewall: it cannot stop a container whose Docker network reaches the internet directly. Egress attestation (`denyDirectEgress: true`) is a claim the host must make true by network topology; the adapter records it as evidence and fails closed when it is absent or malformed. The proxy performs no TLS interception, no DNS rebinding of its own beyond pinning, and no content filtering; audit records contain no secrets. Frozen caps: 32 concurrent connections (hard 256), 64 MiB request/response bytes (hard 1 GiB), 600 s transfer time (hard 1 h), 128 rules (hard 1,024), 5 redirect hops (hard 10).
122
186
 
123
187
  ## Related APIs
124
188
 
@@ -126,5 +190,7 @@ Docker sandbox containment—not command regexes—enforces filesystem/network/p
126
190
  - [Workflows](workflows.md): `runWorkflow` / `resumeWorkflow` / `startWorkflowBackground` composition for coding tasks
127
191
  - [Host security guide](host-security.md)
128
192
  - [Performance limits](performance.md)
193
+ - [Forge integration](forge-integration.md): GitHub adapter whose mutations can be routed through the egress proxy
129
194
  - [Tool execution primitives](tool-execution-primitives.md)
130
195
  - [Security/auth/trust](settings-auth-trust-security.md)
196
+ - [ACP coding-host interop](acp.md): when an ACP client supplies fs/terminal methods and MCP servers, the same containment story applies at the protocol boundary — client paths/dirs pass host seams, MCP servers need host `select` approval (never auto-connect), updates are redacted and capped, and mode switches only narrow or host-authorized widen.
@@ -68,10 +68,12 @@ createDefaultRetryPolicy(options?: DefaultRetryPolicyOptions): RetryPolicy
68
68
  | `maxAttempts` | Total provider-turn attempts; defaults to `3`. |
69
69
  | `baseDelayMs` | First retry delay; defaults to `100`. |
70
70
  | `maxDelayMs` | Backoff cap; defaults to `1000`. |
71
+ | `jitter` | Symmetric jitter fraction on computed delays; defaults to `0.25` (±25%). Set `0` for exact delays. |
72
+ | `random` | Random source for jitter (tests); defaults to `Math.random`. |
71
73
  | `secrets` | Exact known secret strings to redact from retry errors/events. |
72
74
  | `metadata` | Explicit host metadata for retry policy context. |
73
75
 
74
- `RunOptions.retry: false` disables configured retry for that run. Default classification retries generic transient codes/messages such as `ETIMEDOUT`, `ECONNRESET`, `429`, `500`, `502`, `503`, `504`, `timeout`, `rate_limit`, and `temporarily_unavailable`; aborts and non-transient errors fail closed.
76
+ `RunOptions.retry: false` disables configured retry for that run. Default classification retries generic transient codes/messages such as `ETIMEDOUT`, `ECONNRESET`, `429`, `500`, `502`, `503`, `504`, `timeout`, `rate_limit`, and `temporarily_unavailable`; aborts and non-transient errors fail closed. Delays are exponential (`baseDelayMs * 2^(attempt-1)`, capped at `maxDelayMs`) with symmetric jitter, so concurrent sessions do not retry in lockstep during a shared outage. When `ErrorInfo.retryAfterMs` is set — first-party HTTP providers populate it from the `Retry-After` response header via `httpStatusError()` — the hint wins over computed backoff, jitter still applies, and the result is always capped at `maxDelayMs` so a hostile or huge hint cannot pin a run.
75
77
 
76
78
  `CompactionEntryData` is stored in `SessionEntry.data` for compaction entries:
77
79
 
@@ -149,7 +151,7 @@ Retry policies are ordinary `RetryPolicy` implementations and can be registered
149
151
 
150
152
  Compaction strategies are ordinary `CompactionStrategy` implementations. Extensions can register strategies through the existing compaction strategy contribution registry, but registration is inert until a host explicitly selects and passes a strategy to runtime code. Extensions can also register `compaction` middleware; the runtime calls it only when the agent/session has that middleware registry configured.
151
153
 
152
- The default strategy does not call a provider. Hosts that need model-generated summaries can use the optional [`@arnilo/prism-compaction-llm` package](compaction-llm.md); its `maxOutputTokens`/`maxSummaryTokens` budget is passed through `model.parameters.maxTokens` and first-party providers serialize that to provider output-token fields. Hosts that need prepared source-backed memory without a compaction-time model call can use [`@arnilo/prism-compaction-observational-memory`](compaction-observational-memory.md).
154
+ The default strategy does not call a provider. Hosts that need model-generated summaries can use the optional [`@arnilo/prism-compaction-llm` package](compaction-llm.md); its `maxOutputTokens`/`maxSummaryTokens` budget is passed through `model.parameters.maxTokens` and first-party providers serialize that to provider output-token fields. Coding sessions can select that package's `createCodingCompactionStrategy()` preset for paths, patch intent, checks, plans/todos, blockers, and next verification steps; it remains an ordinary `CompactionStrategy` and does not retain complete diffs or add a coding runtime. Hosts that need prepared source-backed memory without a compaction-time model call can use [`@arnilo/prism-compaction-observational-memory`](compaction-observational-memory.md).
153
155
 
154
156
  ## Security and performance notes
155
157
 
@@ -173,5 +175,6 @@ The default strategy does not call a provider. Hosts that need model-generated s
173
175
  - [Configuration and manifests](configuration-and-manifests.md): `compactionStrategy` and `retryPolicy` manifest contribution kinds.
174
176
  - [Provider layer](provider-layer.md): safe provider error codes used by retry classification.
175
177
  - [Credentials and redaction](credentials-and-redaction.md): exact secret redaction helper used by default compaction and retry error handling.
178
+ - [LLM compaction package](compaction-llm.md): `createCodingCompactionStrategy()` is the thin coding-focused preset; see `examples/coding-compaction.ts` for a network-free mock.
176
179
 
177
180
  Runtime redaction composes with compaction and retry secret lists: configured redactors apply at session serialization boundaries, while compaction/retry `secrets` still redact their local summaries and errors.
@@ -12,6 +12,7 @@ Key exports:
12
12
  | Export | Purpose |
13
13
  | --- | --- |
14
14
  | `createLlmCompactionStrategy(options)` | Returns a provider-backed `CompactionStrategy`. |
15
+ | `createCodingCompactionStrategy(options)` | Fixed `coding` preset over the LLM strategy: prioritizes file paths, patch intent, commands/checks, plan/todos, blockers, and next verification while retaining normal limits and raw history. |
15
16
  | `createLlmCompactionExtension(options)` | Registers the strategy into an explicit extension kernel compaction registry. |
16
17
  | `prepareLlmCompaction(context, options?)` | Splits branch entries into summary input, kept suffix, optional split-turn prefix, file details, and compaction data. |
17
18
  | `findLlmCompactionCutPoint(entries, options?)` | Finds the last entry covered by a summary using approximate token budgets. |
@@ -76,6 +77,23 @@ const strategy = createLlmCompactionStrategy({
76
77
  await session.compact({ strategy, secrets: [apiKey] });
77
78
  ```
78
79
 
80
+ Coding-session example:
81
+
82
+ ```ts
83
+ import { createCodingCompactionStrategy } from "@arnilo/prism-compaction-llm";
84
+
85
+ const strategy = createCodingCompactionStrategy({
86
+ provider: summaryProvider,
87
+ summaryModel: { provider: "openai", model: "gpt-4.1-mini" },
88
+ keepRecentTokens: 20_000,
89
+ maxSummaryTokens: 800,
90
+ customInstructions: "Keep migration blockers prominent.",
91
+ });
92
+ await session.compact({ strategy });
93
+ ```
94
+
95
+ The preset always uses strategy name `coding` and enables existing read/modified-file retention. It adds no provider call, parser, worker, filesystem access, or complete-diff retention beyond `createLlmCompactionStrategy()`.
96
+
79
97
  Credential factory example:
80
98
 
81
99
  ```ts
@@ -108,7 +126,7 @@ Preparation is O(n) over branch entries and uses only arrays, strings, and JSON
108
126
 
109
127
  Provider deltas are redacted while retained and stop at `maxSummaryTokens * 4` UTF-16 code units without splitting a surrogate pair. Provider iteration is closed/aborted on overflow. A derived finite event ceiling also stops endless empty/non-text deltas. Final history/turn/file composition receives the same cap. Provider error events, generator throws, provider-factory failures, and policy failures expose only bounded redacted detail; host abort remains authoritative.
110
128
 
111
- The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. It does not discover credentials, read files, start background jobs, or add provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history or provider output.
129
+ The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. The coding preset makes the same calls and uses the same bounded file-operation preparation. Neither discovers credentials, reads files, starts background jobs, or adds provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history, paths, instructions, or provider output.
112
130
 
113
131
  ## Related APIs
114
132
 
@@ -119,3 +137,4 @@ The strategy makes only the needed provider call(s): one history summary plus on
119
137
  - [Agent/session runtime](agent-session-runtime.md): `AgentSession.compact()` and opt-in auto-compaction.
120
138
  - [Provider layer](provider-layer.md): mock providers and provider request contracts.
121
139
  - [Credentials and redaction](credentials-and-redaction.md): exact known-secret redaction behavior.
140
+ - [Frontend interoperability (AG-UI and ACP)](ag-ui.md): separate optional frontend transport; coding compaction adds no UI protocol dependency.
@@ -8,6 +8,21 @@ Current status: ledger/projection/render/recall utilities, explicit worker runti
8
8
 
9
9
  This package is distinct from `@arnilo/prism-memory` working/semantic memory: observational memory compresses and recalls source-backed observations/reflections; semantic memory retrieves embeddings; working memory stores the current structured profile/state. Hosts may compose both.
10
10
 
11
+ ## Four-layer provider context
12
+
13
+ Observational memory composes four independent layers for long sessions (Mastra-style):
14
+
15
+ | Layer | What it holds | How it is produced |
16
+ | --- | --- | --- |
17
+ | **Recent exact messages** | Last `context.recentMessages` user/assistant/tool entries in branch order (optional `recentMessageMaxTokens` trim, oldest first) | `buildObservationalMemoryContextBlocks()` → `recent-messages` ContextBlock; aligned with compaction `keepRecentEntries` |
18
+ | **Observation log** | Source-backed facts with 12-hex ids and `sourceEntryIds` | Observer worker on eligible unscanned `message` entries after `observation.messageTokens`; coverage advances even on empty passes |
19
+ | **Reflections** | Higher-level summaries over observation ids | Reflector worker on observations after last reflection coverage when `reflection.observationTokens` met |
20
+ | **Raw-source retrieval** | Exact branch messages behind a memory id or cursor page | `recallObservationalMemory()` / `recallObservationalMemoryBranchPage()` / `createRecallMemoryTool()` — exact-id or cursor paging only; no semantic search |
21
+
22
+ Activation is explicit: `createObservationalMemory().attach()` coordinates post-run observe/reflect/drop and `context.compactAfterTokens` compaction. Import and extension `setup` start nothing. Recall, commands, and utilities fail closed on invalid ids, wrong `sessionId`, ambiguous tool input, or oversized pages. Pass `secrets` for exact-value redaction in render/recall/worker paths. Branch isolation: hosts supply current-branch `appendEntry` and `getEntries`; mismatched store/session pairs fail closed after append.
23
+
24
+ See `examples/observational-memory-lifecycle.ts` for attach → turn → projection/recall/page without live credentials.
25
+
11
26
  ## When to use it
12
27
 
13
28
  Use it when a host wants to opt in to long-session memory that records observations/reflections as session custom entries, renders prepared memory during compaction, and supports exact-id recall.
@@ -20,7 +35,7 @@ Memory records use `SessionEntry.kind: "custom"` with `entry.data.type` markers:
20
35
 
21
36
  | Type | Payload |
22
37
  | --- | --- |
23
- | `om.observations.recorded` | `{ observations, coversUpToId? }` |
38
+ | `om.observations.recorded` | `{ observations, coversUpToId? }` — successful observer runs append coverage even when `observations` is empty. |
24
39
  | `om.reflections.recorded` | `{ reflections, coversUpToId? }` |
25
40
  | `om.observations.dropped` | `{ observationIds, coversUpToId? }` |
26
41
  | `om.folded` | Compaction `data.memory` folded details. |
@@ -38,6 +53,10 @@ Worker limits are finite positive safe integers:
38
53
  | `maxWorkerResultBytes` | 64 KiB | 1 MiB | Full tool result and replayed value/error payload |
39
54
  | `maxWorkerMessageBytes` | 1 MiB | 8 MiB | System/prompt plus assistant-call/tool-result transcript |
40
55
  | `maxWorkerErrorBytes` | 1 KiB | 8 KiB | Provider/tool/runtime error text after exact known-secret redaction |
56
+ | Rendered memory projection | — | 256 KiB | `renderObservationalMemory()` / context block text |
57
+ | Folded compaction payload | — | 512 KiB | `data.memory` JSON; strategy trims lowest-relevance observations before failing |
58
+ | Recent-message window | — | 512 KiB | `renderRecentMessageWindow()` hard cap |
59
+ | Recall page size | 20 | 100 | `retrieval.pageLimit` / recall tool `limit` |
41
60
 
42
61
  Direct `runObserver()` / `runReflector()` / `runDropper()` calls retain required `maxTurns` and accept the corresponding shorter worker fields (`maxToolCalls`, `maxResultBytes`, etc.). Named default/hard constants and `resolveMemoryWorkerLimits()` are exported.
43
62
 
@@ -48,20 +67,26 @@ Key exports:
48
67
  | Export | Purpose |
49
68
  | --- | --- |
50
69
  | `foldObservationalMemoryLedger()` | Fold custom memory entries into observations, reflections, drops, and coverage markers. |
70
+ | `isEligibleObservationSourceEntry()` / `eligibleObservationSources()` | Select user/assistant/tool `message` entries for observer input. |
71
+ | `unscannedEntries()` / `observationsUncoveredByReflection()` | Dual coverage helpers for observation scan and reflection windows. |
51
72
  | `buildObservationalMemoryProjection()` | Build active/full/folded projections from current branch entries. |
73
+ | `buildObservationalMemoryContextBlocks()` | Render observational-memory + recent-messages context blocks for provider input. |
74
+ | `selectRecentMessageEntries()` / `renderRecentMessageWindow()` | Bounded exact recent-message suffix; count via `keepRecentEntries`, optional token trim via `estimateEntryTokens`. |
52
75
  | `createFoldedMemoryDetails()` | Create JSON details for compaction `data.memory`. |
53
76
  | `renderObservationalMemory()` | Render reflections and observations into a prepared memory summary. |
54
77
  | `recallObservationalMemory()` | Recover source evidence for a known observation/reflection id from supplied current-branch entries. |
78
+ | `recallObservationalMemoryBranchPage()` | Page eligible user/assistant/tool messages around a cursor entry id (`forward`/`backward`, optional `detail: summary|full`). |
55
79
  | `createMemoryId()` / `isMemoryId()` | Create/check 12-character ids. |
56
80
  | `resolveObservationalMemorySettings()` | Merge `observational-memory` settings with defaults and overrides. |
57
- | `createObservationalMemoryRuntime()` | Explicitly run observer/reflector/dropper workers for a supplied session, owned append callback, and provider. |
81
+ | `createObservationalMemory()` / `attach()` | One activation wires post-run observe/reflect/drop and `compactAfterTokens` compaction; returns proxied session, runtime, context provider, and strategy. |
82
+ | `createObservationalMemoryRuntime()` | Low-level explicit flush for advanced hosts or tests. |
58
83
  | `createObservationalMemoryCompactionStrategy()` | Render existing folded memory as a standard Prism compaction summary with `data.memory`. |
59
84
  | `createObservationalMemoryExtension()` | Inert extension helper that registers the strategy contribution unless disabled. |
60
- | `createRecallMemoryTool()` | Optional exact-id `recall` tool factory backed by host-supplied current-branch entries. |
85
+ | `createRecallMemoryTool()` | Optional `recall` tool factory: exact id lookup or current-branch message paging via host-supplied entries. |
61
86
  | `createMemoryStatusCommand()` / `createMemoryViewCommand()` | Optional `om:status` and `om:view` command factories. |
62
87
  | `createObservationalMemoryCommands()` | Convenience factory returning status and view commands. |
63
88
 
64
- Pure utilities create no events, workers, tools, commands, credentials, or provider requests. `createObservationalMemoryRuntime()` runs workers only when the host explicitly constructs it and calls `flush()`. The compaction strategy is O(n) over supplied entries and makes no provider call. Tool and command factories are inert until a host registers/selects them.
89
+ Pure utilities create no events, workers, tools, commands, credentials, or provider requests. `createObservationalMemoryExtension()` and import alone start nothing. `createObservationalMemory().attach()` runs workers only after proxied `run`/`prompt`/`stream`/`compact` complete (or after `wrapResumeRun` / `wrapResumeStream`). `createObservationalMemoryRuntime().flush()` remains for manual/advanced use. Attached `contextProvider` renders two blocks each turn: `observational-memory` (active reflections/observations aligned to the recent-message boundary) and `recent-messages` (last `keepRecentEntries` message entries in branch order, optionally trimmed by `recentMessageMaxTokens` using `estimateEntryTokens`; oldest dropped first). Compaction uses the same `keepRecentEntries` setting. Observer input includes only eligible `message` entries (`user`, `assistant`, `tool`); memory/compaction/bookkeeping entries advance `coversUpToId` scan coverage without entering the observer prompt. Successful observer/reflector runs append coverage markers even when they record zero facts. Reflection uses only active observations recorded after the last `om.reflections.recorded` entry unless `flush({ fullReflectionRebuild: true })`. Attached `flush()` skips with `run_active` while a proxied run is in flight. The compaction strategy is O(n) over supplied entries and makes no provider call. Tool and command factories are inert until a host registers/selects them.
65
90
 
66
91
  ## Request/response example
67
92
 
@@ -74,6 +99,7 @@ Pure utilities create no events, workers, tools, commands, credentials, or provi
74
99
  ```ts
75
100
  import {
76
101
  buildObservationalMemoryProjection,
102
+ createObservationalMemory,
77
103
  createObservationalMemoryCompactionStrategy,
78
104
  createObservationalMemoryExtension,
79
105
  createObservationalMemoryCommands,
@@ -83,6 +109,18 @@ import {
83
109
  renderObservationalMemory,
84
110
  } from "@arnilo/prism-compaction-observational-memory";
85
111
 
112
+ const om = createObservationalMemory({
113
+ observation: { provider: observerProvider, model: observerModel, messageTokens: 10_000 },
114
+ reflection: { provider: reflectorProvider, model: reflectorModel, observationTokens: 20_000 },
115
+ context: { compactAfterTokens: 81_000, recentMessages: 8 },
116
+ retrieval: { pageLimit: 20 },
117
+ });
118
+ const attached = om.attach(session, {
119
+ appendEntry: (entry, options) => store.append(entry, options),
120
+ sessionModel: agent.config.model,
121
+ });
122
+ await attached.session.run("Continue from prior work");
123
+
86
124
  const entries = await session.entries();
87
125
  const projection = buildObservationalMemoryProjection(entries);
88
126
  const summary = renderObservationalMemory(projection.reflections, projection.observations);
@@ -111,13 +149,19 @@ await kernel.load([createObservationalMemoryExtension({ recallTool: { getEntries
111
149
 
112
150
  ## Extension and configuration notes
113
151
 
114
- Settings are read from the `observational-memory` key only when a host calls `resolveObservationalMemorySettings()` or `runtime.flush()`. Defaults are `observeAfterTokens: 10000`, `reflectAfterTokens: 20000`, `compactAfterTokens: 81000`, `observationsPoolMaxTokens: 20000`, `observationsPoolTargetTokens: 10000`, `agentMaxTurns: 16`, `passive: false`, and `debugLog: false`. `agentMaxTurns` now rejects non-integer/non-finite/out-of-range input (hard 64) instead of flooring/falling back. Runtime `maxWorkerTurns` takes precedence.
152
+ Settings resolve to nested `observation` / `reflection` / `dropper` / `context` / `retrieval` groups via `resolveObservationalMemorySettings()`. Defaults: `observation.messageTokens: 10000`, `reflection.observationTokens: 20000`, `context.compactAfterTokens: 81000`, `context.recentMessages: 8`, `context.observationsPoolMaxTokens: 20000`, `dropper.targetTokens: 10000` (from `context.observationsPoolTargetTokens`), `retrieval.pageLimit: 20`, `agentMaxTurns: 16`, `passive: false`, `debugLog: false`. Optional `context.recentMessageMaxTokens` trims the recent-message context window (oldest first) after the count limit.
153
+
154
+ Legacy flat keys still map for pre-1.0 hosts (`observeAfterTokens` → `observation.messageTokens`, `reflectAfterTokens` → `reflection.observationTokens`, `compactAfterTokens` → `context.compactAfterTokens`, `keepRecentEntries` → `context.recentMessages`, flat `workerModel` → all workers when nested models absent). Conflicting flat+nested values throw.
155
+
156
+ Observer/reflector/dropper may use separate providers, models, instructions, thinking levels, credentials, and `requireExplicitModel`. `dropper.policy: "lowest-relevance"` drops deterministically without a model call; default is `"model"`. Top-level `workerProvider` / `workerModel` remain as deprecated aliases.
157
+
158
+ Token counting uses `estimateEntryTokens()` / `estimateMessageTokens()`.
115
159
 
116
- The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, and `workerProvider`. Model selection uses [use-case model selection](use-case-model-selection.md): pass optional `workerModel` (or settings `workerModel`) to override, and `sessionModel: agent.config.model` so workers fall back to the session model when no worker model is configured. `requireExplicitModel: true` restores the historical `missing_model` skip when no explicit worker model is set. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution. Default credential requests use the **resolved** model's provider id.
160
+ The runtime requires host-supplied `session`, an `appendEntry` callback bound to that session's owning store/branch, and at least one worker provider (`observation.provider` or legacy `workerProvider`). Model selection uses [use-case model selection](use-case-model-selection.md): pass per-worker `model` (or settings `observation.model` / `reflection.model` / `dropper.model`) to override, and `sessionModel: agent.config.model` so workers fall back to the session model when no worker model is configured. `requireExplicitModel: true` restores the historical `missing_model` skip when no explicit worker model is set. It no longer accepts a separate `store` option because mismatched session/store pairs can append memory entries outside the active branch. After each memory append, the runtime checks the appended entry is visible at the session leaf and fails closed/restores the previous checkout if the callback points elsewhere. Optional credential resolution is explicit; missing requested credentials skip worker execution. Default credential requests use the **resolved** model's provider id.
117
161
 
118
- `createObservationalMemoryCompactionStrategy()` keeps recent message entries like the default compaction strategy, renders existing observations/reflections as the summary, and returns a standard Prism compaction entry. Its `data` includes `throughEntryId`, `keepEntryIds`, `strategy`, `trigger`, and `memory: { type: "om.folded", version: 1, fullFold, observations, reflections, droppedObservationIds }`. When active observations exceed `observationsPoolMaxTokens`, it performs a full fold into `data.memory`.
162
+ `createObservationalMemoryCompactionStrategy()` keeps recent message entries like the default compaction strategy, renders existing observations/reflections as the summary, and returns a standard Prism compaction entry. Its `data` includes `throughEntryId`, `keepEntryIds`, `strategy`, `trigger`, and `memory: { type: "om.folded", version: 1, fullFold, observations, reflections, droppedObservationIds }`. When active observations exceed `context.observationsPoolMaxTokens`, it performs a full fold and synchronously trims lowest-relevance observations until the folded payload fits hard byte/token caps (or throws a typed error).
119
163
 
120
- `createRecallMemoryTool()` requires `args.id` to match `^[a-f0-9]{12}$`; invalid ids fail before entry lookup. Recall returns text and structured details for observations/reflections, dropped observations, supporting observations, source entries, and missing source ids. It does not search by topic.
164
+ `createRecallMemoryTool()` accepts either `{ id }` for exact memory recall or `{ cursor, limit?, direction?, detail? }` for current-branch raw-message paging (default limit 20, hard cap 100). Reflection recall resolves supporting observations from the full ledger and reports `droppedSupportingObservationIds` / `missingSupportingObservationIds`; dropped supports still return available raw sources. Invalid ids, ambiguous requests, wrong `sessionId`, missing cursors, non-message cursors, and oversized pages fail closed. It does not search by topic.
121
165
 
122
166
  `createMemoryStatusCommand()` reports recorded/dropped/active/visible observations, recorded/visible reflections, pool token counts, and optional runtime in-flight/last-error state. `createMemoryViewCommand()` renders visible memory by default or full active recorded memory with `{ mode: "full" }`; other modes return `Usage: /om:view [full]`.
123
167
 
@@ -6,7 +6,7 @@
6
6
 
7
7
  ## When to use it
8
8
 
9
- Use context resolution when a host wants project/session/context blocks resolved before prompt composition. Use the skill registry when a host wants explicit progressive skill disclosure. Declarative `AgentDefinition.skills` are inactive unless listed; omitted skills means none unless the host uses the migration-only `activateAllCapabilities: true` option.
9
+ Use context resolution when a host wants project/session/context blocks resolved before prompt composition. Use the skill registry when a host wants explicit progressive skill disclosure: catalog `name` + `description` every turn by default, full `instructions` only after `load_skill` or when the host opts into eager mode. Declarative `AgentDefinition.skills` are inactive unless listed; omitted skills means none unless the host uses the migration-only `activateAllCapabilities: true` option.
10
10
 
11
11
  Do not use these helpers as an agent loop, package discovery mechanism, context cache, token budgeter, retrier, credential resolver, semantic skill ranker, tool activator, or permission system.
12
12
 
@@ -107,7 +107,8 @@ The agent/session runtime resolves skills per run and wires each active skill's
107
107
  | Surface | Config shape | Run override | Active skills |
108
108
  | --- | --- | --- | --- |
109
109
  | Runtime agent | `AgentConfig.skills: SkillRegistry` | `RunOptions.activeSkills: ["brief"]` | Named skills only, resolved with `resolveActiveSkills({ registry, names, tools })`. |
110
- | Runtime agent | `AgentConfig.skills: SkillRegistry` | no `activeSkills` / no `skills` | All registry skills (`SkillRegistry.list()`). |
110
+ | Runtime agent | `AgentConfig.skills: SkillRegistry` | no `activeSkills` / no `skills` / no `activateAllSkills` | No skills active (fail-closed default). |
111
+ | Runtime agent | `AgentConfig.skills: SkillRegistry` | `activateAllSkills: true` (run or agent) | All registry skills (`SkillRegistry.list()`), migration opt-in. |
111
112
  | Runtime agent | `AgentConfig.skills: Skill[]` | `RunOptions.skills: [...]` | Override array only. |
112
113
  | Runtime agent | `AgentConfig.skills: Skill[]` | no `RunOptions.skills` | All configured array skills. |
113
114
  | Declarative definition | `AgentDefinition.skills: ["brief"]` | later runtime `activeSkills` optional | Listed names only. |
@@ -118,13 +119,13 @@ Runtime selection precedence mirrors the other `RunOptions` overrides (`redactor
118
119
 
119
120
  1. `AgentConfig.skills` is a `SkillRegistry` and `RunOptions.activeSkills: readonly string[]` (names) is set → the runtime calls `resolveActiveSkills({ registry, names, tools })`.
120
121
  2. `RunOptions.skills: readonly Skill[]` is set → that array replaces `AgentConfig.skills` for the run. This override exists for the case where `AgentConfig.skills` is a plain `Skill[]` (no registry), so name resolution is impossible.
121
- 3. Neither set → all configured runtime skills are active (current behavior; `SkillRegistry.list()` or the plain array as-is). This is not the declarative default.
122
+ 3. Neither set → no skills active when `AgentConfig.skills` is a `SkillRegistry` (fail-closed). Use `activateAllSkills: true` on the run or agent to restore prior list-all behavior (`SkillRegistry.list()`). Plain `Skill[]` configs still activate every configured array skill. This is not the declarative default.
122
123
 
123
124
  names win when a registry exists. `RunOptions.activeSkills` cannot be used against a plain-array `AgentConfig.skills` — use `RunOptions.skills` instead. Use `RunOptions.skills: []` for an explicit no-skills runtime run.
124
125
 
125
- Each active skill contributes two things the runtime now wires together:
126
+ Each active skill contributes two things the runtime wires together:
126
127
 
127
- - `Skill.instructions` → rendered as system messages by `skillMessages()` (active set only).
128
+ - `Skill` prompt text → rendered as system messages by `skillMessages()` / `skillPromptText()` (active set only). Default `skillsDisclosure: "progressive"` sends `Skill <name>: <description>`; full `instructions` appear only when the skill is in the session `LoadedSkillSet` or disclosure is `"eager"`.
128
129
  - `Skill.context: ContextProvider[]` → collected across active skills (`activeSkills.flatMap(s => s.context ?? [])`), resolved through the existing `resolveContextProviders(...)`, and merged into the request's `context` **after** host `AgentConfig.context` blocks. Inactive skills contribute neither instructions nor context.
129
130
 
130
131
  `toolNames` enforcement is live: because selection routes through `resolveActiveSkills()`, a skill demanding a host-inactive tool throws with `Skill ${name} requires inactive tool: ${missing}` **before the first provider turn** — no provider call, no store write, no partial side effect. This is the fail-fast contract the docs already claimed; the runtime now honors it.
@@ -153,7 +154,80 @@ await session.run(input, { skills: [{ name: "verbose", instructions: "Be verbose
153
154
  await session.run(input, { skills: [] }); // explicit no skills for this run
154
155
  ```
155
156
 
156
- Skill selection grants no tool access and cannot bypass permissions — a skill's `toolNames` can only *require* host-active tools, never activate or grant them. Declarative skills also do not activate themselves by presence in a registry; list names on `AgentDefinition.skills` (or pass runtime `activeSkills`) when wanted. Per-skill token budgeting is deferred; the merge order (host context, then skill context) is the only priority knob today.
157
+ Skill selection grants no tool access and cannot bypass permissions — a skill's `toolNames` can only *require* host-active tools, never activate or grant them. Declarative skills also do not activate themselves by presence in a registry; list names on `AgentDefinition.skills` (or pass runtime `activeSkills`) when wanted.
158
+
159
+ ### Progressive skill disclosure
160
+
161
+ `skillsDisclosure` on `AgentConfig` / `RunOptions` (`"progressive"` default, `"eager"` opt-in; run wins) controls how active skills render in provider input:
162
+
163
+ | Mode | Provider view per active skill |
164
+ | --- | --- |
165
+ | `"progressive"` (default) | `Skill <name>: <description>` (or `(no description)` when empty) |
166
+ | `"eager"` | `Skill <name>:\n<instructions>` every turn (pre-0.0.20 behavior) |
167
+
168
+ Catalog caps: **64** entries default / **256** hard; descriptions **512 B** default / **4 KiB** hard; instruction bodies **32 KiB** default / **256 KiB** hard on load/eager render. Oversize catalog/description/instruction payloads fail closed (`SkillDisclosureError` / `SkillLoadError`).
169
+
170
+ ```ts
171
+ import { assembleProviderInput, createLoadedSkillSet } from "@arnilo/prism";
172
+
173
+ const loaded = createLoadedSkillSet(); // session-owned; not checkpoint-persisted in 0.0.20
174
+ const request = await assembleProviderInput({
175
+ model,
176
+ input: "Hi",
177
+ skills: active,
178
+ skillsDisclosure: "progressive",
179
+ loadedSkills: loaded,
180
+ });
181
+ ```
182
+
183
+ ### On-demand skill load (`load_skill`)
184
+
185
+ Hosts opt in by registering `createLoadSkillTool({ registry, loaded })` on the active tool set. The model calls `load_skill { name }` with an exact registry name; success adds the name to the session `LoadedSkillSet` so later turns include `instructions` under progressive mode. The tool does **not** activate tools, widen permissions, or load skills that were not active for the run.
186
+
187
+ Fail-closed cases: unknown name, inactive skill for the run, inactive required `toolNames`, oversize body, duplicate load, missing session loaded-set wiring. Tool output and errors are size-capped; skill text is untrusted host/extension data.
188
+
189
+ ```ts
190
+ import { createAgent, createLoadSkillTool, createSkillRegistry } from "@arnilo/prism";
191
+
192
+ const registry = createSkillRegistry([ponytail, brief]);
193
+ const loadSkill = createLoadSkillTool({ registry }); // session injects loadedSkills at dispatch
194
+ const agent = createAgent({ model, provider, skills: registry, tools: [loadSkill, /* host */] });
195
+ await agent.createSession().run("…", { activeSkills: ["ponytail"] });
196
+ // Turn 1: catalog only. After load_skill({ name: "ponytail" }), later turns include instructions.
197
+ ```
198
+
199
+ ### Third-party behavior packages (Caveman, Ponytail)
200
+
201
+ `@arnilo/prism-caveman` and `@arnilo/prism-ponytail` register upstream skills into the extension kernel skill registry. Hosts should:
202
+
203
+ 1. `kernel.load([createCavemanExtension(...), createPonytailExtension(...)])` with session `appendEntry` / `getEntries` callbacks.
204
+ 2. Build `createSkillRegistry(kernel.registries.skills.list())` and pass `activeSkills` / `resolveActiveSkills` names.
205
+ 3. Keep `skillsDisclosure: "progressive"` and register `createLoadSkillTool` — full `SKILL.md` bodies stay catalog-only until `load_skill`.
206
+ 4. Select `instructionInjectors: ["caveman-mode", "ponytail-mode"]` (or subset) for mode/level slices **without** forcing `skillsDisclosure: "eager"`.
207
+
208
+ Mode slices and skill bodies are independent: the injector can add `PONYTAIL MODE ACTIVE` while `ponytail-audit` remains catalog-only until loaded. See [Caveman](caveman.md), [Ponytail](ponytail.md), and `examples/caveman-ponytail.ts`.
209
+
210
+ Pure validation without the tool: `resolveSkillLoad({ registry, name, tools, loaded, activeSkillNames })`.
211
+
212
+ ### Context budget priority and skill demotion
213
+
214
+ When `assembleProviderInput` runs with `contextBudget`, `applyContextBudget` evicts droppable sections in layout order. Within `context` blocks and skills, victims sort by ascending `ContextBlock.priority` (missing = **0**), then LIFO within the same priority.
215
+
216
+ Under pressure on a skill with a loaded body, eviction may demote to catalog-only first (`ContextBudgetOmissionKind: "skill_body"`), then remove the skill entirely (`"skills"`). Demoted bodies render as description-only even when the name remains in `LoadedSkillSet`. See [Input and prompt assembly](input-and-prompt-assembly.md).
217
+
218
+ ### Optional tool-result fold
219
+
220
+ `toolResultFold` on `AgentConfig` / `RunOptions` (run wins) is **off** unless the host supplies a `summarize` callback. When enabled, aged large tool-result messages in the **provider view** become a one-line header plus bounded summary text; session store entries stay raw. Defaults: `minAgeTurns` **2**, `minBytes` **4096**, `maxSummaryBytes` **512** (hard **4096**). Summarizer failure keeps the raw tool result (fail closed). Not a second memory system — use observational memory / compaction for durable recall.
221
+
222
+ ```ts
223
+ await session.run("…", {
224
+ toolResultFold: {
225
+ minAgeTurns: 2,
226
+ minBytes: 4_096,
227
+ summarize: async ({ toolCallId, text }) => `ref:${toolCallId} ${text.slice(0, 80)}`,
228
+ },
229
+ });
230
+ ```
157
231
 
158
232
  ### Migration note
159
233
 
@@ -164,12 +238,25 @@ For declarative agents, old configs that omitted `skills` should now add explici
164
238
  resolveAgentDefinition({ name: "doc", model, skills: ["brief"] }, context);
165
239
  ```
166
240
 
167
- Use `activateAllCapabilities: true` only as a temporary all-skills/all-tools compatibility opt-in during migration. Runtime `RunOptions.activeSkills` remains the per-run narrowing tool after an agent has a skill registry configured.
241
+ Runtime hosts that relied on `SkillRegistry.list()` when `activeSkills` was omitted must opt in explicitly:
242
+
243
+ ```ts
244
+ // Restore pre-0.0.20 list-all activation (still subject to progressive disclosure):
245
+ await session.run("Hi", { activateAllSkills: true });
246
+
247
+ // Or restore full instruction bodies every turn:
248
+ const agent = createAgent({ model, provider, skills: registry, skillsDisclosure: "eager" });
249
+ ```
250
+
251
+ Use `activateAllCapabilities: true` only as a temporary all-skills/all-tools compatibility opt-in during migration for **declarative** definitions. Runtime `RunOptions.activeSkills` remains the per-run narrowing tool after an agent has a skill registry configured.
168
252
 
169
253
  ## Security and performance notes
170
254
 
171
255
  - Context providers run sequentially and deterministically in caller order.
172
256
  - Skill registry lookup is `Map`-backed, and selection is linear in requested skills plus active tools. Strict duplicate mode adds one O(1) `Map.has()` check during registration only.
257
+ - Progressive catalog render is O(active skills) with byte/count caps; `load_skill` lookup is O(1). Budget eviction over context/skills is O(n log n) worst case.
258
+ - `load_skill` cannot grant tools; loaded instructions are untrusted text bounded by hard caps. `toolResultFold` summarizer output is untrusted and capped; failures keep raw tool results.
259
+ - Loaded-skill names are session-scoped in memory only in 0.0.20 — not checkpoint-persisted; new sessions start catalog-only until reload.
173
260
  - These helpers perform no provider calls, tool execution, resource loading, package discovery, filesystem/network access, retries, timers, or watchers by themselves.
174
261
  - Context and skill output is host/extension data. Do not include secrets unless the host explicitly accepts that prompt exposure.
175
262
  - Active tools remain host-supplied; skills and middleware do not activate tools or grant permissions. Use `duplicate: "error"` when loading third-party skills to prevent silent name shadowing.
@@ -27,6 +27,7 @@ createContributionRegistries(options?: { duplicate?: "replace" | "error" }): Con
27
27
  | Method | Input | Result |
28
28
  | --- | --- | --- |
29
29
  | `register(key, contribution)` | string key and contribution | Stores/replaces the contribution for that key; throws `Duplicate <label>: <key>` when `duplicate: "error"`. |
30
+ | `unregister(key)` | string key | Removes the contribution; returns `false` when the key was not registered. `providers.unregister(id)` and `models.unregister(provider, model)` mirror this on the specialized registries. |
30
31
  | `get(key)` | string key | Returns the contribution or `undefined`. |
31
32
  | `resolve(key)` | string key | Returns the contribution or throws `Unknown <label>: <key>`. |
32
33
  | `list()` | none | Returns contributions in insertion order. |