@arnilo/prism 0.0.4 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/CHANGELOG.md +46 -1
  2. package/README.md +34 -10
  3. package/dist/agent-loops.d.ts +1 -0
  4. package/dist/agent-loops.js +26 -16
  5. package/dist/agents.js +147 -21
  6. package/dist/cli-init.d.ts +41 -0
  7. package/dist/cli-init.js +390 -0
  8. package/dist/cli-runner.d.ts +7 -1
  9. package/dist/cli-runner.js +13 -1
  10. package/dist/content.d.ts +19 -0
  11. package/dist/content.js +197 -69
  12. package/dist/contracts.d.ts +96 -9
  13. package/dist/contracts.js +8 -0
  14. package/dist/feedback.d.ts +48 -0
  15. package/dist/feedback.js +230 -0
  16. package/dist/ids.d.ts +2 -0
  17. package/dist/ids.js +6 -0
  18. package/dist/index.d.ts +10 -4
  19. package/dist/index.js +6 -3
  20. package/dist/providers/media.d.ts +3 -1
  21. package/dist/providers/media.js +11 -1
  22. package/dist/session-stores.js +2 -3
  23. package/dist/testing/feedback.d.ts +6 -0
  24. package/dist/testing/feedback.js +37 -0
  25. package/dist/testing/persistence-schema.d.ts +48 -10
  26. package/dist/testing/persistence-schema.js +166 -22
  27. package/dist/testing/run-ledger-conformance.js +7 -1
  28. package/dist/thinking.d.ts +42 -0
  29. package/dist/thinking.js +92 -0
  30. package/dist/tools.js +2 -3
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +75 -0
  34. package/docs/agent-events.md +14 -21
  35. package/docs/agent-loops.md +12 -9
  36. package/docs/agent-session-runtime.md +14 -16
  37. package/docs/cli-rpc.md +35 -7
  38. package/docs/coding-agent-tools.md +35 -14
  39. package/docs/coding-security.md +7 -3
  40. package/docs/compaction-llm.md +17 -7
  41. package/docs/compaction-observational-memory.md +30 -4
  42. package/docs/context-and-skills.md +1 -0
  43. package/docs/credential-storage.md +58 -9
  44. package/docs/credentials-and-redaction.md +3 -3
  45. package/docs/database-persistence.md +17 -9
  46. package/docs/evaluations.md +122 -0
  47. package/docs/extensions.md +2 -2
  48. package/docs/host-security.md +26 -5
  49. package/docs/index.md +43 -28
  50. package/docs/mcp-tools.md +74 -13
  51. package/docs/migration.md +177 -3
  52. package/docs/multimodal-content.md +14 -6
  53. package/docs/node-filesystem-config.md +1 -0
  54. package/docs/node-jsonl-session-store.md +5 -4
  55. package/docs/observability.md +14 -6
  56. package/docs/performance.md +209 -0
  57. package/docs/postgres-persistence.md +8 -6
  58. package/docs/provider-caching.md +16 -4
  59. package/docs/provider-conformance.md +40 -1
  60. package/docs/provider-packages.md +62 -3
  61. package/docs/providers/ai-sdk.md +149 -0
  62. package/docs/providers/kimi.md +124 -61
  63. package/docs/providers/neuralwatt.md +19 -13
  64. package/docs/providers/openai.md +56 -13
  65. package/docs/providers/opencode-go.md +118 -30
  66. package/docs/providers/openrouter.md +105 -35
  67. package/docs/providers/zai.md +94 -45
  68. package/docs/public-contracts.md +6 -5
  69. package/docs/rag.md +113 -0
  70. package/docs/release-and-install.md +100 -79
  71. package/docs/review-coverage-2026-07-15.md +193 -0
  72. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  73. package/docs/runs-and-usage.md +42 -5
  74. package/docs/server.md +139 -0
  75. package/docs/settings-auth-trust-security.md +5 -5
  76. package/docs/sqlite-persistence.md +6 -5
  77. package/docs/structured-output.md +1 -1
  78. package/docs/supervisors.md +71 -0
  79. package/docs/thinking-and-reasoning.md +98 -0
  80. package/docs/tool-execution-primitives.md +3 -3
  81. package/docs/tools.md +15 -0
  82. package/docs/use-case-model-selection.md +109 -0
  83. package/docs/workflow-orchestration-primitives.md +20 -3
  84. package/docs/workflows.md +114 -33
  85. package/docs/working-and-semantic-memory.md +170 -0
  86. package/package.json +13 -3
  87. package/templates/init/README.md.tmpl +28 -0
  88. package/templates/init/env.example.tmpl +1 -0
  89. package/templates/init/gitignore.tmpl +11 -0
  90. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  91. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  92. package/templates/init/package.json.tmpl +22 -0
  93. package/templates/init/providers.json +76 -0
  94. package/templates/init/src/agent.ts.tmpl +10 -0
  95. package/templates/init/src/index.ts.tmpl +12 -0
  96. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  97. package/templates/init/tsconfig.json.tmpl +15 -0
package/docs/a2a.md ADDED
@@ -0,0 +1,75 @@
1
+ # A2A interoperability
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-supervisor` implements a bounded text-only subset of Agent2Agent (A2A) protocol 1.0: Agent Cards, JSON-RPC `SendMessage`, `SendStreamingMessage`, `GetExtendedAgentCard`, SSE task updates, ES256 JWS card signatures, and an explicit remote client.
6
+
7
+ ## When to use it
8
+
9
+ Use it to expose one explicitly selected Prism agent at an A2A endpoint or call a known remote A2A agent. Do not use it as endpoint discovery, a generic proxy, credential forwarding, or a replacement for local workflows.
10
+
11
+ ## Inputs / request
12
+
13
+ | API/field | Meaning |
14
+ | --- | --- |
15
+ | `createA2AAgentCard(card)` | Validates/freeze a JSONRPC protocol-1.0 HTTPS text card. |
16
+ | `signA2AAgentCard(card, { privateKey, keyId, expiresAt })` | Adds detached-payload ES256 JWS signature using WebCrypto. |
17
+ | `verifyA2AAgentCard(card, { publicKey, keyId?, now?, maxAgeMs? })` | Pins ES256/key/expiry and verifies canonical unsigned card. |
18
+ | `createA2AHandler({ card, exposure, authorize })` | Web-standard card/JSON-RPC/SSE `Request` to `Response` handler. |
19
+ | `createA2AClient({ endpoint, allowedOrigins })` | Explicit HTTPS remote client with optional card verifier/auth callback. |
20
+ | `A2ALimits` | Request 64 KiB, response 1 MiB, event 64 KiB, stream 10 MiB/10k events, concurrency 16, timeout 120s, card 64 KiB defaults; finite hard caps apply. |
21
+
22
+ ## Outputs / response / events
23
+
24
+ The handler serves `GET /.well-known/agent-card.json` and its configured POST endpoint. JSON-RPC returns `{ result: { task } }` or a bounded error. Streaming returns backpressure-driven SSE task envelopes. Client `send()` maps a terminal remote task to `AgentRunResult`; `stream()` incrementally yields validated/redacted text artifacts. Client SSE accepts LF, CRLF, mixed blank-line separators, comments/unknown fields, and multiline `data:` joined with LF.
25
+
26
+ ## Request/response example
27
+
28
+ ```json
29
+ {"jsonrpc":"2.0","id":1,"method":"SendMessage","params":{"message":{"role":"user","messageId":"m1","parts":[{"text":"Check sources"}]}}}
30
+ ```
31
+
32
+ ## Implementation example
33
+
34
+ ```ts
35
+ import { createA2AClient, createA2AHandler, verifyA2AAgentCard } from "@arnilo/prism-supervisor";
36
+
37
+ const handler = createA2AHandler({
38
+ card,
39
+ exposure: { sessionFactory: ({ ownership }) => agent.createSession({ metadata: ownership }) },
40
+ authorize: ({ request }) => authenticate(request),
41
+ });
42
+
43
+ const client = createA2AClient({
44
+ endpoint: "https://agent.example/a2a/v1",
45
+ allowedOrigins: ["https://agent.example"],
46
+ authorize: () => ({ authorization: `Bearer ${resolveOwnedToken()}` }),
47
+ verifyCard: (remoteCard) => verifyA2AAgentCard(remoteCard, { publicKey, keyId: "agent-key" }),
48
+ });
49
+
50
+ const result = await client.send("Check sources");
51
+ ```
52
+
53
+ ## Extension and configuration notes
54
+
55
+ The package owns no listener or credential store. Mount the handler in a host server and resolve authentication/authorization on every request. Client auth executes only after body/card validation and serialization. Injectable `fetch` supports host transports/tests; redirects are disabled.
56
+
57
+ Only `text` parts are accepted. File/data parts, push notifications, task persistence/query/cancel, gRPC, HTTP+JSON binding, automatic JWK fetching, and endpoint discovery are intentionally absent.
58
+
59
+ ## Security and performance notes
60
+
61
+ - Endpoints and card URLs must be HTTPS and exactly origin-allow-listed before fetch; `redirect: "error"` prevents redirect SSRF.
62
+ - Treat every remote card, error, task, status, artifact, and SSE frame as untrusted. Shape/count/byte/time limits apply before mapping. Streaming keeps raw stream bytes, current frame bytes, and event count as separate existing limits.
63
+ - One fatal streaming UTF-8 decoder is reused across every body chunk and flushed once at EOF. Split multibyte code points are preserved; malformed/truncated UTF-8 fails rather than inserting `U+FFFD` into JSON. A small coalesced line buffer keeps one-byte chunk handling incremental.
64
+ - SSE frames require a terminating blank line. A non-whitespace final partial frame, malformed JSON, missing terminal task, failed/canceled task, or any event after a completed task fails with bounded package-owned text. Existing request/response/event/stream/count/timeout hard caps are unchanged.
65
+ - Card verification pins `alg=ES256`, optional key ID, issue/expiry, optional maximum age, and canonical unsigned-card payload. Hosts provision trusted public keys; remote `jku` is never fetched automatically.
66
+ - Card discovery is public; extended-card and invoke methods call host authorization. Use TLS, rate limits, and replay controls at the host edge.
67
+ - Credentials remain in the client auth callback or server authorizer and never enter cards, messages, events, or metrics.
68
+ - Offline conformance is authoritative. Live endpoints are optional operator smoke tests.
69
+
70
+ ## Related APIs
71
+
72
+ - [Supervisor delegation](supervisors.md): local child boundary.
73
+ - [Web-standard server](server.md): non-A2A Prism routes.
74
+ - [Host security](host-security.md): authentication, SSRF, and untrusted-output policy.
75
+ - [Agent/session runtime](agent-session-runtime.md): mapped local execution/result.
@@ -8,7 +8,7 @@ Events are emitted by the runtime and by loops through `LoopContext.emit`, both
8
8
 
9
9
  ## When to use it
10
10
 
11
- Subscribe via `session.subscribe()` whenever a host needs to observe run progress: render streamed assistant text in a UI, react to tool execution, drive observability/telemetry, or audit artifact validation outcomes. Do not parse provider stream events directly for these — `AgentEvent` is the stable, normalized surface across providers and loops.
11
+ Subscribe via `session.stream()` for a single owned run, or `session.subscribe()` when a host needs a long-lived observer across runs: render streamed assistant text in a UI, react to tool execution, drive observability/telemetry, or audit artifact validation outcomes. Do not parse provider stream events directly for these — `AgentEvent` is the stable, normalized surface across providers and loops.
12
12
 
13
13
  Do not use `AgentEvent` for durable replay (use a `SessionStore`) or for cross-session coordination (the broadcaster is per-session and live-only).
14
14
 
@@ -58,7 +58,7 @@ Agent / turn / message events:
58
58
  | Variant | Fields |
59
59
  | --- | --- |
60
60
  | `agent_started` | `sessionId`, `runId` |
61
- | `agent_finished` | `sessionId`, `runId`, `usage?: Usage` |
61
+ | `agent_finished` | `sessionId`, `runId`, `usage?: Usage` (aggregate of all usage-bearing provider turns) |
62
62
  | `turn_started` / `turn_finished` | `sessionId`, `runId`, `turn: number` |
63
63
  | `message_started` / `message_finished` | `sessionId`, `runId`, `message: Message` |
64
64
  | `message_delta` | `sessionId`, `runId`, `content: ContentBlock` (`tool_call_delta` fragments may appear here for live UI streaming; stored messages use final `tool_call` blocks) |
@@ -100,28 +100,23 @@ Artifact validation/refinement events (emitted only by `generateValidateReviseLo
100
100
  | `artifact_validation_finished` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` |
101
101
  | `artifact_revision_started` | `sessionId`, `runId`, `turn`, `attempt`, `failure: ArtifactValidation` |
102
102
  | `artifact_finished` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (loop ended successfully) |
103
- | `artifact_failed` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (budget exhausted) |
103
+ | `artifact_failed` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (candidate budget exhausted, or `result.metadata.reason === "tool_round_limit"`) |
104
104
 
105
105
  ### Artifact event ordering
106
106
 
107
- A `generateValidateReviseLoop` run emits normal turn/message events for every provider turn, then a strictly ordered artifact sequence, correlated by `runId` / `turn` / `attempt`:
107
+ A call-free candidate in `generateValidateReviseLoop` emits normal turn/message events then a strictly ordered artifact sequence, correlated by `runId` / `turn` / `attempt`:
108
108
 
109
109
  ```
110
- turn_started
111
- → message_started
112
- → message_delta*
113
- → message_finished
114
- → turn_finished
115
- → artifact_validation_started
116
- → artifact_validation_finished
117
- → artifact_revision_started # when a revision will run next
118
- | artifact_finished # loop ended successfully
119
- | artifact_failed # budget exhausted (maxRevisions+1 attempts)
110
+ turn_started → message_started → message_delta* → message_finished → turn_finished
111
+ → artifact_validation_started → artifact_validation_finished
112
+ → artifact_revision_started | artifact_finished | artifact_failed
120
113
  ```
121
114
 
122
- - `attempt` is 1-indexed per validation attempt and equals the provider `turn` within `generateValidateReviseLoop`; it mirrors `retry_scheduled.attempt` and the `tool_execution_*` block/finish pairing.
115
+ With opt-in `toolCalls: "bounded"`, a provider turn containing calls emits its normal assistant envelope followed by existing `tool_execution_*` events and matching persisted tool results; it emits no validation event and the next provider turn consumes that transcript. A post-`maxToolRounds` call emits terminal `artifact_failed` directly after `turn_finished` and has no tool execution event.
116
+
117
+ - `attempt` is 1-indexed per call-free validation candidate. It can differ from provider `turn` when bounded tool calls occur.
123
118
  - Single-shot runs emit zero artifact events.
124
- - **Validation failure triggering a revision is recoverable and never an `error`.** Only terminal budget exhaustion emits `artifact_failed`. The `error` channel is reserved for real failures (provider failures not caught by retry, aborts, etc.), matching the existing convention used by `tool_execution_blocked`.
119
+ - **Validation failure triggering a revision is recoverable and never an `error`.** Terminal candidate-budget or `tool_round_limit` exhaustion emits `artifact_failed`; real failures remain on the `error` channel.
125
120
 
126
121
  ## Request/response example
127
122
 
@@ -173,12 +168,10 @@ const session = createAgent({
173
168
  provider: createMockProvider([providerTextDelta("ok"), providerDone()]),
174
169
  }).createSession();
175
170
 
176
- for await (const event of session.subscribe()) {
171
+ for await (const event of session.stream("draft", { loop: { strategy: "generate-validate-revise", validator, maxRevisions: 3 } })) {
177
172
  if (event.type === "artifact_finished") console.log("artifact ok", event.attempt);
178
173
  if (event.type === "artifact_failed") console.log("artifact exhausted", event.attempt, event.result.errors);
179
174
  }
180
-
181
- await session.run("draft", { loop: { strategy: "generate-validate-revise", validator, maxRevisions: 3 } });
182
175
  ```
183
176
 
184
177
  ## Extension and configuration notes
@@ -195,11 +188,11 @@ await session.run("draft", { loop: { strategy: "generate-validate-revise", valid
195
188
  - Slow consumers are bounded by `SubscribeOptions`. Use `RunLedger` or host storage for durable replay; do not rely on a live subscriber as a queue.
196
189
  - Redaction is exact-string-match only and opt-in via `createSecretRedactor`; values not passed as known secrets are not redacted.
197
190
  - `ArtifactValidation.errors[].message` and `metadata` may echo model text; `redactAgentEvent` walks arbitrary nesting and replaces cyclic references with `"[Circular]"` (WeakSet cycle guard), so secret values in `result`/`failure` are redacted without crashing.
198
- - `artifact_*` events are bounded by `maxRevisions + 1` validation attempts; an always-failing validator cannot loop forever and emits exactly one terminal `artifact_failed`.
191
+ - `artifact_*` validation events are bounded by `maxRevisions + 1` call-free candidates. With opt-in bounded artifact tools, provider turns are additionally bounded by run-global `maxToolRounds` (maximum `1 + maxRevisions + maxToolRounds`); a post-cap call emits exactly one terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"` and has no tool lifecycle event because it never dispatches.
199
192
  - Runtime events contain messages/content only; do not put secrets in prompts, metadata, provider events, session entries, tool results, or artifact validation payloads.
200
193
 
201
194
  ## Related APIs
202
- - [Agent/session runtime](agent-session-runtime.md): `session.subscribe()` and the live event broadcaster.
195
+ - [Agent/session runtime](agent-session-runtime.md): `session.stream()`, `session.subscribe()`, and the live event broadcaster.
203
196
  - [Agent loops](agent-loops.md): `singleShotLoop` and `generateValidateReviseLoop` emit the artifact events.
204
197
  - [Structured output](structured-output.md): `ArtifactValidation` shape threaded through parser/validator/repairer.
205
198
  - [Public contracts](public-contracts.md): full `AgentEvent` union and `ArtifactValidation` contract.
@@ -16,7 +16,7 @@ The `Artifact*` contracts (`ArtifactValidation`, `ArtifactContext`, `ArtifactPar
16
16
 
17
17
  Use the default `singleShotLoop` implicitly whenever you call `session.run()` — no configuration needed. Opt into `generateValidateReviseLoop` when a run should produce an artifact that must satisfy a host-supplied schema before it is considered complete (e.g. structured output, a validated JSON document, a generated file passing lint) and the host wants Prism to drive the revision turns.
18
18
 
19
- Do not use a loop to re-implement provider calls, retry, abort, store, or event emission — those stay runtime-owned and are exposed to the loop only through `LoopContext`. A loop that needs tools in revision turns is out of scope for `generateValidateReviseLoop`; use `singleShotLoop` or supply a custom `AgentLoopStrategy`.
19
+ Do not use a loop to re-implement provider calls, retry, abort, store, or event emission — those stay runtime-owned and are exposed to the loop only through `LoopContext`. Artifact-loop tools stay disabled by default; opt into bounded calls only for host-registered, least-privilege lookup tools that must inform an artifact candidate.
20
20
 
21
21
  ## Inputs / request
22
22
 
@@ -57,6 +57,7 @@ await session.run(input, {
57
57
  parser: hostParser, // optional; default treats assistant text as the value
58
58
  repairer: hostRepairer, // optional; default stringifies validation.errors[].message
59
59
  maxRevisions: 3, // optional; default 3
60
+ toolCalls: "bounded", // optional; default "disabled"; uses RunOptions.maxToolRounds
60
61
  },
61
62
  });
62
63
 
@@ -79,6 +80,8 @@ type AgentLoopOptions =
79
80
  readonly parser?: ArtifactParser<unknown>;
80
81
  readonly repairer?: ArtifactRepairer<unknown>;
81
82
  readonly maxRevisions?: number;
83
+ /** Default "disabled". "bounded" dispatches sequentially up to RunOptions.maxToolRounds. */
84
+ readonly toolCalls?: "disabled" | "bounded";
82
85
  };
83
86
  ```
84
87
 
@@ -108,11 +111,11 @@ Host callback contracts (all generic over host `T`):
108
111
 
109
112
  ## Outputs / response / events
110
113
 
111
- `AgentLoopStrategy.run(ctx)` returns `Promise<Usage | undefined>` — the last provider usage, handed back to the runtime which emits `agent_finished` with it.
114
+ `AgentLoopStrategy.run(ctx)` returns `Promise<Usage | undefined>` as a fallback for custom loops. Core runtime independently accumulates every usage-bearing provider turn in O(turns), persists scoped turn/run rows, and emits `agent_finished` with the aggregate.
112
115
 
113
116
  Events during a loop run are the existing `AgentEvent`s (`turn_started`, `message_started`, `message_delta`, `message_finished`, `turn_finished`, tool-execution events when the loop dispatches tools, `error` on real failures). Both built-in loops emit `turn_started` before each provider turn, `message_finished` for every assistant draft, and `turn_finished` after the assistant draft is appended. First-turn input is appended to live history once, matching the already-persisted user message.
114
117
 
115
- Validation-failure-triggering-a-revision is **not** an `error` event — it is recoverable, like `tool_execution_blocked`. `generateValidateReviseLoop` emits normal turn/message events around each provider turn, then the artifact event sequence `artifact_validation_started` → `artifact_validation_finished` → (`artifact_revision_started`)* → `artifact_finished` (success) | `artifact_failed` (budget exhausted), correlated by `runId`/`turn`/`attempt`; see [Agent events § Artifact event ordering](agent-events.md#artifact-event-ordering). `singleShotLoop` emits zero artifact events. Real failures stay on the `error` channel.
118
+ Validation-failure-triggering-a-revision is **not** an `error` event — it is recoverable, like `tool_execution_blocked`. In bounded artifact mode, a tool-calling provider response emits normal assistant/tool lifecycle events, skips artifact parsing/validation, then the next turn sees its persisted result. `generateValidateReviseLoop` emits artifact events only for call-free candidates: `artifact_validation_started` → `artifact_validation_finished` → (`artifact_revision_started`)* → `artifact_finished` | `artifact_failed`. A request beyond `maxToolRounds` executes nothing and emits terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"`; see [Agent events § Artifact event ordering](agent-events.md#artifact-event-ordering). `singleShotLoop` emits zero artifact events. Real failures stay on the `error` channel.
116
119
 
117
120
  A loop has no path to credentials, provider objects, or unredacted secrets. `LoopContext.generate` receives the already-policy-applied, middleware-run, redacted request; `LoopContext.emit` runs through `redactAgentEvent` with the active `SecretRedactor`.
118
121
 
@@ -198,17 +201,17 @@ await session.run(input, { loop: twoShotLoop });
198
201
  - `RunOptions.loop` wins over `AgentConfig.loop`; when neither is set the runtime uses `singleShotLoop`. This mirrors the other `RunOptions` overrides (`redactor`, `validate`, `activeSkills`).
199
202
  - `{ strategy: "single-shot" }` resolves to the exported `singleShotLoop`; `{ strategy: "generate-validate-revise", ... }` is mapped by `resolveLoop()` to `generateValidateReviseLoop(opts)`. An unknown `strategy` throws before the first turn. Passing an `AgentLoopStrategy` instance bypasses the options form entirely (custom-loop escape hatch).
200
203
  - The loop is resolved once per run inside `RuntimeAgentSession.run()`, after the usual setup (provider/skills/tools resolution, history rebuild, model-change entry, input append, auto-compaction). The runtime's outer try/catch/finally, run-exclusivity, abort bridging, and subscriber close remain in place around `loop.run(ctx)`.
201
- - `LoopContext.assemble(nextInput, toolResults?)` accepts an optional tool-result accumulator so `singleShotLoop` can pass its loop-local `toolResults`; `generateValidateReviseLoop` omits it (no tools in revision turns).
202
- - `maxToolRounds` bounds `singleShotLoop` tool rounds; `toolConcurrency` (default `1`) bounds how many independent tool calls from one provider turn may execute concurrently. Results and transcript rows are still appended in original call order. If any resolved `ToolDefinition` in a turn has `exclusive: true`, that turn uses concurrency `1`; later non-exclusive turns restore configured concurrency.
203
- - `maxRevisions` (default 3) bounds `generateValidateReviseLoop` revision turns. Budget exhaustion ends the loop and returns the last usage; it does not throw.
204
- - A revision cycle appends one assistant draft and one repair user message per revision to the session store, so store entries reflect every attempted draft. The original user input is stored once by the runtime and pushed into loop history once on the first turn.
204
+ - `LoopContext.assemble(nextInput, toolResults?)` accepts an optional tool-result accumulator so `singleShotLoop` can pass its loop-local results. Bounded artifact tools append results directly to shared history, then assemble the next turn with empty new input; no second transcript path exists.
205
+ - `maxToolRounds` bounds both `singleShotLoop` and opt-in bounded artifact tool rounds across the whole run. Artifact mode always dispatches sequentially, regardless of `toolConcurrency`; all dispatches still use existing registry/filter/permission/validator/middleware/redactor/ledger guards.
206
+ - `maxRevisions` (default 3) counts only failed call-free artifact candidates. Bounded artifact runs make at most `1 + maxRevisions + maxToolRounds` provider turns. A tool-round limit is terminal and returns last usage after `artifact_failed`; it does not throw.
207
+ - A revision cycle appends one assistant draft and one repair user message per revision to the session store, so store entries reflect every attempted draft. The original user input is stored once by the runtime and pushed into loop history once on the first turn. Repair messages are assembled as the next provider `nextInput` and only pushed into live history after that revision request has been generated, so the model never receives a duplicated repair instruction.
205
208
 
206
209
  ## Security and performance notes
207
210
 
208
211
  - Loops have no path to credentials, provider objects, or unredacted secrets. `LoopContext.generate` consumes an already-redacted request; `LoopContext.emit` runs through `redactAgentEvent` with the active `SecretRedactor`; `LoopContext.appendMessage` appends a redacted entry.
209
212
  - `ArtifactValidation.errors[].message` may echo model text — `artifact_*` event payloads flow through the same `redactAgentEvent` path as other `AgentEvent`s (see [Agent events](agent-events.md)).
210
- - `generateValidateReviseLoop` makes at most `maxRevisions + 1` provider turns; it cannot loop forever on an always-failing validator. Each revision costs one provider turn plus one store append.
211
- - Parallel tool dispatch uses a bounded worker pool over the calls in one turn; queue depth is `calls.length`, not unbounded. Exclusive turns use the same sequential path. Each call still runs through `dispatchToolCall` (permission + validation + execute). Tool lifecycle events may complete out of order; history/store appends stay in call order.
213
+ - `generateValidateReviseLoop` makes at most `1 + maxRevisions + maxToolRounds` provider turns when bounded tools are enabled (otherwise `maxRevisions + 1`); it cannot loop forever. Each revision costs one provider turn plus one store append.
214
+ - Bounded artifact tool calls run sequentially through `dispatchToolCall` (permission + validation + execute); their assistant call and result are persisted before the next provider request. `singleShotLoop` retains its bounded parallel worker pool and original call-order transcript behavior.
212
215
  - The loop is a plain object/factory; no class hierarchy, no background work, no extra dependencies. `LoopContext` is a single object literal of bound arrows built once per run.
213
216
  - The Synapta-free boundary is guarded by tests: `src/` imports no `synapta*` package, and the `Artifact*`/`AgentLoop*`/`LoopContext` contracts contain no `workflow`/`node`/`step` field names. Hosts supply their own schema; no host domain type is imported by `src/`.
214
217
 
@@ -7,8 +7,9 @@ The agent/session runtime adds the minimal shared SDK surface for running provid
7
7
  - `createAgent(config)`
8
8
  - `createAgentSession(config)`
9
9
  - `agent.createSession(config)`
10
- - `session.run(input, options)`
11
- - `session.prompt(input, options)`
10
+ - `session.run(input, options)` → `AgentRunResult`
11
+ - `session.prompt(input, options)` → `AgentRunResult`
12
+ - `session.stream(input, options)` → owned-run `AsyncIterable<AgentEvent>`
12
13
  - `session.compact(options?)`
13
14
  - `session.subscribe(options?)`
14
15
  - `session.abort()`
@@ -46,7 +47,11 @@ string | Message | readonly Message[]
46
47
 
47
48
  ## Outputs / response / events
48
49
 
49
- `session.subscribe(options?)` returns a live `AsyncIterable<AgentEvent>`. Subscribe before `run()` to observe that run's events. The consumer loop and `session.run()` must run concurrently (e.g. start the `for await` consumer, then `await Promise.all([consumer, session.run("Hi")])`): events are only emitted during a live run, so awaiting the subscribe loop before calling `run()` deadlocks. `SubscribeOptions.maxQueuedEvents` defaults to `1024` (minimum `1`) and caps events queued while the consumer is not awaiting `next()`. `SubscribeOptions.overflow` defaults to `"close"`; it clears queued payload events, delivers one `event_subscriber_overflow` notice to that subscriber, then closes it. `"drop_oldest"` keeps newest events; `"drop_newest"` ignores new events while full.
50
+ `session.run()` / `session.prompt()` resolve to an `AgentRunResult` with `sessionId`, `runId`, `status`, `text`, `content`, optional `message`/`usage`/`leafId`, and terminal `error`/`abortReason` when applicable. Callers may ignore the return value. Failed and aborted runs still emit their terminal events, then reject with `AgentRunError` whose `.result` carries the same shape.
51
+
52
+ `session.stream(input, options?)` subscribes first, starts exactly one run, yields only that run's events, and terminates when the run succeeds, fails, or aborts. Early consumer return aborts the owned run and releases the session. `SubscribeOptions.maxQueuedEvents` / `overflow` may be passed alongside `RunOptions`.
53
+
54
+ `session.subscribe(options?)` remains available for hosts that want a long-lived subscriber across runs. Subscribe before `run()` to observe that run's events. The consumer loop and `session.run()` must run concurrently (e.g. start the `for await` consumer, then `await Promise.all([consumer, session.run("Hi")])`): events are only emitted during a live run, so awaiting the subscribe loop before calling `run()` deadlocks. Prefer `session.stream()` when you only need one run's events. `SubscribeOptions.maxQueuedEvents` defaults to `1024` (minimum `1`) and caps events queued while the consumer is not awaiting `next()`. `SubscribeOptions.overflow` defaults to `"close"`; it clears queued payload events, delivers one `event_subscriber_overflow` notice to that subscriber, then closes it. `"drop_oldest"` keeps newest events; `"drop_newest"` ignores new events while full.
50
55
 
51
56
  For a text-only provider turn, the runtime emits:
52
57
 
@@ -101,29 +106,22 @@ const agent = createAgent({
101
106
  });
102
107
 
103
108
  const session = agent.createSession({ id: "s1" });
104
- const reader = (async () => {
105
- for await (const event of session.subscribe()) console.log(event.type);
106
- })();
109
+ const result = await session.run("Hi", { maxToolRounds: 1, compaction: { thresholdEntries: 20, keepRecentEntries: 6 }, retry: { maxAttempts: 3, baseDelayMs: 50 } });
110
+ console.log(result.text, result.usage?.totalTokens);
111
+
112
+ for await (const event of session.stream("Follow up")) console.log(event.type);
107
113
 
108
- await session.run("Hi", { maxToolRounds: 1, compaction: { thresholdEntries: 20, keepRecentEntries: 6 }, retry: { maxAttempts: 3, baseDelayMs: 50 } });
109
114
  await session.compact({ keepRecentEntries: 4 });
110
115
  const branch = await session.entries();
111
116
  await session.checkout(branch.at(-1)?.id);
112
117
  const clone = await session.clone({ id: "s2" });
113
- await reader;
114
118
  ```
115
119
 
116
120
  ## Extension and configuration notes
117
121
 
118
122
  The runtime calls `assembleProviderInput()` on every turn and uses only runtime-consumed values supplied on `AgentConfig`: `instructions`, `systemPrompt`, `inputBuilder`, `promptBuilder`, `inputLayout`, `context`, selected `skills`, active `tools`, `middleware`, `resourceLoader`, metadata, `compaction`, `retry`, and `RunOptions.model`/`systemPrompt`/`inputLayout`/`compaction`/`retry`. Contributions remain inert until a host passes selected values into the agent config.
119
123
 
120
- `AgentConfig` fields that are host-owned metadata, not runtime work:
121
-
122
- | Field | Runtime behavior |
123
- | --- | --- |
124
- | `extensions` | Preserved on `agent.config` only. `createAgent()` / `session.run()` do not call `setup()`, load packages, or auto-register contributions. Load extensions with `createExtensionKernel()` before building config. |
125
- | `settings` | Preserved on `agent.config` only. The runtime does not call `settings.get()`; hosts or provider packages read settings before passing concrete runtime options. |
126
- | `credentials` | Preserved on `agent.config` only. The runtime does not call `credentials.resolve()`; provider adapters/request policies resolve credentials at the provider edge and pass exact secret values to redaction when needed. |
124
+ `AgentConfig` no longer accepts inert `extensions`, `settings`, or `credentials` fields. Load extensions with `createExtensionKernel()` before building config; read settings in the host before passing concrete runtime options; resolve credentials at the provider edge and pass exact secret values to redaction when needed.
127
125
 
128
126
  The runtime calls `middleware.run("compaction", { context, result })` after a compaction strategy returns and before appending the standard compaction entry. Middleware can adjust the result summary/data, but the runtime still owns store append ordering and branch parent ids.
129
127
 
@@ -131,7 +129,7 @@ Provider request policy application is one ordered in-memory pass per provider t
131
129
 
132
130
  The runtime calls `middleware.run("retry", { context, decision })` after the retry policy decision and before emitting `retry_scheduled`. Middleware can stop retrying or adjust the delay. Retry wraps only the current provider turn, reuses the same assembled request, and never retries after assistant output has been emitted.
133
131
 
134
- `createAgent()` is a thin wrapper over explicit config. It does not load `AgentConfig.extensions`, scan packages, resolve credentials, read settings, call `Extension.setup()`, or consult hidden registries. External `AgentDefinition` implementations can call it from their own `create()` method:
132
+ `createAgent()` is a thin wrapper over explicit config. It does not scan packages, resolve credentials, read settings, call `Extension.setup()`, or consult hidden registries. External `AgentDefinition` implementations can call it from their own `create()` method:
135
133
 
136
134
  ```ts
137
135
  import { createAgent, createContributionRegistries } from "@arnilo/prism";
package/docs/cli-rpc.md CHANGED
@@ -2,23 +2,41 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- The `prism` bin is a thin adapter over `AgentSession`:
5
+ The `prism` bin is a thin adapter over `AgentSession` plus a tiny project scaffold:
6
6
 
7
7
  - `prism -p "prompt"`: print assistant text deltas.
8
8
  - `prism --mode json -p "prompt"`: write one normalized event envelope per line.
9
9
  - `prism --mode rpc`: read LF-delimited JSON requests from stdin and write correlated JSON responses/events to stdout.
10
+ - `prism init <dir>`: create a minimal TypeScript project with one selected provider, `.env.example`, and one offline mock test.
10
11
 
11
- It does not add a TUI, app tools, provider globals, extension discovery, resource discovery, or credential storage.
12
+ It does not add a TUI, app tools, provider globals, extension discovery, resource discovery, or credential storage. `init` uses Node standard-library filesystem APIs and checked-in templates only — no interactive prompts or template-engine dependency.
12
13
 
13
14
  ## When to use it
14
15
 
15
- Use the CLI for terminal smoke tests, scriptable JSON event streams, and simple non-Node clients that can speak newline-delimited JSON.
16
+ Use the CLI for terminal smoke tests, scriptable JSON event streams, simple non-Node clients that can speak newline-delimited JSON, and bootstrapping a tiny host project with `prism init`.
16
17
 
17
18
  Use the SDK directly when an app needs custom providers, tools, resources, credentials, trust prompts, or UI behavior.
18
19
 
19
20
  ## Inputs / request
20
21
 
21
- CLI flags:
22
+ ### `prism init`
23
+
24
+ ```bash
25
+ prism init <dir> [--provider <name>] [--with-workflows] [--with-evals] [--force]
26
+ ```
27
+
28
+ | Flag / arg | Purpose |
29
+ | --- | --- |
30
+ | `<dir>` | Destination directory (created if missing). |
31
+ | `--provider <name>` | `mock` (default), `openai`, `openrouter`, `kimi`, `zai`, `opencode-go`, or `neuralwatt`. |
32
+ | `--with-workflows` | Add `@arnilo/prism-workflows` and `src/workflows-example.ts`. |
33
+ | `--with-evals` | Add `@arnilo/prism-evals` and `src/evals-example.ts`. |
34
+ | `--force` | Overwrite generated files when the destination already exists. |
35
+ | `-h`, `--help` | Print init usage. |
36
+
37
+ Default generation installs only `@arnilo/prism` (mock provider). Selecting a real provider adds exactly one `@arnilo/prism-provider-*` package. Storage, telemetry, memory, and server packages are never added unless a later phase introduces an explicit flag for them. Rerunning without `--force` refuses non-empty destinations and existing generated files. `.env.example` contains placeholders only; `.gitignore` excludes `.env` and local stores.
38
+
39
+ ### Run/RPC CLI flags
22
40
 
23
41
  | Flag | Purpose |
24
42
  | --- | --- |
@@ -127,6 +145,11 @@ Events streamed during a run keep the original prompt request id, even when an `
127
145
  prism --provider mock --model demo -p "Hi"
128
146
  prism --provider mock --mode json -p "Hi"
129
147
  printf '{"id":"1","command":"prompt","params":{"input":"Hi"}}\n' | prism --provider mock --mode rpc
148
+
149
+ prism init my-agent
150
+ prism init my-agent --provider openai
151
+ prism init my-agent --provider openrouter --with-workflows --with-evals
152
+ cd my-agent && npm install && npm test
130
153
  ```
131
154
 
132
155
  Programmatic hosts should use the public runtime directly:
@@ -147,7 +170,9 @@ CLI/RPC are adapters over `AgentSession`. They do not scan packages, import exte
147
170
 
148
171
  RPC `command` executes only explicitly registered `CommandDefinition` values. `setModel` stores a model override for later prompt/follow-up calls. `compact`, `switchSession`, `forkSession`, `cloneSession`, and `checkout` call the existing session APIs.
149
172
 
150
- Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.start`, `workflow.status`, `workflow.list`, `workflow.cancel`, and `workflow.resume` via `createWorkflowCommands({ workflows, checkpoints, runOptions? })`. Pass the returned `CommandDefinition[]` into `runRpcServer({ commands })` the same way as observational-memory commands. Cancel aborts in-process runs through the package active-run registry; orphaned durable checkpoints still marked `running` are fail-closed to `aborted`.
173
+ Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.start`, `workflow.enqueue`, `workflow.replay`, `workflow.status`, `workflow.list`, `workflow.cancel`, and `workflow.resume` via `createWorkflowCommands({ workflows, checkpoints, runOptions? })`. Supplying an ownership-scoped `schedules` service additionally registers `schedule.create`, `schedule.list`, `schedule.pause`, `schedule.resume`, `schedule.trigger`, and `schedule.delete`. Pass the returned `CommandDefinition[]` into `runRpcServer({ commands })` the same way as observational-memory commands. Cancel aborts in-process runs through the package active-run registry; orphaned durable checkpoints still marked `running` are fail-closed to `aborted`.
174
+
175
+ Suspended workflow resume parameters are `{ workflowId, runId, decision: "approve" | "deny", input?, expectedVersion, ownership? }`. Read `expectedVersion` from `workflow.status`/`workflow.list`; stale or duplicate decisions fail checkpoint CAS before node execution. Ordinary recovery resume for failed/aborted runs remains backward-compatible without decision fields.
151
176
 
152
177
  `forkSession` creates another handle for the same `sessionId` and selected `leafId`; it no longer overwrites the parent handle in the RPC map. Keep the returned `handleId` when a UI needs to switch among sibling branches. `switchSession` accepts `handleId` (preferred), `sessionId`, or `id`; with multiple branch handles, use `handleId` to avoid ambiguity. `checkout` requires `params.leafId`, calls `AgentSession.checkout(leafId)`, and keeps the active handle id unchanged while moving that handle to the existing leaf. `messages` returns entries for the active branch path.
153
178
 
@@ -157,7 +182,10 @@ Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.s
157
182
  - No hidden provider, credential, extension, resource, config, settings, or tool globals are created.
158
183
  - No full TUI or sandbox is provided or implied.
159
184
  - JSONL is processed line by line with Node stdlib; no parser dependency, worker, watcher, or queue is added.
160
- - Unknown or malformed CLI/RPC input fails closed.
185
+ - Unknown or malformed CLI/RPC input fails closed. Workflow resume validates decision and positive `expectedVersion`; ownership remains host-selected and checkpoint-enforced.
186
+ - `prism init` refuses non-empty destinations without `--force`, keeps writes inside the destination root, and never executes downloaded code beyond the user's later `npm install`.
187
+ - Generated `.env.example` values are placeholders only; `.gitignore` excludes `.env` and local store files.
188
+ - Default generated install stays small (~27 MB with TypeScript tooling in a clean consumer install versus Mastra's measured 439 MB scaffold); unselected storage/telemetry/eval/workflow packages are omitted.
161
189
  - Branch handles (`handleId`, `sessionId`, `leafId`) are identifiers only; do not encode credentials, tokens, provider objects, or secrets into them.
162
190
  - Do not put resolved credential values, tokens, headers, or secrets in prompts, CLI flags, config, events, or docs examples.
163
191
 
@@ -171,7 +199,7 @@ Optional workflow control (from `@arnilo/prism-workflows`) registers `workflow.s
171
199
  - [Resource loading](resource-loading.md): explicit resource loading primitives.
172
200
  - [Credentials and redaction](credentials-and-redaction.md): secret redaction helpers and credential boundaries.
173
201
  - [Observational memory compaction package](compaction-observational-memory.md): optional `om:status` and `om:view` command factories for explicitly wired hosts.
174
- - [Workflows](workflows.md): optional `createWorkflowCommands()` for start/status/list/cancel/resume over the same RPC `command` seam.
202
+ - [Workflows](workflows.md): optional `createWorkflowCommands()` for direct/background/replay/status/cancel/resume and selected schedule control over the same RPC `command` seam.
175
203
 
176
204
  The CLI records flags but does not auto-load project-local resources, extensions, tools, or config. The two system/project prompt files are the exception: in print/json modes the CLI auto-loads `<workspaceRoot>/AGENTS.md` (trust-gated) and an app-supplied `SYSTEM.md` layer as `AgentConfig.systemPrompt` layers composed with `--system` (base); `--no-agents-md` / `--no-system-md` skip them and `--agents-md-file` / `--system-md-file` override the paths. The CLI does not default `globalRoot` to the user's home directory — pass it from a host adapter or use `--agents-config <path>` for the app-config bundle layout. RPC mode does not auto-read these files (the host owns the session factory). Hosts must make explicit trust and permission decisions before wiring any other local loading.
177
205
 
@@ -15,6 +15,8 @@
15
15
  | `createAllTools(cwd, options?)` | Every tool the package provides (currently identical to `createCodingTools`). |
16
16
  | `detectSupportedImageMimeType(buf)` / `detectSupportedImageMimeTypeFromFile(path)` | Magic-byte image MIME detection (PNG/JPEG/GIF/WebP/BMP) used by `read`. |
17
17
  | `DEFAULT_MAX_IMAGE_BYTES` | Default `read` image size ceiling (10 MB). |
18
+ | `DEFAULT_*` / `HARD_*` coding limit constants | Published text-scan, image, write/edit, shell timeout, display, and total-output ceilings. |
19
+ | `ReadTextOptions` / `ReadTextResult` | Bounded text-page contract required by custom `ReadOperations`. |
18
20
  | `TransformImage` / `TransformImageInput` | Types for the optional `read` `transformImage` callback. |
19
21
  | `withFileMutationQueue(path, fn)` | Per-path serialization primitive re-exported for hosts. |
20
22
 
@@ -65,7 +67,7 @@ Run a shell command and return combined stdout+stderr.
65
67
  | Field | Type | Purpose |
66
68
  | --- | --- | --- |
67
69
  | `command` | `string` | Shell command to execute (required). |
68
- | `timeout` | `number` | Timeout in **seconds** (optional; no default). |
70
+ | `timeout` | `number` | Timeout in **seconds** (optional; defaults to 600, hard maximum 3600). |
69
71
 
70
72
  **Outputs:** a `ToolResult` whose `content[0]` is a `TextContent` with the combined output. Non-zero exit is **not** a tool error: it is returned as a normal result with `[Command exited with code N]` appended to the content and `exitCode` in metadata. Timeout and abort are error results that still carry the partial output captured so far.
71
73
 
@@ -75,9 +77,11 @@ Run a shell command and return combined stdout+stderr.
75
77
  | --- | --- | --- |
76
78
  | `exitCode` | always | Process exit code, or `null` when the process was killed by timeout/abort. |
77
79
  | `truncation` | always | `TruncationResult` from the bounded output accumulator. |
78
- | `fullOutputPath?` | truncated only | Path to the spilled temp file holding the full output. |
80
+ | `fullOutputPath?` | successful and truncated only | Host-owned path to retained output. Failed/aborted/timed-out/output-limited calls remove unpublished spills. |
81
+ | `totalOutputBytes` | shell executed | Raw bytes retained, never above `maxTotalOutputBytes`. |
82
+ | `outputLimitExceeded` / `outputStorageFailed` | shell executed | Attributable resource failure flags. |
79
83
 
80
- Shell resolution honors `options.shellPath` → `SHELL` env → `/bin/bash` → `sh`, and the process group is killed on timeout/abort (`process.kill(-pid)` on Unix, `taskkill /F /T` on Windows).
84
+ Shell resolution honors `options.shellPath` → `SHELL` env → `/bin/bash` → `sh`. The process group is killed on timeout, caller abort, spill failure, or total-output overflow (`process.kill(-pid)` on Unix, `taskkill /F /T` on Windows). Combined output defaults to a 64 MiB total cap (1 GiB hard cap). Spill files use random exclusive creation and Unix mode `0600`; hosts own and must delete a successful result's `fullOutputPath` after consumption.
81
85
 
82
86
  ### `read`
83
87
 
@@ -91,7 +95,7 @@ Read a text or image file.
91
95
  | `offset` | `number` | Line to start reading from (1-indexed). |
92
96
  | `limit` | `number` | Maximum number of lines to read. |
93
97
 
94
- **Outputs:** text files become a single `TextContent`, truncated to `maxLines`/`maxBytes` (defaults 2000 lines / 50 KB) with a `Use offset=N to continue` footer when more remains. Image files (PNG/JPEG/GIF/WebP/BMP by **magic bytes**, not extension) become `[TextContent note, ImageContent]` with base64 `data` and `mimeType`. Oversize images are rejected by `stat` (when available) or `buffer.length` against `maxImageBytes` (default 10 MB) before base64 encoding. An optional `transformImage` callback lets hosts resize or re-encode images without adding image-processing dependencies to the base package. Read failures (missing file, offset beyond end, oversize image, abort) are error results.
98
+ **Outputs:** text files are scanned incrementally until one requested page, `maxLines`/`maxBytes`, EOF, or `maxScanBytes` (default 64 MiB scanned per call; 1 GiB hard cap). The default path never loads the complete file and returns a `Use offset=N to continue` footer when more remains. Exact total line count is reported only when EOF was already reached in the bounded scan. Image files (PNG/JPEG/GIF/WebP/BMP by **magic bytes**, not extension) become `[TextContent note, ImageContent]` with base64 `data` and `mimeType`. Oversize images are rejected by `stat` (when available) or `buffer.length` against `maxImageBytes` (default 10 MB) before base64 encoding. An optional `transformImage` callback lets hosts resize or re-encode images without adding image-processing dependencies to the base package. Read failures (missing file, offset beyond end, oversize image, abort) are error results.
95
99
 
96
100
  `read` tool options (via `createReadTool(cwd, options)` or `ToolsOptions.read`):
97
101
 
@@ -100,8 +104,9 @@ Read a text or image file.
100
104
  | `maxImageBytes` | `DEFAULT_MAX_IMAGE_BYTES` (10 MB) | Reject image reads larger than this many bytes. |
101
105
  | `transformImage` | — | Host callback `( { buffer, mimeType } ) => Promise<Buffer>` run after read, before base64. |
102
106
  | `autoResizeImages` | — | **Deprecated.** Ignored unless `transformImage` is also set (use `transformImage` instead). |
103
- | `maxLines` / `maxBytes` | 2000 / 50 KB | Text head truncation limits. |
104
- | `operations` | local fs | Pluggable `ReadOperations` backend. |
107
+ | `maxLines` / `maxBytes` | 2000 / 50 KiB | Text page display limits (hard: 100,000 / 1 MiB). |
108
+ | `maxScanBytes` | 64 MiB | Raw bytes scanned to reach one page (hard: 1 GiB). |
109
+ | `operations` | local fs | Pluggable bounded `ReadOperations` backend. |
105
110
  | `executionPolicy` | — | Structured pre-execution policy (see [Coding security](coding-security.md)). |
106
111
 
107
112
  ```ts
@@ -133,7 +138,7 @@ Create or overwrite a file, creating parent directories as needed.
133
138
  | `path` | `string` | Path to the file to write (relative or absolute). Required. |
134
139
  | `content` | `string` | Content to write (empty string creates an empty file). Required. |
135
140
 
136
- **Outputs:** a `TextContent` confirmation naming the **absolute path** with UTF-8 byte and line counts (e.g. `Successfully wrote 42 bytes (3 lines) to /abs/path.txt`). Write failures and abort are error results. Empty `content` is valid.
141
+ **Outputs:** a `TextContent` confirmation naming the **absolute path** with UTF-8 byte and line counts (e.g. `Successfully wrote 42 bytes (3 lines) to /abs/path.txt`). `maxInputBytes` defaults to 8 MiB (64 MiB hard cap); oversized UTF-8 input fails before policy evaluation, directory creation, or write. Write failures and abort are error results. Empty `content` is valid.
137
142
 
138
143
  `write` result `metadata`: `{ bytes, lines, path }` (absolute path). Concurrent writes to the same path serialize through `withFileMutationQueue`; writes to different paths run in parallel.
139
144
 
@@ -148,7 +153,7 @@ Precise text replacement in an existing file via exact-then-fuzzy matching.
148
153
  | `path` | `string` | Path to the file to edit. Required. |
149
154
  | `edits` | `Array<{ oldText: string, newText: string }>` | Targeted replacements, each matched against the **original** file (not incrementally). No overlapping/nested edits. Required, non-empty. |
150
155
 
151
- Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse). A BOM is stripped before matching and re-prepended on write; original line endings are restored.
156
+ Each `edits[].oldText` must match a unique, non-overlapping region of the original file. Matching is exact first, then fuzzy (unicode normalization / whitespace collapse). A BOM is stripped before matching and re-prepended on write; original line endings are restored. Defaults reject targets over 8 MiB, aggregate old/new UTF-8 input over 2 MiB, or more than 100 edits (hard caps: 64 MiB, 16 MiB, and 1,000). Stat and bounded read checks run before matching or mutation.
152
157
 
153
158
  **Outputs:** a `TextContent` confirmation (`Successfully replaced N block(s) in {path}.`) plus `metadata`. Any failure — missing/unreadable file, no match, duplicate (non-unique) match, overlap, empty `oldText`, no-op edit, or abort — is an error result, and the file is left **unchanged** (the match runs before the write).
154
159
 
@@ -156,7 +161,7 @@ Each `edits[].oldText` must match a unique, non-overlapping region of the origin
156
161
 
157
162
  ## Outputs / response / events
158
163
 
159
- Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`. Mutating tools (`shell` with same cwd, `write`, `edit`) serialize per realpath through `withFileMutationQueue` so concurrent calls targeting one file do not interleave. The package emits no events of its own; hosts observe tool execution through the normal Prism `AgentEvent` stream via `dispatchToolCall`.
164
+ Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`. `write` and `edit` serialize per realpath through `withFileMutationQueue` so concurrent calls targeting one file do not interleave. `shell` is marked `exclusive`; tool dispatch serializes it at the turn level. The package emits no events of its own; hosts observe tool execution through the normal Prism `AgentEvent` stream via `dispatchToolCall`.
160
165
 
161
166
  ## Request/response example
162
167
 
@@ -208,6 +213,8 @@ const shell = createShellTool("/repo", {
208
213
  shellPath: "/bin/bash",
209
214
  commandPrefix: "set -euo pipefail",
210
215
  maxLines: 500,
216
+ timeout: 600,
217
+ maxTotalOutputBytes: 64 * 1024 * 1024,
211
218
  });
212
219
 
213
220
  const remoteWrite = createWriteTool("/repo", {
@@ -220,20 +227,34 @@ const remoteWrite = createWriteTool("/repo", {
220
227
 
221
228
  ## Extension and configuration notes
222
229
 
223
- - **Pluggable operation backends.** Every tool accepts an `operations` seam so a host can delegate to a remote system (e.g. SSH) while keeping the tool's matching/serialization behavior: `BashOperations` (`shell`), `ReadOperations` (`read`), `WriteOperations` (`write`), `EditOperations` (`edit`).
224
- - **Per-tool options.** `ShellToolOptions` (`shellPath`, `commandPrefix`, `maxLines`, `maxBytes`, `tempFilePrefix`, `operations`, `spawnHook`, `executionPolicy`); `ReadToolOptions` (`operations`, `maxImageBytes`, `transformImage`, `maxLines`, `maxBytes`, `executionPolicy`; `autoResizeImages` deprecated); `WriteToolOptions` (`operations`, `executionPolicy`); `EditToolOptions` (`operations`, `executionPolicy`).
225
- - **Aggregator options.** `ToolsOptions` (`{ shell?, read?, write?, edit? }`) threads each sub-object to the matching tool.
230
+ - **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
231
+ - **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`. Invalid/non-finite/unsafe/above-hard-cap values throw during tool construction; request `timeout` errors before spawn.
232
+ - **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override.
226
233
  - **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
227
234
  - No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
228
235
 
229
236
  ## Security and performance notes
230
237
 
231
- - **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
238
+ - **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
232
239
  - **Non-zero exit is not an error.** A failing command is a normal `shell` result (exit code in metadata); only timeout/abort/spawn failures are error results. Do not assume `error == undefined` means the command succeeded.
233
- - **Bounded output.** `shell`/`read` accumulate output into a rolling tail bounded by `maxLines`/`maxBytes`; oversized output spills to a temp file (`fullOutputPath`), so memory use is bounded regardless of command output size.
240
+ - **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
234
241
  - **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
235
242
  - **Bounded image reads.** `read` rejects images over `maxImageBytes` (default 10 MB) by `stat` before read when possible; MIME is detected from magic bytes only. Optional `transformImage` is host-owned — the base package has no image-processing dependency.
236
243
 
244
+ ### Resource-limit defaults and hard caps
245
+
246
+ | Boundary | Default | Hard cap | Failure point |
247
+ | --- | ---: | ---: | --- |
248
+ | Display lines / bytes | 2,000 / 50 KiB | 100,000 / 1 MiB | tool construction |
249
+ | Text scan per read | 64 MiB | 1 GiB | bounded scan before more input is retained |
250
+ | Image | 10,000,000 bytes | 32 MiB | stat and bounded read before base64/transform result use |
251
+ | Write UTF-8 input | 8 MiB | 64 MiB | before policy/filesystem mutation |
252
+ | Edit target / input / count | 8 MiB / 2 MiB / 100 | 64 MiB / 16 MiB / 1,000 | before target read/matching/write |
253
+ | Shell wall time | 600 seconds | 3,600 seconds | process-tree kill |
254
+ | Shell total stdout+stderr | 64 MiB | 1 GiB | process-tree kill; spill removal |
255
+
256
+ Every configurable value is a positive safe integer; Prism rejects rather than clamps invalid values. Limits control resources, not authority: they do not replace root containment, approval, validation, or a sandbox.
257
+
237
258
  ## Related APIs
238
259
 
239
260
  - [Tools](tools.md): the host-owned tool harness — `createToolRegistry`, `dispatchToolCall`, filtering, and the `ToolDefinition` contract these factories satisfy.
@@ -34,9 +34,11 @@ Prism does **not** claim OS-level isolation unless the host provides a sandbox a
34
34
  | `approvalCacheScope` | `"none"` | Optional `run` or `session` decision cache scope. |
35
35
  | `approvalTimeoutMs` | `30000` | Bound approval wait; caller abort also cancels it. |
36
36
 
37
+ `run` caching keys decisions by the tool execution context's `runId`; `session` uses `sessionId`. Coding tools pass both identities to the policy. A missing/empty identity disables caching for that check rather than creating a global bucket. Identical actions in different runs/sessions never share approvals or denials.
38
+
37
39
  ## Outputs / response / events
38
40
 
39
- `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations` and never grant policy approval themselves.
41
+ `createCodingApprovalPolicy()` returns an `ExecutionPolicy`. Allowed checks return `ExecutionDecision { allowed: true }`; denied checks include a stable reason; shell decisions set `exclusive: true`. Sandbox adapters return coding-agent-compatible `BashOperations`, receive `onData(Buffer)` for ordered stdout/stderr forwarding through the shell tool's existing bounded accumulator, and never grant policy approval themselves.
40
42
 
41
43
  ## Request/response example
42
44
 
@@ -70,11 +72,13 @@ const tools = createCodingTools(workspaceRoot, {
70
72
 
71
73
  ## Extension and configuration notes
72
74
 
73
- Policies are ordinary host values: attach one globally through `createCodingTools()` or per tool. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers. Use run-scoped approval caching unless a wider host identity/lifecycle is explicit.
75
+ Policies are ordinary host values: attach one globally through `createCodingTools()`/`createReadOnlyTools()` or per tool. A per-tool policy overrides the shared policy. `SandboxAdapter` is replaceable and host-owned; approval policy and sandboxing are separate layers.
76
+
77
+ Callback approval remains process-local. For approval that must survive restart, wrap the action in an opted-in workflow `toolNode({ approval: { reason, data?, resumeSchema? } })`. The workflow persists `suspended` state before any tool side effect. After explicit approve, it recomputes the action and invokes this package's current `ExecutionPolicy`; durable approval never populates or bypasses the process-local approval cache. Adapters should emit chunks through `request.onData` as they arrive and honor `request.signal`/`request.timeout`; buffering is unnecessary. Coding-agent composes caller abort with its total-output controller, so ignoring the supplied signal defeats process termination even though Prism stops retaining output at the cap. Default caching is `none`; use run-scoped caching only when repeated approval within one run is desired, and session scope only when that wider lifecycle is intentional.
74
78
 
75
79
  ## Security and performance notes
76
80
 
77
- Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
81
+ Containment resolves symlinks and rejects paths outside roots. Command rules are not a shell parser; shell metacharacters require approval. Approval waits and subprocess execution honor abort/timeouts. Coding-agent resource ceilings independently bound text scans, image/edit target reads, write/edit payloads, edit counts, shell wall time, and retained/spilled output. Those ceilings reduce exhaustion risk but do not grant path/command authority or make an unsandboxed shell safe. Durable workflow denial/cancellation is terminal and attributable; approved resume still fails if roots, command rules, read-only mode, or other policy changed while suspended. Cache keys are fixed-size SHA-256 digests of selected identity plus action shape; caches remain process-local, retain at most 1,000 decisions with oldest-entry eviction, and have no default/global mode. Path checks and cache lookup are local; sandbox latency belongs to the supplied adapter.
78
82
 
79
83
  ## Related APIs
80
84