@arnilo/prism 0.0.7 → 0.0.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +50 -2
- package/README.md +3 -1
- package/dist/agent-loops.js +14 -8
- package/dist/agents.js +37 -3
- package/dist/contracts.d.ts +17 -0
- package/dist/index.d.ts +4 -2
- package/dist/index.js +3 -2
- package/dist/provider-events.d.ts +2 -0
- package/dist/provider-events.js +21 -13
- package/dist/providers/openai-compatible.js +8 -5
- package/dist/providers/transport.d.ts +10 -1
- package/dist/providers/transport.js +24 -8
- package/dist/run-ledger.d.ts +21 -0
- package/dist/run-ledger.js +115 -0
- package/dist/tools.js +2 -0
- package/docs/a2a.md +61 -42
- package/docs/agent-events.md +5 -4
- package/docs/agent-loops.md +2 -2
- package/docs/agent-session-runtime.md +1 -0
- package/docs/browser-automation.md +124 -0
- package/docs/coding-agent-tools.md +111 -14
- package/docs/coding-security.md +84 -11
- package/docs/credential-storage.md +9 -0
- package/docs/database-persistence.md +1 -1
- package/docs/evaluations.md +38 -4
- package/docs/guardrails.md +3 -2
- package/docs/host-security.md +29 -4
- package/docs/index.md +20 -15
- package/docs/mcp-tools.md +29 -5
- package/docs/migration.md +104 -0
- package/docs/observability.md +26 -14
- package/docs/performance.md +54 -0
- package/docs/postgres-persistence.md +1 -0
- package/docs/provider-conformance.md +1 -1
- package/docs/provider-primitives.md +7 -1
- package/docs/providers/kimi.md +16 -2
- package/docs/providers/opencode-go.md +43 -2
- package/docs/release-and-install.md +117 -62
- package/docs/resource-loading.md +4 -0
- package/docs/review-coverage-2026-07-19-phase-3.md +174 -0
- package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
- package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
- package/docs/run-ledger-conformance.md +1 -0
- package/docs/runs-and-usage.md +17 -2
- package/docs/sqlite-persistence.md +1 -0
- package/docs/structured-output.md +2 -2
- package/docs/supervisors.md +2 -2
- package/docs/tools.md +5 -1
- package/docs/web-tools.md +78 -0
- package/docs/workflows.md +2 -0
- package/package.json +6 -4
package/docs/a2a.md
CHANGED
|
@@ -1,75 +1,94 @@
|
|
|
1
|
-
# A2A interoperability
|
|
1
|
+
# A2A 1.0 interoperability
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-supervisor` implements
|
|
5
|
+
`@arnilo/prism-supervisor` implements bounded A2A 1.0 over the JSON-RPC/HTTPS binding. Supported operations: `SendMessage`, `SendStreamingMessage`, `GetTask`, `ListTasks`, `CancelTask`, `SubscribeToTask`, push-notification-config create/get/list/delete, and `GetExtendedAgentCard`. Agent Cards retain explicit ES256 verification. gRPC, HTTP+JSON, discovery registries, automatic JWK/OAuth fetching, and an internal task worker/store are absent.
|
|
6
6
|
|
|
7
7
|
## When to use it
|
|
8
8
|
|
|
9
|
-
Use it to expose
|
|
9
|
+
Use it to expose a selected Prism agent or host-owned durable agent/workflow lifecycle to known A2A peers. Use direct `exposure` for backward-compatible text invocation. Supply `tasks` for durable/rich/reconnect operations and `push` only when host persistence and webhook delivery policy already exist.
|
|
10
10
|
|
|
11
11
|
## Inputs / request
|
|
12
12
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
13
|
+
```ts
|
|
14
|
+
const handler = createA2AHandler({
|
|
15
|
+
card,
|
|
16
|
+
exposure: { sessionFactory }, // text fallback
|
|
17
|
+
authorize: authenticateEveryOperation,
|
|
18
|
+
tasks: durableTaskAdapter, // host-owned start/get/list/cancel/subscribe
|
|
19
|
+
push: pushConfigAdapter, // host-owned config persistence/delivery integration
|
|
20
|
+
parts: {
|
|
21
|
+
allowRaw: true,
|
|
22
|
+
allowData: true,
|
|
23
|
+
allowUrl: true,
|
|
24
|
+
validateUrl: validatePinnedPublicHttpsUrl, // validation only; never fetched
|
|
25
|
+
},
|
|
26
|
+
});
|
|
27
|
+
```
|
|
21
28
|
|
|
22
|
-
|
|
29
|
+
`A2ATaskLifecycle` receives validated messages, exact `A2AAuthorization`, abort signals, bounded pagination, and reconnect cursor. Adapter must map existing durable agent/workflow/checkpoint/persistence operations; Prism creates no worker, queue, task map, or database table. Unknown-owner task/config lookups return `undefined`, producing non-disclosing `TaskNotFoundError` (`-32001`). Missing task/push capability returns `UnsupportedOperationError` (`-32004`).
|
|
23
30
|
|
|
24
|
-
|
|
31
|
+
`A2APart` is an exact one-of:
|
|
25
32
|
|
|
26
|
-
|
|
33
|
+
| Part | Default | Rule |
|
|
34
|
+
| --- | --- | --- |
|
|
35
|
+
| `{ text }` | enabled | bounded UTF-8 text |
|
|
36
|
+
| `{ raw, mediaType?, filename? }` | disabled | strict base64 and decoded-byte cap |
|
|
37
|
+
| `{ data }` | disabled | bounded finite JSON, depth 64/properties 10,000 |
|
|
38
|
+
| `{ url, mediaType?, filename? }` | disabled | credential/fragment-free HTTPS plus required host URL policy; never dereferenced |
|
|
27
39
|
|
|
28
|
-
|
|
29
|
-
{"jsonrpc":"2.0","id":1,"method":"SendMessage","params":{"message":{"role":"user","messageId":"m1","parts":[{"text":"Check sources"}]}}}
|
|
30
|
-
```
|
|
40
|
+
Parts, messages, artifacts, histories, metadata, and aggregate responses are untrusted. Rich content remains in A2A task/message/artifact contracts for host mapping; it is never promoted to system instructions or automatically loaded as a Prism resource.
|
|
31
41
|
|
|
32
42
|
## Implementation example
|
|
33
43
|
|
|
34
44
|
```ts
|
|
35
|
-
import { createA2AClient, createA2AHandler, verifyA2AAgentCard } from "@arnilo/prism-supervisor";
|
|
36
|
-
|
|
37
|
-
const handler = createA2AHandler({
|
|
38
|
-
card,
|
|
39
|
-
exposure: { sessionFactory: ({ ownership }) => agent.createSession({ metadata: ownership }) },
|
|
40
|
-
authorize: ({ request }) => authenticate(request),
|
|
41
|
-
});
|
|
42
|
-
|
|
43
45
|
const client = createA2AClient({
|
|
44
46
|
endpoint: "https://agent.example/a2a/v1",
|
|
45
47
|
allowedOrigins: ["https://agent.example"],
|
|
46
|
-
authorize:
|
|
47
|
-
verifyCard: (
|
|
48
|
+
authorize: ownedAuthHeaders,
|
|
49
|
+
verifyCard: (card) => verifyA2AAgentCard(card, { publicKey, keyId: "agent-key" }),
|
|
48
50
|
});
|
|
51
|
+
const task = await client.getTask("task-1");
|
|
52
|
+
for await (const event of client.subscribeToTask(task.id, { afterEventId: savedCursor })) persistCursor(event.eventId);
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
## Outputs / response / events
|
|
56
|
+
|
|
57
|
+
Streams use ordered SSE frames with `id:` and JSON-RPC `result` containing one `A2ATaskEvent`: full `task`, `statusUpdate`, or `artifactUpdate`. `SubscribeToTask({ id, afterEventId })` passes cursor to durable adapter for authorized bounded replay. Duplicate event IDs are rejected/server-bounded; client de-duplicates repeated IDs. Terminal, `INPUT_REQUIRED`, and `AUTH_REQUIRED` states close streams. String-oriented `client.stream()` reports interrupted states as `ERR_PRISM_A2A_INTERRUPTED`; task APIs preserve status for continuation.
|
|
49
58
|
|
|
50
|
-
|
|
59
|
+
Client APIs:
|
|
60
|
+
|
|
61
|
+
- `send()` / `stream()` preserve text-to-`AgentRunResult` compatibility.
|
|
62
|
+
- `sendMessage()` returns rich/durable `A2ATask`.
|
|
63
|
+
- `getTask()`, `listTasks()`, `cancelTask()`, `subscribeToTask()` operate on durable tasks.
|
|
64
|
+
- `createPushConfig()`, `getPushConfig()`, `listPushConfigs()`, `deletePushConfig()` expose declared push config operations.
|
|
65
|
+
|
|
66
|
+
Every protocol request sends/negotiates `A2A-Version: 1.0`. Client endpoint/card URLs require exact allow-listed HTTPS and `redirect: "error"`. Cards are parsed then optionally verified against host-pinned keys; no key URL is fetched.
|
|
67
|
+
|
|
68
|
+
## Request/response example
|
|
69
|
+
|
|
70
|
+
```json
|
|
71
|
+
{"jsonrpc":"2.0","id":1,"method":"SubscribeToTask","params":{"id":"task-1","afterEventId":"event-42"}}
|
|
51
72
|
```
|
|
52
73
|
|
|
53
74
|
## Extension and configuration notes
|
|
54
75
|
|
|
55
|
-
|
|
76
|
+
Handler requires `card.capabilities.pushNotifications` to exactly match supplied `push`; mismatch fails construction, preserving signed-card integrity and preventing false capability claims. Streaming remains available for direct text invocation. Push adapter owns exact-owner persistence, signing/auth credentials, and network transport. Host explicitly calls `deliverA2APushEvent()` from its durable update path; helper bounds event, timeout (10s default/60s hard), attempts (1 default/3 hard), and passes stable event ID as idempotency key to host `A2APushDelivery`. It starts no hidden sender and performs no network itself. Config handling validates IDs/count/bytes and requires same explicit URL policy used for URL parts. Returned push configs omit token and authentication credentials.
|
|
56
77
|
|
|
57
|
-
|
|
78
|
+
Defaults/hard caps include: request 64 KiB/1 MiB; response 1/8 MiB; event 64 KiB/1 MiB; stream 10/64 MiB and 10k/100k events; replay 1k/10k events; concurrency 16/256; timeout 120s/30m; IDs 256/4096 B; parts 32/256; part/raw 1/8 MiB; data 256 KiB/4 MiB; artifacts 32/256; history/page 100/1000; cursor 4/16 KiB; push configs 10/100. Hosts may narrow limits.
|
|
58
79
|
|
|
59
80
|
## Security and performance notes
|
|
60
81
|
|
|
61
|
-
-
|
|
62
|
-
-
|
|
63
|
-
-
|
|
64
|
-
-
|
|
65
|
-
-
|
|
66
|
-
-
|
|
67
|
-
- Credentials remain in the client auth callback or server authorizer and never enter cards, messages, events, or metrics.
|
|
68
|
-
- Offline conformance is authoritative. Live endpoints are optional operator smoke tests.
|
|
82
|
+
- Authorize every operation; lifecycle/push adapters enforce exact owner again at durable storage boundary. Missing and foreign tasks/configs share `-32001`.
|
|
83
|
+
- URL policy must reject private, loopback, link-local, rebound, redirected, or otherwise disallowed destinations. Package never fetches file URLs. Host push delivery must repeat equivalent checks for every attempt/redirect and process event IDs idempotently.
|
|
84
|
+
- Push token/auth credentials are accepted only into host adapter input and removed from protocol reads/responses. Keep them out of task parts, events, telemetry, ledgers, and errors.
|
|
85
|
+
- Known-secret redaction applies before handler JSON/SSE output. Client redacts mapped text/errors. Raw/data/url content remains explicitly untrusted.
|
|
86
|
+
- Canceled/closed streams abort adapter signal, return iterator, clear timeout, and release concurrency slot. Task/push durability and replay retention belong to host adapter and must remain finite.
|
|
87
|
+
- Default tests use in-memory lifecycle/fake fetch only; no public network.
|
|
69
88
|
|
|
70
89
|
## Related APIs
|
|
71
90
|
|
|
72
|
-
- [Supervisor delegation](supervisors.md)
|
|
73
|
-
- [
|
|
74
|
-
- [
|
|
75
|
-
- [
|
|
91
|
+
- [Supervisor delegation](supervisors.md)
|
|
92
|
+
- [Agent/session runtime](agent-session-runtime.md)
|
|
93
|
+
- [Workflows](workflows.md)
|
|
94
|
+
- [Host security](host-security.md)
|
package/docs/agent-events.md
CHANGED
|
@@ -67,7 +67,7 @@ Agent / turn / message events:
|
|
|
67
67
|
| `message_started` / `message_finished` | `sessionId`, `runId`, `message: Message` |
|
|
68
68
|
| `message_delta` | `sessionId`, `runId`, `content: ContentBlock` (`tool_call_delta` fragments may appear here for live UI streaming; stored messages use final `tool_call` blocks) |
|
|
69
69
|
|
|
70
|
-
`message_delta.content.type === "tool_call_delta"` carries `{ index, id?, name?, argumentsText? }`. Treat it as a streaming fragment. The runtime reconstructs and persists a final `tool_call` before executing tools.
|
|
70
|
+
`message_delta.content.type === "tool_call_delta"` carries `{ index, id?, name?, argumentsText? }`. Treat it as a streaming fragment. The runtime reconstructs and persists a final `tool_call` before executing tools. Deltas missing `id`/`name` at stream end fail the provider turn with `ErrorInfo.code: "incomplete_delta"` (typed `ProviderTransportError`); they never throw a bare `Error`. Malformed JSON with id+name present recovers as a blocked tool result (`invalid_json_arguments`) instead.
|
|
71
71
|
|
|
72
72
|
Tool execution events:
|
|
73
73
|
|
|
@@ -112,7 +112,7 @@ Artifact validation/refinement events (emitted only by `generateValidateReviseLo
|
|
|
112
112
|
| `artifact_validation_finished` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` |
|
|
113
113
|
| `artifact_revision_started` | `sessionId`, `runId`, `turn`, `attempt`, `failure: ArtifactValidation` |
|
|
114
114
|
| `artifact_finished` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (loop ended successfully) |
|
|
115
|
-
| `artifact_failed` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (candidate budget exhausted, or `result.metadata.reason === "
|
|
115
|
+
| `artifact_failed` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (candidate budget exhausted, `result.metadata.reason === "tool_round_limit"`, or `result.metadata.reason === "parse_error"` when the budget was consumed by artifact parse failures) |
|
|
116
116
|
|
|
117
117
|
### Artifact event ordering
|
|
118
118
|
|
|
@@ -127,7 +127,8 @@ turn_started → message_started → message_delta* → message_finished → tur
|
|
|
127
127
|
With opt-in `toolCalls: "bounded"`, a provider turn containing calls emits its normal assistant envelope followed by existing `tool_execution_*` events and matching persisted tool results; it emits no validation event and the next provider turn consumes that transcript. A post-`maxToolRounds` call emits terminal `artifact_failed` directly after `turn_finished` and has no tool execution event.
|
|
128
128
|
|
|
129
129
|
- `attempt` is 1-indexed per call-free validation candidate. It can differ from provider `turn` when bounded tool calls occur.
|
|
130
|
-
-
|
|
130
|
+
- Empty/whitespace-only call-free text (including thinking-only content) emits `artifact_validation_*` with `metadata.reason: "parse_error"` before any host parser runs.
|
|
131
|
+
- Single-shot runs emit zero artifact events. Session runs with `generate-validate-revise` require `artifact_finished` to resolve `succeeded`.
|
|
131
132
|
- **Validation failure triggering a revision is recoverable and never an `error`.** Terminal candidate-budget or `tool_round_limit` exhaustion emits `artifact_failed`; real failures remain on the `error` channel.
|
|
132
133
|
|
|
133
134
|
## Request/response example
|
|
@@ -208,6 +209,6 @@ for await (const event of session.stream("draft", { loop: { strategy: "generate-
|
|
|
208
209
|
- [Agent loops](agent-loops.md): `singleShotLoop` and `generateValidateReviseLoop` emit the artifact events.
|
|
209
210
|
- [Structured output](structured-output.md): `ArtifactValidation` shape threaded through parser/validator/repairer.
|
|
210
211
|
- [Public contracts](public-contracts.md): full `AgentEvent` union and `ArtifactValidation` contract.
|
|
211
|
-
- [Observability](observability.md): `ProviderTurnMetadata
|
|
212
|
+
- [Observability](observability.md): `ProviderTurnMetadata`; optional adapter builds one parented GenAI span tree from metadata-only lifecycle events and ignores message/progress deltas.
|
|
212
213
|
- [Tools](tools.md): `tool_execution_*` variants.
|
|
213
214
|
- [Compaction and retry policies](compaction-and-retry.md): `compaction_*` and `retry_scheduled` variants.
|
package/docs/agent-loops.md
CHANGED
|
@@ -89,7 +89,7 @@ Host callback contracts (all generic over host `T`):
|
|
|
89
89
|
|
|
90
90
|
| Contract | Shape |
|
|
91
91
|
| --- | --- |
|
|
92
|
-
| `ArtifactParser<T>` | `(text: string, ctx: ArtifactContext) => ArtifactParseResult<T> \| Promise<...>` — parse assistant text to a typed value. |
|
|
92
|
+
| `ArtifactParser<T>` | `(text: string, ctx: ArtifactContext) => ArtifactParseResult<T> \| Promise<...>` — parse assistant text to a typed value. Empty/whitespace-only call-free text is rejected before the parser (`metadata.reason: "parse_error"`, message `no artifact text in model output`) so thinking-only/reasoning-only turns cannot succeed via the identity parser. A parse failure (`ok: false` or missing `value`) consumes revision budget exactly like a validation failure: the repairer receives `value: undefined` plus a synthetic failure (`errors[0].message` = the parse error, `metadata.reason: "parse_error"`), and budget exhaustion ends with terminal `artifact_failed`. |
|
|
93
93
|
| `ArtifactValidator<T>` | `(value: T, ctx: ArtifactContext) => ArtifactValidation \| Promise<...>` — return `{ ok: true }` or `{ ok: false, errors }`. |
|
|
94
94
|
| `ArtifactRepairer<T>` | `(value: T \| undefined, failure: ArtifactValidation, ctx: ArtifactContext) => AgentInput \| Promise<...>` — build the revision follow-up input. |
|
|
95
95
|
| `ArtifactValidation` | `{ ok: boolean; errors?: readonly { path?: string; message: string }[]; metadata?: ... }`. |
|
|
@@ -119,7 +119,7 @@ Host callback contracts (all generic over host `T`):
|
|
|
119
119
|
|
|
120
120
|
Events during a loop run are the existing `AgentEvent`s (`turn_started`, `message_started`, `message_delta`, `message_finished`, `turn_finished`, tool-execution events when the loop dispatches tools, `error` on real failures). Both built-in loops emit `turn_started` before each provider turn, `message_finished` for every assistant draft, and `turn_finished` after the assistant draft is appended. First-turn input is appended to live history once, matching the already-persisted user message.
|
|
121
121
|
|
|
122
|
-
Validation-failure-triggering-a-revision is **not** an `error` event — it is recoverable, like `tool_execution_blocked`. In bounded artifact mode, a tool-calling provider response emits normal assistant/tool lifecycle events, skips artifact parsing/validation, then the next turn sees its persisted result. `generateValidateReviseLoop` emits artifact events only for call-free candidates: `artifact_validation_started` → `artifact_validation_finished` → (`artifact_revision_started`)* → `artifact_finished` | `artifact_failed`. A request beyond `maxToolRounds` executes nothing and emits terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"`; see [Agent events § Artifact event ordering](agent-events.md#artifact-event-ordering). `singleShotLoop` emits zero artifact events. Real failures stay on the `error` channel.
|
|
122
|
+
Validation-failure-triggering-a-revision is **not** an `error` event — it is recoverable, like `tool_execution_blocked`. In bounded artifact mode, a tool-calling provider response emits normal assistant/tool lifecycle events, skips artifact parsing/validation, then the next turn sees its persisted result. `generateValidateReviseLoop` emits artifact events only for call-free candidates: `artifact_validation_started` → `artifact_validation_finished` → (`artifact_revision_started`)* → `artifact_finished` | `artifact_failed`. A request beyond `maxToolRounds` executes nothing and emits terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"`; see [Agent events § Artifact event ordering](agent-events.md#artifact-event-ordering). `singleShotLoop` emits zero artifact events. Real failures stay on the `error` channel. Session runs using `generate-validate-revise` resolve `succeeded` only after `artifact_finished`; terminal `artifact_failed` (including empty/thinking-only parse exhaustion) fails the run with `AgentRunError` (`error.code` from `result.metadata.reason`, e.g. `parse_error`).
|
|
123
123
|
|
|
124
124
|
A loop has no path to credentials, provider objects, or unredacted secrets. `LoopContext.generate` receives the already-policy-applied, middleware-run, redacted request; `LoopContext.emit` runs through `redactAgentEvent` with the active `SecretRedactor`.
|
|
125
125
|
|
|
@@ -204,6 +204,7 @@ Per-run options may narrow `limits` and append `guardrails`; they cannot replace
|
|
|
204
204
|
- [Middleware hooks](middleware-hooks.md): hooks that configured assembly/runtime can run.
|
|
205
205
|
- [CLI/RPC](cli-rpc.md): terminal and JSONL adapters over this runtime.
|
|
206
206
|
- [Workflows](workflows.md): optional DAG orchestration that calls `AgentSession.run()` for agent nodes.
|
|
207
|
+
- [A2A interoperability](a2a.md): direct text exposure calls `AgentSession.run()`; durable/rich/reconnect behavior uses host `A2ATaskLifecycle` over existing checkpoints/persistence, never an in-memory runtime cache.
|
|
207
208
|
|
|
208
209
|
`AgentConfig.loop` and `RunOptions.loop` select a replaceable per-run control loop (`singleShotLoop` default, or `generate-validate-revise` with host callbacks); see [Agent loops](agent-loops.md). `RunOptions.loop` wins over `AgentConfig.loop`. Built-in loops emit the same normal turn/message envelope around provider turns, and both add the first run input to live history once after the first provider turn so later turns see the same transcript shape.
|
|
209
210
|
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
# Browser automation
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-browser` exposes exactly four exclusive model-facing tools—`browser_open`, `browser_snapshot`, `browser_act`, and `browser_close`—over a host-supplied Playwright `Browser`. Prism creates one non-persistent `BrowserContext` per run, serializes actions, returns bounded AI-mode accessibility snapshots with snapshot-scoped refs, enforces egress/side-effect/upload/download/screenshot policy, and closes context/pages/listeners/quarantined downloads on close, abort, or manager disposal.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
Use when an agent must interact with JavaScript-heavy or authenticated pages that search/fetch cannot cover. Prefer `@arnilo/prism-web-tools` for ordinary public retrieval. Do not use this package as a browser launcher, MCP proxy, visual planner, CDP console, or persistent profile manager.
|
|
10
|
+
|
|
11
|
+
## Inputs / request
|
|
12
|
+
|
|
13
|
+
| Tool | Model-visible input | Host-only construction input |
|
|
14
|
+
| --- | --- | --- |
|
|
15
|
+
| `browser_open` | optional absolute `http(s)` `url` | host Playwright `Browser` or `BrowserManager`, limits, `ExecutionPolicy`, `networkPolicy`, uploads/downloads |
|
|
16
|
+
| `browser_snapshot` | optional `pageId` | same manager/context |
|
|
17
|
+
| `browser_act` | `action` plus action-specific fields (`target`, `snapshotId`, `url`, `text`, `values`, `paths`, `downloadId`, `dialogResponse`, `pageId`, `clip`, …) | policy checked before side effects |
|
|
18
|
+
| `browser_close` | none | closes only the run-owned context, never the host Browser process |
|
|
19
|
+
|
|
20
|
+
`createBrowserTools({ browser, executionPolicy?, limits?, networkPolicy?, uploads?, downloads?, beforeSideEffect? })` builds the four tools. `createBrowserManager(...)` exposes host lifecycle helpers `closeRun(runId)` / `close()` and `listDownloads(runId)`.
|
|
21
|
+
|
|
22
|
+
Targets accepted by `browser_act`: snapshot `ref`, `role`(+`name`), `label`, `testId`, or `text`. CSS, XPath, selector strings, `page.evaluate`, CDP/devtools, extensions, and persistent/local profiles are unsupported.
|
|
23
|
+
|
|
24
|
+
`browser_act` actions: `navigate`, `click`, `type`, `fill`, `select`, `check`, `uncheck`, `scroll`, `wait`, `dialog`, `select_page`, `upload`, `screenshot`, `download_release`.
|
|
25
|
+
|
|
26
|
+
## Outputs / response / events
|
|
27
|
+
|
|
28
|
+
`browser_open` returns run/page ids and URL. `browser_snapshot` returns `snapshotId`, URL/title, bounded AI-mode aria YAML (`ariaSnapshot({ mode: "ai" })`), ref count, and truncation metadata. Refs are valid only for that snapshot id and become stale after navigation or mutation. `browser_act` returns the action, active page id, and URL; `screenshot` also returns bounded `ImageContent`; `download_release` returns quarantine metadata after host approval. `browser_close` is idempotent. Results mark `trust: "untrusted_external"`; page text must never alter tools, permissions, credentials, or policy.
|
|
29
|
+
|
|
30
|
+
## Request/response example
|
|
31
|
+
|
|
32
|
+
```json
|
|
33
|
+
{
|
|
34
|
+
"tool": "browser_snapshot",
|
|
35
|
+
"arguments": {},
|
|
36
|
+
"result": {
|
|
37
|
+
"snapshotId": "snap_ab12…",
|
|
38
|
+
"pageId": "page_1",
|
|
39
|
+
"url": "https://example.com/",
|
|
40
|
+
"title": "Example",
|
|
41
|
+
"refCount": 12,
|
|
42
|
+
"ariaSnapshot": "- main [ref=e8]:\n - button \"Submit\" [ref=e12]"
|
|
43
|
+
}
|
|
44
|
+
}
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
```json
|
|
48
|
+
{
|
|
49
|
+
"tool": "browser_act",
|
|
50
|
+
"arguments": {
|
|
51
|
+
"action": "click",
|
|
52
|
+
"target": { "ref": "e12" },
|
|
53
|
+
"snapshotId": "snap_ab12…"
|
|
54
|
+
}
|
|
55
|
+
}
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
## Implementation example
|
|
59
|
+
|
|
60
|
+
```ts
|
|
61
|
+
import { chromium } from "playwright-core";
|
|
62
|
+
import {
|
|
63
|
+
createBrowserManager,
|
|
64
|
+
createBrowserTools,
|
|
65
|
+
createSharedSandboxBrowserOptions,
|
|
66
|
+
} from "@arnilo/prism-browser";
|
|
67
|
+
import { assertBrowserSandboxNetwork } from "@arnilo/prism-coding-security";
|
|
68
|
+
|
|
69
|
+
assertBrowserSandboxNetwork({
|
|
70
|
+
mode: "custom",
|
|
71
|
+
name: "prism-egress",
|
|
72
|
+
browserEgress: { proxyEndpoint: "http://127.0.0.1:3128", denyDirectEgress: true },
|
|
73
|
+
});
|
|
74
|
+
|
|
75
|
+
const aligned = createSharedSandboxBrowserOptions({
|
|
76
|
+
workspaceRoot: "/workspace",
|
|
77
|
+
downloadsRoot: "/downloads",
|
|
78
|
+
containedProxyAttestation: {
|
|
79
|
+
proxyEndpoint: "http://127.0.0.1:3128",
|
|
80
|
+
denyDirectEgress: true,
|
|
81
|
+
},
|
|
82
|
+
approveDownloadRelease: async (meta) => meta.bytes < 1_000_000,
|
|
83
|
+
});
|
|
84
|
+
|
|
85
|
+
const browser = await chromium.launch({ headless: true });
|
|
86
|
+
const manager = createBrowserManager({
|
|
87
|
+
browser,
|
|
88
|
+
...aligned,
|
|
89
|
+
limits: { maxPages: 4, maxActions: 100, maxSnapshotBytes: 256 * 1024 },
|
|
90
|
+
});
|
|
91
|
+
const tools = createBrowserTools({ manager, executionPolicy });
|
|
92
|
+
|
|
93
|
+
// On run terminal / abort / cancel:
|
|
94
|
+
await manager.closeRun(runId);
|
|
95
|
+
await manager.close();
|
|
96
|
+
await browser.close();
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
## Extension and configuration notes
|
|
100
|
+
|
|
101
|
+
- Compatibility line: `playwright-core@1.61.0` optional peer. Hosts pin browser binaries/images; Prism package install downloads nothing.
|
|
102
|
+
- Default/hard caps: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2k/10k; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30s/120s; action 10s/60s; wait 30s/120s; run wall 20min/30min; popups 4/16; dialogs 16/64; close grace 5s/30s; network requests 1k/10k; redirects/request 10/32; WebSockets 8/32; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate.
|
|
103
|
+
- Contexts use `serviceWorkers: "block"` and install `BrowserContext.route()` for every visible HTTP(S)/WebSocket request. `acceptDownloads` is enabled only when `downloads` is configured.
|
|
104
|
+
- `networkPolicy` defaults to `requireContainedProxy: true` (fail closed). Hosts must supply `containedProxyAttestation: { proxyEndpoint, denyDirectEgress: true }`. Private/loopback/link-local, `file`/`data`/`blob`/`javascript`/`devtools` schemes are denied by default. Playwright routing is defense in depth — production DNS/private egress is a host firewall/proxy.
|
|
105
|
+
- Uploads require absolute paths under `uploads.roots` (realpath-contained; symlink escapes rejected). Downloads stream into `downloads.quarantine` with SHA-256/MIME/name metadata; `download_release` requires host `approveRelease`. Screenshots return bounded `ImageContent`.
|
|
106
|
+
- Observation (`snapshot`, `wait`, open-without-url, `close`) vs mutation/high-impact (`navigate`, click/form, dialog accept, upload, download release, popup select) is classified for `ExecutionPolicy` / `beforeSideEffect`.
|
|
107
|
+
- `createSharedSandboxBrowserOptions()` aligns browser uploads/downloads with Task 1 sandbox `/workspace` and `/downloads`. `assertBrowserSandboxNetwork()` in `@arnilo/prism-coding-security` fails closed for custom Docker networks without browser egress attestation.
|
|
108
|
+
- Raw CSS is absent from production defaults. Ref resolution uses Playwright’s built-in `aria-ref=` selector with a package-owned snapshot ref table for staleness checks.
|
|
109
|
+
|
|
110
|
+
## Security and performance notes
|
|
111
|
+
|
|
112
|
+
Import is inert. Construction fails clearly when neither `browser` nor `manager` is supplied. Browser installation, launch, version, and control endpoint are host-owned. Prism never exposes `page.evaluate`, init scripts, CDP, extensions, persistent profiles, or model-supplied Playwright launch options. Secrets and storage state must not appear in snapshots, tool results, logs, or checkpoints. Finite caps charge before context/page/action/queue/snapshot/network/artifact retention; snapshots retain no unbounded DOM, console, request, response, or trace history. Unreleased downloads are deleted on context close.
|
|
113
|
+
|
|
114
|
+
Default tests use fake Playwright APIs only. Protected live gate: `PRISM_LIVE_PLAYWRIGHT=1` (or `PRISM_TEST_PLAYWRIGHT=1`) `npm run test:live -w @arnilo/prism-browser` exercises a local loopback hostile HTML fixture for snapshot refs, stale-ref rejection, CSS denial, private/file deny, upload containment, screenshot bounds, and download quarantine/release. Missing browser binaries fail closed when the gate is enabled. Adversarial network-free fixtures live in `eval-fixtures.test.ts`; see [Evaluations](evaluations.md) and `examples/coding-browser-evaluation.ts`.
|
|
115
|
+
|
|
116
|
+
## Related APIs
|
|
117
|
+
|
|
118
|
+
- [Tools](tools.md): registry, exclusive dispatch, validation, and ledger.
|
|
119
|
+
- [Web search, fetch, and extraction](web-tools.md): preferred non-interactive retrieval path.
|
|
120
|
+
- [Guardrails](guardrails.md): untrusted external content handling.
|
|
121
|
+
- [Host security](host-security.md): browser endpoint, approval, egress proxy, and artifact trust boundaries.
|
|
122
|
+
- [Performance and resource limits](performance.md): browser ceilings and charging points.
|
|
123
|
+
- [Coding execution approval and sandboxing](coding-security.md): optional shared disposable sandbox for coding+browser.
|
|
124
|
+
- [Migration](migration.md): additive optional package activation.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem tools as Prism `ToolDefinition` objects. It ships
|
|
5
|
+
`@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem/repository tools as Prism `ToolDefinition` objects. It ships six default coding tools — `shell`, `read`, `write`, `edit`, `repo_list`, `repo_search` — plus an opt-in structured Git/check set (`createGitTools`) for status/diff/branch/worktree/apply/commit/PR-handoff and named checks. Bounded coding-plan/checkpoint helpers compose ordinary workspace Markdown with workflow checkpoint state (references/hashes/summaries/fingerprints only). The tools are **inert** until a host imports them and registers them into a `ToolRegistry`. Behavior for shell/read/write/edit is a behavioral port of the pi coding agent's tools, adapted to Prism's `ToolDefinition` / `ToolResult` contracts (no `@earendil-works/*` or `typebox` dependencies; only `diff` plus the Node standard library). List/search/Git are native Prism tools with no glob/ripgrep/Git-library dependency.
|
|
6
6
|
|
|
7
7
|
| Export | Purpose |
|
|
8
8
|
| --- | --- |
|
|
@@ -10,13 +10,23 @@
|
|
|
10
10
|
| `createReadTool(cwd, options?)` | `read` tool: read a text or image file into `TextContent` / `ImageContent`. |
|
|
11
11
|
| `createWriteTool(cwd, options?)` | `write` tool: create or overwrite a file, creating parent directories. |
|
|
12
12
|
| `createEditTool(cwd, options?)` | `edit` tool: precise exact-then-fuzzy text replacement in an existing file. |
|
|
13
|
-
| `
|
|
14
|
-
| `
|
|
15
|
-
| `
|
|
13
|
+
| `createRepoListTool(cwd, options?)` | `repo_list` tool: bounded deterministic repository listing. |
|
|
14
|
+
| `createRepoSearchTool(cwd, options?)` | `repo_search` tool: bounded literal/regex text search. |
|
|
15
|
+
| `createCodingTools(cwd, options?)` | Default six tools (`shell`, `read`, `write`, `edit`, `repo_list`, `repo_search`). |
|
|
16
|
+
| `createReadOnlyTools(cwd, options?)` | Read-only subset: `read`, `repo_list`, `repo_search`. |
|
|
17
|
+
| `createAllTools(cwd, options?)` | Identical to `createCodingTools` (Git tools remain opt-in via `createGitTools`). |
|
|
18
|
+
| `createGitTools(cwd, options?)` | Opt-in Git tools (`git_status`/`git_diff`/`git_branch`/`git_worktree`/`git_apply`/`git_commit`/`git_pr_handoff`) plus optional `coding_check`. |
|
|
19
|
+
| `createCodingCheckTool(cwd, options)` | Named host-declared checks; model selects only a name. |
|
|
20
|
+
| `createLocalRepositoryOperations(limits?)` | Default streaming Node filesystem backend for list/search. |
|
|
21
|
+
| `createGitOperations(options)` | Typed Git operations backend (argument arrays, safe config, finite output). |
|
|
22
|
+
| `buildCodingCheckpointMetadata` / `validateCodingCheckpointMetadata` / `assertCodingResumeAllowed` | Bounded durable coding-task metadata for workflow `state.coding` (no second runtime). |
|
|
23
|
+
| `writeCodingPlanFile` / `readCodingPlanFile` / `createCodingPlanMarkdown` / `parseCodingPlanTodos` | Workspace plan/todo Markdown helpers with finite byte/todo caps and hash verification. |
|
|
24
|
+
| `fingerprintJson` / `CODING_STATE_KEY` | Stable tool/policy fingerprints and the shared-state key for coding metadata. |
|
|
16
25
|
| `detectSupportedImageMimeType(buf)` / `detectSupportedImageMimeTypeFromFile(path)` | Magic-byte image MIME detection (PNG/JPEG/GIF/WebP/BMP) used by `read`. |
|
|
17
26
|
| `DEFAULT_MAX_IMAGE_BYTES` | Default `read` image size ceiling (10 MB). |
|
|
18
|
-
| `DEFAULT_*` / `HARD_*` coding limit constants | Published text-scan, image, write/edit, shell
|
|
27
|
+
| `DEFAULT_*` / `HARD_*` coding limit constants | Published text-scan, image, write/edit, shell, repository, Git, check, handoff, and plan/checkpoint ceilings. |
|
|
19
28
|
| `ReadTextOptions` / `ReadTextResult` | Bounded text-page contract required by custom `ReadOperations`. |
|
|
29
|
+
| `RepositoryOperations` / `RepositoryLimitOptions` | Pluggable list/search backend and finite caps. |
|
|
20
30
|
| `TransformImage` / `TransformImageInput` | Types for the optional `read` `transformImage` callback. |
|
|
21
31
|
| `withFileMutationQueue(path, fn)` | Per-path serialization primitive re-exported for hosts. |
|
|
22
32
|
|
|
@@ -55,6 +65,7 @@ const tools = createCodingTools(workspaceRoot, {
|
|
|
55
65
|
| `read` | `read` |
|
|
56
66
|
| `write` | `write` |
|
|
57
67
|
| `edit` | `edit` |
|
|
68
|
+
| `repo_list` / `repo_search` | _(native; no pi equivalent)_ |
|
|
58
69
|
|
|
59
70
|
## Inputs / request
|
|
60
71
|
|
|
@@ -159,6 +170,80 @@ Each `edits[].oldText` must match a unique, non-overlapping region of the origin
|
|
|
159
170
|
|
|
160
171
|
`edit` result `metadata`: `{ diff, patch, firstChangedLine }` — a display-oriented diff, a standard unified patch, and the first changed line in the new file. These are host-readable; the model only sees the short confirmation (keeps model context small).
|
|
161
172
|
|
|
173
|
+
### `repo_list`
|
|
174
|
+
|
|
175
|
+
List repository entries with deterministic relative paths. Uses Node `opendir`/`lstat` only — no glob dependency. Does not follow symlinks; rejects path escapes outside the workspace root. Hidden names and excluded basenames (default `.git`, `node_modules`, `dist`) are skipped unless `includeHidden` is set / host `exclude` is overridden.
|
|
176
|
+
|
|
177
|
+
**Inputs:**
|
|
178
|
+
|
|
179
|
+
| Field | Type | Purpose |
|
|
180
|
+
| --- | --- | --- |
|
|
181
|
+
| `path` | `string` | Workspace-relative directory or file to list (default root). |
|
|
182
|
+
| `includeHidden` | `boolean` | Include dot names (default false). |
|
|
183
|
+
| `maxDepth` | `number` | Directory depth cap (default 32, hard 128). |
|
|
184
|
+
| `maxResults` | `number` | Page size (default 1,000, hard 10,000). |
|
|
185
|
+
| `offset` | `number` | Entries to skip before retaining (default 0). |
|
|
186
|
+
|
|
187
|
+
**Outputs:** text lines `kind\trelative/path[\tsize]` plus metadata (`truncated`, `truncatedBy`, `nextOffset`, `entries`, scan counts). Continue with `offset=nextOffset` when truncated by results.
|
|
188
|
+
|
|
189
|
+
### `repo_search`
|
|
190
|
+
|
|
191
|
+
Search text files under the workspace. Default mode is literal substring match; `mode: "regex"` enables length-bounded regular expressions. Binary files (NUL in a bounded prefix) and oversize files are skipped. Aggregate scanned bytes, matches, line bytes, pattern bytes, and wall time are finite.
|
|
192
|
+
|
|
193
|
+
**Inputs:**
|
|
194
|
+
|
|
195
|
+
| Field | Type | Purpose |
|
|
196
|
+
| --- | --- | --- |
|
|
197
|
+
| `query` | `string` | Literal or regex pattern (required). |
|
|
198
|
+
| `path` | `string` | Workspace-relative start path. |
|
|
199
|
+
| `mode` | `"literal" \| "regex"` | Default `literal`. |
|
|
200
|
+
| `caseSensitive` | `boolean` | Default false. |
|
|
201
|
+
| `includeHidden` | `boolean` | Default false. |
|
|
202
|
+
| `context` | `number` | Context lines before/after each match (default 5, hard 20). |
|
|
203
|
+
| `maxMatches` | `number` | Match cap (default 1,000, hard 10,000). |
|
|
204
|
+
|
|
205
|
+
**Outputs:** ripgrep-like lines `path:line:column:text` with optional `path-` / `path+` context, plus metadata (`matches`, `truncated`, scan/skip counts).
|
|
206
|
+
|
|
207
|
+
### Structured Git tools (`createGitTools`)
|
|
208
|
+
|
|
209
|
+
Opt-in tools over a host-pinned Git executable (`gitPath`, default `/usr/bin/git`) or sandbox `execFile`. Every invocation uses argument arrays with safe config (`core.hooksPath=/dev/null`, empty credential helper, pager disabled, `GIT_TERMINAL_PROMPT=0`). Shell is never used internally. Git tools are **not** included in `createCodingTools()` / `createAllTools()`.
|
|
210
|
+
|
|
211
|
+
| Tool | Purpose |
|
|
212
|
+
| --- | --- |
|
|
213
|
+
| `git_status` | `status --porcelain=v2 -z --branch` → structured branch + entries + `dirty`. |
|
|
214
|
+
| `git_diff` | Bounded `--no-ext-diff --no-textconv` diff; oversized output may spill via `artifactWriter`. |
|
|
215
|
+
| `git_branch` | `validate` / `list` / `create` / `switch` with `git check-ref-format --branch`. Switch refuses unrelated dirty trees unless `createCheckpoint=true`. |
|
|
216
|
+
| `git_worktree` | `list` / `add` / `remove` within finite worktree caps. |
|
|
217
|
+
| `git_apply` | `check` / `apply` / `reverse`; always `--check` before mutating apply. Apply requires clean/checkpoint; failures restore. |
|
|
218
|
+
| `git_commit` | Explicit-path `add` + `commit --no-verify -F <tempfile>`; requires host `commitIdentity`. Allows dirty entries that are exactly the requested paths; unrelated dirt requires checkpoint. Never pushes. |
|
|
219
|
+
| `git_pr_handoff` | Bounded `{ base, head, commits, changedPaths, diffstat, checks, artifact? }` for host PR creation. Never authenticates or opens a PR. |
|
|
220
|
+
| `coding_check` | Included when `checks` are declared: model selects only a name; executable/args/env are host-fixed. |
|
|
221
|
+
|
|
222
|
+
```ts
|
|
223
|
+
import { createGitTools } from "@arnilo/prism-coding-agent";
|
|
224
|
+
|
|
225
|
+
const gitTools = createGitTools(workspaceRoot, {
|
|
226
|
+
gitPath: "/usr/bin/git",
|
|
227
|
+
commitIdentity: { name: "Prism Bot", email: "bot@example.com" },
|
|
228
|
+
checks: {
|
|
229
|
+
test: { file: "/usr/bin/npm", args: ["test"] },
|
|
230
|
+
},
|
|
231
|
+
});
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
### Durable coding plans and checkpoints
|
|
235
|
+
|
|
236
|
+
There is no `CodingRun`, todo database, or second approval engine. Persist executable plan/todos as ordinary workspace Markdown (for example `plans/<task>.md`) and store only bounded metadata under workflow `state.coding`:
|
|
237
|
+
|
|
238
|
+
| Field group | Stored in checkpoint | Not stored |
|
|
239
|
+
| --- | --- | --- |
|
|
240
|
+
| Plan / workspace export / patch artifacts | URI + SHA-256 + byte count | File contents, credentials, raw command output |
|
|
241
|
+
| Branch / worktree / base | Paths and ref names | Full diffs |
|
|
242
|
+
| Named checks | Name + exit code + short summary | Full stdout/stderr |
|
|
243
|
+
| Fingerprints | Workflow revision, definition hash, tool/policy fingerprints, optional image digest | Browser storage state, secrets, env |
|
|
244
|
+
|
|
245
|
+
Use `writeCodingPlanFile` / `readCodingPlanFile` for the workspace artifact, `buildCodingCheckpointMetadata` before `ctx.updateState({ coding })`, and `assertCodingResumeAllowed` before import/resume. Wrong owner/revision/hash/fingerprint fails closed. See `examples/durable-coding-workflow.ts` for a network-free plan → branch → edit → check → approval → handoff composition over `runWorkflow` / `resumeWorkflow` / `startWorkflowBackground`.
|
|
246
|
+
|
|
162
247
|
## Outputs / response / events
|
|
163
248
|
|
|
164
249
|
Every tool returns a `ToolResult` with `toolCallId`, `name`, `content` (`readonly ContentBlock[]`), optional `error`, and optional `metadata`. `write` and `edit` serialize per realpath through `withFileMutationQueue` so concurrent calls targeting one file do not interleave. `shell` is marked `exclusive`; tool dispatch serializes it at the turn level. The package emits no events of its own; hosts observe tool execution through the normal Prism `AgentEvent` stream via `dispatchToolCall`.
|
|
@@ -197,10 +282,10 @@ Minimal drop-in for any Prism app:
|
|
|
197
282
|
import { createToolRegistry } from "@arnilo/prism";
|
|
198
283
|
import { createCodingTools, createReadOnlyTools } from "@arnilo/prism-coding-agent";
|
|
199
284
|
|
|
200
|
-
// Full coding set (shell + read + write + edit) against the project root:
|
|
285
|
+
// Full coding set (shell + read + write + edit + repo_list + repo_search) against the project root:
|
|
201
286
|
const tools = createToolRegistry(createCodingTools(process.cwd()));
|
|
202
287
|
|
|
203
|
-
// Or a read-only set for inspection-only agents:
|
|
288
|
+
// Or a read-only set for inspection-only agents (read + repo_list + repo_search):
|
|
204
289
|
const ro = createToolRegistry(createReadOnlyTools(process.cwd()));
|
|
205
290
|
```
|
|
206
291
|
|
|
@@ -227,17 +312,18 @@ const remoteWrite = createWriteTool("/repo", {
|
|
|
227
312
|
|
|
228
313
|
## Extension and configuration notes
|
|
229
314
|
|
|
230
|
-
- **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
|
|
231
|
-
- **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits
|
|
232
|
-
- **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override.
|
|
315
|
+
- **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. Custom `RepositoryOperations` must honor depth/entry/file/match/scan/time caps and abort. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
|
|
316
|
+
- **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`; list/search accept `repository` limits and shared aggregator `ToolsOptions.repository`.
|
|
317
|
+
- **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit?, list?, search?, repository? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override. Read-only membership is deliberately `read` + `repo_list` + `repo_search` (0.0.9 behavior change).
|
|
318
|
+
- **Sandbox composition.** Prefer `@arnilo/prism-coding-security` `createSandboxCodingComposition(cwd, { workspaceMode, sandbox, ... })` (or tools-only wrappers). `workspaceMode` is required: `"sandbox"` keeps shell/read/write/edit/list/search on one disposable tree; `"host"` runs against host cwd and never claims containment. Mixed sandbox-shell + host-FS wiring throws unless `allowMixedWorkspaceWiring: true`. Same-tree Git: `createGitTools(composition.workspaceRoot, { execFile: sandbox.execFile, commitIdentity })`.
|
|
233
319
|
- **`ToolsOptions`** and the per-tool option types are exported from the package barrel for host configuration.
|
|
234
320
|
- No auto-discovery or manifest registration: import and register explicitly. This package registers no extensions and owns no globals (the mutation queue is a process-wide per-path map — see `ponytail:` note in the source).
|
|
235
321
|
|
|
236
322
|
## Security and performance notes
|
|
237
323
|
|
|
238
|
-
- **Host shell/filesystem access.** These tools run real commands and read/write real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
|
|
324
|
+
- **Host shell/filesystem access.** These tools run real commands and read/write/list/search real files. They provide **no sandbox**. Gate them with Prism `PermissionPolicy` / `ToolValidator` / trust policies before registering them for any provider turn. Shared `executionPolicy` applies to both full and read-only aggregators before filesystem/process side effects. See [Host security guide](host-security.md) and [Security/auth/trust](settings-auth-trust-security.md).
|
|
239
325
|
- **Non-zero exit is not an error.** A failing command is a normal `shell` result (exit code in metadata); only timeout/abort/spawn failures are error results. Do not assume `error == undefined` means the command succeeded.
|
|
240
|
-
- **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
|
|
326
|
+
- **Bounded I/O.** `read` streams one page and bounds scan bytes; image/edit reads use stat plus a shared cap-enforcing reader; write/edit inputs are measured before mutation. `repo_list`/`repo_search` stream walks and charge depth/entry/file/match/scan/time before retention. Structured Git tools use argument arrays with finite output/path/ref/message/patch caps, disable hooks/credential prompts/external diff by default, and never push or open PRs. `shell` retains only a rolling display tail and synchronously spills accepted raw chunks so stream backpressure cannot grow heap; wall time and total raw output remain finite.
|
|
241
327
|
- **Per-path serialization.** Concurrent mutations to the same file serialize; concurrent mutations to different files do not block each other. The queue is a process-wide map — across sessions in one process, same-path writes still serialize (upgrade path: scope per registry if throughput matters).
|
|
242
328
|
- **Bounded image reads.** `read` rejects images over `maxImageBytes` (default 10 MB) by `stat` before read when possible; MIME is detected from magic bytes only. Optional `transformImage` is host-owned — the base package has no image-processing dependency.
|
|
243
329
|
|
|
@@ -252,8 +338,19 @@ const remoteWrite = createWriteTool("/repo", {
|
|
|
252
338
|
| Edit target / input / count | 8 MiB / 2 MiB / 100 | 64 MiB / 16 MiB / 1,000 | before target read/matching/write |
|
|
253
339
|
| Shell wall time | 600 seconds | 3,600 seconds | process-tree kill |
|
|
254
340
|
| Shell total stdout+stderr | 64 MiB | 1 GiB | process-tree kill; spill removal |
|
|
255
|
-
|
|
256
|
-
|
|
341
|
+
| Repo depth / entries / files / page | 32 / 10,000 / 10,000 / 1,000 | 128 / 100,000 / 100,000 / 10,000 | before descending/retaining next entry |
|
|
342
|
+
| Search scan / file / matches | 64 MiB / 8 MiB / 1,000 | 1 GiB / 64 MiB / 10,000 | before next file/match retention |
|
|
343
|
+
| Search pattern / line / context / time | 512 B / 50 KiB / 5 / 30 s | 4 KiB / 1 MiB / 20 / 300 s | before regex compile / line retain / deadline |
|
|
344
|
+
| Git paths / refs / message | 1,000 / 1 KiB / 64 KiB | 10,000 / 4 KiB / 256 KiB | before process/temp-file creation |
|
|
345
|
+
| Git output / diff lines / changed files / patch | 4 MiB / 10,000 / 1,000 / 16 MiB | 64 MiB / 100,000 / 10,000 / 64 MiB | stream before retain; artifact spill optional |
|
|
346
|
+
| Worktrees | 4 | 16 | before add |
|
|
347
|
+
| Named checks (names / concurrency / time / lines / output) | 8 / 1 / 10 min / 2,000 / 4 MiB | 32 / 4 / 60 min / 100,000 / 64 MiB | construction / before start / line retention |
|
|
348
|
+
| PR handoff JSON / commits | 256 KiB / 100 | 1 MiB / 1,000 | before result exposure |
|
|
349
|
+
| Plan markdown / todos / todo text | 256 KiB / 1,000 / 512 B | 1 MiB / 10,000 / 4 KiB | before write/parse/checkpoint |
|
|
350
|
+
| Coding checkpoint metadata / artifact refs / artifact bytes | 64 KiB / 16 / 256 MiB | 512 KiB / 64 / 2 GiB | before state save / resume verify |
|
|
351
|
+
| Check summary text | 1 KiB | 8 KiB | before checkpoint retention |
|
|
352
|
+
|
|
353
|
+
Every configurable value is a positive safe integer (context may be zero); Prism rejects rather than clamps invalid values. Limits control resources, not authority: they do not replace root containment, approval, validation, or a sandbox.
|
|
257
354
|
|
|
258
355
|
## Related APIs
|
|
259
356
|
|