@arnilo/prism 0.0.7 → 0.0.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/CHANGELOG.md +50 -2
  2. package/README.md +3 -1
  3. package/dist/agent-loops.js +14 -8
  4. package/dist/agents.js +37 -3
  5. package/dist/contracts.d.ts +17 -0
  6. package/dist/index.d.ts +4 -2
  7. package/dist/index.js +3 -2
  8. package/dist/provider-events.d.ts +2 -0
  9. package/dist/provider-events.js +21 -13
  10. package/dist/providers/openai-compatible.js +8 -5
  11. package/dist/providers/transport.d.ts +10 -1
  12. package/dist/providers/transport.js +24 -8
  13. package/dist/run-ledger.d.ts +21 -0
  14. package/dist/run-ledger.js +115 -0
  15. package/dist/tools.js +2 -0
  16. package/docs/a2a.md +61 -42
  17. package/docs/agent-events.md +5 -4
  18. package/docs/agent-loops.md +2 -2
  19. package/docs/agent-session-runtime.md +1 -0
  20. package/docs/browser-automation.md +124 -0
  21. package/docs/coding-agent-tools.md +111 -14
  22. package/docs/coding-security.md +84 -11
  23. package/docs/credential-storage.md +9 -0
  24. package/docs/database-persistence.md +1 -1
  25. package/docs/evaluations.md +38 -4
  26. package/docs/guardrails.md +3 -2
  27. package/docs/host-security.md +29 -4
  28. package/docs/index.md +20 -15
  29. package/docs/mcp-tools.md +29 -5
  30. package/docs/migration.md +104 -0
  31. package/docs/observability.md +26 -14
  32. package/docs/performance.md +54 -0
  33. package/docs/postgres-persistence.md +1 -0
  34. package/docs/provider-conformance.md +1 -1
  35. package/docs/provider-primitives.md +7 -1
  36. package/docs/providers/kimi.md +16 -2
  37. package/docs/providers/opencode-go.md +43 -2
  38. package/docs/release-and-install.md +117 -62
  39. package/docs/resource-loading.md +4 -0
  40. package/docs/review-coverage-2026-07-19-phase-3.md +174 -0
  41. package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
  42. package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
  43. package/docs/run-ledger-conformance.md +1 -0
  44. package/docs/runs-and-usage.md +17 -2
  45. package/docs/sqlite-persistence.md +1 -0
  46. package/docs/structured-output.md +2 -2
  47. package/docs/supervisors.md +2 -2
  48. package/docs/tools.md +5 -1
  49. package/docs/web-tools.md +78 -0
  50. package/docs/workflows.md +2 -0
  51. package/package.json +6 -4
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. `RunFeedbackStore` is the separate post-run seam for immutable ratings, comments, tags, and evaluation links. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
5
+ `RunLedger` is the host-implemented, write-only seam Prism uses to durably persist run metadata, agent events, tool calls, and usage during a `session.run()`. `createBatchedRunLedger()` is an explicit optional wrapper; direct ledger writes remain default. `RunFeedbackStore` is the separate post-run seam for immutable ratings, comments, tags, and evaluation links. The runtime calls the adapter as each record becomes available; the adapter decides how to write it (SQL insert, NoSQL put, JSONL append, time-series batch, etc.).
6
6
 
7
7
  APIs:
8
8
 
@@ -152,7 +152,7 @@ const page = await feedback.query({ runId: result.runId, tenantId: "t1", userId:
152
152
  await feedback.delete({ id: "fb_1", tenantId: "t1", userId: "u1" });
153
153
  ```
154
154
 
155
- Default/hard bounds: comment 4/16 KiB, tags 16/64, scorer/evaluation IDs 16/64 each, metadata 16/64 KiB, query page 100/500; tags are 64 characters and identifiers 128. The store redacts comment/tags/metadata after run ownership validation and before persistence. IDs are linked, not scorer payloads. `ProductionPersistenceStore.feedback?` exposes this capability; first-party SQLite/PostgreSQL adapters implement it in schema migration `003_run_feedback` and reject missing/cross-owned runs.
155
+ Default/hard bounds: comment 4/16 KiB, tags 16/64, scorer/evaluation IDs 16/64 each, metadata 16/64 KiB, query page 100/500; tags are 64 characters and identifiers 128. `@arnilo/prism-evals` may read `queryRuns/queryEvents/queryToolCalls/queryUsage` only through an explicit owner/session/run-scoped trace resolver with finite cursor pages and aggregate bytes. The store redacts comment/tags/metadata after run ownership validation and before persistence. IDs are linked, not scorer payloads. `ProductionPersistenceStore.feedback?` exposes this capability; first-party SQLite/PostgreSQL adapters implement it in schema migration `003_run_feedback` and reject missing/cross-owned runs.
156
156
 
157
157
  ## Status transitions
158
158
 
@@ -287,6 +287,21 @@ console.log(cacheUsageReport(aggregate?.usage));
287
287
  - **Tenant isolation.** `OwnershipScope` fields are copied from the active ownership scope, but the runtime does not enforce tenant isolation for ledger rows. Feedback is stricter: append/query/delete require tenant plus account/user, and first-party stores compare the exact scope to the linked run.
288
288
  - **Feedback privacy.** Comments/tags/metadata can contain PII. Configure a feedback redactor, apply retention, and call owned `delete()` for erasure. Never copy comments or tag values into metric labels.
289
289
 
290
+ ## Optional batching and durability
291
+
292
+ ```ts
293
+ const ledger = createBatchedRunLedger(store, {
294
+ maxBatchEntries: 128,
295
+ maxBatchBytes: 512 * 1024,
296
+ maxDelayMs: 25,
297
+ durability: "flush_on_terminal",
298
+ });
299
+ ```
300
+
301
+ Modes: `write_through` acknowledges each target write; `flush_on_terminal` buffers but runtime awaits terminal flush; `buffered` acknowledges enqueue only and requires host `flush()` for durability. `status()` distinguishes accepted/flushed/buffered counts. Defaults/hard caps: 128/4,096 batch entries, 512 KiB/8 MiB batch bytes, 25 ms/60 s delay; buffered count/bytes apply backpressure before enqueue. FIFO spans all record kinds. Inputs are already runtime-redacted. Flush errors propagate and retain the failing record for retry. `dispose({ flush: false })` clears memory but deliberately loses unflushed records—same crash-before-flush ceiling as process failure.
302
+
303
+ Runtime session snapshots cache one leaf/generation for at most one second. Successful append, compaction append, checkout, and durable resume invalidate; failed append does not advance leaf/cache. Cache is session-local and never shared across ownership/session/branch.
304
+
290
305
  ## Related APIs
291
306
 
292
307
  - [Performance limits](performance.md): batching, cursor keys, and production sizing assumptions.
@@ -108,6 +108,7 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
108
108
  - **File ownership.** Create database files on a host-controlled path with restrictive permissions (`0600` default on Unix via `fileMode`).
109
109
  - **No path interpolation.** The adapter opens exactly the caller-supplied `filename`; it does not expand environment variables or discover paths.
110
110
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
111
+ - **Optional batching.** SQLite remains write-through by default. Hosts may wrap its ledger with core `createBatchedRunLedger()`; `flush_on_terminal` preserves terminal acknowledgement, while `buffered` explicitly risks crash-before-flush loss.
111
112
  - **WAL + busy timeout.** WAL is enabled by default; busy timeout defaults to 5 seconds. This meets the Plan 056 local workload target but SQLite still serializes writers — prefer PostgreSQL for high write concurrency.
112
113
  - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans. Startup validation reads SQLite catalog/PRAGMA metadata only, never application rows.
113
114
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns on run and ownership tables participate in query filters; hosts must still scope writes correctly.
@@ -231,9 +231,9 @@ Key cross-seam points:
231
231
  - `generate-validate-revise` is selected via `AgentConfig.loop` / `RunOptions.loop` (`RunOptions.loop` wins). See [Agent loops](agent-loops.md). `resolveLoop()` maps the options form to the factory; an unknown `strategy` throws before the first turn; a custom `AgentLoopStrategy` instance bypasses the options form.
232
232
  - Native structured output uses provider-neutral `StructuredOutputOptions` on `ProviderRequestOptions` / loop options. Capable OpenAI-family providers map to JSON-schema wire fields; unsupported models fail before fetch unless the host sets `structuredOutputMode: "artifact-loop"` and relies on parser/validator/repairer only.
233
233
  - `validateStructuredOutputOptions()` enforces JSON-safe schemas, forbidden prototype-pollution keys, and a 64 KiB schema size cap.
234
- - The default parser treats assistant text as the value (`{ ok: true, value: text }`); supply a host parser whenever `T` is not `string`.
234
+ - The default parser treats non-empty assistant text as the value (`{ ok: true, value: text }`); empty/whitespace-only call-free text is a `parse_error` before the parser. Supply a host parser whenever `T` is not `string`.
235
235
  - The default repairer builds a user message from `validation.errors[].message`; supply a host repairer for schema-specific guidance.
236
- - `maxRevisions` (default 3) bounds revision turns; budget exhaustion ends the loop and emits `artifact_failed` (it does not throw).
236
+ - `maxRevisions` (default 3) bounds revision turns; budget exhaustion ends the loop and emits `artifact_failed`. Session runs then fail with `AgentRunError` unless `artifact_finished` occurred (direct `loop.run` still returns usage without throwing).
237
237
  - Tools are inert in artifact turns unless `loop.toolCalls: "bounded"` is explicit. Bounded mode uses run-global `maxToolRounds`, dispatches calls sequentially through normal runtime guards, skips parser/validator for tool-calling responses, and permits at most `1 + maxRevisions + maxToolRounds` provider turns. An extra tool response yields terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"` and executes nothing.
238
238
 
239
239
  ## Security and performance notes
@@ -21,7 +21,7 @@ Use a supervisor when a host or agent must choose a child dynamically. Use `@arn
21
21
 
22
22
  ## Outputs / response / events
23
23
 
24
- `delegate()` returns the child's `AgentRunResult` or throws its `AgentRunError`/a supervisor denial or limit error. `subscribe()` emits bounded `delegation_started`, `delegation_finished`, `delegation_rejected`, and `delegation_error` metadata events.
24
+ `delegate()` returns the child's `AgentRunResult` or throws its `AgentRunError`/a supervisor denial or limit error. `subscribe()` emits bounded `delegation_started`, `delegation_finished`, `delegation_rejected`, and `delegation_error` metadata events. Hosts may project those events through observability `handleDelegation()` using the parent Prism run ID; no OpenTelemetry dependency enters this package.
25
25
 
26
26
  ## Request/response example
27
27
 
@@ -65,7 +65,7 @@ Child factories resolve their own providers/credentials and construct context/me
65
65
 
66
66
  ## Related APIs
67
67
 
68
- - [A2A interoperability](a2a.md): remote protocol boundary.
68
+ - [A2A interoperability](a2a.md): separate remote protocol boundary. `A2ATaskLifecycle` adapts host durable agent/workflow state directly; it does not route A2A execution through local supervisor child planning.
69
69
  - [Workflows](workflows.md): preferred deterministic orchestration.
70
70
  - [Working and semantic memory](working-and-semantic-memory.md): child scope construction.
71
71
  - [Host security](host-security.md): permission and credential boundaries.
package/docs/tools.md CHANGED
@@ -10,6 +10,8 @@ APIs:
10
10
  - `filterTools()`
11
11
  - `dispatchToolCall()`
12
12
  - `ToolFilter`, `ToolFilterInput`, `ToolValidator`, `DispatchToolCallOptions`
13
+ - Optional [Web search, fetch, and extraction](web-tools.md): three narrow host-selected `ToolDefinition`s with untrusted bounded outputs.
14
+ - Optional [Browser automation](browser-automation.md): four exclusive Playwright tools over a host-supplied browser with run-owned contexts and snapshot refs.
13
15
 
14
16
  ## When to use it
15
17
 
@@ -90,6 +92,8 @@ When `options.ledger` is set, `dispatchToolCall()` also appends a `ToolCallRecor
90
92
 
91
93
  Blocked reasons are `unknown_tool`, `tool_denied`, `invalid_arguments`, `permission_denied`, and `validation_failed`. Progress snapshots reuse status `started` because the tool call is still in flight.
92
94
 
95
+ Malformed streamed tool-call JSON (id+name present, arguments not a JSON object) does not fail the provider turn. Providers emit a tool call with `argumentsError`; `dispatchToolCall` blocks with reason `invalid_arguments` and `error.code: "invalid_json_arguments"`, persists a failed tool result, and never calls `execute()`. The model can self-correct within `maxToolRounds`/`maxTurns`.
96
+
93
97
  ## Request/response example
94
98
 
95
99
  ```json
@@ -260,7 +264,7 @@ createJsonSchemaToolArgumentValidator({
260
264
  - [Credentials and redaction](credentials-and-redaction.md): redaction helpers used for tool execution errors.
261
265
  - [Observational memory compaction package](compaction-observational-memory.md): optional exact-id recall tool factory.
262
266
  - [Tool execution primitives](tool-execution-primitives.md): JSON Schema adapter, parallelism, MCP bridge, and execution-policy designs.
263
- - [MCP client bridge](mcp-tools.md): optional `@arnilo/prism-mcp` remote tool mapping.
267
+ - [MCP client bridge](mcp-tools.md): optional remote tool mapping plus separate bounded resource/prompt facades; non-tool MCP capabilities never bypass tool dispatch by masquerading as `ToolDefinition`.
264
268
  - [Coding agent tools](coding-agent-tools.md): optional first-party `@arnilo/prism-coding-agent` `shell`/`read`/`write`/`edit` tools a host registers into this harness.
265
269
 
266
270
  `DispatchToolCallOptions.trust` and `.permission` run before validation or `execute()`; denial emits `tool_execution_blocked`. Middleware cannot bypass either guard. `AgentConfig.validator`/`RunOptions.validate` run after these guards; their output is redacted through the active `SecretRedactor`. `createSecureAgent()` requires all three seams plus non-empty schemas and durable pre-tool approval. Prism does not sandbox tools. See [Security/auth/trust](settings-auth-trust-security.md).
@@ -0,0 +1,78 @@
1
+ # Web search, fetch, and extraction
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-web-tools` provides three separate bounded `ToolDefinition`s: host-selected Brave or Exa `web_search`, Firecrawl Markdown `web_fetch`, and Firecrawl JSON `web_extract`. Package uses native `fetch`; no vendor SDK or browser is installed.
6
+
7
+ ## When to use it
8
+
9
+ Use when agent needs explicit public-web discovery or host-approved document retrieval/extraction. Keep search separate from fetch/extract so model cannot select provider, credential, API origin, extraction schema, or cost path.
10
+
11
+ ## Inputs / request
12
+
13
+ | Tool | Model-visible input | Host-only construction input |
14
+ | --- | --- | --- |
15
+ | `web_search` | `query`, optional `count` | exactly one `createBraveSearch()` or `createExaSearch()`, credential source, limits |
16
+ | `web_fetch` | one absolute public `url` | `createFirecrawlFetch()`, credential source, SSRF/DNS policy, limits |
17
+ | `web_extract` | bounded `urls` array | `createFirecrawlExtractor()`, fixed JSON Schema, validator, credential source, SSRF/DNS policy, limits |
18
+
19
+ Credentials may be explicit callbacks or core `CredentialResolver`s. They resolve at request edge as Brave `subscription_token`, Exa `api_key`, or Firecrawl `api_key`. `allowedOrigins` must contain fixed provider API origin; adapters reject redirects. Firecrawl target URLs reject userinfo, non-HTTP(S), private literals, and hostnames denied by `SsrfPolicy`; supply `validateUrl` for host DNS/rebinding/egress checks before handing URL to Firecrawl.
20
+
21
+ ## Outputs / response / events
22
+
23
+ Search results retain bounded `title`, canonical fragment-free `url`, `snippet`, highlights, provider result ID, publication/retrieval time, returned request/cost/rate facts, and stable citation identity. Citation is `web:<provider>:<sourceId>` when provider supplies ID, otherwise SHA-256 of canonical URL.
24
+
25
+ Fetch returns bounded Markdown and selected attribution. Extract validates host schema before I/O and validates returned JSON through host `ToolArgumentValidator`. Every result has `untrusted: true`; tool result metadata is `trust: "untrusted_external"`. Missing provider facts remain absent—Prism never guesses billing or freshness.
26
+
27
+ ## Request/response example
28
+
29
+ ```json
30
+ {
31
+ "tool": "web_search",
32
+ "arguments": { "query": "Prism TypeScript SDK", "count": 5 },
33
+ "result": {
34
+ "provider": "brave",
35
+ "untrusted": true,
36
+ "results": [{ "citationId": "web:brave:…", "url": "https://example.com/", "title": "Example" }]
37
+ }
38
+ }
39
+ ```
40
+
41
+ ## Implementation example
42
+
43
+ ```ts
44
+ import { createEnvCredentialResolver } from "@arnilo/prism";
45
+ import { createJsonSchemaArgumentValidator } from "@arnilo/prism-tool-validator-json-schema";
46
+ import { createBraveSearch, createFirecrawlExtractor, createFirecrawlFetch, createWebTools } from "@arnilo/prism-web-tools";
47
+
48
+ const credentials = createEnvCredentialResolver(process.env, {
49
+ "brave:subscription_token": "BRAVE_SEARCH_TOKEN",
50
+ "firecrawl:api_key": "FIRECRAWL_API_KEY",
51
+ });
52
+ const schema = { type: "object", properties: { title: { type: "string" } }, required: ["title"], additionalProperties: false };
53
+ const tools = createWebTools({
54
+ search: createBraveSearch({ credentials }),
55
+ fetch: createFirecrawlFetch({ credentials, validateUrl: publicDnsPolicy }),
56
+ extract: createFirecrawlExtractor({ credentials, schema, validator: createJsonSchemaArgumentValidator(), validateUrl: publicDnsPolicy }),
57
+ });
58
+ ```
59
+
60
+ ## Extension and configuration notes
61
+
62
+ Hosts substitute `createExaSearch()` for Brave; no runtime/model routing exists. Root export includes all adapters; `./brave`, `./exa`, and `./firecrawl` subpaths support atomic imports. Official vendor MCP servers are prototypes only: use hardened MCP origin/auth/capability policy, never generic remote passthrough.
63
+
64
+ Default/hard limits: query 4/16 KiB; results 10/20; URLs 5/20; request 256 KiB/1 MiB; response and aggregate 2/16 MiB; Markdown 1/8 MiB; extraction 256 KiB/1 MiB; schema 64/256 KiB; JSON depth 64/128 and properties 10k/100k; retries 2/4; rate delay 5/60 seconds; concurrency 4/16 active plus the same bounded waiting queue; polling 20/100; wall time 60 seconds/30 minutes. Hosts may only narrow or raise within hard caps.
65
+
66
+ ## Security and performance notes
67
+
68
+ Provider credentials never enter tool schemas/results, prompts, telemetry, URLs, or errors. Error text excludes remote bodies. Search snippets, Markdown, and extracted JSON are prompt-injection-capable data: never concatenate them into system instructions or use them to modify tools, permissions, credentials, trust, routing, or schemas. Firecrawl fetches target URLs remotely; Prism cannot claim target DNS pinning after handoff. Use controlled host fetch when that guarantee is required.
69
+
70
+ Default tests use injected fake fetch and make no public request. Restricted smoke: `PRISM_LIVE_WEB=1 npm run test:live -w @arnilo/prism-web-tools` plus least-privilege provider environment credential. Prefer these tools over `@arnilo/prism-browser` for ordinary public retrieval; use browser automation only for interactive/authenticated/JavaScript-heavy work behind a host egress proxy. Arbitrary HTML execution, model-selected providers, automatic OAuth forwarding, and generic web/MCP passthrough are unsupported.
71
+
72
+ ## Related APIs
73
+
74
+ - [Tools](tools.md): registry, validation, permission, trust, guardrails, and ledger dispatch.
75
+ - [Credential storage](credential-storage.md): explicit resolver composition and environment mapping.
76
+ - [Host security](host-security.md): SSRF, untrusted-content, and secret boundaries.
77
+ - [MCP tools](mcp-tools.md): hardened prototype path for official vendor MCP servers.
78
+ - [Performance and resource limits](performance.md): operational ceilings and benchmark evidence.
package/docs/workflows.md CHANGED
@@ -290,6 +290,7 @@ Use workflows for known, durable, replayable graphs. Use optional supervisor del
290
290
  - [Agent/session runtime](agent-session-runtime.md): `AgentSession.run()`/`stream()`, abort, subscribe
291
291
  - [Guardrails](guardrails.md): `RunWorkflowOptions.guardrails` routes tool nodes through core dispatch before policy and side effects.
292
292
  - [Supervisor delegation](supervisors.md): bounded dynamic child selection.
293
+ - [A2A interoperability](a2a.md): hosts may adapt existing exact-owner workflow status/list/cancel/checkpoint/event surfaces to `A2ATaskLifecycle`; A2A package adds no workflow worker, queue, or schema.
293
294
  - [Agent events](agent-events.md): core `AgentEvent` wrapped by `agent_event`
294
295
  - [Session stores and branching](session-stores-and-branching.md): session `leafId` reuse on resume
295
296
  - [CLI/RPC](cli-rpc.md): host control seam; wire `createWorkflowCommands()` into `runRpcServer`
@@ -298,4 +299,5 @@ Use workflows for known, durable, replayable graphs. Use optional supervisor del
298
299
  - [PostgreSQL persistence](postgres-persistence.md): durable `persistence.checkpoints`
299
300
  - [Observability](observability.md): exporting workflow/agent events
300
301
  - [Coding execution approval and sandboxing](coding-security.md): `ExecutionPolicy` for tool nodes
302
+ - [Coding agent tools](coding-agent-tools.md): opt-in `createGitTools()` / `git_pr_handoff` produce bounded host-owned PR payloads; durable coding plans/todos are workspace Markdown plus `state.coding` metadata helpers — workflows may compose them for restart/resume/background branches but Prism never pushes or opens PRs
301
303
  - [Release and install](release-and-install.md): atomic and profile installs
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@arnilo/prism",
3
- "version": "0.0.7",
3
+ "version": "0.0.10",
4
4
  "description": "Agent harness for AI providers, agents, sessions, and tools.",
5
5
  "type": "module",
6
6
  "main": "./dist/index.js",
@@ -110,14 +110,16 @@
110
110
  "packages/session-store-*",
111
111
  "packages/credentials-node",
112
112
  "packages/mcp",
113
+ "packages/evals",
113
114
  "packages/coding-agent",
114
115
  "packages/coding-security",
115
116
  "packages/workflows",
116
- "packages/evals",
117
117
  "packages/memory",
118
118
  "packages/rag",
119
119
  "packages/server",
120
120
  "packages/supervisor",
121
+ "packages/web-tools",
122
+ "packages/browser",
121
123
  "packages/prism-*"
122
124
  ],
123
125
  "scripts": {
@@ -133,8 +135,8 @@
133
135
  "sdk:ready": "npm run typecheck && npm test && npm run pack:dry-run"
134
136
  },
135
137
  "devDependencies": {
136
- "typescript": "^5.7.0",
137
- "@types/node": "^22.0.0"
138
+ "typescript": "^7.0.2",
139
+ "@types/node": "^26.1.1"
138
140
  },
139
141
  "engines": {
140
142
  "node": ">=20"