@arnilo/prism 0.0.4 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/CHANGELOG.md +46 -1
  2. package/README.md +34 -10
  3. package/dist/agent-loops.d.ts +1 -0
  4. package/dist/agent-loops.js +26 -16
  5. package/dist/agents.js +147 -21
  6. package/dist/cli-init.d.ts +41 -0
  7. package/dist/cli-init.js +390 -0
  8. package/dist/cli-runner.d.ts +7 -1
  9. package/dist/cli-runner.js +13 -1
  10. package/dist/content.d.ts +19 -0
  11. package/dist/content.js +197 -69
  12. package/dist/contracts.d.ts +96 -9
  13. package/dist/contracts.js +8 -0
  14. package/dist/feedback.d.ts +48 -0
  15. package/dist/feedback.js +230 -0
  16. package/dist/ids.d.ts +2 -0
  17. package/dist/ids.js +6 -0
  18. package/dist/index.d.ts +10 -4
  19. package/dist/index.js +6 -3
  20. package/dist/providers/media.d.ts +3 -1
  21. package/dist/providers/media.js +11 -1
  22. package/dist/session-stores.js +2 -3
  23. package/dist/testing/feedback.d.ts +6 -0
  24. package/dist/testing/feedback.js +37 -0
  25. package/dist/testing/persistence-schema.d.ts +48 -10
  26. package/dist/testing/persistence-schema.js +166 -22
  27. package/dist/testing/run-ledger-conformance.js +7 -1
  28. package/dist/thinking.d.ts +42 -0
  29. package/dist/thinking.js +92 -0
  30. package/dist/tools.js +2 -3
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +75 -0
  34. package/docs/agent-events.md +14 -21
  35. package/docs/agent-loops.md +12 -9
  36. package/docs/agent-session-runtime.md +14 -16
  37. package/docs/cli-rpc.md +35 -7
  38. package/docs/coding-agent-tools.md +35 -14
  39. package/docs/coding-security.md +7 -3
  40. package/docs/compaction-llm.md +17 -7
  41. package/docs/compaction-observational-memory.md +30 -4
  42. package/docs/context-and-skills.md +1 -0
  43. package/docs/credential-storage.md +58 -9
  44. package/docs/credentials-and-redaction.md +3 -3
  45. package/docs/database-persistence.md +17 -9
  46. package/docs/evaluations.md +122 -0
  47. package/docs/extensions.md +2 -2
  48. package/docs/host-security.md +26 -5
  49. package/docs/index.md +43 -28
  50. package/docs/mcp-tools.md +74 -13
  51. package/docs/migration.md +177 -3
  52. package/docs/multimodal-content.md +14 -6
  53. package/docs/node-filesystem-config.md +1 -0
  54. package/docs/node-jsonl-session-store.md +5 -4
  55. package/docs/observability.md +14 -6
  56. package/docs/performance.md +209 -0
  57. package/docs/postgres-persistence.md +8 -6
  58. package/docs/provider-caching.md +16 -4
  59. package/docs/provider-conformance.md +40 -1
  60. package/docs/provider-packages.md +62 -3
  61. package/docs/providers/ai-sdk.md +149 -0
  62. package/docs/providers/kimi.md +124 -61
  63. package/docs/providers/neuralwatt.md +19 -13
  64. package/docs/providers/openai.md +56 -13
  65. package/docs/providers/opencode-go.md +118 -30
  66. package/docs/providers/openrouter.md +105 -35
  67. package/docs/providers/zai.md +94 -45
  68. package/docs/public-contracts.md +6 -5
  69. package/docs/rag.md +113 -0
  70. package/docs/release-and-install.md +100 -79
  71. package/docs/review-coverage-2026-07-15.md +193 -0
  72. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  73. package/docs/runs-and-usage.md +42 -5
  74. package/docs/server.md +139 -0
  75. package/docs/settings-auth-trust-security.md +5 -5
  76. package/docs/sqlite-persistence.md +6 -5
  77. package/docs/structured-output.md +1 -1
  78. package/docs/supervisors.md +71 -0
  79. package/docs/thinking-and-reasoning.md +98 -0
  80. package/docs/tool-execution-primitives.md +3 -3
  81. package/docs/tools.md +15 -0
  82. package/docs/use-case-model-selection.md +109 -0
  83. package/docs/workflow-orchestration-primitives.md +20 -3
  84. package/docs/workflows.md +114 -33
  85. package/docs/working-and-semantic-memory.md +170 -0
  86. package/package.json +13 -3
  87. package/templates/init/README.md.tmpl +28 -0
  88. package/templates/init/env.example.tmpl +1 -0
  89. package/templates/init/gitignore.tmpl +11 -0
  90. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  91. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  92. package/templates/init/package.json.tmpl +22 -0
  93. package/templates/init/providers.json +76 -0
  94. package/templates/init/src/agent.ts.tmpl +10 -0
  95. package/templates/init/src/index.ts.tmpl +12 -0
  96. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  97. package/templates/init/tsconfig.json.tmpl +15 -0
package/docs/server.md ADDED
@@ -0,0 +1,139 @@
1
+ # Web-standard server handler
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-server` exposes explicitly selected agents and workflows through one framework-free `(Request) => Promise<Response>` handler. It supports direct agent results, bounded agent/workflow SSE, durable workflow start/enqueue/status/cancel/resume/replay, ownership-scoped schedules, host authorization, ownership propagation, redaction, and resource ceilings.
6
+
7
+ No listener starts on import. Empty `agents`/`workflows` maps expose nothing. Authentication, authorization, route selection, durable stores, TLS, rate limiting, and framework/serverless adaptation remain host-owned.
8
+
9
+ ## When to use it
10
+
11
+ Use it when a Node 20, serverless, worker, or framework host already speaks Web `Request`/`Response` and needs a small Prism API boundary. Wrap it in the platform's native adapter rather than adding Express, Fastify, Hono, Koa, Nest, or Next to Prism.
12
+
13
+ Use `AgentSession` or workflow APIs directly for in-process applications. Do not treat this package as an auth provider, user database, firewall, durable agent-result store, or public listener.
14
+
15
+ ## Inputs / request
16
+
17
+ ```ts
18
+ const handler = createPrismHandler({
19
+ agents?: Record<string, Agent | PrismAgentExposure>,
20
+ workflows?: Record<string, PrismWorkflowExposure>,
21
+ schedules?: WorkflowSchedules | ((authorization, signal) => WorkflowSchedules),
22
+ authorize: async ({ request, operation, capabilityId }) => false | {
23
+ ownership: { tenantId?: string; accountId?: string; userId?: string },
24
+ metadata?: Record<string, unknown>,
25
+ },
26
+ basePath?: "/prism",
27
+ allowedHosts?: string[],
28
+ allowedOrigins?: string[],
29
+ redactor?: SecretRedactor,
30
+ limits?: PrismServerLimits,
31
+ disconnectAborts?: boolean,
32
+ });
33
+ ```
34
+
35
+ At least one non-empty ownership field must come from `authorize()`. Request JSON never chooses ownership.
36
+
37
+ | Method and route | Authorization operation | Body |
38
+ | --- | --- | --- |
39
+ | `POST /prism/agents/:id/runs` | `agent.run` | `{ "input": string | Message | Message[] }` |
40
+ | `POST /prism/agents/:id/stream` | `agent.stream` | same; SSE response |
41
+ | `POST /prism/workflows/:id/runs` | `workflow.run` | `{ "input": unknown, "runId"?: string }` |
42
+ | `POST /prism/workflows/:id/stream` | `workflow.stream` | same; SSE response |
43
+ | `POST /prism/workflows/:id/enqueue` | `workflow.enqueue` | `{ "input": unknown, "runId"?: string }`; returns `202` queued handle |
44
+ | `GET /prism/workflows/:id/runs/:runId` | `workflow.status` | none |
45
+ | `DELETE /prism/workflows/:id/runs/:runId` | `workflow.cancel` | none |
46
+ | `POST /prism/workflows/:id/runs/:runId/resume` | `workflow.resume` | `{ "decision": "approve" | "deny", "input"?: unknown, "expectedVersion": number }` |
47
+ | `POST /prism/workflows/:id/runs/:runId/replay` | `workflow.replay` | `{ "fromNodeId": string, "runId"?: string }` |
48
+ | `POST /prism/schedules/:id` | `schedule.create` | `{ "workflowId", "nextRunAt", "input"?, "intervalMs"?, "calculatorId"?, "paused"?, "metadata"? }` |
49
+ | `GET /prism/schedules?status=&cursor=&limit=` | `schedule.list` | none |
50
+ | `POST /prism/schedules/:id/pause` | `schedule.pause` | `{}` |
51
+ | `POST /prism/schedules/:id/resume` | `schedule.resume` | `{ "nextRunAt"?: string }` |
52
+ | `POST /prism/schedules/:id/trigger` | `schedule.trigger` | `{ "idempotencyKey": string }` |
53
+ | `DELETE /prism/schedules/:id` | `schedule.delete` | none |
54
+
55
+ POST routes require `Content-Type: application/json`. Capability/run IDs are bounded URL-safe identifiers. A custom `PrismAgentExposure.sessionFactory` can build sessions from authorized host context; otherwise an `Agent` creates a fresh session.
56
+
57
+ ## Outputs / response / events
58
+
59
+ Direct routes return bounded JSON. Stream routes return `text/event-stream`; every event is one `data: <AgentEvent|WorkflowEvent>` frame. Status returns the ownership-scoped durable checkpoint record. Resume uses Phase 8 expected-version CAS. Cancel aborts active work or marks eligible durable checkpoints aborted.
60
+
61
+ Errors use `{ "error": { "code", "message" } }`. Unknown routes/capabilities are `404`, authorization/policy denial `403`, malformed input `400`, unsupported content type `415`, body overflow `413`, concurrency overflow `429`, and result overflow `507`. Unexpected errors are generic and never include stacks.
62
+
63
+ ## Request/response example
64
+
65
+ ```json
66
+ {
67
+ "request": { "method": "POST", "path": "/prism/agents/support/runs", "body": { "input": "Summarize this" } },
68
+ "response": { "status": "succeeded", "sessionId": "...", "runId": "...", "text": "Summary" }
69
+ }
70
+ ```
71
+
72
+ ## Implementation example
73
+
74
+ ```ts
75
+ import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
76
+ import { createPrismHandler } from "@arnilo/prism-server";
77
+
78
+ const agent = createAgent({
79
+ model: { provider: "mock", model: "offline" },
80
+ provider: createMockProvider([providerTextDelta("ready"), providerDone()]),
81
+ });
82
+
83
+ const handler = createPrismHandler({
84
+ agents: { support: agent },
85
+ authorize: async ({ request }) => request.headers.get("authorization") === "Bearer host-validated"
86
+ ? { ownership: { tenantId: "tenant-1", userId: "user-1" } }
87
+ : false,
88
+ allowedHosts: ["api.example.test"],
89
+ allowedOrigins: ["https://app.example.test"],
90
+ });
91
+
92
+ // Cloudflare/Bun/Deno-style: export { handler as fetch }.
93
+ // Node/framework hosts adapt their request to Web Request and return Web Response.
94
+ ```
95
+
96
+ ## Extension and configuration notes
97
+
98
+ - `basePath` defaults to `/prism`; URL root exposure is rejected.
99
+ - Agent maps and workflow maps are immutable host selections. No registry/package discovery runs.
100
+ - Workflow exposure requires its existing `WorkflowCheckpointAdapter`; no server-owned database exists.
101
+ - Schedule exposure is optional and may be one service or an authorization-selected resolver. Returned service ownership must exactly match authorized tenant/account/user scope; otherwise request is forbidden.
102
+ - `PrismWorkflowExposure.runOptions` can supply agent/tool/policy/resume-validator wiring. Server-owned ownership, signal, checkpoint, redactor, run ID, and event bus fields cannot be overridden.
103
+ - Host/origin checks and CORS headers activate only when their allow-lists are configured. Hosts still own reverse-proxy trust and canonical host handling.
104
+
105
+ Default/hard ceilings:
106
+
107
+ | Limit | Default | Hard cap |
108
+ | --- | ---: | ---: |
109
+ | JSON request | 64 KiB | 1 MiB |
110
+ | direct response | 1 MiB | 8 MiB |
111
+ | SSE event | 64 KiB | 1 MiB |
112
+ | SSE total | 10 MiB | 64 MiB |
113
+ | SSE event count | 10,000 | 100,000 |
114
+ | concurrent runs | 16 | 256 |
115
+ | subscriber queue | 128 | 4,096 |
116
+ | request/run timeout | 120 s | 30 min |
117
+
118
+ ## Security and performance notes
119
+
120
+ - `authorize()` is required and runs for every matched operation before capability lookup or body execution. Return `false` on missing/invalid credentials. Do not trust caller ownership fields.
121
+ - Use authorization metadata only for non-secret audit context. Never put credentials in metadata, input, route IDs, run IDs, checkpoints, events, or responses.
122
+ - Configure `SecretRedactor` before runs. Redaction matches known secrets; it is not DLP.
123
+ - Agent tools and workflow tool nodes still need their own `PermissionPolicy`, `ToolValidator`, and `ExecutionPolicy`. HTTP authorization does not replace side-effect policy.
124
+ - Host and origin allow-lists are exact string matches. Configure reverse-proxy normalization, TLS, rate limiting, IP policy, CSRF/cookie policy, and authentication outside Prism.
125
+ - SSE uses bounded upstream subscriber queues. Consumer cancellation aborts owned work by default and releases concurrency; set `disconnectAborts: false` only when the host deliberately owns background completion.
126
+ - Source inputs/resource URLs remain host responsibilities and use existing resource/media SSRF policies. Server package does not fetch URLs.
127
+ - Schedule routes never accept ownership from JSON. Services carry mandatory ownership and explicit workflow/calculator registries; route authorization cannot broaden either. Replay applies workflow ownership/hash/approval checks.
128
+ - No agent status/reconnect store is invented. Durable reconnect/status/resume is the workflow path; persistent agent run querying remains a host persistence API.
129
+
130
+ A2A routes are not added to `createPrismHandler()`. Install `@arnilo/prism-supervisor` and explicitly mount `createA2AHandler()` when protocol interoperability is required; this keeps cards and remote invoke absent from ordinary Prism servers.
131
+
132
+ ## Related APIs
133
+
134
+ - [Agent/session runtime](agent-session-runtime.md): direct result and event stream semantics.
135
+ - [Workflows](workflows.md): durable checkpoints, status, cancellation, and exact-once resume.
136
+ - [MCP client and server exposure](mcp-tools.md): selected MCP capabilities and web-standard MCP transport.
137
+ - [Host security guide](host-security.md): remote-boundary checklist.
138
+ - [A2A interoperability](a2a.md): separately mounted A2A 1.0 handler/client.
139
+ - [Release and install](release-and-install.md): optional package installation and profiles.
@@ -18,7 +18,7 @@ Use these APIs when a host wants one explicit place to compose settings, resolve
18
18
  - Node-only subpaths: `@arnilo/prism/node/settings` for caller-named JSON settings files and `@arnilo/prism/node/trust` for explicit trusted path roots with symlink-aware realpath checks.
19
19
 
20
20
  ## Outputs / response / events
21
- Settings and credential helpers return existing `SettingsProvider` and `CredentialResolver` contracts. `AgentConfig.settings` and `AgentConfig.credentials` are host-owned metadata for compatibility; `createAgent()` / `session.run()` do not call `settings.get()` or `credentials.resolve()`. Permission denial blocks tool execution, extension setup, and resource loader calls before side effects. A configured `AgentConfig.redactor` or `RunOptions.redactor` redacts provider requests, emitted `AgentEvent` payloads, stored `SessionEntry` values, and runtime `InstructionContext` input/history seen by instruction injectors.
21
+ Settings and credential helpers return existing `SettingsProvider` and `CredentialResolver` contracts. These seams are host-owned outside `AgentConfig`; `createAgent()` / `session.run()` do not call `settings.get()` or `credentials.resolve()`. Permission denial blocks tool execution, extension setup, and resource loader calls before side effects. A configured `AgentConfig.redactor` or `RunOptions.redactor` redacts provider requests, emitted `AgentEvent` payloads, stored `SessionEntry` values, and runtime `InstructionContext` input/history seen by instruction injectors.
22
22
 
23
23
  ## Request/response example
24
24
  ```ts
@@ -47,17 +47,17 @@ const apiKey = await resolveCredentialValue(credentials, { name: "api", provider
47
47
 
48
48
  const agent = createAgent({
49
49
  model: { provider: "demo", model: "model" },
50
- // host-owned metadata; runtime does not read/resolve these fields
51
- settings,
52
- credentials,
50
+ // Resolve credentials at the provider edge; register known secrets for redaction.
53
51
  redactor: apiKey ? createSecretRedactor([apiKey]) : undefined,
54
52
  });
53
+ void settings;
54
+ void credentials;
55
55
  void trust;
56
56
  void agent;
57
57
  ```
58
58
 
59
59
  ## Extension and configuration notes
60
- Root imports stay filesystem-free. Node settings files are caller-named and read once; optional missing files are skipped. Trust storage, prompts, approval UI, OAuth token storage, environment-variable selection, and persistent credentials belong in the host or an extension package. For Node.js hosts, [`@arnilo/prism-credentials-node`](credential-storage.md) provides encrypted-file and system-keychain backends. Passing `settings` / `credentials` on `AgentConfig` does not wire hidden runtime reads; hosts pass concrete values or resolvers to the provider/request edge that needs them.
60
+ Root imports stay filesystem-free. Node settings files are caller-named and read once; optional missing files are skipped. Trust storage, prompts, approval UI, OAuth token storage, environment-variable selection, and persistent credentials belong in the host or an extension package. For Node.js hosts, [`@arnilo/prism-credentials-node`](credential-storage.md) provides encrypted-file and system-keychain backends. Pass concrete settings values or credential resolvers to the provider/request edge that needs them; do not place them on `AgentConfig`.
61
61
 
62
62
  ## Security and performance notes
63
63
  Prism does not sandbox host tools or extensions. Prism does not read environment variables, keychains, user config files, package manifests, resources, settings providers, credential resolvers, or project-local extensions unless the host explicitly wires those operations. Redaction is exact known-secret replacement only; it is not secret detection. Permission and trust checks are one operation per guarded call and add no workers, watchers, retries, network, or filesystem scans.
@@ -37,6 +37,7 @@ import { createSqlitePersistence } from "@arnilo/prism-session-store-sqlite";
37
37
  | `filename` | `string` | SQLite database path. Use `:memory:` for ephemeral tests. |
38
38
  | `wal` | `boolean` | Enable WAL journal mode. Defaults to `true`. |
39
39
  | `busyTimeoutMs` | `number` | SQLite `busy_timeout` in milliseconds. Defaults to `5000`. |
40
+ | `feedbackRedactor` | `SecretRedactor` | Optional redaction for feedback comment/tags/metadata before insert. |
40
41
  | `fileMode` | `number` | Unix file mode for newly created database files. Defaults to `0o600`. |
41
42
  | `database` | `Database` | Advanced: supply an existing `better-sqlite3` handle (caller owns lifecycle). |
42
43
 
@@ -51,11 +52,11 @@ import { createSqlitePersistence } from "@arnilo/prism-session-store-sqlite";
51
52
  | `SessionStore.readBranchPath` | Recursive ancestor query from `leafId` (or latest leaf) in root→leaf order. |
52
53
  | `RunLedger.append*` | Inserts run/event/tool/usage rows; events receive monotonic per-run `sequence` values. |
53
54
  | `ProductionPersistenceStore.query*` | Parameterized cursor pagination on indexed columns. |
54
- | `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, and bounded pagination. |
55
+ | `checkpoints` | Generic versioned `CheckpointStore` backed by `prism_checkpoints`; ownership, CAS/fencing checks, bounded pagination, and workflow suspended/denied/schedule/state/replay values without a schema migration. |
55
56
  | `leases` | Atomic `LeaseStore` backed by `prism_leases`; database-clock expiry, opaque renew/release token, monotonic takeover fence. |
56
57
  | `close()` | Closes the underlying database when the adapter opened it. |
57
58
 
58
- Migrations run automatically on open and are idempotent across reopen.
59
+ Migrations run automatically on open and are idempotent across reopen. Under the SQLite migration transaction, startup checks ordered contract name/version/SHA-256 rows plus full schema-v3 PRAGMA/catalog shape (all required tables, columns/types/nullability/defaults, PK/unique/FK keys, and named indexes) before any runtime write. A complete legacy 0.0.5 history with all `checksum` values `NULL` is shape-verified then backfilled transactionally once. Unknown, duplicate, out-of-order, partial-legacy, checksum, or shape drift rejects open; restore or apply reviewed DDL rather than editing migration rows.
59
60
 
60
61
  ## Request/response example
61
62
 
@@ -98,7 +99,7 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
98
99
  - The package is optional and workspace-local; `@arnilo/prism` core has no SQLite dependency.
99
100
  - Hosts choose the database path and own backup, retention enforcement, and filesystem permissions.
100
101
  - `SessionAppendOptions` idempotency rows are durable in `prism_session_append_idempotency` and survive reopen.
101
- - Schema version **1** (`001_init`) matches `@arnilo/prism/testing/persistence-schema` — PostgreSQL adapters share the same model with dialect-local DDL.
102
+ - Schema version **3** applies `001_init`, additive `002_usage_scope`, and `003_run_feedback`. Migration 003 adds immutable `prism_run_feedback` rows with run FK/cascade deletion and owner/run/trace cursor indexes. `persistence.feedback` validates exact run ownership, bounds/redacts through optional `feedbackRedactor`, queries bounded pages, and deletes only exact-owned IDs. PostgreSQL shares the same model with dialect-local DDL.
102
103
  - Pass an existing `better-sqlite3` `Database` via `database` when your host already manages connections.
103
104
 
104
105
  ## Security and performance notes
@@ -108,7 +109,7 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
108
109
  - **No path interpolation.** The adapter opens exactly the caller-supplied `filename`; it does not expand environment variables or discover paths.
109
110
  - **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
110
111
  - **WAL + busy timeout.** WAL is enabled by default; busy timeout defaults to 5 seconds. This meets the Plan 056 local workload target but SQLite still serializes writers — prefer PostgreSQL for high write concurrency.
111
- - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans.
112
+ - **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid whole-database scans. Startup validation reads SQLite catalog/PRAGMA metadata only, never application rows.
112
113
  - **Tenant isolation.** `tenant_id` / `account_id` / `user_id` columns on run and ownership tables participate in query filters; hosts must still scope writes correctly.
113
114
 
114
115
  ## Related APIs
@@ -118,5 +119,5 @@ For resume/timeline flows, use `queryRuns`, `queryEvents`, `queryToolCalls`, and
118
119
  - [Run ledger conformance](run-ledger-conformance.md): `assertRunLedgerConforms` / `runRunLedgerConformance`.
119
120
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): package matrix and threat model.
120
121
  - [Node JSONL session store](node-jsonl-session-store.md): dev-only single-process alternative.
121
- - [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` for durable multi-process workflow execution.
122
+ - [Workflows](workflows.md): adapt `persistence.checkpoints` and pass `persistence.leases` to `createWorkflowCoordinator()` and `createWorkflowSchedules()` for durable background execution and schedules.
122
123
  - [Migration guide](migration.md): moving from JSONL/in-memory to database-backed persistence.
@@ -234,7 +234,7 @@ Key cross-seam points:
234
234
  - The default parser treats assistant text as the value (`{ ok: true, value: text }`); supply a host parser whenever `T` is not `string`.
235
235
  - The default repairer builds a user message from `validation.errors[].message`; supply a host repairer for schema-specific guidance.
236
236
  - `maxRevisions` (default 3) bounds revision turns; budget exhaustion ends the loop and emits `artifact_failed` (it does not throw).
237
- - Tools are not dispatched in revision turns. Hosts needing tools in artifact turns use `singleShotLoop` or a custom loop.
237
+ - Tools are inert in artifact turns unless `loop.toolCalls: "bounded"` is explicit. Bounded mode uses run-global `maxToolRounds`, dispatches calls sequentially through normal runtime guards, skips parser/validator for tool-calling responses, and permits at most `1 + maxRevisions + maxToolRounds` provider turns. An extra tool response yields terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"` and executes nothing.
238
238
 
239
239
  ## Security and performance notes
240
240
 
@@ -0,0 +1,71 @@
1
+ # Supervisor delegation
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-supervisor` adds optional runtime-selected delegation to an explicit local child allow-list. It returns normal `AgentRunResult` values and does not modify core `createAgent()` or deterministic workflows.
6
+
7
+ ## When to use it
8
+
9
+ Use a supervisor when a host or agent must choose a child dynamically. Use `@arnilo/prism-workflows` for known DAGs, durable checkpoints, schedules, replay, or human suspension.
10
+
11
+ ## Inputs / request
12
+
13
+ | API/field | Meaning |
14
+ | --- | --- |
15
+ | `createSupervisor({ ownership, children })` | Creates one ownership-scoped supervisor. |
16
+ | `SupervisorChild.createAgent(context)` | Child-owned factory; receives derived resource/thread IDs, narrowed permission, abort signal, and nested `delegate`. |
17
+ | `delegate({ childId, input, threadId?, limits?, signal? })` | Invokes one allow-listed child. Input is text and byte-bounded. |
18
+ | `hooks.before` | May reject, modify redacted input, or narrow limits/policy. |
19
+ | `hooks.after` | Observes redacted terminal summary; failures cannot alter settled result. |
20
+ | `limits` | Depth 4/16, active children 4/32, input 64 KiB/1 MiB, steps 8/64, tools 32/256, tokens 20k/1m, timeout 60s/30m, event queue 128/4096 default/hard. |
21
+
22
+ ## Outputs / response / events
23
+
24
+ `delegate()` returns the child's `AgentRunResult` or throws its `AgentRunError`/a supervisor denial or limit error. `subscribe()` emits bounded `delegation_started`, `delegation_finished`, `delegation_rejected`, and `delegation_error` metadata events.
25
+
26
+ ## Request/response example
27
+
28
+ ```json
29
+ {"childId":"research","input":"Check primary sources","limits":{"maxTokens":4000}}
30
+ ```
31
+
32
+ ## Implementation example
33
+
34
+ ```ts
35
+ import { createSupervisor } from "@arnilo/prism-supervisor";
36
+
37
+ const supervisor = createSupervisor({
38
+ ownership: { tenantId: "tenant", userId: "user" },
39
+ permission: parentPolicy,
40
+ children: {
41
+ research: {
42
+ permission: readOnlyPolicy,
43
+ createAgent: ({ resourceId, threadId, permission, delegate }) =>
44
+ createResearchAgent({ resourceId, threadId, permission, delegate }),
45
+ },
46
+ },
47
+ hooks: { before: ({ input }) => ({ input, limits: { maxTokens: 4000 } }) },
48
+ });
49
+
50
+ const result = await supervisor.delegate({ childId: "research", input: "Check sources" });
51
+ ```
52
+
53
+ ## Extension and configuration notes
54
+
55
+ Child factories resolve their own providers/credentials and construct context/memory using the supplied IDs. Parent, child, returned-agent, budget, and hook permission policies are AND-composed. Child/request/hook limits can only lower inherited limits. A nested factory can call the supplied `delegate()`; immutable path state rejects cycles and depth overflow.
56
+
57
+ ## Security and performance notes
58
+
59
+ - Child IDs are explicit; no package/provider discovery occurs.
60
+ - `resourceId` and `threadId` include supervisor/delegation/child identity. Do not replace them with parent memory IDs.
61
+ - Tool budget is checked before side effects. Token usage is enforced on terminal aggregate usage and can exceed by at most one provider turn because providers report tokens after generation.
62
+ - Abort and timeout cover hooks, child creation, nested delegation, and the run. Host child code must cooperate with `AbortSignal`.
63
+ - Redaction applies before hook input, run metadata/results, completion hooks, and events. Child credentials are never supplied in delegation context.
64
+ - Static workflows remain smaller and more reproducible for known graphs.
65
+
66
+ ## Related APIs
67
+
68
+ - [A2A interoperability](a2a.md): remote protocol boundary.
69
+ - [Workflows](workflows.md): preferred deterministic orchestration.
70
+ - [Working and semantic memory](working-and-semantic-memory.md): child scope construction.
71
+ - [Host security](host-security.md): permission and credential boundaries.
@@ -0,0 +1,98 @@
1
+ # Thinking and reasoning
2
+
3
+ ## What it does
4
+
5
+ Prism keeps thinking/reasoning **provider-owned on the wire** while giving hosts one portable way to set effort per turn. Model defaults live on `ModelConfig.compat` (and `capabilities.reasoning` where declared). Per-turn overrides live on `ProviderRequestOptions.compat` and win through existing `mergeProviderRequestOptions`. Shared helpers map a portable `ThinkingLevel` into the official compat fields each family already reads — they do **not** invent a second options tree.
6
+
7
+ ## When to use it
8
+
9
+ - Session runs: pass `providerOptions.compat` (or `applyThinkingLevel`) on `RunOptions`.
10
+ - Use-case workers (LLM compaction, observational memory): pass `thinkingLevel`; packages map it into `compat` via the shared helpers.
11
+ - Provider authors: keep reading official fields from `options.compat` / `model.compat`; add package-local escape hatches only when the official API has unique knobs.
12
+
13
+ ## Contract
14
+
15
+ | Layer | Surface |
16
+ | --- | --- |
17
+ | Model default | `ModelConfig.compat` (+ `capabilities.reasoning` when the model can reason) |
18
+ | Per-turn override | `ProviderRequestOptions.compat` (request wins over model via merge) |
19
+ | Portable level | `ThinkingLevel`: `none` \| `minimal` \| `low` \| `medium` \| `high` \| `xhigh` \| `max` |
20
+ | Helpers | `thinkingCompatFor`, `applyThinkingLevel`, `thinkingFamilyForModel`, `isThinkingLevel`, `normalizeThinkingLevel`, `THINKING_LEVELS` |
21
+ | Not used | Inert `options.extra.thinkingLevel` — providers do not read `extra` for effort |
22
+
23
+ ```ts
24
+ import { applyThinkingLevel, thinkingCompatFor, thinkingFamilyForModel } from "@arnilo/prism";
25
+
26
+ // Per-turn override on a session run (OpenAI / OpenRouter family)
27
+ await session.run(input, {
28
+ providerOptions: applyThinkingLevel(undefined, "low", "openai_reasoning"),
29
+ });
30
+
31
+ // Equivalent explicit compat
32
+ await session.run(input, {
33
+ providerOptions: { compat: thinkingCompatFor("openai_reasoning", "low") },
34
+ // → { reasoning: { effort: "low" } }
35
+ });
36
+
37
+ // Use-case worker: family from model metadata (or pass an explicit family)
38
+ const family = thinkingFamilyForModel(model);
39
+ await runObserver({
40
+ ...,
41
+ providerOptions: applyThinkingLevel(base, "low", family === "noop" ? "reasoning_effort" : family),
42
+ });
43
+ ```
44
+
45
+ ## Compat families
46
+
47
+ Core maps only shapes shared by ≥2 packages (or an explicit no-op). Unique knobs stay package-local.
48
+
49
+ | Family | Compat patch | Used by (official fields) |
50
+ | --- | --- | --- |
51
+ | `openai_reasoning` | `{ reasoning: { effort } }` | OpenAI Responses `reasoning.effort`; OpenRouter `reasoning.effort` |
52
+ | `reasoning_effort` | `{ reasoning_effort }` | Z.AI `reasoning_effort`; NeuralWatt `reasoning_effort`; Kimi K3 `reasoning_effort` |
53
+ | `thinking_type` | `{ thinking: { type: "enabled" \| "disabled" } }` | Z.AI `thinking.type`; Kimi K2.x `thinking.type` (`none` → `disabled`) |
54
+ | `noop` | `{}` | AI SDK / host-owned adapters — effort is host-model settings |
55
+
56
+ `applyThinkingLevel` defaults `family` to `reasoning_effort` when omitted. For `openai_reasoning`, an existing `compat.reasoning.summary` (or other reasoning keys) is preserved when merging `effort`.
57
+
58
+ ### Recommended family by first-party package
59
+
60
+ | Package | Recommended family | Notes |
61
+ | --- | --- | --- |
62
+ | `@arnilo/prism-provider-openai` | `openai_reasoning` | First-class body `reasoning` from model + per-turn compat merge; `summary`/`mode`/`context` via compat |
63
+ | `@arnilo/prism-provider-openrouter` | `openai_reasoning` | First-class `resolveOpenRouterReasoning` merge; prefer `reasoning` object over legacy `reasoning_effort` shorthand; `preserveThinking` replays as body `reasoning` |
64
+ | `@arnilo/prism-provider-zai` | `reasoning_effort` (+ optional `thinking_type`) | Official `thinking` / `reasoning_effort` / `tool_stream` / `clear_thinking`; Preserved Thinking via `reasoning_content` |
65
+ | `@arnilo/prism-provider-neuralwatt` | `reasoning_effort` | Budgets / `preserve_thinking` / `clear_thinking` / `chat_template_kwargs` stay package-local on `compat` |
66
+ | `@arnilo/prism-provider-kimi` | K3: `reasoning_effort`; K2.x: `thinking_type` | K2.7-code thinking is always on; do not send conflicting `thinking` + `reasoning_effort` |
67
+ | `@arnilo/prism-provider-opencode-go` | Anthropic route: thinking blocks (`thinking_type` family); OpenAI route: `reasoning_content` preserve + optional `thinking`/`reasoning_effort`/`reasoning` passthrough | Official dual endpoints; MiniMax/Qwen → Anthropic, others → OpenAI |
68
+ | `@arnilo/prism-provider-ai-sdk` | `noop` | Host `LanguageModelV4` owns reasoning settings |
69
+
70
+ `thinkingFamilyForModel` infers family from existing `compat` shape, then safe provider heuristics (`openai*` → `openai_reasoning`, `neuralwatt` → `reasoning_effort`), then `capabilities.reasoning` → `reasoning_effort`, else `noop`. Docs and packages may map other provider ids explicitly; core avoids provider-specific literals beyond those heuristics.
71
+
72
+ ## Merge order
73
+
74
+ 1. `ModelConfig.compat` / model defaults inside the provider
75
+ 2. `ProviderRequestOptions.compat` from agent / session policies
76
+ 3. Per-turn `RunOptions.providerOptions` or use-case `applyThinkingLevel` patch (wins)
77
+
78
+ Providers already prefer `request.options.compat.*` over `request.model.compat.*`.
79
+
80
+ ## Use-case workers
81
+
82
+ LLM compaction and observational memory accept `thinkingLevel?: string`. They call `applyThinkingLevel` into `compat` (not `extra.thinkingLevel`). When model inference returns `noop`, an explicit `thinkingLevel` still falls back to `reasoning_effort` so the host setting is never inert. Model selection for those workers (including session-model fallback) is documented in [Use-case model selection](use-case-model-selection.md).
83
+
84
+ ## Non-reasoning models
85
+
86
+ - Helper with `noop`: returns options unchanged — no invented body fields.
87
+ - Helper with a real family on a model that rejects the field: provider/API error — hosts should gate on `capabilities.reasoning` or package docs.
88
+ - `thinking_type` + `none` sets `{ type: "disabled" }`; other levels set `{ type: "enabled" }` without encoding effort (compose with `reasoning_effort` when the API supports both).
89
+
90
+ ## Related pages
91
+
92
+ - [Use-case model selection](use-case-model-selection.md) — session vs worker/summary model binding
93
+ - [Provider packages](provider-packages.md) — package boundaries and discovery
94
+ - [Provider caching](provider-caching.md) — cache retention can disable thinking on some providers (e.g. Z.AI when `cacheRetention: "none"`)
95
+ - [Provider request policies](provider-request-policies.md) — `mergeProviderRequestOptions`
96
+ - [Agent/session runtime](agent-session-runtime.md) — prior-reasoning preservation across turns
97
+ - Per-provider pages under [docs/providers](providers/)
98
+ - Evidence matrix: [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md)
@@ -123,8 +123,8 @@ Package performs **no** `PermissionPolicy`, `ToolValidator`, or trust checks of
123
123
  | --- | --- |
124
124
  | `ToolDefinition.parameters` | Stored and forwarded to providers; **not validated** by core |
125
125
  | `ToolValidator` | Host function hook; Phase 25 threads through agent runtime |
126
- | Standards-based schema validation | **Not shipped** — capability gap C-001 |
127
- | Schema compile cache | **None** — every dispatch would re-validate if host validator is naive |
126
+ | Standards-based schema validation | Optional `@arnilo/prism-tool-validator-json-schema`; host wires it through `ToolValidator` |
127
+ | Schema compile cache | Adapter-owned finite LRU; core never compiles schemas |
128
128
 
129
129
  ### MCP mapping (shipped — Task 3)
130
130
 
@@ -186,7 +186,7 @@ import { createJsonSchemaToolArgumentValidator } from "@arnilo/prism-tool-valida
186
186
  createAgent({ model, validator: createJsonSchemaToolArgumentValidator() });
187
187
  ```
188
188
 
189
- **Cache key:** stable `JSON.stringify(schema)` in adapter-owned `Map`. **Bounds:** configurable depth/properties/string/array limits before Ajv validation. **Security:** remote `$ref` rejected; prototype-pollution keys rejected in schemas and instances.
189
+ **Cache key:** stable `JSON.stringify(schema)` in an adapter-owned 256-entry LRU (hard cap 1,024); eviction removes the matching Ajv schema. **Bounds:** schemas default to 256 KiB, depth 64, 10,000 properties/keywords, and 128 refs (hard 1 MiB/128/100,000/1,024); instance depth/properties/string/array limits remain configurable. Every limit rejects non-finite, unsafe, zero/negative, and above-hard values. **Security:** only fragment-local `$ref` is accepted; prototype-pollution keys, cycles, and non-finite schema numbers reject before Ajv compilation.
190
190
 
191
191
  ### Task 2 — Parallel tool execution — **shipped**
192
192
 
package/docs/tools.md CHANGED
@@ -154,6 +154,10 @@ const agent = createAgent({ model, provider, tools: activeTools, permission, val
154
154
 
155
155
  Need different tools for one request? Build a short-lived agent/session with a narrower registry, or block extra calls with `PermissionPolicy` / `RunOptions.validate`. No extra per-run tool API exists yet; add one only when host apps need it.
156
156
 
157
+ ### Artifact-loop tools
158
+
159
+ `generate-validate-revise` treats provider tools as inert by default. Set `loop.toolCalls: "bounded"` and `RunOptions.maxToolRounds` only when an artifact needs a host-owned lookup before its next candidate. Each response with one-or-more calls consumes one shared round, dispatches calls sequentially through this exact `dispatchToolCall()` path, persists assistant-call then result transcript rows, and skips artifact parsing/validation for that response. A post-limit call executes nothing; the loop emits `artifact_failed` with `metadata.reason: "tool_round_limit"`. Tools do not consume `maxRevisions`, and tool schemas/context never grant authority.
160
+
157
161
  ### Runtime-supplied validators
158
162
 
159
163
  `AgentConfig.validator?` and `RunOptions.validate?` expose the same `ToolValidator` seam that `DispatchToolCallOptions.validate` already uses. The runtime threads `validate: RunOptions.validate ?? AgentConfig.validator` into every `dispatchToolCall` it issues during the tool loop, so an app can supply argument validation without taking ownership of dispatch itself. `RunOptions.validate` overrides `AgentConfig.validator` on a per-run basis (RunOptions wins). When neither is set, dispatch runs unmodified.
@@ -229,6 +233,17 @@ await session.run(input, {
229
233
  - Contribution registration and registry/filter calls do not perform provider calls, credential resolution, resource loading, network, filesystem discovery, or tool execution.
230
234
  - Dispatch performs explicit in-memory checks and executes only the selected host-active tool; it adds no retries, queues, timers, or new dependencies.
231
235
 
236
+ ## JSON Schema validator limits
237
+
238
+ Core stores `ToolDefinition.parameters` but does not compile schemas. Hosts that install `@arnilo/prism-tool-validator-json-schema` receive pre-Ajv schema limits: 256 KiB bytes, depth 64, 10,000 properties/keywords, 128 refs, and a 256-entry LRU compiled cache by default. All reject invalid values and have finite hard ceilings. Only fragment-local `$ref` values are accepted; non-local refs, cycles, forbidden keys, and non-finite schema numbers fail before tool execution.
239
+
240
+ ```ts
241
+ createJsonSchemaToolArgumentValidator({
242
+ maxSchemaBytes: 256 * 1024,
243
+ maxCompiledSchemas: 256,
244
+ });
245
+ ```
246
+
232
247
  ## Related APIs
233
248
 
234
249
  - [Agent/session runtime](agent-session-runtime.md): dispatches complete provider tool calls through the host-active tool harness and returns tool results on the next provider turn.
@@ -0,0 +1,109 @@
1
+ # Use-case model selection
2
+
3
+ ## What it does
4
+
5
+ Prism separates the **session chat model** (`AgentConfig.model` / `RunOptions.model`) from **use-case models** used by background or adjacent LLM jobs (observational memory workers, LLM compaction summarizers, declarative agents, supervisor children, evals). Hosts bind `{ model?, provider?, providerOptions?, thinkingLevel? }` per use case. When the use-case omits `model`, resolution falls back to the active session model. Workers never write `model_change` session entries for their own jobs.
6
+
7
+ ## When to use it
8
+
9
+ - Observational memory should run a cheaper/faster model than the chat session (or inherit the session model when unset).
10
+ - LLM compaction should summarize with an explicit `summaryModel`, falling back to a host-supplied session `model`.
11
+ - Declarative agents, supervisor children, evals, and RPC/CLI runs already own their models — document them as use-case sites that stay separate from a parent session.
12
+ - Memory/RAG `Embedder` selection is related but **not** a chat `ModelConfig` binding.
13
+
14
+ ## Contract
15
+
16
+ | Layer | Surface |
17
+ | --- | --- |
18
+ | Binding | `UseCaseModelBinding`: `{ model?, provider?, providerOptions?, thinkingLevel?, requireExplicitModel? }` |
19
+ | Resolver | `resolveUseCaseModel({ configured, sessionModel, requireExplicitModel?, … })` → `{ model, source }` or `undefined` |
20
+ | Binding helper | `resolveUseCaseModelBinding(binding, sessionModel)` |
21
+ | Credential id | `useCaseCredentialProviderId(resolved, binding?)` — always the **resolved** `model.provider` |
22
+ | Escape hatch | `requireExplicitModel: true` skips session fallback (OM historical `missing_model`) |
23
+
24
+ ```ts
25
+ import { resolveUseCaseModel, applyThinkingLevel, thinkingFamilyForModel } from "@arnilo/prism";
26
+
27
+ // Prefer an explicit worker; otherwise inherit the session/agent model.
28
+ const resolved = resolveUseCaseModel({
29
+ configured: settings.workerModel, // optional use-case ModelConfig
30
+ sessionModel: agent.config.model, // host-supplied; AgentSession does not expose agent
31
+ thinkingLevel: settings.thinkingLevel,
32
+ });
33
+ if (!resolved) {
34
+ // skip — neither configured nor session model (or requireExplicitModel)
35
+ }
36
+
37
+ const family = thinkingFamilyForModel(resolved.model);
38
+ const providerOptions = resolved.thinkingLevel
39
+ ? applyThinkingLevel(resolved.providerOptions, resolved.thinkingLevel, family === "noop" ? "reasoning_effort" : family)
40
+ : resolved.providerOptions;
41
+ ```
42
+
43
+ ### Precedence
44
+
45
+ 1. `configured` / `binding.model` → `source: "configured"`
46
+ 2. Else `sessionModel` when `requireExplicitModel` is not set → `source: "session"`
47
+ 3. Else `undefined` (package skips or throws)
48
+
49
+ Resolution is O(1) and network-free. It does not mutate session history.
50
+
51
+ ## Binding sites
52
+
53
+ | Site | How hosts bind | Session fallback |
54
+ | --- | --- | --- |
55
+ | Observational memory | `workerModel` / settings `workerModel` + runtime `sessionModel` | Yes — pass `sessionModel: agent.config.model`; `requireExplicitModel` restores skip |
56
+ | LLM compaction | `summaryModel` with `model` as fallback slot | Yes — `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` |
57
+ | `RunOptions.model` | Per-run override on the **session** | N/A — this *is* the session/run model (writes `model_change`) |
58
+ | Declarative `AgentDefinition` | Definition `model` / registry resolve | Definition-scoped (independent agent) |
59
+ | Supervisor children | Child `createSession` / child `AgentConfig.model` | Independent child session |
60
+ | Evals / workflows / RPC / CLI | Caller `runOptions.model` | Caller-owned |
61
+ | Structured output | Reuses session/run model | Same as run |
62
+ | Memory / RAG | Host `Embedder` | Not a chat model — see [Working and semantic memory](working-and-semantic-memory.md) |
63
+
64
+ ## Observational memory
65
+
66
+ ```ts
67
+ createObservationalMemoryRuntime({
68
+ session,
69
+ appendEntry: (entry) => store.append(entry),
70
+ workerProvider,
71
+ sessionModel: agent.config.model, // enables fallback when workerModel unset
72
+ // workerModel: { provider: "neuralwatt", model: "glm-5.2-fast" }, // optional override
73
+ overrides: { thinkingLevel: "low", observeAfterTokens: 1 },
74
+ });
75
+ ```
76
+
77
+ - Default: no `workerModel` + `sessionModel` set → workers use the session model.
78
+ - Explicit `workerModel` (or settings `workerModel`) always wins.
79
+ - `requireExplicitModel: true` (runtime or settings) → `skipped: "missing_model"` when no worker model, even if `sessionModel` is set.
80
+ - Neither worker nor session model → `skipped: "missing_model"`.
81
+ - Default credential request uses the **resolved** model's `provider` id.
82
+
83
+ ## LLM compaction
84
+
85
+ ```ts
86
+ createLlmCompactionStrategy({
87
+ provider: summaryProvider,
88
+ summaryModel: { provider: "example", model: "cheap-summary" }, // optional
89
+ model: agent.config.model, // session fallback when summaryModel omitted
90
+ thinkingLevel: "low",
91
+ });
92
+ ```
93
+
94
+ `summaryModel` wins; otherwise `model` is required. Thinking maps into `compat` via `applyThinkingLevel` ([Thinking and reasoning](thinking-and-reasoning.md)).
95
+
96
+ ## Security
97
+
98
+ - Credential requests for worker calls must target the **resolved** model’s provider — not ambient session credentials for a different provider unless the host wires that explicitly.
99
+ - Pass known secrets into worker/compaction options so prompts, ledger custom entries, and errors stay redacted.
100
+ - Background workers must not append `model_change` entries or otherwise rewrite the chat session’s model timeline.
101
+
102
+ ## Related pages
103
+
104
+ - [Thinking and reasoning](thinking-and-reasoning.md) — per-turn `thinkingLevel` → `compat`
105
+ - [Observational memory compaction package](compaction-observational-memory.md)
106
+ - [LLM compaction package](compaction-llm.md)
107
+ - [Agent/session runtime](agent-session-runtime.md) — `RunOptions.model` / `model_change`
108
+ - [Working and semantic memory](working-and-semantic-memory.md) — `Embedder` (non-chat)
109
+ - Evidence matrix: [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md)