@arnilo/prism 0.0.5 → 0.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/CHANGELOG.md +39 -1
  2. package/dist/agent-loops.d.ts +1 -0
  3. package/dist/agent-loops.js +27 -16
  4. package/dist/agent-run-lifecycle.d.ts +28 -0
  5. package/dist/agent-run-lifecycle.js +33 -0
  6. package/dist/agent-run-state.d.ts +53 -0
  7. package/dist/agent-run-state.js +127 -0
  8. package/dist/agents.d.ts +3 -1
  9. package/dist/agents.js +337 -46
  10. package/dist/contracts.d.ts +205 -3
  11. package/dist/contracts.js +4 -0
  12. package/dist/guardrails.d.ts +25 -0
  13. package/dist/guardrails.js +133 -0
  14. package/dist/ids.d.ts +2 -0
  15. package/dist/ids.js +6 -0
  16. package/dist/index.d.ts +17 -3
  17. package/dist/index.js +10 -3
  18. package/dist/input.js +2 -0
  19. package/dist/resources.js +2 -1
  20. package/dist/run-limits.d.ts +34 -0
  21. package/dist/run-limits.js +163 -0
  22. package/dist/secure-agent.d.ts +3 -0
  23. package/dist/secure-agent.js +63 -0
  24. package/dist/session-stores.js +2 -3
  25. package/dist/testing/persistence-schema.d.ts +45 -7
  26. package/dist/testing/persistence-schema.js +138 -24
  27. package/dist/thinking.d.ts +42 -0
  28. package/dist/thinking.js +92 -0
  29. package/dist/tools.d.ts +10 -2
  30. package/dist/tools.js +56 -7
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +4 -2
  34. package/docs/agent-events.md +23 -16
  35. package/docs/agent-loops.md +19 -8
  36. package/docs/agent-session-runtime.md +33 -1
  37. package/docs/coding-agent-tools.md +33 -12
  38. package/docs/coding-security.md +2 -2
  39. package/docs/compaction-llm.md +17 -7
  40. package/docs/compaction-observational-memory.md +28 -4
  41. package/docs/credential-storage.md +58 -9
  42. package/docs/credentials-and-redaction.md +1 -1
  43. package/docs/database-persistence.md +8 -3
  44. package/docs/guardrails.md +75 -0
  45. package/docs/host-security.md +16 -8
  46. package/docs/index.md +26 -22
  47. package/docs/mcp-tools.md +32 -12
  48. package/docs/migration.md +164 -2
  49. package/docs/node-filesystem-config.md +1 -0
  50. package/docs/node-jsonl-session-store.md +5 -4
  51. package/docs/postgres-persistence.md +3 -3
  52. package/docs/provider-caching.md +16 -4
  53. package/docs/provider-conformance.md +39 -1
  54. package/docs/provider-packages.md +60 -3
  55. package/docs/providers/ai-sdk.md +36 -0
  56. package/docs/providers/kimi.md +124 -61
  57. package/docs/providers/neuralwatt.md +19 -13
  58. package/docs/providers/openai.md +56 -13
  59. package/docs/providers/opencode-go.md +118 -30
  60. package/docs/providers/openrouter.md +105 -35
  61. package/docs/providers/zai.md +94 -45
  62. package/docs/release-and-install.md +47 -49
  63. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  64. package/docs/runs-and-usage.md +30 -3
  65. package/docs/server.md +5 -2
  66. package/docs/sqlite-persistence.md +2 -2
  67. package/docs/structured-output.md +1 -1
  68. package/docs/thinking-and-reasoning.md +98 -0
  69. package/docs/tool-execution-primitives.md +3 -3
  70. package/docs/tools.md +21 -1
  71. package/docs/use-case-model-selection.md +109 -0
  72. package/docs/workflow-orchestration-primitives.md +1 -0
  73. package/docs/workflows.md +18 -10
  74. package/docs/working-and-semantic-memory.md +1 -0
  75. package/package.json +2 -2
@@ -0,0 +1,109 @@
1
+ # Use-case model selection
2
+
3
+ ## What it does
4
+
5
+ Prism separates the **session chat model** (`AgentConfig.model` / `RunOptions.model`) from **use-case models** used by background or adjacent LLM jobs (observational memory workers, LLM compaction summarizers, declarative agents, supervisor children, evals). Hosts bind `{ model?, provider?, providerOptions?, thinkingLevel? }` per use case. When the use-case omits `model`, resolution falls back to the active session model. Workers never write `model_change` session entries for their own jobs.
6
+
7
+ ## When to use it
8
+
9
+ - Observational memory should run a cheaper/faster model than the chat session (or inherit the session model when unset).
10
+ - LLM compaction should summarize with an explicit `summaryModel`, falling back to a host-supplied session `model`.
11
+ - Declarative agents, supervisor children, evals, and RPC/CLI runs already own their models — document them as use-case sites that stay separate from a parent session.
12
+ - Memory/RAG `Embedder` selection is related but **not** a chat `ModelConfig` binding.
13
+
14
+ ## Contract
15
+
16
+ | Layer | Surface |
17
+ | --- | --- |
18
+ | Binding | `UseCaseModelBinding`: `{ model?, provider?, providerOptions?, thinkingLevel?, requireExplicitModel? }` |
19
+ | Resolver | `resolveUseCaseModel({ configured, sessionModel, requireExplicitModel?, … })` → `{ model, source }` or `undefined` |
20
+ | Binding helper | `resolveUseCaseModelBinding(binding, sessionModel)` |
21
+ | Credential id | `useCaseCredentialProviderId(resolved, binding?)` — always the **resolved** `model.provider` |
22
+ | Escape hatch | `requireExplicitModel: true` skips session fallback (OM historical `missing_model`) |
23
+
24
+ ```ts
25
+ import { resolveUseCaseModel, applyThinkingLevel, thinkingFamilyForModel } from "@arnilo/prism";
26
+
27
+ // Prefer an explicit worker; otherwise inherit the session/agent model.
28
+ const resolved = resolveUseCaseModel({
29
+ configured: settings.workerModel, // optional use-case ModelConfig
30
+ sessionModel: agent.config.model, // host-supplied; AgentSession does not expose agent
31
+ thinkingLevel: settings.thinkingLevel,
32
+ });
33
+ if (!resolved) {
34
+ // skip — neither configured nor session model (or requireExplicitModel)
35
+ }
36
+
37
+ const family = thinkingFamilyForModel(resolved.model);
38
+ const providerOptions = resolved.thinkingLevel
39
+ ? applyThinkingLevel(resolved.providerOptions, resolved.thinkingLevel, family === "noop" ? "reasoning_effort" : family)
40
+ : resolved.providerOptions;
41
+ ```
42
+
43
+ ### Precedence
44
+
45
+ 1. `configured` / `binding.model` → `source: "configured"`
46
+ 2. Else `sessionModel` when `requireExplicitModel` is not set → `source: "session"`
47
+ 3. Else `undefined` (package skips or throws)
48
+
49
+ Resolution is O(1) and network-free. It does not mutate session history.
50
+
51
+ ## Binding sites
52
+
53
+ | Site | How hosts bind | Session fallback |
54
+ | --- | --- | --- |
55
+ | Observational memory | `workerModel` / settings `workerModel` + runtime `sessionModel` | Yes — pass `sessionModel: agent.config.model`; `requireExplicitModel` restores skip |
56
+ | LLM compaction | `summaryModel` with `model` as fallback slot | Yes — `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` |
57
+ | `RunOptions.model` | Per-run override on the **session** | N/A — this *is* the session/run model (writes `model_change`) |
58
+ | Declarative `AgentDefinition` | Definition `model` / registry resolve | Definition-scoped (independent agent) |
59
+ | Supervisor children | Child `createSession` / child `AgentConfig.model` | Independent child session |
60
+ | Evals / workflows / RPC / CLI | Caller `runOptions.model` | Caller-owned |
61
+ | Structured output | Reuses session/run model | Same as run |
62
+ | Memory / RAG | Host `Embedder` | Not a chat model — see [Working and semantic memory](working-and-semantic-memory.md) |
63
+
64
+ ## Observational memory
65
+
66
+ ```ts
67
+ createObservationalMemoryRuntime({
68
+ session,
69
+ appendEntry: (entry) => store.append(entry),
70
+ workerProvider,
71
+ sessionModel: agent.config.model, // enables fallback when workerModel unset
72
+ // workerModel: { provider: "neuralwatt", model: "glm-5.2-fast" }, // optional override
73
+ overrides: { thinkingLevel: "low", observeAfterTokens: 1 },
74
+ });
75
+ ```
76
+
77
+ - Default: no `workerModel` + `sessionModel` set → workers use the session model.
78
+ - Explicit `workerModel` (or settings `workerModel`) always wins.
79
+ - `requireExplicitModel: true` (runtime or settings) → `skipped: "missing_model"` when no worker model, even if `sessionModel` is set.
80
+ - Neither worker nor session model → `skipped: "missing_model"`.
81
+ - Default credential request uses the **resolved** model's `provider` id.
82
+
83
+ ## LLM compaction
84
+
85
+ ```ts
86
+ createLlmCompactionStrategy({
87
+ provider: summaryProvider,
88
+ summaryModel: { provider: "example", model: "cheap-summary" }, // optional
89
+ model: agent.config.model, // session fallback when summaryModel omitted
90
+ thinkingLevel: "low",
91
+ });
92
+ ```
93
+
94
+ `summaryModel` wins; otherwise `model` is required. Thinking maps into `compat` via `applyThinkingLevel` ([Thinking and reasoning](thinking-and-reasoning.md)).
95
+
96
+ ## Security
97
+
98
+ - Credential requests for worker calls must target the **resolved** model’s provider — not ambient session credentials for a different provider unless the host wires that explicitly.
99
+ - Pass known secrets into worker/compaction options so prompts, ledger custom entries, and errors stay redacted.
100
+ - Background workers must not append `model_change` entries or otherwise rewrite the chat session’s model timeline.
101
+
102
+ ## Related pages
103
+
104
+ - [Thinking and reasoning](thinking-and-reasoning.md) — per-turn `thinkingLevel` → `compat`
105
+ - [Observational memory compaction package](compaction-observational-memory.md)
106
+ - [LLM compaction package](compaction-llm.md)
107
+ - [Agent/session runtime](agent-session-runtime.md) — `RunOptions.model` / `model_change`
108
+ - [Working and semantic memory](working-and-semantic-memory.md) — `Embedder` (non-chat)
109
+ - Evidence matrix: [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md)
@@ -387,6 +387,7 @@ const review = functionNode({ execute: async (ctx) => lint(ctx.upstream.draft) }
387
387
 
388
388
  const workflow = defineWorkflow({
389
389
  id: "research-draft-review",
390
+ revision: "2026-07-19.1",
390
391
  nodes: { research, draft, review },
391
392
  edges: [
392
393
  ["research", "draft"],
package/docs/workflows.md CHANGED
@@ -29,29 +29,32 @@ Use `createWorkflowCoordinator()` when multiple processes share SQLite/PostgreSQ
29
29
 
30
30
  ## Inputs / request
31
31
 
32
- `defineWorkflow({ id, nodes, edges, limits? })`:
32
+ `defineWorkflow({ id, revision, nodes, edges, limits? })`:
33
33
 
34
34
  | Field | Notes |
35
35
  | --- | --- |
36
36
  | `id` | Stable workflow id (required) |
37
+ | `revision` | Non-empty host-authored definition revision (required); parent and nested revisions enter `definitionHash` |
37
38
  | `nodes` | Record of node definitions (`kind` + typed fields) |
38
39
  | `edges` | `[from, to]` pairs; must be acyclic; unknown ids rejected |
39
- | `limits.maxNodes` | Default 1000 |
40
- | `limits.maxFanOut` | Default 64 |
41
- | `limits.maxConcurrency` | Default 8 |
42
- | `limits.maxNodeOutputBytes` | Default 4 MiB |
43
- | `limits.maxCheckpointBytes` | Default 1 MiB |
40
+ | `limits.maxNodes` | Default 1,000 / hard cap 10,000 |
41
+ | `limits.maxFanOut` | Default 64 / hard cap 1,024 |
42
+ | `limits.maxConcurrency` | Default 8 / hard cap 256 |
43
+ | `limits.maxNodeOutputBytes` | Default 4 MiB / hard cap 16 MiB |
44
+ | `limits.maxCheckpointBytes` | Default 1 MiB / hard cap 8 MiB |
44
45
  | `limits.maxNestedDepth` / hard cap | 8 / 32; inherited by child workflows |
45
46
  | `limits.maxStateBytes` / hard cap | 64 KiB / 512 KiB |
46
47
  | `limits.maxStateHistory` / hard cap | 32 / 128 state snapshots; updates stop before evidence would be discarded |
47
48
  | `limits.maxReplayDepth` / hard cap | 8 / 32 lineage generations |
48
49
  | `state.initial` / `state.schema` | Initial shared JSON object and optional host-validated schema |
49
50
 
51
+ All workflow limits and runtime `concurrency` reject non-safe integers, zero, negatives, NaN, `Infinity`, and values above the named hard cap. Node retries allow 0–100; an explicit node timeout allows 1–86,400,000 ms. Omitting `timeoutMs` remains an explicit host choice.
52
+
50
53
  `runWorkflow(workflow, input, options?)`:
51
54
 
52
55
  | Option | Notes |
53
56
  | --- | --- |
54
- | `concurrency` | Worker pool size (capped by workflow/global limits) |
57
+ | `concurrency` | Worker pool size; positive safe integer, hard cap 256, and capped by the workflow limit |
55
58
  | `checkpoints` | `WorkflowCheckpointAdapter` for save/load/list |
56
59
  | `agentFactory` | `(agentName) => AgentSession` for agent nodes |
57
60
  | `tools` | Tool registry/lookup for tool nodes |
@@ -100,6 +103,7 @@ Package-local `WorkflowEvent` types: `workflow_started`, `workflow_suspended`, `
100
103
  ```json
101
104
  {
102
105
  "id": "research-draft",
106
+ "revision": "2026-07-19.1",
103
107
  "nodes": ["research", "draft"],
104
108
  "edges": [["research", "draft"]],
105
109
  "limits": { "maxNodes": 256, "maxFanOut": 32, "maxConcurrency": 4 }
@@ -159,6 +163,7 @@ const publish = functionNode({
159
163
 
160
164
  const workflow = defineWorkflow({
161
165
  id: "research-draft",
166
+ revision: "2026-07-19.1",
162
167
  nodes: { research, draft, publish },
163
168
  edges: [["research", "draft"], ["draft", "publish"]],
164
169
  limits: { maxNodes: 256, maxFanOut: 32, maxConcurrency: 4 },
@@ -209,6 +214,7 @@ if (result.status === "suspended") {
209
214
  await cancelWorkflowRun({
210
215
  workflowId: workflow.id,
211
216
  runId: result.runId,
217
+ workflow,
212
218
  checkpoints,
213
219
  ownership: { tenantId: "t1" },
214
220
  });
@@ -258,15 +264,16 @@ runRpcServer({
258
264
 
259
265
  ## Security and performance notes
260
266
 
261
- - Definitions fail closed on cycles, unknown edges, self-edges, and `maxNodes` overflow.
262
- - Fan-out length is bounded by `maxFanOut`; concurrency by `maxConcurrency`.
267
+ - Definitions require a non-empty host-authored `revision` and fail closed on cycles, unknown edges, self-edges, invalid limits, and `maxNodes` overflow. Revision and every nested revision enter the deterministic definition hash; hosts must bump revision when function/tool behavior changes.
268
+ - Fan-out length is bounded by `maxFanOut`; concurrency by `maxConcurrency`; every count/byte/runtime option has a finite hard cap.
263
269
  - Node outputs, shared state/history, schedule input/records, and checkpoints are byte/count/depth bounded. Checkpoint size remains the final aggregate ceiling.
264
270
  - Event buses use a bounded buffer (default 2048) with `close` / `drop_oldest` / `drop_newest` overflow.
265
271
  - Checkpoints redact suspension/resume payloads via `SecretRedactor` / `secrets` before save; resume rejects tenant, schema, definition-hash, and expected-version mismatch.
266
272
  - Suspension requires a checkpoint adapter, consumes no worker/polling slot, and is ignored by distributed coordinators until explicit resume.
267
273
  - Concurrent resumes race on checkpoint CAS before node execution; one wins and stale/duplicate reviewers fail closed. Approved tool nodes then re-run current `ExecutionPolicy`, so durable approval cannot grant stale permissions.
268
274
  - `toolNode({ approval: { reason, data?, resumeSchema? } })` suspends before tool execution. Denial is terminal `denied`; no tool side effect occurs.
269
- - `cancelWorkflowRun` aborts local runs immediately and writes a durable cancellation request for a remotely leased run; workers check it during lease renewal.
275
+ - `cancelWorkflowRun` requires the current workflow definition and exact tenant/account/user ownership. It verifies recursive definition hash before abort/mutation, then aborts local runs or writes a durable cancellation request for remotely leased work. Tenant-only or missing ownership cannot cancel a more-specific owned run.
276
+ - Active registry identity includes workflow ID, run ID, and exact ownership. Exact duplicates fail instead of overwriting; distinct owners remain isolated in lookup/list/cancel/unregister.
270
277
  - Tool nodes attach `workflowId` / `nodeId` on `ExecutionAction.metadata` for approval/audit context.
271
278
  - Nested workflows inherit host registries/policies and cannot inject broader tools, agents, ownership, or credentials. Nested depth is inherited; child suspension bubbles to the parent review cursor.
272
279
  - Replay source ownership/hash/status/node eligibility are checked before a new checkpoint is created. Source records are immutable, lineage is bounded, and copied approval-bearing paths are rejected.
@@ -281,6 +288,7 @@ Use workflows for known, durable, replayable graphs. Use optional supervisor del
281
288
  - Examples: `examples/workflow-research-and-review.ts`, `examples/workflow-parallel-research.ts`, `examples/workflow-tool-approval.ts`, `examples/workflow-multimodal-document.ts`, `examples/workflow-sqlite-resume.ts`, `examples/workflow-postgres-resume.ts`, `examples/workflow-event-sink.ts`, `examples/workflow-rpc-cancel.ts`, `examples/workflow-distributed-coordinator.ts` — offline runnable demos; PostgreSQL safely skips unless `PRISM_TEST_POSTGRES_URL` is set.
282
289
  - [Workflow orchestration primitives](workflow-orchestration-primitives.md): Task 0–1 inventory and locked adapter contracts
283
290
  - [Agent/session runtime](agent-session-runtime.md): `AgentSession.run()`/`stream()`, abort, subscribe
291
+ - [Guardrails](guardrails.md): `RunWorkflowOptions.guardrails` routes tool nodes through core dispatch before policy and side effects.
284
292
  - [Supervisor delegation](supervisors.md): bounded dynamic child selection.
285
293
  - [Agent events](agent-events.md): core `AgentEvent` wrapped by `agent_event`
286
294
  - [Session stores and branching](session-stores-and-branching.md): session `leafId` reuse on resume
@@ -152,6 +152,7 @@ await runMemoryConformance(() => ({
152
152
  - Configure `secrets` / `redactor` so memory text and metadata cannot persist or inject raw canaries.
153
153
  - Injected context is inert text — it cannot grant tools or permissions.
154
154
  - Hard caps: top-K ≤ 32, messageRange ≤ 4, embed batch ≤ 128, injected tokens ≤ 8000, payload/working-memory byte limits enforced.
155
+ - Every embedding is a non-empty finite number vector. `embedBatched()`, in-memory `VectorStore` upserts/queries, and PostgreSQL/pgvector parameters reject NaN, ±Infinity, non-numbers, and wrong configured dimensions before similarity scoring or SQL. Custom adapters can call `assertFiniteVector(vector, label, expectedLength?)` at their trust boundary.
155
156
  - Default `remember()` does not block agent completion; pass `{ wait: true }` when indexing must finish first.
156
157
  - PostgreSQL live suite is gated by `PRISM_TEST_POSTGRES_URL` and requires the `vector` extension.
157
158
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@arnilo/prism",
3
- "version": "0.0.5",
3
+ "version": "0.0.7",
4
4
  "description": "Agent harness for AI providers, agents, sessions, and tools.",
5
5
  "type": "module",
6
6
  "main": "./dist/index.js",
@@ -92,7 +92,7 @@
92
92
  }
93
93
  },
94
94
  "bin": {
95
- "prism": "./dist/cli.js"
95
+ "prism": "dist/cli.js"
96
96
  },
97
97
  "files": [
98
98
  "dist",