@arnilo/prism 0.0.5 → 0.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -1
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +27 -16
- package/dist/agent-run-lifecycle.d.ts +28 -0
- package/dist/agent-run-lifecycle.js +33 -0
- package/dist/agent-run-state.d.ts +53 -0
- package/dist/agent-run-state.js +127 -0
- package/dist/agents.d.ts +3 -1
- package/dist/agents.js +337 -46
- package/dist/contracts.d.ts +205 -3
- package/dist/contracts.js +4 -0
- package/dist/guardrails.d.ts +25 -0
- package/dist/guardrails.js +133 -0
- package/dist/ids.d.ts +2 -0
- package/dist/ids.js +6 -0
- package/dist/index.d.ts +17 -3
- package/dist/index.js +10 -3
- package/dist/input.js +2 -0
- package/dist/resources.js +2 -1
- package/dist/run-limits.d.ts +34 -0
- package/dist/run-limits.js +163 -0
- package/dist/secure-agent.d.ts +3 -0
- package/dist/secure-agent.js +63 -0
- package/dist/session-stores.js +2 -3
- package/dist/testing/persistence-schema.d.ts +45 -7
- package/dist/testing/persistence-schema.js +138 -24
- package/dist/thinking.d.ts +42 -0
- package/dist/thinking.js +92 -0
- package/dist/tools.d.ts +10 -2
- package/dist/tools.js +56 -7
- package/dist/use-case-model.d.ts +63 -0
- package/dist/use-case-model.js +52 -0
- package/docs/a2a.md +4 -2
- package/docs/agent-events.md +23 -16
- package/docs/agent-loops.md +19 -8
- package/docs/agent-session-runtime.md +33 -1
- package/docs/coding-agent-tools.md +33 -12
- package/docs/coding-security.md +2 -2
- package/docs/compaction-llm.md +17 -7
- package/docs/compaction-observational-memory.md +28 -4
- package/docs/credential-storage.md +58 -9
- package/docs/credentials-and-redaction.md +1 -1
- package/docs/database-persistence.md +8 -3
- package/docs/guardrails.md +75 -0
- package/docs/host-security.md +16 -8
- package/docs/index.md +26 -22
- package/docs/mcp-tools.md +32 -12
- package/docs/migration.md +164 -2
- package/docs/node-filesystem-config.md +1 -0
- package/docs/node-jsonl-session-store.md +5 -4
- package/docs/postgres-persistence.md +3 -3
- package/docs/provider-caching.md +16 -4
- package/docs/provider-conformance.md +39 -1
- package/docs/provider-packages.md +60 -3
- package/docs/providers/ai-sdk.md +36 -0
- package/docs/providers/kimi.md +124 -61
- package/docs/providers/neuralwatt.md +19 -13
- package/docs/providers/openai.md +56 -13
- package/docs/providers/opencode-go.md +118 -30
- package/docs/providers/openrouter.md +105 -35
- package/docs/providers/zai.md +94 -45
- package/docs/release-and-install.md +47 -49
- package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
- package/docs/runs-and-usage.md +30 -3
- package/docs/server.md +5 -2
- package/docs/sqlite-persistence.md +2 -2
- package/docs/structured-output.md +1 -1
- package/docs/thinking-and-reasoning.md +98 -0
- package/docs/tool-execution-primitives.md +3 -3
- package/docs/tools.md +21 -1
- package/docs/use-case-model-selection.md +109 -0
- package/docs/workflow-orchestration-primitives.md +1 -0
- package/docs/workflows.md +18 -10
- package/docs/working-and-semantic-memory.md +1 -0
- package/package.json +2 -2
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# Use-case model selection
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Prism separates the **session chat model** (`AgentConfig.model` / `RunOptions.model`) from **use-case models** used by background or adjacent LLM jobs (observational memory workers, LLM compaction summarizers, declarative agents, supervisor children, evals). Hosts bind `{ model?, provider?, providerOptions?, thinkingLevel? }` per use case. When the use-case omits `model`, resolution falls back to the active session model. Workers never write `model_change` session entries for their own jobs.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
- Observational memory should run a cheaper/faster model than the chat session (or inherit the session model when unset).
|
|
10
|
+
- LLM compaction should summarize with an explicit `summaryModel`, falling back to a host-supplied session `model`.
|
|
11
|
+
- Declarative agents, supervisor children, evals, and RPC/CLI runs already own their models — document them as use-case sites that stay separate from a parent session.
|
|
12
|
+
- Memory/RAG `Embedder` selection is related but **not** a chat `ModelConfig` binding.
|
|
13
|
+
|
|
14
|
+
## Contract
|
|
15
|
+
|
|
16
|
+
| Layer | Surface |
|
|
17
|
+
| --- | --- |
|
|
18
|
+
| Binding | `UseCaseModelBinding`: `{ model?, provider?, providerOptions?, thinkingLevel?, requireExplicitModel? }` |
|
|
19
|
+
| Resolver | `resolveUseCaseModel({ configured, sessionModel, requireExplicitModel?, … })` → `{ model, source }` or `undefined` |
|
|
20
|
+
| Binding helper | `resolveUseCaseModelBinding(binding, sessionModel)` |
|
|
21
|
+
| Credential id | `useCaseCredentialProviderId(resolved, binding?)` — always the **resolved** `model.provider` |
|
|
22
|
+
| Escape hatch | `requireExplicitModel: true` skips session fallback (OM historical `missing_model`) |
|
|
23
|
+
|
|
24
|
+
```ts
|
|
25
|
+
import { resolveUseCaseModel, applyThinkingLevel, thinkingFamilyForModel } from "@arnilo/prism";
|
|
26
|
+
|
|
27
|
+
// Prefer an explicit worker; otherwise inherit the session/agent model.
|
|
28
|
+
const resolved = resolveUseCaseModel({
|
|
29
|
+
configured: settings.workerModel, // optional use-case ModelConfig
|
|
30
|
+
sessionModel: agent.config.model, // host-supplied; AgentSession does not expose agent
|
|
31
|
+
thinkingLevel: settings.thinkingLevel,
|
|
32
|
+
});
|
|
33
|
+
if (!resolved) {
|
|
34
|
+
// skip — neither configured nor session model (or requireExplicitModel)
|
|
35
|
+
}
|
|
36
|
+
|
|
37
|
+
const family = thinkingFamilyForModel(resolved.model);
|
|
38
|
+
const providerOptions = resolved.thinkingLevel
|
|
39
|
+
? applyThinkingLevel(resolved.providerOptions, resolved.thinkingLevel, family === "noop" ? "reasoning_effort" : family)
|
|
40
|
+
: resolved.providerOptions;
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
### Precedence
|
|
44
|
+
|
|
45
|
+
1. `configured` / `binding.model` → `source: "configured"`
|
|
46
|
+
2. Else `sessionModel` when `requireExplicitModel` is not set → `source: "session"`
|
|
47
|
+
3. Else `undefined` (package skips or throws)
|
|
48
|
+
|
|
49
|
+
Resolution is O(1) and network-free. It does not mutate session history.
|
|
50
|
+
|
|
51
|
+
## Binding sites
|
|
52
|
+
|
|
53
|
+
| Site | How hosts bind | Session fallback |
|
|
54
|
+
| --- | --- | --- |
|
|
55
|
+
| Observational memory | `workerModel` / settings `workerModel` + runtime `sessionModel` | Yes — pass `sessionModel: agent.config.model`; `requireExplicitModel` restores skip |
|
|
56
|
+
| LLM compaction | `summaryModel` with `model` as fallback slot | Yes — `resolveUseCaseModel({ configured: summaryModel, sessionModel: model })` |
|
|
57
|
+
| `RunOptions.model` | Per-run override on the **session** | N/A — this *is* the session/run model (writes `model_change`) |
|
|
58
|
+
| Declarative `AgentDefinition` | Definition `model` / registry resolve | Definition-scoped (independent agent) |
|
|
59
|
+
| Supervisor children | Child `createSession` / child `AgentConfig.model` | Independent child session |
|
|
60
|
+
| Evals / workflows / RPC / CLI | Caller `runOptions.model` | Caller-owned |
|
|
61
|
+
| Structured output | Reuses session/run model | Same as run |
|
|
62
|
+
| Memory / RAG | Host `Embedder` | Not a chat model — see [Working and semantic memory](working-and-semantic-memory.md) |
|
|
63
|
+
|
|
64
|
+
## Observational memory
|
|
65
|
+
|
|
66
|
+
```ts
|
|
67
|
+
createObservationalMemoryRuntime({
|
|
68
|
+
session,
|
|
69
|
+
appendEntry: (entry) => store.append(entry),
|
|
70
|
+
workerProvider,
|
|
71
|
+
sessionModel: agent.config.model, // enables fallback when workerModel unset
|
|
72
|
+
// workerModel: { provider: "neuralwatt", model: "glm-5.2-fast" }, // optional override
|
|
73
|
+
overrides: { thinkingLevel: "low", observeAfterTokens: 1 },
|
|
74
|
+
});
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
- Default: no `workerModel` + `sessionModel` set → workers use the session model.
|
|
78
|
+
- Explicit `workerModel` (or settings `workerModel`) always wins.
|
|
79
|
+
- `requireExplicitModel: true` (runtime or settings) → `skipped: "missing_model"` when no worker model, even if `sessionModel` is set.
|
|
80
|
+
- Neither worker nor session model → `skipped: "missing_model"`.
|
|
81
|
+
- Default credential request uses the **resolved** model's `provider` id.
|
|
82
|
+
|
|
83
|
+
## LLM compaction
|
|
84
|
+
|
|
85
|
+
```ts
|
|
86
|
+
createLlmCompactionStrategy({
|
|
87
|
+
provider: summaryProvider,
|
|
88
|
+
summaryModel: { provider: "example", model: "cheap-summary" }, // optional
|
|
89
|
+
model: agent.config.model, // session fallback when summaryModel omitted
|
|
90
|
+
thinkingLevel: "low",
|
|
91
|
+
});
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
`summaryModel` wins; otherwise `model` is required. Thinking maps into `compat` via `applyThinkingLevel` ([Thinking and reasoning](thinking-and-reasoning.md)).
|
|
95
|
+
|
|
96
|
+
## Security
|
|
97
|
+
|
|
98
|
+
- Credential requests for worker calls must target the **resolved** model’s provider — not ambient session credentials for a different provider unless the host wires that explicitly.
|
|
99
|
+
- Pass known secrets into worker/compaction options so prompts, ledger custom entries, and errors stay redacted.
|
|
100
|
+
- Background workers must not append `model_change` entries or otherwise rewrite the chat session’s model timeline.
|
|
101
|
+
|
|
102
|
+
## Related pages
|
|
103
|
+
|
|
104
|
+
- [Thinking and reasoning](thinking-and-reasoning.md) — per-turn `thinkingLevel` → `compat`
|
|
105
|
+
- [Observational memory compaction package](compaction-observational-memory.md)
|
|
106
|
+
- [LLM compaction package](compaction-llm.md)
|
|
107
|
+
- [Agent/session runtime](agent-session-runtime.md) — `RunOptions.model` / `model_change`
|
|
108
|
+
- [Working and semantic memory](working-and-semantic-memory.md) — `Embedder` (non-chat)
|
|
109
|
+
- Evidence matrix: [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md)
|
|
@@ -387,6 +387,7 @@ const review = functionNode({ execute: async (ctx) => lint(ctx.upstream.draft) }
|
|
|
387
387
|
|
|
388
388
|
const workflow = defineWorkflow({
|
|
389
389
|
id: "research-draft-review",
|
|
390
|
+
revision: "2026-07-19.1",
|
|
390
391
|
nodes: { research, draft, review },
|
|
391
392
|
edges: [
|
|
392
393
|
["research", "draft"],
|
package/docs/workflows.md
CHANGED
|
@@ -29,29 +29,32 @@ Use `createWorkflowCoordinator()` when multiple processes share SQLite/PostgreSQ
|
|
|
29
29
|
|
|
30
30
|
## Inputs / request
|
|
31
31
|
|
|
32
|
-
`defineWorkflow({ id, nodes, edges, limits? })`:
|
|
32
|
+
`defineWorkflow({ id, revision, nodes, edges, limits? })`:
|
|
33
33
|
|
|
34
34
|
| Field | Notes |
|
|
35
35
|
| --- | --- |
|
|
36
36
|
| `id` | Stable workflow id (required) |
|
|
37
|
+
| `revision` | Non-empty host-authored definition revision (required); parent and nested revisions enter `definitionHash` |
|
|
37
38
|
| `nodes` | Record of node definitions (`kind` + typed fields) |
|
|
38
39
|
| `edges` | `[from, to]` pairs; must be acyclic; unknown ids rejected |
|
|
39
|
-
| `limits.maxNodes` | Default
|
|
40
|
-
| `limits.maxFanOut` | Default 64 |
|
|
41
|
-
| `limits.maxConcurrency` | Default 8 |
|
|
42
|
-
| `limits.maxNodeOutputBytes` | Default 4 MiB |
|
|
43
|
-
| `limits.maxCheckpointBytes` | Default 1 MiB |
|
|
40
|
+
| `limits.maxNodes` | Default 1,000 / hard cap 10,000 |
|
|
41
|
+
| `limits.maxFanOut` | Default 64 / hard cap 1,024 |
|
|
42
|
+
| `limits.maxConcurrency` | Default 8 / hard cap 256 |
|
|
43
|
+
| `limits.maxNodeOutputBytes` | Default 4 MiB / hard cap 16 MiB |
|
|
44
|
+
| `limits.maxCheckpointBytes` | Default 1 MiB / hard cap 8 MiB |
|
|
44
45
|
| `limits.maxNestedDepth` / hard cap | 8 / 32; inherited by child workflows |
|
|
45
46
|
| `limits.maxStateBytes` / hard cap | 64 KiB / 512 KiB |
|
|
46
47
|
| `limits.maxStateHistory` / hard cap | 32 / 128 state snapshots; updates stop before evidence would be discarded |
|
|
47
48
|
| `limits.maxReplayDepth` / hard cap | 8 / 32 lineage generations |
|
|
48
49
|
| `state.initial` / `state.schema` | Initial shared JSON object and optional host-validated schema |
|
|
49
50
|
|
|
51
|
+
All workflow limits and runtime `concurrency` reject non-safe integers, zero, negatives, NaN, `Infinity`, and values above the named hard cap. Node retries allow 0–100; an explicit node timeout allows 1–86,400,000 ms. Omitting `timeoutMs` remains an explicit host choice.
|
|
52
|
+
|
|
50
53
|
`runWorkflow(workflow, input, options?)`:
|
|
51
54
|
|
|
52
55
|
| Option | Notes |
|
|
53
56
|
| --- | --- |
|
|
54
|
-
| `concurrency` | Worker pool size
|
|
57
|
+
| `concurrency` | Worker pool size; positive safe integer, hard cap 256, and capped by the workflow limit |
|
|
55
58
|
| `checkpoints` | `WorkflowCheckpointAdapter` for save/load/list |
|
|
56
59
|
| `agentFactory` | `(agentName) => AgentSession` for agent nodes |
|
|
57
60
|
| `tools` | Tool registry/lookup for tool nodes |
|
|
@@ -100,6 +103,7 @@ Package-local `WorkflowEvent` types: `workflow_started`, `workflow_suspended`, `
|
|
|
100
103
|
```json
|
|
101
104
|
{
|
|
102
105
|
"id": "research-draft",
|
|
106
|
+
"revision": "2026-07-19.1",
|
|
103
107
|
"nodes": ["research", "draft"],
|
|
104
108
|
"edges": [["research", "draft"]],
|
|
105
109
|
"limits": { "maxNodes": 256, "maxFanOut": 32, "maxConcurrency": 4 }
|
|
@@ -159,6 +163,7 @@ const publish = functionNode({
|
|
|
159
163
|
|
|
160
164
|
const workflow = defineWorkflow({
|
|
161
165
|
id: "research-draft",
|
|
166
|
+
revision: "2026-07-19.1",
|
|
162
167
|
nodes: { research, draft, publish },
|
|
163
168
|
edges: [["research", "draft"], ["draft", "publish"]],
|
|
164
169
|
limits: { maxNodes: 256, maxFanOut: 32, maxConcurrency: 4 },
|
|
@@ -209,6 +214,7 @@ if (result.status === "suspended") {
|
|
|
209
214
|
await cancelWorkflowRun({
|
|
210
215
|
workflowId: workflow.id,
|
|
211
216
|
runId: result.runId,
|
|
217
|
+
workflow,
|
|
212
218
|
checkpoints,
|
|
213
219
|
ownership: { tenantId: "t1" },
|
|
214
220
|
});
|
|
@@ -258,15 +264,16 @@ runRpcServer({
|
|
|
258
264
|
|
|
259
265
|
## Security and performance notes
|
|
260
266
|
|
|
261
|
-
- Definitions fail closed on cycles, unknown edges, self-edges, and `maxNodes` overflow.
|
|
262
|
-
- Fan-out length is bounded by `maxFanOut`; concurrency by `maxConcurrency
|
|
267
|
+
- Definitions require a non-empty host-authored `revision` and fail closed on cycles, unknown edges, self-edges, invalid limits, and `maxNodes` overflow. Revision and every nested revision enter the deterministic definition hash; hosts must bump revision when function/tool behavior changes.
|
|
268
|
+
- Fan-out length is bounded by `maxFanOut`; concurrency by `maxConcurrency`; every count/byte/runtime option has a finite hard cap.
|
|
263
269
|
- Node outputs, shared state/history, schedule input/records, and checkpoints are byte/count/depth bounded. Checkpoint size remains the final aggregate ceiling.
|
|
264
270
|
- Event buses use a bounded buffer (default 2048) with `close` / `drop_oldest` / `drop_newest` overflow.
|
|
265
271
|
- Checkpoints redact suspension/resume payloads via `SecretRedactor` / `secrets` before save; resume rejects tenant, schema, definition-hash, and expected-version mismatch.
|
|
266
272
|
- Suspension requires a checkpoint adapter, consumes no worker/polling slot, and is ignored by distributed coordinators until explicit resume.
|
|
267
273
|
- Concurrent resumes race on checkpoint CAS before node execution; one wins and stale/duplicate reviewers fail closed. Approved tool nodes then re-run current `ExecutionPolicy`, so durable approval cannot grant stale permissions.
|
|
268
274
|
- `toolNode({ approval: { reason, data?, resumeSchema? } })` suspends before tool execution. Denial is terminal `denied`; no tool side effect occurs.
|
|
269
|
-
- `cancelWorkflowRun` aborts local runs
|
|
275
|
+
- `cancelWorkflowRun` requires the current workflow definition and exact tenant/account/user ownership. It verifies recursive definition hash before abort/mutation, then aborts local runs or writes a durable cancellation request for remotely leased work. Tenant-only or missing ownership cannot cancel a more-specific owned run.
|
|
276
|
+
- Active registry identity includes workflow ID, run ID, and exact ownership. Exact duplicates fail instead of overwriting; distinct owners remain isolated in lookup/list/cancel/unregister.
|
|
270
277
|
- Tool nodes attach `workflowId` / `nodeId` on `ExecutionAction.metadata` for approval/audit context.
|
|
271
278
|
- Nested workflows inherit host registries/policies and cannot inject broader tools, agents, ownership, or credentials. Nested depth is inherited; child suspension bubbles to the parent review cursor.
|
|
272
279
|
- Replay source ownership/hash/status/node eligibility are checked before a new checkpoint is created. Source records are immutable, lineage is bounded, and copied approval-bearing paths are rejected.
|
|
@@ -281,6 +288,7 @@ Use workflows for known, durable, replayable graphs. Use optional supervisor del
|
|
|
281
288
|
- Examples: `examples/workflow-research-and-review.ts`, `examples/workflow-parallel-research.ts`, `examples/workflow-tool-approval.ts`, `examples/workflow-multimodal-document.ts`, `examples/workflow-sqlite-resume.ts`, `examples/workflow-postgres-resume.ts`, `examples/workflow-event-sink.ts`, `examples/workflow-rpc-cancel.ts`, `examples/workflow-distributed-coordinator.ts` — offline runnable demos; PostgreSQL safely skips unless `PRISM_TEST_POSTGRES_URL` is set.
|
|
282
289
|
- [Workflow orchestration primitives](workflow-orchestration-primitives.md): Task 0–1 inventory and locked adapter contracts
|
|
283
290
|
- [Agent/session runtime](agent-session-runtime.md): `AgentSession.run()`/`stream()`, abort, subscribe
|
|
291
|
+
- [Guardrails](guardrails.md): `RunWorkflowOptions.guardrails` routes tool nodes through core dispatch before policy and side effects.
|
|
284
292
|
- [Supervisor delegation](supervisors.md): bounded dynamic child selection.
|
|
285
293
|
- [Agent events](agent-events.md): core `AgentEvent` wrapped by `agent_event`
|
|
286
294
|
- [Session stores and branching](session-stores-and-branching.md): session `leafId` reuse on resume
|
|
@@ -152,6 +152,7 @@ await runMemoryConformance(() => ({
|
|
|
152
152
|
- Configure `secrets` / `redactor` so memory text and metadata cannot persist or inject raw canaries.
|
|
153
153
|
- Injected context is inert text — it cannot grant tools or permissions.
|
|
154
154
|
- Hard caps: top-K ≤ 32, messageRange ≤ 4, embed batch ≤ 128, injected tokens ≤ 8000, payload/working-memory byte limits enforced.
|
|
155
|
+
- Every embedding is a non-empty finite number vector. `embedBatched()`, in-memory `VectorStore` upserts/queries, and PostgreSQL/pgvector parameters reject NaN, ±Infinity, non-numbers, and wrong configured dimensions before similarity scoring or SQL. Custom adapters can call `assertFiniteVector(vector, label, expectedLength?)` at their trust boundary.
|
|
155
156
|
- Default `remember()` does not block agent completion; pass `{ wait: true }` when indexing must finish first.
|
|
156
157
|
- PostgreSQL live suite is gated by `PRISM_TEST_POSTGRES_URL` and requires the `vector` extension.
|
|
157
158
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@arnilo/prism",
|
|
3
|
-
"version": "0.0.
|
|
3
|
+
"version": "0.0.7",
|
|
4
4
|
"description": "Agent harness for AI providers, agents, sessions, and tools.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "./dist/index.js",
|
|
@@ -92,7 +92,7 @@
|
|
|
92
92
|
}
|
|
93
93
|
},
|
|
94
94
|
"bin": {
|
|
95
|
-
"prism": "
|
|
95
|
+
"prism": "dist/cli.js"
|
|
96
96
|
},
|
|
97
97
|
"files": [
|
|
98
98
|
"dist",
|