@arnilo/prism 0.0.3 → 0.0.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +40 -0
- package/README.md +62 -26
- package/dist/agent-loops.d.ts +8 -1
- package/dist/agent-loops.js +57 -11
- package/dist/agents.js +212 -32
- package/dist/checkpoints.d.ts +11 -0
- package/dist/checkpoints.js +144 -0
- package/dist/cli-init.d.ts +41 -0
- package/dist/cli-init.js +390 -0
- package/dist/cli-runner.d.ts +7 -1
- package/dist/cli-runner.js +13 -1
- package/dist/compaction.js +9 -1
- package/dist/content.d.ts +121 -0
- package/dist/content.js +538 -0
- package/dist/contracts.d.ts +236 -11
- package/dist/contracts.js +8 -0
- package/dist/event-multiplexer.d.ts +23 -0
- package/dist/event-multiplexer.js +136 -0
- package/dist/execution-policy.d.ts +28 -0
- package/dist/execution-policy.js +24 -0
- package/dist/feedback.d.ts +48 -0
- package/dist/feedback.js +230 -0
- package/dist/index.d.ts +20 -6
- package/dist/index.js +13 -5
- package/dist/input.js +11 -1
- package/dist/leases.d.ts +8 -0
- package/dist/leases.js +111 -0
- package/dist/node/agent-definitions.js +3 -5
- package/dist/node/config.d.ts +1 -0
- package/dist/node/config.js +5 -3
- package/dist/node/contribution-discovery.js +5 -8
- package/dist/node/session-store-jsonl.js +8 -5
- package/dist/node/settings.js +2 -2
- package/dist/node/trust.js +2 -4
- package/dist/observability.d.ts +3 -0
- package/dist/observability.js +18 -0
- package/dist/providers/media.d.ts +44 -0
- package/dist/providers/media.js +126 -0
- package/dist/providers/openai-compatible.js +18 -119
- package/dist/providers/openai-primitives.d.ts +9 -0
- package/dist/providers/openai-primitives.js +129 -0
- package/dist/providers/transport.d.ts +40 -0
- package/dist/providers/transport.js +221 -0
- package/dist/redaction.js +40 -13
- package/dist/resources.d.ts +5 -0
- package/dist/resources.js +4 -0
- package/dist/structured-output.d.ts +11 -0
- package/dist/structured-output.js +59 -0
- package/dist/testing/feedback.d.ts +6 -0
- package/dist/testing/feedback.js +37 -0
- package/dist/testing/persistence-schema.d.ts +102 -0
- package/dist/testing/persistence-schema.js +487 -0
- package/dist/testing/provider-conformance.js +10 -1
- package/dist/testing/run-ledger-conformance.d.ts +33 -0
- package/dist/testing/run-ledger-conformance.js +178 -0
- package/dist/testing/session-store-conformance.d.ts +16 -0
- package/dist/testing/session-store-conformance.js +73 -0
- package/dist/tools.d.ts +17 -0
- package/dist/tools.js +29 -2
- package/docs/a2a.md +73 -0
- package/docs/agent-events.md +17 -10
- package/docs/agent-loops.md +11 -5
- package/docs/agent-session-runtime.md +15 -16
- package/docs/cli-rpc.md +36 -5
- package/docs/coding-agent-tools.md +43 -9
- package/docs/coding-security.md +88 -0
- package/docs/compaction-observational-memory.md +2 -0
- package/docs/context-and-skills.md +1 -0
- package/docs/credential-storage.md +177 -0
- package/docs/credentials-and-redaction.md +4 -3
- package/docs/database-persistence.md +52 -7
- package/docs/evaluations.md +122 -0
- package/docs/extensions.md +2 -2
- package/docs/host-security.md +33 -2
- package/docs/index.md +46 -18
- package/docs/input-and-prompt-assembly.md +6 -5
- package/docs/mcp-tools.md +184 -0
- package/docs/middleware-hooks.md +2 -0
- package/docs/migration.md +51 -28
- package/docs/model-registry.md +5 -3
- package/docs/multimodal-content.md +156 -0
- package/docs/observability.md +171 -0
- package/docs/performance.md +249 -1
- package/docs/persistence-credentials-multimodality-primitives.md +303 -0
- package/docs/postgres-persistence.md +143 -0
- package/docs/provider-conformance.md +18 -0
- package/docs/provider-layer.md +1 -1
- package/docs/provider-packages.md +2 -0
- package/docs/provider-primitives.md +281 -0
- package/docs/providers/ai-sdk.md +113 -0
- package/docs/providers/kimi.md +1 -0
- package/docs/providers/neuralwatt.md +1 -0
- package/docs/providers/openai-compatible.md +2 -1
- package/docs/providers/openai.md +8 -1
- package/docs/providers/opencode-go.md +1 -0
- package/docs/providers/openrouter.md +1 -0
- package/docs/providers/zai.md +1 -0
- package/docs/public-contracts.md +13 -5
- package/docs/rag.md +113 -0
- package/docs/release-and-install.md +237 -30
- package/docs/resource-loading.md +14 -4
- package/docs/review-coverage-2026-07-14.md +260 -0
- package/docs/review-coverage-2026-07-15.md +193 -0
- package/docs/run-ledger-conformance.md +96 -0
- package/docs/runs-and-usage.md +43 -4
- package/docs/server.md +139 -0
- package/docs/session-store-conformance.md +16 -0
- package/docs/session-stores-and-branching.md +1 -0
- package/docs/settings-auth-trust-security.md +6 -5
- package/docs/sqlite-persistence.md +123 -0
- package/docs/structured-output.md +9 -0
- package/docs/supervisors.md +71 -0
- package/docs/tool-conformance.md +1 -0
- package/docs/tool-execution-primitives.md +374 -0
- package/docs/tools.md +39 -1
- package/docs/workflow-orchestration-primitives.md +581 -0
- package/docs/workflow-tui-primitives.md +5 -0
- package/docs/workflows.md +293 -0
- package/docs/working-and-semantic-memory.md +169 -0
- package/package.json +43 -5
- package/templates/init/README.md.tmpl +28 -0
- package/templates/init/env.example.tmpl +1 -0
- package/templates/init/gitignore.tmpl +11 -0
- package/templates/init/optional/evals-example.ts.tmpl +17 -0
- package/templates/init/optional/workflows-example.ts.tmpl +27 -0
- package/templates/init/package.json.tmpl +22 -0
- package/templates/init/providers.json +76 -0
- package/templates/init/src/agent.ts.tmpl +10 -0
- package/templates/init/src/index.ts.tmpl +12 -0
- package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
- package/templates/init/tsconfig.json.tmpl +15 -0
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# Observability
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
Prism exposes provider and tool timing through stable, metadata-only `AgentEvent` variants. Hosts subscribe via `session.subscribe()` or persist events through `RunLedger`. Core helpers build `ProviderTurnMetadata` and classify HTTP failures without echoing prompts, tool arguments, or credentials.
|
|
6
|
+
|
|
7
|
+
Optional package `@arnilo/prism-observability-opentelemetry` maps those events to OpenTelemetry spans and low-cardinality metrics. OpenTelemetry is **not** a dependency of `@arnilo/prism`.
|
|
8
|
+
|
|
9
|
+
APIs:
|
|
10
|
+
|
|
11
|
+
- `ProviderTurnMetadata`, `ToolExecutionMetadata` on `AgentEvent`
|
|
12
|
+
- `createProviderTurnMetadata()`, `readProviderHttpStatus()` in `@arnilo/prism`
|
|
13
|
+
- `createOpenTelemetryInstrumentation()`, `wrapOpenTelemetryApi()`, `createInMemoryTelemetry()` in `@arnilo/prism-observability-opentelemetry`
|
|
14
|
+
- `handleRunFeedback()` / `handleEvaluation()` for explicit safe post-run projection
|
|
15
|
+
|
|
16
|
+
## When to use it
|
|
17
|
+
|
|
18
|
+
Use agent events when you need run-scoped latency, retry attempt numbers, token/cache usage, tool duration, or error classification in-process or through your own exporter.
|
|
19
|
+
|
|
20
|
+
Use the OpenTelemetry adapter when you already run the OpenTelemetry SDK and want spans/metrics without forking the runtime.
|
|
21
|
+
|
|
22
|
+
Do not parse raw provider SSE for timing — provider packages normalize stream events; the session emits `provider_turn_*` once per `generate()` attempt.
|
|
23
|
+
|
|
24
|
+
## Inputs / request
|
|
25
|
+
|
|
26
|
+
Core metadata helpers:
|
|
27
|
+
|
|
28
|
+
```ts
|
|
29
|
+
import { createProviderTurnMetadata, readProviderHttpStatus } from "@arnilo/prism";
|
|
30
|
+
|
|
31
|
+
const metadata = createProviderTurnMetadata(request, providerId, { attempt: 2, latencyMs: 120 });
|
|
32
|
+
const httpStatus = readProviderHttpStatus(errorInfo);
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
New agent event variants (metadata only):
|
|
36
|
+
|
|
37
|
+
| Variant | When | Key fields |
|
|
38
|
+
| --- | --- | --- |
|
|
39
|
+
| `provider_turn_started` | Before each provider `generate()` attempt | `turn`, `metadata: ProviderTurnMetadata` |
|
|
40
|
+
| `provider_turn_finished` | After success or failure of that attempt | `metadata` (includes `latencyMs`, optional `httpStatus`), `usage?`, `error?` |
|
|
41
|
+
|
|
42
|
+
`ToolExecutionMetadata` on terminal tool events:
|
|
43
|
+
|
|
44
|
+
| Field | Meaning |
|
|
45
|
+
| --- | --- |
|
|
46
|
+
| `durationMs` | Wall time from dispatch start to finish/block/error |
|
|
47
|
+
| `status` | `finished` \| `error` \| `blocked` |
|
|
48
|
+
|
|
49
|
+
OpenTelemetry adapter:
|
|
50
|
+
|
|
51
|
+
```ts
|
|
52
|
+
import { trace, metrics } from "@opentelemetry/api";
|
|
53
|
+
import { createOpenTelemetryInstrumentation, wrapOpenTelemetryApi } from "@arnilo/prism-observability-opentelemetry";
|
|
54
|
+
|
|
55
|
+
const { tracer, meter } = wrapOpenTelemetryApi(trace.getTracer("app"), metrics.getMeter("app"));
|
|
56
|
+
const telemetry = createOpenTelemetryInstrumentation({ tracer, meter, onExporterError: console.error });
|
|
57
|
+
|
|
58
|
+
const detach = telemetry.attachSession(session);
|
|
59
|
+
// or: for await (const event of session.subscribe()) telemetry.handleAgentEvent(event);
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Set `enabled: false` or omit `tracer`/`meter` for a no-op adapter. Feedback handlers accept only `runId`, rating/score, booleans, bounded counts, and fixed status — never comment, tag values, scorer/evaluation IDs, or arbitrary metadata.
|
|
63
|
+
|
|
64
|
+
## Outputs / response / events
|
|
65
|
+
|
|
66
|
+
Provider turn metadata fields:
|
|
67
|
+
|
|
68
|
+
| Field | Source |
|
|
69
|
+
| --- | --- |
|
|
70
|
+
| `providerId` | Active provider id |
|
|
71
|
+
| `model` | `ProviderRequest.model` |
|
|
72
|
+
| `requestId` | `request.metadata.requestId` or `request.options.sessionId` |
|
|
73
|
+
| `attempt` | Retry attempt (1-based) |
|
|
74
|
+
| `latencyMs` | Set on `provider_turn_finished` |
|
|
75
|
+
| `httpStatus` | Numeric `ErrorInfo.code` when present |
|
|
76
|
+
| `rateLimitRemaining` / `rateLimitResetMs` | Reserved for provider adapters (optional) |
|
|
77
|
+
|
|
78
|
+
OpenTelemetry mapping (when enabled):
|
|
79
|
+
|
|
80
|
+
| Agent event | Span | Metric labels |
|
|
81
|
+
| --- | --- | --- |
|
|
82
|
+
| `agent_started` / `agent_finished` / run `error` | `prism.agent.run` | `prism.run.tokens` on successful aggregate usage |
|
|
83
|
+
| `provider_turn_*` | `prism.provider.turn` | `provider_id`, `outcome` on duration histogram |
|
|
84
|
+
| `tool_execution_*` (terminal) | `prism.tool.execute` when started | `status` on duration histogram |
|
|
85
|
+
| `provider_turn_finished` usage | span attributes | `prism.provider.tokens` (`kind`: input/output/cache_*`) |
|
|
86
|
+
| `agent_finished` aggregate usage | — | `prism.run.tokens` (`kind`: input/output) |
|
|
87
|
+
| `handleRunFeedback` | active-run `prism.run.feedback` event or ended-run span | `prism.run.feedback` (`rating`, `linked_evaluation`) |
|
|
88
|
+
| `handleEvaluation` | active-run `prism.run.evaluation` event or ended-run span | `prism.run.evaluation` (`status`) |
|
|
89
|
+
|
|
90
|
+
High-cardinality identifiers (`sessionId`, `runId`, `requestId`, `toolCallId`) are **span attributes only**, never metric labels.
|
|
91
|
+
|
|
92
|
+
## Request/response example
|
|
93
|
+
|
|
94
|
+
```json
|
|
95
|
+
{
|
|
96
|
+
"type": "provider_turn_finished",
|
|
97
|
+
"sessionId": "sess_01J...",
|
|
98
|
+
"runId": "run_01J...",
|
|
99
|
+
"turn": 1,
|
|
100
|
+
"metadata": {
|
|
101
|
+
"providerId": "openai",
|
|
102
|
+
"model": { "provider": "openai", "model": "gpt-4.1" },
|
|
103
|
+
"requestId": "sess_01J...",
|
|
104
|
+
"attempt": 2,
|
|
105
|
+
"latencyMs": 842,
|
|
106
|
+
"httpStatus": 503
|
|
107
|
+
},
|
|
108
|
+
"error": { "message": "upstream unavailable", "code": 503 }
|
|
109
|
+
}
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
```json
|
|
113
|
+
{
|
|
114
|
+
"type": "tool_execution_finished",
|
|
115
|
+
"sessionId": "sess_01J...",
|
|
116
|
+
"runId": "run_01J...",
|
|
117
|
+
"result": { "toolCallId": "call_1", "name": "echo" },
|
|
118
|
+
"metadata": { "durationMs": 12, "status": "finished" }
|
|
119
|
+
}
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
## Implementation example
|
|
123
|
+
|
|
124
|
+
```ts
|
|
125
|
+
import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
|
|
126
|
+
import { createInMemoryTelemetry, createOpenTelemetryInstrumentation } from "@arnilo/prism-observability-opentelemetry";
|
|
127
|
+
|
|
128
|
+
const memory = createInMemoryTelemetry();
|
|
129
|
+
const telemetry = createOpenTelemetryInstrumentation({ tracer: memory.tracer, meter: memory.meter });
|
|
130
|
+
|
|
131
|
+
const session = createAgent({
|
|
132
|
+
model: { provider: "mock", model: "demo" },
|
|
133
|
+
provider: createMockProvider([providerTextDelta("hi"), providerDone()]),
|
|
134
|
+
}).createSession();
|
|
135
|
+
|
|
136
|
+
const detach = telemetry.attachSession(session);
|
|
137
|
+
const result = await session.run("hello");
|
|
138
|
+
detach();
|
|
139
|
+
telemetry.handleRunFeedback({ runId: result.runId, rating: 1, hasComment: true, tagCount: 1, scorerCount: 1, evaluationCount: 1 });
|
|
140
|
+
telemetry.handleEvaluation({ runId: result.runId, status: "scored", score: 0.9, hasReason: true });
|
|
141
|
+
|
|
142
|
+
console.log(memory.spans.map((span) => span.name));
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
## Extension and configuration notes
|
|
146
|
+
|
|
147
|
+
- Events flow through `redactAgentEvent` before subscribers and ledger writes — configure `createSecretRedactor` on the agent/run.
|
|
148
|
+
- `retry_scheduled` still signals backoff; each retry attempt emits its own `provider_turn_*` pair with `metadata.attempt`.
|
|
149
|
+
- NeuralWatt `neuralwatt:telemetry` provider events remain package-local; hosts may forward numeric cost/energy into custom metrics.
|
|
150
|
+
- `@arnilo/prism-observability-opentelemetry` is optional and included through `@arnilo/prism-sdk` and `@arnilo/prism-all`; instrumentation remains disabled until a host configures it.
|
|
151
|
+
- Exporter failures are isolated: instrumentation catches tracer/meter errors and invokes `onExporterError` without affecting the run, feedback persistence, or evaluation scoring.
|
|
152
|
+
- Run `error` events close every outstanding span attributable to that run. Detaching a session closes any remaining session spans; repeated terminal events are idempotent and cannot end a span twice.
|
|
153
|
+
- Disabled instrumentation performs no per-delta span work (`enabled: false` or missing tracer/meter).
|
|
154
|
+
|
|
155
|
+
## Security and performance notes
|
|
156
|
+
|
|
157
|
+
- Default events are metadata-only — no prompts, streamed deltas, tool arguments, or credentials.
|
|
158
|
+
- Opt-in content in other event types (`message_delta`, tool `result`) is still subject to `redactAgentEvent`.
|
|
159
|
+
- Metric labels stay low-cardinality (`provider_id`, `outcome`, `status`, token `kind`, feedback rating bucket/link presence); never use `sessionId`/`runId`, comments, tag values, scorer/evaluation IDs, or arbitrary metadata as labels. Provider-turn and run-total tokens use distinct instruments, so one counter cannot double count both scopes.
|
|
160
|
+
- Target overhead when enabled is under 5% excluding exporter I/O; disabled hooks allocate no spans.
|
|
161
|
+
- Provider transport limits and redaction order are documented in [Provider primitives](provider-primitives.md).
|
|
162
|
+
|
|
163
|
+
## Related APIs
|
|
164
|
+
- [Evaluations](evaluations.md): optional scorers can link scores to run/session/trace IDs from agent events.
|
|
165
|
+
|
|
166
|
+
- [Agent events](agent-events.md): full `AgentEvent` union and subscriber semantics.
|
|
167
|
+
- [Runs and usage ledger](runs-and-usage.md): durable `AgentEventRecord` persistence.
|
|
168
|
+
- [Middleware hooks](middleware-hooks.md): transform boundaries alongside event subscribers.
|
|
169
|
+
- [Provider primitives](provider-primitives.md): frozen observability contract for Plan 054.
|
|
170
|
+
- [Credentials and redaction](credentials-and-redaction.md): secret redaction before events and ledger rows.
|
|
171
|
+
- [Workflows](workflows.md): package-local `WorkflowEvent` stream that can wrap redacted `AgentEvent`s from agent nodes.
|
package/docs/performance.md
CHANGED
|
@@ -7,6 +7,7 @@ This page states Prism runtime limits that keep slow consumers and long sessions
|
|
|
7
7
|
Current surfaces:
|
|
8
8
|
|
|
9
9
|
- `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
|
|
10
|
+
- Bounded provider transport primitives (`readSseEvents`, `readBoundedResponseText`) used by every first-party provider package — see [Provider primitives](provider-primitives.md).
|
|
10
11
|
- `SessionStore.readBranchPath(query)` for branch reads that avoid full-session scans.
|
|
11
12
|
- `ProductionPersistenceStore` cursor queries for entries, events, runs, tool calls, and usage.
|
|
12
13
|
- JSONL and memory stores documented as development/local adapters, not production multi-writer stores.
|
|
@@ -112,10 +113,257 @@ const store = {
|
|
|
112
113
|
|
|
113
114
|
- Overflow events contain only counts and policy, never message text, tool arguments, prompts, provider payloads, or credentials.
|
|
114
115
|
- Runtime event payloads can be large (`Message`, content deltas, tool results, summaries, artifact metadata). Size queues by events and keep payload size in mind.
|
|
116
|
+
- `toolConcurrency` on the single-shot loop bounds in-flight tool dispatches per provider turn to `min(toolConcurrency, calls.length)`. Independent slow tools can overlap; transcript appends remain ordered. Default `1` preserves sequential behavior.
|
|
117
|
+
- `read` image bounds: default `maxImageBytes` is 10 MB (`DEFAULT_MAX_IMAGE_BYTES`). Oversize images are rejected by `stat` before read when possible; hosts may supply `transformImage` for resize/re-encode without adding image-processing deps to the base package.
|
|
115
118
|
- Live subscriber queues are bounded by default. Durable replay belongs to host storage.
|
|
116
119
|
- `SessionStore.list(sessionId)` is a full-session read. It is fine for memory/JSONL development stores, but production adapters should use `readBranchPath` for provider context and branch views.
|
|
117
120
|
- The JSONL store rereads/parses the file for validation/list/get and serializes appends only within one process. It has no cross-process lock, pagination, migrations, tenant isolation, or retention.
|
|
118
121
|
- Recommended database indexes: session id, run id, parent id, branch leaf id, timestamps, tenant/account/user, event type, entry kind, `(run_id, sequence)` for event timelines, and `(run_id, recorded_at, id)` for usage. Allocate event `sequence` per run for stable timeline pagination.
|
|
122
|
+
- Provider SSE parsing defaults: 256 KiB per completed event, 512 KiB incomplete buffer, 64 KiB error response bodies. Override per call via `BoundedStreamLimits` on `@arnilo/prism/providers/transport`.
|
|
123
|
+
|
|
124
|
+
### Provider-phase benchmark snapshot (2026-07-14)
|
|
125
|
+
|
|
126
|
+
Node v24.18.0, Linux x86_64, AMD Ryzen 9 PRO 7940HS; local synthetic streams, no network or exporter I/O:
|
|
127
|
+
|
|
128
|
+
| Case | Result |
|
|
129
|
+
| --- | --- |
|
|
130
|
+
| `readSseEvents`, 16 MiB total as 4 KiB events | 380 MiB/s, +1.7 MiB end-of-run heap delta |
|
|
131
|
+
| 100 provider deltas with 1 ms transport delay, no telemetry | 107.97 ms median |
|
|
132
|
+
| Same run, disabled adapter attached | 107.25 ms median (-0.67%, noise) |
|
|
133
|
+
| Same run, enabled no-op exporter | 107.22 ms median (-0.70%, noise) |
|
|
134
|
+
|
|
135
|
+
Nine measured runs per telemetry mode after warm-up; table reports median. A zero-I/O burst of 5,000 deltas measured 1.06 ms without telemetry and 2.09 ms with the adapter: about 1 ms absolute adapter cost, but a large percentage against an unrealistically tiny baseline. No span is created for message deltas. Add subscriber-side event filtering only if measured high-frequency in-memory streams make that ceiling material.
|
|
136
|
+
|
|
137
|
+
Configured overflow behavior is enforced by `src/__tests__/provider-transport.test.ts` for event, incomplete-buffer, response-body, argument, and abort limits.
|
|
138
|
+
|
|
139
|
+
### 0.0.4 release audit snapshot (2026-07-14)
|
|
140
|
+
|
|
141
|
+
Node v24.18.0 on the same Linux x86_64 / Ryzen 9 PRO 7940HS host. Results are medians of 7-9 warm runs unless the row describes file/database appends. Synthetic operations use local memory/files only. These numbers are release ceilings and comparison points, not cross-machine guarantees.
|
|
142
|
+
|
|
143
|
+
| Surface | Workload | Result | 0.0.4 release threshold |
|
|
144
|
+
| --- | --- | --- | --- |
|
|
145
|
+
| Run ledger | One mock run, 500 text deltas, 510 total records | 1.19 ms with in-memory ledger vs 0.17 ms without; +1.02 ms absolute; event append max concurrency 1 | < 10 ms with zero-I/O adapter; event appends remain serialized |
|
|
146
|
+
| JSONL store | 500 sequential label appends, including fail-closed reread/validation | 141.10 ms; 3,544 appends/s | < 500 ms; development/single-process only |
|
|
147
|
+
| JSON Schema compile cache | 5,000 validations through one warm adapter | 4.97 ms; 0.99 µs/validation | < 25 µs/validation |
|
|
148
|
+
| JSON Schema cold compile | 100 new adapters + first validation | 249.89 ms; 2.50 ms/compile | Warm cache must remain at least 20x faster than cold compile |
|
|
149
|
+
| Parallel tools | Six independent 20 ms calls | concurrency 1: 121.12 ms; concurrency 2: 60.92 ms; 1.99x speedup | concurrency 2 < 75% of sequential; configured worker cap remains enforced |
|
|
150
|
+
| SQLite session store | 1,000 sequential transactional label appends | 31.84 ms; 31,405 appends/s | < 250 ms on local SSD/tmp storage |
|
|
151
|
+
| Secret redaction | 10,000 shallow objects containing one known secret | 4.79 ms; 2.09 million objects/s | < 25 ms |
|
|
152
|
+
| Credential KDF | Default scrypt + AES-256-GCM encryption | 48.09 ms median | 20-250 ms; security floor stays `N >= 16,384` |
|
|
153
|
+
| Workflow runner | Existing bounded 1,000-node DAG fixture | 27.68 ms in aggregate release gate | < 1 s; no rescan failure |
|
|
154
|
+
|
|
155
|
+
Provider SSE remained at the frozen 380 MiB/s / +1.7 MiB heap snapshot. Media and MCP retain 10 MB defaults and finite timeout/total-byte guards; their malicious-input and oversize fixtures pass. PostgreSQL latency remains environment-dependent and is gated by transactional conformance in CI rather than a hardware-specific wall-clock assertion.
|
|
156
|
+
|
|
157
|
+
The ledger percentage overhead is intentionally not a threshold: its no-ledger baseline is below 1 ms, making the percentage unstable while absolute added latency remains about 1 ms. JSONL's append path is intentionally O(n²) across repeated appends because it rereads for corruption/conflict checks; move production or high-volume workloads to SQLite/PostgreSQL rather than weakening validation.
|
|
158
|
+
|
|
159
|
+
### 0.0.5 Phase 0 baseline (2026-07-15)
|
|
160
|
+
|
|
161
|
+
Scope froze at commit `f5128a816ae204c52f3e2f089de71c99bd5de6d4`. Measurement host: Node v24.18.0, npm 11.16.0, Linux 7.1.3 x86_64, AMD Ryzen 9 PRO 7940HS (16 logical CPUs). Supported package runtime remains Node >=20. These are dated local comparison points, not portable CI wall-clock assertions.
|
|
162
|
+
|
|
163
|
+
| Surface | Workload | Result |
|
|
164
|
+
| --- | --- | --- |
|
|
165
|
+
| Network-free tests | `npm test` | 25.750 s; 1,475 tests, 1,450 pass, 25 explicit live skips, 0 fail |
|
|
166
|
+
| Release readiness | `npm run sdk:ready` | 54.341 s; typecheck, tests, examples, builds, and 24 dry-run packs pass |
|
|
167
|
+
| Provider/agent stream | One mock run with 5,000 one-character text deltas and a concurrently drained 8,192-event subscriber | 3.78 ms median |
|
|
168
|
+
| Tool dispatch | Six independent 20 ms tools, concurrency 1 | 121.05 ms median |
|
|
169
|
+
| Tool dispatch | Same calls, concurrency 2 | 60.65 ms median (2.00x speedup) |
|
|
170
|
+
| Workflow runner | Existing bounded 1,000-node chain, configured concurrency 8 | 9.66 ms median |
|
|
171
|
+
| Package artifacts | All 24 dry-run tarballs | 542,993 packed bytes; 2,084,900 unpacked bytes aggregate |
|
|
172
|
+
| Root artifact | `@arnilo/prism@0.0.4` dry-run tarball | 346.0 kB packed; 1.3 MB unpacked; 196 files |
|
|
173
|
+
| Installed workspace | Current root `node_modules` | 72 MiB |
|
|
174
|
+
|
|
175
|
+
Synthetic stream/tool/workflow values are medians of seven measured runs after one warm-up and contain no network, database, or exporter I/O. The temporary benchmark reused public `AgentSession`, `dispatchToolCallsInOrder`, and `@arnilo/prism-workflows` APIs; it was not added to CI because this phase records a baseline rather than creating hardware-sensitive tests.
|
|
176
|
+
|
|
177
|
+
Repository size at the same commit, counted from `src/` and `packages/` while excluding `dist/`:
|
|
178
|
+
|
|
179
|
+
| Area | Files | Lines |
|
|
180
|
+
| --- | ---: | ---: |
|
|
181
|
+
| Production TypeScript | 189 | 26,828 |
|
|
182
|
+
| Test TypeScript | 144 | 23,535 |
|
|
183
|
+
| Documentation Markdown | 70 | 12,662 |
|
|
184
|
+
| Numbered plans | 58 | 24,270 |
|
|
185
|
+
| TypeScript examples | 39 | 3,134 |
|
|
186
|
+
|
|
187
|
+
Prism has no project generator before Phase 5, so a generated-Prism-project install/build size is **not applicable** at this baseline. The closest current install figure is the 72 MiB development workspace; it is not a scaffold target. The comparison Mastra default scaffold measured during the review used 439 MB `node_modules`, 300 MB build output, and 427 installed packages. Phase 5 must establish a real generated Prism project baseline and keep unselected storage, telemetry, eval, memory, server, and workflow dependencies absent.
|
|
188
|
+
|
|
189
|
+
See [Review coverage — 2026-07-15](review-coverage-2026-07-15.md) for scope, primitive, package, and threat-boundary ownership.
|
|
190
|
+
|
|
191
|
+
### 0.0.5 Phase 2 verification (2026-07-15)
|
|
192
|
+
|
|
193
|
+
Same Phase 0 host and seven-run warm benchmark. Runtime correctness changes stayed inside frozen ceilings:
|
|
194
|
+
|
|
195
|
+
| Surface | Result |
|
|
196
|
+
| --- | --- |
|
|
197
|
+
| Network-free tests | 27.992 s; 1,485 tests, 1,460 pass, 25 explicit live skips, 0 fail |
|
|
198
|
+
| `npm run sdk:ready` | 55.598 s; typecheck, examples, tests, builds, and all 24 dry-run packs pass |
|
|
199
|
+
| Provider/agent stream, 5,000 deltas | 3.54 ms median (Phase 0: 3.78 ms) |
|
|
200
|
+
| Six 20 ms tools, concurrency 1 / 2 | 121.22 ms / 60.63 ms (2.00x speedup retained) |
|
|
201
|
+
| Workflow 1,000-node chain | 10.31 ms median (well below 1 s ceiling) |
|
|
202
|
+
| Root dry-run tarball | 361.2 kB packed, 1.3 MB unpacked, 197 files |
|
|
203
|
+
|
|
204
|
+
Usage aggregation performs one constant-size accumulator update per terminal provider turn. Telemetry retains only active span metadata and removes every terminal/detached entry. Complete media resolution is sequential, rejects item count and inline estimates before I/O, and retains at most the request budget plus one per-item-bounded candidate before failing an aggregate overflow. Sandbox output still streams into the existing bounded `OutputAccumulator`; no adapter-side response buffer was added.
|
|
205
|
+
|
|
206
|
+
### 0.0.5 Phase 4 verification (2026-07-15)
|
|
207
|
+
|
|
208
|
+
Optional `@arnilo/prism-evals` adds package-local scoring without changing core run latency. Validation stayed within the frozen release gate:
|
|
209
|
+
|
|
210
|
+
| Surface | Result |
|
|
211
|
+
| --- | --- |
|
|
212
|
+
| Network-free tests | 1,503 tests, 1,478 pass, 25 explicit live skips, 0 fail |
|
|
213
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 25 dry-run packs pass |
|
|
214
|
+
| Evals dry-run tarball | 35.4 kB unpacked package payload |
|
|
215
|
+
| Profile bundles | unchanged; evals remains opt-in until size/use review |
|
|
216
|
+
|
|
217
|
+
Experiment concurrency is capped at 32 workers and defaults to 1. Scorers operate on `AgentRunResult` references plus dataset item metadata rather than duplicating event ledgers.
|
|
218
|
+
|
|
219
|
+
### 0.0.5 Phase 5 verification (2026-07-15)
|
|
220
|
+
|
|
221
|
+
`prism init` lands as a stdlib-only CLI subcommand with checked-in templates under `templates/init/`.
|
|
222
|
+
|
|
223
|
+
| Surface | Result |
|
|
224
|
+
| --- | --- |
|
|
225
|
+
| Default generated sources | 8 files / ~3.3 KB |
|
|
226
|
+
| Default clean consumer install (`@arnilo/prism` + TypeScript tooling) | ~27.5 MB `node_modules` |
|
|
227
|
+
| Mastra comparator | 439 MB install / 300 MB build / 427 packages |
|
|
228
|
+
| Default dependencies | `@arnilo/prism` only; no storage, telemetry, eval, memory, server, or workflow packages unless `--with-*` / provider flags select them |
|
|
229
|
+
| Offline proof | packed core tarball → `npm install` → `npm run typecheck` → `npm test` (mock provider) |
|
|
230
|
+
|
|
231
|
+
### 0.0.5 Phase 6 verification (2026-07-15)
|
|
232
|
+
|
|
233
|
+
Optional `@arnilo/prism-provider-ai-sdk` adapts AI SDK `LanguageModelV4` streams to Prism without adding an AI SDK dependency to core.
|
|
234
|
+
|
|
235
|
+
| Surface | Result |
|
|
236
|
+
| --- | --- |
|
|
237
|
+
| Supported specification | `@ai-sdk/provider@^4` (`LanguageModelV4`) |
|
|
238
|
+
| Adapter behavior | incremental stream translation; unsupported content fails before `doStream`; abort owned by Prism `request.signal` |
|
|
239
|
+
| Network-free tests | 1,522 tests, 1,497 pass, 25 explicit live skips, 0 fail |
|
|
240
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 26 dry-run packs pass |
|
|
241
|
+
| AI SDK adapter dry-run tarball | 6.5 kB packed / 22.5 kB unpacked / 16 files |
|
|
242
|
+
| Profile bundles | unchanged; AI SDK adapter remains opt-in until size/use review |
|
|
243
|
+
| Publishable graph | 26 packages |
|
|
244
|
+
|
|
245
|
+
### 0.0.5 Phase 7 verification (2026-07-15)
|
|
246
|
+
|
|
247
|
+
Optional `@arnilo/prism-memory` adds working memory and semantic recall without changing core session stores.
|
|
248
|
+
|
|
249
|
+
| Surface | Result |
|
|
250
|
+
| --- | --- |
|
|
251
|
+
| Contracts | package-owned `Embedder`, `VectorStore`, `WorkingMemoryStore`, `createMemory` |
|
|
252
|
+
| Adapters | in-memory reference + PostgreSQL/pgvector production path |
|
|
253
|
+
| Injection | existing `ContextProvider` seam; opt-in working-memory processor |
|
|
254
|
+
| Profile bundles | unchanged; memory remains opt-in until size/use review |
|
|
255
|
+
| Publishable graph | 27 packages |
|
|
256
|
+
| Network-free tests | 1,538 tests, 1,513 pass, 25 explicit live skips, 0 fail |
|
|
257
|
+
| `npm run sdk:ready` | pass |
|
|
258
|
+
| Memory dry-run tarball | 17.9 kB packed / 76.6 kB unpacked / 32 files |
|
|
259
|
+
|
|
260
|
+
### 0.0.5 Phase 8 verification (2026-07-15)
|
|
261
|
+
|
|
262
|
+
Durable human suspension extends existing workflow checkpoint JSON/CAS; no worker polling loop, package, dependency, or database migration was added.
|
|
263
|
+
|
|
264
|
+
| Surface | Result |
|
|
265
|
+
| --- | --- |
|
|
266
|
+
| Focused workflow suite | 43 tests pass, 0 fail |
|
|
267
|
+
| Network-free tests | 1,547 tests, 1,522 pass, 25 explicit live skips, 0 fail |
|
|
268
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 27 dry-run packs pass |
|
|
269
|
+
| Workflow dry-run tarball | 25.7 kB packed / 121.6 kB unpacked / 34 files |
|
|
270
|
+
| Coordinator behavior | `suspended` absent from queued/running poll; zero worker/lease retained |
|
|
271
|
+
| Storage | existing bounded checkpoint JSON/category; no SQLite/PostgreSQL migration |
|
|
272
|
+
|
|
273
|
+
### 0.0.5 Phase 9 verification (2026-07-16)
|
|
274
|
+
|
|
275
|
+
Optional `@arnilo/prism-rag` reuses Phase 7 vector contracts and adds no core path, parser dependency, network loader, or profile activation.
|
|
276
|
+
|
|
277
|
+
| Surface | Result |
|
|
278
|
+
| --- | --- |
|
|
279
|
+
| Focused RAG suite | 9 tests pass, 0 fail |
|
|
280
|
+
| Network-free tests | 1,561 tests, 1,536 pass, 25 explicit live skips, 0 fail |
|
|
281
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 28 dry-run packs pass |
|
|
282
|
+
| RAG dry-run tarball | 9.0 kB packed / 34.6 kB unpacked / 22 files |
|
|
283
|
+
| Index bounds | chunk/document/count/metadata caps; embed batches default 32, hard 128 |
|
|
284
|
+
| Retrieval bounds | top-K default 5/hard 32; candidates default 20/hard 128; result 64/512 KiB; context 2,000/8,000 estimated tokens |
|
|
285
|
+
| Profile bundles | unchanged; RAG and memory remain explicit opt-ins |
|
|
286
|
+
|
|
287
|
+
### 0.0.5 Phase 10 verification (2026-07-16)
|
|
288
|
+
|
|
289
|
+
Optional `@arnilo/prism-server` and MCP server-direction APIs compose existing agent/workflow/tool/SDK primitives; no core path, framework/listener, auth provider, database, or profile activation was added.
|
|
290
|
+
|
|
291
|
+
| Surface | Result |
|
|
292
|
+
| --- | --- |
|
|
293
|
+
| Focused server suites | 6 Web handler tests + 4 MCP server tests pass; existing 12 MCP client tests remain green |
|
|
294
|
+
| Network-free tests | 1,576 tests, 1,551 pass, 25 explicit live skips, 0 fail |
|
|
295
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 29 dry-run packs pass |
|
|
296
|
+
| Server dry-run tarball | 8.4 kB packed / 34.4 kB unpacked / 12 files |
|
|
297
|
+
| MCP dry-run tarball | 11.6 kB packed / 45.0 kB unpacked / 20 files |
|
|
298
|
+
| Web handler bounds | request 64 KiB, result 1 MiB, event 64 KiB, stream 10 MiB/10k events, queue 128, concurrency 16, timeout 120 s by default; all have hard caps |
|
|
299
|
+
| MCP server bounds | call result 1 MiB, calls 16, timeout 60 s; HTTP request 1 MiB, response 2 MiB, requests 32 by default; all have hard caps |
|
|
300
|
+
| Profile bundles | unchanged; server remains explicit opt-in |
|
|
301
|
+
|
|
302
|
+
### 0.0.5 Phase 11 verification (2026-07-16)
|
|
303
|
+
|
|
304
|
+
Workflow schedules, background runs, composition, state, and replay reuse the existing workflow package plus generic checkpoint/lease stores. No package, runtime dependency, SQL migration, listener, cron parser, or auto-started worker was added.
|
|
305
|
+
|
|
306
|
+
| Surface | Result |
|
|
307
|
+
| --- | --- |
|
|
308
|
+
| Focused workflow/server suites | 54 workflow tests + 8 Web handler tests pass, 0 fail |
|
|
309
|
+
| Network-free tests | 1,589 tests, 1,564 pass, 25 explicit live skips, 0 fail |
|
|
310
|
+
| `npm run sdk:ready` | typecheck, examples, tests, builds, and all 29 dry-run packs pass |
|
|
311
|
+
| Workflow dry-run tarball | 34.7 kB packed / 171.5 kB unpacked / 38 files |
|
|
312
|
+
| Server dry-run tarball | 9.9 kB packed / 45.2 kB unpacked / 12 files |
|
|
313
|
+
| Synthetic schedule bound | 100 in-memory creates: 1.28 ms; scan 100 / claim+enqueue 16 due fires: 6.64 ms |
|
|
314
|
+
| Synthetic composition/replay | depth-8 nested run: 2.73 ms; 100-node source: 35.42 ms; replay 50 nodes: 21.17 ms |
|
|
315
|
+
| State/replay ceilings | state 64/512 KiB; history 32/128; nested depth 8/32; replay depth 8/32 default/hard |
|
|
316
|
+
| Schedule ceilings | page 100/500; claims 16/256; input 256 KiB/1 MiB; 1s idle timer; 30s fire lease defaults |
|
|
317
|
+
|
|
318
|
+
Synthetic timings are one local Node v24.18.0 run over memory adapters with no network/database I/O; finite limits and behavior tests, not wall-clock numbers, are CI gates.
|
|
319
|
+
|
|
320
|
+
### 0.0.5 Phase 12 verification (2026-07-16)
|
|
321
|
+
|
|
322
|
+
Run feedback adds no package or runtime dependency. Memory/SQLite/PostgreSQL implementations share bounded append/query/delete semantics; OTel projection accepts only fixed scalar metadata.
|
|
323
|
+
|
|
324
|
+
| Surface | Result |
|
|
325
|
+
| --- | --- |
|
|
326
|
+
| Focused feedback/eval/SQLite/OTel tests | 35 tests pass, 0 fail; PostgreSQL DDL suite passes and live feedback conformance is env-gated |
|
|
327
|
+
| Synthetic memory feedback | 1,000 bounded appends: 3.82 ms; 100 filtered 100-row queries over 1,000 records: 11.53 ms |
|
|
328
|
+
| Core dry-run tarball | 398.1 kB packed / 1.4 MB unpacked / 219 files |
|
|
329
|
+
| Evals dry-run tarball | 9.8 kB packed / 38.4 kB unpacked / 26 files |
|
|
330
|
+
| OTel dry-run tarball | 6.3 kB packed / 26.5 kB unpacked / 8 files |
|
|
331
|
+
| SQLite/PostgreSQL tarballs | 17.7/18.1 kB packed; 89.8/89.8 kB unpacked |
|
|
332
|
+
| Feedback limits | comment 4/16 KiB; tags 16/64; links 16/64; metadata 16/64 KiB; pages 100/500 default/hard |
|
|
333
|
+
|
|
334
|
+
Metrics came from one local Node v24.18.0 memory-adapter run. SQL correctness/indexing/migration behavior and hard bounds are gates; local timings are not release thresholds.
|
|
335
|
+
|
|
336
|
+
### 0.0.5 Phase 13 verification (2026-07-16)
|
|
337
|
+
|
|
338
|
+
Supervisor/A2A stays in one optional zero-runtime-dependency package; core and profile bundles gained no import, listener, worker, protocol SDK, or network activation.
|
|
339
|
+
|
|
340
|
+
| Surface | Result |
|
|
341
|
+
| --- | --- |
|
|
342
|
+
| Focused supervisor/A2A suite | 11 tests pass, 0 fail; local delegation, policy/budget/abort/redaction, card signatures, server/client/stream bounds |
|
|
343
|
+
| Synthetic local delegation | 100 sequential mock child results: 11.83 ms |
|
|
344
|
+
| Synthetic in-process A2A | 100 card discovery + JSON-RPC mock round trips: 34.17 ms |
|
|
345
|
+
| Supervisor dry-run tarball | 15.3 kB packed / 69.4 kB unpacked / 22 files |
|
|
346
|
+
| Local hard ceilings | depth 16; active 32; message 1 MiB; steps 64; tools 256; tokens 1m; timeout 30m; event queue 4096 |
|
|
347
|
+
| A2A hard ceilings | request/card/event 1 MiB; response 8 MiB; stream 64 MiB/100k events; concurrency 256; timeout 30m |
|
|
348
|
+
|
|
349
|
+
Timings are one local Node v24.18.0 run over mock agents and an in-process fetch adapter. Bounds, protocol validation, signature/auth/origin checks, and offline behavior tests are release gates; timings are not thresholds.
|
|
350
|
+
|
|
351
|
+
### 0.0.5 Phase 14 release-candidate verification (2026-07-16)
|
|
352
|
+
|
|
353
|
+
| Surface | Result |
|
|
354
|
+
| --- | --- |
|
|
355
|
+
| Default network-free test | 32.247 s, below 60 s budget |
|
|
356
|
+
| Full SDK readiness | 70.560 s; build/typecheck/examples/tests/30 pack dry-runs |
|
|
357
|
+
| Test matrix | 1,618 total; 1,593 pass; 25 explicit live skips; 0 fail |
|
|
358
|
+
| Node compatibility | Node 20.20.2 imports 44 built root/package export targets; Node 24.18.0 runs full matrix |
|
|
359
|
+
| PostgreSQL/pgvector | 29 live checks pass in fresh `pgvector/pgvector:pg16` container |
|
|
360
|
+
| Packed artifact set | 30 tarballs / 699 files; post-bundle snapshot ~690.6 kB packed / 2.64 MB unpacked |
|
|
361
|
+
| Core artifact | post-bundle snapshot ~403.7 kB packed / 1.46 MB unpacked / 221 files |
|
|
362
|
+
| Generated default project | under 50 KiB source and under 50 MiB installed; packed-core typecheck/test pass |
|
|
363
|
+
| Fresh packed journey | 30 packages install/import and Phase 1-13 optional composition pass in ~8.0 s |
|
|
364
|
+
| Registry/publish preview | 30/30 versions available; 30/30 dependency-ordered provenance dry-runs pass |
|
|
365
|
+
|
|
366
|
+
No performance ceiling was raised. Core grew from Phase 0's 346.0 kB packed baseline to ~403.7 kB after documented APIs/templates, while the full package set remains ~690.6 kB packed. Follow-up review includes all six Phase 4-13 capability packages through `prism-all` and AI SDK interoperability through `prism-providers`; focused base/code/SDK profiles remain unchanged and no capability auto-activates. Manifest tarballs remain tiny: providers 1.4 kB and all 1.6 kB packed.
|
|
119
367
|
|
|
120
368
|
## Related APIs
|
|
121
369
|
|
|
@@ -124,4 +372,4 @@ const store = {
|
|
|
124
372
|
- [Session stores](session-stores.md): `SessionStore.readBranchPath` and dev-vs-production branch reads.
|
|
125
373
|
- [Database persistence](database-persistence.md): cursor queries, reference schema, indexes, and event sequence guidance.
|
|
126
374
|
- [Runs and usage ledger](runs-and-usage.md): durable event, tool-call, and usage persistence.
|
|
127
|
-
- [
|
|
375
|
+
- [Provider primitives](provider-primitives.md): bounded SSE/error-body limits for first-party providers.
|