@arnilo/prism 0.0.96 → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +290 -2
- package/README.md +17 -3
- package/dist/agent-definitions.js +2 -3
- package/dist/agent-event-source.d.ts +11 -0
- package/dist/agent-event-source.js +512 -0
- package/dist/agent-loops.d.ts +5 -0
- package/dist/agent-loops.js +99 -14
- package/dist/agent-run-lifecycle.d.ts +5 -2
- package/dist/agent-run-lifecycle.js +18 -2
- package/dist/agent-run-state.d.ts +27 -1
- package/dist/agent-run-state.js +113 -7
- package/dist/agents.d.ts +3 -1
- package/dist/agents.js +1255 -129
- package/dist/artifacts.d.ts +132 -0
- package/dist/artifacts.js +44 -0
- package/dist/cache-helpers.js +18 -9
- package/dist/checkpoints.d.ts +4 -0
- package/dist/checkpoints.js +17 -9
- package/dist/cli-init.js +3 -7
- package/dist/cli-runner.d.ts +2 -6
- package/dist/cli-runner.js +71 -33
- package/dist/compaction.js +5 -4
- package/dist/config.js +7 -4
- package/dist/content.js +26 -24
- package/dist/context-budget.d.ts +67 -0
- package/dist/context-budget.js +288 -0
- package/dist/contracts.d.ts +590 -8
- package/dist/contracts.js +142 -1
- package/dist/contribution-parsing.js +6 -2
- package/dist/contributions.d.ts +2 -0
- package/dist/contributions.js +3 -0
- package/dist/conversations.d.ts +50 -0
- package/dist/conversations.js +98 -0
- package/dist/credentials.d.ts +22 -2
- package/dist/credentials.js +18 -3
- package/dist/devices.d.ts +94 -0
- package/dist/devices.js +138 -0
- package/dist/event-multiplexer.js +18 -4
- package/dist/extensions.d.ts +18 -1
- package/dist/extensions.js +79 -6
- package/dist/feedback.js +12 -10
- package/dist/guardrails.d.ts +1 -1
- package/dist/guardrails.js +26 -17
- package/dist/identity.d.ts +92 -0
- package/dist/identity.js +265 -0
- package/dist/index.d.ts +94 -72
- package/dist/index.js +48 -36
- package/dist/input.d.ts +10 -1
- package/dist/input.js +152 -52
- package/dist/instruction-injection.d.ts +1 -1
- package/dist/middleware.js +9 -1
- package/dist/models.d.ts +2 -0
- package/dist/models.js +3 -0
- package/dist/node/agent-definitions.js +16 -8
- package/dist/node/contribution-discovery.d.ts +1 -2
- package/dist/node/contribution-discovery.js +3 -3
- package/dist/node/session-store-jsonl.js +13 -7
- package/dist/node/settings.d.ts +1 -1
- package/dist/node/settings.js +1 -1
- package/dist/node/system-project-prompts.js +2 -4
- package/dist/node/trust.js +1 -1
- package/dist/persistence-lifecycle.d.ts +103 -0
- package/dist/persistence-lifecycle.js +202 -0
- package/dist/provider-events.d.ts +1 -0
- package/dist/provider-events.js +6 -1
- package/dist/provider-request-policy.js +3 -4
- package/dist/providers/media.d.ts +1 -1
- package/dist/providers/openai-compatible.d.ts +46 -1
- package/dist/providers/openai-compatible.js +123 -53
- package/dist/providers/openai-primitives.js +10 -7
- package/dist/providers/transport.d.ts +6 -0
- package/dist/providers/transport.js +21 -0
- package/dist/providers.d.ts +2 -0
- package/dist/providers.js +3 -0
- package/dist/redaction.d.ts +1 -0
- package/dist/redaction.js +26 -9
- package/dist/resources.d.ts +2 -2
- package/dist/resources.js +2 -2
- package/dist/retry.d.ts +5 -0
- package/dist/retry.js +8 -1
- package/dist/rpc.js +55 -11
- package/dist/run-ledger.d.ts +6 -0
- package/dist/run-ledger.js +16 -13
- package/dist/run-limits.js +49 -10
- package/dist/secure-agent.js +8 -2
- package/dist/security.js +7 -2
- package/dist/session-stores.d.ts +7 -2
- package/dist/session-stores.js +195 -21
- package/dist/skill-disclosure.d.ts +35 -0
- package/dist/skill-disclosure.js +101 -0
- package/dist/skill-load.d.ts +25 -0
- package/dist/skill-load.js +112 -0
- package/dist/structured-output.d.ts +5 -1
- package/dist/structured-output.js +20 -2
- package/dist/system-prompts.js +7 -2
- package/dist/testing/agent-event-source-conformance.d.ts +4 -0
- package/dist/testing/agent-event-source-conformance.js +54 -0
- package/dist/testing/compaction-conformance.js +5 -1
- package/dist/testing/extension-conformance.js +15 -3
- package/dist/testing/feedback.d.ts +1 -3
- package/dist/testing/feedback.js +1 -1
- package/dist/testing/persistence-schema.d.ts +2 -2
- package/dist/testing/persistence-schema.js +280 -35
- package/dist/testing/provider-conformance.js +3 -3
- package/dist/testing/run-ledger-conformance.js +1 -1
- package/dist/testing/session-store-conformance.d.ts +6 -0
- package/dist/testing/session-store-conformance.js +37 -2
- package/dist/testing/tool-conformance.js +30 -5
- package/dist/testing/tool-effect-store-conformance.d.ts +9 -0
- package/dist/testing/tool-effect-store-conformance.js +85 -0
- package/dist/thinking.js +4 -1
- package/dist/tool-effects.d.ts +15 -0
- package/dist/tool-effects.js +352 -0
- package/dist/tool-result-fold.d.ts +40 -0
- package/dist/tool-result-fold.js +176 -0
- package/dist/tools.d.ts +8 -3
- package/dist/tools.js +248 -13
- package/docs/0.1.0-readiness.md +215 -0
- package/docs/a2a.md +33 -2
- package/docs/acp.md +152 -0
- package/docs/ag-ui-adoption.md +77 -0
- package/docs/ag-ui.md +225 -0
- package/docs/agent-events.md +34 -3
- package/docs/agent-identity.md +144 -0
- package/docs/agent-loops.md +17 -2
- package/docs/agent-session-runtime.md +21 -4
- package/docs/browser-automation.md +5 -0
- package/docs/caveman.md +129 -0
- package/docs/cli-rpc.md +3 -6
- package/docs/coding-agent-tools.md +229 -25
- package/docs/coding-security.md +77 -11
- package/docs/compaction-and-retry.md +5 -2
- package/docs/compaction-llm.md +20 -1
- package/docs/compaction-observational-memory.md +52 -8
- package/docs/context-and-skills.md +94 -7
- package/docs/contribution-registries.md +1 -0
- package/docs/conversations.md +135 -0
- package/docs/credential-storage.md +34 -1
- package/docs/credentials-and-redaction.md +11 -1
- package/docs/database-persistence.md +27 -7
- package/docs/device-adapters.md +97 -0
- package/docs/enterprise-postgres-state.md +178 -0
- package/docs/evaluations.md +14 -1
- package/docs/extensions.md +4 -1
- package/docs/forge-integration.md +113 -0
- package/docs/guardrails.md +16 -2
- package/docs/host-security.md +35 -4
- package/docs/index.md +69 -37
- package/docs/input-and-prompt-assembly.md +8 -7
- package/docs/language-intelligence.md +162 -0
- package/docs/mcp-tools.md +62 -5
- package/docs/middleware-hooks.md +2 -2
- package/docs/migration.md +427 -2
- package/docs/model-routing.md +111 -0
- package/docs/multimodal-content.md +8 -5
- package/docs/node-jsonl-session-store.md +1 -1
- package/docs/observability.md +2 -0
- package/docs/openapi-tools.md +56 -0
- package/docs/performance.md +282 -0
- package/docs/policy-and-audit.md +171 -0
- package/docs/ponytail.md +127 -0
- package/docs/postgres-persistence.md +8 -4
- package/docs/process-sessions.md +147 -0
- package/docs/provider-caching.md +13 -1
- package/docs/provider-conformance.md +29 -5
- package/docs/provider-packages.md +43 -2
- package/docs/provider-request-policies.md +2 -0
- package/docs/providers/ai-sdk.md +24 -7
- package/docs/providers/alibaba.md +179 -0
- package/docs/providers/anthropic.md +93 -0
- package/docs/providers/azure.md +74 -0
- package/docs/providers/bedrock.md +72 -0
- package/docs/providers/google.md +89 -0
- package/docs/providers/ollama.md +166 -0
- package/docs/providers/openai-compatible.md +31 -2
- package/docs/providers/openai.md +24 -5
- package/docs/providers/openrouter.md +2 -0
- package/docs/providers/vertex.md +71 -0
- package/docs/public-contracts.md +68 -4
- package/docs/rag.md +41 -12
- package/docs/release-and-install.md +362 -208
- package/docs/resource-loading.md +3 -0
- package/docs/runs-and-usage.md +3 -0
- package/docs/server.md +44 -6
- package/docs/session-store-conformance.md +2 -0
- package/docs/session-stores.md +41 -2
- package/docs/sqlite-persistence.md +11 -3
- package/docs/structured-output.md +7 -1
- package/docs/supervisors.md +8 -0
- package/docs/tool-effects.md +95 -0
- package/docs/tools.md +5 -0
- package/docs/work-artifacts-and-review.md +102 -0
- package/docs/work-connectors.md +32 -0
- package/docs/work-tools.md +137 -0
- package/docs/workflows.md +6 -0
- package/docs/working-and-semantic-memory.md +40 -7
- package/package.json +30 -7
- package/templates/init/providers.json +22 -0
- package/docs/review-coverage-2026-07-14.md +0 -260
- package/docs/review-coverage-2026-07-15.md +0 -193
- package/docs/review-coverage-2026-07-17-provider-validation.md +0 -192
- package/docs/review-coverage-2026-07-19-phase-3.md +0 -174
- package/docs/review-coverage-2026-07-20-phase-4.md +0 -175
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Model routing
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-model-router` is an optional governance facade over an existing `ProviderResolver`. It enforces allow-lists, residency, token/cost budgets, rate limits, circuit breaking, and bounded fallbacks before provider selection, and emits redacted selection diagnostics. It does not implement a second provider runtime.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
Use it for enterprise hosts that must deny models/regions before any provider call and attribute selection for audit. Skip it when a plain `createProviderResolver` allow-list is enough.
|
|
10
|
+
|
|
11
|
+
Do not put secrets, prompts, or raw OpenRouter keys into diagnostics. Do not honor `compat.openRouterRouting` unless `allowOpenRouterRouting: true`.
|
|
12
|
+
|
|
13
|
+
## Inputs / request
|
|
14
|
+
|
|
15
|
+
| API / field | Meaning |
|
|
16
|
+
| --- | --- |
|
|
17
|
+
| `createModelRouter({ resolver, stateStore?, ... })` | Wraps host `ProviderResolver`; omit `stateStore` for in-process memory state or pass durable async state. |
|
|
18
|
+
| `allowList.providers` / `allowList.models` | Exact provider id / model id or `provider/model` |
|
|
19
|
+
| `allowedResidencies` | Request residency must match when configured |
|
|
20
|
+
| `budgets` / per-call `maxTokens` / `maxCostUsd` | Finite non-negative ceilings; `recordUsage` charges |
|
|
21
|
+
| `rateLimit` | Per identity+model key window |
|
|
22
|
+
| `circuit` | Failure threshold + cooldown; keys capped |
|
|
23
|
+
| `fallbacks` | Ordered candidates after primary; total attempts capped |
|
|
24
|
+
| `allowOpenRouterRouting` | Default `false`; when false, routing metadata is stripped |
|
|
25
|
+
| `onDiagnostics` | Optional redacted hook (e.g. policy ledger evidence ref) |
|
|
26
|
+
| `router.resolve({ model, identity?, residency?, ... })` | Rich async selection |
|
|
27
|
+
| `router.providerSource` | Sync facade only for memory state; with `stateStore` it throws `ERR_PRISM_MODEL_ROUTER_ASYNC_STATE` rather than bypass durable checks. |
|
|
28
|
+
|
|
29
|
+
Frozen caps (default / hard): attempts `3 / 8`, circuit keys `1,024 / 16,384`, diagnostics `8 KiB / 64 KiB`.
|
|
30
|
+
|
|
31
|
+
## Outputs / response / events
|
|
32
|
+
|
|
33
|
+
- `ModelRouterResolveResult` — selected `provider` + possibly stripped `model`, `diagnostics`, and `providerRequestPolicy`.
|
|
34
|
+
- Deny throws `ModelRouterError` with code + redacted `diagnostics` (allow-list/residency/budget fail closed without calling resolver).
|
|
35
|
+
- `await recordOutcome({ identity, success, circuitProbeToken? })` opens/closes circuits; `await recordUsage({ identity, ... })` advances budgets. Pass the probe token returned by `resolve` for a half-open outcome.
|
|
36
|
+
|
|
37
|
+
## Request/response example
|
|
38
|
+
|
|
39
|
+
```json
|
|
40
|
+
{
|
|
41
|
+
"outcome": "allow",
|
|
42
|
+
"selectedProvider": "openrouter",
|
|
43
|
+
"selectedModel": "auto",
|
|
44
|
+
"attempts": [
|
|
45
|
+
{ "provider": "openai", "model": "gpt-4o", "outcome": "circuit_open", "reason": "circuit_open" },
|
|
46
|
+
{ "provider": "openrouter", "model": "auto", "outcome": "selected" }
|
|
47
|
+
],
|
|
48
|
+
"identityRefs": { "tenantId": "t1", "principalId": "a1", "principalKind": "agent" },
|
|
49
|
+
"openRouterRoutingHonored": false,
|
|
50
|
+
"residency": "eu"
|
|
51
|
+
}
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
## Implementation example
|
|
55
|
+
|
|
56
|
+
```ts
|
|
57
|
+
import { createAgent, createProviderResolver } from "@arnilo/prism";
|
|
58
|
+
import { createModelRouter } from "@arnilo/prism-model-router";
|
|
59
|
+
import { createPostgresEnterpriseState } from "@arnilo/prism-enterprise-postgres";
|
|
60
|
+
|
|
61
|
+
const enterprise = await createPostgresEnterpriseState({ pool, schema: "prism" });
|
|
62
|
+
const router = createModelRouter({
|
|
63
|
+
resolver: createProviderResolver(providers),
|
|
64
|
+
stateStore: enterprise.modelRouter,
|
|
65
|
+
allowList: { providers: ["openai", "openrouter"] },
|
|
66
|
+
allowedResidencies: ["eu"],
|
|
67
|
+
fallbacks: [{ provider: "openrouter", model: "auto" }],
|
|
68
|
+
allowOpenRouterRouting: false,
|
|
69
|
+
});
|
|
70
|
+
|
|
71
|
+
const { provider, model, providerRequestPolicy } = await router.resolve({
|
|
72
|
+
model: sessionModel,
|
|
73
|
+
identity,
|
|
74
|
+
residency: "eu",
|
|
75
|
+
maxCostUsd: 0.25,
|
|
76
|
+
});
|
|
77
|
+
|
|
78
|
+
const agent = createAgent({
|
|
79
|
+
model,
|
|
80
|
+
provider,
|
|
81
|
+
providerRequestPolicies: [providerRequestPolicy],
|
|
82
|
+
});
|
|
83
|
+
await router.recordUsage({ identity, provider: provider.id, model: model.model, tokens: 500 });
|
|
84
|
+
await router.recordOutcome({ identity, provider: provider.id, model: model.model, success: true });
|
|
85
|
+
await enterprise.close();
|
|
86
|
+
|
|
87
|
+
// `router.providerSource` is unavailable with durable state; resolve before provider I/O.
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
## Extension and configuration notes
|
|
91
|
+
|
|
92
|
+
Router is optional. Chain returned `providerRequestPolicy` with other `ProviderRequestPolicy` values. Wire `onDiagnostics` to `@arnilo/prism-policy` when audit export is required. OpenRouter package behavior is unchanged; routing metadata participates only when this gate allows it.
|
|
93
|
+
|
|
94
|
+
## Security and performance notes
|
|
95
|
+
|
|
96
|
+
- Allow-list and residency denies never call the underlying resolver.
|
|
97
|
+
- Without `stateStore`, budget/rate/circuit state is process-local, memory-capped, and oldest keys evict. It is not a cross-replica production path.
|
|
98
|
+
- With `stateStore: createPostgresEnterpriseState(...).modelRouter`, rate/budget updates and circuit probes are atomic across replicas, use database time, and are owner/principal/provider/model scoped. Router calls become asynchronous and require verified identity.
|
|
99
|
+
- Diagnostics carry identity refs and attempt outcomes only — no prompts/secrets. Durable state stores at most bounded numeric/timestamp/token material, never prompts or credentials.
|
|
100
|
+
- Selection is O(attempts × state operations); no provider network I/O happens inside state updates. Recorded 0.0.23 PostgreSQL p95 point operations stayed under 50 ms and cursor/cleanup pages under 100 ms on the documented fixture.
|
|
101
|
+
- Raising hard caps requires a reviewed release update with tests and docs.
|
|
102
|
+
|
|
103
|
+
## Related APIs
|
|
104
|
+
|
|
105
|
+
- [Provider layer](provider-layer.md)
|
|
106
|
+
- [Provider request policies](provider-request-policies.md)
|
|
107
|
+
- [OpenRouter](providers/openrouter.md)
|
|
108
|
+
- [Policy and audit](policy-and-audit.md)
|
|
109
|
+
- [Agent identity](agent-identity.md)
|
|
110
|
+
- [Enterprise PostgreSQL state](enterprise-postgres-state.md): durable router state, migration, cleanup, and ownership requirements.
|
|
111
|
+
- Package README: [`@arnilo/prism-model-router`](../packages/model-router/README.md)
|
|
@@ -49,11 +49,13 @@ Known `ModelCapabilities.input` tags are exported as `MODEL_INPUT_CAPABILITIES`:
|
|
|
49
49
|
|
|
50
50
|
| Tag | Block type | First-party mapping (declared capability required) |
|
|
51
51
|
| --- | --- | --- |
|
|
52
|
-
| `text` | `text` (default) | All providers |
|
|
53
|
-
| `image` | `image` | OpenAI Responses
|
|
54
|
-
| `audio` | `audio` | OpenAI Responses (`input_audio`) |
|
|
55
|
-
| `file` | `file` | OpenAI Responses (`input_file`); Anthropic
|
|
56
|
-
| `document` | `document` | OpenAI Responses (`input_file`); OpenCode Go Anthropic route;
|
|
52
|
+
| `text` | `text` (default) | All first-party providers; Azure/Bedrock/Vertex use their host-selected OpenAI-compatible endpoint/model. |
|
|
53
|
+
| `image` | `image` | OpenAI Responses; Anthropic; Google; Kimi; Z.AI; OpenRouter; OpenCode Go OpenAI route; Alibaba; Ollama; NeuralWatt. Enterprise OpenAI-compatible packages require the host model/endpoint to declare and accept image input. |
|
|
54
|
+
| `audio` | `audio` | OpenAI Responses (`input_audio`) and Google `generateContent` inline data. OpenAI Realtime instead receives `RealtimeSession.sendAudio()` chunks, not an `audio` `ContentBlock`. |
|
|
55
|
+
| `file` | `file` | OpenAI Responses (`input_file`); Anthropic/Kimi/OpenCode Go Anthropic route accept PDF file/document forms; Google maps inline file data. |
|
|
56
|
+
| `document` | `document` | OpenAI Responses (`input_file`); Anthropic/Kimi/OpenCode Go Anthropic route map PDF; Google maps inline document data. |
|
|
57
|
+
|
|
58
|
+
The AI SDK adapter maps declared user text/image/audio/file/document blocks (and assistant text/image/file/document) to AI SDK file parts; `resourceUri` remains host-resolved before `doStream`. Its output `file`, `reasoning-file`, and `source` parts are deliberately rejected as `unsupported_mapping`, not converted to trusted Prism content. Provider capability metadata is the gate—this matrix never upgrades a model that does not declare the matching input tag.
|
|
57
59
|
|
|
58
60
|
## Outputs / response / events
|
|
59
61
|
|
|
@@ -136,6 +138,7 @@ try {
|
|
|
136
138
|
- Local filesystem paths should use trust policies such as `createPathTrustPolicy()` before exposing URIs to loaders.
|
|
137
139
|
- Provider upload/create/delete lifecycles are provider-package-local. `@arnilo/prism-provider-openai` inlines files under 4 MiB as `data:<mediaType>;base64,...` `file_data`, otherwise uses a bounded per-run upload cache and best-effort `DELETE /v1/files` cleanup after each stream.
|
|
138
140
|
- Shared wire helpers live in `@arnilo/prism/providers/media` (`resolveProviderMediaMessages`, `serializeOpenAIResponsesInputFile`, `serializePdfDocumentWireBlock`, `createBoundedUploadCache`). OpenAI Responses, Kimi, and OpenCode Go Anthropic routes resolve their complete media collection once before serialization or upload.
|
|
141
|
+
- OpenAI Realtime audio is a bidirectional `RealtimeSession` stream, not a `ContentBlock`: provide host-captured `Uint8Array` chunks with `sendAudio()` and consume untrusted `audio_delta` / transcript events. It has a fixed 256 events/s, 1 MiB/s, and 600 s default ceiling.
|
|
139
142
|
|
|
140
143
|
## Security and performance notes
|
|
141
144
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
The optional `@arnilo/prism/node/session-store-jsonl` subpath stores `SessionEntry` records in a caller-named JSONL file: one JSON object per line.
|
|
5
|
+
The optional `@arnilo/prism/node/session-store-jsonl` subpath stores `SessionEntry` records in a caller-named JSONL file: one JSON object per line. `searchSessions` is unsupported and throws `SessionSearchUnsupportedError` (use memory linear mode or a DB adapter for search).
|
|
6
6
|
|
|
7
7
|
APIs:
|
|
8
8
|
|
package/docs/observability.md
CHANGED
|
@@ -167,12 +167,14 @@ console.log(traceId, memory.spans.map((span) => span.name));
|
|
|
167
167
|
## Security and performance notes
|
|
168
168
|
|
|
169
169
|
- Default events are metadata-only — no prompts, streamed deltas, tool arguments, or credentials.
|
|
170
|
+
- Use `identityTelemetryAttributes(identity)` when attaching enterprise identity to run metadata or OTel attributes; it emits `prism.identity.*` refs only (tenant/principal/scope counts), never credential secrets or raw tokens.
|
|
170
171
|
- Opt-in content in other event types (`message_delta`, tool `result`) is still subject to `redactAgentEvent`.
|
|
171
172
|
- Metric labels stay low-cardinality (`gen_ai.operation.name`, `gen_ai.provider.name`, token type, controlled outcome/status, feedback rating bucket/link presence); never use session/run/request/call IDs, model output, comments, tag values, scorer/evaluation IDs, or arbitrary metadata as labels. Token usage is recorded once at provider operation scope.
|
|
172
173
|
- Target overhead when enabled is under 5% excluding exporter I/O; disabled hooks allocate no spans.
|
|
173
174
|
- Provider transport limits and redaction order are documented in [Provider primitives](provider-primitives.md).
|
|
174
175
|
|
|
175
176
|
## Related APIs
|
|
177
|
+
- [Agent identity](agent-identity.md): redacted identity attribute helper for telemetry.
|
|
176
178
|
- [Evaluations](evaluations.md): optional scorers can link scores to run/session/trace IDs from agent events.
|
|
177
179
|
|
|
178
180
|
- [Agent events](agent-events.md): full `AgentEvent` union and subscriber semantics.
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# OpenAPI tools adapter (`@arnilo/prism-openapi-tools`)
|
|
2
|
+
|
|
3
|
+
Optional `createOpenApiTools` compiles host-selected OpenAPI 3.1 operations into bounded Prism `ToolDefinition`s. Zero dependencies (native fetch + WebCrypto-free); the compile step is pure and separated from the runtime executor.
|
|
4
|
+
|
|
5
|
+
## When to use it
|
|
6
|
+
|
|
7
|
+
Hosts that already expose a JSON API with an OpenAPI 3.1 document and want the agent to call a **fixed, host-chosen subset** of it — never model-driven discovery, never a raw method/path passthrough. For vendor web search/extraction use `@arnilo/prism-web-tools`; for M365/GWS use `@arnilo/prism-work-tools`; this adapter is for arbitrary host APIs.
|
|
8
|
+
|
|
9
|
+
## Usage
|
|
10
|
+
|
|
11
|
+
```ts
|
|
12
|
+
import { createOpenApiTools } from "@arnilo/prism-openapi-tools";
|
|
13
|
+
|
|
14
|
+
const tools = createOpenApiTools({
|
|
15
|
+
document, // OpenAPI 3.1 document (JSON string or parsed object)
|
|
16
|
+
operations: ["getCustomer", "createCase"], // only these operationIds compile
|
|
17
|
+
server: "https://api.example.com/v1", // pinned base URL
|
|
18
|
+
credentials: async ({ operationId }) => ({
|
|
19
|
+
headers: { authorization: `Bearer ${await hostToken(operationId)}` },
|
|
20
|
+
}),
|
|
21
|
+
policy: ({ operationId, args }, context) => {
|
|
22
|
+
if (operationId === "createCase" && !context.identity) throw new Error("identity required");
|
|
23
|
+
},
|
|
24
|
+
redactor, // applied to response text before it enters the tool result
|
|
25
|
+
pagination: { pageParam: "page", pageSizeParam: "limit", pageSize: 50, nextPath: "next", itemsPath: "items" },
|
|
26
|
+
idempotencyKeyHeader: true, // forward the core Idempotency-Key on mutating requests
|
|
27
|
+
});
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
Register the returned tools with `createToolRegistry` (or pass them to the MCP bridge); validate arguments with `createJsonSchemaToolArgumentValidator` as usual.
|
|
31
|
+
|
|
32
|
+
## Compile-time guarantees
|
|
33
|
+
|
|
34
|
+
- **Allow-list only**: operations not listed never compile; unknown ids throw `ERR_PRISM_OPENAPI_OPERATION_UNKNOWN`. No generic arbitrary-request escape hatch.
|
|
35
|
+
- **Origin pinned**: the `server` option is the only authority. Every document/path/operation `servers` entry must resolve to the pinned origin (relative URLs resolve against it), else `ERR_PRISM_OPENAPI_SERVER_DRIFT`. https only; http allowed only for loopback hosts.
|
|
36
|
+
- **Bounded schemas**: internal `$ref`s resolve into self-contained argument schemas (path + query + header + body in one object, `additionalProperties: false`). Cycles, depth (`maxSchemaDepth`), ref count (`maxRefs`), external refs, cookie parameters, non-JSON request bodies, and duplicate argument names all fail closed (`ERR_PRISM_OPENAPI_SCHEMA_BOUNDS`).
|
|
37
|
+
- **Effects**: GET/HEAD/OPTIONS/TRACE compile as `{ kind: "none", idempotency: "none" }`; POST/PUT/PATCH/DELETE as `{ kind: "external_mutation", idempotency: "required" }` — the core run loop gates approval ("Tool side effect requires approval") and deduplicates via the `ToolEffectStore` (mutating tools require a verified identity and an effect store at dispatch; retries never re-execute a completed effect). Optional `idempotencyKeyHeader: true` also forwards the core key as `Idempotency-Key` for APIs that honor it.
|
|
38
|
+
|
|
39
|
+
## Runtime guarantees
|
|
40
|
+
|
|
41
|
+
- Request body and response bounded (`maxBodyBytes`, `maxResponseBytes`); oversized responses fail closed (`ERR_PRISM_OPENAPI_RESPONSE_BOUNDS`).
|
|
42
|
+
- Retries only on transport errors and 5xx, bounded by `maxRetries` (default 0, hard 3); transport failures after exhaustion throw `ERR_PRISM_OPENAPI_RETRY_EXHAUSTED`; 4xx never retried.
|
|
43
|
+
- Optional cursor pagination applies only to operations whose compiled query parameters include `pageParam`; bounded by `maxPages` and `maxPaginationItems`.
|
|
44
|
+
- Credentials come only from the host `credentials` resolver (headers/query merged per call), never from the document or options; the optional `redactor` runs over response text before the result is built, so echoed secrets are stripped.
|
|
45
|
+
- Responses are untrusted data: results carry an "UNTRUSTED EXTERNAL API CONTENT" marker and `metadata.trust: "untrusted_external"`; redirects are never followed (`redirect: "manual"`); caller aborts propagate without retry.
|
|
46
|
+
|
|
47
|
+
## Limits
|
|
48
|
+
|
|
49
|
+
Defaults and hard caps (frozen in `scripts/phase11-freeze-manifest.json`): `maxDocumentBytes` 2 MiB/16 MiB, `maxOperations` 256/1024, `maxSchemaDepth` 32/128, `maxRefs` 1024/8192, `maxBodyBytes` 1 MiB/16 MiB, `maxResponseBytes` 1 MiB/16 MiB, `maxPages` 20/100, `maxPaginationItems` 1000/10000, `maxRetries` 0/3. Invalid limits throw `ERR_PRISM_OPENAPI_DOCUMENT_BOUNDS`.
|
|
50
|
+
|
|
51
|
+
## Related
|
|
52
|
+
|
|
53
|
+
- [Tools](tools.md): registry, dispatch, validation
|
|
54
|
+
- [Recoverable tool effects](tool-effects.md): approval + idempotency contracts
|
|
55
|
+
- [Host security guide](host-security.md): permission, trust, validation checklist
|
|
56
|
+
- Package README: [`@arnilo/prism-openapi-tools`](../packages/prism-openapi-tools/README.md)
|
package/docs/performance.md
CHANGED
|
@@ -6,6 +6,267 @@ Evaluation defaults are finite: 100 trace rows × 20 pages and 4 MiB aggregate t
|
|
|
6
6
|
|
|
7
7
|
This page states Prism runtime limits that keep slow consumers and long sessions from becoming unbounded memory or latency problems.
|
|
8
8
|
|
|
9
|
+
## Release 0.1.0 capacity envelopes (frozen performance contract)
|
|
10
|
+
|
|
11
|
+
`scripts/benchmark-0.1.0.mjs` composes the six phase benchmark scripts
|
|
12
|
+
(0.0.23–0.0.28) into one 0.1.0 capacity envelope; the merged evidence is
|
|
13
|
+
checked in as `scripts/benchmark-0.1.0.json` and re-gated on every `npm test`
|
|
14
|
+
by `scripts/benchmark-0.1.0.test.mjs` against the Task 0 freeze-manifest
|
|
15
|
+
capacity contract (`scripts/phase12-freeze-manifest.json`): a row that drifts
|
|
16
|
+
above its frozen p95 ceiling, a startup import above 250 ms, or a root pack
|
|
17
|
+
row beyond its ±5% diet tolerance fails the gate.
|
|
18
|
+
|
|
19
|
+
**Methodology.** Each leg is the same fixture as the phase benchmark that
|
|
20
|
+
introduced it (warmups and measured operations per leg are recorded in the
|
|
21
|
+
JSON `legs` array): in-process fakes and loopback fixture servers for the
|
|
22
|
+
network-free legs, disposable PostgreSQL 16 schema for the protected legs.
|
|
23
|
+
Measured on Node v24.18.0 / Linux x64 (local hardware; values are environment
|
|
24
|
+
evidence, not universal SLOs). Regenerate with:
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json
|
|
28
|
+
PRISM_TEST_POSTGRES_URL="postgresql://…" node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json # adds protected legs
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
**Pass/fail thresholds.** Network-free rows fail above the frozen ceiling in
|
|
32
|
+
the table below; protected PostgreSQL rows fail above their per-phase
|
|
33
|
+
budgets.json ceilings (50/100 ms per the approved budget contract); startup
|
|
34
|
+
import fails above `startupImportMsCeiling` (250 ms); root packed bytes and
|
|
35
|
+
file count fail above baseline × 1.05. Labels: **network-free** = runs in
|
|
36
|
+
`npm test` evidence, no network; **protected** = requires live PostgreSQL.
|
|
37
|
+
|
|
38
|
+
| Envelope | Recorded p95 ms | Ceiling ms | Source leg | Label |
|
|
39
|
+
| --- | ---: | ---: | --- | --- |
|
|
40
|
+
| oidcVerifyCacheHitMs | 0.151 | 5 | enterprise adapters (0.0.28) | network-free |
|
|
41
|
+
| oidcVerifyCacheMissMs | 0.351 | 100 | enterprise adapters (0.0.28) | network-free |
|
|
42
|
+
| policyDecisionMs | 0.052 | 100 | enterprise adapters (0.0.28) | network-free |
|
|
43
|
+
| mcpDiscoveryRoundTripMs | 1.563 | 250 | enterprise adapters (0.0.28) | network-free |
|
|
44
|
+
| mcpAuthHandshakeMs | 6.909 | 2,000 | enterprise adapters (0.0.28) | network-free |
|
|
45
|
+
| openapiToolCallMs | 0.011 | 1,000 | enterprise adapters (0.0.28) | network-free |
|
|
46
|
+
| artifactPut1MiBMs | 9.337 | 2,000 | enterprise adapters (0.0.28) | network-free |
|
|
47
|
+
| artifactPresignMs | 0.381 | 100 | enterprise adapters (0.0.28) | network-free |
|
|
48
|
+
| decisionApply | 4.484 | 5 | durable loops/HITL (0.0.25) | network-free |
|
|
49
|
+
| stickyMatch | 0.343 | 5 | durable loops/HITL (0.0.25) | network-free |
|
|
50
|
+
| snapshotCaptureRestore | 7.814 | 20 | durable loops/HITL (0.0.25) | network-free |
|
|
51
|
+
| a2uiPaint | 0.330 | 10 | durable loops/HITL (0.0.25) | network-free |
|
|
52
|
+
| enumerationList | 363.610 | 2,000 | coding/process/forge/egress (0.0.26) | network-free |
|
|
53
|
+
| processChunkPage | 0.048 | 10 | coding/process/forge/egress (0.0.26) | network-free |
|
|
54
|
+
| lspDiagnosticNormalize | 0.610 | 100 | coding/process/forge/egress (0.0.26) | network-free |
|
|
55
|
+
| forgePagination | 144.868 | 10,000 | coding/process/forge/egress (0.0.26) | network-free |
|
|
56
|
+
| proxyDownload | 86.879 | 30,000 | coding/process/forge/egress (0.0.26) | network-free |
|
|
57
|
+
| rendererStreamOps | 1.647 | 100 | coding/process/forge/egress (0.0.26) | network-free |
|
|
58
|
+
| agUiMapperSync | 33.059 | 100 | coding/process/forge/egress (0.0.26) | network-free |
|
|
59
|
+
| fsReadWriteRoundTripMs | 0.210 | 250 | ACP (0.0.27) | network-free |
|
|
60
|
+
| modeSwitchMs | 0.185 | 250 | ACP (0.0.27) | network-free |
|
|
61
|
+
| terminalChunkAckMs | 0.049 | 1,000 | ACP (0.0.27) | network-free |
|
|
62
|
+
| promptFirstUpdateMs | 0.082 | 2,000 | ACP (0.0.27) | network-free |
|
|
63
|
+
| promptEndMs | 0.094 | 30,000 | ACP (0.0.27) | network-free |
|
|
64
|
+
| policyAppend | 0.684 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
65
|
+
| policyQuery | 1.274 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
66
|
+
| evaluationAppend | 0.721 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
67
|
+
| evaluationQuery | 0.918 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
68
|
+
| workClaimComplete | 1.938 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
69
|
+
| workContention | 5.002 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
70
|
+
| routerRateContention | 17.582 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
71
|
+
| routerBudgetContention | 6.018 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
72
|
+
| routerCircuitContention | 27.098 | 50 | enterprise PostgreSQL (0.0.23) | protected |
|
|
73
|
+
| cleanupBatch | 3.509 | 100 | enterprise PostgreSQL (0.0.23) | protected |
|
|
74
|
+
| eventAppend | 1.553 | 50 | distributed events (0.0.24) | protected |
|
|
75
|
+
| eventPage | 3.187 | 50 | distributed events (0.0.24) | protected |
|
|
76
|
+
| effectClaimTransition | 2.642 | 50 | distributed events (0.0.24) | protected |
|
|
77
|
+
| eventCleanup | 1.255 | 100 | distributed events (0.0.24) | protected |
|
|
78
|
+
| effectCleanup | 3.174 | 100 | distributed events (0.0.24) | protected |
|
|
79
|
+
| reconnectCatchup | 8.374 | 100 | distributed events (0.0.24) | protected |
|
|
80
|
+
|
|
81
|
+
Install/startup rows (same helpers as the budget gate — no duplicate
|
|
82
|
+
measurement): startup import 41.7 ms (ceiling 250 ms); root packed 711,755
|
|
83
|
+
bytes vs baseline 678,541 (+5%, tolerance 5%); root file count 295 vs 293
|
|
84
|
+
(+5%). Storage-growth rows and query plans from the protected legs are in the
|
|
85
|
+
recorded JSON (`storageBeforeCleanup` / `storageAfterCleanup` per leg).
|
|
86
|
+
|
|
87
|
+
Conformance companions: `scripts/phase8–11-conformance.test.mjs` plus the
|
|
88
|
+
Task 3 packed-install journeys and Task 4 restart-recovery evidence (see
|
|
89
|
+
[`docs/0.1.0-readiness.md`](./0.1.0-readiness.md)).
|
|
90
|
+
|
|
91
|
+
## Release 0.0.28 enterprise auth, policy, MCP OAuth, API, and artifact adapters
|
|
92
|
+
|
|
93
|
+
`node scripts/benchmark-0.0.28.mjs` is network-free (in-process fake JWKS/OPA/API fetches plus loopback fixture servers for the authorization server, Prism MCP server, and S3-compatible object store). Checked `scripts/benchmark-0.0.28.json` (Node v24.18.0/Linux x64): 20 warmups, 100 measured ops per seam.
|
|
94
|
+
|
|
95
|
+
| Scenario | Recorded p95 ms | Ceiling |
|
|
96
|
+
| --- | ---: | ---: |
|
|
97
|
+
| OIDC verify (warm JWKS cache) | 0.178 | 5 |
|
|
98
|
+
| OIDC verify (TTL-expired JWKS refetch) | 0.366 | 100 |
|
|
99
|
+
| OPA policy decision (fake endpoint) | 0.063 | 100 |
|
|
100
|
+
| MCP OAuth discovery round trip | 2.196 | 250 |
|
|
101
|
+
| MCP OAuth interactive handshake (PKCE + token + authorized connect) | 6.359 | 2,000 |
|
|
102
|
+
| OpenAPI tool call (compiled operation, fake API) | 0.023 | 1,000 |
|
|
103
|
+
| Artifact 1 MiB body put (fake object store) | 11.810 | 2,000 |
|
|
104
|
+
| Artifact presign | 0.630 | 100 |
|
|
105
|
+
|
|
106
|
+
Conformance: `scripts/phase11-conformance.test.mjs` (5 network-free cases: composed OIDC → OPA ledger → MCP OAuth tool → OpenAPI side effect → artifact body + signed delivery, adapter-absent baseline, hostile origins and limit ladder, redaction sweep). Values are environment evidence, not universal SLOs.
|
|
107
|
+
|
|
108
|
+
## Release 0.0.25 durable loops and human-in-the-loop
|
|
109
|
+
|
|
110
|
+
`node scripts/benchmark-0.0.25.mjs` is network-free (in-memory checkpoint store). Checked `scripts/benchmark-0.0.25.json` (Node v24.18.0/Linux x64): 20 warmups, 100 measured ops, 32 pending decisions, ~250 KiB snapshot, 64 A2UI ops/message.
|
|
111
|
+
|
|
112
|
+
| Scenario | Recorded p95 ms | Ceiling |
|
|
113
|
+
| --- | ---: | ---: |
|
|
114
|
+
| Decision apply (batch CAS) | 3.913 | 5 |
|
|
115
|
+
| Sticky match | 0.407 | 5 |
|
|
116
|
+
| Snapshot capture/restore | 6.742 | 20 |
|
|
117
|
+
| A2UI paint | 0.348 | 10 |
|
|
118
|
+
|
|
119
|
+
Conformance: `scripts/phase8-conformance.test.mjs` (8 network-free cases). Values are environment evidence, not universal SLOs.
|
|
120
|
+
|
|
121
|
+
## Release 0.0.26 coding intelligence, processes, forge, and egress
|
|
122
|
+
|
|
123
|
+
`node scripts/benchmark-0.0.26.mjs` is network-free (fake LSP/forge/proxy, synthetic 100k-file repo, real process spill). Checked `scripts/benchmark-0.0.26.json` (Node v24.18.0/Linux x64): 5 warmups, 20 measured ops, 100k enumeration files, 1 GiB process spill, 1,000 LSP diagnostics, 100 forge pages × 100 items, 64 MiB proxy download.
|
|
124
|
+
|
|
125
|
+
| Scenario | Recorded p95 ms | Ceiling |
|
|
126
|
+
| --- | ---: | ---: |
|
|
127
|
+
| Git-aware enumeration (100k-file repo, ≤ 2 git invocations, 10k results cap) | 299.166 | 2,000 |
|
|
128
|
+
| Process chunk page (50 KiB pages over 1 GiB spill, 64 MiB retained) | 0.051 | 10 |
|
|
129
|
+
| LSP diagnostic normalization (1,000 diagnostics at hard per-file cap) | 0.210 | 100 |
|
|
130
|
+
| Forge pagination (100 pages × 100 check-runs, deduped) | 144.233 | 10,000 |
|
|
131
|
+
| Proxy download (64 MiB at default response cap, resident buffering ≤ 2× maxBytes) | 93.667 | 30,000 |
|
|
132
|
+
| Renderer stream (1,000-op A2UI surface as 16×64-op batches + full tree render) | 2.000 | 100 |
|
|
133
|
+
| AG-UI mapper sync path (4,000 events through the async pipeline, sync hooks only) | 30.924 | 100 |
|
|
134
|
+
|
|
135
|
+
Conformance: `scripts/phase9-conformance.test.mjs` (8 network-free cases: composed enumeration → LSP rename → process → forge → egress, symlink/ignore escape, LSP URI escape, process ownership, forge cross-tenant + token hygiene, egress private/metadata bypass, limit ladder, packed example). Values are environment evidence, not universal SLOs.
|
|
136
|
+
|
|
137
|
+
## Release 0.0.24 distributed events and tool effects
|
|
138
|
+
|
|
139
|
+
`node scripts/benchmark-0.0.24.mjs` is an explicit protected PostgreSQL benchmark behind `PRISM_TEST_POSTGRES_URL`. Checked `scripts/benchmark-0.0.24.json` (Node v24.18.0/Linux x64, PostgreSQL 16.14): 10 tenants × 10 principals × 1,000 events/owner, 16 producers/subscribers, 100 warmups, 1,000 measured ops, 10,000-event sustained replay, 100-row cleanup.
|
|
140
|
+
|
|
141
|
+
| Scenario | Recorded p95 ms | Ceiling |
|
|
142
|
+
| --- | ---: | ---: |
|
|
143
|
+
| Event append / page | 1.502 / 3.103 | 50 / 100 |
|
|
144
|
+
| Effect claim+transition / cleanup | 3.084 / 3.242 | 50 / 100 |
|
|
145
|
+
| Event cleanup / reconnect catch-up | 1.370 / 7.883 | 100 / 2000 |
|
|
146
|
+
|
|
147
|
+
Sustained replay delivered 160,000 subscriber-events at 101.34 events/s. Five `EXPLAIN (ANALYZE, BUFFERS, FORMAT JSON)` plans used named indexes with no sequential scans. Process conformance (`scripts/phase7-conformance.test.mjs`) covers 16-process producers, `LISTEN` backend kill + poll catch-up, and pending/dispatched effect crash windows. Values are environment evidence, not universal SLOs.
|
|
148
|
+
|
|
149
|
+
## Release 0.0.23 enterprise PostgreSQL evidence
|
|
150
|
+
|
|
151
|
+
`node scripts/benchmark-0.0.23.mjs` is an explicit protected PostgreSQL benchmark, not part of `npm test` or `sdk:ready`. It requires `PRISM_TEST_POSTGRES_URL`, creates/drops an isolated schema, and checks frozen p95 ceilings from `scripts/budgets.json`. The checked `scripts/benchmark-0.0.23.json` evidence was recorded on Node v24.18.0/Linux x64 with `postgres:16-alpine`: 10 tenants × 10 principals × 1,000 policy/evaluation rows, 10,000 router keys, 16 pool clients, 100 warmups, 1,000 measured operations, and 100-row cleanup batches.
|
|
152
|
+
|
|
153
|
+
| Scenario | Recorded p95 ms | Ceiling |
|
|
154
|
+
| --- | ---: | ---: |
|
|
155
|
+
| Policy append / query | 0.747 / 1.479 | 50 / 100 |
|
|
156
|
+
| Evaluation append / query | 0.698 / 0.963 | 50 / 100 |
|
|
157
|
+
| Work claim+complete / contention | 1.892 / 4.162 | 50 / 50 |
|
|
158
|
+
| Router rate / budget / circuit contention | 12.011 / 6.715 / 28.410 | 50 / 50 / 50 |
|
|
159
|
+
| Explicit cleanup batch | 2.981 | 100 |
|
|
160
|
+
|
|
161
|
+
The same run accepted 1,000 rate claims, accumulated 16,000 budget tokens, granted 1,000 circuit probes, and verified 14 named-index `EXPLAIN (ANALYZE, BUFFERS, FORMAT JSON)` plans with no sequential scans. Before cleanup it recorded 101,100 policy rows (68,517,888 bytes), 101,100 evaluation rows (97,296,384 bytes), and 121,100 router-rate rows; cleanup removed exactly 100,000 expired rate rows. PostgreSQL relation allocation does not necessarily shrink after `DELETE` under MVCC, so row removal—not immediate file shrink—is the cleanup assertion. Values are dated environment evidence, not universal production SLOs; size pools, partitions, retention, and cleanup frequency from host measurements.
|
|
162
|
+
|
|
163
|
+
## Release 0.0.16 performance budgets and artifact diet
|
|
164
|
+
|
|
165
|
+
Release 0.0.16 is a simplification/readiness release: it added no performance-affecting code, so the six network-free scenario medians are held at the 0.0.15 baseline and the win is a smaller published artifact. Budgets live in `scripts/budgets.json` (measured baselines + tolerance) and are enforced two ways:
|
|
166
|
+
|
|
167
|
+
- **Fast gate (every `npm test`)** — `scripts/budget-gate.test.mjs` re-packs the root tarball (`npm pack --dry-run --json`) and fails if packed bytes, unpacked bytes, or file count exceed baseline + 5%, and fails if cold-process `import('./dist/index.js')` exceeds the 250 ms sanity ceiling. Negative fixtures prove an inflated/regressed value fails.
|
|
168
|
+
- **Release evidence runner** — `node scripts/benchmark-0.0.16.mjs` re-measures root pack + startup, spawns `benchmark-0.0.15.mjs` for the six scenario medians (reused unchanged), compares every value to `budgets.json` (throughput floor / latency ceiling at ±25%), prints the evidence report below, and exits non-zero on any regression.
|
|
169
|
+
|
|
170
|
+
**Artifact diet (the 0.0.16 finding).** The Task 1 tarball deny list dropped the historical `docs/review-coverage-*.md` (11 files, 283,022 bytes) from the root package: the root tarball went from **659,478 packed / 2,310,686 unpacked / 281 files** (0.0.15) to a budgeted **≈575,680 packed / 2,043,402 unpacked / 270 files**. The per-release `scripts/benchmark-0.0.*.mjs` history never shipped in artifacts (root `files` is `dist`/`docs`/`templates`/`CHANGELOG.md` only — zero `scripts/` entries packed), so no archive move was needed; `benchmark-0.0.16.mjs` consolidates the current evidence behind one budget-gating runner.
|
|
171
|
+
|
|
172
|
+
**Recorded budgets (`scripts/budgets.json`, measured 2026-07-26, Node v24.18.0, Linux x86_64):**
|
|
173
|
+
|
|
174
|
+
| Budget | Baseline | Tolerance |
|
|
175
|
+
| --- | --- | --- |
|
|
176
|
+
| Root packed bytes | 575,680 | +5% |
|
|
177
|
+
| Root unpacked bytes | 2,043,402 | +5% |
|
|
178
|
+
| Root file count | 270 | +5% |
|
|
179
|
+
| Aggregate packed bytes (44 manifests, reference only) | 1,217,694 | +10% |
|
|
180
|
+
| Startup `import('./dist/index.js')` | ~38 ms | ceiling 250 ms |
|
|
181
|
+
| Six scenario medians (below) | 0.0.15 baseline | ±25% |
|
|
182
|
+
|
|
183
|
+
**0.0.16 measured evidence** (`node scripts/benchmark-0.0.16.mjs`, 100 iterations each, network-free, 0 backpressure / 0 resource-limit signals; all 22 budget checks passed):
|
|
184
|
+
|
|
185
|
+
| Scenario | throughput/s | p50 ms | p95 ms |
|
|
186
|
+
| --- | --- | --- | --- |
|
|
187
|
+
| openai-hosted-continuation | 5,514.9 | 0.1305 | 0.2735 |
|
|
188
|
+
| openai-realtime-envelope | 900.5 | 1.1277 | 1.2002 |
|
|
189
|
+
| ai-sdk-v4-stream-mapping | 23,850.4 | 0.0225 | 0.0795 |
|
|
190
|
+
| provider-package-metadata | 54,097.0 | 0.0066 | 0.0386 |
|
|
191
|
+
| rag-parse-replace-rerank-retrieve | 5,176.6 | 0.1428 | 0.3671 |
|
|
192
|
+
| memory-retention-export-rebuild | 13,952.1 | 0.0470 | 0.1339 |
|
|
193
|
+
|
|
194
|
+
Root startup measured ≈37.7 ms (ceiling 250 ms). Timing is machine-dependent, so medians carry a wide ±25% band and are release evidence rather than tight cross-machine guarantees; the deterministic artifact-size gate is the hard CI tripwire. Raise the baselines in `scripts/budgets.json` after a deliberate, reviewed performance change.
|
|
195
|
+
|
|
196
|
+
## Release 0.0.15 provider, RAG, and memory evidence
|
|
197
|
+
|
|
198
|
+
Run `node scripts/benchmark-0.0.15.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.15.test.mjs`. Default mode is network-free: fake Responses SSE/WebSocket transports, a fake AI SDK v4 model, zero-fetch provider-package registration, hash embeddings, in-memory RAG replacement/reranking/retrieval/status, and in-memory memory retention/export/rebuild.
|
|
199
|
+
|
|
200
|
+
Scenarios: `openai-hosted-continuation`, `openai-realtime-envelope`, `ai-sdk-v4-stream-mapping`, `provider-package-metadata`, `rag-parse-replace-rerank-retrieve`, and `memory-retention-export-rebuild`.
|
|
201
|
+
|
|
202
|
+
Every row reports throughput, p50/p95 latency, heap, disk, queue/backpressure, and safety signals. `resourceLimitSignals` must be zero: hosted calls remain provider-owned, continuation stops after its finite path, Realtime credentials are absent from events, provider setup does not resolve credentials, retrieved RAG content stays inert, and memory export redacts the fixture secret. These are behavior/bound gates; host-local timings are comparison evidence, not portable release thresholds.
|
|
203
|
+
|
|
204
|
+
| Resource | Default / hard |
|
|
205
|
+
| --- | --- |
|
|
206
|
+
| OpenAI continuation hops | 8 |
|
|
207
|
+
| Realtime audio events / bytes per second | 64 / 256 · 1 MiB / 8 MiB |
|
|
208
|
+
| RAG document bytes | 1 MiB / 8 MiB |
|
|
209
|
+
| RAG rerank input / time / active calls | 64 KiB / 256 KiB · 2 s / 10 s · 2 / 8 |
|
|
210
|
+
| RAG ingestion-status page | 50 / 200 |
|
|
211
|
+
| Memory retention batch | 500 / 5,000 |
|
|
212
|
+
| Memory export | 100 / 200 entries · 4 MiB / 32 MiB · 10 s / 60 s |
|
|
213
|
+
| Memory rebuild | 32 / 128 entries · 10 s / 60 s |
|
|
214
|
+
|
|
215
|
+
This task adds no package or runtime dependency: package/install delta is zero and the frozen graph remains 43 publishable manifests. Credentialed protocol checks are documented in the [0.0.15 protected live-canary matrix](release-and-install.md#015-protected-live-canary-matrix); they never run in this benchmark, `npm test`, or `sdk:ready`.
|
|
216
|
+
|
|
217
|
+
2026-07-26 baseline: Node v24.18.0, Linux x64, 100 iterations/scenario, network=false, credentials=false.
|
|
218
|
+
|
|
219
|
+
| Scenario | ops/s | p95 ms | heap bytes | backpressure | resource limits |
|
|
220
|
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
221
|
+
| OpenAI hosted continuation | 4,885 | 0.2924 | 15,475,920 | 0 | 0 |
|
|
222
|
+
| OpenAI Realtime envelope | 858 | 1.3130 | 12,092,232 | 0 | 0 |
|
|
223
|
+
| AI SDK v4 mapping | 27,580 | 0.0573 | 14,848,288 | 0 | 0 |
|
|
224
|
+
| Provider package metadata | 79,823 | 0.0288 | 16,077,080 | 0 | 0 |
|
|
225
|
+
| RAG lifecycle/reranking | 5,324 | 0.3586 | 13,894,576 | 0 | 0 |
|
|
226
|
+
| Memory lifecycle | 12,892 | 0.1608 | 13,923,936 | 0 | 0 |
|
|
227
|
+
|
|
228
|
+
These values are dated local comparison evidence, not portable thresholds.
|
|
229
|
+
|
|
230
|
+
## Release 0.0.12 frontend interoperability caps and evidence
|
|
231
|
+
|
|
232
|
+
`@arnilo/prism-ag-ui` uses finite handler/projection limits, all defaults / hard: request 64 KiB / 1 MiB; input 128 / 1024 messages and 64 KiB / 1 MiB text; event 64 KiB / 1 MiB; error 8 KiB / 64 KiB; replay cursor 4 / 16 KiB; replay page 100 / 500; subscriber queue 128 / 4096; stream 10,000 / 100,000 events and 10 / 64 MiB; request wall time 120 seconds / 30 minutes. Tool arguments/results/progress, frontend tools, and mutable frontend state default to zero exposure; hosts may only add bounded safe projection.
|
|
233
|
+
|
|
234
|
+
Reconnect is one ownership-scoped redacted durable page plus an optional bounded live subscriber. It is at-least-once at a page boundary, never a polling loop or terminal-run rerun. ACP uses the same event/byte/queue caps. Coding compaction reuses LLM summary/reserve/error/file-operation bounds (16,384 / 131,072 summary and reserve tokens; 1 / 8 KiB summary errors) and makes no additional provider call.
|
|
235
|
+
|
|
236
|
+
Run `node scripts/benchmark-0.0.12.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.12.test.mjs`. Default mode is network-free and reports mapper/handler/replay throughput and p50/p95, peak emitted queue rows, event bytes, heap, and coding-preparation overhead. Bounds and hostile-input fixtures—not these host-local timings—are release gates.
|
|
237
|
+
|
|
238
|
+
2026-07-22 baseline: Node v24.18.0, Linux x64, 100 iterations/scenario, network=false, credentials=false.
|
|
239
|
+
|
|
240
|
+
| Scenario | mode | ops/s | p95 ms | heap bytes | peak queue events | event bytes | cost USD | backpressure | resource limits |
|
|
241
|
+
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
|
242
|
+
| AG-UI mapper | in-process | 23,561 | 0.0398 | 12,029,512 | 2 | 166 | 0 | 0 | 0 |
|
|
243
|
+
| AG-UI handler | web-in-process | 1,401 | 2.2651 | 19,545,416 | 5 | 508 | 0 | 0 | 0 |
|
|
244
|
+
| AG-UI replay | memory-page | 6,094 | 0.3047 | 19,098,240 | 2 | 243 | 0 | 0 | 0 |
|
|
245
|
+
| Coding compaction preparation | in-process | 75,515 | 0.0291 | 20,585,408 | 1 | 208 | 0 | 0 | 0 |
|
|
246
|
+
|
|
247
|
+
No network, credentials, provider summary call, durable database, or live subscriber is involved. These values are dated local comparison evidence, not portable thresholds.
|
|
248
|
+
|
|
249
|
+
## Release 0.0.11 session search / context budget / steer caps
|
|
250
|
+
|
|
251
|
+
Finite caps (defaults / hard) — full matrix in [Phase 6 evidence](review-coverage-2026-07-22-phase-6.md):
|
|
252
|
+
|
|
253
|
+
| Resource | Default / hard |
|
|
254
|
+
| --- | --- |
|
|
255
|
+
| Session search page | 20 / 100 |
|
|
256
|
+
| Search query string | 4 KiB / 16 KiB |
|
|
257
|
+
| Search snippet | 512 B / 4 KiB |
|
|
258
|
+
| Memory linear sessions / entries / bytes | 1000/5000 · 10000/50000 · 8 MiB/64 MiB |
|
|
259
|
+
| FTS candidates | 1000 / 5000 |
|
|
260
|
+
| Context budget tokens / bytes | caller-set / hard 2_000_000 tokens · 32 MiB |
|
|
261
|
+
| Context omission rows | 256 / 1024 |
|
|
262
|
+
| Pending steers | 8 messages / 64 KiB |
|
|
263
|
+
|
|
264
|
+
Run `node scripts/benchmark-0.0.11.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.11.test.mjs`. Default mode is network-free: memory-linear `searchSessions` (label + query) plus assembler `contextBudget` eviction/fit. Emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals. Search never default-scans an unbounded store; budget fails closed on mandatory prefix overflow; steer overflow fails closed. These are evidence fields, not CI timing gates.
|
|
265
|
+
|
|
266
|
+
## Release 0.0.10 reproducible workspace-mode evidence
|
|
267
|
+
|
|
268
|
+
Run `node scripts/benchmark-0.0.10.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.10.test.mjs`. Default mode is network-free: host-composition write/read/list plus sandbox-fake composition write/read/list/search (in-memory `DisposableSandbox`). Emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals. Optional `PRISM_BENCH_DOCKER=1` (with `PRISM_TEST_DOCKER_*`) appends real local Docker composition rows. Unified workspace mode reuses existing sandbox/repo hard caps and adds no unbounded host↔container sync. These are evidence fields, not CI timing gates.
|
|
269
|
+
|
|
9
270
|
## Release 0.0.9 reproducible coding/browser evidence
|
|
10
271
|
|
|
11
272
|
Run `node scripts/benchmark-0.0.9.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.9.test.mjs`. Default mode is network-free fake/in-process only and emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals for repository list/search, Git status, and browser open/snapshot/action/close. Optional `PRISM_BENCH_DOCKER=1` (with `PRISM_TEST_DOCKER_*`) and `PRISM_BENCH_PLAYWRIGHT=1` append real local Docker / protected Playwright rows. These are evidence fields, not CI timing gates.
|
|
@@ -54,6 +315,8 @@ Durable coding plan/checkpoint defaults/hard caps: plan Markdown 256 KiB/1 MiB;
|
|
|
54
315
|
|
|
55
316
|
Browser automation defaults/hard caps from `@arnilo/prism-browser`: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2,000/10,000; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30 s/120 s; action 10 s/60 s; wait 30 s/120 s; run wall 20 min/30 min; popups 4/16; dialogs 16/64; listeners 64/256; action input 64 KiB/256 KiB; close grace 5 s/30 s; network requests 1,000/10,000 with 10/32 redirects per request and 8/32 WebSockets; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate. Caps charge before context/page/action/queue/snapshot/network/artifact retention. Host supplies Playwright and egress proxy attestation; package import launches nothing.
|
|
56
317
|
|
|
318
|
+
0.0.14 co-work defaults/hard caps (frozen in [Phase 9 evidence](review-coverage-2026-07-25-phase-9.md)): conversation thread list pages 50/200, active branches per thread 16/64, replay/export page 100/500 events; artifact revisions per artifact 32/128, artifacts per thread 64/256, metadata record 8/64 KiB, preview 16/64 KiB, citations 32/128 (2/8 KiB each), delivery-link TTL 5 min/24 h, delivery token 4/16 KiB, compare exactly 2 revisions; memory retention batch 500/5000; proactive capability TTL 24 h/31 d, capability token record 16 KiB; browser checkpoint URL 8 KiB/16 KiB, domain-state hash 256 B/1 KiB, host-data ref 2 KiB/8 KiB, 16/64 checkpoints per run; device stream chunk 1 MiB/8 MiB, concurrent device sessions per identity 1/4 (device wall/turns/tool calls consume shared `RunLimits`). All caps charge before persist/emit and fail closed on overflow. Benchmark placeholder: `node scripts/benchmark-0.0.14.mjs` (release Task 12) reports conversation replay, memory injection/consent, artifact revision/delivery, AG-UI co-work mapping, and connector refresh overhead against these budgets.
|
|
319
|
+
|
|
57
320
|
Current surfaces:
|
|
58
321
|
|
|
59
322
|
- `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
|
|
@@ -415,6 +678,25 @@ Timings are one local Node v24.18.0 run over mock agents and an in-process fetch
|
|
|
415
678
|
|
|
416
679
|
No performance ceiling was raised. Core grew from Phase 0's 346.0 kB packed baseline to ~403.7 kB after documented APIs/templates, while the full package set remains ~690.6 kB packed. Follow-up review includes all six Phase 4-13 capability packages through `prism-all` and AI SDK interoperability through `prism-providers`; focused base/code/SDK profiles remain unchanged and no capability auto-activates. Manifest tarballs remain tiny: providers 1.4 kB and all 1.6 kB packed.
|
|
417
680
|
|
|
681
|
+
### 0.0.13 Phase 8 server deployment seams (2026-07-23)
|
|
682
|
+
|
|
683
|
+
Optional health/drain/rate-limit/replay/deployment-lease helpers on `@arnilo/prism-server`. No listener, queue adapter, or concurrency hard-cap raise.
|
|
684
|
+
|
|
685
|
+
| Surface | Result |
|
|
686
|
+
| --- | --- |
|
|
687
|
+
| Focused server suite | existing handler tests + 4 deployment seam tests pass |
|
|
688
|
+
| Health body | default 4 KiB / hard 64 KiB; detail requires authorize |
|
|
689
|
+
| Drain admit cutoff | default 30 s / hard 5 min; admits reject immediately on `beginDrain` |
|
|
690
|
+
| Replay page / cursor | 100 / 4 KiB default; 500 / 16 KiB hard |
|
|
691
|
+
| Concurrent runs | unchanged 16 / 256 |
|
|
692
|
+
| Queues | absent; use `createWorkflowCoordinator` polling until measured need |
|
|
693
|
+
|
|
694
|
+
### 0.0.13 Phase 8 identity, policy, router, and work connectors (2026-07-24)
|
|
695
|
+
|
|
696
|
+
Enterprise governance and connector caps (defaults / hard). Timings: `node scripts/benchmark-0.0.13.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.13.test.mjs`. Default mode is network-free and reports identity/policy/router/work-connector/deployment throughput and p50/p95 with frozen budget refs in the report JSON. Bounds and hostile-input fixtures—not these host-local timings—are release gates.
|
|
697
|
+
|
|
698
|
+
Offline behavior tests (identity propagation, policy export, router deny paths, fake CLI argv) are release gates; live tenant canaries remain operator-gated.
|
|
699
|
+
|
|
418
700
|
## Related APIs
|
|
419
701
|
|
|
420
702
|
- [Agent events](agent-events.md): `SubscribeOptions` and `event_subscriber_overflow` event details.
|