@arnilo/prism 0.0.7 → 0.0.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +21 -0
- package/README.md +3 -1
- package/dist/agent-loops.js +7 -5
- package/dist/agents.js +21 -2
- package/dist/contracts.d.ts +15 -0
- package/dist/index.d.ts +3 -1
- package/dist/index.js +2 -1
- package/dist/run-ledger.d.ts +21 -0
- package/dist/run-ledger.js +115 -0
- package/docs/a2a.md +61 -42
- package/docs/agent-events.md +2 -2
- package/docs/agent-loops.md +1 -1
- package/docs/agent-session-runtime.md +1 -0
- package/docs/credential-storage.md +9 -0
- package/docs/database-persistence.md +1 -1
- package/docs/evaluations.md +26 -3
- package/docs/guardrails.md +1 -1
- package/docs/host-security.md +25 -2
- package/docs/index.md +14 -12
- package/docs/mcp-tools.md +29 -5
- package/docs/migration.md +36 -0
- package/docs/observability.md +26 -14
- package/docs/performance.md +25 -0
- package/docs/postgres-persistence.md +1 -0
- package/docs/providers/kimi.md +16 -2
- package/docs/providers/opencode-go.md +43 -2
- package/docs/release-and-install.md +73 -62
- package/docs/resource-loading.md +4 -0
- package/docs/review-coverage-2026-07-19-phase-3.md +174 -0
- package/docs/run-ledger-conformance.md +1 -0
- package/docs/runs-and-usage.md +17 -2
- package/docs/sqlite-persistence.md +1 -0
- package/docs/supervisors.md +2 -2
- package/docs/tools.md +2 -1
- package/docs/web-tools.md +78 -0
- package/docs/workflows.md +1 -0
- package/package.json +2 -1
package/docs/evaluations.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, live post-run scoring, and
|
|
5
|
+
`@arnilo/prism-evals` adds optional deterministic scorers, immutable datasets, bounded persistence-trace grading, explicit host model judges, pairwise comparisons, CI thresholds, live post-run scoring, and batch experiments over `AgentRunResult`. Scores are finite numbers in `[0, 1]` with optional reason/metadata and linkage to run/session/trace/experiment IDs.
|
|
6
6
|
|
|
7
7
|
## When to use it
|
|
8
8
|
|
|
@@ -18,6 +18,10 @@ Use this package when a host needs offline quality checks or sampled live scorin
|
|
|
18
18
|
| `runExperiment` | `agent`, dataset, scorers, bounded `concurrency`, optional store/ownership |
|
|
19
19
|
| `createMemoryEvaluationStore` | optional seed records |
|
|
20
20
|
| `appendEvaluationFeedback` | `RunFeedbackStore`, `EvaluationStore`, feedback fields, and 1–64 known evaluation IDs |
|
|
21
|
+
| `createPersistenceTraceResolver` | explicit `ProductionPersistenceStore`, exact session/run/ownership, page/byte bounds |
|
|
22
|
+
| `createModelJudge` | host judge callback, stable rubric/version, timeout/attempt/output bounds |
|
|
23
|
+
| `runComparison` | immutable dataset, 2–8 named candidates by default, pairwise scorers |
|
|
24
|
+
| `assertEvaluationThreshold` / `serializeEvaluationReport` | mean/failure/per-scorer gates and bounded redacted JSON |
|
|
21
25
|
|
|
22
26
|
## Outputs / response / events
|
|
23
27
|
|
|
@@ -112,11 +116,30 @@ console.log(report.aggregate.meanScore, linked.evaluationIds);
|
|
|
112
116
|
- Scorers receive result/item data only. Credentials, tools, and workspace access are not provided unless the host deliberately closes over them.
|
|
113
117
|
- Records pass through `SecretRedactor` / `secrets` before store append.
|
|
114
118
|
- Queries filter by ownership scope. Feedback linkage additionally requires tenant plus account/user and the feedback store re-verifies the run.
|
|
115
|
-
- Experiment concurrency defaults to `1` and is capped at `32`.
|
|
119
|
+
- Experiment concurrency defaults to `1` and is capped at `32`. Datasets cap at 10,000 items.
|
|
120
|
+
- Trace reads default to 100 rows × 20 pages with a 4 MiB aggregate cap (hard: 1,000 × 100 and 32 MiB). Repeated/missing cursors, identity drift, ownership drift, and overflow fail closed before scoring.
|
|
121
|
+
- Model judges are host callbacks, not providers: Prism passes rubric/version plus bounded target only—never credential resolvers, tools, or workspace. Defaults are one attempt, 30 seconds, and 16 KiB output; failures become redacted evaluation records.
|
|
122
|
+
- Pairwise candidates are sorted by name, executed once per item, compared in stable item/pair/scorer order, and record ties/failures without choosing a winner. Candidate and scorer outputs have byte caps.
|
|
123
|
+
- `assertEvaluationThreshold()` throws `ERR_PRISM_EVAL_THRESHOLD`; an uncaught error gives CI a non-zero exit. Keep model-judge/live gates credential-gated and outside the network-free default suite. `serializeEvaluationReport()` bounds/redacts checked-in artifacts.
|
|
124
|
+
|
|
125
|
+
## Trace, judge, comparison, and CI example
|
|
126
|
+
|
|
127
|
+
```ts
|
|
128
|
+
const traceResolver = createPersistenceTraceResolver(persistence);
|
|
129
|
+
const judge = createModelJudge({
|
|
130
|
+
id: "quality", rubric: "Score factual quality from 0 to 1", rubricVersion: "2026-07-20",
|
|
131
|
+
judge: hostStructuredJudge,
|
|
132
|
+
});
|
|
133
|
+
const evaluations = await scoreRun({ result, scorers: [judge], traceResolver, ownership });
|
|
134
|
+
const comparison = await runComparison({ dataset, candidates: { baseline, candidate }, scorers: [preference] });
|
|
135
|
+
assertEvaluationThreshold(report, { minimumMean: 0.9, maximumFailures: 0 });
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
`traceResolver` is explicit; no arbitrary run search occurs. `baseline`/`candidate` are host functions returning `AgentRunResult`. See `examples/evaluation-gate.ts` for a network-free gate.
|
|
116
139
|
|
|
117
140
|
## Related APIs
|
|
118
141
|
|
|
119
142
|
- [Agent/session runtime](agent-session-runtime.md): `AgentRunResult` and `session.run()`
|
|
120
143
|
- [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
|
|
121
|
-
- [Observability](observability.md):
|
|
144
|
+
- [Observability](observability.md): use `onTraceReference` or bounded `traceId(runId)` to supply `ScoreRunOptions.traceId`; evaluation telemetry emits no reason/explanation content
|
|
122
145
|
- [Release and install](release-and-install.md): optional package install
|
package/docs/guardrails.md
CHANGED
|
@@ -30,7 +30,7 @@ Decisions are `allow`, `block`, `tripwire`, or `interrupt`. Evaluation defaults
|
|
|
30
30
|
|
|
31
31
|
## Outputs / response / events
|
|
32
32
|
|
|
33
|
-
Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
|
|
33
|
+
Every evaluated guard produces a redacted `guardrail_decision` `AgentEvent` with a bounded `GuardrailRecord`. Optional OpenTelemetry instrumentation records only controlled stage/action on a short run-child span; guardrail name, reason, and metadata are excluded. An input or output terminal decision rejects the run with `GuardrailError`; `tripwire` stops remaining evaluation. A tool-input or tool-output `block` returns a redacted blocked `ToolResult`; a `tripwire` rejects the enclosing run. `interrupt` is reserved for durable runs and currently fails closed with `ERR_PRISM_GUARDRAIL_INTERRUPT_UNAVAILABLE`.
|
|
34
34
|
|
|
35
35
|
Ordering is fixed:
|
|
36
36
|
|
package/docs/host-security.md
CHANGED
|
@@ -30,6 +30,7 @@ Start from explicit host inputs. Do not let runtime code discover security state
|
|
|
30
30
|
| Remote media policy | public/default pinned DNS or explicit trusted transport | `SsrfPolicy`, `resolveMediaContentBlock()` |
|
|
31
31
|
| Durable history | host database adapter | `SessionStore`, `assertSessionStoreConforms()` |
|
|
32
32
|
| Durable audit | host ledger adapter | `RunLedger`, `redactRunLedgerRecord()` |
|
|
33
|
+
| Telemetry | host OpenTelemetry SDK/exporter | metadata-only adapter, controlled metric labels, `onTraceReference` |
|
|
33
34
|
| Durable interruption | host checkpoint + session stores, exact ownership | `RunOptions.runState`, `resumeAgentRun()`, `createAgentRunLifecycle()`, `createSecureAgent()` |
|
|
34
35
|
| Extensions | explicit package imports only | `createExtensionKernel`, `ExtensionAPI` |
|
|
35
36
|
| Remote agent/workflow API | host authentication + ownership mapping | `@arnilo/prism-server`, `createPrismHandler()` |
|
|
@@ -120,6 +121,8 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
120
121
|
- Use `createContributionRegistries({ duplicate: "error" })` and prefixed names for third-party packages to prevent silent shadowing.
|
|
121
122
|
- Extension contributions are inert until selected. Loading an extension package runs its `setup(api)` code, so hosts should load only trusted packages or isolate untrusted code outside Prism.
|
|
122
123
|
- Skills and instruction injectors grant no tools, permissions, validators, or resource access. Host-active tools and permission policies still decide execution.
|
|
124
|
+
- Optional ledger batching accepts runtime-redacted records only. Prefer `flush_on_terminal`; `buffered` explicitly permits crash-before-flush loss. Flush failures propagate; hosts can call `dispose({ flush: false })` to clear queued objects when deliberately discarding an aborted buffered workload.
|
|
125
|
+
- Session snapshot cache holds one session-local leaf for at most one second and invalidates after committed mutation, checkout, compaction, and resume; it never crosses session/branch/ownership.
|
|
123
126
|
- For production persistence, implement a database-backed `SessionStore`/`RunLedger`, run `assertSessionStoreConforms()` against the store, and follow the database schema guidance. Do not ship provider instances, credential resolvers, or secrets into durable rows.
|
|
124
127
|
|
|
125
128
|
## Security and performance notes
|
|
@@ -130,8 +133,9 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
130
133
|
- Known secrets must be passed into redactors before data is emitted or persisted. Redact again in host adapters if they transform records after Prism redaction.
|
|
131
134
|
- Tool `parameters` metadata is not validated by default. Add a `ToolValidator`, use `createToolParameterValidator()` with a schema adapter, or install `@arnilo/prism-tool-validator-json-schema` before side effects. Its untrusted-schema adapter rejects non-local refs, forbidden keys/cycles/non-finite values and bounds bytes/depth/properties/keywords/refs plus its LRU cache before Ajv compilation; do not raise caps above documented hard limits.
|
|
132
135
|
- Treat embeddings as untrusted numeric input. `@arnilo/prism-memory` rejects empty, non-number, NaN, and infinite vectors before in-memory similarity or pgvector parameters; custom `Embedder`/`VectorStore` implementations must retain the same boundary.
|
|
136
|
+
- Evaluation trace readers require exact supplied ownership plus session/run identity, reject cursor/identity drift, and redact before bounded scorer/judge input. Model-judge callbacks receive no credential resolver, tools, or workspace; keep live judges outside default CI and redact report artifacts.
|
|
133
137
|
- Prism-generated session/run/tool/workflow/evaluation IDs use Node cryptographic UUIDs. Keep host-provided IDs authorization-scoped and validate them as untrusted identifiers; do not substitute timestamps or `Math.random()` for durable/security-relevant IDs.
|
|
134
|
-
- MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands, requires per-
|
|
138
|
+
- MCP client tools from `@arnilo/prism-mcp` are untrusted remote servers. Stdio remains an explicit host executable. Streamable HTTP requires exact HTTPS origins, rejects credentials/fragments/redirects/private or mixed DNS, pins a validated address on every SDK request/reconnect, and bounds each response; plaintext is explicit loopback-only development mode. Discovery has finite page/tool/cursor/metadata/schema totals and commits atomically. Every result branch shares byte/depth/property bounds before core dispatch; supply a known-secret `SecretRedactor`, `PermissionPolicy`, and `ToolValidator` there. MCP server direction exposes only passed tools/commands/resources/prompts, requires per-operation `authorize`, and retains core gates. Sampling, roots, model/credential selection, and elicitation consent stay host-owned; URL elicitation is never opened automatically. Stateful web mode requires host `resolveAuthInfo` plus `resolveIdentity`, exact origin policy, and binds every POST/GET/DELETE/SSE request to one non-secret principal; mismatches return 404. Handler still needs TLS and edge rate limiting. See [MCP client/server exposure](mcp-tools.md).
|
|
135
139
|
- `@arnilo/prism-server` exposes no agent/workflow by default and requires `authorize()` for every matched operation. Derive complete tenant/account/user ownership from validated host identity, never request JSON. Workflow active identity and cancellation compare exact ownership; a tenant-only scope intentionally cannot cancel a checkpoint/run carrying account or user identity. Pass the current explicitly revised workflow definition so recursive hash mismatch fails before abort or durable mutation. Configure exact host/origin allow-lists where needed, wire redaction before execution, retain tool/workflow policy checks, and adapt the Web handler behind host TLS/rate limits. Disconnect abort is default; persistent reconnect/status belongs to durable workflow checkpoints, not an invented in-memory agent result cache.
|
|
136
140
|
- Coding tools from `@arnilo/prism-coding-agent` accept an optional `ExecutionPolicy` checked inside each tool before side effects; shared policy propagation includes `createReadOnlyTools()`. They enforce finite text-scan/image/edit/write/shell limits, a 600-second default shell wall time, and a 64 MiB default total-output ceiling. Successful truncated shell output leaves a host-owned exclusive `0600` temp file; delete `metadata.fullOutputPath` after use. Error/abort/timeout/overflow removes unpublished spills. Custom read/edit/shell backends must honor supplied caps/signals. Use `@arnilo/prism-coding-security` for path roots, command rules, and identity-scoped approval caching. Limits are not containment: Prism provides no OS sandbox unless the host supplies one.
|
|
137
141
|
- `@arnilo/prism-credentials-node` rejects oversized/malformed envelopes and excessive scrypt work before KDF allocation, uses async scrypt, and requires restrictive existing/new Unix vault modes. Keep vault ownership and parent-directory access host-controlled; review before `chmod 600`, never auto-weaken a file policy. Keychain calls use abort-aware native async work with finite timeout/payload caps and sanitized errors. OS prompts, service availability, and whether a native backend promptly honors cancellation remain host/platform boundaries; no plaintext fallback is attempted.
|
|
@@ -160,7 +164,26 @@ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy
|
|
|
160
164
|
- Keep depth, active children, input, turn/tool/token, timeout, and queue ceilings finite; propagate abort through nested calls.
|
|
161
165
|
- Expose A2A only behind per-request authentication/authorization, TLS, edge rate limits, and replay policy. Public card discovery grants no invoke access.
|
|
162
166
|
- Remote A2A endpoints/card URLs require exact HTTPS origin allow-lists and redirect rejection. Pin ES256 card keys/expiry; never auto-fetch untrusted `jku`.
|
|
163
|
-
- Treat cards, task status, errors, artifacts, and SSE frames as untrusted bounded input and redact before logs/hooks/events.
|
|
167
|
+
- Treat cards, rich parts, task status/history, errors, artifacts, and SSE replay frames as untrusted bounded input and redact before logs/hooks/events. URL parts require host public/pinned-network policy and are never auto-fetched. Durable task/push adapters repeat exact-owner checks; foreign/missing records share not-found responses.
|
|
168
|
+
- Push delivery remains host-owned: validate every attempt/redirect against SSRF/rebinding policy, cap retries/time/output, authenticate webhook payloads, deduplicate event IDs, and keep token/auth credentials out of configs returned over A2A. Client streaming uses fatal UTF-8 decoding and rejects partial/post-terminal frames.
|
|
169
|
+
|
|
170
|
+
## Web research boundaries
|
|
171
|
+
|
|
172
|
+
- Construct `@arnilo/prism-web-tools` with one host-selected Brave or Exa adapter; never expose adapter/provider/credential/schema selection to model arguments.
|
|
173
|
+
- Provider API origins are fixed exact HTTPS origins and redirects fail. Credentials resolve immediately before I/O; remote bodies and secrets are excluded from errors/results/telemetry.
|
|
174
|
+
- Firecrawl targets reject userinfo, non-HTTP(S), private literals, and policy-denied hosts. Supply `validateUrl` for host DNS/rebinding/egress checks. Firecrawl performs remote retrieval, so Prism cannot pin target DNS after handoff.
|
|
175
|
+
- Treat every snippet, highlight, Markdown byte, metadata field, and extracted JSON value as prompt-injection-capable untrusted data. Never elevate it into system instructions or let it modify tools, permissions, trust, credentials, routing, or extraction schema.
|
|
176
|
+
- Keep counts/bytes/retries/rate delays/polling/concurrency/wall time finite. Live credentials belong only in explicit protected `PRISM_LIVE_WEB=1` runs.
|
|
177
|
+
|
|
178
|
+
## Supply-chain and live-canary boundaries
|
|
179
|
+
|
|
180
|
+
- Require `security / codeql`, `security / supply-chain`, PR dependency review, release readiness, and PostgreSQL integration in protected-branch rules. Enable GitHub secret scanning and push protection as repository settings; checked-in workflows cannot enable those service controls.
|
|
181
|
+
- Actions are pinned to full commit revisions. Dependabot proposes weekly npm/action revision changes; review upstream release notes before merge rather than replacing pins with moving tags.
|
|
182
|
+
- `scripts/verify-sbom.mjs` accepts only bounded SPDX 2.3 inventory with exact checked-in permissive licenses. Any missing/new expression fails until reviewed; do not widen policy merely to unblock CI.
|
|
183
|
+
- `scripts/scan-secrets.mjs` checks tracked source and unpacked public tarballs for high-confidence credential/private-key forms without printing matched values. It complements GitHub secret scanning; it is not entropy scanning or DLP.
|
|
184
|
+
- Tag publication alone receives npm/OIDC/attestation permissions. Untrusted pull-request code receives no canary, npm, or OIDC secret and no workflow uses `pull_request_target`.
|
|
185
|
+
- Scheduled/manual canaries run only in protected `live-canaries` environment. Use dedicated read-only/low-quota credentials and provider account spend limits. Runner performs four probes, at most one MCP cleanup, one provider output token, one Brave result, 64-KiB responses, and finite timeouts; report excludes endpoints, headers, bodies, credentials, and MCP session IDs.
|
|
186
|
+
- Live endpoint operators own TLS, egress allow-lists, account-dollar budget, cleanup beyond MCP session DELETE, and revocation. Failed canaries log only operation kind plus status/timeout; inspect provider-side audit logs for details.
|
|
164
187
|
|
|
165
188
|
## Related APIs
|
|
166
189
|
|
package/docs/index.md
CHANGED
|
@@ -10,11 +10,11 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
10
10
|
- [Agent definitions](agent-definitions.md): resolve declarative `AgentDefinition` values via `resolveAgentDefinition`, and turn app-config `<configRoot>/agents/<name>/AGENT.md` bundles into runnable agents via `discoverAgentBundles` / `resolveAgentBundle` (explicit tool/skill activation by name, fail-closed omitted capabilities, migration-only `activateAllCapabilities`, strict duplicate scope checks, configurable prompt layers, no auto-discovery).
|
|
11
11
|
- [Agent loops](agent-loops.md): replaceable per-run control loops — `singleShotLoop` default and opt-in bounded artifact-loop tool rounds with host-supplied `validator`/`parser`/`repairer` callbacks.
|
|
12
12
|
- [Guardrails](guardrails.md): typed fail-closed input/output/tool checks with buffered provider output and redacted decision records.
|
|
13
|
-
- [Agent events](agent-events.md):
|
|
14
|
-
- [Observability](observability.md):
|
|
15
|
-
- [Evaluations](evaluations.md):
|
|
16
|
-
- [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence
|
|
17
|
-
- [Performance limits](performance.md): bounded live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
|
|
13
|
+
- [Agent events](agent-events.md): redacted lifecycle stream used by UIs, ledgers, and metadata-only parented telemetry; message/progress deltas never create spans.
|
|
14
|
+
- [Observability](observability.md): OTel GenAI agent/provider/tool hierarchy, host context parenting, bounded trace linkage, safe evaluation events, controlled metrics, and exporter isolation.
|
|
15
|
+
- [Evaluations](evaluations.md): deterministic and bounded trace/model-judge/pairwise scoring, CI thresholds, OTel trace-reference linkage, and ID-only linkage to immutable owned run feedback.
|
|
16
|
+
- [Runs and usage ledger](runs-and-usage.md): durable run/event/tool/usage persistence, optional bounded FIFO durability policies, session snapshot caching, and immutable run/trace feedback.
|
|
17
|
+
- [Performance limits](performance.md): bounded evaluation traces/judges/reports, security scan/live-canary backstops, live subscriber queues, branch-read pagination expectations, JSONL/dev-store limits, and production sizing assumptions.
|
|
18
18
|
- [Structured output](structured-output.md): the `Artifact*` seam plus provider-native `StructuredOutputOptions` / `structuredOutputMode` for capable models.
|
|
19
19
|
|
|
20
20
|
## Compaction/session memory
|
|
@@ -27,7 +27,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
27
27
|
- [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention, and NoSQL mapping.
|
|
28
28
|
- [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, and transactionally verified/backfilled migration-v3 metadata.
|
|
29
29
|
- [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
|
|
30
|
-
- [Migration guide](migration.md): 0.0.3 compatibility
|
|
30
|
+
- [Migration guide](migration.md): 0.0.3 compatibility through 0.0.8 telemetry/evaluation, MCP/A2A, web research, ledger batching, and release-security changes.
|
|
31
31
|
- [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety.
|
|
32
32
|
- [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
|
|
33
33
|
|
|
@@ -57,7 +57,8 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
57
57
|
- [Tools](tools.md): register host-owned active tools with replace-or-error duplicate policy, apply exact allow/deny filtering, dispatch normal or opt-in bounded artifact-loop calls, and optionally bound untrusted JSON Schema compilation.
|
|
58
58
|
- [Tool execution primitives](tool-execution-primitives.md): finite JSON Schema LRU validation, exclusive-aware bounded parallel dispatch, MCP bridge mapping, coding execution policy, and image-read bounds.
|
|
59
59
|
- [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
|
|
60
|
-
- [MCP client bridge and server exposure](mcp-tools.md):
|
|
60
|
+
- [MCP client bridge and server exposure](mcp-tools.md): SDK-1.29.0 bounded tools/resources/prompts, host-owned roots/sampling/elicitation, exact-origin DNS-pinned client transport, and principal-bound opt-in Streamable HTTP sessions.
|
|
61
|
+
- [Web search, fetch, and extraction](web-tools.md): optional host-selected Brave/Exa discovery and Firecrawl Markdown/schema tools with native fetch, stable citations, late credentials, finite limits, and explicit untrusted-content boundaries.
|
|
61
62
|
- [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, and `edit` definitions with streamed text pages, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
|
|
62
63
|
- [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, and abort-aware streaming sandbox adapters for coding tools.
|
|
63
64
|
|
|
@@ -77,8 +78,8 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
77
78
|
- [Web-standard server handler](server.md): optional framework-free authorized direct/SSE agent, explicitly selected durable agent lifecycle, and durable workflow routes with explicit bounds and zero default exposure.
|
|
78
79
|
|
|
79
80
|
## Multi-agent and interoperability
|
|
80
|
-
- [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, and
|
|
81
|
-
- [A2A interoperability](a2a.md):
|
|
81
|
+
- [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, finite budgets, host-projected delegation telemetry, and separate A2A durable adapter boundary.
|
|
82
|
+
- [A2A interoperability](a2a.md): A2A 1.0 JSON-RPC/HTTPS cards plus host-owned durable task get/list/cancel/subscribe, bounded rich parts/replay, principal-scoped push configs, and exact-origin verified client.
|
|
82
83
|
|
|
83
84
|
## CLI/RPC
|
|
84
85
|
- [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
|
|
@@ -87,7 +88,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
87
88
|
- [Workflow/TUI scope](workflow-tui-primitives.md): records why 0.0.5 ships workflow APIs/RPC control but no interactive terminal UI.
|
|
88
89
|
|
|
89
90
|
## Security and credentials
|
|
90
|
-
- [Host security guide](host-security.md): fail-closed checklist for
|
|
91
|
+
- [Host security guide](host-security.md): fail-closed checklist for supply-chain/attestation/canary isolation, bounded credentials, JSON/schema/vector/crypto, MCP/A2A/web remote boundaries, untrusted external content, settings, redaction, trust roots, workflow ownership, coding I/O, permissions, persistence, extensions, and tool validation.
|
|
91
92
|
- [Security/auth/trust](settings-auth-trust-security.md): settings providers, credential helpers, trust/permission policies, redaction controls, host-owned settings/credentials wiring outside `AgentConfig`, and security-boundary hardening summary.
|
|
92
93
|
- [Credentials and redaction](credentials-and-redaction.md): compose explicit credential resolver order, use caller-supplied env objects/OAuth refresh helpers, resolve credentials only at the provider edge, and redact known secret values.
|
|
93
94
|
- [Credential storage](credential-storage.md): optional `@arnilo/prism-credentials-node` adapter with strict bounded AES-GCM envelopes, async finite scrypt, restrictive Unix files, and abort-aware bounded system-keychain calls.
|
|
@@ -96,14 +97,15 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
96
97
|
- Provider test doubles: `createMockProvider()` and provider event helpers are documented on the canonical Provider layer page above.
|
|
97
98
|
- [Provider conformance](provider-conformance.md): run network-free provider adapter assertions (stream order, abort, tool-call reconstruction, cache usage, content coverage, protected header ownership, secret leak) from `@arnilo/prism/testing/provider-conformance`.
|
|
98
99
|
- [Session store conformance](session-store-conformance.md): assert any `SessionStore` adapter satisfies append/idempotency/conflict/branch invariants from `@arnilo/prism/testing/session-store-conformance`.
|
|
99
|
-
- [Run ledger conformance](run-ledger-conformance.md): assert durable run/event/tool/usage writes and reopen survival. Run-feedback stores use `@arnilo/prism/testing/feedback` for append/query/delete/ownership linkage conformance.
|
|
100
|
+
- [Run ledger conformance](run-ledger-conformance.md): assert durable run/event/tool/usage writes and reopen survival; batch-wrapper FIFO/bounds/flush checks remain separate. Run-feedback stores use `@arnilo/prism/testing/feedback` for append/query/delete/ownership linkage conformance.
|
|
100
101
|
- [Compaction conformance](compaction-conformance.md): assert any `CompactionStrategy` returns a non-empty redacted summary and observes abort from `@arnilo/prism/testing/compaction-conformance`.
|
|
101
102
|
- [Tool conformance](tool-conformance.md): assert the tool-dispatch blocked-reason matrix (unknown/denied/invalid/permission/validator) and success path from `@arnilo/prism/testing/tool-conformance`.
|
|
102
103
|
- [Extension conformance](extension-conformance.md): assert an `Extension` setup runs, contributions stay inert, and setup errors are redacted or rethrown from `@arnilo/prism/testing/extension-conformance`.
|
|
103
104
|
- `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, cache-aware prompt assembly, NeuralWatt agent run, stores/branching, compaction, observational-memory recall, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
|
|
104
105
|
|
|
105
106
|
## Release and install
|
|
106
|
-
- [Release and install](release-and-install.md):
|
|
107
|
+
- [Release and install](release-and-install.md): 31-package graph, install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, and protected live canaries.
|
|
108
|
+
- [Review coverage (2026-07-19 Phase 3)](review-coverage-2026-07-19-phase-3.md): Plan 070 evidence freeze — exact protocol/vendor references, capability/primitive/limit matrices, supported boundaries, and 0.0.8 release evidence.
|
|
107
109
|
- [Review coverage (2026-07-17 provider validation)](review-coverage-2026-07-17-provider-validation.md): Plan 067 evidence freeze — P0–P2 re-verification owners, seven first-party provider packages mapped to official-doc URLs, Pi secondary refs, cache/thinking/discovery surfaces, credential canaries, and use-case model-binding inventory.
|
|
108
110
|
- [Review coverage (2026-07-15)](review-coverage-2026-07-15.md): frozen 0.0.5 finding/feature ownership, existing-primitive inventory, package decisions, threat boundaries, exclusions, and measured Phase 0 baseline.
|
|
109
111
|
- [Review coverage (2026-07-14)](review-coverage-2026-07-14.md): traceability matrix linking review findings and bug-report fixes to plan tasks, tests, and documentation for release 0.0.4.
|
package/docs/mcp-tools.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-mcp` has two explicit directions. Its client bridge connects hosts to remote [Model Context Protocol](https://modelcontextprotocol.io) servers and maps discovered tools to ordinary `ToolDefinition`s. Its server API registers selected Prism `ToolDefinition` and `CommandDefinition` values on the official SDK `McpServer`, with required authorization and a bounded optional Web-standard Streamable HTTP handler. The package
|
|
5
|
+
`@arnilo/prism-mcp` has two explicit directions. Its client bridge connects hosts to remote [Model Context Protocol](https://modelcontextprotocol.io) servers and maps discovered tools to ordinary `ToolDefinition`s. Its server API registers selected Prism `ToolDefinition` and `CommandDefinition` values on the official SDK `McpServer`, with required authorization and a bounded optional Web-standard Streamable HTTP handler. The package pins `@modelcontextprotocol/sdk` **1.29.0** (MCP protocol negotiation remains SDK-owned) and adds no MCP branch to core Prism.
|
|
6
6
|
|
|
7
7
|
Primary API:
|
|
8
8
|
|
|
@@ -19,7 +19,21 @@ await bridge.refresh(); // re-list after notifications or TTL expiry
|
|
|
19
19
|
await bridge.close(); // close client + transport
|
|
20
20
|
```
|
|
21
21
|
|
|
22
|
-
Advanced hosts that manage their own `Client` + `Transport` can call `attachMcpToolBridge(
|
|
22
|
+
Advanced hosts that manage their own `Client` + `Transport` can call `attachMcpToolBridge()` or `attachMcpCapabilities()` after connect. `connectMcpCapabilities()` keeps resources/prompts as host-facing facades rather than converting them into model tools, and declares roots/sampling/elicitation only when callbacks are supplied.
|
|
23
|
+
|
|
24
|
+
```ts
|
|
25
|
+
const bridge = await connectMcpCapabilities({
|
|
26
|
+
serverId: "research",
|
|
27
|
+
transport: { type: "streamable-http", url, allowedOrigins: [origin] },
|
|
28
|
+
roots: () => [{ uri: "file:///workspace", name: "workspace" }],
|
|
29
|
+
sampling: hostSampling, // host selects model/provider/credentials
|
|
30
|
+
elicitation: hostElicitation, // URL mode returns approval; Prism never opens/fetches URL
|
|
31
|
+
});
|
|
32
|
+
await bridge.listResources();
|
|
33
|
+
await bridge.getPrompt("review", { topic: "security" });
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
Server capability matrix for SDK 1.29.0: tools/resources/prompts and their list-change notifications are supported through official registrations; roots/sampling/form+URL elicitation are supported as explicit client callbacks. Missing server resources/prompts throw `McpUnsupportedCapabilityError` with `ERR_PRISM_MCP_UNSUPPORTED_CAPABILITY`. Resource/prompt results and sampling/elicitation inputs/results are bounded JSON. Accepted form/URL elicitation requires host-only `humanInteraction: true`; bridge strips marker before protocol output and fails closed when absent. Automatic root discovery/consent, model selection, credential resolution, URL navigation, generic command proxying, and custom JSON-RPC are unsupported.
|
|
23
37
|
|
|
24
38
|
Server direction:
|
|
25
39
|
|
|
@@ -41,10 +55,13 @@ const handleMcp = await createPrismMcpWebHandler(server, {
|
|
|
41
55
|
resolveAuthInfo: authenticateRequest,
|
|
42
56
|
allowedHosts: ["api.example.test"],
|
|
43
57
|
allowedOrigins: ["https://app.example.test"],
|
|
58
|
+
// Omit these two for bounded stateless JSON mode.
|
|
59
|
+
sessionIdGenerator: crypto.randomUUID,
|
|
60
|
+
resolveIdentity: (_request, auth) => auth ? { id: validatedPrincipalId(auth) } : false,
|
|
44
61
|
});
|
|
45
62
|
```
|
|
46
63
|
|
|
47
|
-
`McpServer.connect(transport)` remains available for SDK stdio or in-memory transports. The helper uses SDK `WebStandardStreamableHTTPServerTransport
|
|
64
|
+
`McpServer.connect(transport)` remains available for SDK stdio or in-memory transports. The helper uses SDK `WebStandardStreamableHTTPServerTransport`; it does not start a listener. Default remains bounded stateless JSON-response mode. Supplying `sessionIdGenerator` enables SDK `MCP-Session-Id` POST/GET/DELETE/SSE lifecycle and requires exact `allowedOrigins` plus host `resolveIdentity`. Every request re-authenticates, and a different principal receives non-disclosing 404. SDK owns protocol-version/session headers and SSE semantics. SDK 1.29.0's in-memory event store is not enabled, so `Last-Event-ID` replay is explicitly unsupported; reconnect starts only through SDK-supported active session GET.
|
|
48
65
|
|
|
49
66
|
## When to use it
|
|
50
67
|
|
|
@@ -160,6 +177,7 @@ Plaintext is accepted only when `allowLoopbackHttp: true`, the URL hostname is l
|
|
|
160
177
|
| Option | Default | Purpose |
|
|
161
178
|
| --- | --- | --- |
|
|
162
179
|
| `tools` / `commands` | empty | Explicit allow-list; zero default exposure |
|
|
180
|
+
| `resources` / `prompts` | empty | Static URI/name registrations with bounded host callbacks and per-read/get authorization |
|
|
163
181
|
| `agentRuns` | empty | Explicit `{ [agentId]: { lifecycle } }` map; registers `agent.<id>.status` and `agent.<id>.resume` only |
|
|
164
182
|
| `authorize` | required | Per-call host authz using SDK auth/session metadata |
|
|
165
183
|
| `permission` / `validate` / `redactor` | none | Core tool-dispatch gates and known-secret redaction |
|
|
@@ -167,7 +185,7 @@ Plaintext is accepted only when `allowLoopbackHttp: true`, the URL hostname is l
|
|
|
167
185
|
| `maxConcurrentCalls` | 16 (256 hard) | Bound active tool/command execution |
|
|
168
186
|
| `callTimeoutMs` | 60 s (30 min hard) | Abort and return timed-out calls |
|
|
169
187
|
|
|
170
|
-
Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard), 32 concurrent requests (512 hard),
|
|
188
|
+
Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard), 32 concurrent requests (512 hard), 60 s timeout (30 min hard), and 32 sessions (512 hard). Stateful mode is intentionally one official SDK transport/session lineage per handler; use one handler/server instance per independently hosted endpoint when multi-tenant transport isolation is required. It parses bounded JSON before passing `parsedBody` to the SDK transport. `allowedHosts`/`allowedOrigins` activate SDK DNS-rebinding checks only when explicitly configured. Authentication data comes only from host `resolveAuthInfo()`.
|
|
171
189
|
|
|
172
190
|
## Security and performance notes
|
|
173
191
|
|
|
@@ -183,7 +201,8 @@ Web handler defaults: 1 MiB request (8 MiB hard), 2 MiB response (16 MiB hard),
|
|
|
183
201
|
| Accidental server exposure | Empty default arrays/maps, duplicate-name rejection, explicit tools/commands/lifecycle only |
|
|
184
202
|
| Agent lifecycle data leak or cross-tenant resume | `agentRuns` requires exact tenant plus account/user ownership; core lifecycle returns public redacted state only and CAS-resumes with current agent/revision |
|
|
185
203
|
| Unbounded MCP HTTP | Bounded pre-parsed JSON, response bytes, concurrent requests, call timeout, SDK web-standard transport |
|
|
186
|
-
| Cross-tenant operation | Authorizer derives ownership from validated auth and passes it to tool dispatch
|
|
204
|
+
| Cross-tenant operation | Authorizer derives ownership from validated auth and passes it to tool/resource/prompt dispatch; stateful handler binds session to stable validated principal on every request; never trust arguments as identity |
|
|
205
|
+
| Sampling / elicitation authority | Host callbacks alone choose model/provider/credentials or obtain consent; bounded URL elicitation is returned to host UI and never fetched/opened automatically; auth tokens never enter callback params/results |
|
|
187
206
|
|
|
188
207
|
For durable lifecycle exposure, construct `createAgentRunLifecycle({ checkpoints, resolveAgent })` in core, then pass selected entries as `agentRuns: { support: { lifecycle } }`. MCP registers two tools: `agent.support.status` accepts `{ runId, sessionId? }`; `agent.support.resume` accepts `{ runId, sessionId?, decision, expectedVersion }`. Do not expose an agent without durable checkpoints and a restart-safe `SessionStore`; no lifecycle tool appears by default.
|
|
189
208
|
|
|
@@ -191,9 +210,14 @@ MCP output is untrusted. Register bridge tools through core dispatch with a `Sec
|
|
|
191
210
|
|
|
192
211
|
Discovery validation is atomic: cursor/page/tool/name/description/schema failures reject `refresh()` and preserve the previous immutable tool-array reference. The bridge intentionally uses raw SDK `request()` for `tools/list` and `tools/call`; this avoids eager Ajv compilation/validation of untrusted remote output schemas. Host `ToolValidator` remains the argument-validation owner.
|
|
193
212
|
|
|
213
|
+
## Vendor web MCP prototype boundary
|
|
214
|
+
|
|
215
|
+
Official Exa/Firecrawl MCP servers may be tested only as explicit hardened prototypes: pin endpoint/origin/auth, inspect declared capabilities, allow-list individual tools/resources, retain all MCP bounds, and never expose generic remote passthrough. Production web research uses direct host-selected `@arnilo/prism-web-tools` adapters so provider choice, credentials, schema, and costs remain outside model control.
|
|
216
|
+
|
|
194
217
|
## Related APIs
|
|
195
218
|
|
|
196
219
|
- [Tools](tools.md): registry, dispatch, validation
|
|
220
|
+
- [Web search, fetch, and extraction](web-tools.md): preferred direct bounded Brave/Exa/Firecrawl production path
|
|
197
221
|
- [Tool execution primitives](tool-execution-primitives.md): Plan 055 design and conformance matrix
|
|
198
222
|
- [Host security guide](host-security.md): permission, trust, validation checklist
|
|
199
223
|
- [Web-standard server handler](server.md): agent/workflow HTTP routes and shared remote-boundary rules
|
package/docs/migration.md
CHANGED
|
@@ -7,6 +7,42 @@ Prism 0.0.6 preserves documented 0.0.3 agent construction except for two intenti
|
|
|
7
7
|
1. **`session.run()` / `session.prompt()` return `AgentRunResult`** and `session.stream()` starts one owned run after subscribing. Callers that ignored the previous `Promise<void>` keep working; failed/aborted runs reject with `AgentRunError` (`.result` attached).
|
|
8
8
|
2. **`AgentConfig.extensions` / `settings` / `credentials` are removed.** Wire extensions through `createExtensionKernel()`, read settings in the host, and pass credential resolvers to the provider edge.
|
|
9
9
|
|
|
10
|
+
## 0.0.7 → 0.0.8 release overview
|
|
11
|
+
|
|
12
|
+
All 31 first-party manifests and exact internal ranges move together to `0.0.8`; mixed first-party versions are unsupported. Core remains dependency-free at runtime and existing low-level agent/session APIs remain compatible. New telemetry, evaluation, MCP, A2A, ledger batching, and web research surfaces are opt-in. Release CI now requires CodeQL, dependency/license/SBOM/secret checks, packed-artifact attestations, PostgreSQL integration, and protected live-canary prerequisites; no tag or publication is automatic from this migration.
|
|
13
|
+
|
|
14
|
+
## 0.0.7 → 0.0.8 evaluations and ledger operation
|
|
15
|
+
|
|
16
|
+
`@arnilo/prism-evals` adds owner-scoped trace resolution, optional host model judges, deterministic pairwise reports, and `assertEvaluationThreshold()` without changing stored evaluation schemas. Hosts select all judge/provider credentials and should version rubrics. Core adds optional `createBatchedRunLedger()`; direct ledgers remain write-through. Choose `flush_on_terminal` only after accepting bounded pre-flush crash loss, and call `dispose()` during shutdown. Runtime snapshot caching is session/leaf-local and requires no persistence migration.
|
|
17
|
+
|
|
18
|
+
## 0.0.7 → 0.0.8 web research tools
|
|
19
|
+
|
|
20
|
+
Install `@arnilo/prism-web-tools` explicitly (or through `@arnilo/prism-all`) to add web capability; core and existing profiles remain inert. Select Brave or Exa at construction, provide Firecrawl separately for Markdown/schema extraction, and register returned `web_search`/`web_fetch`/`web_extract` tools through normal permission/trust/validation dispatch. Provider selection, credentials, target DNS policy, and extraction schema are host-only. All returned content is marked untrusted; no browser or vendor SDK is added.
|
|
21
|
+
|
|
22
|
+
## 0.0.7 → 0.0.8 A2A durable tasks
|
|
23
|
+
|
|
24
|
+
Existing text `createA2AHandler({ exposure })`, `client.send()`, and `client.stream()` remain compatible. Add host `tasks` to enable `GetTask`/`ListTasks`/`CancelTask`/`SubscribeToTask`, rich parts, interrupted states, and replay cursors; no task store or migration is created. Add host `push` for push-config CRUD and matching card capability. Raw/data/URL parts are disabled until selected in `parts`; URL/push endpoints additionally require host URL policy and are never fetched by part parsing. Push delivery/retries/idempotency remain host-owned.
|
|
25
|
+
|
|
26
|
+
## 0.0.7 → 0.0.8 MCP capabilities and sessions
|
|
27
|
+
|
|
28
|
+
`@arnilo/prism-mcp` now pins official SDK 1.29.0. Existing `connectMcpTools()` and stateless web handlers remain compatible. Use `connectMcpCapabilities()` for bounded resources/prompts and explicit roots/sampling/elicitation callbacks. Server resources/prompts must be selected explicitly and authorize every operation. Stateful Streamable HTTP additionally requires `sessionIdGenerator`, exact `allowedOrigins`, and host `resolveIdentity`; omission preserves stateless mode. `Last-Event-ID` replay is not enabled. Missing capability calls fail with `ERR_PRISM_MCP_UNSUPPORTED_CAPABILITY`.
|
|
29
|
+
|
|
30
|
+
## 0.0.7 → 0.0.8 OpenTelemetry adapter
|
|
31
|
+
|
|
32
|
+
The optional observability package now emits OTel GenAI names and units instead of independent `prism.agent.run` / `prism.provider.turn` / `prism.tool.execute` spans and millisecond metrics. Update dashboards to `invoke_agent prism`, `chat {model}`, `execute_tool {tool}`, `gen_ai.*.duration` (seconds), and `gen_ai.client.token.usage`. Pass `{ context, trace }` as third `wrapOpenTelemetryApi()` argument for native parent context, and use `onTraceReference` or `traceId(runId)` for evaluation linkage. Core APIs and persistence schemas are unchanged.
|
|
33
|
+
|
|
34
|
+
## 0.0.7 → 0.0.8 Kimi provider alignment
|
|
35
|
+
|
|
36
|
+
`@arnilo/prism-provider-kimi` now matches the official contracts: featured Coding `k3` defaults to `reasoning_effort: "high"` (Open Platform `kimi-k3` keeps `"max"`); featured context windows use the official `262_144` for 256K-class models; the featured Moonshot catalog adds `kimi-k2.7-code-highspeed`, `kimi-k2.6`, and `kimi-k2.5` (K2.5 intentionally without Preserved Thinking). Provider-owned compat keys (`route`, `preserveThinking`, `preserve_thinking`) are stripped before the opaque compat spread and no longer leak into request bodies. The Coding route additionally sends provider-owned `x-api-key` and `anthropic-version: 2023-06-01` headers per the official third-party setup. Streams emit `done` only on protocol completion evidence (`message_stop` on the Coding route, `[DONE]` + `finish_reason` on the Moonshot route); truncated streams now surface as run failures.
|
|
37
|
+
|
|
38
|
+
## 0.0.7 → 0.0.8 artifact-loop parse failures
|
|
39
|
+
|
|
40
|
+
`generateValidateReviseLoop` no longer returns silently on artifact parse failure. A parser returning `{ ok: false }` (or no `value`) now consumes revision budget exactly like a validation failure: the repairer receives `value: undefined` plus a synthetic failure (`metadata.reason: "parse_error"`), and exhaustion ends with terminal `artifact_failed`. Host repairers must already tolerate `value: undefined` per the `ArtifactRepairer` contract; runs that previously ended after one silent parse failure now spend up to `maxRevisions` repair turns first.
|
|
41
|
+
|
|
42
|
+
## 0.0.7 → 0.0.8 OpenCode Go provider fixes
|
|
43
|
+
|
|
44
|
+
`@arnilo/prism-provider-opencode-go` no longer infers `structuredOutput: "json_schema"` from OpenAI-compatible routing alone. Only verified models (`mimo-v2.5`, `mimo-v2.5-pro`) advertise it; other OpenAI-route models (for example `deepseek-v4-pro`) now use the artifact-loop parsing/validation path, and requests that still pass `options.structuredOutput` for an unverified model fail before dispatch with `unsupported_model`. Hosts with their own verification evidence can set the capability explicitly through `defineOpenCodeGoModel({ capabilities })`. The Anthropic route additionally sends provider-owned `x-api-key` and `anthropic-version: 2023-06-01` headers alongside Bearer, fixing HTTP 401 on MiniMax/Qwen models; caller headers cannot override them. Streams now emit `done` only on protocol completion evidence (`[DONE]` plus a terminal `finish_reason` on the OpenAI route, `message_stop` on the Anthropic route) with no dangling tool-call accumulators; truncated connections and incomplete tool calls terminate with an `error` event, so hosts may see previously silent truncations surface as run failures.
|
|
45
|
+
|
|
10
46
|
## 0.0.6 → 0.0.7 secure run lifecycle
|
|
11
47
|
|
|
12
48
|
`createAgent()` remains backward-compatible. Version 0.0.7 adds opt-in typed `Guardrails` (`input`, provider `output`, `toolInput`, `toolOutput`) and narrowing-only `RunLimits`. Output guardrails and configured output-token/total-token/cost limits buffer provider output before exposure; blocked content is neither emitted nor persisted. A breach emits one redacted `run_limit_exceeded` event and rejects with `AgentRunError.result.limit`.
|
package/docs/observability.md
CHANGED
|
@@ -52,8 +52,17 @@ OpenTelemetry adapter:
|
|
|
52
52
|
import { trace, metrics } from "@opentelemetry/api";
|
|
53
53
|
import { createOpenTelemetryInstrumentation, wrapOpenTelemetryApi } from "@arnilo/prism-observability-opentelemetry";
|
|
54
54
|
|
|
55
|
-
const { tracer, meter } = wrapOpenTelemetryApi(
|
|
56
|
-
|
|
55
|
+
const { tracer, meter } = wrapOpenTelemetryApi(
|
|
56
|
+
trace.getTracer("app"),
|
|
57
|
+
metrics.getMeter("app"),
|
|
58
|
+
{ context, trace },
|
|
59
|
+
);
|
|
60
|
+
const telemetry = createOpenTelemetryInstrumentation({
|
|
61
|
+
tracer,
|
|
62
|
+
meter,
|
|
63
|
+
onTraceReference: ({ runId, traceId }) => saveRunTrace(runId, traceId),
|
|
64
|
+
onExporterError: console.error,
|
|
65
|
+
});
|
|
57
66
|
|
|
58
67
|
const detach = telemetry.attachSession(session);
|
|
59
68
|
// or: for await (const event of session.subscribe()) telemetry.handleAgentEvent(event);
|
|
@@ -79,13 +88,13 @@ OpenTelemetry mapping (when enabled):
|
|
|
79
88
|
|
|
80
89
|
| Agent event | Span | Metric labels |
|
|
81
90
|
| --- | --- | --- |
|
|
82
|
-
| `agent_started` /
|
|
83
|
-
| `provider_turn_*` | `
|
|
84
|
-
| `tool_execution_*`
|
|
85
|
-
| `
|
|
86
|
-
| `
|
|
87
|
-
| `handleRunFeedback` | active-run `prism.run.feedback` event or ended-run span | `prism.run.feedback`
|
|
88
|
-
| `handleEvaluation` | active-run `
|
|
91
|
+
| `agent_started` / terminal event | `invoke_agent prism` (`INTERNAL`) | `gen_ai.invoke_agent.duration` |
|
|
92
|
+
| `provider_turn_*` | `chat {model}` (`CLIENT`) | `gen_ai.client.operation.duration`, `gen_ai.client.token.usage` |
|
|
93
|
+
| `tool_execution_*` | `execute_tool {tool}` (`INTERNAL`) when started | `gen_ai.execute_tool.duration` |
|
|
94
|
+
| `guardrail_decision` | `prism.guardrail.evaluate` child (`INTERNAL`) | none |
|
|
95
|
+
| `handleDelegation()` | `prism.agent.delegate` child (`INTERNAL`) | none |
|
|
96
|
+
| `handleRunFeedback` | active-run `prism.run.feedback` event or ended-run span | `prism.run.feedback` |
|
|
97
|
+
| `handleEvaluation` | active-run `gen_ai.evaluation.result` event or ended-run span | `prism.run.evaluation` (`status`) |
|
|
89
98
|
|
|
90
99
|
High-cardinality identifiers (`sessionId`, `runId`, `requestId`, `toolCallId`) are **span attributes only**, never metric labels.
|
|
91
100
|
|
|
@@ -135,11 +144,11 @@ const session = createAgent({
|
|
|
135
144
|
|
|
136
145
|
const detach = telemetry.attachSession(session);
|
|
137
146
|
const result = await session.run("hello");
|
|
147
|
+
const traceId = telemetry.traceId(result.runId); // or persist onTraceReference immediately
|
|
138
148
|
detach();
|
|
139
149
|
telemetry.handleRunFeedback({ runId: result.runId, rating: 1, hasComment: true, tagCount: 1, scorerCount: 1, evaluationCount: 1 });
|
|
140
|
-
telemetry.handleEvaluation({ runId: result.runId, status: "scored", score: 0.9, hasReason: true });
|
|
141
|
-
|
|
142
|
-
console.log(memory.spans.map((span) => span.name));
|
|
150
|
+
telemetry.handleEvaluation({ runId: result.runId, name: "citation", status: "scored", score: 0.9, hasReason: true });
|
|
151
|
+
console.log(traceId, memory.spans.map((span) => span.name));
|
|
143
152
|
```
|
|
144
153
|
|
|
145
154
|
## Extension and configuration notes
|
|
@@ -149,14 +158,17 @@ console.log(memory.spans.map((span) => span.name));
|
|
|
149
158
|
- NeuralWatt `neuralwatt:telemetry` provider events remain package-local; hosts may forward numeric cost/energy into custom metrics.
|
|
150
159
|
- `@arnilo/prism-observability-opentelemetry` is optional and included through `@arnilo/prism-sdk` and `@arnilo/prism-all`; instrumentation remains disabled until a host configures it.
|
|
151
160
|
- Exporter failures are isolated: instrumentation catches tracer/meter errors and invokes `onExporterError` without affecting the run, feedback persistence, or evaluation scoring.
|
|
152
|
-
-
|
|
161
|
+
- Trace grading uses `createPersistenceTraceResolver()` with explicit session/run/ownership and finite pages/bytes. Judge reasons remain evaluation data; `gen_ai.evaluation.result` receives only name, finite score, controlled status, and reason-presence.
|
|
162
|
+
- Run spans parent provider, tool, guardrail, and explicit delegation spans. Pass `{ context, trace }` to `wrapOpenTelemetryApi()` for native parent context creation; `parentContext` can attach the run to host ambient/remote context.
|
|
163
|
+
- `onTraceReference` receives `{ runId, traceId }` when a run starts. `traceId(runId)` keeps only the newest 1,024 mappings by default (`maxTraceReferences`, hard cap 10,000); durable linkage remains host-owned.
|
|
164
|
+
- Run `error`, suspension, denial, and detach close every attributable span. Repeated terminal events are idempotent and cannot end a span twice.
|
|
153
165
|
- Disabled instrumentation performs no per-delta span work (`enabled: false` or missing tracer/meter).
|
|
154
166
|
|
|
155
167
|
## Security and performance notes
|
|
156
168
|
|
|
157
169
|
- Default events are metadata-only — no prompts, streamed deltas, tool arguments, or credentials.
|
|
158
170
|
- Opt-in content in other event types (`message_delta`, tool `result`) is still subject to `redactAgentEvent`.
|
|
159
|
-
- Metric labels stay low-cardinality (`
|
|
171
|
+
- Metric labels stay low-cardinality (`gen_ai.operation.name`, `gen_ai.provider.name`, token type, controlled outcome/status, feedback rating bucket/link presence); never use session/run/request/call IDs, model output, comments, tag values, scorer/evaluation IDs, or arbitrary metadata as labels. Token usage is recorded once at provider operation scope.
|
|
160
172
|
- Target overhead when enabled is under 5% excluding exporter I/O; disabled hooks allocate no spans.
|
|
161
173
|
- Provider transport limits and redaction order are documented in [Provider primitives](provider-primitives.md).
|
|
162
174
|
|
package/docs/performance.md
CHANGED
|
@@ -1,9 +1,34 @@
|
|
|
1
1
|
# Performance limits
|
|
2
2
|
|
|
3
|
+
Evaluation defaults are finite: 100 trace rows × 20 pages and 4 MiB aggregate trace data; one model-judge attempt with 30-second/16-KiB bounds; 8 comparison candidates, 1-MiB candidate results, 10,000 dataset items, and 4-MiB serialized reports. Hard caps are exported by `@arnilo/prism-evals`; overflow fails rather than truncating grading evidence.
|
|
4
|
+
|
|
3
5
|
## What it does
|
|
4
6
|
|
|
5
7
|
This page states Prism runtime limits that keep slow consumers and long sessions from becoming unbounded memory or latency problems.
|
|
6
8
|
|
|
9
|
+
## Release 0.0.8 reproducible synthetic evidence
|
|
10
|
+
|
|
11
|
+
Run `node scripts/benchmark-0.0.8.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000. Script uses no network/credentials and emits environment, throughput, p50/p95 latency, heap, synthetic disk bytes, zero external cost, and backpressure signals. These are evidence fields, not CI timing gates.
|
|
12
|
+
|
|
13
|
+
2026-07-20 baseline: Node v24.18.0, Linux x64, 1,000 operations/scenario.
|
|
14
|
+
|
|
15
|
+
| Scenario | ops/s | p95 ms | heap bytes | disk bytes | cost USD | backpressure |
|
|
16
|
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
|
17
|
+
| provider envelope | 675,430 | 0.0026 | 10,974,296 | 0 | 0 | 0 |
|
|
18
|
+
| actual `createBatchedRunLedger` enqueue/flush | 514,493 | 0.0023 | 12,601,840 | 0 | 0 | 7 |
|
|
19
|
+
| one-entry snapshot-cache hit | 4,582,216 | 0.0001 | 10,072,656 | 0 | 0 | 0 |
|
|
20
|
+
| actual in-memory OTel agent span start/end | 295,372 | 0.0049 | 10,357,048 | 0 | 0 | 0 |
|
|
21
|
+
| PostgreSQL-ledger-shaped file workload | 494,403 | 0.0009 | 11,584,760 | 54,890 | 0 | 7 |
|
|
22
|
+
| MCP envelope | 1,243,254 | 0.0008 | 12,915,288 | 0 | 0 | 0 |
|
|
23
|
+
| A2A envelope | 961,492 | 0.0007 | 10,425,024 | 0 | 0 | 0 |
|
|
24
|
+
| web-tools envelope | 1,388,694 | 0.0007 | 11,734,208 | 0 | 0 | 0 |
|
|
25
|
+
|
|
26
|
+
Ledger and OTel rows exercise shipped implementations; cache row isolates the runtime's one-entry lookup shape. Provider/PostgreSQL/MCP/A2A/web rows remain local serialization/file envelopes and prove repeatability/schema/backpressure instrumentation only—not external latency, throughput, or billing. Real PostgreSQL correctness runs in protected CI; provider/MCP/A2A/web timings and costs remain explicit protected live-canary/release-host evidence because this release-candidate host has no credentials/endpoints. No live claim is inferred from skipped gates.
|
|
27
|
+
|
|
28
|
+
Security automation is isolated from `npm test`: CodeQL/supply-chain jobs have 10-minute backstops, dependency review and live workflow have 5-minute job backstops, live probe step has 3 minutes, SBOM is capped at 16 MiB/10,000 packages, packed release/security tarballs at 128 MiB aggregate, secret scan at 100,000 files/16 MiB each, and retained security/canary reports expire after 7 days. Live canaries issue four probes plus at most one MCP cleanup, cap responses at 64 KiB and requests at 15 seconds (30 seconds hard), and never enter `sdk:ready`.
|
|
29
|
+
|
|
30
|
+
Web tools default/hard ceilings are query 4/16 KiB, results 10/20, URLs 5/20, request 256 KiB/1 MiB, response/aggregate 2/16 MiB, Markdown 1/8 MiB, extraction 256 KiB/1 MiB, schema 64/256 KiB, concurrency 4/16, retries 2/4, polling 20/100, and wall time 60 seconds/30 minutes. Bounds charge before request, retention, retry, or polling; overflow fails rather than truncating citation/extraction evidence.
|
|
31
|
+
|
|
7
32
|
Current surfaces:
|
|
8
33
|
|
|
9
34
|
- `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
|
|
@@ -125,6 +125,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
|
|
|
125
125
|
- **Identifier validation.** Configurable `schema` is validated and double-quoted; table names are fixed constants in adapter SQL.
|
|
126
126
|
- **TLS and credentials.** Configure via `pg` `Pool` / `PoolConfig`; the adapter does not read environment variables unless the host passes them into `connectionString` or `poolConfig`.
|
|
127
127
|
- **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
|
|
128
|
+
- **Optional batching.** PostgreSQL remains write-through by default. Hosts may wrap its ledger with core `createBatchedRunLedger()`; the bounded FIFO retains a failed head record and propagates flush failure instead of silently acknowledging it.
|
|
128
129
|
- **Bounded pool.** Adapter-owned pools default to `max: 10`. Hosts with heavy concurrency should supply their own pool sizing.
|
|
129
130
|
- **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid sequential scans.
|
|
130
131
|
- **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once. Startup catalog reads are bounded metadata queries, not application-row scans.
|
package/docs/providers/kimi.md
CHANGED
|
@@ -70,7 +70,7 @@ Unsupported block placements or unclaimed images fail before fetch.
|
|
|
70
70
|
| --- | --- | --- |
|
|
71
71
|
| Base URL | `https://api.kimi.com/coding` | `https://api.moonshot.ai/v1` (or `.cn`) |
|
|
72
72
|
| Wire API | Anthropic `/messages` | OpenAI `/chat/completions` |
|
|
73
|
-
| Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k3` (+ discovery) |
|
|
73
|
+
| Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k2.7-code-highspeed`, `kimi-k2.6`, `kimi-k2.5`, `kimi-k3` (+ discovery) |
|
|
74
74
|
| Discovery | No public list API — curated featured aliases | Official `GET /v1/models` via `listKimiModels()` |
|
|
75
75
|
| Cache | Implicit by default; opt-in Anthropic `cache_control` | Implicit only — never emits Anthropic `cache_control` |
|
|
76
76
|
| Thinking | Block replay + body `thinking` / `reasoning_effort` | `reasoning_content` replay + body `thinking` / `reasoning_effort` |
|
|
@@ -86,12 +86,26 @@ Official fields (Open Platform docs; Coding docs for `k3` effort mapping):
|
|
|
86
86
|
|
|
87
87
|
| Model family | Official control | Prism `compat` |
|
|
88
88
|
| --- | --- | --- |
|
|
89
|
-
| K3 / Coding `k3` | top-level `reasoning_effort
|
|
89
|
+
| K3 / Coding `k3` | top-level `reasoning_effort`: `"low"`/`"high"`/`"max"` (Open Platform default `"max"`; Kimi Code default `"high"`) | `compat.reasoning_effort` — use Task 4 family `reasoning_effort` |
|
|
90
90
|
| K2.7-code / Coding | thinking always on; Preserved Thinking always on | omit `thinking` by default; `preserveThinking: true` for replay; do not send `disabled` |
|
|
91
91
|
| K2.6 / K2.5 | `thinking.type` enabled/disabled; K2.6 optional `keep: "all"` | `compat.thinking` — Task 4 family `thinking_type` |
|
|
92
92
|
|
|
93
93
|
Per-turn `ProviderRequestOptions.compat` wins over `ModelConfig.compat`. Helpers:
|
|
94
94
|
`kimiThinking`, `kimiReasoningEffort`, `kimiPreserveThinking`.
|
|
95
|
+
`stripKimiThinkingCompat` removes provider-owned routing/serialization keys
|
|
96
|
+
(`route`, `preserveThinking`, `preserve_thinking`, thinking/effort keys) before
|
|
97
|
+
the opaque compat spread, so they never leak into wire bodies.
|
|
98
|
+
|
|
99
|
+
Featured context windows follow the official docs exactly (`262_144` for the
|
|
100
|
+
256K-class models, `1_048_576` for K3). Both stream parsers emit `done` only on
|
|
101
|
+
protocol completion evidence — Coding route: `message_stop` with all `tool_use`
|
|
102
|
+
blocks closed; Moonshot route: `[DONE]` plus a terminal `finish_reason` with no
|
|
103
|
+
dangling tool calls. Truncated streams terminate with an `error` event instead.
|
|
104
|
+
|
|
105
|
+
The Coding route authenticates with provider-owned `authorization: Bearer`,
|
|
106
|
+
`x-api-key`, and `anthropic-version: 2023-06-01` headers (official third-party
|
|
107
|
+
setup uses `ANTHROPIC_API_KEY` semantics); caller-supplied headers cannot
|
|
108
|
+
override them.
|
|
95
109
|
|
|
96
110
|
## Request/response example
|
|
97
111
|
|