@arnilo/prism 0.0.7 → 0.0.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,27 @@ All notable changes to this project will be documented in this file.
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
 
8
+ ## Unreleased
9
+
10
+ ## [0.0.8] - 2026-07-20
11
+
12
+ ### Added
13
+
14
+ - Added OpenTelemetry GenAI agent/provider/tool hierarchy, context propagation, delegation/guardrail spans, bounded trace references, and evaluation linkage.
15
+ - Added bounded evaluation trace resolution, host model judges, deterministic pairwise reports, serialized artifacts, and CI threshold assertions.
16
+ - Added MCP resources/prompts/roots/sampling/elicitation plus principal-bound Streamable HTTP sessions on pinned SDK 1.29.0, and full A2A 1.0 durable task/rich-part/reconnect/push interoperability.
17
+ - Added immutable-revision CodeQL/dependency/SBOM/license/secret/attestation release gates, weekly dependency updates, and protected bounded provider/MCP/A2A/web live canaries.
18
+ - Added optional `@arnilo/prism-web-tools` with bounded host-selected Brave/Exa search, Firecrawl Markdown/schema extraction, stable citations, late credentials, and explicit untrusted-content results.
19
+ - Added optional `createBatchedRunLedger()` with bounded FIFO/backpressure, explicit durability/flush status, terminal acknowledgement, and documented buffered crash-loss semantics.
20
+ - Added one-leaf, one-second runtime session snapshot caching with mutation/checkout/resume invalidation and reproducible network-free 0.0.8 performance evidence.
21
+ - Versioned all 31 first-party manifests and exact internal ranges to 0.0.8; no tag or publication was created.
22
+
23
+ ### Fixed
24
+
25
+ - `generateValidateReviseLoop` routes artifact parse failures through the revision budget (`metadata.reason: "parse_error"`, repairer receives `value: undefined`) instead of returning silently after one provider turn.
26
+ - `@arnilo/prism-provider-opencode-go` Anthropic route sends provider-owned `x-api-key` and `anthropic-version: 2023-06-01` headers alongside Bearer, fixing HTTP 401 on MiniMax/Qwen models; `structuredOutput: "json_schema"` is no longer inferred from OpenAI routing alone (verified models only), fixing HTTP 400 on `deepseek-v4-pro`; both stream parsers require protocol completion evidence and fail truncated streams with a terminal `error` instead of a false `done`.
27
+ - `@arnilo/prism-provider-kimi` aligns with official contracts: featured Coding `k3` defaults `reasoning_effort: "high"`, 256K-class context windows use the exact `262_144`, the featured Moonshot catalog adds `kimi-k2.7-code-highspeed`/`kimi-k2.6`/`kimi-k2.5`, routing keys (`route`, `preserve_thinking`) no longer leak into wire bodies, the Coding route sends provider-owned `x-api-key`/`anthropic-version` headers, and both stream parsers fail truncated streams instead of emitting `done`.
28
+
8
29
  ## [0.0.7] - 2026-07-19
9
30
 
10
31
  ### Added
package/README.md CHANGED
@@ -53,6 +53,7 @@ npm install @arnilo/prism-sdk @arnilo/prism-provider-openai # application profi
53
53
  npm install @arnilo/prism-all # every first-party package
54
54
  npm install @arnilo/prism-server @arnilo/prism-workflows # optional Web API boundary
55
55
  npm install @arnilo/prism-supervisor # optional local delegation + A2A 1.0
56
+ npm install @arnilo/prism-web-tools # optional bounded Brave/Exa/Firecrawl research
56
57
  ```
57
58
 
58
59
  See [docs/release-and-install.md](docs/release-and-install.md) for install
@@ -157,6 +158,7 @@ printf '{"id":"1","command":"prompt","params":{"input":"Hi"}}\n' \
157
158
  | `@arnilo/prism-mcp` | MCP client/tool bridge |
158
159
  | `@arnilo/prism-workflows` | bounded DAG workflows, durable suspend/resume, schedules/background runs, composition/state/replay, and multi-process coordination |
159
160
  | `@arnilo/prism-supervisor` | bounded local child delegation and A2A 1.0 interoperability |
161
+ | `@arnilo/prism-web-tools` | host-selected bounded Brave/Exa search and Firecrawl Markdown/schema extraction |
160
162
  | `@arnilo/prism-observability-opentelemetry` | optional OpenTelemetry adapter |
161
163
  | `@arnilo/prism-credentials-node` | encrypted-file and keychain credentials |
162
164
  | `@arnilo/prism-session-store-sqlite` | SQLite persistence/checkpoints/leases/owned run feedback |
@@ -166,7 +168,7 @@ printf '{"id":"1","command":"prompt","params":{"input":"Hi"}}\n' \
166
168
  | `@arnilo/prism-base` | profile: core + compaction + JSON Schema validation |
167
169
  | `@arnilo/prism-code` | profile: base + coding tools/security + MCP |
168
170
  | `@arnilo/prism-sdk` | profile: base + workflows + MCP + credentials + OpenTelemetry |
169
- | `@arnilo/prism-all` | every first-party package, including both persistence adapters |
171
+ | `@arnilo/prism-all` | every first-party package, including both persistence adapters and web tools |
170
172
 
171
173
  ## Scripts
172
174
 
@@ -114,12 +114,14 @@ export function generateValidateReviseLoop(opts) {
114
114
  const parsed = opts.parser
115
115
  ? await opts.parser(text, artifactCtx)
116
116
  : { ok: true, value: text };
117
- // Parse failure ends the loop silently (terminal parse errors stay on `error`).
118
- if (!parsed.ok || parsed.value === undefined)
119
- return usage;
117
+ // Parse failure consumes revision budget like a validation failure; the
118
+ // repairer receives `undefined` value plus a synthetic parse issue.
119
+ const parseFailure = !parsed.ok || parsed.value === undefined
120
+ ? { ok: false, errors: [{ path: "$", message: parsed.error ?? "artifact parse failed" }], metadata: { reason: "parse_error" } }
121
+ : undefined;
120
122
  const attempt = ++attempts;
121
123
  ctx.emit({ type: "artifact_validation_started", sessionId: ctx.sessionId, runId: ctx.runId, turn, attempt });
122
- const result = await opts.validator(parsed.value, artifactCtx);
124
+ const result = parseFailure ?? await opts.validator(parsed.value, artifactCtx);
123
125
  ctx.emit({ type: "artifact_validation_finished", sessionId: ctx.sessionId, runId: ctx.runId, turn, attempt, result });
124
126
  if (result.ok) {
125
127
  ctx.emit({ type: "artifact_finished", sessionId: ctx.sessionId, runId: ctx.runId, turn, attempt, result });
@@ -130,7 +132,7 @@ export function generateValidateReviseLoop(opts) {
130
132
  return usage;
131
133
  }
132
134
  ctx.emit({ type: "artifact_revision_started", sessionId: ctx.sessionId, runId: ctx.runId, turn, attempt, failure: result });
133
- const repair = await repairer(parsed.value, result, artifactCtx);
135
+ const repair = await repairer(parseFailure ? undefined : parsed.value, result, artifactCtx);
134
136
  const repairMessages = inputMessages(repair).map((m) => ({ ...m, id: randomId("msg") }));
135
137
  for (const message of repairMessages)
136
138
  await ctx.appendMessage(message);
package/dist/agents.js CHANGED
@@ -12,6 +12,7 @@ import { errorToErrorInfo, redactAgentEvent, redactProviderRequest, redactRunLed
12
12
  import { composeSystemPrompt, mergeSystemPromptConfig } from "./system-prompts.js";
13
13
  import { createDefaultRetryPolicy, waitForRetry } from "./retry.js";
14
14
  import { createMemorySessionStore, createSessionEntry, getSessionBranchEntries, rebuildSessionContext } from "./session-stores.js";
15
+ import { isFlushableRunLedger } from "./run-ledger.js";
15
16
  import { createToolRegistry, dispatchToolCall } from "./tools.js";
16
17
  import { RunLimitError, RunLimitTracker, resolveRunLimits } from "./run-limits.js";
17
18
  import { agentFingerprint, initialAgentRunState, loadAgentRunState, publicState, saveAgentRunState, validateRunStateOptions } from "./agent-run-state.js";
@@ -103,6 +104,8 @@ class RuntimeAgentSession {
103
104
  activeLoopTurn = 1;
104
105
  ledgerChain = Promise.resolve();
105
106
  ledgerFailure;
107
+ snapshotGeneration = 0;
108
+ snapshotCache;
106
109
  constructor(config) {
107
110
  this.id = config.id ?? randomId("session");
108
111
  this.agent = config.agent;
@@ -168,6 +171,8 @@ class RuntimeAgentSession {
168
171
  this.activeIdempotencyKey = options.idempotencyKey ?? this.agent.config.idempotencyKey;
169
172
  this.activeGuardrails = mergeGuardrails(this.agent.config.guardrails, options.guardrails);
170
173
  this.activeDurable = resumed ?? (durableOptions ? { options: durableOptions, version: 0 } : undefined);
174
+ if (resumed)
175
+ this.invalidateSnapshot();
171
176
  const model = options.model ?? this.agent.config.model;
172
177
  const startedAt = new Date().toISOString();
173
178
  let runError;
@@ -442,6 +447,8 @@ class RuntimeAgentSession {
442
447
  ...this.activeOwnership,
443
448
  };
444
449
  await this.activeLedger.appendRun(redactRunLedgerRecord(finishRecord, this.activeRedactor));
450
+ if (isFlushableRunLedger(this.activeLedger) && this.activeLedger.durability === "flush_on_terminal")
451
+ await this.activeLedger.flush();
445
452
  }
446
453
  }
447
454
  finally {
@@ -570,6 +577,7 @@ class RuntimeAgentSession {
570
577
  }
571
578
  async checkout(leafId) {
572
579
  this.currentLeafId = leafId;
580
+ this.invalidateSnapshot();
573
581
  await this.rebuildHistory();
574
582
  }
575
583
  fork(options = {}) {
@@ -853,6 +861,11 @@ class RuntimeAgentSession {
853
861
  idempotencyKey: this.activeIdempotencyKey,
854
862
  });
855
863
  this.currentLeafId = redacted.id;
864
+ this.invalidateSnapshot();
865
+ }
866
+ invalidateSnapshot() {
867
+ this.snapshotGeneration += 1;
868
+ this.snapshotCache = undefined;
856
869
  }
857
870
  redact(value) {
858
871
  return this.activeRedactor?.redact(value) ?? value;
@@ -864,10 +877,16 @@ class RuntimeAgentSession {
864
877
  this.history = (await this.snapshot()).messages.slice();
865
878
  }
866
879
  async snapshot() {
880
+ const now = performance.now();
881
+ const cached = this.snapshotCache;
882
+ if (cached && cached.leafId === this.currentLeafId && cached.generation === this.snapshotGeneration && cached.expiresAt > now)
883
+ return cached.value;
867
884
  const reader = this.branchReader();
868
- return reader
869
- ? rebuildSessionContext(reader, { sessionId: this.id, leafId: this.currentLeafId })
885
+ const value = reader
886
+ ? await rebuildSessionContext(reader, { sessionId: this.id, leafId: this.currentLeafId })
870
887
  : rebuildSessionContext(await this.store.list(this.id), { leafId: this.currentLeafId });
888
+ this.snapshotCache = { leafId: this.currentLeafId, generation: this.snapshotGeneration, expiresAt: now + 1_000, value };
889
+ return value;
871
890
  }
872
891
  }
873
892
  class EventSubscriber {
@@ -1278,6 +1278,21 @@ export interface RunLedger {
1278
1278
  }
1279
1279
  /** Union of records that may be handed to a {@link RunLedger}. */
1280
1280
  export type RunLedgerRecord = RunRecord | AgentEventRecord | ToolCallRecord | UsageRecord;
1281
+ export type RunLedgerDurability = "write_through" | "flush_on_terminal" | "buffered";
1282
+ export interface RunLedgerFlushResult {
1283
+ readonly accepted: number;
1284
+ readonly flushed: number;
1285
+ readonly buffered: number;
1286
+ }
1287
+ /** Optional durability seam implemented by bounded ledger adapters. */
1288
+ export interface FlushableRunLedger extends RunLedger {
1289
+ readonly durability: RunLedgerDurability;
1290
+ flush(): Promise<RunLedgerFlushResult>;
1291
+ status(): RunLedgerFlushResult;
1292
+ dispose(options?: {
1293
+ readonly flush?: boolean;
1294
+ }): Promise<void>;
1295
+ }
1281
1296
  /** Immutable human feedback linked to an existing owned run/trace and optional evaluations. */
1282
1297
  export interface RunFeedbackRecord extends OwnershipScope {
1283
1298
  readonly id: string;
package/dist/index.d.ts CHANGED
@@ -2,6 +2,8 @@ export type * from "./contracts.js";
2
2
  export type { RunLimitCounters, RunLimitName, SecureAgentOptions } from "./contracts.js";
3
3
  export { isSessionEntryKind, SESSION_APPEND_CONFLICT_CODE, SESSION_ENTRY_KINDS, SESSION_ENTRY_SCHEMA_VERSION, SessionAppendConflictError, isSessionAppendConflict, AgentRunError, AgentRunStateError } from "./contracts.js";
4
4
  export { createAgent, createAgentSession, resumeAgentRun } from "./agents.js";
5
+ export { createBatchedRunLedger, isFlushableRunLedger, DEFAULT_LEDGER_BATCH_ENTRIES, HARD_LEDGER_BATCH_ENTRIES, DEFAULT_LEDGER_BATCH_BYTES, HARD_LEDGER_BATCH_BYTES, DEFAULT_LEDGER_BATCH_DELAY_MS, HARD_LEDGER_BATCH_DELAY_MS, } from "./run-ledger.js";
6
+ export type { BatchedRunLedgerOptions } from "./run-ledger.js";
5
7
  export { createSecureAgent } from "./secure-agent.js";
6
8
  export { createMemoryRunFeedbackStore, prepareRunFeedback, requireRunFeedbackOwnership, runFeedbackPageLimit, RunFeedbackError, } from "./feedback.js";
7
9
  export type { MemoryRunFeedbackStoreOptions, PrepareRunFeedbackOptions, RunFeedbackLimits, RunFeedbackRun, RunFeedbackRunResolver, } from "./feedback.js";
@@ -81,5 +83,5 @@ export type { DispatchToolCallOptions, ToolArgumentValidationError, ToolArgument
81
83
  export type { DuplicateRegistrationOptions, DuplicateRegistrationPolicy } from "./registry-options.js";
82
84
  export { dispatchToolCallsInOrder, generateValidateReviseLoop, isAgentLoopOptions, resolveLoop, resolveToolConcurrency, singleShotLoop } from "./agent-loops.js";
83
85
  export declare const name = "prism";
84
- export declare const version = "0.0.7";
86
+ export declare const version = "0.0.8";
85
87
  export declare const description = "Agent harness for AI providers, agents, sessions, and tools.";
package/dist/index.js CHANGED
@@ -1,5 +1,6 @@
1
1
  export { isSessionEntryKind, SESSION_APPEND_CONFLICT_CODE, SESSION_ENTRY_KINDS, SESSION_ENTRY_SCHEMA_VERSION, SessionAppendConflictError, isSessionAppendConflict, AgentRunError, AgentRunStateError } from "./contracts.js";
2
2
  export { createAgent, createAgentSession, resumeAgentRun } from "./agents.js";
3
+ export { createBatchedRunLedger, isFlushableRunLedger, DEFAULT_LEDGER_BATCH_ENTRIES, HARD_LEDGER_BATCH_ENTRIES, DEFAULT_LEDGER_BATCH_BYTES, HARD_LEDGER_BATCH_BYTES, DEFAULT_LEDGER_BATCH_DELAY_MS, HARD_LEDGER_BATCH_DELAY_MS, } from "./run-ledger.js";
3
4
  export { createSecureAgent } from "./secure-agent.js";
4
5
  export { createMemoryRunFeedbackStore, prepareRunFeedback, requireRunFeedbackOwnership, runFeedbackPageLimit, RunFeedbackError, } from "./feedback.js";
5
6
  export { CHECKPOINT_CONFLICT_CODE, CheckpointConflictError, createMemoryCheckpointStore } from "./checkpoints.js";
@@ -44,6 +45,6 @@ export { assertGuardrailsAllowed, GuardrailError, MAX_GUARDRAIL_CONCURRENCY, run
44
45
  export { createRunLimitTracker, DEFAULT_RUN_LIMITS, HARD_MAX_RUN_COST, HARD_RUN_LIMITS, RunLimitError, RunLimitTracker, resolveRunLimits } from "./run-limits.js";
45
46
  export { dispatchToolCallsInOrder, generateValidateReviseLoop, isAgentLoopOptions, resolveLoop, resolveToolConcurrency, singleShotLoop } from "./agent-loops.js";
46
47
  export const name = "prism";
47
- export const version = "0.0.7";
48
+ export const version = "0.0.8";
48
49
  export const description = "Agent harness for AI providers, agents, sessions, and tools.";
49
50
  //# sourceMappingURL=index.js.map
@@ -0,0 +1,21 @@
1
+ import type { FlushableRunLedger, RunLedger, RunLedgerDurability } from "./contracts.js";
2
+ export declare const DEFAULT_LEDGER_BATCH_ENTRIES = 128;
3
+ export declare const HARD_LEDGER_BATCH_ENTRIES = 4096;
4
+ export declare const DEFAULT_LEDGER_BATCH_BYTES: number;
5
+ export declare const HARD_LEDGER_BATCH_BYTES: number;
6
+ export declare const DEFAULT_LEDGER_BATCH_DELAY_MS = 25;
7
+ export declare const HARD_LEDGER_BATCH_DELAY_MS = 60000;
8
+ export interface BatchedRunLedgerOptions {
9
+ readonly maxBatchEntries?: number;
10
+ readonly maxBatchBytes?: number;
11
+ readonly maxBufferedEntries?: number;
12
+ readonly maxBufferedBytes?: number;
13
+ readonly maxDelayMs?: number;
14
+ readonly durability?: RunLedgerDurability;
15
+ }
16
+ /**
17
+ * Wrap any RunLedger with one bounded FIFO. Inputs must already be redacted, as required by RunLedger.
18
+ * `buffered` may lose accepted records on process crash; call `flush()` for acknowledgement.
19
+ */
20
+ export declare function createBatchedRunLedger(target: RunLedger, options?: BatchedRunLedgerOptions): FlushableRunLedger;
21
+ export declare function isFlushableRunLedger(ledger: RunLedger): ledger is FlushableRunLedger;
@@ -0,0 +1,115 @@
1
+ export const DEFAULT_LEDGER_BATCH_ENTRIES = 128;
2
+ export const HARD_LEDGER_BATCH_ENTRIES = 4096;
3
+ export const DEFAULT_LEDGER_BATCH_BYTES = 512 * 1024;
4
+ export const HARD_LEDGER_BATCH_BYTES = 8 * 1024 * 1024;
5
+ export const DEFAULT_LEDGER_BATCH_DELAY_MS = 25;
6
+ export const HARD_LEDGER_BATCH_DELAY_MS = 60_000;
7
+ function integer(value, fallback, hard, name) {
8
+ const selected = value ?? fallback;
9
+ if (!Number.isInteger(selected) || selected < 1 || selected > hard)
10
+ throw new RangeError(`${name} must be an integer in [1, ${hard}]`);
11
+ return selected;
12
+ }
13
+ function terminal(record) {
14
+ return record.status !== undefined && record.status !== "queued" && record.status !== "running";
15
+ }
16
+ /**
17
+ * Wrap any RunLedger with one bounded FIFO. Inputs must already be redacted, as required by RunLedger.
18
+ * `buffered` may lose accepted records on process crash; call `flush()` for acknowledgement.
19
+ */
20
+ export function createBatchedRunLedger(target, options = {}) {
21
+ const maxBatchEntries = integer(options.maxBatchEntries, DEFAULT_LEDGER_BATCH_ENTRIES, HARD_LEDGER_BATCH_ENTRIES, "maxBatchEntries");
22
+ const maxBatchBytes = integer(options.maxBatchBytes, DEFAULT_LEDGER_BATCH_BYTES, HARD_LEDGER_BATCH_BYTES, "maxBatchBytes");
23
+ const maxBufferedEntries = integer(options.maxBufferedEntries, Math.min(HARD_LEDGER_BATCH_ENTRIES, maxBatchEntries * 2), HARD_LEDGER_BATCH_ENTRIES, "maxBufferedEntries");
24
+ const maxBufferedBytes = integer(options.maxBufferedBytes, Math.min(HARD_LEDGER_BATCH_BYTES, maxBatchBytes * 2), HARD_LEDGER_BATCH_BYTES, "maxBufferedBytes");
25
+ const maxDelayMs = integer(options.maxDelayMs, DEFAULT_LEDGER_BATCH_DELAY_MS, HARD_LEDGER_BATCH_DELAY_MS, "maxDelayMs");
26
+ const durability = options.durability ?? "flush_on_terminal";
27
+ const queue = [];
28
+ let bufferedBytes = 0;
29
+ let accepted = 0;
30
+ let flushed = 0;
31
+ let timer;
32
+ let flushChain = Promise.resolve();
33
+ let disposed = false;
34
+ const status = () => ({ accepted, flushed, buffered: queue.length });
35
+ const cancelTimer = () => { if (timer)
36
+ clearTimeout(timer); timer = undefined; };
37
+ const schedule = () => {
38
+ if (timer || disposed || queue.length === 0)
39
+ return;
40
+ timer = setTimeout(() => { timer = undefined; void flush().catch(() => undefined); }, maxDelayMs);
41
+ timer.unref?.();
42
+ };
43
+ const write = (item) => {
44
+ if (item.kind === "run")
45
+ return target.appendRun(item.record);
46
+ if (item.kind === "event")
47
+ return target.appendEvent(item.record);
48
+ if (item.kind === "tool")
49
+ return target.appendToolCall(item.record);
50
+ return target.appendUsage(item.record);
51
+ };
52
+ const flush = () => {
53
+ cancelTimer();
54
+ const operation = flushChain.then(async () => {
55
+ let entries = 0;
56
+ let bytes = 0;
57
+ while (queue.length) {
58
+ const item = queue[0];
59
+ if (entries && (entries >= maxBatchEntries || bytes + item.bytes > maxBatchBytes)) {
60
+ entries = 0;
61
+ bytes = 0;
62
+ }
63
+ await write(item);
64
+ queue.shift();
65
+ bufferedBytes -= item.bytes;
66
+ flushed += 1;
67
+ entries += 1;
68
+ bytes += item.bytes;
69
+ }
70
+ return status();
71
+ });
72
+ flushChain = operation.then(() => undefined, () => undefined);
73
+ return operation;
74
+ };
75
+ const enqueue = async (item) => {
76
+ if (disposed)
77
+ throw new Error("batched run ledger is disposed");
78
+ const bytes = Buffer.byteLength(JSON.stringify(item.record));
79
+ if (bytes > maxBatchBytes || bytes > maxBufferedBytes)
80
+ throw new RangeError("run ledger record exceeds byte limit");
81
+ if (queue.length >= maxBufferedEntries || bufferedBytes + bytes > maxBufferedBytes)
82
+ await flush();
83
+ queue.push({ ...item, bytes });
84
+ bufferedBytes += bytes;
85
+ accepted += 1;
86
+ if (durability === "write_through" || queue.length >= maxBatchEntries || bufferedBytes >= maxBatchBytes || (item.kind === "run" && terminal(item.record) && durability === "flush_on_terminal"))
87
+ await flush();
88
+ else
89
+ schedule();
90
+ };
91
+ return {
92
+ durability,
93
+ appendRun: (record) => enqueue({ kind: "run", record }),
94
+ appendEvent: (record) => enqueue({ kind: "event", record }),
95
+ appendToolCall: (record) => enqueue({ kind: "tool", record }),
96
+ appendUsage: (record) => enqueue({ kind: "usage", record }),
97
+ flush,
98
+ status,
99
+ async dispose(disposeOptions = {}) {
100
+ disposed = true;
101
+ cancelTimer();
102
+ if (disposeOptions.flush !== false)
103
+ await flush();
104
+ else {
105
+ await flushChain;
106
+ queue.length = 0;
107
+ bufferedBytes = 0;
108
+ }
109
+ },
110
+ };
111
+ }
112
+ export function isFlushableRunLedger(ledger) {
113
+ return "flush" in ledger && typeof ledger.flush === "function";
114
+ }
115
+ //# sourceMappingURL=run-ledger.js.map
package/docs/a2a.md CHANGED
@@ -1,75 +1,94 @@
1
- # A2A interoperability
1
+ # A2A 1.0 interoperability
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-supervisor` implements a bounded text-only subset of Agent2Agent (A2A) protocol 1.0: Agent Cards, JSON-RPC `SendMessage`, `SendStreamingMessage`, `GetExtendedAgentCard`, SSE task updates, ES256 JWS card signatures, and an explicit remote client.
5
+ `@arnilo/prism-supervisor` implements bounded A2A 1.0 over the JSON-RPC/HTTPS binding. Supported operations: `SendMessage`, `SendStreamingMessage`, `GetTask`, `ListTasks`, `CancelTask`, `SubscribeToTask`, push-notification-config create/get/list/delete, and `GetExtendedAgentCard`. Agent Cards retain explicit ES256 verification. gRPC, HTTP+JSON, discovery registries, automatic JWK/OAuth fetching, and an internal task worker/store are absent.
6
6
 
7
7
  ## When to use it
8
8
 
9
- Use it to expose one explicitly selected Prism agent at an A2A endpoint or call a known remote A2A agent. Do not use it as endpoint discovery, a generic proxy, credential forwarding, or a replacement for local workflows.
9
+ Use it to expose a selected Prism agent or host-owned durable agent/workflow lifecycle to known A2A peers. Use direct `exposure` for backward-compatible text invocation. Supply `tasks` for durable/rich/reconnect operations and `push` only when host persistence and webhook delivery policy already exist.
10
10
 
11
11
  ## Inputs / request
12
12
 
13
- | API/field | Meaning |
14
- | --- | --- |
15
- | `createA2AAgentCard(card)` | Validates/freeze a JSONRPC protocol-1.0 HTTPS text card. |
16
- | `signA2AAgentCard(card, { privateKey, keyId, expiresAt })` | Adds detached-payload ES256 JWS signature using WebCrypto. |
17
- | `verifyA2AAgentCard(card, { publicKey, keyId?, now?, maxAgeMs? })` | Pins ES256/key/expiry and verifies canonical unsigned card. |
18
- | `createA2AHandler({ card, exposure, authorize })` | Web-standard card/JSON-RPC/SSE `Request` to `Response` handler. |
19
- | `createA2AClient({ endpoint, allowedOrigins })` | Explicit HTTPS remote client with optional card verifier/auth callback. |
20
- | `A2ALimits` | Request 64 KiB, response 1 MiB, event 64 KiB, stream 10 MiB/10k events, concurrency 16, timeout 120s, card 64 KiB defaults; finite hard caps apply. |
13
+ ```ts
14
+ const handler = createA2AHandler({
15
+ card,
16
+ exposure: { sessionFactory }, // text fallback
17
+ authorize: authenticateEveryOperation,
18
+ tasks: durableTaskAdapter, // host-owned start/get/list/cancel/subscribe
19
+ push: pushConfigAdapter, // host-owned config persistence/delivery integration
20
+ parts: {
21
+ allowRaw: true,
22
+ allowData: true,
23
+ allowUrl: true,
24
+ validateUrl: validatePinnedPublicHttpsUrl, // validation only; never fetched
25
+ },
26
+ });
27
+ ```
21
28
 
22
- ## Outputs / response / events
29
+ `A2ATaskLifecycle` receives validated messages, exact `A2AAuthorization`, abort signals, bounded pagination, and reconnect cursor. Adapter must map existing durable agent/workflow/checkpoint/persistence operations; Prism creates no worker, queue, task map, or database table. Unknown-owner task/config lookups return `undefined`, producing non-disclosing `TaskNotFoundError` (`-32001`). Missing task/push capability returns `UnsupportedOperationError` (`-32004`).
23
30
 
24
- The handler serves `GET /.well-known/agent-card.json` and its configured POST endpoint. JSON-RPC returns `{ result: { task } }` or a bounded error. Streaming returns backpressure-driven SSE task envelopes. Client `send()` maps a terminal remote task to `AgentRunResult`; `stream()` incrementally yields validated/redacted text artifacts. Client SSE accepts LF, CRLF, mixed blank-line separators, comments/unknown fields, and multiline `data:` joined with LF.
31
+ `A2APart` is an exact one-of:
25
32
 
26
- ## Request/response example
33
+ | Part | Default | Rule |
34
+ | --- | --- | --- |
35
+ | `{ text }` | enabled | bounded UTF-8 text |
36
+ | `{ raw, mediaType?, filename? }` | disabled | strict base64 and decoded-byte cap |
37
+ | `{ data }` | disabled | bounded finite JSON, depth 64/properties 10,000 |
38
+ | `{ url, mediaType?, filename? }` | disabled | credential/fragment-free HTTPS plus required host URL policy; never dereferenced |
27
39
 
28
- ```json
29
- {"jsonrpc":"2.0","id":1,"method":"SendMessage","params":{"message":{"role":"user","messageId":"m1","parts":[{"text":"Check sources"}]}}}
30
- ```
40
+ Parts, messages, artifacts, histories, metadata, and aggregate responses are untrusted. Rich content remains in A2A task/message/artifact contracts for host mapping; it is never promoted to system instructions or automatically loaded as a Prism resource.
31
41
 
32
42
  ## Implementation example
33
43
 
34
44
  ```ts
35
- import { createA2AClient, createA2AHandler, verifyA2AAgentCard } from "@arnilo/prism-supervisor";
36
-
37
- const handler = createA2AHandler({
38
- card,
39
- exposure: { sessionFactory: ({ ownership }) => agent.createSession({ metadata: ownership }) },
40
- authorize: ({ request }) => authenticate(request),
41
- });
42
-
43
45
  const client = createA2AClient({
44
46
  endpoint: "https://agent.example/a2a/v1",
45
47
  allowedOrigins: ["https://agent.example"],
46
- authorize: () => ({ authorization: `Bearer ${resolveOwnedToken()}` }),
47
- verifyCard: (remoteCard) => verifyA2AAgentCard(remoteCard, { publicKey, keyId: "agent-key" }),
48
+ authorize: ownedAuthHeaders,
49
+ verifyCard: (card) => verifyA2AAgentCard(card, { publicKey, keyId: "agent-key" }),
48
50
  });
51
+ const task = await client.getTask("task-1");
52
+ for await (const event of client.subscribeToTask(task.id, { afterEventId: savedCursor })) persistCursor(event.eventId);
53
+ ```
54
+
55
+ ## Outputs / response / events
56
+
57
+ Streams use ordered SSE frames with `id:` and JSON-RPC `result` containing one `A2ATaskEvent`: full `task`, `statusUpdate`, or `artifactUpdate`. `SubscribeToTask({ id, afterEventId })` passes cursor to durable adapter for authorized bounded replay. Duplicate event IDs are rejected/server-bounded; client de-duplicates repeated IDs. Terminal, `INPUT_REQUIRED`, and `AUTH_REQUIRED` states close streams. String-oriented `client.stream()` reports interrupted states as `ERR_PRISM_A2A_INTERRUPTED`; task APIs preserve status for continuation.
49
58
 
50
- const result = await client.send("Check sources");
59
+ Client APIs:
60
+
61
+ - `send()` / `stream()` preserve text-to-`AgentRunResult` compatibility.
62
+ - `sendMessage()` returns rich/durable `A2ATask`.
63
+ - `getTask()`, `listTasks()`, `cancelTask()`, `subscribeToTask()` operate on durable tasks.
64
+ - `createPushConfig()`, `getPushConfig()`, `listPushConfigs()`, `deletePushConfig()` expose declared push config operations.
65
+
66
+ Every protocol request sends/negotiates `A2A-Version: 1.0`. Client endpoint/card URLs require exact allow-listed HTTPS and `redirect: "error"`. Cards are parsed then optionally verified against host-pinned keys; no key URL is fetched.
67
+
68
+ ## Request/response example
69
+
70
+ ```json
71
+ {"jsonrpc":"2.0","id":1,"method":"SubscribeToTask","params":{"id":"task-1","afterEventId":"event-42"}}
51
72
  ```
52
73
 
53
74
  ## Extension and configuration notes
54
75
 
55
- The package owns no listener or credential store. Mount the handler in a host server and resolve authentication/authorization on every request. Client auth executes only after body/card validation and serialization. Injectable `fetch` supports host transports/tests; redirects are disabled.
76
+ Handler requires `card.capabilities.pushNotifications` to exactly match supplied `push`; mismatch fails construction, preserving signed-card integrity and preventing false capability claims. Streaming remains available for direct text invocation. Push adapter owns exact-owner persistence, signing/auth credentials, and network transport. Host explicitly calls `deliverA2APushEvent()` from its durable update path; helper bounds event, timeout (10s default/60s hard), attempts (1 default/3 hard), and passes stable event ID as idempotency key to host `A2APushDelivery`. It starts no hidden sender and performs no network itself. Config handling validates IDs/count/bytes and requires same explicit URL policy used for URL parts. Returned push configs omit token and authentication credentials.
56
77
 
57
- Only `text` parts are accepted. File/data parts, push notifications, task persistence/query/cancel, gRPC, HTTP+JSON binding, automatic JWK fetching, and endpoint discovery are intentionally absent.
78
+ Defaults/hard caps include: request 64 KiB/1 MiB; response 1/8 MiB; event 64 KiB/1 MiB; stream 10/64 MiB and 10k/100k events; replay 1k/10k events; concurrency 16/256; timeout 120s/30m; IDs 256/4096 B; parts 32/256; part/raw 1/8 MiB; data 256 KiB/4 MiB; artifacts 32/256; history/page 100/1000; cursor 4/16 KiB; push configs 10/100. Hosts may narrow limits.
58
79
 
59
80
  ## Security and performance notes
60
81
 
61
- - Endpoints and card URLs must be HTTPS and exactly origin-allow-listed before fetch; `redirect: "error"` prevents redirect SSRF.
62
- - Treat every remote card, error, task, status, artifact, and SSE frame as untrusted. Shape/count/byte/time limits apply before mapping. Streaming keeps raw stream bytes, current frame bytes, and event count as separate existing limits.
63
- - One fatal streaming UTF-8 decoder is reused across every body chunk and flushed once at EOF. Split multibyte code points are preserved; malformed/truncated UTF-8 fails rather than inserting `U+FFFD` into JSON. A small coalesced line buffer keeps one-byte chunk handling incremental.
64
- - SSE frames require a terminating blank line. A non-whitespace final partial frame, malformed JSON, missing terminal task, failed/canceled task, or any event after a completed task fails with bounded package-owned text. Existing request/response/event/stream/count/timeout hard caps are unchanged.
65
- - Card verification pins `alg=ES256`, optional key ID, issue/expiry, optional maximum age, and canonical unsigned-card payload. Hosts provision trusted public keys; remote `jku` is never fetched automatically.
66
- - Card discovery is public; extended-card and invoke methods call host authorization. Use TLS, rate limits, and replay controls at the host edge.
67
- - Credentials remain in the client auth callback or server authorizer and never enter cards, messages, events, or metrics.
68
- - Offline conformance is authoritative. Live endpoints are optional operator smoke tests.
82
+ - Authorize every operation; lifecycle/push adapters enforce exact owner again at durable storage boundary. Missing and foreign tasks/configs share `-32001`.
83
+ - URL policy must reject private, loopback, link-local, rebound, redirected, or otherwise disallowed destinations. Package never fetches file URLs. Host push delivery must repeat equivalent checks for every attempt/redirect and process event IDs idempotently.
84
+ - Push token/auth credentials are accepted only into host adapter input and removed from protocol reads/responses. Keep them out of task parts, events, telemetry, ledgers, and errors.
85
+ - Known-secret redaction applies before handler JSON/SSE output. Client redacts mapped text/errors. Raw/data/url content remains explicitly untrusted.
86
+ - Canceled/closed streams abort adapter signal, return iterator, clear timeout, and release concurrency slot. Task/push durability and replay retention belong to host adapter and must remain finite.
87
+ - Default tests use in-memory lifecycle/fake fetch only; no public network.
69
88
 
70
89
  ## Related APIs
71
90
 
72
- - [Supervisor delegation](supervisors.md): local child boundary.
73
- - [Web-standard server](server.md): non-A2A Prism routes.
74
- - [Host security](host-security.md): authentication, SSRF, and untrusted-output policy.
75
- - [Agent/session runtime](agent-session-runtime.md): mapped local execution/result.
91
+ - [Supervisor delegation](supervisors.md)
92
+ - [Agent/session runtime](agent-session-runtime.md)
93
+ - [Workflows](workflows.md)
94
+ - [Host security](host-security.md)
@@ -112,7 +112,7 @@ Artifact validation/refinement events (emitted only by `generateValidateReviseLo
112
112
  | `artifact_validation_finished` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` |
113
113
  | `artifact_revision_started` | `sessionId`, `runId`, `turn`, `attempt`, `failure: ArtifactValidation` |
114
114
  | `artifact_finished` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (loop ended successfully) |
115
- | `artifact_failed` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (candidate budget exhausted, or `result.metadata.reason === "tool_round_limit"`) |
115
+ | `artifact_failed` | `sessionId`, `runId`, `turn`, `attempt`, `result: ArtifactValidation` (candidate budget exhausted, `result.metadata.reason === "tool_round_limit"`, or `result.metadata.reason === "parse_error"` when the budget was consumed by artifact parse failures) |
116
116
 
117
117
  ### Artifact event ordering
118
118
 
@@ -208,6 +208,6 @@ for await (const event of session.stream("draft", { loop: { strategy: "generate-
208
208
  - [Agent loops](agent-loops.md): `singleShotLoop` and `generateValidateReviseLoop` emit the artifact events.
209
209
  - [Structured output](structured-output.md): `ArtifactValidation` shape threaded through parser/validator/repairer.
210
210
  - [Public contracts](public-contracts.md): full `AgentEvent` union and `ArtifactValidation` contract.
211
- - [Observability](observability.md): `ProviderTurnMetadata`, OpenTelemetry adapter package.
211
+ - [Observability](observability.md): `ProviderTurnMetadata`; optional adapter builds one parented GenAI span tree from metadata-only lifecycle events and ignores message/progress deltas.
212
212
  - [Tools](tools.md): `tool_execution_*` variants.
213
213
  - [Compaction and retry policies](compaction-and-retry.md): `compaction_*` and `retry_scheduled` variants.
@@ -89,7 +89,7 @@ Host callback contracts (all generic over host `T`):
89
89
 
90
90
  | Contract | Shape |
91
91
  | --- | --- |
92
- | `ArtifactParser<T>` | `(text: string, ctx: ArtifactContext) => ArtifactParseResult<T> \| Promise<...>` — parse assistant text to a typed value. |
92
+ | `ArtifactParser<T>` | `(text: string, ctx: ArtifactContext) => ArtifactParseResult<T> \| Promise<...>` — parse assistant text to a typed value. A parse failure (`ok: false` or missing `value`) consumes revision budget exactly like a validation failure: the repairer receives `value: undefined` plus a synthetic failure (`errors[0].message` = the parse error, `metadata.reason: "parse_error"`), and budget exhaustion ends with terminal `artifact_failed`. |
93
93
  | `ArtifactValidator<T>` | `(value: T, ctx: ArtifactContext) => ArtifactValidation \| Promise<...>` — return `{ ok: true }` or `{ ok: false, errors }`. |
94
94
  | `ArtifactRepairer<T>` | `(value: T \| undefined, failure: ArtifactValidation, ctx: ArtifactContext) => AgentInput \| Promise<...>` — build the revision follow-up input. |
95
95
  | `ArtifactValidation` | `{ ok: boolean; errors?: readonly { path?: string; message: string }[]; metadata?: ... }`. |
@@ -204,6 +204,7 @@ Per-run options may narrow `limits` and append `guardrails`; they cannot replace
204
204
  - [Middleware hooks](middleware-hooks.md): hooks that configured assembly/runtime can run.
205
205
  - [CLI/RPC](cli-rpc.md): terminal and JSONL adapters over this runtime.
206
206
  - [Workflows](workflows.md): optional DAG orchestration that calls `AgentSession.run()` for agent nodes.
207
+ - [A2A interoperability](a2a.md): direct text exposure calls `AgentSession.run()`; durable/rich/reconnect behavior uses host `A2ATaskLifecycle` over existing checkpoints/persistence, never an in-memory runtime cache.
207
208
 
208
209
  `AgentConfig.loop` and `RunOptions.loop` select a replaceable per-run control loop (`singleShotLoop` default, or `generate-validate-revise` with host callbacks); see [Agent loops](agent-loops.md). `RunOptions.loop` wins over `AgentConfig.loop`. Built-in loops emit the same normal turn/message envelope around provider turns, and both add the first run input to live history once after the first provider turn so later turns see the same transcript shape.
209
210
 
@@ -218,9 +218,18 @@ const providers = createOpenAIProviderPackage({ apiKey });
218
218
  - Never log passphrases, derived keys, or decrypted credential payloads.
219
219
  - Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
220
220
 
221
+ ## MCP authentication boundary
222
+
223
+ MCP credentials remain host inputs: resolve them before constructing client `requestInit` or inside server `resolveAuthInfo`. Stateful server `resolveIdentity` receives validated SDK auth metadata only to derive a stable non-secret principal ID. Never copy access/refresh tokens into MCP resource/prompt/sampling/elicitation payloads, telemetry, errors, or session bindings; Prism does not refresh or persist MCP OAuth automatically.
224
+
225
+ ## Web adapter credential boundary
226
+
227
+ `@arnilo/prism-web-tools` accepts explicit callbacks or `CredentialResolver`. Brave resolves `subscription_token`; Exa and Firecrawl resolve `api_key` immediately before each fixed-origin request. Keys never enter tool arguments/results, URLs, provider metadata, errors, telemetry, or prompts. Use separate least-privilege credentials and do not forward MCP/provider tokens between adapters.
228
+
221
229
  ## Related APIs
222
230
 
223
231
  - [Credentials and redaction](credentials-and-redaction.md): core resolver helpers and `refreshOAuthCredential()`
232
+ - [Web search, fetch, and extraction](web-tools.md): late-bound Brave/Exa/Firecrawl credentials
224
233
  - [Security/auth/trust](settings-auth-trust-security.md): host-owned settings/credentials boundaries
225
234
  - [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 threat model and conformance matrix rows 7–10
226
235
  - `@arnilo/prism`: `CredentialResolver`, `OAuthCredentialStore`, `createMemoryCredentialStore()`
@@ -241,7 +241,7 @@ Minimum production guidance:
241
241
 
242
242
  - **Branch context:** implement `SessionStore.readBranchPath(query)` with an ancestor query / recursive CTE. Treat `SessionStore.list(sessionId)` as an O(n) development fallback only.
243
243
  - **Cursor pagination:** every `query*` method should honor `cursor`, `limit`, and `order`. Encode cursors from indexed columns such as `(timestamp, id)`, `(started_at, id)`, `(recorded_at, id)`, or `(run_id, sequence)`; never use offset pagination for long sessions.
244
- - **Batch appends:** `SessionStore.append()` is single-entry because the runtime advances one branch leaf at a time. Hosts may batch inside their DB/ledger adapters for `RunLedger` rows, but the adapter must preserve per-run event order and must not acknowledge writes before durable enqueue/commit.
244
+ - **Batch appends:** `SessionStore.append()` stays single-entry because runtime advances one branch leaf at a time. Optional `createBatchedRunLedger()` wraps any ledger with bounded FIFO/backpressure and explicit `write_through`, `flush_on_terminal`, or crash-loss-capable `buffered` acknowledgement semantics; SQLite/PostgreSQL defaults remain direct durable writes.
245
245
  - **Event sequence allocation:** allocate a monotonic `sequence` per `run_id` when inserting `prism_agent_events`. Use it with `run_id` for stable event timeline pagination when timestamps collide.
246
246
  - **Run/event/usage query shapes:** runs page by `(session_id, started_at, id)` or `(branch_id, started_at, id)`; events page by `(run_id, sequence)` or `(session_id, timestamp, id)`; usage pages by `(run_id, recorded_at, id)` or `(session_id, recorded_at, id)`.
247
247
  - **Host-owned sizing:** hosts own connection pools, transaction timeouts, page-size caps, queue/batch size, retention jobs, partitioning, and tenant/account/user isolation. Prism does not guess production limits.