@arnilo/prism 0.0.3 → 0.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/CHANGELOG.md +40 -0
  2. package/README.md +62 -26
  3. package/dist/agent-loops.d.ts +8 -1
  4. package/dist/agent-loops.js +57 -11
  5. package/dist/agents.js +212 -32
  6. package/dist/checkpoints.d.ts +11 -0
  7. package/dist/checkpoints.js +144 -0
  8. package/dist/cli-init.d.ts +41 -0
  9. package/dist/cli-init.js +390 -0
  10. package/dist/cli-runner.d.ts +7 -1
  11. package/dist/cli-runner.js +13 -1
  12. package/dist/compaction.js +9 -1
  13. package/dist/content.d.ts +121 -0
  14. package/dist/content.js +538 -0
  15. package/dist/contracts.d.ts +236 -11
  16. package/dist/contracts.js +8 -0
  17. package/dist/event-multiplexer.d.ts +23 -0
  18. package/dist/event-multiplexer.js +136 -0
  19. package/dist/execution-policy.d.ts +28 -0
  20. package/dist/execution-policy.js +24 -0
  21. package/dist/feedback.d.ts +48 -0
  22. package/dist/feedback.js +230 -0
  23. package/dist/index.d.ts +20 -6
  24. package/dist/index.js +13 -5
  25. package/dist/input.js +11 -1
  26. package/dist/leases.d.ts +8 -0
  27. package/dist/leases.js +111 -0
  28. package/dist/node/agent-definitions.js +3 -5
  29. package/dist/node/config.d.ts +1 -0
  30. package/dist/node/config.js +5 -3
  31. package/dist/node/contribution-discovery.js +5 -8
  32. package/dist/node/session-store-jsonl.js +8 -5
  33. package/dist/node/settings.js +2 -2
  34. package/dist/node/trust.js +2 -4
  35. package/dist/observability.d.ts +3 -0
  36. package/dist/observability.js +18 -0
  37. package/dist/providers/media.d.ts +44 -0
  38. package/dist/providers/media.js +126 -0
  39. package/dist/providers/openai-compatible.js +18 -119
  40. package/dist/providers/openai-primitives.d.ts +9 -0
  41. package/dist/providers/openai-primitives.js +129 -0
  42. package/dist/providers/transport.d.ts +40 -0
  43. package/dist/providers/transport.js +221 -0
  44. package/dist/redaction.js +40 -13
  45. package/dist/resources.d.ts +5 -0
  46. package/dist/resources.js +4 -0
  47. package/dist/structured-output.d.ts +11 -0
  48. package/dist/structured-output.js +59 -0
  49. package/dist/testing/feedback.d.ts +6 -0
  50. package/dist/testing/feedback.js +37 -0
  51. package/dist/testing/persistence-schema.d.ts +102 -0
  52. package/dist/testing/persistence-schema.js +487 -0
  53. package/dist/testing/provider-conformance.js +10 -1
  54. package/dist/testing/run-ledger-conformance.d.ts +33 -0
  55. package/dist/testing/run-ledger-conformance.js +178 -0
  56. package/dist/testing/session-store-conformance.d.ts +16 -0
  57. package/dist/testing/session-store-conformance.js +73 -0
  58. package/dist/tools.d.ts +17 -0
  59. package/dist/tools.js +29 -2
  60. package/docs/a2a.md +73 -0
  61. package/docs/agent-events.md +17 -10
  62. package/docs/agent-loops.md +11 -5
  63. package/docs/agent-session-runtime.md +15 -16
  64. package/docs/cli-rpc.md +36 -5
  65. package/docs/coding-agent-tools.md +43 -9
  66. package/docs/coding-security.md +88 -0
  67. package/docs/compaction-observational-memory.md +2 -0
  68. package/docs/context-and-skills.md +1 -0
  69. package/docs/credential-storage.md +177 -0
  70. package/docs/credentials-and-redaction.md +4 -3
  71. package/docs/database-persistence.md +52 -7
  72. package/docs/evaluations.md +122 -0
  73. package/docs/extensions.md +2 -2
  74. package/docs/host-security.md +33 -2
  75. package/docs/index.md +46 -18
  76. package/docs/input-and-prompt-assembly.md +6 -5
  77. package/docs/mcp-tools.md +184 -0
  78. package/docs/middleware-hooks.md +2 -0
  79. package/docs/migration.md +51 -28
  80. package/docs/model-registry.md +5 -3
  81. package/docs/multimodal-content.md +156 -0
  82. package/docs/observability.md +171 -0
  83. package/docs/performance.md +249 -1
  84. package/docs/persistence-credentials-multimodality-primitives.md +303 -0
  85. package/docs/postgres-persistence.md +143 -0
  86. package/docs/provider-conformance.md +18 -0
  87. package/docs/provider-layer.md +1 -1
  88. package/docs/provider-packages.md +2 -0
  89. package/docs/provider-primitives.md +281 -0
  90. package/docs/providers/ai-sdk.md +113 -0
  91. package/docs/providers/kimi.md +1 -0
  92. package/docs/providers/neuralwatt.md +1 -0
  93. package/docs/providers/openai-compatible.md +2 -1
  94. package/docs/providers/openai.md +8 -1
  95. package/docs/providers/opencode-go.md +1 -0
  96. package/docs/providers/openrouter.md +1 -0
  97. package/docs/providers/zai.md +1 -0
  98. package/docs/public-contracts.md +13 -5
  99. package/docs/rag.md +113 -0
  100. package/docs/release-and-install.md +237 -30
  101. package/docs/resource-loading.md +14 -4
  102. package/docs/review-coverage-2026-07-14.md +260 -0
  103. package/docs/review-coverage-2026-07-15.md +193 -0
  104. package/docs/run-ledger-conformance.md +96 -0
  105. package/docs/runs-and-usage.md +43 -4
  106. package/docs/server.md +139 -0
  107. package/docs/session-store-conformance.md +16 -0
  108. package/docs/session-stores-and-branching.md +1 -0
  109. package/docs/settings-auth-trust-security.md +6 -5
  110. package/docs/sqlite-persistence.md +123 -0
  111. package/docs/structured-output.md +9 -0
  112. package/docs/supervisors.md +71 -0
  113. package/docs/tool-conformance.md +1 -0
  114. package/docs/tool-execution-primitives.md +374 -0
  115. package/docs/tools.md +39 -1
  116. package/docs/workflow-orchestration-primitives.md +581 -0
  117. package/docs/workflow-tui-primitives.md +5 -0
  118. package/docs/workflows.md +293 -0
  119. package/docs/working-and-semantic-memory.md +169 -0
  120. package/package.json +43 -5
  121. package/templates/init/README.md.tmpl +28 -0
  122. package/templates/init/env.example.tmpl +1 -0
  123. package/templates/init/gitignore.tmpl +11 -0
  124. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  125. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  126. package/templates/init/package.json.tmpl +22 -0
  127. package/templates/init/providers.json +76 -0
  128. package/templates/init/src/agent.ts.tmpl +10 -0
  129. package/templates/init/src/index.ts.tmpl +12 -0
  130. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  131. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -0,0 +1,178 @@
1
+ // ponytail: dependency-free conformance helper for the RunLedger adapter contract.
2
+ // Database-backed adapters call this once (or via runRunLedgerConformance factory)
3
+ // to assert durable run/event/tool/usage writes, per-run ordering, tenant
4
+ // isolation, and restart idempotency before shipping dialect-local SQL.
5
+ const NOW = "2026-01-01T00:00:00.000Z";
6
+ /**
7
+ * Assert that a `RunLedger` implementation satisfies the write contract:
8
+ * all record kinds round-trip via optional read callbacks, per-run event order
9
+ * is preserved, and tenant-scoped rows store `tenant_id` when
10
+ * `exerciseTenantIsolation` is enabled.
11
+ */
12
+ export async function assertRunLedgerConforms(fixture, options = {}) {
13
+ const sessionId = options.sessionId ?? "ledger-conformance";
14
+ const runId = options.runId ?? "run-conformance";
15
+ const tenantId = options.tenantId ?? "tenant-a";
16
+ const scope = {
17
+ tenantId,
18
+ accountId: options.accountId ?? "account-a",
19
+ userId: options.userId ?? "user-a",
20
+ };
21
+ const runStart = {
22
+ id: runId,
23
+ sessionId,
24
+ status: "running",
25
+ startedAt: NOW,
26
+ provider: "mock",
27
+ model: { provider: "mock", model: "demo" },
28
+ idempotencyKey: "run-key-1",
29
+ ...scope,
30
+ };
31
+ const runFinish = {
32
+ ...runStart,
33
+ status: "succeeded",
34
+ finishedAt: "2026-01-01T00:00:01.000Z",
35
+ };
36
+ const eventA = {
37
+ id: "event-a",
38
+ sessionId,
39
+ runId,
40
+ type: "agent_started",
41
+ timestamp: NOW,
42
+ event: { type: "agent_started", sessionId, runId },
43
+ redacted: false,
44
+ ...scope,
45
+ };
46
+ const eventB = {
47
+ id: "event-b",
48
+ sessionId,
49
+ runId,
50
+ type: "turn_started",
51
+ timestamp: "2026-01-01T00:00:00.500Z",
52
+ event: { type: "turn_started", sessionId, runId, turn: 1 },
53
+ redacted: false,
54
+ ...scope,
55
+ };
56
+ const tool = {
57
+ id: "tool-row-1",
58
+ sessionId,
59
+ runId,
60
+ toolCallId: "call-1",
61
+ name: "echo",
62
+ arguments: { msg: "hi" },
63
+ status: "finished",
64
+ result: { toolCallId: "call-1", name: "echo", value: "hi" },
65
+ startedAt: NOW,
66
+ finishedAt: "2026-01-01T00:00:00.800Z",
67
+ redacted: false,
68
+ ...scope,
69
+ };
70
+ const usage = {
71
+ id: "usage-1",
72
+ sessionId,
73
+ runId,
74
+ scope: "provider_turn",
75
+ turn: 1,
76
+ attempt: 2,
77
+ usage: { inputTokens: 3, outputTokens: 5, totalTokens: 8 },
78
+ recordedAt: "2026-01-01T00:00:01.000Z",
79
+ ...scope,
80
+ };
81
+ await fixture.ledger.appendRun(runStart);
82
+ await fixture.ledger.appendEvent(eventA);
83
+ await fixture.ledger.appendEvent(eventB);
84
+ await fixture.ledger.appendToolCall(tool);
85
+ await fixture.ledger.appendUsage(usage);
86
+ await fixture.ledger.appendRun(runFinish);
87
+ if (fixture.readRuns) {
88
+ const runs = await fixture.readRuns();
89
+ const snapshots = runs.filter((row) => row.id === runId);
90
+ if (!snapshots.some((row) => row.status === "succeeded")) {
91
+ throw new Error("RunLedger must persist the terminal RunRecord");
92
+ }
93
+ const succeeded = snapshots.find((row) => row.status === "succeeded");
94
+ if (succeeded.startedAt !== runStart.startedAt) {
95
+ throw new Error("RunLedger must preserve startedAt on the terminal RunRecord");
96
+ }
97
+ if (snapshots.length > 1 && !snapshots.some((row) => row.status === "running")) {
98
+ throw new Error("RunLedger must persist the running RunRecord when multiple snapshots are stored");
99
+ }
100
+ }
101
+ if (fixture.readEvents) {
102
+ const events = await fixture.readEvents();
103
+ const forRun = events.filter((row) => row.runId === runId);
104
+ const ids = forRun.map((row) => row.id);
105
+ const aIndex = ids.indexOf("event-a");
106
+ const bIndex = ids.indexOf("event-b");
107
+ if (aIndex < 0 || bIndex < 0)
108
+ throw new Error("RunLedger dropped appended AgentEventRecord rows");
109
+ if (aIndex > bIndex)
110
+ throw new Error("RunLedger must preserve per-run event append order");
111
+ }
112
+ if (fixture.readToolCalls) {
113
+ const toolCalls = await fixture.readToolCalls();
114
+ if (!toolCalls.some((row) => row.toolCallId === "call-1")) {
115
+ throw new Error("RunLedger must persist ToolCallRecord rows");
116
+ }
117
+ }
118
+ if (fixture.readUsage) {
119
+ const usageRows = await fixture.readUsage();
120
+ const storedUsage = usageRows.find((row) => row.id === "usage-1");
121
+ if (!storedUsage)
122
+ throw new Error("RunLedger must persist UsageRecord rows");
123
+ if (storedUsage.scope !== "provider_turn" || storedUsage.turn !== 1 || storedUsage.attempt !== 2) {
124
+ throw new Error("RunLedger must preserve UsageRecord scope, turn, and attempt");
125
+ }
126
+ }
127
+ if (options.exerciseTenantIsolation && fixture.readRuns) {
128
+ const otherRun = {
129
+ id: "run-tenant-b",
130
+ sessionId: `${sessionId}-tenant-b`,
131
+ status: "running",
132
+ startedAt: NOW,
133
+ provider: "mock",
134
+ tenantId: "tenant-b",
135
+ accountId: "account-b",
136
+ userId: "user-b",
137
+ };
138
+ await fixture.ledger.appendRun(otherRun);
139
+ const runs = await fixture.readRuns();
140
+ const stored = runs.find((row) => row.id === "run-tenant-b");
141
+ if (!stored || stored.tenantId !== "tenant-b") {
142
+ throw new Error("RunLedger must persist tenant_id on records for tenant-scoped isolation");
143
+ }
144
+ const scoped = runs.filter((row) => row.tenantId === tenantId);
145
+ if (!scoped.some((row) => row.id === runId)) {
146
+ throw new Error("RunLedger read path must return tenant-a rows when queried for conformance session");
147
+ }
148
+ }
149
+ }
150
+ /**
151
+ * Factory-based conformance entry point for durable adapters. Invokes
152
+ * `assertRunLedgerConforms` against a fresh fixture and optionally reopens via
153
+ * the same factory to assert writes survive process/database reopen.
154
+ */
155
+ export async function runRunLedgerConformance(factory, options = {}) {
156
+ const first = await factory();
157
+ await assertRunLedgerConforms(first, options);
158
+ if (!options.exerciseReopen)
159
+ return;
160
+ const reopened = await factory();
161
+ if (!reopened.readRuns && !reopened.readEvents) {
162
+ throw new Error("exerciseReopen requires readRuns or readEvents on the fixture");
163
+ }
164
+ if (reopened.readRuns) {
165
+ const runs = await reopened.readRuns();
166
+ const runId = options.runId ?? "run-conformance";
167
+ if (!runs.some((row) => row.id === runId)) {
168
+ throw new Error("RunLedger rows did not survive adapter reopen");
169
+ }
170
+ }
171
+ if (reopened.readEvents) {
172
+ const events = await reopened.readEvents();
173
+ if (!events.some((row) => row.id === "event-a")) {
174
+ throw new Error("RunLedger events did not survive adapter reopen");
175
+ }
176
+ }
177
+ }
178
+ //# sourceMappingURL=run-ledger-conformance.js.map
@@ -2,13 +2,23 @@ import type { SessionStore } from "../contracts.js";
2
2
  export interface SessionStoreConformanceOptions {
3
3
  /** Stable session id used for the conformance run; defaults to "conformance". */
4
4
  readonly sessionId?: string;
5
+ /** Secondary session id for branch-isolation probes; defaults to "conformance-other". */
6
+ readonly otherSessionId?: string;
5
7
  /**
6
8
  * When true, also exercises the optional `readBranchPath` branch-reader path
7
9
  * and asserts it returns the ancestor chain in root-to-leaf order. Skipped
8
10
  * when the store does not implement `readBranchPath`.
9
11
  */
10
12
  readonly exerciseReadBranchPath?: boolean;
13
+ /** When true, appends concurrent children of the same parent (fork allowed). */
14
+ readonly exerciseConcurrentParentAppend?: boolean;
15
+ /**
16
+ * When true, the factory is invoked again after writes to assert durable state
17
+ * survives reopen (database adapters only).
18
+ */
19
+ readonly exerciseReopen?: boolean;
11
20
  }
21
+ export type SessionStoreConformanceFactory = () => SessionStore | Promise<SessionStore>;
12
22
  /**
13
23
  * Assert that a `SessionStore` implementation satisfies the core adapter
14
24
  * contract: round-trip append/list, duplicate-entry-id rejection,
@@ -18,3 +28,9 @@ export interface SessionStoreConformanceOptions {
18
28
  * violation; returns silently when the store conforms.
19
29
  */
20
30
  export declare function assertSessionStoreConforms(store: SessionStore, options?: SessionStoreConformanceOptions): Promise<void>;
31
+ /**
32
+ * Factory-based conformance entry point for durable adapters. Calls the factory
33
+ * to obtain a store, runs the full contract, and optionally reopens through the
34
+ * same factory to assert idempotency rows and entries survive restart.
35
+ */
36
+ export declare function runSessionStoreConformance(factory: SessionStoreConformanceFactory, options?: SessionStoreConformanceOptions): Promise<void>;
@@ -63,6 +63,79 @@ export async function assertSessionStoreConforms(store, options = {}) {
63
63
  throw new Error(`readBranchPath must return the ancestor chain root→leaf in order; got ${JSON.stringify(ids)}`);
64
64
  }
65
65
  }
66
+ await assertSessionStoreBranchIsolation(store, options);
67
+ if (options.exerciseConcurrentParentAppend) {
68
+ await assertConcurrentParentAppendAllowed(store, sessionId, now, make);
69
+ }
70
+ }
71
+ /**
72
+ * Factory-based conformance entry point for durable adapters. Calls the factory
73
+ * to obtain a store, runs the full contract, and optionally reopens through the
74
+ * same factory to assert idempotency rows and entries survive restart.
75
+ */
76
+ export async function runSessionStoreConformance(factory, options = {}) {
77
+ const store = await factory();
78
+ await assertSessionStoreConforms(store, options);
79
+ if (!options.exerciseReopen)
80
+ return;
81
+ const sessionId = options.sessionId ?? "conformance";
82
+ const listed = await store.list(sessionId);
83
+ const parent = listed.at(-1);
84
+ await store.append({
85
+ id: "reopen-target",
86
+ parentId: parent?.id,
87
+ sessionId,
88
+ timestamp: "2026-01-01T00:00:01.000Z",
89
+ kind: "label",
90
+ label: "reopen",
91
+ }, { idempotencyKey: "reopen-idem", expectedParentId: parent?.id });
92
+ const before = (await store.list(sessionId)).map((entry) => entry.id);
93
+ const reopened = await factory();
94
+ const after = (await reopened.list(sessionId)).map((entry) => entry.id);
95
+ if (before.join("\u0000") !== after.join("\u0000")) {
96
+ throw new Error("Session entries did not survive adapter reopen");
97
+ }
98
+ await reject(() => reopened.append({
99
+ id: "reopen-idem-dup",
100
+ parentId: parent?.id,
101
+ sessionId,
102
+ timestamp: "2026-01-01T00:00:01.000Z",
103
+ kind: "label",
104
+ label: "dup",
105
+ }, { idempotencyKey: "reopen-idem", expectedParentId: parent?.id }), (error) => isSessionAppendConflict(error) && error.conflict.idempotencyDuplicate === true, "Restarted store must still deduplicate an exact idempotency retry");
106
+ }
107
+ async function assertSessionStoreBranchIsolation(store, options) {
108
+ const sessionId = options.sessionId ?? "conformance";
109
+ const otherSessionId = options.otherSessionId ?? `${sessionId}-other`;
110
+ const entry = {
111
+ id: "isolation-entry",
112
+ sessionId: otherSessionId,
113
+ timestamp: "2026-01-01T00:00:00.000Z",
114
+ kind: "label",
115
+ label: "isolated",
116
+ };
117
+ await store.append(entry);
118
+ const primary = await store.list(sessionId);
119
+ if (primary.some((row) => row.id === "isolation-entry")) {
120
+ throw new Error("SessionStore leaked entries across session ids");
121
+ }
122
+ }
123
+ async function assertConcurrentParentAppendAllowed(store, sessionId, now, make) {
124
+ const forkRoot = make("fork-root");
125
+ await store.append(forkRoot);
126
+ const childA = make("fork-a", forkRoot.id);
127
+ const childB = make("fork-b", forkRoot.id);
128
+ const results = await Promise.allSettled([
129
+ store.append(childA, { expectedParentId: forkRoot.id }),
130
+ store.append(childB, { expectedParentId: forkRoot.id }),
131
+ ]);
132
+ const succeeded = results.filter((result) => result.status === "fulfilled").length;
133
+ if (succeeded === 0)
134
+ throw new Error("Concurrent append to an existing parent rejected both writers; at least one fork child must succeed");
135
+ const listed = await store.list(sessionId);
136
+ if (!listed.some((entry) => entry.id === "fork-a") && !listed.some((entry) => entry.id === "fork-b")) {
137
+ throw new Error("Concurrent parent append wrote no child entries");
138
+ }
66
139
  }
67
140
  function assertIds(entries, expected, message) {
68
141
  const actual = entries.map((entry) => entry.id);
package/dist/tools.d.ts CHANGED
@@ -9,6 +9,23 @@ export interface ToolFilter {
9
9
  }
10
10
  export type ToolFilterInput = ToolFilter | readonly ToolFilter[];
11
11
  export type ToolValidator = (tool: ToolDefinition, args: JsonObject, context: ToolExecutionContext) => void | string | ErrorInfo | Promise<void | string | ErrorInfo>;
12
+ export interface ToolArgumentValidationError {
13
+ readonly path?: string;
14
+ readonly message: string;
15
+ }
16
+ export interface ToolArgumentValidationResult {
17
+ readonly ok: boolean;
18
+ readonly errors?: readonly ToolArgumentValidationError[];
19
+ }
20
+ export interface ToolArgumentValidator {
21
+ validate(schema: JsonObject, value: unknown): ToolArgumentValidationResult;
22
+ }
23
+ export interface ToolParameterValidatorOptions {
24
+ /** When a tool omits `parameters`. Default `"allow"` preserves pre-validation behavior. */
25
+ readonly missingSchema?: "allow" | "reject";
26
+ }
27
+ /** Wrap a schema adapter as the existing `ToolValidator` seam used by dispatch and the agent runtime. */
28
+ export declare function createToolParameterValidator(validator: ToolArgumentValidator, options?: ToolParameterValidatorOptions): ToolValidator;
12
29
  export interface DispatchToolCallOptions {
13
30
  readonly call: ToolCallContent;
14
31
  readonly registry: ToolRegistry;
package/dist/tools.js CHANGED
@@ -2,6 +2,26 @@ import { isJsonObject } from "./config.js";
2
2
  import { errorToErrorInfo, redactRunLedgerRecord, redactSecrets } from "./redaction.js";
3
3
  import { assertCanRegister } from "./registry-options.js";
4
4
  import { assertPermission } from "./security.js";
5
+ /** Wrap a schema adapter as the existing `ToolValidator` seam used by dispatch and the agent runtime. */
6
+ export function createToolParameterValidator(validator, options = {}) {
7
+ const missingSchema = options.missingSchema ?? "allow";
8
+ return (tool, args) => {
9
+ if (!tool.parameters) {
10
+ if (missingSchema === "reject")
11
+ return `Tool ${tool.name} has no parameters schema`;
12
+ return undefined;
13
+ }
14
+ const result = validator.validate(tool.parameters, args);
15
+ if (result.ok)
16
+ return undefined;
17
+ return formatToolArgumentValidationErrors(tool.name, result.errors);
18
+ };
19
+ }
20
+ function formatToolArgumentValidationErrors(toolName, errors) {
21
+ if (!errors?.length)
22
+ return `Tool arguments failed validation: ${toolName}`;
23
+ return errors.map((error) => (error.path ? `${error.path}: ${error.message}` : error.message)).join("; ");
24
+ }
5
25
  export function createToolRegistry(tools = [], options = {}) {
6
26
  const byName = new Map();
7
27
  const registry = {
@@ -32,6 +52,9 @@ export function filterTools(tools, filter) {
32
52
  const allows = filters.map((item) => item.allow?.length ? new Set(item.allow) : undefined).filter((item) => Boolean(item));
33
53
  return tools.filter((tool) => !denied.has(tool.name) && allows.every((allow) => allow.has(tool.name)));
34
54
  }
55
+ function toolExecutionMetadata(startedAt, status) {
56
+ return { durationMs: Math.max(0, Date.now() - Date.parse(startedAt)), status };
57
+ }
35
58
  export async function dispatchToolCall(options) {
36
59
  const secrets = options.secrets ?? [];
37
60
  const startedAt = new Date().toISOString();
@@ -80,7 +103,8 @@ export async function dispatchToolCall(options) {
80
103
  const mediatedResult = await (options.middleware?.run("tool_result", raw) ?? raw);
81
104
  const result = options.redactor?.redact(mediatedResult) ?? mediatedResult;
82
105
  const finishedAt = new Date().toISOString();
83
- await options.emit?.({ type: "tool_execution_finished", sessionId: context.sessionId, runId: context.runId, result });
106
+ const metadata = toolExecutionMetadata(startedAt, "finished");
107
+ await options.emit?.({ type: "tool_execution_finished", sessionId: context.sessionId, runId: context.runId, result, metadata });
84
108
  await appendToolCallRecord(options, "finished", mediatedCall, startedAt, { finishedAt, result });
85
109
  return result;
86
110
  }
@@ -88,7 +112,8 @@ export async function dispatchToolCall(options) {
88
112
  const info = errorToErrorInfo(error, secrets);
89
113
  const result = { toolCallId: mediatedCall.id, name: mediatedCall.name, error: info };
90
114
  const finishedAt = new Date().toISOString();
91
- await options.emit?.({ type: "tool_execution_error", sessionId: context.sessionId, runId: context.runId, call: mediatedCall, error: info });
115
+ const metadata = toolExecutionMetadata(startedAt, "error");
116
+ await options.emit?.({ type: "tool_execution_error", sessionId: context.sessionId, runId: context.runId, call: mediatedCall, error: info, metadata });
92
117
  await appendToolCallRecord(options, "error", mediatedCall, startedAt, { finishedAt, result });
93
118
  return result;
94
119
  }
@@ -105,6 +130,7 @@ async function checkCall(call, options, startedAt) {
105
130
  return undefined;
106
131
  }
107
132
  async function blocked(call, context, reason, error, options, startedAt) {
133
+ const metadata = toolExecutionMetadata(startedAt, "blocked");
108
134
  await options.emit?.({
109
135
  type: "tool_execution_blocked",
110
136
  sessionId: context.sessionId,
@@ -113,6 +139,7 @@ async function blocked(call, context, reason, error, options, startedAt) {
113
139
  name: call.name,
114
140
  reason,
115
141
  error,
142
+ metadata,
116
143
  });
117
144
  const finishedAt = new Date().toISOString();
118
145
  const result = { toolCallId: call.id, name: call.name, error };
package/docs/a2a.md ADDED
@@ -0,0 +1,73 @@
1
+ # A2A interoperability
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-supervisor` implements a bounded text-only subset of Agent2Agent (A2A) protocol 1.0: Agent Cards, JSON-RPC `SendMessage`, `SendStreamingMessage`, `GetExtendedAgentCard`, SSE task updates, ES256 JWS card signatures, and an explicit remote client.
6
+
7
+ ## When to use it
8
+
9
+ Use it to expose one explicitly selected Prism agent at an A2A endpoint or call a known remote A2A agent. Do not use it as endpoint discovery, a generic proxy, credential forwarding, or a replacement for local workflows.
10
+
11
+ ## Inputs / request
12
+
13
+ | API/field | Meaning |
14
+ | --- | --- |
15
+ | `createA2AAgentCard(card)` | Validates/freeze a JSONRPC protocol-1.0 HTTPS text card. |
16
+ | `signA2AAgentCard(card, { privateKey, keyId, expiresAt })` | Adds detached-payload ES256 JWS signature using WebCrypto. |
17
+ | `verifyA2AAgentCard(card, { publicKey, keyId?, now?, maxAgeMs? })` | Pins ES256/key/expiry and verifies canonical unsigned card. |
18
+ | `createA2AHandler({ card, exposure, authorize })` | Web-standard card/JSON-RPC/SSE `Request` to `Response` handler. |
19
+ | `createA2AClient({ endpoint, allowedOrigins })` | Explicit HTTPS remote client with optional card verifier/auth callback. |
20
+ | `A2ALimits` | Request 64 KiB, response 1 MiB, event 64 KiB, stream 10 MiB/10k events, concurrency 16, timeout 120s, card 64 KiB defaults; finite hard caps apply. |
21
+
22
+ ## Outputs / response / events
23
+
24
+ The handler serves `GET /.well-known/agent-card.json` and its configured POST endpoint. JSON-RPC returns `{ result: { task } }` or a bounded error. Streaming returns backpressure-driven SSE task envelopes. Client `send()` maps a terminal remote task to `AgentRunResult`; `stream()` yields validated/redacted text artifacts.
25
+
26
+ ## Request/response example
27
+
28
+ ```json
29
+ {"jsonrpc":"2.0","id":1,"method":"SendMessage","params":{"message":{"role":"user","messageId":"m1","parts":[{"text":"Check sources"}]}}}
30
+ ```
31
+
32
+ ## Implementation example
33
+
34
+ ```ts
35
+ import { createA2AClient, createA2AHandler, verifyA2AAgentCard } from "@arnilo/prism-supervisor";
36
+
37
+ const handler = createA2AHandler({
38
+ card,
39
+ exposure: { sessionFactory: ({ ownership }) => agent.createSession({ metadata: ownership }) },
40
+ authorize: ({ request }) => authenticate(request),
41
+ });
42
+
43
+ const client = createA2AClient({
44
+ endpoint: "https://agent.example/a2a/v1",
45
+ allowedOrigins: ["https://agent.example"],
46
+ authorize: () => ({ authorization: `Bearer ${resolveOwnedToken()}` }),
47
+ verifyCard: (remoteCard) => verifyA2AAgentCard(remoteCard, { publicKey, keyId: "agent-key" }),
48
+ });
49
+
50
+ const result = await client.send("Check sources");
51
+ ```
52
+
53
+ ## Extension and configuration notes
54
+
55
+ The package owns no listener or credential store. Mount the handler in a host server and resolve authentication/authorization on every request. Client auth executes only after body/card validation and serialization. Injectable `fetch` supports host transports/tests; redirects are disabled.
56
+
57
+ Only `text` parts are accepted. File/data parts, push notifications, task persistence/query/cancel, gRPC, HTTP+JSON binding, automatic JWK fetching, and endpoint discovery are intentionally absent.
58
+
59
+ ## Security and performance notes
60
+
61
+ - Endpoints and card URLs must be HTTPS and exactly origin-allow-listed before fetch; `redirect: "error"` prevents redirect SSRF.
62
+ - Treat every remote card, error, task, status, artifact, and SSE frame as untrusted. Shape/count/byte/time limits apply before mapping.
63
+ - Card verification pins `alg=ES256`, optional key ID, issue/expiry, optional maximum age, and canonical unsigned-card payload. Hosts provision trusted public keys; remote `jku` is never fetched automatically.
64
+ - Card discovery is public; extended-card and invoke methods call host authorization. Use TLS, rate limits, and replay controls at the host edge.
65
+ - Credentials remain in the client auth callback or server authorizer and never enter cards, messages, events, or metrics.
66
+ - Offline conformance is authoritative. Live endpoints are optional operator smoke tests.
67
+
68
+ ## Related APIs
69
+
70
+ - [Supervisor delegation](supervisors.md): local child boundary.
71
+ - [Web-standard server](server.md): non-A2A Prism routes.
72
+ - [Host security](host-security.md): authentication, SSRF, and untrusted-output policy.
73
+ - [Agent/session runtime](agent-session-runtime.md): mapped local execution/result.
@@ -8,7 +8,7 @@ Events are emitted by the runtime and by loops through `LoopContext.emit`, both
8
8
 
9
9
  ## When to use it
10
10
 
11
- Subscribe via `session.subscribe()` whenever a host needs to observe run progress: render streamed assistant text in a UI, react to tool execution, drive observability/telemetry, or audit artifact validation outcomes. Do not parse provider stream events directly for these — `AgentEvent` is the stable, normalized surface across providers and loops.
11
+ Subscribe via `session.stream()` for a single owned run, or `session.subscribe()` when a host needs a long-lived observer across runs: render streamed assistant text in a UI, react to tool execution, drive observability/telemetry, or audit artifact validation outcomes. Do not parse provider stream events directly for these — `AgentEvent` is the stable, normalized surface across providers and loops.
12
12
 
13
13
  Do not use `AgentEvent` for durable replay (use a `SessionStore`) or for cross-session coordination (the broadcaster is per-session and live-only).
14
14
 
@@ -40,6 +40,7 @@ The `AgentEvent` union (grouped by concern):
40
40
  | --- | --- |
41
41
  | Agent lifecycle | `agent_started`, `agent_finished` |
42
42
  | Turns | `turn_started`, `turn_finished` |
43
+ | Provider turns | `provider_turn_started`, `provider_turn_finished` |
43
44
  | Assistant messages | `message_started`, `message_delta`, `message_finished` |
44
45
  | Tool execution | `tool_execution_started`, `tool_execution_progress`, `tool_execution_finished`, `tool_execution_error`, `tool_execution_blocked` |
45
46
  | Queue/subscribers | `queue_updated`, `event_subscriber_overflow` |
@@ -57,7 +58,7 @@ Agent / turn / message events:
57
58
  | Variant | Fields |
58
59
  | --- | --- |
59
60
  | `agent_started` | `sessionId`, `runId` |
60
- | `agent_finished` | `sessionId`, `runId`, `usage?: Usage` |
61
+ | `agent_finished` | `sessionId`, `runId`, `usage?: Usage` (aggregate of all usage-bearing provider turns) |
61
62
  | `turn_started` / `turn_finished` | `sessionId`, `runId`, `turn: number` |
62
63
  | `message_started` / `message_finished` | `sessionId`, `runId`, `message: Message` |
63
64
  | `message_delta` | `sessionId`, `runId`, `content: ContentBlock` (`tool_call_delta` fragments may appear here for live UI streaming; stored messages use final `tool_call` blocks) |
@@ -70,11 +71,11 @@ Tool execution events:
70
71
  | --- | --- |
71
72
  | `tool_execution_started` | `sessionId`, `runId`, `call: ToolCallContent` |
72
73
  | `tool_execution_progress` | `sessionId`, `runId`, `toolCallId`, `name`, `progress?`, `metadata?` |
73
- | `tool_execution_finished` | `sessionId`, `runId`, `result: ToolResult` |
74
- | `tool_execution_error` | `sessionId`, `runId`, `call: ToolCallContent`, `error: ErrorInfo` |
75
- | `tool_execution_blocked` | `sessionId`, `runId`, `toolCallId`, `name`, `reason: string`, `error: ErrorInfo` |
74
+ | `tool_execution_finished` | `sessionId`, `runId`, `result: ToolResult`, `metadata: ToolExecutionMetadata` |
75
+ | `tool_execution_error` | `sessionId`, `runId`, `call: ToolCallContent`, `error: ErrorInfo`, `metadata: ToolExecutionMetadata` |
76
+ | `tool_execution_blocked` | `sessionId`, `runId`, `toolCallId`, `name`, `reason: string`, `error: ErrorInfo`, `metadata: ToolExecutionMetadata` |
76
77
 
77
- Queue / subscriber / compaction / retry events:
78
+ Queue / subscriber / compaction / retry / provider events:
78
79
 
79
80
  | Variant | Fields |
80
81
  | --- | --- |
@@ -84,6 +85,13 @@ Queue / subscriber / compaction / retry events:
84
85
  | `compaction_finished` | `sessionId`, `runId?`, `summary: string` |
85
86
  | `retry_scheduled` | `sessionId`, `runId`, `attempt: number`, `delayMs: number`, `error: ErrorInfo` |
86
87
 
88
+ Provider turn events (metadata only — see [Observability](observability.md)):
89
+
90
+ | Variant | Fields |
91
+ | --- | --- |
92
+ | `provider_turn_started` | `sessionId`, `runId`, `turn`, `metadata: ProviderTurnMetadata` |
93
+ | `provider_turn_finished` | `sessionId`, `runId`, `turn`, `metadata` (includes `latencyMs` on finish), `usage?`, `error?` |
94
+
87
95
  Artifact validation/refinement events (emitted only by `generateValidateReviseLoop`; `singleShotLoop` emits zero artifact events):
88
96
 
89
97
  | Variant | Fields |
@@ -165,12 +173,10 @@ const session = createAgent({
165
173
  provider: createMockProvider([providerTextDelta("ok"), providerDone()]),
166
174
  }).createSession();
167
175
 
168
- for await (const event of session.subscribe()) {
176
+ for await (const event of session.stream("draft", { loop: { strategy: "generate-validate-revise", validator, maxRevisions: 3 } })) {
169
177
  if (event.type === "artifact_finished") console.log("artifact ok", event.attempt);
170
178
  if (event.type === "artifact_failed") console.log("artifact exhausted", event.attempt, event.result.errors);
171
179
  }
172
-
173
- await session.run("draft", { loop: { strategy: "generate-validate-revise", validator, maxRevisions: 3 } });
174
180
  ```
175
181
 
176
182
  ## Extension and configuration notes
@@ -191,9 +197,10 @@ await session.run("draft", { loop: { strategy: "generate-validate-revise", valid
191
197
  - Runtime events contain messages/content only; do not put secrets in prompts, metadata, provider events, session entries, tool results, or artifact validation payloads.
192
198
 
193
199
  ## Related APIs
194
- - [Agent/session runtime](agent-session-runtime.md): `session.subscribe()` and the live event broadcaster.
200
+ - [Agent/session runtime](agent-session-runtime.md): `session.stream()`, `session.subscribe()`, and the live event broadcaster.
195
201
  - [Agent loops](agent-loops.md): `singleShotLoop` and `generateValidateReviseLoop` emit the artifact events.
196
202
  - [Structured output](structured-output.md): `ArtifactValidation` shape threaded through parser/validator/repairer.
197
203
  - [Public contracts](public-contracts.md): full `AgentEvent` union and `ArtifactValidation` contract.
204
+ - [Observability](observability.md): `ProviderTurnMetadata`, OpenTelemetry adapter package.
198
205
  - [Tools](tools.md): `tool_execution_*` variants.
199
206
  - [Compaction and retry policies](compaction-and-retry.md): `compaction_*` and `retry_scheduled` variants.
@@ -46,7 +46,7 @@ const agent = createAgent({
46
46
  model,
47
47
  provider,
48
48
  // optional default loop for this agent:
49
- loop: { strategy: "single-shot" },
49
+ loop: { strategy: "single-shot", toolConcurrency: 4 },
50
50
  });
51
51
 
52
52
  // RunOptions.loop overrides per request.
@@ -68,7 +68,11 @@ await session.run(input, { loop: myCustomLoop });
68
68
 
69
69
  ```ts
70
70
  type AgentLoopOptions =
71
- | { readonly strategy: "single-shot" }
71
+ | {
72
+ readonly strategy: "single-shot";
73
+ /** Independent tool calls per turn run concurrently up to this limit. Default `1`. */
74
+ readonly toolConcurrency?: number;
75
+ }
72
76
  | {
73
77
  readonly strategy: "generate-validate-revise";
74
78
  readonly validator: ArtifactValidator<unknown>;
@@ -95,7 +99,7 @@ Host callback contracts (all generic over host `T`):
95
99
  | --- | --- |
96
100
  | `sessionId`, `runId`, `metadata`, `signal` | Run identity and abort. |
97
101
  | `history: Message[]` | Live mutable history — the loop pushes assistant and repair messages directly. |
98
- | `input`, `inputMessages`, `maxToolRounds` | First-turn input, the redacted input messages, and the tool-round budget (single-shot parity hooks). |
102
+ | `input`, `inputMessages`, `maxToolRounds`, `toolConcurrency` | First-turn input, redacted input messages, tool-round budget, and per-turn parallel dispatch limit (`toolConcurrency` default `1`). |
99
103
  | `assemble(nextInput, toolResults?)` | Wraps `assembleProviderInput()` with resolved skills/tools/context/system prompt/provider options. |
100
104
  | `generate(request)` | Wraps provider request policies + `provider_request` middleware + `generateWithRetry()`; returns `ProviderTurnResult`. |
101
105
  | `dispatchToolCall(call)` | Wraps `dispatchToolCall()` with resolved registry/middleware/permission/redactor/validate. |
@@ -104,7 +108,7 @@ Host callback contracts (all generic over host `T`):
104
108
 
105
109
  ## Outputs / response / events
106
110
 
107
- `AgentLoopStrategy.run(ctx)` returns `Promise<Usage | undefined>` — the last provider usage, handed back to the runtime which emits `agent_finished` with it.
111
+ `AgentLoopStrategy.run(ctx)` returns `Promise<Usage | undefined>` as a fallback for custom loops. Core runtime independently accumulates every usage-bearing provider turn in O(turns), persists scoped turn/run rows, and emits `agent_finished` with the aggregate.
108
112
 
109
113
  Events during a loop run are the existing `AgentEvent`s (`turn_started`, `message_started`, `message_delta`, `message_finished`, `turn_finished`, tool-execution events when the loop dispatches tools, `error` on real failures). Both built-in loops emit `turn_started` before each provider turn, `message_finished` for every assistant draft, and `turn_finished` after the assistant draft is appended. First-turn input is appended to live history once, matching the already-persisted user message.
110
114
 
@@ -195,7 +199,8 @@ await session.run(input, { loop: twoShotLoop });
195
199
  - `{ strategy: "single-shot" }` resolves to the exported `singleShotLoop`; `{ strategy: "generate-validate-revise", ... }` is mapped by `resolveLoop()` to `generateValidateReviseLoop(opts)`. An unknown `strategy` throws before the first turn. Passing an `AgentLoopStrategy` instance bypasses the options form entirely (custom-loop escape hatch).
196
200
  - The loop is resolved once per run inside `RuntimeAgentSession.run()`, after the usual setup (provider/skills/tools resolution, history rebuild, model-change entry, input append, auto-compaction). The runtime's outer try/catch/finally, run-exclusivity, abort bridging, and subscriber close remain in place around `loop.run(ctx)`.
197
201
  - `LoopContext.assemble(nextInput, toolResults?)` accepts an optional tool-result accumulator so `singleShotLoop` can pass its loop-local `toolResults`; `generateValidateReviseLoop` omits it (no tools in revision turns).
198
- - `maxToolRounds` bounds `singleShotLoop` tool rounds; `maxRevisions` (default 3) bounds `generateValidateReviseLoop` revision turns. Budget exhaustion ends the loop and returns the last usage; it does not throw.
202
+ - `maxToolRounds` bounds `singleShotLoop` tool rounds; `toolConcurrency` (default `1`) bounds how many independent tool calls from one provider turn may execute concurrently. Results and transcript rows are still appended in original call order. If any resolved `ToolDefinition` in a turn has `exclusive: true`, that turn uses concurrency `1`; later non-exclusive turns restore configured concurrency.
203
+ - `maxRevisions` (default 3) bounds `generateValidateReviseLoop` revision turns. Budget exhaustion ends the loop and returns the last usage; it does not throw.
199
204
  - A revision cycle appends one assistant draft and one repair user message per revision to the session store, so store entries reflect every attempted draft. The original user input is stored once by the runtime and pushed into loop history once on the first turn.
200
205
 
201
206
  ## Security and performance notes
@@ -203,6 +208,7 @@ await session.run(input, { loop: twoShotLoop });
203
208
  - Loops have no path to credentials, provider objects, or unredacted secrets. `LoopContext.generate` consumes an already-redacted request; `LoopContext.emit` runs through `redactAgentEvent` with the active `SecretRedactor`; `LoopContext.appendMessage` appends a redacted entry.
204
209
  - `ArtifactValidation.errors[].message` may echo model text — `artifact_*` event payloads flow through the same `redactAgentEvent` path as other `AgentEvent`s (see [Agent events](agent-events.md)).
205
210
  - `generateValidateReviseLoop` makes at most `maxRevisions + 1` provider turns; it cannot loop forever on an always-failing validator. Each revision costs one provider turn plus one store append.
211
+ - Parallel tool dispatch uses a bounded worker pool over the calls in one turn; queue depth is `calls.length`, not unbounded. Exclusive turns use the same sequential path. Each call still runs through `dispatchToolCall` (permission + validation + execute). Tool lifecycle events may complete out of order; history/store appends stay in call order.
206
212
  - The loop is a plain object/factory; no class hierarchy, no background work, no extra dependencies. `LoopContext` is a single object literal of bound arrows built once per run.
207
213
  - The Synapta-free boundary is guarded by tests: `src/` imports no `synapta*` package, and the `Artifact*`/`AgentLoop*`/`LoopContext` contracts contain no `workflow`/`node`/`step` field names. Hosts supply their own schema; no host domain type is imported by `src/`.
208
214