@arnilo/prism 0.0.8 → 0.0.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +30 -3
- package/dist/agent-loops.js +7 -3
- package/dist/agents.js +16 -1
- package/dist/contracts.d.ts +2 -0
- package/dist/index.d.ts +2 -2
- package/dist/index.js +2 -2
- package/dist/provider-events.d.ts +2 -0
- package/dist/provider-events.js +21 -13
- package/dist/providers/openai-compatible.js +8 -5
- package/dist/providers/transport.d.ts +10 -1
- package/dist/providers/transport.js +24 -8
- package/dist/tools.js +2 -0
- package/docs/agent-events.md +3 -2
- package/docs/agent-loops.md +2 -2
- package/docs/browser-automation.md +124 -0
- package/docs/coding-agent-tools.md +111 -14
- package/docs/coding-security.md +84 -11
- package/docs/evaluations.md +13 -2
- package/docs/guardrails.md +2 -1
- package/docs/host-security.md +4 -2
- package/docs/index.md +10 -7
- package/docs/migration.md +68 -0
- package/docs/performance.md +29 -0
- package/docs/provider-conformance.md +1 -1
- package/docs/provider-primitives.md +7 -1
- package/docs/release-and-install.md +89 -45
- package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
- package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
- package/docs/structured-output.md +2 -2
- package/docs/tools.md +3 -0
- package/docs/web-tools.md +1 -1
- package/docs/workflows.md +1 -0
- package/package.json +5 -4
package/CHANGELOG.md
CHANGED
|
@@ -1,11 +1,38 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
-
All notable changes to this project will be documented in this file.
|
|
4
|
-
|
|
5
3
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
6
4
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
5
|
|
|
8
|
-
|
|
6
|
+
All notable changes to this project will be documented in this file.
|
|
7
|
+
|
|
8
|
+
## [0.0.10] - 2026-07-21
|
|
9
|
+
|
|
10
|
+
### Changed
|
|
11
|
+
|
|
12
|
+
- Coding harness workspace modes (Phase 5): required `workspaceMode` on `@arnilo/prism-coding-security` composition; sandbox mode unifies shell/FS on one disposable tree; host mode never claims containment; fail-closed mixed wiring + `allowMixedWorkspaceWiring` escape hatch; import/export tree identity; `scripts/benchmark-0.0.10.mjs` evidence.
|
|
13
|
+
- Versioned all 32 first-party manifests and exact internal ranges from the post-ship `0.0.96` graph to `0.0.10` for the roadmap Phase 5 release line.
|
|
14
|
+
|
|
15
|
+
## [0.0.96] - 2026-07-21
|
|
16
|
+
|
|
17
|
+
### Changed
|
|
18
|
+
|
|
19
|
+
- Package graph and runtime version pins bumped from 0.0.9 to 0.0.96 for a clean publish tag after the mistaken `v0.0.95` tag and TypeScript 7 / workspace-order CI fixes.
|
|
20
|
+
|
|
21
|
+
## [0.0.9] - 2026-07-21
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- Production coding and browser execution for Release 0.0.9: disposable Docker sandbox, bounded native repository list/search, structured Git/named checks/PR handoff, durable coding-plan/checkpoint composition, and optional `@arnilo/prism-browser` with egress/side-effect/upload/download/screenshot policy.
|
|
26
|
+
- Versioned all 32 first-party manifests and exact internal ranges to 0.0.9 (adds `@arnilo/prism-browser` to the publishable graph; browser stays out of `@arnilo/prism-code` and activates only through explicit install or `@arnilo/prism-all`).
|
|
27
|
+
- Added network-free coding/browser adversarial evaluation fixtures, `scripts/benchmark-0.0.9.mjs`, and protected Docker/Playwright gates via `.github/workflows/sandbox-browser.yml`.
|
|
28
|
+
- Office execution remains outside Prism packaging by product decision (host-selected skills/instructions only).
|
|
29
|
+
- `tryParseJsonObjectArguments` and `toolCallFromArgumentsText` for recoverable streamed tool-call argument parsing.
|
|
30
|
+
|
|
31
|
+
### Fixed
|
|
32
|
+
|
|
33
|
+
- Malformed streamed tool-call arguments (id+name present) become failed/`tool_execution_blocked` tool results (`invalid_arguments` / `invalid_json_arguments`) instead of terminal `ProviderTransportError`, so models can self-correct within existing turn budgets.
|
|
34
|
+
- Incomplete tool-call deltas (missing id/name) fail with typed `ProviderTransportError` / `ErrorInfo.code: "incomplete_delta"` instead of a bare `Error("Incomplete tool call delta...")`; openai-compatible streams no longer emit `done` alongside leftover incomplete deltas.
|
|
35
|
+
- Empty/whitespace-only call-free artifact candidates (including thinking-only output) are `parse_error` through the revision budget; `generate-validate-revise` session runs no longer resolve `succeeded` without `artifact_finished`.
|
|
9
36
|
|
|
10
37
|
## [0.0.8] - 2026-07-20
|
|
11
38
|
|
package/dist/agent-loops.js
CHANGED
|
@@ -111,9 +111,13 @@ export function generateValidateReviseLoop(opts) {
|
|
|
111
111
|
.filter((b) => b.type === "text")
|
|
112
112
|
.map((b) => b.text)
|
|
113
113
|
.join("");
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
114
|
+
// Empty/whitespace-only call-free output is a parse failure (thinking-only
|
|
115
|
+
// models must not succeed with an empty artifact via the identity parser).
|
|
116
|
+
const parsed = text.trim() === ""
|
|
117
|
+
? { ok: false, error: "no artifact text in model output" }
|
|
118
|
+
: opts.parser
|
|
119
|
+
? await opts.parser(text, artifactCtx)
|
|
120
|
+
: { ok: true, value: text };
|
|
117
121
|
// Parse failure consumes revision budget like a validation failure; the
|
|
118
122
|
// repairer receives `undefined` value plus a synthetic parse issue.
|
|
119
123
|
const parseFailure = !parsed.ok || parsed.value === undefined
|
package/dist/agents.js
CHANGED
|
@@ -274,6 +274,8 @@ class RuntimeAgentSession {
|
|
|
274
274
|
};
|
|
275
275
|
// ponytail: LoopContext binds existing private helpers; loop orchestrates only.
|
|
276
276
|
let assembledTurn = false;
|
|
277
|
+
let artifactFinished = false;
|
|
278
|
+
let artifactFailedInfo;
|
|
277
279
|
const ctx = {
|
|
278
280
|
sessionId: this.id,
|
|
279
281
|
runId,
|
|
@@ -368,6 +370,16 @@ class RuntimeAgentSession {
|
|
|
368
370
|
emit: (event) => {
|
|
369
371
|
if (event.type === "turn_started")
|
|
370
372
|
this.activeLoopTurn = event.turn;
|
|
373
|
+
if (event.type === "artifact_finished")
|
|
374
|
+
artifactFinished = true;
|
|
375
|
+
if (event.type === "artifact_failed") {
|
|
376
|
+
const first = event.result.errors?.[0];
|
|
377
|
+
const reason = event.result.metadata?.reason;
|
|
378
|
+
artifactFailedInfo = {
|
|
379
|
+
message: first?.message ?? "artifact failed",
|
|
380
|
+
code: typeof reason === "string" || typeof reason === "number" ? reason : "artifact_failed",
|
|
381
|
+
};
|
|
382
|
+
}
|
|
371
383
|
this.emit(event);
|
|
372
384
|
},
|
|
373
385
|
};
|
|
@@ -380,6 +392,9 @@ class RuntimeAgentSession {
|
|
|
380
392
|
});
|
|
381
393
|
}
|
|
382
394
|
const loopUsage = await loop.run(ctx);
|
|
395
|
+
if (loop.name === "generate-validate-revise" && !artifactFinished) {
|
|
396
|
+
throw Object.assign(new Error(artifactFailedInfo?.message ?? "artifact loop ended without a validated artifact"), { name: "ArtifactFailed", code: artifactFailedInfo?.code ?? "artifact_failed" });
|
|
397
|
+
}
|
|
383
398
|
usage = runUsage.value() ?? loopUsage;
|
|
384
399
|
if (usage && this.activeLedger) {
|
|
385
400
|
const usageRecord = {
|
|
@@ -999,7 +1014,7 @@ function policyList(policies) {
|
|
|
999
1014
|
return "apply" in policies ? [policies] : policies;
|
|
1000
1015
|
}
|
|
1001
1016
|
function errorFromInfo(error) {
|
|
1002
|
-
return Object.assign(new Error(error.message), { name: error.name ?? "Error", cause: error.cause });
|
|
1017
|
+
return Object.assign(new Error(error.message), { name: error.name ?? "Error", cause: error.cause, code: error.code });
|
|
1003
1018
|
}
|
|
1004
1019
|
class ProviderTurnFailure extends Error {
|
|
1005
1020
|
info;
|
package/dist/contracts.d.ts
CHANGED
|
@@ -49,6 +49,8 @@ export interface ToolCallContent {
|
|
|
49
49
|
readonly id: string;
|
|
50
50
|
readonly name: string;
|
|
51
51
|
readonly arguments: JsonObject;
|
|
52
|
+
/** Set when streamed arguments failed JSON parse; dispatch blocks without execute(). */
|
|
53
|
+
readonly argumentsError?: ErrorInfo;
|
|
52
54
|
}
|
|
53
55
|
export interface ToolResultContent {
|
|
54
56
|
readonly type: "tool_result";
|
package/dist/index.d.ts
CHANGED
|
@@ -60,7 +60,7 @@ export { createMockProvider } from "./mock-provider.js";
|
|
|
60
60
|
export { createMemorySessionStore, createSessionEntry, getSessionBranchEntries, listSessionBranches, rebuildSessionContext } from "./session-stores.js";
|
|
61
61
|
export type { CreateSessionEntryOptions, SessionBranch, SessionBranchOptions, SessionContextSnapshot } from "./session-stores.js";
|
|
62
62
|
export type { MockProviderOptions } from "./mock-provider.js";
|
|
63
|
-
export { providerContentDelta, providerDone, providerError, providerTextDelta, providerThinkingDelta, providerToolCall, providerToolCallDelta, providerUsage, toolCallContent, } from "./provider-events.js";
|
|
63
|
+
export { providerContentDelta, providerDone, providerError, providerTextDelta, providerThinkingDelta, providerToolCall, providerToolCallDelta, providerUsage, toolCallContent, toolCallFromArgumentsText, } from "./provider-events.js";
|
|
64
64
|
export type { ProviderResolver } from "./contracts.js";
|
|
65
65
|
export { createProviderRegistry, createProviderResolver } from "./providers.js";
|
|
66
66
|
export type { ProviderRegistry, ProviderRegistryOptions } from "./providers.js";
|
|
@@ -83,5 +83,5 @@ export type { DispatchToolCallOptions, ToolArgumentValidationError, ToolArgument
|
|
|
83
83
|
export type { DuplicateRegistrationOptions, DuplicateRegistrationPolicy } from "./registry-options.js";
|
|
84
84
|
export { dispatchToolCallsInOrder, generateValidateReviseLoop, isAgentLoopOptions, resolveLoop, resolveToolConcurrency, singleShotLoop } from "./agent-loops.js";
|
|
85
85
|
export declare const name = "prism";
|
|
86
|
-
export declare const version = "0.0.
|
|
86
|
+
export declare const version = "0.0.10";
|
|
87
87
|
export declare const description = "Agent harness for AI providers, agents, sessions, and tools.";
|
package/dist/index.js
CHANGED
|
@@ -33,7 +33,7 @@ export { createMiddlewareRegistry } from "./middleware.js";
|
|
|
33
33
|
export { assembleProviderInput, createDefaultInputBuilder, createDefaultPromptBuilder, renderPromptTemplate, resolveContextProviders } from "./input.js";
|
|
34
34
|
export { createMockProvider } from "./mock-provider.js";
|
|
35
35
|
export { createMemorySessionStore, createSessionEntry, getSessionBranchEntries, listSessionBranches, rebuildSessionContext } from "./session-stores.js";
|
|
36
|
-
export { providerContentDelta, providerDone, providerError, providerTextDelta, providerThinkingDelta, providerToolCall, providerToolCallDelta, providerUsage, toolCallContent, } from "./provider-events.js";
|
|
36
|
+
export { providerContentDelta, providerDone, providerError, providerTextDelta, providerThinkingDelta, providerToolCall, providerToolCallDelta, providerUsage, toolCallContent, toolCallFromArgumentsText, } from "./provider-events.js";
|
|
37
37
|
export { createProviderRegistry, createProviderResolver } from "./providers.js";
|
|
38
38
|
export { createSecretRedactor, errorToErrorInfo, redactAgentEvent, redactMessage, redactProviderRequest, redactRunLedgerRecord, redactSecrets, redactSessionEntry } from "./redaction.js";
|
|
39
39
|
export { assertPermission, assertTrusted, checkPermission, createStaticPermissionPolicy, createStaticTrustPolicy, denialToErrorInfo, isTrusted, PermissionDeniedError, TrustDeniedError } from "./security.js";
|
|
@@ -45,6 +45,6 @@ export { assertGuardrailsAllowed, GuardrailError, MAX_GUARDRAIL_CONCURRENCY, run
|
|
|
45
45
|
export { createRunLimitTracker, DEFAULT_RUN_LIMITS, HARD_MAX_RUN_COST, HARD_RUN_LIMITS, RunLimitError, RunLimitTracker, resolveRunLimits } from "./run-limits.js";
|
|
46
46
|
export { dispatchToolCallsInOrder, generateValidateReviseLoop, isAgentLoopOptions, resolveLoop, resolveToolConcurrency, singleShotLoop } from "./agent-loops.js";
|
|
47
47
|
export const name = "prism";
|
|
48
|
-
export const version = "0.0.
|
|
48
|
+
export const version = "0.0.10";
|
|
49
49
|
export const description = "Agent harness for AI providers, agents, sessions, and tools.";
|
|
50
50
|
//# sourceMappingURL=index.js.map
|
|
@@ -15,3 +15,5 @@ export declare function providerUsage(usage: Usage): ProviderEvent;
|
|
|
15
15
|
export declare function providerDone(usage?: Usage): ProviderEvent;
|
|
16
16
|
export declare function providerError(error: unknown, secrets?: readonly (string | undefined)[]): ProviderEvent;
|
|
17
17
|
export declare function toolCallContent(id: string, name: string, args?: JsonObject): ToolCallContent;
|
|
18
|
+
/** Build a tool call from streamed arguments text; malformed JSON becomes a blocked call (no throw). */
|
|
19
|
+
export declare function toolCallFromArgumentsText(id: string, name: string, argumentsText: string): ToolCallContent;
|
package/dist/provider-events.js
CHANGED
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
import { ProviderTransportError, tryParseJsonObjectArguments } from "./providers/transport.js";
|
|
1
2
|
import { errorToErrorInfo } from "./redaction.js";
|
|
2
3
|
export function providerTextDelta(text) {
|
|
3
4
|
return { type: "content_delta", content: { type: "text", text } };
|
|
@@ -32,9 +33,10 @@ export function reconstructToolCallDeltas(events) {
|
|
|
32
33
|
partials.set(event.index, partial);
|
|
33
34
|
}
|
|
34
35
|
return [...partials.entries()].sort(([a], [b]) => a - b).map(([index, partial]) => {
|
|
35
|
-
if (!partial.id || !partial.name)
|
|
36
|
-
throw new
|
|
37
|
-
|
|
36
|
+
if (!partial.id || !partial.name) {
|
|
37
|
+
throw new ProviderTransportError("incomplete_delta", `Incomplete tool call delta at index ${index}`);
|
|
38
|
+
}
|
|
39
|
+
return toolCallFromArgumentsText(partial.id, partial.name, partial.argumentsText);
|
|
38
40
|
});
|
|
39
41
|
}
|
|
40
42
|
export function providerUsage(usage) {
|
|
@@ -50,15 +52,21 @@ export function providerError(error, secrets = []) {
|
|
|
50
52
|
export function toolCallContent(id, name, args = {}) {
|
|
51
53
|
return { type: "tool_call", id, name, arguments: args };
|
|
52
54
|
}
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
55
|
+
/** Build a tool call from streamed arguments text; malformed JSON becomes a blocked call (no throw). */
|
|
56
|
+
export function toolCallFromArgumentsText(id, name, argumentsText) {
|
|
57
|
+
const parsed = tryParseJsonObjectArguments(argumentsText, { toolName: name });
|
|
58
|
+
if (parsed.ok)
|
|
59
|
+
return toolCallContent(id, name, parsed.value);
|
|
60
|
+
return {
|
|
61
|
+
type: "tool_call",
|
|
62
|
+
id,
|
|
63
|
+
name,
|
|
64
|
+
arguments: {},
|
|
65
|
+
argumentsError: {
|
|
66
|
+
name: parsed.error.name,
|
|
67
|
+
message: parsed.error.message,
|
|
68
|
+
code: parsed.error.code,
|
|
69
|
+
},
|
|
70
|
+
};
|
|
63
71
|
}
|
|
64
72
|
//# sourceMappingURL=provider-events.js.map
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
import { resolveCredentialValue } from "../credentials.js";
|
|
2
|
-
import { providerDone, providerError, providerTextDelta, providerThinkingDelta, providerToolCall, providerToolCallDelta, providerUsage,
|
|
2
|
+
import { providerDone, providerError, providerTextDelta, providerThinkingDelta, providerToolCall, providerToolCallDelta, providerUsage, toolCallFromArgumentsText, } from "../provider-events.js";
|
|
3
3
|
import { assertOpenAIChatMessage, applyOpenAIChatStructuredOutput, mapOpenAIChatUsage, serializeOpenAIChatMessage, serializeOpenAITool, } from "./openai-primitives.js";
|
|
4
|
-
import {
|
|
4
|
+
import { ProviderTransportError, readBoundedResponseText, readSseData, } from "./transport.js";
|
|
5
5
|
import { assertStructuredOutputRequestSupported } from "../structured-output.js";
|
|
6
6
|
export function createOpenAICompatibleProvider(options) {
|
|
7
7
|
const providerId = options.id ?? "openai-compatible";
|
|
@@ -64,10 +64,13 @@ export function createOpenAICompatibleProvider(options) {
|
|
|
64
64
|
}
|
|
65
65
|
}
|
|
66
66
|
}
|
|
67
|
+
const incomplete = [...tools.entries()].find(([, call]) => !call.id || !call.name);
|
|
68
|
+
if (incomplete) {
|
|
69
|
+
yield providerError(new ProviderTransportError("incomplete_delta", `Incomplete tool call delta at index ${incomplete[0]}`), secrets);
|
|
70
|
+
return;
|
|
71
|
+
}
|
|
67
72
|
for (const call of tools.values()) {
|
|
68
|
-
|
|
69
|
-
yield providerToolCall(toolCallContent(call.id, call.name, parseJsonObjectArguments(call.argumentsText, { toolName: call.name })));
|
|
70
|
-
}
|
|
73
|
+
yield providerToolCall(toolCallFromArgumentsText(call.id, call.name, call.argumentsText));
|
|
71
74
|
}
|
|
72
75
|
yield providerDone();
|
|
73
76
|
}
|
|
@@ -14,7 +14,7 @@ export interface SseEvent {
|
|
|
14
14
|
readonly data: string;
|
|
15
15
|
readonly comments?: readonly string[];
|
|
16
16
|
}
|
|
17
|
-
export type ProviderTransportErrorCode = "sse_buffer_overflow" | "sse_event_overflow" | "response_body_overflow" | "aborted" | "invalid_json_arguments";
|
|
17
|
+
export type ProviderTransportErrorCode = "sse_buffer_overflow" | "sse_event_overflow" | "response_body_overflow" | "aborted" | "invalid_json_arguments" | "incomplete_delta";
|
|
18
18
|
export declare class ProviderTransportError extends Error {
|
|
19
19
|
readonly code: ProviderTransportErrorCode;
|
|
20
20
|
readonly limitBytes?: number;
|
|
@@ -36,5 +36,14 @@ export declare function readSseEvents(body: ReadableStream<Uint8Array>, options?
|
|
|
36
36
|
export declare function readSseData(body: ReadableStream<Uint8Array>, options?: ReadSseEventsOptions): AsyncGenerator<string>;
|
|
37
37
|
/** Read a response body with a hard byte ceiling; redacts optional secrets. */
|
|
38
38
|
export declare function readBoundedResponseText(response: Response, options?: ReadBoundedResponseTextOptions): Promise<string>;
|
|
39
|
+
export type ParseJsonObjectArgumentsResult = {
|
|
40
|
+
readonly ok: true;
|
|
41
|
+
readonly value: JsonObject;
|
|
42
|
+
} | {
|
|
43
|
+
readonly ok: false;
|
|
44
|
+
readonly error: ProviderTransportError;
|
|
45
|
+
};
|
|
46
|
+
/** Parse streamed tool arguments without throwing; prefer this for recoverable tool-call recovery. */
|
|
47
|
+
export declare function tryParseJsonObjectArguments(text: string, options?: ParseJsonObjectArgumentsOptions): ParseJsonObjectArgumentsResult;
|
|
39
48
|
/** Parse streamed tool arguments as a JSON object; throws {@link ProviderTransportError} on invalid input. */
|
|
40
49
|
export declare function parseJsonObjectArguments(text: string, options?: ParseJsonObjectArgumentsOptions): JsonObject;
|
|
@@ -197,25 +197,41 @@ export async function readBoundedResponseText(response, options) {
|
|
|
197
197
|
}
|
|
198
198
|
}
|
|
199
199
|
}
|
|
200
|
-
/** Parse streamed tool arguments
|
|
201
|
-
export function
|
|
200
|
+
/** Parse streamed tool arguments without throwing; prefer this for recoverable tool-call recovery. */
|
|
201
|
+
export function tryParseJsonObjectArguments(text, options) {
|
|
202
202
|
const maxBytes = options?.maxBytes ?? DEFAULT_MAX_ARGUMENT_BYTES;
|
|
203
203
|
const suffix = options?.toolName ? ` for tool ${options.toolName}` : "";
|
|
204
204
|
if (!text)
|
|
205
|
-
return {};
|
|
205
|
+
return { ok: true, value: {} };
|
|
206
206
|
if (byteLength(text) > maxBytes) {
|
|
207
|
-
|
|
207
|
+
return {
|
|
208
|
+
ok: false,
|
|
209
|
+
error: new ProviderTransportError("invalid_json_arguments", `Tool arguments${suffix} exceeded ${maxBytes} bytes`, maxBytes),
|
|
210
|
+
};
|
|
208
211
|
}
|
|
209
212
|
let parsed;
|
|
210
213
|
try {
|
|
211
214
|
parsed = JSON.parse(text);
|
|
212
215
|
}
|
|
213
|
-
catch
|
|
214
|
-
|
|
216
|
+
catch {
|
|
217
|
+
return {
|
|
218
|
+
ok: false,
|
|
219
|
+
error: new ProviderTransportError("invalid_json_arguments", `Invalid tool arguments JSON${suffix}`),
|
|
220
|
+
};
|
|
215
221
|
}
|
|
216
222
|
if (!parsed || typeof parsed !== "object" || Array.isArray(parsed)) {
|
|
217
|
-
|
|
223
|
+
return {
|
|
224
|
+
ok: false,
|
|
225
|
+
error: new ProviderTransportError("invalid_json_arguments", `Tool arguments${suffix} must be a JSON object`),
|
|
226
|
+
};
|
|
218
227
|
}
|
|
219
|
-
return parsed;
|
|
228
|
+
return { ok: true, value: parsed };
|
|
229
|
+
}
|
|
230
|
+
/** Parse streamed tool arguments as a JSON object; throws {@link ProviderTransportError} on invalid input. */
|
|
231
|
+
export function parseJsonObjectArguments(text, options) {
|
|
232
|
+
const result = tryParseJsonObjectArguments(text, options);
|
|
233
|
+
if (!result.ok)
|
|
234
|
+
throw result.error;
|
|
235
|
+
return result.value;
|
|
220
236
|
}
|
|
221
237
|
//# sourceMappingURL=transport.js.map
|
package/dist/tools.js
CHANGED
|
@@ -176,6 +176,8 @@ async function checkCall(call, options, startedAt) {
|
|
|
176
176
|
return blocked(call, context, "unknown_tool", { message: `Unknown tool: ${call.name}` }, options, startedAt);
|
|
177
177
|
if (filterTools([tool], options.filter).length === 0)
|
|
178
178
|
return blocked(call, context, "tool_denied", { message: `Tool denied: ${call.name}` }, options, startedAt);
|
|
179
|
+
if (call.argumentsError)
|
|
180
|
+
return blocked(call, context, "invalid_arguments", call.argumentsError, options, startedAt);
|
|
179
181
|
if (!isJsonObject(call.arguments))
|
|
180
182
|
return blocked(call, context, "invalid_arguments", { message: "Tool arguments must be a JSON object" }, options, startedAt);
|
|
181
183
|
return undefined;
|
package/docs/agent-events.md
CHANGED
|
@@ -67,7 +67,7 @@ Agent / turn / message events:
|
|
|
67
67
|
| `message_started` / `message_finished` | `sessionId`, `runId`, `message: Message` |
|
|
68
68
|
| `message_delta` | `sessionId`, `runId`, `content: ContentBlock` (`tool_call_delta` fragments may appear here for live UI streaming; stored messages use final `tool_call` blocks) |
|
|
69
69
|
|
|
70
|
-
`message_delta.content.type === "tool_call_delta"` carries `{ index, id?, name?, argumentsText? }`. Treat it as a streaming fragment. The runtime reconstructs and persists a final `tool_call` before executing tools.
|
|
70
|
+
`message_delta.content.type === "tool_call_delta"` carries `{ index, id?, name?, argumentsText? }`. Treat it as a streaming fragment. The runtime reconstructs and persists a final `tool_call` before executing tools. Deltas missing `id`/`name` at stream end fail the provider turn with `ErrorInfo.code: "incomplete_delta"` (typed `ProviderTransportError`); they never throw a bare `Error`. Malformed JSON with id+name present recovers as a blocked tool result (`invalid_json_arguments`) instead.
|
|
71
71
|
|
|
72
72
|
Tool execution events:
|
|
73
73
|
|
|
@@ -127,7 +127,8 @@ turn_started → message_started → message_delta* → message_finished → tur
|
|
|
127
127
|
With opt-in `toolCalls: "bounded"`, a provider turn containing calls emits its normal assistant envelope followed by existing `tool_execution_*` events and matching persisted tool results; it emits no validation event and the next provider turn consumes that transcript. A post-`maxToolRounds` call emits terminal `artifact_failed` directly after `turn_finished` and has no tool execution event.
|
|
128
128
|
|
|
129
129
|
- `attempt` is 1-indexed per call-free validation candidate. It can differ from provider `turn` when bounded tool calls occur.
|
|
130
|
-
-
|
|
130
|
+
- Empty/whitespace-only call-free text (including thinking-only content) emits `artifact_validation_*` with `metadata.reason: "parse_error"` before any host parser runs.
|
|
131
|
+
- Single-shot runs emit zero artifact events. Session runs with `generate-validate-revise` require `artifact_finished` to resolve `succeeded`.
|
|
131
132
|
- **Validation failure triggering a revision is recoverable and never an `error`.** Terminal candidate-budget or `tool_round_limit` exhaustion emits `artifact_failed`; real failures remain on the `error` channel.
|
|
132
133
|
|
|
133
134
|
## Request/response example
|
package/docs/agent-loops.md
CHANGED
|
@@ -89,7 +89,7 @@ Host callback contracts (all generic over host `T`):
|
|
|
89
89
|
|
|
90
90
|
| Contract | Shape |
|
|
91
91
|
| --- | --- |
|
|
92
|
-
| `ArtifactParser<T>` | `(text: string, ctx: ArtifactContext) => ArtifactParseResult<T> \| Promise<...>` — parse assistant text to a typed value. A parse failure (`ok: false` or missing `value`) consumes revision budget exactly like a validation failure: the repairer receives `value: undefined` plus a synthetic failure (`errors[0].message` = the parse error, `metadata.reason: "parse_error"`), and budget exhaustion ends with terminal `artifact_failed`. |
|
|
92
|
+
| `ArtifactParser<T>` | `(text: string, ctx: ArtifactContext) => ArtifactParseResult<T> \| Promise<...>` — parse assistant text to a typed value. Empty/whitespace-only call-free text is rejected before the parser (`metadata.reason: "parse_error"`, message `no artifact text in model output`) so thinking-only/reasoning-only turns cannot succeed via the identity parser. A parse failure (`ok: false` or missing `value`) consumes revision budget exactly like a validation failure: the repairer receives `value: undefined` plus a synthetic failure (`errors[0].message` = the parse error, `metadata.reason: "parse_error"`), and budget exhaustion ends with terminal `artifact_failed`. |
|
|
93
93
|
| `ArtifactValidator<T>` | `(value: T, ctx: ArtifactContext) => ArtifactValidation \| Promise<...>` — return `{ ok: true }` or `{ ok: false, errors }`. |
|
|
94
94
|
| `ArtifactRepairer<T>` | `(value: T \| undefined, failure: ArtifactValidation, ctx: ArtifactContext) => AgentInput \| Promise<...>` — build the revision follow-up input. |
|
|
95
95
|
| `ArtifactValidation` | `{ ok: boolean; errors?: readonly { path?: string; message: string }[]; metadata?: ... }`. |
|
|
@@ -119,7 +119,7 @@ Host callback contracts (all generic over host `T`):
|
|
|
119
119
|
|
|
120
120
|
Events during a loop run are the existing `AgentEvent`s (`turn_started`, `message_started`, `message_delta`, `message_finished`, `turn_finished`, tool-execution events when the loop dispatches tools, `error` on real failures). Both built-in loops emit `turn_started` before each provider turn, `message_finished` for every assistant draft, and `turn_finished` after the assistant draft is appended. First-turn input is appended to live history once, matching the already-persisted user message.
|
|
121
121
|
|
|
122
|
-
Validation-failure-triggering-a-revision is **not** an `error` event — it is recoverable, like `tool_execution_blocked`. In bounded artifact mode, a tool-calling provider response emits normal assistant/tool lifecycle events, skips artifact parsing/validation, then the next turn sees its persisted result. `generateValidateReviseLoop` emits artifact events only for call-free candidates: `artifact_validation_started` → `artifact_validation_finished` → (`artifact_revision_started`)* → `artifact_finished` | `artifact_failed`. A request beyond `maxToolRounds` executes nothing and emits terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"`; see [Agent events § Artifact event ordering](agent-events.md#artifact-event-ordering). `singleShotLoop` emits zero artifact events. Real failures stay on the `error` channel.
|
|
122
|
+
Validation-failure-triggering-a-revision is **not** an `error` event — it is recoverable, like `tool_execution_blocked`. In bounded artifact mode, a tool-calling provider response emits normal assistant/tool lifecycle events, skips artifact parsing/validation, then the next turn sees its persisted result. `generateValidateReviseLoop` emits artifact events only for call-free candidates: `artifact_validation_started` → `artifact_validation_finished` → (`artifact_revision_started`)* → `artifact_finished` | `artifact_failed`. A request beyond `maxToolRounds` executes nothing and emits terminal `artifact_failed` with `result.metadata.reason === "tool_round_limit"`; see [Agent events § Artifact event ordering](agent-events.md#artifact-event-ordering). `singleShotLoop` emits zero artifact events. Real failures stay on the `error` channel. Session runs using `generate-validate-revise` resolve `succeeded` only after `artifact_finished`; terminal `artifact_failed` (including empty/thinking-only parse exhaustion) fails the run with `AgentRunError` (`error.code` from `result.metadata.reason`, e.g. `parse_error`).
|
|
123
123
|
|
|
124
124
|
A loop has no path to credentials, provider objects, or unredacted secrets. `LoopContext.generate` receives the already-policy-applied, middleware-run, redacted request; `LoopContext.emit` runs through `redactAgentEvent` with the active `SecretRedactor`.
|
|
125
125
|
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
# Browser automation
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-browser` exposes exactly four exclusive model-facing tools—`browser_open`, `browser_snapshot`, `browser_act`, and `browser_close`—over a host-supplied Playwright `Browser`. Prism creates one non-persistent `BrowserContext` per run, serializes actions, returns bounded AI-mode accessibility snapshots with snapshot-scoped refs, enforces egress/side-effect/upload/download/screenshot policy, and closes context/pages/listeners/quarantined downloads on close, abort, or manager disposal.
|
|
6
|
+
|
|
7
|
+
## When to use it
|
|
8
|
+
|
|
9
|
+
Use when an agent must interact with JavaScript-heavy or authenticated pages that search/fetch cannot cover. Prefer `@arnilo/prism-web-tools` for ordinary public retrieval. Do not use this package as a browser launcher, MCP proxy, visual planner, CDP console, or persistent profile manager.
|
|
10
|
+
|
|
11
|
+
## Inputs / request
|
|
12
|
+
|
|
13
|
+
| Tool | Model-visible input | Host-only construction input |
|
|
14
|
+
| --- | --- | --- |
|
|
15
|
+
| `browser_open` | optional absolute `http(s)` `url` | host Playwright `Browser` or `BrowserManager`, limits, `ExecutionPolicy`, `networkPolicy`, uploads/downloads |
|
|
16
|
+
| `browser_snapshot` | optional `pageId` | same manager/context |
|
|
17
|
+
| `browser_act` | `action` plus action-specific fields (`target`, `snapshotId`, `url`, `text`, `values`, `paths`, `downloadId`, `dialogResponse`, `pageId`, `clip`, …) | policy checked before side effects |
|
|
18
|
+
| `browser_close` | none | closes only the run-owned context, never the host Browser process |
|
|
19
|
+
|
|
20
|
+
`createBrowserTools({ browser, executionPolicy?, limits?, networkPolicy?, uploads?, downloads?, beforeSideEffect? })` builds the four tools. `createBrowserManager(...)` exposes host lifecycle helpers `closeRun(runId)` / `close()` and `listDownloads(runId)`.
|
|
21
|
+
|
|
22
|
+
Targets accepted by `browser_act`: snapshot `ref`, `role`(+`name`), `label`, `testId`, or `text`. CSS, XPath, selector strings, `page.evaluate`, CDP/devtools, extensions, and persistent/local profiles are unsupported.
|
|
23
|
+
|
|
24
|
+
`browser_act` actions: `navigate`, `click`, `type`, `fill`, `select`, `check`, `uncheck`, `scroll`, `wait`, `dialog`, `select_page`, `upload`, `screenshot`, `download_release`.
|
|
25
|
+
|
|
26
|
+
## Outputs / response / events
|
|
27
|
+
|
|
28
|
+
`browser_open` returns run/page ids and URL. `browser_snapshot` returns `snapshotId`, URL/title, bounded AI-mode aria YAML (`ariaSnapshot({ mode: "ai" })`), ref count, and truncation metadata. Refs are valid only for that snapshot id and become stale after navigation or mutation. `browser_act` returns the action, active page id, and URL; `screenshot` also returns bounded `ImageContent`; `download_release` returns quarantine metadata after host approval. `browser_close` is idempotent. Results mark `trust: "untrusted_external"`; page text must never alter tools, permissions, credentials, or policy.
|
|
29
|
+
|
|
30
|
+
## Request/response example
|
|
31
|
+
|
|
32
|
+
```json
|
|
33
|
+
{
|
|
34
|
+
"tool": "browser_snapshot",
|
|
35
|
+
"arguments": {},
|
|
36
|
+
"result": {
|
|
37
|
+
"snapshotId": "snap_ab12…",
|
|
38
|
+
"pageId": "page_1",
|
|
39
|
+
"url": "https://example.com/",
|
|
40
|
+
"title": "Example",
|
|
41
|
+
"refCount": 12,
|
|
42
|
+
"ariaSnapshot": "- main [ref=e8]:\n - button \"Submit\" [ref=e12]"
|
|
43
|
+
}
|
|
44
|
+
}
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
```json
|
|
48
|
+
{
|
|
49
|
+
"tool": "browser_act",
|
|
50
|
+
"arguments": {
|
|
51
|
+
"action": "click",
|
|
52
|
+
"target": { "ref": "e12" },
|
|
53
|
+
"snapshotId": "snap_ab12…"
|
|
54
|
+
}
|
|
55
|
+
}
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
## Implementation example
|
|
59
|
+
|
|
60
|
+
```ts
|
|
61
|
+
import { chromium } from "playwright-core";
|
|
62
|
+
import {
|
|
63
|
+
createBrowserManager,
|
|
64
|
+
createBrowserTools,
|
|
65
|
+
createSharedSandboxBrowserOptions,
|
|
66
|
+
} from "@arnilo/prism-browser";
|
|
67
|
+
import { assertBrowserSandboxNetwork } from "@arnilo/prism-coding-security";
|
|
68
|
+
|
|
69
|
+
assertBrowserSandboxNetwork({
|
|
70
|
+
mode: "custom",
|
|
71
|
+
name: "prism-egress",
|
|
72
|
+
browserEgress: { proxyEndpoint: "http://127.0.0.1:3128", denyDirectEgress: true },
|
|
73
|
+
});
|
|
74
|
+
|
|
75
|
+
const aligned = createSharedSandboxBrowserOptions({
|
|
76
|
+
workspaceRoot: "/workspace",
|
|
77
|
+
downloadsRoot: "/downloads",
|
|
78
|
+
containedProxyAttestation: {
|
|
79
|
+
proxyEndpoint: "http://127.0.0.1:3128",
|
|
80
|
+
denyDirectEgress: true,
|
|
81
|
+
},
|
|
82
|
+
approveDownloadRelease: async (meta) => meta.bytes < 1_000_000,
|
|
83
|
+
});
|
|
84
|
+
|
|
85
|
+
const browser = await chromium.launch({ headless: true });
|
|
86
|
+
const manager = createBrowserManager({
|
|
87
|
+
browser,
|
|
88
|
+
...aligned,
|
|
89
|
+
limits: { maxPages: 4, maxActions: 100, maxSnapshotBytes: 256 * 1024 },
|
|
90
|
+
});
|
|
91
|
+
const tools = createBrowserTools({ manager, executionPolicy });
|
|
92
|
+
|
|
93
|
+
// On run terminal / abort / cancel:
|
|
94
|
+
await manager.closeRun(runId);
|
|
95
|
+
await manager.close();
|
|
96
|
+
await browser.close();
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
## Extension and configuration notes
|
|
100
|
+
|
|
101
|
+
- Compatibility line: `playwright-core@1.61.0` optional peer. Hosts pin browser binaries/images; Prism package install downloads nothing.
|
|
102
|
+
- Default/hard caps: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2k/10k; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30s/120s; action 10s/60s; wait 30s/120s; run wall 20min/30min; popups 4/16; dialogs 16/64; close grace 5s/30s; network requests 1k/10k; redirects/request 10/32; WebSockets 8/32; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate.
|
|
103
|
+
- Contexts use `serviceWorkers: "block"` and install `BrowserContext.route()` for every visible HTTP(S)/WebSocket request. `acceptDownloads` is enabled only when `downloads` is configured.
|
|
104
|
+
- `networkPolicy` defaults to `requireContainedProxy: true` (fail closed). Hosts must supply `containedProxyAttestation: { proxyEndpoint, denyDirectEgress: true }`. Private/loopback/link-local, `file`/`data`/`blob`/`javascript`/`devtools` schemes are denied by default. Playwright routing is defense in depth — production DNS/private egress is a host firewall/proxy.
|
|
105
|
+
- Uploads require absolute paths under `uploads.roots` (realpath-contained; symlink escapes rejected). Downloads stream into `downloads.quarantine` with SHA-256/MIME/name metadata; `download_release` requires host `approveRelease`. Screenshots return bounded `ImageContent`.
|
|
106
|
+
- Observation (`snapshot`, `wait`, open-without-url, `close`) vs mutation/high-impact (`navigate`, click/form, dialog accept, upload, download release, popup select) is classified for `ExecutionPolicy` / `beforeSideEffect`.
|
|
107
|
+
- `createSharedSandboxBrowserOptions()` aligns browser uploads/downloads with Task 1 sandbox `/workspace` and `/downloads`. `assertBrowserSandboxNetwork()` in `@arnilo/prism-coding-security` fails closed for custom Docker networks without browser egress attestation.
|
|
108
|
+
- Raw CSS is absent from production defaults. Ref resolution uses Playwright’s built-in `aria-ref=` selector with a package-owned snapshot ref table for staleness checks.
|
|
109
|
+
|
|
110
|
+
## Security and performance notes
|
|
111
|
+
|
|
112
|
+
Import is inert. Construction fails clearly when neither `browser` nor `manager` is supplied. Browser installation, launch, version, and control endpoint are host-owned. Prism never exposes `page.evaluate`, init scripts, CDP, extensions, persistent profiles, or model-supplied Playwright launch options. Secrets and storage state must not appear in snapshots, tool results, logs, or checkpoints. Finite caps charge before context/page/action/queue/snapshot/network/artifact retention; snapshots retain no unbounded DOM, console, request, response, or trace history. Unreleased downloads are deleted on context close.
|
|
113
|
+
|
|
114
|
+
Default tests use fake Playwright APIs only. Protected live gate: `PRISM_LIVE_PLAYWRIGHT=1` (or `PRISM_TEST_PLAYWRIGHT=1`) `npm run test:live -w @arnilo/prism-browser` exercises a local loopback hostile HTML fixture for snapshot refs, stale-ref rejection, CSS denial, private/file deny, upload containment, screenshot bounds, and download quarantine/release. Missing browser binaries fail closed when the gate is enabled. Adversarial network-free fixtures live in `eval-fixtures.test.ts`; see [Evaluations](evaluations.md) and `examples/coding-browser-evaluation.ts`.
|
|
115
|
+
|
|
116
|
+
## Related APIs
|
|
117
|
+
|
|
118
|
+
- [Tools](tools.md): registry, exclusive dispatch, validation, and ledger.
|
|
119
|
+
- [Web search, fetch, and extraction](web-tools.md): preferred non-interactive retrieval path.
|
|
120
|
+
- [Guardrails](guardrails.md): untrusted external content handling.
|
|
121
|
+
- [Host security](host-security.md): browser endpoint, approval, egress proxy, and artifact trust boundaries.
|
|
122
|
+
- [Performance and resource limits](performance.md): browser ceilings and charging points.
|
|
123
|
+
- [Coding execution approval and sandboxing](coding-security.md): optional shared disposable sandbox for coding+browser.
|
|
124
|
+
- [Migration](migration.md): additive optional package activation.
|