@arnilo/prism 0.0.8 → 0.0.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +44 -3
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +44 -5
- package/dist/agents.js +141 -7
- package/dist/context-budget.d.ts +63 -0
- package/dist/context-budget.js +235 -0
- package/dist/contracts.d.ts +107 -0
- package/dist/contracts.js +77 -0
- package/dist/index.d.ts +7 -5
- package/dist/index.js +5 -4
- package/dist/input.d.ts +3 -0
- package/dist/input.js +71 -28
- package/dist/node/session-store-jsonl.js +4 -1
- package/dist/provider-events.d.ts +2 -0
- package/dist/provider-events.js +21 -13
- package/dist/providers/openai-compatible.js +8 -5
- package/dist/providers/transport.d.ts +10 -1
- package/dist/providers/transport.js +24 -8
- package/dist/rpc.js +13 -2
- package/dist/session-stores.d.ts +7 -2
- package/dist/session-stores.js +174 -4
- package/dist/structured-output.d.ts +5 -1
- package/dist/structured-output.js +18 -0
- package/dist/testing/persistence-schema.d.ts +1 -1
- package/dist/testing/persistence-schema.js +8 -2
- package/dist/testing/session-store-conformance.d.ts +6 -0
- package/dist/testing/session-store-conformance.js +36 -1
- package/dist/tools.js +2 -0
- package/docs/agent-events.md +3 -2
- package/docs/agent-loops.md +10 -3
- package/docs/agent-session-runtime.md +5 -1
- package/docs/browser-automation.md +124 -0
- package/docs/cli-rpc.md +2 -1
- package/docs/coding-agent-tools.md +178 -14
- package/docs/coding-security.md +84 -11
- package/docs/evaluations.md +13 -2
- package/docs/guardrails.md +2 -1
- package/docs/host-security.md +4 -2
- package/docs/index.md +19 -15
- package/docs/input-and-prompt-assembly.md +4 -1
- package/docs/migration.md +84 -0
- package/docs/node-jsonl-session-store.md +1 -1
- package/docs/performance.md +46 -0
- package/docs/postgres-persistence.md +3 -3
- package/docs/provider-conformance.md +1 -1
- package/docs/provider-packages.md +1 -1
- package/docs/provider-primitives.md +7 -1
- package/docs/providers/anthropic.md +92 -0
- package/docs/providers/google.md +87 -0
- package/docs/public-contracts.md +4 -0
- package/docs/release-and-install.md +197 -64
- package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
- package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
- package/docs/review-coverage-2026-07-22-phase-6.md +209 -0
- package/docs/session-store-conformance.md +2 -0
- package/docs/session-stores.md +40 -1
- package/docs/sqlite-persistence.md +3 -3
- package/docs/structured-output.md +9 -3
- package/docs/tools.md +3 -0
- package/docs/web-tools.md +1 -1
- package/docs/workflows.md +3 -0
- package/package.json +6 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,11 +1,52 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
-
All notable changes to this project will be documented in this file.
|
|
4
|
-
|
|
5
3
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
6
4
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
5
|
|
|
8
|
-
|
|
6
|
+
All notable changes to this project will be documented in this file.
|
|
7
|
+
|
|
8
|
+
## [Unreleased]
|
|
9
|
+
|
|
10
|
+
## [0.0.11] - 2026-07-22
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- Coding harness fundamentals for 0.0.11 (Plan 074): bounded `SessionIndex`/`searchSessions` (SQLite/Postgres FTS migration 004; memory linear|unsupported; JSONL unsupported), assembler `contextBudget` + omission reports, `@arnilo/prism-provider-anthropic` + `@arnilo/prism-provider-google`, mid-run `AgentSession.steer` / RPC steer, coding-agent `runCodingGoalVerify` and opt-in `ask_user_decision` (multi/free-text/durable suspend glue).
|
|
15
|
+
- Opt-in `structuredOutputTiming: "final-turn-only"` on `generate-validate-revise` (default `"every-turn"`): tool-eligible turns omit native schema so models can call tools; artifact/revision turns attach schema and withdraw tools.
|
|
16
|
+
|
|
17
|
+
### Changed
|
|
18
|
+
|
|
19
|
+
- Versioned all 34 first-party manifests and exact internal ranges to `0.0.11` (adds `@arnilo/prism-provider-anthropic` + `@arnilo/prism-provider-google` to the publishable graph and `@arnilo/prism-providers` umbrella).
|
|
20
|
+
- Network-free search/budget evidence: `scripts/benchmark-0.0.11.mjs`.
|
|
21
|
+
|
|
22
|
+
## [0.0.10] - 2026-07-21
|
|
23
|
+
|
|
24
|
+
### Changed
|
|
25
|
+
|
|
26
|
+
- Coding harness workspace modes (Phase 5): required `workspaceMode` on `@arnilo/prism-coding-security` composition; sandbox mode unifies shell/FS on one disposable tree; host mode never claims containment; fail-closed mixed wiring + `allowMixedWorkspaceWiring` escape hatch; import/export tree identity; `scripts/benchmark-0.0.10.mjs` evidence.
|
|
27
|
+
- Versioned all 32 first-party manifests and exact internal ranges from the post-ship `0.0.96` graph to `0.0.10` for the roadmap Phase 5 release line.
|
|
28
|
+
|
|
29
|
+
## [0.0.96] - 2026-07-21
|
|
30
|
+
|
|
31
|
+
### Changed
|
|
32
|
+
|
|
33
|
+
- Package graph and runtime version pins bumped from 0.0.9 to 0.0.96 for a clean publish tag after the mistaken `v0.0.95` tag and TypeScript 7 / workspace-order CI fixes.
|
|
34
|
+
|
|
35
|
+
## [0.0.9] - 2026-07-21
|
|
36
|
+
|
|
37
|
+
### Added
|
|
38
|
+
|
|
39
|
+
- Production coding and browser execution for Release 0.0.9: disposable Docker sandbox, bounded native repository list/search, structured Git/named checks/PR handoff, durable coding-plan/checkpoint composition, and optional `@arnilo/prism-browser` with egress/side-effect/upload/download/screenshot policy.
|
|
40
|
+
- Versioned all 32 first-party manifests and exact internal ranges to 0.0.9 (adds `@arnilo/prism-browser` to the publishable graph; browser stays out of `@arnilo/prism-code` and activates only through explicit install or `@arnilo/prism-all`).
|
|
41
|
+
- Added network-free coding/browser adversarial evaluation fixtures, `scripts/benchmark-0.0.9.mjs`, and protected Docker/Playwright gates via `.github/workflows/sandbox-browser.yml`.
|
|
42
|
+
- Office execution remains outside Prism packaging by product decision (host-selected skills/instructions only).
|
|
43
|
+
- `tryParseJsonObjectArguments` and `toolCallFromArgumentsText` for recoverable streamed tool-call argument parsing.
|
|
44
|
+
|
|
45
|
+
### Fixed
|
|
46
|
+
|
|
47
|
+
- Malformed streamed tool-call arguments (id+name present) become failed/`tool_execution_blocked` tool results (`invalid_arguments` / `invalid_json_arguments`) instead of terminal `ProviderTransportError`, so models can self-correct within existing turn budgets.
|
|
48
|
+
- Incomplete tool-call deltas (missing id/name) fail with typed `ProviderTransportError` / `ErrorInfo.code: "incomplete_delta"` instead of a bare `Error("Incomplete tool call delta...")`; openai-compatible streams no longer emit `done` alongside leftover incomplete deltas.
|
|
49
|
+
- Empty/whitespace-only call-free artifact candidates (including thinking-only output) are `parse_error` through the revision budget; `generate-validate-revise` session runs no longer resolve `succeeded` without `artifact_finished`.
|
|
9
50
|
|
|
10
51
|
## [0.0.8] - 2026-07-20
|
|
11
52
|
|
package/dist/agent-loops.d.ts
CHANGED
|
@@ -7,6 +7,7 @@ export declare function generateValidateReviseLoop(opts: {
|
|
|
7
7
|
readonly repairer?: ArtifactRepairer<unknown>;
|
|
8
8
|
readonly maxRevisions?: number;
|
|
9
9
|
readonly toolCalls?: "disabled" | "bounded";
|
|
10
|
+
readonly structuredOutputTiming?: "every-turn" | "final-turn-only";
|
|
10
11
|
}): AgentLoopStrategy;
|
|
11
12
|
export declare function resolveToolConcurrency(options: {
|
|
12
13
|
loop?: AgentLoopStrategy | AgentLoopOptions;
|
package/dist/agent-loops.js
CHANGED
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
import { inputMessages } from "./input.js";
|
|
2
2
|
import { createId } from "./ids.js";
|
|
3
|
+
import { artifactStructuredOutputRequest, withoutStructuredOutput } from "./structured-output.js";
|
|
3
4
|
function throwIfAborted(signal) {
|
|
4
5
|
if (signal.aborted)
|
|
5
6
|
throw signal.reason instanceof Error ? signal.reason : new Error("Agent run aborted");
|
|
@@ -23,6 +24,7 @@ export const singleShotLoop = {
|
|
|
23
24
|
let nextInput = ctx.input;
|
|
24
25
|
for (let turn = 1;; turn += 1) {
|
|
25
26
|
throwIfAborted(ctx.signal);
|
|
27
|
+
await ctx.applyPendingSteers?.();
|
|
26
28
|
ctx.emit({ type: "turn_started", sessionId: ctx.sessionId, runId: ctx.runId, turn });
|
|
27
29
|
const request = await ctx.assemble(nextInput, undefined, turn);
|
|
28
30
|
throwIfAborted(ctx.signal);
|
|
@@ -37,8 +39,14 @@ export const singleShotLoop = {
|
|
|
37
39
|
ctx.emit({ type: "message_finished", sessionId: ctx.sessionId, runId: ctx.runId, message });
|
|
38
40
|
}
|
|
39
41
|
ctx.emit({ type: "turn_finished", sessionId: ctx.sessionId, runId: ctx.runId, turn });
|
|
40
|
-
if (calls.length === 0 || toolRounds >= ctx.maxToolRounds)
|
|
42
|
+
if (calls.length === 0 || toolRounds >= ctx.maxToolRounds) {
|
|
43
|
+
// Soft-interrupt / late steer: keep same run going when queue still has text.
|
|
44
|
+
if (await ctx.applyPendingSteers?.()) {
|
|
45
|
+
nextInput = [];
|
|
46
|
+
continue;
|
|
47
|
+
}
|
|
41
48
|
break;
|
|
49
|
+
}
|
|
42
50
|
toolRounds += 1;
|
|
43
51
|
await dispatchToolCallsInOrder(calls, ctx);
|
|
44
52
|
nextInput = [];
|
|
@@ -61,6 +69,7 @@ function defaultRepairer() {
|
|
|
61
69
|
export function generateValidateReviseLoop(opts) {
|
|
62
70
|
const max = opts.maxRevisions ?? 3;
|
|
63
71
|
const repairer = opts.repairer ?? defaultRepairer();
|
|
72
|
+
const finalOnly = opts.structuredOutputTiming === "final-turn-only" && opts.toolCalls === "bounded";
|
|
64
73
|
return {
|
|
65
74
|
name: "generate-validate-revise",
|
|
66
75
|
async run(ctx) {
|
|
@@ -69,10 +78,24 @@ export function generateValidateReviseLoop(opts) {
|
|
|
69
78
|
let pendingHistory = [];
|
|
70
79
|
let toolRounds = 0;
|
|
71
80
|
let attempts = 0;
|
|
81
|
+
let artifactPhase = !finalOnly;
|
|
82
|
+
let savedSchema;
|
|
72
83
|
for (let turn = 1; attempts <= max; turn += 1) {
|
|
73
84
|
throwIfAborted(ctx.signal);
|
|
85
|
+
await ctx.applyPendingSteers?.();
|
|
74
86
|
ctx.emit({ type: "turn_started", sessionId: ctx.sessionId, runId: ctx.runId, turn });
|
|
75
|
-
|
|
87
|
+
let request = await ctx.assemble(nextInput, undefined, turn);
|
|
88
|
+
if (request.options?.structuredOutput)
|
|
89
|
+
savedSchema ??= request.options.structuredOutput;
|
|
90
|
+
if (finalOnly) {
|
|
91
|
+
if (!artifactPhase && toolRounds < ctx.maxToolRounds) {
|
|
92
|
+
request = withoutStructuredOutput(request);
|
|
93
|
+
}
|
|
94
|
+
else {
|
|
95
|
+
artifactPhase = true;
|
|
96
|
+
request = artifactStructuredOutputRequest(request, savedSchema);
|
|
97
|
+
}
|
|
98
|
+
}
|
|
76
99
|
throwIfAborted(ctx.signal);
|
|
77
100
|
const { content, calls, messageId, started, usage: turnUsage } = await ctx.generate(request);
|
|
78
101
|
usage = turnUsage ?? usage;
|
|
@@ -100,6 +123,17 @@ export function generateValidateReviseLoop(opts) {
|
|
|
100
123
|
nextInput = [];
|
|
101
124
|
continue;
|
|
102
125
|
}
|
|
126
|
+
// Soft-interrupt / late steer before treating empty/final output as artifact work.
|
|
127
|
+
if (await ctx.applyPendingSteers?.()) {
|
|
128
|
+
nextInput = [];
|
|
129
|
+
continue;
|
|
130
|
+
}
|
|
131
|
+
// Call-free during tool phase → one more turn with schema on / tools off.
|
|
132
|
+
if (finalOnly && !artifactPhase) {
|
|
133
|
+
artifactPhase = true;
|
|
134
|
+
nextInput = [];
|
|
135
|
+
continue;
|
|
136
|
+
}
|
|
103
137
|
const artifactCtx = {
|
|
104
138
|
sessionId: ctx.sessionId,
|
|
105
139
|
runId: ctx.runId,
|
|
@@ -111,9 +145,13 @@ export function generateValidateReviseLoop(opts) {
|
|
|
111
145
|
.filter((b) => b.type === "text")
|
|
112
146
|
.map((b) => b.text)
|
|
113
147
|
.join("");
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
148
|
+
// Empty/whitespace-only call-free output is a parse failure (thinking-only
|
|
149
|
+
// models must not succeed with an empty artifact via the identity parser).
|
|
150
|
+
const parsed = text.trim() === ""
|
|
151
|
+
? { ok: false, error: "no artifact text in model output" }
|
|
152
|
+
: opts.parser
|
|
153
|
+
? await opts.parser(text, artifactCtx)
|
|
154
|
+
: { ok: true, value: text };
|
|
117
155
|
// Parse failure consumes revision budget like a validation failure; the
|
|
118
156
|
// repairer receives `undefined` value plus a synthetic parse issue.
|
|
119
157
|
const parseFailure = !parsed.ok || parsed.value === undefined
|
|
@@ -210,6 +248,7 @@ export function resolveLoop(options, config) {
|
|
|
210
248
|
repairer: loop.repairer,
|
|
211
249
|
maxRevisions: loop.maxRevisions,
|
|
212
250
|
toolCalls: loop.toolCalls,
|
|
251
|
+
structuredOutputTiming: loop.structuredOutputTiming,
|
|
213
252
|
});
|
|
214
253
|
}
|
|
215
254
|
throw new Error(`Unknown agent loop strategy: ${strategy}`);
|
package/dist/agents.js
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import { AgentRunError, AgentRunStateError } from "./contracts.js";
|
|
1
|
+
import { AgentRunError, AgentRunStateError, DEFAULT_MAX_PENDING_STEER_BYTES, DEFAULT_MAX_PENDING_STEERS, } from "./contracts.js";
|
|
2
2
|
import { resolveLoop, resolveToolConcurrency } from "./agent-loops.js";
|
|
3
3
|
import { createId } from "./ids.js";
|
|
4
4
|
import { assertGuardrailsAllowed, GuardrailError, runGuardrails } from "./guardrails.js";
|
|
@@ -91,6 +91,11 @@ class RuntimeAgentSession {
|
|
|
91
91
|
currentLeafId;
|
|
92
92
|
history = [];
|
|
93
93
|
activeRun;
|
|
94
|
+
activeRunId;
|
|
95
|
+
activeProviderTurnAbort;
|
|
96
|
+
pendingSoftInterrupt = false;
|
|
97
|
+
pendingSteers = [];
|
|
98
|
+
pendingSteerBytes = 0;
|
|
94
99
|
activeRedactor;
|
|
95
100
|
activeProvider;
|
|
96
101
|
activeLedger;
|
|
@@ -124,6 +129,31 @@ class RuntimeAgentSession {
|
|
|
124
129
|
async run(input, options = {}) {
|
|
125
130
|
return this.runInternal(input, options, randomId("run"));
|
|
126
131
|
}
|
|
132
|
+
steer(input, options = {}) {
|
|
133
|
+
if (!this.activeRun || !this.activeRunId)
|
|
134
|
+
throw new Error("Agent session has no active run to steer");
|
|
135
|
+
const messages = inputToMessages(input).map((message) => this.redact(message));
|
|
136
|
+
if (messages.length === 0)
|
|
137
|
+
throw new Error("steer requires non-empty input");
|
|
138
|
+
let addBytes = 0;
|
|
139
|
+
for (const message of messages)
|
|
140
|
+
addBytes += messageTextBytes(message);
|
|
141
|
+
if (this.pendingSteers.length + messages.length > DEFAULT_MAX_PENDING_STEERS) {
|
|
142
|
+
throw new Error(`steer queue exceeds max pending messages (${DEFAULT_MAX_PENDING_STEERS})`);
|
|
143
|
+
}
|
|
144
|
+
if (this.pendingSteerBytes + addBytes > DEFAULT_MAX_PENDING_STEER_BYTES) {
|
|
145
|
+
throw new Error(`steer queue exceeds max pending bytes (${DEFAULT_MAX_PENDING_STEER_BYTES})`);
|
|
146
|
+
}
|
|
147
|
+
this.pendingSteers.push(...messages);
|
|
148
|
+
this.pendingSteerBytes += addBytes;
|
|
149
|
+
this.emit({ type: "queue_updated", sessionId: this.id, runId: this.activeRunId, size: this.pendingSteers.length });
|
|
150
|
+
if (options.softInterrupt) {
|
|
151
|
+
if (this.activeProviderTurnAbort)
|
|
152
|
+
this.activeProviderTurnAbort.abort(new SteerSoftInterrupt());
|
|
153
|
+
else
|
|
154
|
+
this.pendingSoftInterrupt = true;
|
|
155
|
+
}
|
|
156
|
+
}
|
|
127
157
|
async resumeDurable(state, runState, ownership) {
|
|
128
158
|
return this.runInternal(state.input ?? [], { runState, ownership }, state.runId, { options: runState, state, version: state.version });
|
|
129
159
|
}
|
|
@@ -165,6 +195,10 @@ class RuntimeAgentSession {
|
|
|
165
195
|
const controller = new AbortController();
|
|
166
196
|
const cleanupSignal = bridgeAbort(options.signal, controller);
|
|
167
197
|
this.activeRun = controller;
|
|
198
|
+
this.activeRunId = runId;
|
|
199
|
+
this.pendingSteers = [];
|
|
200
|
+
this.pendingSteerBytes = 0;
|
|
201
|
+
this.pendingSoftInterrupt = false;
|
|
168
202
|
this.activeRedactor = options.redactor ?? this.agent.config.redactor;
|
|
169
203
|
this.activeLedger = options.runLedger ?? this.agent.config.runLedger;
|
|
170
204
|
this.activeOwnership = options.ownership ?? this.agent.config.ownership;
|
|
@@ -274,6 +308,8 @@ class RuntimeAgentSession {
|
|
|
274
308
|
};
|
|
275
309
|
// ponytail: LoopContext binds existing private helpers; loop orchestrates only.
|
|
276
310
|
let assembledTurn = false;
|
|
311
|
+
let artifactFinished = false;
|
|
312
|
+
let artifactFailedInfo;
|
|
277
313
|
const ctx = {
|
|
278
314
|
sessionId: this.id,
|
|
279
315
|
runId,
|
|
@@ -325,7 +361,15 @@ class RuntimeAgentSession {
|
|
|
325
361
|
assembledTurn = false;
|
|
326
362
|
const policyResult = await this.applyProviderRequestPolicies(request, runId, options, metadata, controller.signal);
|
|
327
363
|
const middlewareRequest = await this.agent.config.middleware?.run("provider_request", policyResult.request) ?? policyResult.request;
|
|
328
|
-
|
|
364
|
+
try {
|
|
365
|
+
return await this.generateWithRetry(this.redactProviderRequest(middlewareRequest), runId, options, controller.signal, policyResult.secrets, this.activeLoopTurn, recordProviderUsage);
|
|
366
|
+
}
|
|
367
|
+
catch (error) {
|
|
368
|
+
if (isSteerSoftInterrupt(error)) {
|
|
369
|
+
return { content: [], calls: [], started: false, usage: undefined };
|
|
370
|
+
}
|
|
371
|
+
throw error;
|
|
372
|
+
}
|
|
329
373
|
},
|
|
330
374
|
isToolCallExclusive: (call) => registry.get(call.name)?.exclusive === true,
|
|
331
375
|
dispatchToolCall: (call) => dispatchToolCall({
|
|
@@ -365,9 +409,21 @@ class RuntimeAgentSession {
|
|
|
365
409
|
validate,
|
|
366
410
|
}),
|
|
367
411
|
appendMessage: (message) => this.appendMessage(message, runId),
|
|
412
|
+
hasPendingSteers: () => this.pendingSteers.length > 0,
|
|
413
|
+
applyPendingSteers: () => this.applyPendingSteers(runId, metadata, controller.signal),
|
|
368
414
|
emit: (event) => {
|
|
369
415
|
if (event.type === "turn_started")
|
|
370
416
|
this.activeLoopTurn = event.turn;
|
|
417
|
+
if (event.type === "artifact_finished")
|
|
418
|
+
artifactFinished = true;
|
|
419
|
+
if (event.type === "artifact_failed") {
|
|
420
|
+
const first = event.result.errors?.[0];
|
|
421
|
+
const reason = event.result.metadata?.reason;
|
|
422
|
+
artifactFailedInfo = {
|
|
423
|
+
message: first?.message ?? "artifact failed",
|
|
424
|
+
code: typeof reason === "string" || typeof reason === "number" ? reason : "artifact_failed",
|
|
425
|
+
};
|
|
426
|
+
}
|
|
371
427
|
this.emit(event);
|
|
372
428
|
},
|
|
373
429
|
};
|
|
@@ -380,6 +436,9 @@ class RuntimeAgentSession {
|
|
|
380
436
|
});
|
|
381
437
|
}
|
|
382
438
|
const loopUsage = await loop.run(ctx);
|
|
439
|
+
if (loop.name === "generate-validate-revise" && !artifactFinished) {
|
|
440
|
+
throw Object.assign(new Error(artifactFailedInfo?.message ?? "artifact loop ended without a validated artifact"), { name: "ArtifactFailed", code: artifactFailedInfo?.code ?? "artifact_failed" });
|
|
441
|
+
}
|
|
383
442
|
usage = runUsage.value() ?? loopUsage;
|
|
384
443
|
if (usage && this.activeLedger) {
|
|
385
444
|
const usageRecord = {
|
|
@@ -428,6 +487,11 @@ class RuntimeAgentSession {
|
|
|
428
487
|
finally {
|
|
429
488
|
if (this.activeRun === controller)
|
|
430
489
|
this.activeRun = undefined;
|
|
490
|
+
this.activeRunId = undefined;
|
|
491
|
+
this.activeProviderTurnAbort = undefined;
|
|
492
|
+
this.pendingSoftInterrupt = false;
|
|
493
|
+
this.pendingSteers = [];
|
|
494
|
+
this.pendingSteerBytes = 0;
|
|
431
495
|
try {
|
|
432
496
|
await this.drainLedger();
|
|
433
497
|
if (this.activeLedger) {
|
|
@@ -678,7 +742,7 @@ class RuntimeAgentSession {
|
|
|
678
742
|
return await this.generateProviderTurn(request, runId, signal, secrets, turn, attempt, recordUsage);
|
|
679
743
|
}
|
|
680
744
|
catch (error) {
|
|
681
|
-
if (error instanceof GuardrailError)
|
|
745
|
+
if (error instanceof GuardrailError || isSteerSoftInterrupt(error))
|
|
682
746
|
throw error;
|
|
683
747
|
const failure = error instanceof ProviderTurnFailure ? error : undefined;
|
|
684
748
|
const info = failure ? redactSecrets(failure.info, secrets) : errorToErrorInfo(error, secrets);
|
|
@@ -730,9 +794,18 @@ class RuntimeAgentSession {
|
|
|
730
794
|
usageRecorded = true;
|
|
731
795
|
await recordUsage?.(usage, turn, attempt);
|
|
732
796
|
};
|
|
797
|
+
const turnAbort = new AbortController();
|
|
798
|
+
const cleanupTurn = bridgeAbort(signal, turnAbort);
|
|
799
|
+
this.activeProviderTurnAbort = turnAbort;
|
|
800
|
+
if (this.pendingSoftInterrupt) {
|
|
801
|
+
this.pendingSoftInterrupt = false;
|
|
802
|
+
turnAbort.abort(new SteerSoftInterrupt());
|
|
803
|
+
}
|
|
804
|
+
const turnRequest = { ...request, signal: turnAbort.signal };
|
|
733
805
|
try {
|
|
734
|
-
|
|
735
|
-
|
|
806
|
+
throwIfAborted(turnAbort.signal);
|
|
807
|
+
for await (const event of this.activeProvider.generate(turnRequest)) {
|
|
808
|
+
throwIfAborted(turnAbort.signal);
|
|
736
809
|
this.activeLimits.charge("maxResponseBytes", jsonBytes(event));
|
|
737
810
|
if (event.type === "error")
|
|
738
811
|
throw new ProviderTurnFailure(event.error, started);
|
|
@@ -776,7 +849,7 @@ class RuntimeAgentSession {
|
|
|
776
849
|
stage: "output",
|
|
777
850
|
guardrails: this.activeGuardrails,
|
|
778
851
|
value: { content, calls, messageId, started, usage },
|
|
779
|
-
context: { sessionId: this.id, runId, metadata: this.activeMetadata ?? {}, signal },
|
|
852
|
+
context: { sessionId: this.id, runId, metadata: this.activeMetadata ?? {}, signal: turnAbort.signal },
|
|
780
853
|
redactor: this.activeRedactor,
|
|
781
854
|
emit: (event) => this.emit(event),
|
|
782
855
|
}));
|
|
@@ -796,6 +869,19 @@ class RuntimeAgentSession {
|
|
|
796
869
|
return { content, calls, messageId, started, usage };
|
|
797
870
|
}
|
|
798
871
|
catch (error) {
|
|
872
|
+
if (isSteerSoftInterrupt(error) || isSteerSoftInterrupt(turnAbort.signal.reason)) {
|
|
873
|
+
await recordTurnUsage();
|
|
874
|
+
const latencyMs = Math.round(performance.now() - startedAt);
|
|
875
|
+
this.emit({
|
|
876
|
+
type: "provider_turn_finished",
|
|
877
|
+
sessionId: this.id,
|
|
878
|
+
runId,
|
|
879
|
+
turn,
|
|
880
|
+
metadata: buildMetadata({ latencyMs }),
|
|
881
|
+
usage,
|
|
882
|
+
});
|
|
883
|
+
throw new SteerSoftInterrupt();
|
|
884
|
+
}
|
|
799
885
|
const latencyMs = Math.round(performance.now() - startedAt);
|
|
800
886
|
const info = error instanceof ProviderTurnFailure ? redactSecrets(error.info, secrets) : errorToErrorInfo(error, secrets);
|
|
801
887
|
await recordTurnUsage();
|
|
@@ -812,6 +898,33 @@ class RuntimeAgentSession {
|
|
|
812
898
|
throw error;
|
|
813
899
|
throw new ProviderTurnFailure(info, started);
|
|
814
900
|
}
|
|
901
|
+
finally {
|
|
902
|
+
cleanupTurn();
|
|
903
|
+
if (this.activeProviderTurnAbort === turnAbort)
|
|
904
|
+
this.activeProviderTurnAbort = undefined;
|
|
905
|
+
}
|
|
906
|
+
}
|
|
907
|
+
async applyPendingSteers(runId, metadata, signal) {
|
|
908
|
+
if (this.pendingSteers.length === 0)
|
|
909
|
+
return false;
|
|
910
|
+
const drained = this.pendingSteers.splice(0);
|
|
911
|
+
this.pendingSteerBytes = 0;
|
|
912
|
+
this.emit({ type: "queue_updated", sessionId: this.id, runId, size: 0 });
|
|
913
|
+
for (const message of drained) {
|
|
914
|
+
throwIfAborted(signal);
|
|
915
|
+
const inputGuardrails = await runGuardrails({
|
|
916
|
+
stage: "input",
|
|
917
|
+
guardrails: this.activeGuardrails,
|
|
918
|
+
value: [message],
|
|
919
|
+
context: { sessionId: this.id, runId, metadata, signal },
|
|
920
|
+
redactor: this.activeRedactor,
|
|
921
|
+
emit: (event) => this.emit(event),
|
|
922
|
+
});
|
|
923
|
+
assertGuardrailsAllowed(inputGuardrails);
|
|
924
|
+
this.history.push(message);
|
|
925
|
+
await this.appendMessage(message, runId);
|
|
926
|
+
}
|
|
927
|
+
return true;
|
|
815
928
|
}
|
|
816
929
|
async applyProviderRequestPolicies(request, runId, options, metadata, signal) {
|
|
817
930
|
const policies = [...policyList(this.agent.config.providerRequestPolicies), ...policyList(options.providerRequestPolicies)];
|
|
@@ -972,6 +1085,27 @@ function inputToMessages(input) {
|
|
|
972
1085
|
return [input];
|
|
973
1086
|
return [...input];
|
|
974
1087
|
}
|
|
1088
|
+
const steerTextEncoder = new TextEncoder();
|
|
1089
|
+
function messageTextBytes(message) {
|
|
1090
|
+
let total = 0;
|
|
1091
|
+
for (const block of message.content) {
|
|
1092
|
+
if (block.type === "text")
|
|
1093
|
+
total += steerTextEncoder.encode(block.text).byteLength;
|
|
1094
|
+
}
|
|
1095
|
+
return total;
|
|
1096
|
+
}
|
|
1097
|
+
const STEER_SOFT_INTERRUPT_CODE = "steer_soft_interrupt";
|
|
1098
|
+
class SteerSoftInterrupt extends Error {
|
|
1099
|
+
code = STEER_SOFT_INTERRUPT_CODE;
|
|
1100
|
+
constructor() {
|
|
1101
|
+
super("Provider turn soft-interrupted by steer");
|
|
1102
|
+
this.name = "SteerSoftInterrupt";
|
|
1103
|
+
}
|
|
1104
|
+
}
|
|
1105
|
+
function isSteerSoftInterrupt(error) {
|
|
1106
|
+
return error instanceof SteerSoftInterrupt
|
|
1107
|
+
|| (typeof error === "object" && error !== null && error.code === STEER_SOFT_INTERRUPT_CODE);
|
|
1108
|
+
}
|
|
975
1109
|
function finalAssistantMessage(history) {
|
|
976
1110
|
for (let index = history.length - 1; index >= 0; index -= 1) {
|
|
977
1111
|
const message = history[index];
|
|
@@ -999,7 +1133,7 @@ function policyList(policies) {
|
|
|
999
1133
|
return "apply" in policies ? [policies] : policies;
|
|
1000
1134
|
}
|
|
1001
1135
|
function errorFromInfo(error) {
|
|
1002
|
-
return Object.assign(new Error(error.message), { name: error.name ?? "Error", cause: error.cause });
|
|
1136
|
+
return Object.assign(new Error(error.message), { name: error.name ?? "Error", cause: error.cause, code: error.code });
|
|
1003
1137
|
}
|
|
1004
1138
|
class ProviderTurnFailure extends Error {
|
|
1005
1139
|
info;
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
import type { ContextBlock, InputAssemblyLayout, Message, ProviderRequest, Skill, ToolDefinition } from "./contracts.js";
|
|
2
|
+
/** Assembler-time input budget. At least one max required when present. */
|
|
3
|
+
export interface ContextBudget {
|
|
4
|
+
readonly maxInputTokens?: number;
|
|
5
|
+
readonly maxInputBytes?: number;
|
|
6
|
+
readonly reportOmissions?: boolean;
|
|
7
|
+
}
|
|
8
|
+
export type ContextBudgetOmissionKind = "skills" | "context" | "history" | "tool_results" | "summaries" | "attachments" | "tools";
|
|
9
|
+
export interface ContextBudgetOmission {
|
|
10
|
+
readonly kind: ContextBudgetOmissionKind;
|
|
11
|
+
readonly id?: string;
|
|
12
|
+
readonly tokenEstimate: number;
|
|
13
|
+
readonly byteLength: number;
|
|
14
|
+
}
|
|
15
|
+
export interface ContextBudgetReport {
|
|
16
|
+
readonly omitted: readonly ContextBudgetOmission[];
|
|
17
|
+
readonly keptTokens: number;
|
|
18
|
+
readonly keptBytes: number;
|
|
19
|
+
readonly maxInputTokens?: number;
|
|
20
|
+
readonly maxInputBytes?: number;
|
|
21
|
+
readonly truncated: boolean;
|
|
22
|
+
}
|
|
23
|
+
export interface ContextBudgetMessageGroups {
|
|
24
|
+
readonly instructions: readonly Message[];
|
|
25
|
+
readonly summaries: readonly Message[];
|
|
26
|
+
readonly history: readonly Message[];
|
|
27
|
+
readonly input: readonly Message[];
|
|
28
|
+
readonly attachments: readonly Message[];
|
|
29
|
+
readonly toolResults: readonly Message[];
|
|
30
|
+
}
|
|
31
|
+
export declare const CONTEXT_BUDGET_REPORT_METADATA_KEY: "contextBudgetReport";
|
|
32
|
+
export declare const HARD_MAX_CONTEXT_BUDGET_TOKENS = 2000000;
|
|
33
|
+
export declare const HARD_MAX_CONTEXT_BUDGET_BYTES: number;
|
|
34
|
+
export declare const DEFAULT_MAX_CONTEXT_BUDGET_OMISSIONS = 256;
|
|
35
|
+
export declare const HARD_MAX_CONTEXT_BUDGET_OMISSIONS = 1024;
|
|
36
|
+
export declare const CONTEXT_BUDGET_ERROR_CODE: "context_budget_exceeded";
|
|
37
|
+
export declare class ContextBudgetError extends Error {
|
|
38
|
+
readonly code: "context_budget_exceeded";
|
|
39
|
+
constructor(message?: string);
|
|
40
|
+
}
|
|
41
|
+
export declare function isContextBudgetError(error: unknown): error is ContextBudgetError;
|
|
42
|
+
/** UTF-16 code units / 4. Estimate only — not billing. */
|
|
43
|
+
export declare function estimateTextTokens(text: string): number;
|
|
44
|
+
export declare function estimateTextBytes(text: string): number;
|
|
45
|
+
export declare function estimateMessageTokens(message: Message): number;
|
|
46
|
+
export declare function estimateMessageBytes(message: Message): number;
|
|
47
|
+
export declare function estimateAssemblyTokens(messages: readonly Message[]): number;
|
|
48
|
+
export declare function resolveContextBudget(budget: ContextBudget): Required<Pick<ContextBudget, "reportOmissions">> & ContextBudget;
|
|
49
|
+
export declare function getContextBudgetReport(request: ProviderRequest): ContextBudgetReport | undefined;
|
|
50
|
+
export declare function applyContextBudget(options: {
|
|
51
|
+
readonly groups: ContextBudgetMessageGroups;
|
|
52
|
+
readonly context?: readonly ContextBlock[];
|
|
53
|
+
readonly skills?: readonly Skill[];
|
|
54
|
+
readonly tools?: readonly ToolDefinition[];
|
|
55
|
+
readonly budget: ContextBudget;
|
|
56
|
+
readonly layout?: InputAssemblyLayout;
|
|
57
|
+
}): {
|
|
58
|
+
readonly groups: ContextBudgetMessageGroups;
|
|
59
|
+
readonly context: readonly ContextBlock[];
|
|
60
|
+
readonly skills: readonly Skill[];
|
|
61
|
+
readonly tools: readonly ToolDefinition[] | undefined;
|
|
62
|
+
readonly report: ContextBudgetReport;
|
|
63
|
+
};
|