@orkestrel/scaffold 0.0.67 → 0.0.68
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/bin/main.js +67 -44
- package/dist/bin/main.js.map +1 -1
- package/dist/host/agents/templates/brief.md +9 -0
- package/dist/host/claude/agents/orkestrel.md +4 -4
- package/dist/host/claude/rules/names.md +15 -0
- package/dist/host/claude/rules/tests.md +33 -4
- package/dist/host/claude/rules/workspace.md +14 -2
- package/dist/host/dotfiles/prettierignore +3 -0
- package/dist/host/guides/README.md +65 -0
- package/dist/host/guides/abort.md +169 -0
- package/dist/host/guides/agent.md +1509 -0
- package/dist/host/guides/brief.md +1266 -0
- package/dist/host/guides/browser.md +2200 -0
- package/dist/host/guides/budget.md +196 -0
- package/dist/host/guides/codec.md +519 -0
- package/dist/host/guides/console.md +785 -0
- package/dist/host/guides/contract.md +1193 -0
- package/dist/host/guides/csv.md +541 -0
- package/dist/host/guides/database.md +2518 -0
- package/dist/host/guides/emitter.md +233 -0
- package/dist/host/guides/form.md +1791 -0
- package/dist/host/guides/html.md +717 -0
- package/dist/host/guides/indexeddb.md +505 -0
- package/dist/host/guides/interpret.md +1029 -0
- package/dist/host/guides/lsp.md +515 -0
- package/dist/host/guides/markdown.md +964 -0
- package/dist/host/guides/mcp.md +5554 -0
- package/dist/host/guides/middleware.md +927 -0
- package/dist/host/guides/msg.md +440 -0
- package/dist/host/guides/ndjson.md +120 -0
- package/dist/host/guides/ollama.md +380 -0
- package/dist/host/guides/pool.md +280 -0
- package/dist/host/guides/probe.md +1210 -0
- package/dist/host/guides/process.md +1620 -0
- package/dist/host/guides/program.md +1110 -0
- package/dist/host/guides/qualifier.md +854 -0
- package/dist/host/guides/queue.md +370 -0
- package/dist/host/guides/rater.md +330 -0
- package/dist/host/guides/reason.md +1122 -0
- package/dist/host/guides/relation.md +373 -0
- package/dist/host/guides/router.md +753 -0
- package/dist/host/guides/scaffold.md +192 -31
- package/dist/host/guides/sea.md +383 -0
- package/dist/host/guides/server.md +752 -0
- package/dist/host/guides/sqlite.md +330 -0
- package/dist/host/guides/sse.md +187 -0
- package/dist/host/guides/supervisor.md +4890 -0
- package/dist/host/guides/table.md +1556 -0
- package/dist/host/guides/template.md +280 -0
- package/dist/host/guides/terminal.md +1145 -0
- package/dist/host/guides/test.md +2969 -0
- package/dist/host/guides/timeout.md +252 -0
- package/dist/host/guides/tool.md +311 -0
- package/dist/host/guides/toolbox.md +1038 -0
- package/dist/host/guides/websocket.md +282 -0
- package/dist/host/guides/worker.md +615 -0
- package/dist/host/guides/workflow.md +1507 -0
- package/dist/host/guides/workspace.md +595 -0
- package/dist/host/manifest.json +1218 -10
- package/dist/host/tests/policy.test.ts +279 -2
- package/dist/host/tests/setupPolicy.ts +437 -6
- package/dist/src/core/index.cjs +38 -16
- package/dist/src/core/index.cjs.map +1 -1
- package/dist/src/core/index.d.cts +33 -9
- package/dist/src/core/index.d.ts +33 -9
- package/dist/src/core/index.js +37 -17
- package/dist/src/core/index.js.map +1 -1
- package/dist/src/server/index.cjs +1750 -1567
- package/dist/src/server/index.cjs.map +1 -1
- package/dist/src/server/index.d.cts +106 -24
- package/dist/src/server/index.d.ts +106 -24
- package/dist/src/server/index.js +1751 -1570
- package/dist/src/server/index.js.map +1 -1
- package/package.json +3 -3
|
@@ -0,0 +1,1509 @@
|
|
|
1
|
+
# Agent
|
|
2
|
+
|
|
3
|
+
> The conversation runtime for the `@orkestrel` line: the `ProviderInterface` inference
|
|
4
|
+
> boundary with the host-independent HTTP engine and browser relay that implement it, the
|
|
5
|
+
> conversation layer that feeds it — messages, compaction, instructions, scopes, and prompt
|
|
6
|
+
> assembly — and the bounded context → provider → tools → repeat loop that carries a turn to
|
|
7
|
+
> its end.
|
|
8
|
+
|
|
9
|
+
An agent is a conversation with a model and the loop that carries it forward. A `Conversation` holds the history — a live tail of immutable messages plus the sections older turns were compacted into. An `AgentContext` assembles that history into the next prompt, folding in the instructions and the active workspace and applying the active scope. An `Agent` drives that prompt through a provider, dispatches whatever tools the model asks for, feeds the results back, and repeats until the model stops. Everything else in this module either configures those nouns or observes them. Source: [`src/core`](../src/core). Published through `@orkestrel/agent`.
|
|
10
|
+
|
|
11
|
+
Hand a provider a conversation and get back one assembled `ProviderResult` (`generate`), or a live stream of channel-tagged `ProviderDelta`s that returns that same assembled result when it ends (`stream`). Reasoning separation, the authority gate, and durable jobs sit around that boundary. The model itself is the one thing this package does not supply. `ProviderInterface` stays the boundary any backend can satisfy, and `AgentProvider` is the host-independent HTTP engine a concrete provider extends rather than rewrites: it owns the deadline, the transport, the bounded error read, the framing loop, reasoning separation, and result assembly, and the subclass fills in one vendor's wire. `RelayProvider` and `createRelay` carry that same boundary across your own server, so a browser drives a model it holds no credential for. There is no hidden global state, a plugin lifecycle, a prompt-template DSL, or an implicit memory store. This is a kit of composable primitives: the loop is the convenient way to use them, not the only one, and a caller that would rather bound and drive a provider by hand can skip it entirely.
|
|
12
|
+
|
|
13
|
+
Tools and files are borrowed, not owned. Callable tools come from [`@orkestrel/tool`](tool.md): the loop advertises their definitions to the model, dispatches the calls that come back, and feeds each `ToolResult` in as a tool message. A tool is loop machinery — it is never rendered into the prompt. Documents come from [`@orkestrel/workspace`](workspace.md): the context renders the active workspace into every turn, split by carrier — text as fenced reference blocks in the system message, images attached to the last user turn. That split is this package's own product policy, decided here because only the prompt-assembly layer knows what a turn looks like.
|
|
14
|
+
|
|
15
|
+
A turn is bounded and always terminates. One `AbortSignal` — a cancel, a [timeout](timeout.md), and a [budget](budget.md) folded together through `AbortSignal.any` — bounds the whole run, and tool iteration is capped at `limit`. A cancel is not an error: it commits a partial `AgentResult` that resolves, so only a genuine provider or tool failure rejects. `generate` and `stream` share one private run, so the one-shot result can never diverge from the live stream, and a buggy observer cannot corrupt either, because the emitter isolates a listener's throw.
|
|
16
|
+
|
|
17
|
+
## Surface
|
|
18
|
+
|
|
19
|
+
The agent-owned surface: the inference boundary and the HTTP engine behind it, the relay that carries that boundary to a browser, the conversation layer, the context and its managers, the loop, the authority gate, and the durable-job bridge. Tool and workspace entities belong to their originating packages and are consumed directly — never re-exported here — and one vendor's wire belongs to the concrete provider that extends `AgentProvider`.
|
|
20
|
+
|
|
21
|
+
A provider turns a conversation (plus optional tools) into a turn: `generate` resolves the assembled `ProviderResult` (content + any tool calls + any usage); `stream` yields channel-tagged `ProviderDelta`s as they arrive (`content` for answer text, `thinking` for live reasoning) and returns the same assembled result when the stream completes, so a caller can render tokens / reasoning live and still get the full outcome. Both bound the call with an `AbortSignal`:
|
|
22
|
+
|
|
23
|
+
```ts
|
|
24
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
25
|
+
import { createAbort } from '@orkestrel/abort'
|
|
26
|
+
|
|
27
|
+
declare const provider: ProviderInterface // any concrete implementation supplied by the host app
|
|
28
|
+
const abort = createAbort()
|
|
29
|
+
const messages = [{ id: '1', role: 'user', content: 'Say hello.' }] as const
|
|
30
|
+
|
|
31
|
+
const result = await provider.generate(messages, abort.signal)
|
|
32
|
+
result.content // the assembled content
|
|
33
|
+
result.usage // { prompt, completion, total } — folds into a token budget
|
|
34
|
+
|
|
35
|
+
const generator = provider.stream(messages, abort.signal)
|
|
36
|
+
let step = await generator.next()
|
|
37
|
+
while (!step.done) {
|
|
38
|
+
if (step.value.channel === 'content') process.stdout.write(step.value.text)
|
|
39
|
+
if (step.value.channel === 'thinking') process.stderr.write(step.value.text)
|
|
40
|
+
step = await generator.next()
|
|
41
|
+
}
|
|
42
|
+
const streamed = step.value // the assembled ProviderResult (content === the joined content deltas)
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Pass `tools` (a non-empty `ToolDefinition[]`) to advertise callable tools for the turn; when the model calls one, `result.tools` is a `ToolCall[]` (each with a guaranteed `id`, the tool `name`, and parsed `arguments`). Aborting a `stream` mid-flight throws a `ProviderAbortError` whose `partial` holds whatever streamed before the cancel.
|
|
46
|
+
|
|
47
|
+
A tool is a JSON-Schema-described callable from [`@orkestrel/tool`](tool.md), and the loop needs exactly this from its registry: `definitions()` advertises the tools to the model, and `execute` dispatches the `ToolCall`s that come back. What returns is a `ToolResult` discriminated on `success` — an unknown name and a throwing handler both arrive as the failure arm rather than as an exception, and a batch isolates each call from its siblings, which is what lets the loop hand every outcome back to the model and let it react:
|
|
48
|
+
|
|
49
|
+
```ts
|
|
50
|
+
import { createTool, createToolManager } from '@orkestrel/tool'
|
|
51
|
+
|
|
52
|
+
const tools = createToolManager()
|
|
53
|
+
tools.add([
|
|
54
|
+
createTool({
|
|
55
|
+
name: 'add',
|
|
56
|
+
description: 'Add two numbers',
|
|
57
|
+
parameters: { type: 'object', properties: { a: { type: 'number' }, b: { type: 'number' } } },
|
|
58
|
+
execute: (args) => Number(args.a) + Number(args.b), // narrow the model-supplied unknown
|
|
59
|
+
}),
|
|
60
|
+
createTool({ name: 'now', execute: () => Date.now() }),
|
|
61
|
+
])
|
|
62
|
+
|
|
63
|
+
const definitions = tools.definitions() // hand these to provider.generate / .stream as `tools`
|
|
64
|
+
const results = await tools.execute([
|
|
65
|
+
{ id: '1', name: 'add', arguments: { a: 2, b: 3 } }, // → { success: true, id: '1', name: 'add', value: 5 }
|
|
66
|
+
{ id: '2', name: 'ghost', arguments: {} }, // → { success: false, id: '2', name: 'ghost', error: 'tool not found: ghost' }
|
|
67
|
+
])
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
Contained failure is the registry's contract, not a limitation of it: in-process code that wants a typed error calls the tool itself — `tools.tool(name)` then `tool.execute(args)` inside its own `try`/`catch`. Registration, advertising, dispatch, and error containment are documented in [`tool.md`](tool.md).
|
|
71
|
+
|
|
72
|
+
Collect a turn's conversation in an `AgentContext`. Add turns through `context.messages` — the active conversation's live tail, always present, satisfying `MessageManagerInterface` by minting each `id` on `add` and keeping stored messages immutable and in insertion order — then `build()` the provider input: `[systemMessage?, ...messages]`. `context.tools` sits beside them, but it is a different kind of thing: the other managers assemble prompt text, while the tool registry exists so the loop can advertise definitions and dispatch calls. Its contents reach the model as the `tools` argument, never as a message:
|
|
73
|
+
|
|
74
|
+
```ts
|
|
75
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
76
|
+
import { createAgentContext } from '@orkestrel/agent'
|
|
77
|
+
import { createAbort } from '@orkestrel/abort'
|
|
78
|
+
import { createToolManager } from '@orkestrel/tool'
|
|
79
|
+
|
|
80
|
+
declare const provider: ProviderInterface
|
|
81
|
+
const abort = createAbort()
|
|
82
|
+
const context = createAgentContext({ system: 'You are concise.', tools: createToolManager() })
|
|
83
|
+
context.messages.add([
|
|
84
|
+
{ role: 'user', content: 'What is 2 + 3?' }, // the `id` is minted by add, not supplied
|
|
85
|
+
{ role: 'user', content: 'Reply with just the number.' },
|
|
86
|
+
])
|
|
87
|
+
|
|
88
|
+
const input = context.build() // [{ role: 'system', content: 'You are concise.' }, …the two user turns]
|
|
89
|
+
const definitions = context.tools.definitions() // tools reach the provider here, NOT in `input`
|
|
90
|
+
const result = await provider.generate(input, abort.signal, definitions)
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
`context.messages.add` mints each message's `id` (a random UUID) and returns the created message(s); `build()` is computed fresh on every call, so it always reflects the current conversation. Without a system prompt, `build()` is only the conversation, and no tool's name, description, or parameter schema ever appears in its output.
|
|
94
|
+
|
|
95
|
+
Drive the whole turn with an `Agent` (`createAgent`) — it composes the provider, its `AgentContext`, and the tool registry into the bounded context → provider → tools → repeat loop. Seed the conversation through `agent.context.messages`, then either `generate()` for a one-shot `AgentResult` or `stream()` for a live `AgentChunk` stream (`token` answer deltas, `think` reasoning deltas, `tool` dispatches, `usage`) whose `result` resolves the same `AgentResult`. `generate` drains that same stream, so they can't diverge:
|
|
96
|
+
|
|
97
|
+
```ts
|
|
98
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
99
|
+
import { createAgent } from '@orkestrel/agent'
|
|
100
|
+
import { createTokenBudget } from '@orkestrel/budget'
|
|
101
|
+
import { createTool, createToolManager } from '@orkestrel/tool'
|
|
102
|
+
|
|
103
|
+
declare const provider: ProviderInterface
|
|
104
|
+
const tools = createToolManager()
|
|
105
|
+
tools.add(createTool({ name: 'add', execute: (args) => Number(args.a) + Number(args.b) }))
|
|
106
|
+
|
|
107
|
+
const agent = createAgent(provider, {
|
|
108
|
+
system: 'You are concise.',
|
|
109
|
+
tools,
|
|
110
|
+
limit: 4, // cap tool iterations
|
|
111
|
+
timeout: 30_000, // wall-clock deadline for the whole turn
|
|
112
|
+
budget: createTokenBudget({ max: 50_000, scope: 'total' }), // cost ceiling
|
|
113
|
+
})
|
|
114
|
+
agent.context.messages.add({ role: 'user', content: 'Use the add tool to add 2 and 3.' })
|
|
115
|
+
|
|
116
|
+
const stream = agent.stream()
|
|
117
|
+
for await (const chunk of stream.events) {
|
|
118
|
+
if (chunk.category === 'token') process.stdout.write(chunk.content) // live deltas
|
|
119
|
+
if (chunk.category === 'think') process.stderr.write(chunk.content) // live reasoning
|
|
120
|
+
if (chunk.category === 'tool') log(chunk.call, chunk.result) // a dispatched tool + its result
|
|
121
|
+
}
|
|
122
|
+
const result = await stream.result // { content, usage?, partial } — usage summed across the turn
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Both `generate` and `stream` accept optional per-run `AgentRunOptions` — `think` and `schema` (forwarded to the provider as `ProviderStreamOptions`) plus `limit` / `timeout` / `budget` / `signal`, each overriding its `AgentOptions` construction default for this run only. Omitting one keeps the constructed default, so a caller that passes no options gets the agent it configured. A per-run `signal` composes with (never replaces) a constructed `signal` — either aborting cancels the run; a per-run `budget` is `start()`ed for that run and is the one the loop charges, leaving a constructed `budget` untouched:
|
|
126
|
+
|
|
127
|
+
```ts
|
|
128
|
+
const agent = createAgent(provider, { tools, limit: 10, timeout: 60_000 }) // construction defaults
|
|
129
|
+
agent.context.messages.add({ role: 'user', content: 'Summarize this doc.' })
|
|
130
|
+
|
|
131
|
+
// A tighter, structured-output run -- overrides limit + timeout, adds a schema, for THIS call only.
|
|
132
|
+
const result = await agent.generate({
|
|
133
|
+
limit: 2,
|
|
134
|
+
timeout: 5_000,
|
|
135
|
+
schema: { type: 'object', properties: { summary: { type: 'string' } } },
|
|
136
|
+
})
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
`schema`, like `think`, rides into `provider.stream` as a `ProviderStreamOptions`: the loop composes both into one options object, omitting whichever key is unset, and passes no options object at all when neither is present — a provider that never received one still never does.
|
|
140
|
+
|
|
141
|
+
The turn is bounded by one cancel folded from the external `signal` + the `timeout` deadline + the `budget` signal through `AbortSignal.any`; `agent.abort(reason)` (or `stream.abort(reason)`) fires it. A cancel — external, deadline, budget, or `abort()` — commits a partial result: the `result` promise resolves with `{ partial: true, content: <what accumulated> }`, never rejects (a cancel is not an error); only a genuine provider / tool error rejects. An optional `scheduler.yield`s between turns; tool iteration is capped at `limit` (default `DEFAULT_AGENT_LIMIT`). Exhausting `limit` while the model still holds unresolved tool intent (it requested tools on the very last allowed turn) is a distinct, non-cancel cause of `partial: true` — it fires an `exhaust` event (the turns reached) instead of `abort`. A natural finish on the last allowed turn, or `limit: 0` (which never enters the loop), stays `partial: false`. `agent.status` transitions `idle` → `running` → `done` / `error`.
|
|
142
|
+
|
|
143
|
+
Gate the model's tool calls with an optional `Authority` (`createAuthority`) — a synchronous policy gate the loop consults before each call runs, passed through `AgentOptions.authority`. It walks ordered `rules` first-match-wins (a matched rule allows unless its `allowed` is `false`), falling back to a configurable default when none match — allow-unmatched by default (a denylist), or deny-by-default when its `fallback` denies (an allowlist). A denied call is never executed: the loop synthesizes the failure arm of `ToolResult` (`error: 'denied: <reason>'`) and feeds it back as a `tool` chunk and a tool message, so no handler runs, no budget is spent, and the model still sees what happened and can choose something else. An allowed call dispatches normally, and with no `authority` set every call dispatches:
|
|
144
|
+
|
|
145
|
+
```ts
|
|
146
|
+
import { createAgent, createAuthority } from '@orkestrel/agent'
|
|
147
|
+
import { createTool, createToolManager } from '@orkestrel/tool'
|
|
148
|
+
|
|
149
|
+
const tools = createToolManager()
|
|
150
|
+
tools.add([
|
|
151
|
+
createTool({ name: 'add', execute: (args) => Number(args.a) + Number(args.b) }),
|
|
152
|
+
createTool({ name: 'delete', execute: (args) => drop(args.id) }),
|
|
153
|
+
])
|
|
154
|
+
|
|
155
|
+
// A denylist: deny `delete`, allow everything else (the default allow fallback).
|
|
156
|
+
const authority = createAuthority({
|
|
157
|
+
rules: [
|
|
158
|
+
{
|
|
159
|
+
match: (c) => c.call.name === 'delete',
|
|
160
|
+
zone: 'restricted',
|
|
161
|
+
allowed: false,
|
|
162
|
+
reason: 'read-only mode',
|
|
163
|
+
},
|
|
164
|
+
],
|
|
165
|
+
})
|
|
166
|
+
const agent = createAgent(provider, { tools, authority })
|
|
167
|
+
agent.context.messages.add({ role: 'user', content: 'Delete record 42.' })
|
|
168
|
+
// When the model calls `delete`, the loop feeds back { error: 'denied: read-only mode' } — never runs it.
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
Alongside the conversation store sits the standalone `InstructionManager` a richer context assembles a prompt from — named directives, keyed by `name`, listed by descending `priority`. It mirrors the registry shape — `add` (one or a batch) mints each `id` and overwrites a same-key entry (last write wins), an `instruction(name)` / `instructions()` accessor pair, `remove` (one or a batch) / `clear` / `count` — holds immutable entries, and is observable (`emitter` with an `add` / `remove` / `clear` event map, wired through the reserved `on` option; an `error` option receives a listener's throw). It carries the **build-contract** members a context's assembly step calls: `open` (the section header text, `'## Instructions'`) and `render(instruction)` (per-item rendering — the instruction's `content`):
|
|
172
|
+
|
|
173
|
+
```ts
|
|
174
|
+
import { createInstructionManager } from '@orkestrel/agent'
|
|
175
|
+
|
|
176
|
+
const instructions = createInstructionManager()
|
|
177
|
+
const safety = instructions.add({
|
|
178
|
+
name: 'safety',
|
|
179
|
+
content: 'Refuse unsafe requests.',
|
|
180
|
+
priority: 10,
|
|
181
|
+
})
|
|
182
|
+
instructions.open // '## Instructions'
|
|
183
|
+
instructions.render(safety) // 'Refuse unsafe requests.'
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
Documents reach a turn one way only: through the active workspace. A [`@orkestrel/workspace`](workspace.md) workspace is a flat map of immutable files, and it takes no position on prompts — deciding how a file becomes part of a turn is this package's job, and the decision is a split by carrier. A text file renders as a fenced reference block in a `## Workspace` system section, where the model can read it as quoted material. An image file cannot be text, so its `base64` payload attaches to the last user message instead, which is where a vision model looks. A message carries that payload on its optional `images` field — `Message` and `MessageInput` both accept `readonly images?: readonly string[]` — and a vision-capable provider forwards it onto the wire (an empty or absent array is never sent). It is input-only; `ProviderResult` is unchanged.
|
|
187
|
+
|
|
188
|
+
```ts
|
|
189
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
190
|
+
|
|
191
|
+
declare const provider: ProviderInterface // a vision-capable model
|
|
192
|
+
const result = await provider.generate(
|
|
193
|
+
[{ id: '1', role: 'user', content: 'Describe this image.', images: ['<payload>'] }],
|
|
194
|
+
abort.signal,
|
|
195
|
+
)
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
`AgentContext` wires the instruction manager and the workspace registry in. Beyond `system`, `messages`, and `tools`, a context exposes its own `instructions` manager and `workspaces` registry — pass pre-built ones through `AgentContextOptions`, or fresh empty ones are created — and `build()` folds them into the turn. The assembly order is one leading `system` message holding the system prompt, then the non-empty instructions block (its `description` header followed by every item's `format`), then the active workspace's text files under a `## Workspace` header, joined by blank lines; then the conversation. With no instructions, no active workspace, and no scope, that reduces to exactly the lean `[systemMessage?, ...messages]`. The carrier split shows here: text rides the system block, image data rides the last user message.
|
|
199
|
+
|
|
200
|
+
````ts
|
|
201
|
+
import { createAgentContext } from '@orkestrel/agent'
|
|
202
|
+
|
|
203
|
+
const context = createAgentContext({ system: 'You are a code reviewer.' })
|
|
204
|
+
context.instructions.add({ name: 'tone', content: 'Be terse.', priority: 10 })
|
|
205
|
+
context.workspaces.add().write('src/main.ts', 'export const x = 1') // the active workspace
|
|
206
|
+
context.messages.add({ role: 'user', content: 'Review this.' })
|
|
207
|
+
|
|
208
|
+
const input = context.build()
|
|
209
|
+
// input[0] = { role: 'system', content:
|
|
210
|
+
// 'You are a code reviewer.\n\n## Instructions\n\nBe terse.\n\n## Workspace\n\nFile: src/main.ts\n```typescript\nexport const x = 1\n```' }
|
|
211
|
+
// input[1] = { role: 'user', content: 'Review this.' }
|
|
212
|
+
````
|
|
213
|
+
|
|
214
|
+
### Conversations & compaction
|
|
215
|
+
|
|
216
|
+
Above the flat `MessageManagerInterface` sits the `Conversation` (`createConversation` / a `ConversationManager`) — it owns its messages directly and compacts older ones into summarized `sections` so a long history fits a turn's context window without discarding the originals. Append turns through the conversation's own `add` (the live uncompacted tail; `message` / `messages` / `remove` / `clear` / `count` round it out); `compact()` folds the older live messages into a summarized `Section` (retaining their originals), regenerates the conversation rollup `summary`, and shrinks `view()` — the model input, where each section becomes one summary message followed by the live tail. Compaction is driven by a provider-agnostic `ConversationSummaryHandler` seam (`(messages) => Promise<string>`) the agent runtime supplies, so a `compact()` without one throws a `ConversationError`. `keep` retains a recent tail (default `DEFAULT_CONVERSATION_KEEP` = `0`, fold all); `rehydrate(id)` / `search(query)` read the retained originals:
|
|
217
|
+
|
|
218
|
+
```ts
|
|
219
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
220
|
+
import { createConversation } from '@orkestrel/agent'
|
|
221
|
+
|
|
222
|
+
declare const provider: ProviderInterface // any concrete implementation supplied by the host app
|
|
223
|
+
// The summarizer seam — built from the provider by the runtime; core stays provider-agnostic.
|
|
224
|
+
// Append the instruction as the FINAL user turn: a chat model emits nothing when the prompt
|
|
225
|
+
// ends on an assistant turn, so a leading-system instruction is unreliable.
|
|
226
|
+
const conversation = createConversation({
|
|
227
|
+
summarize: async (messages) =>
|
|
228
|
+
(
|
|
229
|
+
await provider.generate(
|
|
230
|
+
[
|
|
231
|
+
...messages,
|
|
232
|
+
{ id: 's', role: 'user', content: 'Summarize the conversation so far concisely.' },
|
|
233
|
+
],
|
|
234
|
+
AbortSignal.timeout(30_000),
|
|
235
|
+
)
|
|
236
|
+
).content,
|
|
237
|
+
keep: 2, // retain the two most recent turns verbatim on each compaction
|
|
238
|
+
})
|
|
239
|
+
conversation.add([
|
|
240
|
+
{ role: 'user', content: 'My name is Ada.' },
|
|
241
|
+
{ role: 'assistant', content: 'Nice to meet you, Ada.' },
|
|
242
|
+
{ role: 'user', content: 'What did I say my name was?' },
|
|
243
|
+
])
|
|
244
|
+
|
|
245
|
+
const section = await conversation.compact() // folds the older turns → a summarized section
|
|
246
|
+
conversation.view() // [<section summary message>, ...the retained recent tail] — the model input
|
|
247
|
+
conversation.summary // the regenerated rollup (a summary-of-summaries over all sections)
|
|
248
|
+
conversation.search('ada') // case-insensitive across sections' originals + the live tail
|
|
249
|
+
section && conversation.rehydrate(section.id) // the section's full original messages (a pure read)
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
Register a conversation in a `ConversationManager` and pass that registry through `AgentContextOptions.conversations` to make it the message source. `context.messages` then is the active conversation's live tail, and `build()` folds its `view()`; the conversation owns message inclusion through compaction, which is why a scope never filters messages. When the registry is omitted, the context creates one with an active default conversation.
|
|
253
|
+
|
|
254
|
+
Pass that registry through `AgentOptions.conversations` together with an `AgentOptions.window` context [`Budget`](budget.md) to enable automatic compaction. Each turn the loop estimates the current full prompt against the window and, when it reaches the ceiling, compacts the active summarizable conversation before continuing on the rebuilt view. Omit `window` and the loop does not auto-compact. See the [Contract](#contract) (the automatic-compaction clause) for the exact trigger and single-level limitation.
|
|
255
|
+
|
|
256
|
+
```ts
|
|
257
|
+
import { createAgent, createConversationManager, estimateMessages } from '@orkestrel/agent'
|
|
258
|
+
import { createBudget } from '@orkestrel/budget'
|
|
259
|
+
|
|
260
|
+
const conversations = createConversationManager({ summarize, keep: 2 })
|
|
261
|
+
const conversation = conversations.add()
|
|
262
|
+
const agent = createAgent(provider, {
|
|
263
|
+
conversations,
|
|
264
|
+
// A context Budget: consumer = a token estimator, max = the context window. The loop measures
|
|
265
|
+
// the CURRENT FULL prompt against it each turn; when the prompt reaches the window it compacts
|
|
266
|
+
// + continues on the rebuilt (smaller) view (compact-and-continue), never aborts.
|
|
267
|
+
window: createBudget({ max: 8_000, consumer: estimateMessages }),
|
|
268
|
+
})
|
|
269
|
+
agent.context.messages.add({ role: 'user', content: 'Hi' })
|
|
270
|
+
await agent.generate() // folds older turns into a section mid-run when the prompt reaches the window, then continues
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
##### One agent, many conversations (switching the active conversation)
|
|
274
|
+
|
|
275
|
+
`agent.context.conversations` is the structural message-source registry — supplied at construction and never reassigned; switch its active conversation with `conversations.switch(id)` to switch the agent's message source. `context.messages` is dynamic: it always points at the current active conversation's live tail (the same reference, no duplication) and follows a switch. The registry always has an active conversation (a default is added when it has none). This is the real app pattern: one `Agent` over a `ConversationManager` of threads, switching the active conversation per request — not an agent per thread. Each conversation accumulates its own history and compacts independently (one thread's sections never leak into another). The agent reads `context.conversations` / `context.messages` fresh on each run, so switching between runs works:
|
|
276
|
+
|
|
277
|
+
```ts
|
|
278
|
+
import { createAgent, createConversationManager, estimateMessages } from '@orkestrel/agent'
|
|
279
|
+
import { createBudget } from '@orkestrel/budget'
|
|
280
|
+
|
|
281
|
+
const threads = createConversationManager({ summarize, keep: 2 }) // its defaults flow into each thread
|
|
282
|
+
const agent = createAgent(provider, {
|
|
283
|
+
conversations: threads, // the agent's message source
|
|
284
|
+
window: createBudget({ max: 8_000, consumer: estimateMessages }),
|
|
285
|
+
})
|
|
286
|
+
|
|
287
|
+
// Per request: make the request's thread active, append the user turn, run.
|
|
288
|
+
async function handle(threadId: string, text: string): Promise<string> {
|
|
289
|
+
if (threads.conversation(threadId) === undefined) threads.add({ id: threadId })
|
|
290
|
+
threads.switch(threadId) // SWITCH — context.messages now IS this thread's tail
|
|
291
|
+
agent.context.messages.add({ role: 'user', content: text })
|
|
292
|
+
return (await agent.generate()).content
|
|
293
|
+
}
|
|
294
|
+
|
|
295
|
+
await handle('user-1', 'Hi, I am Ada.') // thread user-1 accumulates + compacts on its own
|
|
296
|
+
await handle('user-2', 'What is 2 + 2?') // thread user-2 is fully independent
|
|
297
|
+
await handle('user-1', 'What did I say my name was?') // back to user-1 — its own history is intact
|
|
298
|
+
```
|
|
299
|
+
|
|
300
|
+
> **Concurrency caveat.** Switch the active conversation between runs, never during one (the loop reads the active conversation at run entry and drives it through to the end). The framework ships the switch mechanism; the app owns concurrency policy — for threads that must run concurrently, use a separate `Agent` per concurrent thread (each agent is cheap; they can share the provider and tool registry). Switching mid-flight would repoint the live run's message source under it.
|
|
301
|
+
|
|
302
|
+
##### Production behaviors of automatic compaction
|
|
303
|
+
|
|
304
|
+
Auto-compaction (the `window` budget) is hardened for a long-running app:
|
|
305
|
+
|
|
306
|
+
- **Pre-first-turn + run-entry reset.** The budget check runs before the first provider request and between turns — so a resumed or already-long conversation whose initial prompt already exceeds the window compacts immediately (not only after a tool turn). The `window` budget is reset at run entry, so no stale measurement carries across runs or a conversation switch.
|
|
307
|
+
- **Non-fatal, observable summarizer failure.** If the automatic `compact()`'s summarizer throws, the agent run does not crash: the loop skips compaction that turn, surfaces the error as a `fault` event (so the failure is observable, never silently lost), and continues (the over-window prompt proceeds to the provider). Only the agent's auto path is resilient — a manual `conversation.compact()` you call yourself still propagates its error.
|
|
308
|
+
- **Futile-compaction guard (the single-level limit).** If `compact()` folds nothing (returns `undefined`) while the prompt is still over the window — that is, the section summaries alone already exceed it — the loop stops auto-compacting for the rest of that run (a per-run latch), avoiding per-turn churn. The over-window prompt then proceeds to the provider, which surfaces a genuine context-length error if it truly cannot fit — the real limit. Compaction is single-level: it folds the live tail, never the existing sections.
|
|
309
|
+
|
|
310
|
+
```ts
|
|
311
|
+
agent.emitter.on('fault', (error) =>
|
|
312
|
+
log('auto-compaction summarizer failed (run continues)', error),
|
|
313
|
+
)
|
|
314
|
+
```
|
|
315
|
+
|
|
316
|
+
### Scoping a turn
|
|
317
|
+
|
|
318
|
+
A `Scope` (`createScope` / a `ScopeManager`) is a named allow-list filter the context applies at `build()` time and at the loop's tool-advertise step. It carries an optional `instructions` / `tools` / `files` list keyed by each category's identity — `instructions` (by `name`), `tools` (by `name`), `files` (the active workspace's files, by `path`) — each three-way: `undefined` ⇒ no constraint (all pass), `[]` ⇒ none pass, a non-empty list ⇒ only the listed keys. Conversation messages are not scoped — the active conversation owns message inclusion through compaction (`view()`), so there is no `messages` allow-list. Apply the active filter through `context.apply(scope)`; call `context.apply(undefined)` to remove filtering. The readonly `context.scope` getter reports the current filter, and `build()` reflects whatever scope is active when it runs (recomputed fresh each call). `narrow(config)` composes a tighter child scope by set-intersection (an `undefined` side imposes no constraint), so narrowing can only tighten — a parent-excluded key never returns:
|
|
319
|
+
|
|
320
|
+
```ts
|
|
321
|
+
import { createAgent, createScope } from '@orkestrel/agent'
|
|
322
|
+
|
|
323
|
+
const agent = createAgent(provider, { tools }) // tools holds `search` + `delete`
|
|
324
|
+
agent.context.instructions.add([
|
|
325
|
+
{ name: 'safety', content: 'Refuse unsafe requests.' },
|
|
326
|
+
{ name: 'verbose', content: 'Explain every step.' },
|
|
327
|
+
])
|
|
328
|
+
// This turn: only the `safety` instruction, and only the `search` tool.
|
|
329
|
+
agent.context.apply(
|
|
330
|
+
createScope({
|
|
331
|
+
name: 'read-only',
|
|
332
|
+
instructions: ['safety'],
|
|
333
|
+
tools: ['search'],
|
|
334
|
+
}),
|
|
335
|
+
)
|
|
336
|
+
const result = await agent.generate()
|
|
337
|
+
```
|
|
338
|
+
|
|
339
|
+
A scoped-out tool is filtered out of the `definitions()` the loop advertises, which is the only place a tool ever reaches the model — so it is neither described nor callable, not merely hidden. An `undefined` scope, or an `undefined` `tools` list, advertises every registered tool; an empty `tools` list (`[]`) advertises none. A `ScopeManager` is the optional reuse registry for named scopes (keyed by a minted `id`, so two scopes may share a `name`); it is observable like the other managers.
|
|
340
|
+
|
|
341
|
+
### Customizing the format (the cascade)
|
|
342
|
+
|
|
343
|
+
Each context section frames as `[open, ...items.map(render), close]` — a top line rendered once before the items, each item's text, and a bottom line rendered once after — with empty / absent slots dropped and the survivors blank-line (`\n\n`) joined. The slots resolve independently through a cascade, most-specific-first; each level is optional, what you omit falls through to the next, and omitting everything leaves each section on its manager's built-in framing — only the header and the items, with no closing line. From most to least specific:
|
|
344
|
+
|
|
345
|
+
1. **Item override** — `override?: string` on a single `InstructionInput`: a fully-rendered string for that item, round-tripped onto the stored entity. Beats everything for that item's `render`.
|
|
346
|
+
2. **Manager-options override** — `format?: ContextSectionFormat<…>` on a manager's `Options` (an `{ open?; render?; close? }` trio): a per-section open / item-render / close override for that whole manager. Beats the provider default + the built-in. A manager exposes it as the `readonly format` accessor, and its own `open` / `render(item)` already consult that override's `open` / `render` (so a manager used standalone renders with it).
|
|
347
|
+
3. **Provider default** — `format?: ContextFormat` on a `ProviderInterface` (keyed by section kind — the `instructions` section): the model's preferred framing. Optional — an agnostic provider supplies none, and the agent loop passes `provider.format` (often `undefined`) into `build()`.
|
|
348
|
+
4. **Built-in** — the manager's hardcoded `open` getter + `render(item)` method (`## Instructions` + the content) — the floor for `open` and `render`. There is no built-in `close`: an unset `close` yields no closing line.
|
|
349
|
+
|
|
350
|
+
So for a section kind `K`, manager `M`, and a provider format `F`: **open** = `M.format?.open ?? F?.[K]?.open ?? M.open` (manager-options > provider > built-in — the leading text has no per-item level); **per item** `I` = `I.override ?? M.format?.render?.(I) ?? F?.[K]?.render?.(I) ?? M.render(I)` (item > manager-options > provider > built-in); **close** = `M.format?.close ?? F?.[K]?.close` (manager-options > provider, no built-in ⇒ no closing line when unset). `open`, the rendering, and `close` resolve independently, so an override may set only the open, only the rendering, only the close, or any mix — and `open` + `close` together wrap the whole group. (The `## Workspace` text section has no cascade level of its own — it renders with the fixed `renderFencedFile` framing.)
|
|
351
|
+
|
|
352
|
+
```ts
|
|
353
|
+
import { createAgentContext, createInstructionManager } from '@orkestrel/agent'
|
|
354
|
+
|
|
355
|
+
// Manager-options override — wrap the instructions as a closed XML group for this manager.
|
|
356
|
+
const instructions = createInstructionManager({
|
|
357
|
+
format: {
|
|
358
|
+
open: '<rules>',
|
|
359
|
+
render: (one) => `<rule>${one.content}</rule>`,
|
|
360
|
+
close: '</rules>',
|
|
361
|
+
},
|
|
362
|
+
})
|
|
363
|
+
const context = createAgentContext({ instructions })
|
|
364
|
+
context.instructions.add({ name: 'tone', content: 'Be terse.' })
|
|
365
|
+
// An item override beats the manager `render` for THAT item only:
|
|
366
|
+
context.instructions.add({
|
|
367
|
+
name: 'raw',
|
|
368
|
+
content: 'ignored',
|
|
369
|
+
format: '<rule priority="high">Escalate.</rule>',
|
|
370
|
+
})
|
|
371
|
+
|
|
372
|
+
context.build()
|
|
373
|
+
// system block instructions section (the group wrapped by open + close):
|
|
374
|
+
// '<rules>\n\n<rule>Be terse.</rule>\n\n<rule priority="high">Escalate.</rule>\n\n</rules>'
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
A provider declares its framing default by exposing `format` on its `ProviderInterface`; the `Agent` passes it into `build()` automatically. Because it is optional, no provider is forced to supply one — omitting it leaves every section on the managers' built-in framing.
|
|
378
|
+
|
|
379
|
+
### The HTTP provider engine
|
|
380
|
+
|
|
381
|
+
`ProviderInterface` is the boundary; `AgentProvider` is the engine behind it. Every HTTP-backed provider needs the same machinery — a deadline, a transport, an authorization hook, a bounded error read, a chunk decoder, reasoning separation, and result assembly — and rewriting that per vendor is how a fleet of providers drifts apart. `AgentProvider` owns all of it and leaves exactly one thing open: the wire. A subclass supplies `name` and the wire seams — `frame()` returns fresh framing state for the call, `body(request)` projects a `ProviderRequest` onto the vendor's request shape, `read(record)` decodes one framed record into a `ProviderIncrement`, and `finish(parser)` returns whatever the parser still held at end of input. The engine is generic over `TRecord`: the record type a call's `frame()` parser emits and `read` consumes, defaulting to `Readonly<Record<string, unknown>>` for a wire whose records are JSON objects. The constructor switches decide the rest: `split` separates in-content `<think>` reasoning from the answer (default `true`), and `strict` requires the wire to carry a settled `result` record rather than assembling one at end of input (default `false`).
|
|
382
|
+
|
|
383
|
+
```ts
|
|
384
|
+
import type { AgentProviderInput, ProviderInterface } from '@orkestrel/agent'
|
|
385
|
+
import { createAbort } from '@orkestrel/abort'
|
|
386
|
+
|
|
387
|
+
declare function createTextProvider(options: AgentProviderInput): ProviderInterface
|
|
388
|
+
declare function token(signal: AbortSignal): Promise<string>
|
|
389
|
+
const abort = createAbort()
|
|
390
|
+
const messages = [{ id: '1', role: 'user', content: 'Say hello.' }] as const
|
|
391
|
+
|
|
392
|
+
const provider = createTextProvider({
|
|
393
|
+
url: 'https://api.example',
|
|
394
|
+
path: '/generate', // appended to `url` on every call
|
|
395
|
+
timeout: 30_000, // the base's own deadline; DEFAULT_PROVIDER_TIMEOUT when omitted
|
|
396
|
+
headers: async (signal) => ({ authorization: await token(signal) }), // awaited inside that deadline
|
|
397
|
+
split: true, // route <think> spans to `thinking`, yield the clean answer
|
|
398
|
+
strict: false, // assemble at end of input instead of requiring a settled record
|
|
399
|
+
})
|
|
400
|
+
const result = await provider.generate(messages, abort.signal)
|
|
401
|
+
result.content // the assembled answer, with any <think> span routed to result.thinking
|
|
402
|
+
```
|
|
403
|
+
|
|
404
|
+
`AgentProvider` is `abstract`; the members a subclass fills are listed under [`## Methods`](#methods).
|
|
405
|
+
|
|
406
|
+
The engine's failures arrive as one error class. A `ProviderError` carries a machine-readable `code` and, for an HTTP failure alone, the response `status`; `isProviderError` narrows a caught value so a caller can branch on the code:
|
|
407
|
+
|
|
408
|
+
- `'HTTP'` — a non-OK response. The message is `provider error: <status>`, with ` - <excerpt>` appended only when the response body carried text. The excerpt is read to at most `MAX_ERROR_BODY_LENGTH` bytes and the remainder of the body is cancelled, so a stalled error body cannot hold the call open past its deadline.
|
|
409
|
+
- `'PROTOCOL'` — a successful response with no body, a wire record the concrete provider refuses, or a `strict` stream that ended without a settled result.
|
|
410
|
+
- `'PROVIDER'` — an upstream failure a relay reported in an `error` frame.
|
|
411
|
+
|
|
412
|
+
A cancel is not in that taxonomy: the call's own deadline and the caller's signal are folded into one bound, and a cancel throws `ProviderAbortError` carrying the partial assembled so far.
|
|
413
|
+
|
|
414
|
+
### The relay
|
|
415
|
+
|
|
416
|
+
A browser must not hold a model credential, so this package ships the hop rather than the credential. `createRelay` returns a `RelayHandler` — a plain `(request: Request) => Promise<Response>` you mount on any fetch-standard router — that authorizes the request, validates its JSON body against `providerRequestContract`, and streams one upstream `provider.stream` call back as newline-delimited `RelayFrame` records under `RELAY_CONTENT_TYPE`. `createRelayProvider` is the other end: a `ProviderInterface` the browser drives exactly like a local one, which posts the request and decodes those frames back into deltas and the settled result. `RelayStream` is the response half `createRelay` composes, exported for a host that mounts its own route.
|
|
417
|
+
|
|
418
|
+
```ts
|
|
419
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
420
|
+
import { createRelay, createRelayProvider } from '@orkestrel/agent'
|
|
421
|
+
// The browser application supplies this parser dependency.
|
|
422
|
+
import { createNDJSONParser } from '@orkestrel/ndjson'
|
|
423
|
+
|
|
424
|
+
declare const upstream: ProviderInterface // the server-side provider holding the credential
|
|
425
|
+
|
|
426
|
+
// On the server: the route, the authorization decision, and the byte budget.
|
|
427
|
+
const handler = createRelay({
|
|
428
|
+
provider: upstream,
|
|
429
|
+
authorize: (request) => request.headers.get('authorization') === 'Bearer session-token',
|
|
430
|
+
limit: 65_536, // a body at or above this answers 413
|
|
431
|
+
})
|
|
432
|
+
|
|
433
|
+
// In the browser: a provider like any other, reached over that route.
|
|
434
|
+
const browser = createRelayProvider({
|
|
435
|
+
url: 'https://app.example/relay',
|
|
436
|
+
parser: createNDJSONParser,
|
|
437
|
+
headers: () => ({ authorization: 'Bearer session-token' }),
|
|
438
|
+
})
|
|
439
|
+
```
|
|
440
|
+
|
|
441
|
+
The frame vocabulary is `RelayFrame`: a `ProviderDelta` (`content` or `thinking`) per streamed delta, one `{ channel: 'result', result }` when the turn settles, `{ channel: 'abort', partial }` when the upstream call was cancelled, and `{ channel: 'error', message }` for anything else — always the fixed `RELAY_PROVIDER_MESSAGE` text, so an upstream failure's own message never reaches the browser. A refusal never becomes a frame: the handler answers `401` when `authorize` refuses or throws, `400` when the body is missing, unreadable, or rejected by the contract, `413` when the body reaches the byte limit, and `502` when the upstream call cannot be constructed. Each refusal carries no body and reaches the browser as a `ProviderError` with the `HTTP` code and that status.
|
|
442
|
+
|
|
443
|
+
The wire is strictly narrower than the domain on purpose. `providerRequestContract` and `relayFrameContract` are compiled from the shapes in [Shapes and contracts](#shapes-and-contracts), and `RelayProvider.body` refuses a request the JSON wire cannot carry — a function-valued tool argument, a parameter schema holding one — before it fetches. The body it sends is an owned snapshot of that projection read through property descriptors, so a serializer reachable only through a `get` trap or a prototype is never consulted; an own function-valued property such as a `toJSON` method is a value outside JSON and is refused before fetching, with the clone's failure as the refusal's `cause`. A `ToolCall.caller` is local context and never crosses the hop.
|
|
444
|
+
|
|
445
|
+
### Factories
|
|
446
|
+
|
|
447
|
+
| API | Kind | Summary |
|
|
448
|
+
| --------------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
449
|
+
| `createConversation` | function | Creates a conversation — a `ConversationInterface` grouping messages above a flat message store it owns directly, with compaction into summarized sections, a regenerated rollup `summary`, on-demand `rehydrate`, and substring `search`, driven by a provider-agnostic `ConversationSummaryHandler` seam. |
|
|
450
|
+
| `createConversationManager` | function | Creates a conversation registry — a `ConversationManagerInterface` holding `ConversationInterface`s keyed by their `id`, in insertion order, with an active pointer: the id-keyed store over the conversation layer plus the `active` / `switch` seam the context renders. `add` auto-activates the first conversation and flows the registry's default `summarize` / `keep` into every conversation it creates. |
|
|
451
|
+
| `createMemoryConversationStore` | function | Creates the in-memory conversation store — a `ConversationStoreInterface` backed by a process-lifetime `Map` of `ConversationSnapshot`s keyed by conversation id, the default backing for the durable `ConversationManagerInterface.open` / `ConversationManagerInterface.save` seam. The exact twin of `createMemoryWorkspaceStore`. |
|
|
452
|
+
| `createDatabaseConversationStore` | function | Creates a `DatabaseConversationStore` over any `DriverInterface`, defaulting to `createMemoryDriver()` — the durable, driver-pluggable backing for the conversation persistence seam, holding each snapshot as one opaque JSON column and standing as the opt-in twin of `createMemoryConversationStore`. The exact twin of `createDatabaseWorkspaceStore`. |
|
|
453
|
+
| `createInstruction` | function | Creates an instruction — an immutable `InstructionInterface` (a named directive) from its `name` / `content` and optional `priority`, the `id` minted at construction. |
|
|
454
|
+
| `createInstructionManager` | function | Creates an instruction registry — an `InstructionManagerInterface` holding immutable instructions keyed by `name`, listed by descending `priority`. |
|
|
455
|
+
| `createScope` | function | Creates a named scope — an immutable `ScopeInterface` from its `name` and its per-category allow-lists, the `id` minted at construction. |
|
|
456
|
+
| `createScopeManager` | function | Creates a scope registry — a `ScopeManagerInterface` holding immutable scopes keyed by their minted `id`, in insertion order. |
|
|
457
|
+
| `createAgentContext` | function | Creates a richer turn context — an `AgentContextInterface` assembling a provider request from the optional system prompt, the instruction registry, the workspace registry (the only document channel), the conversation registry that is its `messages` source, the tool registry, and the active scope, which `build()` folds into the next turn's input. |
|
|
458
|
+
| `createAgent` | function | Creates an agent loop — an `AgentInterface` composing a `ProviderInterface`, its `AgentContextInterface`, and a tool registry into a bounded context → provider → tools → repeat turn, exposed as a one-shot `generate` and a live `stream`. |
|
|
459
|
+
| `createAuthority` | function | Creates a policy gate — an `AuthorityInterface` the agent loop consults before each tool call runs, evaluating the ordered rules first-match-wins and falling back to the configured default when none match. |
|
|
460
|
+
| `createRelay` | function | Creates an authorized relay handler that validates a bounded JSON request before streaming. |
|
|
461
|
+
| `createRelayProvider` | function | Creates a provider that carries calls through a relay endpoint. |
|
|
462
|
+
| `createThinkSplitter` | function | Creates a fresh stream-stateful `<think>` separator — a `ThinkSplitterInterface` that splits a thinking model's in-content `<think>…</think>` reasoning spans away from the answer, delta by delta, so a provider yields clean content alone and surfaces the accumulated reasoning as `ProviderResult.thinking`. One splitter serves one stream. |
|
|
463
|
+
| `createChannel` | function | Creates an empty unbounded async channel — a `ChannelInterface` a producer writes values into (`push`) and ends (`close` / `fail`) regardless of consumption, while a consumer reads them back live through `drain`. |
|
|
464
|
+
| `createAgentRegistry` | function | Creates an agent registry — an `AgentRegistryInterface` holding the named pools of live, non-serializable pieces (providers, tools, authorities, schedulers) that a serializable `AgentJobInput`'s names resolve against, and `build`ing a seeded, signal-wired `AgentInterface` from a job. |
|
|
465
|
+
| `createAgentQueue` | function | Creates a durable, bounded-concurrency agent-job queue — a `QueueInterface` over serializable `AgentJobInput`s that composes `createQueue`: each job is rehydrated through the `registry` into a live `AgentInterface`, run to its `AgentResult`, and subjected to the partial-as-configurable-failure policy. |
|
|
466
|
+
| `createAgentRunner` | function | Creates an agent-job runner — a `RunnerInterface` over serializable `AgentJobInput`s that composes `createRunner` (one-shot, ordered, fail-fast), each unit rehydrated through the `registry` and subjected to the partial policy. The runner also carries sub-agent fan-out: a parent job's handler can `controller.spawn(childJob)`. |
|
|
467
|
+
|
|
468
|
+
### Classes
|
|
469
|
+
|
|
470
|
+
| API | Kind | Summary |
|
|
471
|
+
| --------------------------- | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
472
|
+
| `AgentProvider` | class | Implements bounded HTTP streaming and result assembly behind concrete wire seams. |
|
|
473
|
+
| `RelayProvider` | class | Carries provider calls over an authenticated NDJSON relay endpoint. |
|
|
474
|
+
| `RelayStream` | class | Streams a provider call as validated NDJSON frames under response backpressure. |
|
|
475
|
+
| `Conversation` | class | Represents a conversation — a live uncompacted tail of messages it owns directly above a flat message store, plus compacted, summarized `Section`s, a regenerated rollup `summary`, and a `summarizable` flag, with on-demand `rehydrate` and substring `search`, driven by a provider-agnostic `ConversationSummaryHandler` seam so `core` never imports a provider. Observable through its own `emitter`. |
|
|
476
|
+
| `ConversationManager` | class | Registers `Conversation`s keyed by `id`, in insertion order, with an active pointer — the id-keyed store over the conversation layer, the `active` / `switch` seam the `AgentContext` renders, and the durable `open` / `save` store seam. Event-free (a registry, like `WorkspaceManager`); the observability lives on each `Conversation`. |
|
|
477
|
+
| `MemoryConversationStore` | class | Implements the `ConversationStoreInterface` in memory — a process-lifetime `Map` of `ConversationSnapshot`s keyed by conversation id, the default store `createMemoryConversationStore` builds and the default backing for `open` / `save`. The exact twin of `MemoryWorkspaceStore`. |
|
|
478
|
+
| `DatabaseConversationStore` | class | Backs a `ConversationStoreInterface` with one table of the `databases` layer — a conversation's durable state is a row holding the snapshot as one opaque JSON column, narrowed back on `get` by `isConversationSnapshot`, so persistence reduces to keyed point-access (`get` / `set` / `delete`) over a `TableInterface`. The driver-pluggable twin of the plain-`Map` `MemoryConversationStore`, and the exact twin of `DatabaseWorkspaceStore`. |
|
|
479
|
+
| `Instruction` | class | Represents an immutable named directive — an `InstructionInterface` assembled once from its input (`name` / `content`, an optional `priority` defaulting to `0`), the `id` minted at construction. |
|
|
480
|
+
| `InstructionManager` | class | Registers the immutable `Instruction`s a richer context assembles a directives block from — keyed by `name` so a re-`add` overwrites, last write wins, and listed by descending `priority`, carrying the `open` / `render` build contract and an observable `emitter`. |
|
|
481
|
+
| `Scope` | class | Represents a named, immutable filter over a richer context's items — an optional allow-list per category (`instructions` / `tools` / `files`), each keyed by that category's identity (an instruction's `name`, a tool's `name`, a workspace file's `path`) and read as an allow-list: `undefined` lets everything pass, `[]` lets nothing pass, and a non-empty list passes the listed keys alone. `narrow` composes a tighter child by set intersection. |
|
|
482
|
+
| `ScopeManager` | class | Registers the named filters a richer context reuses — immutable `Scope`s keyed by their minted `id`, in insertion order, where `create` always mints and stores rather than overwriting, and an observable `emitter` reports each change. |
|
|
483
|
+
| `AgentContext` | class | Assembles a provider request from the richer turn context — the optional system prompt, the observable context managers (instructions / workspaces), the `ConversationManagerInterface` message source whose active conversation is `messages`, the `ToolManagerInterface` registry, and an active `ScopeInterface` changed through `AgentContextInterface.apply`. `build()` folds the scoped managers and the active workspace into one system block, then the conversation, and never reads `tools`. |
|
|
484
|
+
| `Agent` | class | Composes a `ProviderInterface`, an `AgentContext`, and a `ToolManagerInterface` into a bounded context → provider → tools → repeat turn, exposed as both a one-shot `generate` and a live `stream` that share one private run — bounded by the run `signal`, the `timeout`, and the `budget` folded through `AbortSignal.any`, paced by `scheduler`, with tool iteration capped at `limit`. |
|
|
485
|
+
| `Authority` | class | Gates the agent loop's tool calls — the synchronous policy consulted before each call runs, turning one `AuthorityContext` into an `AuthorityDecision` by walking the ordered rules first-match-wins and falling back to a configurable default, which allows an unmatched call unless its `fallback` denies. |
|
|
486
|
+
| `AgentRegistry` | class | Makes a durable, JSON-serializable `AgentJobInput` runnable — holds the named pools of live, non-serializable pieces (providers, tools, authorities, schedulers), throws on a name absent from its pool, and `build`s a seeded, signal-wired `Agent` from a job's names and data. |
|
|
487
|
+
| `Channel` | class | Buffers chunks in a minimal unbounded async channel — the eager pump writes them in (`push`) and ends it (`close` / `fail`) regardless of consumption, while a consumer reads them back live through the `drain` async-iterator. Decoupling write from read is what lets a producer make progress without a consumer pulling, and it is why an agent's `result` settles whether or not its `events` are drained. |
|
|
488
|
+
| `ThinkSplitter` | class | Feeds raw content deltas through a tiny stream-stateful state machine that routes everything inside a `<think>…</think>` span to `thinking` and returns everything outside it as clean content, so a provider yields the answer alone and surfaces the reasoning as `ProviderResult.thinking`. A tag split across deltas is held until disambiguated, `flush()` settles the stream end, and one splitter serves one stream. |
|
|
489
|
+
|
|
490
|
+
### Constants
|
|
491
|
+
|
|
492
|
+
A `Shape` cell holds the constant's declared type.
|
|
493
|
+
|
|
494
|
+
| API | Kind | Shape | Summary |
|
|
495
|
+
| --------------------------- | ----- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
496
|
+
| `CONVERSATION_RECAP_PREFIX` | const | `string` | Names the framing label a `ConversationInterface`'s `view()` prefixes onto each compacted section's summary so a small model reads it as a condensed recap of earlier turns — the lean `'[Summary of earlier messages] '` marker, never a literal assistant turn to echo or treat as the live answer. |
|
|
497
|
+
| `DEFAULT_AGENT_LIMIT` | const | `number` | Caps an `AgentInterface` turn's tool iterations by default — `10` context → provider → tools cycles before the loop stops, so a model that keeps requesting tools can never loop forever. Overridable per agent through `AgentOptions.limit`. |
|
|
498
|
+
| `DEFAULT_AUTHORITY_ZONE` | const | `string` | Names the zone an `AuthorityInterface`'s default fallback `AuthorityDecision` carries — `'default'`, the classification for a tool call that matched no rule. Paired with the default `allowed: true` fallback, an unmatched call is allowed under this zone, so a rules list of denials acts as a denylist; a caller wanting deny-by-default supplies an `allowed: false` `fallback` of their own (see `AuthorityOptions`). |
|
|
499
|
+
| `DEFAULT_CONVERSATION_KEEP` | const | `number` | Sets the default number of recent live messages a `ConversationInterface`'s `compact()` retains verbatim — `0`, so a manual `compact()` folds every current live message into one summarized section and keeps no tail. A caller retains a recent tail by passing `keep` (on `ConversationOptions`, `ConversationManagerOptions`, or per-fold through `CompactOptions`), folding only the older `count - keep` messages and leaving the most recent `keep` live for the next turn. Overridable everywhere `keep` is accepted. |
|
|
500
|
+
| `THINK_OPEN` | const | `string` | Names the opening tag a `ThinkSplitter` recognizes as the start of an in-content reasoning span — `'<think>'`, the de-facto wire convention thinking models (qwen3, DeepSeek-R1 family) emit their chain-of-thought under when a daemon renders it inline instead of on a separate wire field. Paired with `THINK_CLOSE`. |
|
|
501
|
+
| `THINK_CLOSE` | const | `string` | Names the closing tag that ends a `THINK_OPEN` reasoning span — `'</think>'`. A span the stream never closes (the model was cut off mid-reasoning) is treated as thinking to its end, and `ThinkSplitterInterface.flush` settles it. |
|
|
502
|
+
| `WORKSPACE_SECTION_HEADER` | const | `string` | Names the section header `AgentContext`'s `build()` renders the active workspace's text files under — `'## Workspace'`, the leading line of the dedicated workspace block in the system message and the carrier-split counterpart to the documents and images section headers. |
|
|
503
|
+
| `MESSAGE_TOKEN_OVERHEAD` | const | `number` | Estimates the per-message role and framing overhead `estimateMessages` adds on top of a message's content estimate — `4` tokens for the fixed wire framing every conversation turn carries (its role tag, its delimiters) that `estimateTokens`'s content-only heuristic does not otherwise capture. |
|
|
504
|
+
| `IMAGE_TOKEN_ESTIMATE` | const | `number` | Names the coarse, deliberately approximate per-image token cost `estimateMessages` charges for each attached image — `512`, because a base64 payload's length is no reliable token proxy. |
|
|
505
|
+
| `DEFAULT_PROVIDER_TIMEOUT` | const | `number` | Holds the default provider deadline in milliseconds — `120_000`, the wall-clock bound a call runs under when `AgentProviderInput.timeout` is omitted, folded with the caller's signal so whichever trips first cancels the call. |
|
|
506
|
+
| `MAX_ERROR_BODY_LENGTH` | const | `number` | Bounds the decoded error excerpt's input in bytes — `2048`, the leading bytes of a non-OK response body handed to the decoder before the read cancels the remainder, so a `ProviderError` message never carries a longer excerpt. |
|
|
507
|
+
| `DEFAULT_RELAY_LIMIT` | const | `number` | Holds the default relay request limit in bytes — `1_048_576`, the byte budget a relay applies to an inbound body when `RelayOptions.limit` is omitted, refusing a body that reaches it. |
|
|
508
|
+
| `RELAY_CONTENT_TYPE` | const | `string` | Names the relay's newline-delimited JSON content type — `'application/x-ndjson; charset=utf-8'`, the header a relay response carries beside `cache-control: no-store`. |
|
|
509
|
+
| `RELAY_PROVIDER_MESSAGE` | const | `string` | Names the public message for an unexpected upstream relay failure — `'relay provider failed'`, the fixed text every `error` frame carries, so an upstream failure's own message never reaches the browser. |
|
|
510
|
+
| `UNAUTHORIZED_RELAY_STATUS` | const | `number` | Names the status a relay answers when the `authorize` callback returns anything but `true` or throws — `401`, carried with no body and reaching the browser as a `ProviderError` with the `HTTP` code. |
|
|
511
|
+
| `INVALID_RELAY_STATUS` | const | `number` | Names the status a relay answers for a body that is missing, unreadable, or rejected by `providerRequestContract` — `400`, carried with no body and reaching the browser as a `ProviderError` with the `HTTP` code. |
|
|
512
|
+
| `OVERSIZED_RELAY_STATUS` | const | `number` | Names the status a relay answers for a request body at or above its byte budget — `413`, carried with no body and answered for an aborted inbound read as well. |
|
|
513
|
+
| `UPSTREAM_RELAY_STATUS` | const | `number` | Names the status a relay answers when the upstream provider call cannot be constructed — `502`, carried with no body after `provider.stream` was entered and threw before returning its iterator. |
|
|
514
|
+
|
|
515
|
+
### Shapes and contracts
|
|
516
|
+
|
|
517
|
+
The JSON wire projections the relay validates against, and the contracts compiled from them. On a shape row a `Shape` cell holds the projected record's members as bare names in braces, `?` marking an optional member and a union's arms escaped as `\|`; on a contract row it holds the domain type that contract narrows to. Each projection is strictly narrower than the domain type it mirrors: a value JSON cannot carry is refused rather than silently dropped, and a `ToolCall`'s `caller` is local context that never appears here.
|
|
518
|
+
|
|
519
|
+
| API | Kind | Shape | Summary |
|
|
520
|
+
| ------------------------- | ----- | ----------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- |
|
|
521
|
+
| `toolCallShape` | const | `{ id, name, arguments }` | Describes a tool call's JSON wire projection. |
|
|
522
|
+
| `messageShape` | const | `{ id, role, content, calls?, images? }` | Describes a conversation message's JSON wire projection. |
|
|
523
|
+
| `providerRequestShape` | const | `{ messages, tools?, options? }` | Describes a provider request's JSON wire projection. |
|
|
524
|
+
| `providerResultShape` | const | `{ content, thinking?, tools?, usage? }` | Describes a provider result's JSON wire projection. |
|
|
525
|
+
| `relayFrameShape` | const | `{ channel: 'content' \| 'thinking', text } \| { channel: 'result', result } \| { channel: 'abort', partial } \| { channel: 'error', message }` | Describes the channel-discriminated JSON relay wire projection. |
|
|
526
|
+
| `messageContract` | const | `Message` | Validates and projects conversation messages at a JSON wire boundary. |
|
|
527
|
+
| `providerRequestContract` | const | `ProviderRequest` | Validates and projects provider requests at a JSON wire boundary. |
|
|
528
|
+
| `providerResultContract` | const | `ProviderResult` | Validates and projects provider results at a JSON wire boundary. |
|
|
529
|
+
| `relayFrameContract` | const | `RelayFrame` | Validates and projects channel-discriminated relay frames at a JSON wire boundary. |
|
|
530
|
+
|
|
531
|
+
A contract's `is` narrows an unknown record to its wire shape, and its `parse` projects one — returning the value stripped to that shape, or `undefined` when the value is invalid — so the same declaration guards an inbound body and shapes a frame written back out:
|
|
532
|
+
|
|
533
|
+
```ts
|
|
534
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
535
|
+
import { providerRequestContract, relayFrameContract } from '@orkestrel/agent'
|
|
536
|
+
|
|
537
|
+
declare const upstream: ProviderInterface
|
|
538
|
+
declare const body: unknown
|
|
539
|
+
declare const signal: AbortSignal
|
|
540
|
+
declare function write(line: string): void
|
|
541
|
+
|
|
542
|
+
if (providerRequestContract.is(body)) {
|
|
543
|
+
const result = await upstream.generate(body.messages, signal, body.tools, body.options)
|
|
544
|
+
const frame = relayFrameContract.parse({ channel: 'result', result })
|
|
545
|
+
write(`${JSON.stringify(frame)}\n`) // the newline-delimited record a relay writes back
|
|
546
|
+
}
|
|
547
|
+
relayFrameContract.is({ channel: 'error', message: 'relay provider failed' }) // true
|
|
548
|
+
relayFrameContract.is({ channel: 'error', message: 'oops', code: 'X' }) // false — an extra member is refused
|
|
549
|
+
relayFrameContract.parse({ channel: 'error', message: 'oops', code: 'X' }) // { channel: 'error', message: 'oops' } — parse projects the extra member away
|
|
550
|
+
```
|
|
551
|
+
|
|
552
|
+
### Helpers
|
|
553
|
+
|
|
554
|
+
| API | Kind | Summary |
|
|
555
|
+
| ---------------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
556
|
+
| `agentResultToJSON` | function | Projects an unknown value onto a fresh, exact `JSONValue` representation of an `AgentResult` — capturing each structural field once through a total boundary, accepting conforming accessors and inherited properties, preserving finite negative and fractional usage counts, dropping extras, and resolving `undefined` for a malformed field, a non-finite usage number, a throwing getter, or a hostile or revoked proxy. |
|
|
557
|
+
| `filterAllowList` | function | Filters a list of items by a `ScopeInterface` allow-list of keys — `undefined` passes everything, `[]` passes nothing, and a non-empty list passes the listed keys alone, order preserved. The pure, total set-membership primitive the context's build step and the agent loop's tool-advertise step apply a scope through. |
|
|
558
|
+
| `estimateTokens` | function | Estimates the context-token footprint of a string — the deterministic `ceil(length / 4)` character heuristic `estimateMessages` sums over a conversation's messages (the default context-budget estimator). |
|
|
559
|
+
| `estimateMessages` | function | Estimates the context-token footprint of a batch of messages — each message's content plus `MESSAGE_TOKEN_OVERHEAD`, a tool-call JSON estimate, and `IMAGE_TOKEN_ESTIMATE` for each attached image. The default `consumer` estimator for an agent's context budget (the `AgentOptions` `window`), total and never throwing, and a deliberate provider-agnostic approximation rather than an exact tokenizer count. |
|
|
560
|
+
| `sanitizeToken` | function | Sanitizes one reported token count into a safe non-negative integer — a non-finite or non-positive value becomes `0`, and a positive fractional value floors down. |
|
|
561
|
+
| `sanitizeUsage` | function | Sanitizes a `TokenUsage` into safe, non-negative integers — the guard an agent's abort-usage path applies to a provider's partial usage before it is charged against a budget or folded into the run total. |
|
|
562
|
+
| `settleAgentJob` | function | Runs one rehydrated agent and applies the partial-as-configurable-failure policy — a partial run throws an `AgentJobError` unless the `partial` policy allows it, and a natural finish resolves. The shared job-handler step `createAgentQueue` and `createAgentRunner` both settle each job through, so the policy can never diverge between them. |
|
|
563
|
+
| `handleAgentQueueJob` | function | Handles one queued agent job by rehydrating it through a registry with the queue attempt's signal, then applying the shared partial-result policy. |
|
|
564
|
+
| `handleAgentRunnerJob` | function | Handles one runner agent job by fanning out its declared children, rehydrating the parent through a registry with the controller signal, and applying the shared partial-result policy. |
|
|
565
|
+
| `renderFencedFile` | function | Renders a path-addressed text body as a fenced reference block — a `File: <path>` label line over a language-tagged fence, the framing an `AgentContext`'s active-workspace text-file render emits. |
|
|
566
|
+
| `joinThinking` | function | Joins the reasoning a run's provider calls separated from the answer — the first call seeds the accumulation, a later call appends blank-line separated so each turn's reasoning stays readable. |
|
|
567
|
+
| `sumUsage` | function | Adds two `TokenUsage` values field by field — the running total an agent run keeps across its provider calls. |
|
|
568
|
+
| `assembleResult` | function | Assembles the settled `AgentResult` from a run's `RunOutcome` — `thinking` and `usage` are carried only when the run surfaced them, and the loop-internal `exhausted` flag is left out. |
|
|
569
|
+
| `denyCall` | function | Synthesizes the denial `ToolResult` an authority-blocked call is fed back with — the call's `id` / `name` keyed back, carrying a denial `error` instead of a value. |
|
|
570
|
+
| `renderSection` | function | Renders one context section — the resolved `open`, each item's rendering, and the resolved `close` when one exists, blank-line joined; `undefined` when the section has no items. |
|
|
571
|
+
| `resolveOpen` | function | Resolves one section's open text through the format cascade — manager-options override > provider default > built-in header. |
|
|
572
|
+
| `resolveClose` | function | Resolves one section's close text through the format cascade — manager-options override > provider default; `undefined` when neither sets one, because there is no built-in close. |
|
|
573
|
+
| `resolveItem` | function | Resolves one item's rendering through the format cascade — item override > manager-options override > provider default > built-in rendering. |
|
|
574
|
+
| `attachImages` | function | Copies a message with image data merged onto its `images` — the message's own images first, then the attached data, carrying `calls` only when present and never mutating the original. |
|
|
575
|
+
| `attachUserImages` | function | Attaches image data to a conversation's last user message — the turn a vision provider reads images off — as a new array with that one message replaced by its carrying copy, and unchanged when there is no data or no user turn. |
|
|
576
|
+
| `collectImageData` | function | Collects the `base64` payload of the image files in a workspace file list — the data an agent context attaches to the last user message. |
|
|
577
|
+
| `buildSummaryMessage` | function | Builds the raw synthetic summary message for one compacted section — role `'assistant'`, the section's stable `id`, and its `summary` verbatim as content. |
|
|
578
|
+
| `buildRecapMessage` | function | Builds the framed recap message for one compacted section — the same role and stable `id` as `buildSummaryMessage`, with the content prefixed by `CONVERSATION_RECAP_PREFIX`. |
|
|
579
|
+
| `buildProviderResult` | function | Assembles a provider result with only populated optional fields. |
|
|
580
|
+
| `readText` | function | Reads a UTF-8 prefix of a byte stream and cancels its remainder. |
|
|
581
|
+
| `readChunks` | function | Decodes UTF-8 chunks with a final flush and releases the stream on every exit. |
|
|
582
|
+
| `intersectKeys` | function | Intersects two scope category lists under the "`undefined` is the universal set" rule — a fresh copy that can only tighten, and the primitive a scope narrows through. |
|
|
583
|
+
|
|
584
|
+
Project an agent result at its originating package before carrying it through a JSON boundary:
|
|
585
|
+
|
|
586
|
+
```ts
|
|
587
|
+
import type { AgentResult } from '@orkestrel/agent'
|
|
588
|
+
import { agentResultToJSON } from '@orkestrel/agent'
|
|
589
|
+
|
|
590
|
+
declare const result: AgentResult
|
|
591
|
+
const portable = agentResultToJSON(result)
|
|
592
|
+
if (portable === undefined) throw new Error('invalid agent result')
|
|
593
|
+
JSON.stringify(portable)
|
|
594
|
+
```
|
|
595
|
+
|
|
596
|
+
The queue and runner factories bind their named handlers to a registry and partial policy; callers composing the lower-level substrates can do the same:
|
|
597
|
+
|
|
598
|
+
```ts
|
|
599
|
+
import type { AgentRegistryInterface } from '@orkestrel/agent'
|
|
600
|
+
import { handleAgentQueueJob, handleAgentRunnerJob, sanitizeToken } from '@orkestrel/agent'
|
|
601
|
+
|
|
602
|
+
declare const registry: AgentRegistryInterface
|
|
603
|
+
|
|
604
|
+
const tokens = sanitizeToken(12.7) // 12
|
|
605
|
+
const queueHandler = handleAgentQueueJob.bind(undefined, registry, false)
|
|
606
|
+
const runnerHandler = handleAgentRunnerJob.bind(undefined, registry, false)
|
|
607
|
+
```
|
|
608
|
+
|
|
609
|
+
### Validators
|
|
610
|
+
|
|
611
|
+
In a guard table a `Shape` cell holds the type the guard narrows to. Each guard reads an `unknown`, returns `false` off-shape, and never throws. An error guard stays in the Errors table beside the error it narrows.
|
|
612
|
+
|
|
613
|
+
| API | Kind | Shape | Summary |
|
|
614
|
+
| ------------------------ | -------- | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
615
|
+
| `isMessage` | function | `Message` | Checks whether a value satisfies the domain conversation-message contract. |
|
|
616
|
+
| `isSection` | function | `Section` | Checks whether an `unknown` is structurally a `Section` record — a `string` `id` and `summary` beside a `messages` array of valid `Message`s, the per-section step of the `isConversationSnapshot` read-boundary narrow. Total, never throwing, and never an assertion. |
|
|
617
|
+
| `isConversationSnapshot` | function | `ConversationSnapshot` | Narrows an `unknown` to a `ConversationSnapshot` — a `string` `id`, an optional `string` `summary`, and valid `sections` and `messages` arrays; the total boundary guard for an untrusted snapshot read (a storage row a `DatabaseConversationStore` reads back from its opaque JSON column, a snapshot loaded from disk), never throwing. The exact analogue of `isWorkspaceSnapshot`. |
|
|
618
|
+
|
|
619
|
+
A `DatabaseConversationStore` reads its snapshot column back as `unknown` and narrows it through the last of them, so a malformed blob resolves `undefined` rather than a broken conversation:
|
|
620
|
+
|
|
621
|
+
```ts
|
|
622
|
+
import { isConversationSnapshot } from '@orkestrel/agent'
|
|
623
|
+
|
|
624
|
+
declare const row: unknown
|
|
625
|
+
const snapshot = isConversationSnapshot(row) ? row : undefined
|
|
626
|
+
```
|
|
627
|
+
|
|
628
|
+
### Errors
|
|
629
|
+
|
|
630
|
+
| API | Kind | Summary |
|
|
631
|
+
| ---------------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
632
|
+
| `ProviderAbortError` | class | Reports a provider stream cancelled mid-flight by its bound signal — thrown by a `ProviderInterface`'s `stream`, carrying the `ProviderResult` assembled from whatever streamed before the cancel and the machine-readable `code` `'ABORT'`. |
|
|
633
|
+
| `isProviderAbortError` | function | Narrows an unknown caught value to a `ProviderAbortError` through `instanceof`, so a `catch` can recover its `partial` result. |
|
|
634
|
+
| `ProviderError` | class | Reports a coded provider failure with its HTTP status and underlying cause when available. |
|
|
635
|
+
| `isProviderError` | function | Narrows a caught value to the provider failure class through instanceof. |
|
|
636
|
+
| `AgentJobError` | class | Reports an `AgentInterface` run that ended `AgentResult.partial` under a `partial` policy of `false` (the default) — thrown by an agent-job handler (a `createAgentQueue` / `createAgentRunner` job), carrying the partial `AgentResult` so the failure stays inspectable, and the machine-readable `code` `'PARTIAL'`. |
|
|
637
|
+
| `isAgentJobError` | function | Narrows an unknown caught value to an `AgentJobError` through `instanceof`, so a `catch` can recover its `partial` result. |
|
|
638
|
+
| `ConversationError` | class | Reports a conversation with no `ConversationSummaryHandler` to fold its messages with, or with a `sections` cap below `1` — thrown by a `ConversationInterface`'s `compact()` or its construction, carrying the machine-readable `code` `'SUMMARIZER' \| 'SECTIONS'`. |
|
|
639
|
+
| `isConversationError` | function | Narrows an unknown caught value to a `ConversationError` through `instanceof`, so a `catch` can branch on its `code`. |
|
|
640
|
+
| `AgentError` | class | Reports a concurrent run that would corrupt shared per-agent accounting, or a rehydration name absent from its registry pool — thrown synchronously by an `AgentInterface`'s `stream()` (and so by `generate()`, which calls it) and by an `AgentRegistryInterface`'s accessors, carrying the machine-readable `code` `'CONCURRENCY' \| 'REGISTRY'`. Synchronous means a fire-and-forget `agent.generate().catch(…)` never catches it: `await` the call inside `try`/`catch`, or wrap the call expression itself. |
|
|
641
|
+
| `isAgentError` | function | Narrows an unknown caught value to an `AgentError` through `instanceof`, so a `catch` can branch on its `code`. |
|
|
642
|
+
|
|
643
|
+
### Types
|
|
644
|
+
|
|
645
|
+
A `Shape` cell holds an interface's data members as bare names in braces, `?` marking an optional member and `plus` introducing its call-signature members, and a type alias's own type literal with a union's arms escaped as `\|`. An extended interface's name comes before `plus`, with the members it adds after.
|
|
646
|
+
|
|
647
|
+
| Type | Kind | Shape | Summary |
|
|
648
|
+
| ------------------------------- | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
649
|
+
| `MessageRole` | type | `'system' \| 'user' \| 'assistant' \| 'tool'` | Names the role a `Message` plays in a conversation turn. |
|
|
650
|
+
| `Message` | interface | `{ id, role, content, calls?, images? }` | Represents one conversation turn fed to a `ProviderInterface` — a stored, identified message. |
|
|
651
|
+
| `MessageInput` | interface | `{ role, content, calls?, images? }` | Carries the minimal data needed to author a `Message` — the `id` is assigned by the layer that stores it, so a caller supplies only role / content (and, for a replayed assistant turn, its `calls`). |
|
|
652
|
+
| `ProviderResult` | interface | `{ content, thinking?, tools?, usage? }` | Holds a single inference turn's structured outcome — the assembled assistant content, any reasoning the provider separated from it, any tool calls the model requested, and the token usage it reported. |
|
|
653
|
+
| `ProviderDelta` | type | `{ channel: 'content', text } \| { channel: 'thinking', text }` | Represents one streamed delta a `ProviderInterface`'s `stream` yields — a unit tagged by the channel it belongs to, so the agent loop can re-surface answer content and live reasoning separately as it pumps. |
|
|
654
|
+
| `ProviderStreamOptions` | interface | `{ think?, schema? }` | Carries the per-call options threaded into a `ProviderInterface`'s `generate` / `stream` — the bag a caller passes to influence one inference call without reconfiguring the provider instance. |
|
|
655
|
+
| `ProviderInterface` | interface | `{ id, name, format? } plus generate, stream` | Defines the pluggable LLM inference boundary — the one contract every agent chunk depends on. A provider turns a conversation (plus optional tools) into either a single assembled `ProviderResult` (`generate`) or a stream of `ProviderDelta`s that returns the assembled result (`stream`). |
|
|
656
|
+
| `ProviderParserInterface` | interface | `{} plus parse, clear` | Defines the structural framing seam supplied by a concrete provider. |
|
|
657
|
+
| `ProviderRequest` | interface | `{ messages, tools?, options? }` | Carries the conversation and per-call configuration sent to a provider. |
|
|
658
|
+
| `ProviderIncrement` | interface | `{ content, thinking, tools, usage?, result? }` | Holds the decoded contribution of a wire record to a provider turn. |
|
|
659
|
+
| `ProviderOptions` | interface | `{ timeout?, fetch?, headers?, format? }` | Configures a provider's deadline, transport, headers, and context framing. |
|
|
660
|
+
| `AgentProviderInput` | interface | `ProviderOptions plus { url, path?, split?, strict? }` | Configures the HTTP destination and stream assembly of a provider base. |
|
|
661
|
+
| `AgentProviderInterface` | interface | `ProviderInterface plus frame, body, read, finish` | Defines the wire-specific seams of the shared HTTP provider engine. |
|
|
662
|
+
| `ProviderErrorCode` | type | `'HTTP' \| 'PROTOCOL' \| 'PROVIDER'` | Names the machine-readable provider failure conditions. |
|
|
663
|
+
| `ProviderErrorOptions` | interface | `{ status?, cause? }` | Carries a provider failure's HTTP status and underlying cause. |
|
|
664
|
+
| `TextRead` | interface | `{ text, complete }` | Carries a decoded stream prefix and whether the stream ended within its byte budget. |
|
|
665
|
+
| `RelayFrame` | type | `ProviderDelta \| { channel: 'result', result } \| { channel: 'abort', partial } \| { channel: 'error', message }` | Carries a relay delta, settled result, remote abort, or remote failure. |
|
|
666
|
+
| `RelayHandler` | type | `(request: Request) => Promise<Response>` | Defines a host-independent relay request handler. |
|
|
667
|
+
| `RelayOptions` | interface | `{ provider, authorize, limit? }` | Configures the upstream provider, mandatory authorization, and request byte limit. |
|
|
668
|
+
| `RelayProviderOptions` | interface | `ProviderOptions plus { url, parser }` | Configures a relay destination and its fresh structural parser factory. |
|
|
669
|
+
| `RelayStreamOptions` | interface | `{ provider, request, signal }` | Carries the upstream call and cancellation bound of a relay response stream. |
|
|
670
|
+
| `ThinkSplitterInterface` | interface | `{ content, thinking } plus split, flush` | Splits a thinking model's in-content `<think>…</think>` reasoning spans away from the answer, delta by delta with per-stream state, so a provider yields clean content alone and surfaces the reasoning as `ProviderResult.thinking`. |
|
|
671
|
+
| `ContextSectionFormat` | interface | `{ open?, render?, close? }` | Overrides one context section's format — an `open` / `render` / `close` trio that frames a section in the `AgentContext` build cascade: a top line rendered once before the items, a per-item rendering, and a bottom line rendered once after the items. |
|
|
672
|
+
| `ContextFormat` | interface | `{ instructions? }` | Holds a provider's optional context-framing default, keyed by section kind — the framing a model prefers (for example XML tags against Markdown headers), declared by a `ProviderInterface` that opts in. |
|
|
673
|
+
| `ContextSectionSourceInterface` | interface | `{ open, format } plus render` | Exposes the manager surface one context section's format cascade reads — its built-in `open` / `render`, plus the raw options override the cascade layers a provider default beneath; `InstructionManagerInterface` satisfies it structurally. |
|
|
674
|
+
| `MessageManagerInterface` | interface | `{ count } plus add, message, messages, remove, clear` | Stores immutable `Message`s in insertion order and mints each `id` on `add` — the message-store contract `AgentContextInterface.messages` is typed to, which the active `ConversationInterface` satisfies structurally. |
|
|
675
|
+
| `InstructionInterface` | interface | `{ id, name, content, priority, override? }` | Represents an immutable instruction — a named directive a richer context places between the system prompt and the conversation, ordered by descending `priority`. |
|
|
676
|
+
| `InstructionInput` | interface | `{ name, content, priority?, override? }` | Carries the minimal data to author an `InstructionInterface` — the `id` is minted by the `InstructionManagerInterface` that stores it, so a caller supplies only `name` / `content` (and an optional `priority`, defaulting to `0`). |
|
|
677
|
+
| `InstructionManagerEventMap` | type | `{ add, remove, clear }` | Maps the push observation surface of an `InstructionManagerInterface` — the mutation moments a fire-and-forget observer subscribes to through `manager.emitter.on`. |
|
|
678
|
+
| `InstructionManagerOptions` | interface | `{ on?, error?, format? }` | Configures `createInstructionManager` — the reserved `on` hooks plus an optional per-section format override. |
|
|
679
|
+
| `InstructionManagerInterface` | interface | `{ emitter, count, open, format } plus add, instruction, instructions, render, remove, clear` | Registers `InstructionInterface`s keyed by `name` — `add` (one or a batch) mints each `id` and overwrites a same-name instruction, last write wins, while `instructions()` lists them sorted by descending `priority` and stable for ties. |
|
|
680
|
+
| `ScopeFilter` | interface | `{ instructions?, tools?, files? }` | Lists the per-category allow-lists a `ScopeInterface` carries — an optional `readonly string[]` for `instructions`, for `tools`, and for `files`, each keyed by that category's identity (an instruction's `name`, a tool's `name`, a workspace file's `path`) and read as an allow-list: `undefined` lets everything pass, `[]` lets nothing pass, and a non-empty list passes the listed keys alone. |
|
|
681
|
+
| `ScopeInput` | interface | `ScopeFilter plus { name }` | Carries the data to author a `ScopeInterface` — a `ScopeFilter` plus the required `name` (a human label; the `id` is minted by the layer that stores it). |
|
|
682
|
+
| `ScopeInterface` | interface | `ScopeFilter plus { id, name } plus narrow` | Represents a named, immutable filter over a richer context's items — the per-category allow-lists (`ScopeFilter`) plus an `id` / `name`, and a `narrow` that composes a tighter child by set intersection. |
|
|
683
|
+
| `ScopeManagerEventMap` | type | `{ create, remove, clear }` | Maps the push observation surface of a `ScopeManagerInterface` — analogous to `InstructionManagerEventMap`, but keyed by the minted `id` and carrying `create` (a scope always mints, never overwrites) rather than `add`. |
|
|
684
|
+
| `ScopeManagerOptions` | interface | `{ on?, error? }` | Configures `createScopeManager` — the reserved `on` hooks: initial listeners for the manager's `ScopeManagerEventMap`, wired at construction. |
|
|
685
|
+
| `ScopeManagerInterface` | interface | `{ emitter, count } plus create, scope, scopes, remove, clear` | Registers reusable `ScopeInterface`s keyed by their minted `id` — `create` mints + stores one (never overwrites), `scopes()` lists them in insertion order. |
|
|
686
|
+
| `AgentContextOptions` | interface | `{ system?, tools?, instructions?, workspaces?, scope?, conversations? }` | Configures `createAgentContext` — the optional system prompt plus the pre-built managers to reuse: an `instructions` registry, a `workspaces` registry (the only document channel), a `conversations` registry (the message source), a `tools` registry (the loop's advertise and dispatch surface), and an initial `scope`. |
|
|
687
|
+
| `AgentContextInterface` | interface | `{ system, instructions, workspaces, messages, conversations, tools, scope } plus apply, build` | Assembles a turn's provider input from the system prompt + the context managers + the conversation, applying the active scope per category. |
|
|
688
|
+
| `AgentStatus` | type | `'idle' \| 'running' \| 'done' \| 'error'` | Names the lifecycle state of an `AgentInterface` turn — `idle` before a run, `running` while the loop is in flight, then the settled `done` (a normal finish or a cancel) or `error` (a genuine provider / tool failure). |
|
|
689
|
+
| `AgentChunk` | type | `{ category: 'token', content } \| { category: 'think', content } \| { category: 'tool', call, result } \| { category: 'usage', usage }` | Represents a streamed step of an agent turn — the union the loop yields as it runs, discriminated by the `category` of step it carries, and the pull surface beside the push `AgentEventMap`. |
|
|
690
|
+
| `AgentEventMap` | type | `{ start, turn, tool, usage, deny, finish, error, abort, exhaust, fault }` | Maps the push observation surface of an `AgentInterface` — the lifecycle, usage, and tool moments a fire-and-forget observer (logging, metrics, tracing) subscribes to, beside the pull `AgentChunk` stream. |
|
|
691
|
+
| `AgentResult` | interface | `{ content, thinking?, usage?, partial }` | Holds the settled outcome of an agent turn — the assembled assistant `content`, the `usage` summed across the turn's provider calls, and whether it was committed `partial`. |
|
|
692
|
+
| `RunOutcome` | interface | `{ content, thinking, usage, partial, exhausted }` | Holds the immutable per-run outcome an `AgentInterface`'s loop settles on — the value its run returns, assembled from there into the `AgentResult` its `stream`'s `result` promise resolves. |
|
|
693
|
+
| `ChannelInterface` | interface | `{} plus push, close, fail, drain` | Buffers values in an unbounded async channel — a producer writes them in (`push`) and ends it (`close` / `fail`) regardless of consumption, while a consumer reads them back live through `drain`. |
|
|
694
|
+
| `StreamInterface` | interface | `{ events, result } plus abort` | Pairs a live event stream with the eventual settled result and a cancel — the generic pull/streaming handle a long-running operation hands back. |
|
|
695
|
+
| `AgentStreamInterface` | type | `StreamInterface<AgentChunk, AgentResult>` | Names the agent turn's live handle — a `StreamInterface` of `AgentChunk`s resolving an `AgentResult`. |
|
|
696
|
+
| `AgentOptions` | interface | `{ on?, error?, system?, tools?, instructions?, workspaces?, scope?, limit?, timeout?, budget?, scheduler?, signal?, authority?, conversations?, window?, strict? }` | Configures `createAgent` — the loop's bounds and pacing, the reserved `on` hooks, the construction-time context wiring (`instructions` / `workspaces` / `scope`), the `conversations` registry that is the message source, the context `window` budget that opts into automatic compaction of the active conversation, and the `strict` switch that aborts the run on an automatic-compaction summarizer failure instead of the lenient default. |
|
|
697
|
+
| `AgentRunOptions` | interface | `{ think?, schema?, limit?, timeout?, budget?, signal? }` | Carries the per-run override bag an `AgentInterface`'s `generate` / `stream` accepts — each member overrides the matching `AgentOptions` value for one run, where `think` and `schema` forward to the provider call and `signal` composes with the constructed one. |
|
|
698
|
+
| `AgentInterface` | interface | `{ emitter, id, status, context } plus generate, stream, abort` | Composes a `ProviderInterface`, an `AgentContextInterface`, and a `ToolManagerInterface` into a bounded context → provider → tools → repeat turn. |
|
|
699
|
+
| `AuthorityContext` | interface | `{ call }` | Carries what an `AuthorityInterface` evaluates for one tool call — the call under consideration. |
|
|
700
|
+
| `AuthorityDecision` | interface | `{ zone, allowed, reason? }` | Holds an `AuthorityInterface`'s verdict on one tool call. |
|
|
701
|
+
| `AuthorityRule` | interface | `{ match, zone, allowed?, reason? }` | Represents one ordered policy rule an `AuthorityInterface` evaluates. |
|
|
702
|
+
| `AuthorityOptions` | interface | `{ rules?, fallback? }` | Configures `createAuthority` — the ordered rules and the no-match fallback. |
|
|
703
|
+
| `AuthorityInterface` | interface | `{} plus evaluate` | Gates each tool call before it runs — the synchronous policy that turns one `AuthorityContext` into an `AuthorityDecision`. |
|
|
704
|
+
| `AgentJobInput` | interface | `{ provider, messages, system?, tools?, authority?, scheduler?, limit?, timeout?, budget?, children? }` | Represents a JSON-serializable agent job — the descriptor a durable queue or runner runs. Its non-serializable pieces (the provider, tools, authority, scheduler) are referenced by name and resolved to live objects through an `AgentRegistryInterface` at handler time, while its data fields (the seed `messages`, `system`, `limit`, `timeout`, and a token `budget` ceiling) carry directly. |
|
|
705
|
+
| `AgentRegistryInterface` | interface | `{} plus provider, tool, authority, scheduler, build` | Resolves an `AgentJobInput`'s names to the live, non-serializable pieces and rehydrates a seeded, signal-wired `AgentInterface` — the bridge that makes a durable, serializable job runnable. |
|
|
706
|
+
| `AgentRegistryOptions` | interface | `{ providers, tools?, authorities?, schedulers?, store? }` | Configures `createAgentRegistry` — the named pools of live, non-serializable pieces an `AgentJobInput`'s names resolve against, plus the optional durable `store` every built agent's conversation manager shares. |
|
|
707
|
+
| `AgentQueueOptions` | interface | `{ registry, partial?, concurrency?, retries?, timeout?, store? }` | Configures `createAgentQueue` — the registry that rehydrates jobs, the partial-result policy, and the substrate knobs threaded into the backing `createQueue`. |
|
|
708
|
+
| `AgentRunnerOptions` | interface | `{ registry, partial?, concurrency?, retries?, timeout? }` | Configures `createAgentRunner` — the registry that rehydrates jobs, the partial-result policy, and the substrate knobs threaded into the backing `createRunner`. |
|
|
709
|
+
| `ConversationSummaryHandler` | type | `(messages: readonly Message[]) => Promise<string>` | Summarizes a conversation, provider-agnostically — the seam the agent runtime supplies so core never imports a provider. Given the folded messages, it resolves their digest, the model-written summary used to summarize a compacted `Section` and to regenerate a `ConversationInterface`'s rollup `summary`. |
|
|
710
|
+
| `Section` | interface | `{ id, summary, messages }` | Holds a slice of folded messages digested into a summary — the unit of compaction a `ConversationInterface` produces when it `compact`s its live tail. |
|
|
711
|
+
| `ConversationEventMap` | type | `{ compact, summary, rehydrate, collapse }` | Maps the push observation surface of a `ConversationInterface` — the compaction moments a fire-and-forget observer subscribes to through `conversation.emitter.on`. |
|
|
712
|
+
| `ConversationOptions` | interface | `{ id?, on?, error?, summarize?, keep?, sections?, snapshot? }` | Configures `createConversation` — the optional `id`, the reserved `on` hooks, the provider-agnostic `summarize` seam, the retained-tail size, an optional cap on the compacted `sections` list, and a `ConversationSnapshot` to hydrate from. |
|
|
713
|
+
| `CompactOptions` | interface | `{ keep?, sections? }` | Configures one `ConversationInterface.compact` call — the retained-tail size, the `sections` cap, or both, overridden for one fold. |
|
|
714
|
+
| `ConversationReferenceOptions` | interface | `{ label?, summary?, messages? }` | Configures `ConversationInterface.reference` — how to render one conversation as a self-labeled, fenced provenance block to pull into another conversation by writing it to the active context's active workspace: `label` defaults to the `id`, `summary` defaults to `true`, and `messages` are cherry-picked excerpts defaulting to none. |
|
|
715
|
+
| `ConversationInterface` | interface | `{ id, emitter, summary, sections, summarizable, count } plus add, message, messages, remove, clear, view, compact, rehydrate, search, reference, snapshot` | Groups messages above the flat `MessageManagerInterface` — a live uncompacted tail plus compacted, summarized `Section`s and a conversation rollup `summary`, with on-demand `rehydrate`, substring `search`, a cross-conversation `reference`, and a JSON `snapshot`, driven by a provider-agnostic `ConversationSummaryHandler` seam; `summarizable` reports whether that seam was supplied, and the agent loop gates automatic compaction on it. |
|
|
716
|
+
| `ConversationInput` | interface | `{ id?, summarize?, keep?, sections?, on?, snapshot? }` | Carries the data to author a `ConversationInterface` through a `ConversationManagerInterface` — the optional `id`, a `summarize` override, a `keep` override, a `sections` cap override, the reserved `on` hooks, and a `ConversationSnapshot` to hydrate from. |
|
|
717
|
+
| `ConversationManagerOptions` | interface | `{ summarize?, keep?, sections?, store? }` | Configures `createConversationManager` — the default `ConversationSummaryHandler`, retained-tail size, and `sections` cap the conversations it creates inherit, plus the optional durable `store` backing `open` / `save`. |
|
|
718
|
+
| `ConversationManagerInterface` | interface | `{ count, active } plus conversation, conversations, add, switch, open, save, remove, clear` | Registers `ConversationInterface`s keyed by their `id`, in insertion order, with an active pointer — the id-keyed store over the conversation layer, the `active` / `switch` seam the `AgentContextInterface` renders, and the durable `open` / `save` store seam. Event-free (a registry, like `WorkspaceManagerInterface`); the observability lives on each `ConversationInterface`. |
|
|
719
|
+
| `ConversationSnapshot` | interface | `{ id, summary?, sections, messages }` | Holds a JSON-serializable snapshot of a conversation's state — its `id`, the rollup `summary`, the compacted `sections`, and the live tail `messages` — the durable payload the `ConversationStoreInterface` persists. The exact analogue of `WorkspaceSnapshot`. |
|
|
720
|
+
| `ConversationStoreInterface` | interface | `{} plus get, set, delete` | Persists a `ConversationSnapshot` durably — the async `get` / `set` / `delete` primitives, keyed by a conversation id and holding no expiry, the exact analogue of `WorkspaceStoreInterface`. |
|
|
721
|
+
| `ConversationSnapshotRow` | interface | `{ id, snapshot }` | Represents one row of the table a `DatabaseConversationStore` persists — a conversation `id` plus its `ConversationSnapshot` held as one opaque JSON column, read back as `unknown` and narrowed on `get`. The exact analogue of `WorkspaceSnapshotRow`. |
|
|
722
|
+
|
|
723
|
+
Agent-owned readonly data members stay in the preceding Surface tables; their call-signature methods are documented under [`## Methods`](#methods). Tool contracts resolve to [`tool.md`](tool.md) and workspace contracts to [`workspace.md`](workspace.md) — neither dependency surface is duplicated or re-exported here. Note where the boundary falls inside the context: `instructions`, `conversations`, and `workspaces` are the managers `build()` renders a prompt from, while `tools` is loop machinery for advertising and dispatch and is never read by `build()` at all.
|
|
724
|
+
|
|
725
|
+
## Methods
|
|
726
|
+
|
|
727
|
+
The tables list every public call-signature member of `ProviderInterface`, `AgentProviderInterface`, `ProviderParserInterface`, `RelayProvider`, `ThinkSplitterInterface`, `MessageManagerInterface`, `InstructionManagerInterface`, `ContextSectionSourceInterface`, `ScopeInterface`, `ScopeManagerInterface`, `AgentContextInterface`, `AgentInterface`, `StreamInterface`, `ChannelInterface`, `AuthorityInterface`, `AgentRegistryInterface`, `ConversationInterface`, `ConversationManagerInterface`, `ConversationStoreInterface`, `MemoryConversationStore`, and `DatabaseConversationStore`. Their readonly data members remain Surface rows. `AgentProvider`, `ThinkSplitter`, `InstructionManager`, `Scope`, `ScopeManager`, `AgentContext`, `Agent`, `Authority`, `AgentRegistry`, `Conversation`, and `ConversationManager` implement their interfaces exactly, so the tables also describe those classes' instance methods. `RelayProvider` and the store classes keep explicit tables because their class names have no same-name interface contract. `MessageManagerInterface` has no separate concrete class here: the active `Conversation` satisfies it structurally. `RelayStream` exposes `response` alone, a data member, so it keeps its Surface row and takes no table. Tool and workspace methods live in their dependency guides.
|
|
728
|
+
|
|
729
|
+
#### `ProviderInterface`
|
|
730
|
+
|
|
731
|
+
`generate` produces one complete turn; `stream` yields `ProviderDelta`s and returns the assembled result. Both take the conversation, a bounding `AbortSignal`, optional `tools`, and optional per-call `ProviderStreamOptions`.
|
|
732
|
+
|
|
733
|
+
| Method | Returns | Summary |
|
|
734
|
+
| ---------- | ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
735
|
+
| `generate` | `Promise<ProviderResult>` | Generates one complete turn — resolves the assembled `ProviderResult`. |
|
|
736
|
+
| `stream` | `AsyncGenerator<ProviderDelta, ProviderResult>` | Streams one turn — yields channel-tagged `content` / `thinking` `ProviderDelta`s as they arrive and returns the assembled `ProviderResult` (the concatenated content, any separated reasoning, any tool calls, and any usage) when the stream completes. A mid-stream abort throws a `ProviderAbortError` carrying the partial result. |
|
|
737
|
+
|
|
738
|
+
#### `AgentProviderInterface`
|
|
739
|
+
|
|
740
|
+
The seams a subclass fills, beneath the boundary members it inherits. `AgentProviderInterface` extends `ProviderInterface`, so `generate` and `stream` are part of it and repeat here with their boundary contracts; `AgentProvider` implements each of them once, for every subclass. A subclass writes `frame` / `body` / `read` / `finish` and nothing else. The `id` / `name` / `format` data members stay Surface rows.
|
|
741
|
+
|
|
742
|
+
| Method | Returns | Summary |
|
|
743
|
+
| ---------- | ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
744
|
+
| `generate` | `Promise<ProviderResult>` | Generates one complete turn — resolves the assembled `ProviderResult`. |
|
|
745
|
+
| `stream` | `AsyncGenerator<ProviderDelta, ProviderResult>` | Streams one turn — yields channel-tagged `content` / `thinking` `ProviderDelta`s as they arrive and returns the assembled `ProviderResult` (the concatenated content, any separated reasoning, any tool calls, and any usage) when the stream completes. A mid-stream abort throws a `ProviderAbortError` carrying the partial result. |
|
|
746
|
+
| `frame` | `ProviderParserInterface<TRecord>` | Creates fresh framing state for a call. |
|
|
747
|
+
| `body` | `object` | Projects a request to the concrete protocol's serializable body. |
|
|
748
|
+
| `read` | `ProviderIncrement` | Decodes a framed record into its contribution to the turn. |
|
|
749
|
+
| `finish` | `readonly TRecord[]` | Returns records retained at end of input before the parser is cleared. |
|
|
750
|
+
|
|
751
|
+
#### `ProviderParserInterface`
|
|
752
|
+
|
|
753
|
+
The framing seam a concrete provider hands the engine, one fresh instance for each call, so no call inherits another's half-read record. The engine feeds each decoded chunk to `parse` and clears the parser on every exit.
|
|
754
|
+
|
|
755
|
+
| Method | Returns | Summary |
|
|
756
|
+
| ------- | -------------------- | --------------------------------------------- |
|
|
757
|
+
| `parse` | `readonly TRecord[]` | Parses a decoded chunk into complete records. |
|
|
758
|
+
| `clear` | `void` | Clears retained framing state. |
|
|
759
|
+
|
|
760
|
+
#### `RelayProvider`
|
|
761
|
+
|
|
762
|
+
The relay wire over `AgentProvider`: it fills every seam and inherits `generate` / `stream` from the base, so a browser drives it exactly like a local provider. Its `name` reports `'relay'` and stays a data member.
|
|
763
|
+
|
|
764
|
+
| Method | Returns | Summary |
|
|
765
|
+
| -------- | -------------------------------------------------- | --------------------------------------------------------------------------------- |
|
|
766
|
+
| `frame` | `ProviderParserInterface` | Creates fresh framing state for each response. |
|
|
767
|
+
| `body` | `object` | Projects declared request fields and refuses values the JSON wire cannot carry. |
|
|
768
|
+
| `read` | `ProviderIncrement` | Validates a relay frame and translates its channel into the shared stream engine. |
|
|
769
|
+
| `finish` | `ReadonlyArray<Readonly<Record<string, unknown>>>` | Recovers an unterminated final frame by completing its NDJSON line. |
|
|
770
|
+
|
|
771
|
+
#### `ThinkSplitterInterface`
|
|
772
|
+
|
|
773
|
+
The stream-stateful `<think>…</think>` separator a provider routes raw content deltas through, so it yields clean content and surfaces the reasoning as `ProviderResult.thinking`. The `content` / `thinking` data members (the authoritative clean-content + reasoning accumulations) stay Surface rows — `content` matters because some chat templates pre-seed `<think>` into the prompt scaffold (the qwen3 shape), so only a bare `</think>` ever appears on the wire: before any tag event, that bare close reclassifies everything surfaced so far into `thinking` (one-shot — afterwards a bare close is plain text), correcting `content` retroactively where the already-returned deltas cannot be recalled. One splitter serves one stream — create a fresh one for each call (`createThinkSplitter`).
|
|
774
|
+
|
|
775
|
+
| Method | Returns | Summary |
|
|
776
|
+
| ------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
777
|
+
| `split` | `string` | Feeds one raw delta and returns the clean, non-think content to surface for it (possibly `''`) — a tag split across deltas is held until disambiguated, never leaked as content and never mis-eaten as thinking. |
|
|
778
|
+
| `flush` | `string` | Settles the stream end — a held partial tag that never completed returns as the final content delta, and an unclosed think span's tail lands on `thinking`. |
|
|
779
|
+
|
|
780
|
+
#### `MessageManagerInterface`
|
|
781
|
+
|
|
782
|
+
The immutable conversation store. `add` mints each message's `id` and carries batch overloads (one input → one message, a batch → the array); `remove` carries batch overloads (one or a list). The `count` data member stays a Surface row.
|
|
783
|
+
|
|
784
|
+
| Method | Returns | Summary |
|
|
785
|
+
| ---------- | -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
786
|
+
| `add` | `Message` / `readonly Message[]` | Stores one `MessageInput`, or a batch — mints each message's `id` and returns the created message or messages; a stored message is immutable. |
|
|
787
|
+
| `message` | `Message \| undefined` | Looks up one stored message by id (`undefined` when absent). |
|
|
788
|
+
| `messages` | `readonly Message[]` | Lists every stored message, in insertion order. |
|
|
789
|
+
| `remove` | `boolean` | Removes one message by id, or a batch — `true` only when every supplied id was removed. |
|
|
790
|
+
| `clear` | `void` | Removes every stored message. |
|
|
791
|
+
|
|
792
|
+
#### `InstructionManagerInterface`
|
|
793
|
+
|
|
794
|
+
The name-keyed instruction registry a richer context renders a directives block from. `add` mints each `id` and carries batch overloads (a re-`add` of the same name overwrites it, last write wins); `remove` carries batch overloads. The `emitter` / `count` / `open` / `format` data members stay Surface rows.
|
|
795
|
+
|
|
796
|
+
| Method | Returns | Summary |
|
|
797
|
+
| -------------- | ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- |
|
|
798
|
+
| `add` | `InstructionInterface` / `readonly InstructionInterface[]` | Adds one `InstructionInput`, or a batch — mints each `id`; a re-`add` of the same name overwrites it, last write wins. |
|
|
799
|
+
| `instruction` | `InstructionInterface \| undefined` | Looks up one instruction by name (`undefined` when absent). |
|
|
800
|
+
| `instructions` | `readonly InstructionInterface[]` | Lists every instruction, sorted by descending `priority` (stable for equal priorities). |
|
|
801
|
+
| `render` | `string` | Renders one instruction for the prompt — its `content`. |
|
|
802
|
+
| `remove` | `boolean` | Removes one instruction by name, or a batch — `true` only when every supplied name was removed. |
|
|
803
|
+
| `clear` | `void` | Removes every instruction. |
|
|
804
|
+
|
|
805
|
+
#### `ContextSectionSourceInterface`
|
|
806
|
+
|
|
807
|
+
The manager surface one section's format cascade reads. `render` is its only method — the `open` (the built-in header) and `format` (the raw manager-options override) data members stay Surface rows. An `InstructionManagerInterface` satisfies it structurally, which is what lets `resolveOpen` / `resolveClose` / `resolveItem` stay independent of which manager supplies the section.
|
|
808
|
+
|
|
809
|
+
| Method | Returns | Summary |
|
|
810
|
+
| -------- | -------- | ------------------------------------------------------------------------------------------------------------------------ |
|
|
811
|
+
| `render` | `string` | Renders one section item, already resolved against the manager-options override and otherwise on the built-in rendering. |
|
|
812
|
+
|
|
813
|
+
#### `ScopeInterface`
|
|
814
|
+
|
|
815
|
+
The named, immutable allow-list filter. `narrow` is the only method — the `id` / `name` data members and the per-category allow-lists stay Surface rows.
|
|
816
|
+
|
|
817
|
+
| Method | Returns | Summary |
|
|
818
|
+
| -------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
819
|
+
| `narrow` | `ScopeInterface` | Composes a tighter child scope — each category is the set intersection of this scope's list and `config`'s (an `undefined` side imposing no constraint), returned as a new scope that leaves this one unchanged. |
|
|
820
|
+
|
|
821
|
+
#### `ScopeManagerInterface`
|
|
822
|
+
|
|
823
|
+
The id-keyed registry of reusable named scopes. `create` mints + stores a scope (always adds — never overwrites, since two scopes may share a `name`); `remove` carries batch overloads. The `emitter` / `count` data members stay Surface rows.
|
|
824
|
+
|
|
825
|
+
| Method | Returns | Summary |
|
|
826
|
+
| -------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
|
|
827
|
+
| `create` | `ScopeInterface` | Mints a scope from a `ScopeInput` (an `id` plus the per-category allow-lists) and stores it — always adds, never overwrites. |
|
|
828
|
+
| `scope` | `ScopeInterface \| undefined` | Looks up one scope by id (`undefined` when absent). |
|
|
829
|
+
| `scopes` | `readonly ScopeInterface[]` | Lists every scope, in insertion order. |
|
|
830
|
+
| `remove` | `boolean` | Removes one scope by id, or a batch — `true` only when every supplied id was removed. |
|
|
831
|
+
| `clear` | `void` | Removes every scope. |
|
|
832
|
+
|
|
833
|
+
#### `AgentContextInterface`
|
|
834
|
+
|
|
835
|
+
The richer turn context. `apply` changes the active per-turn filter and `build` assembles the provider input. The `system` / `instructions` / `messages` / `tools` / `scope` / `workspaces` / `conversations` readonly data members stay Surface rows.
|
|
836
|
+
|
|
837
|
+
| Method | Returns | Summary |
|
|
838
|
+
| ------- | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
839
|
+
| `apply` | `void` | Applies the given scope as the active per-turn filter; passing `undefined` explicitly removes filtering. |
|
|
840
|
+
| `build` | `readonly Message[]` | Builds the provider input for the next turn: a leading `system` message folding the prompt, the scope-filtered instructions (each section's header and each item's rendering resolved through the format cascade), and the active workspace's scope-filtered (`scope.files`) text files as fenced reference blocks in a `## Workspace` section, then the active conversation's `view()`, with the active workspace's image files' `base64` payload attached to the last user message. Takes an optional `format` — typically `provider.format`, the provider level of the cascade — and omitting it with no overrides set renders each section on its manager's built-in framing. The `system` message is prepended only when some part of it exists, the workspace render covers the active workspace alone, tools are advertised structurally rather than in the prompt, and the input is built fresh on each call. |
|
|
841
|
+
|
|
842
|
+
#### `AgentInterface`
|
|
843
|
+
|
|
844
|
+
The bounded agent loop. `generate` and `stream` share one private run (`generate` drains the same stream `stream` exposes, so they can't diverge); `abort` cancels the in-flight turn. The `emitter` / `id` / `status` / `context` data members stay Surface rows (`emitter` is a `readonly` accessor — a property, not a method).
|
|
845
|
+
|
|
846
|
+
| Method | Returns | Summary |
|
|
847
|
+
| ---------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
848
|
+
| `generate` | `Promise<AgentResult>` | Runs the turn to completion, discarding the live chunks — drains the shared stream and resolves the settled `AgentResult` (`partial: true` when cancelled). |
|
|
849
|
+
| `stream` | `AgentStreamInterface` | Runs the turn as a live stream — iterate `events` for `AgentChunk`s and `await result` for the settled outcome; `result` resolves partial on a cancel and rejects on a genuine error. |
|
|
850
|
+
| `abort` | `void` | Cancels the in-flight turn — fires the turn's signal; the `result` settles `partial: true` with whatever content accumulated. |
|
|
851
|
+
|
|
852
|
+
#### `StreamInterface`
|
|
853
|
+
|
|
854
|
+
The generic live handle pairs its `events` and `result` data members with a cancellation method. `AgentStreamInterface` specializes it for `AgentChunk` and `AgentResult`.
|
|
855
|
+
|
|
856
|
+
| Method | Returns | Summary |
|
|
857
|
+
| ------- | ------- | --------------------------------------------------------- |
|
|
858
|
+
| `abort` | `void` | Cancels the in-flight operation — fires its bound signal. |
|
|
859
|
+
|
|
860
|
+
#### `ChannelInterface`
|
|
861
|
+
|
|
862
|
+
The unbounded async channel. A producer writes with `push` and ends it with `close` or `fail`; a consumer reads it back live with `drain`. Write and read are decoupled, so the producer never waits for a consumer — an agent's eager pump writes each chunk into one, which is why the run's `result` settles whether or not `events` is ever drained. It carries no data members.
|
|
863
|
+
|
|
864
|
+
| Method | Returns | Summary |
|
|
865
|
+
| ------- | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
|
|
866
|
+
| `push` | `void` | Writes one value — buffered, then handed to a parked consumer; a value pushed at an already-parked reader is delivered, never dropped. |
|
|
867
|
+
| `close` | `void` | Ends the channel normally — a draining consumer returns once the buffer is empty. |
|
|
868
|
+
| `fail` | `void` | Ends the channel with a failure — a draining consumer throws it once the buffer is empty; the first failure wins. |
|
|
869
|
+
| `drain` | `AsyncGenerator<T, void>` | Reads the values back live, in write order — returning on `close` and throwing on `fail`. |
|
|
870
|
+
|
|
871
|
+
Buffered values are always delivered before the end is reported, so a `close` or `fail` arriving alongside the last values still hands them over first:
|
|
872
|
+
|
|
873
|
+
```ts
|
|
874
|
+
import { createChannel } from '@orkestrel/agent'
|
|
875
|
+
|
|
876
|
+
const channel = createChannel<number>()
|
|
877
|
+
channel.push(1)
|
|
878
|
+
channel.close()
|
|
879
|
+
for await (const value of channel.drain()) {
|
|
880
|
+
value // 1
|
|
881
|
+
}
|
|
882
|
+
|
|
883
|
+
const failing = createChannel<number>()
|
|
884
|
+
failing.push(2)
|
|
885
|
+
failing.fail(new Error('upstream died')) // the 2 is delivered, then the drain throws
|
|
886
|
+
```
|
|
887
|
+
|
|
888
|
+
#### `AuthorityInterface`
|
|
889
|
+
|
|
890
|
+
The synchronous policy gate the agent loop consults before each tool call. `evaluate` is the only method — it has no data members.
|
|
891
|
+
|
|
892
|
+
| Method | Returns | Summary |
|
|
893
|
+
| ---------- | ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
894
|
+
| `evaluate` | `AuthorityDecision` | Evaluates one tool call against the ordered rules — returns the first matching rule's verdict, which allows unless `allowed: false`, or the fallback when none match. |
|
|
895
|
+
|
|
896
|
+
#### `AgentRegistryInterface`
|
|
897
|
+
|
|
898
|
+
The job-rehydration bridge. `provider` / `tool` / `authority` / `scheduler` resolve a name against their pool (throwing `unknown <category>: <name>` on a miss); `build` rehydrates a seeded, signal-wired agent from a serializable job. It has no data members.
|
|
899
|
+
|
|
900
|
+
| Method | Returns | Summary |
|
|
901
|
+
| ----------- | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
902
|
+
| `provider` | `ProviderInterface` | Resolves a registered `ProviderInterface` by name — throws `unknown provider: <name>` when absent. |
|
|
903
|
+
| `tool` | `ToolInterface` | Resolves a registered `ToolInterface` by name — throws `unknown tool: <name>` when absent. |
|
|
904
|
+
| `authority` | `AuthorityInterface` | Resolves a registered `AuthorityInterface` by name — throws `unknown authority: <name>` when absent. |
|
|
905
|
+
| `scheduler` | `SchedulerInterface` | Resolves a registered `SchedulerInterface` by name — throws `unknown scheduler: <name>` when absent. |
|
|
906
|
+
| `build` | `AgentInterface` | Rehydrates a live, seeded `AgentInterface` from a serializable `AgentJobInput` — resolving its names, rebuilding its token budget, seeding its conversation, and wiring `signal`; a name absent from its pool throws. |
|
|
907
|
+
|
|
908
|
+
#### `ConversationInterface`
|
|
909
|
+
|
|
910
|
+
A conversation that owns its live message tail directly (the flat store verbs folded in, like a `Workspace` owns its files). `add` mints each message's `id` and stores it (batch overloads); `message` / `messages` look up the live tail; `remove` / `clear` drop from it. `view` is the model input; `compact` folds the older live messages into a summarized `Section` (regenerating the rollup, emitting `summary` then `compact`); `rehydrate` / `search` read the retained originals; `reference` renders this conversation as a provenance-labeled block to pull into another (a pure string, no model call). The `id` / `emitter` / `summary` / `sections` / `count` data members stay Surface rows (`emitter` is a `readonly` accessor — a property, not a method).
|
|
911
|
+
|
|
912
|
+
| Method | Returns | Summary |
|
|
913
|
+
| ----------- | -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
914
|
+
| `add` | `Message` / `readonly Message[]` | Appends one `MessageInput` to the live tail, or a batch — mints each message's `id` (a random UUID) and returns the created message or messages; a stored message is immutable. |
|
|
915
|
+
| `message` | `Message \| undefined` | Looks up one live message by id (`undefined` when absent). |
|
|
916
|
+
| `messages` | `readonly Message[]` | Lists every live, uncompacted message in the tail, in insertion order. |
|
|
917
|
+
| `remove` | `boolean` | Removes one live message by id, or a batch, from the tail — `true` only when every supplied id was removed. |
|
|
918
|
+
| `clear` | `void` | Empties the live tail, leaving the compacted `sections` untouched. |
|
|
919
|
+
| `view` | `readonly Message[]` | Builds the model input for the next turn — each section as one synthetic recap message, its summary prefixed with `CONVERSATION_RECAP_PREFIX` so a small model reads it as a recap rather than a literal turn, then the live tail verbatim; the rollup `summary` is not injected. |
|
|
920
|
+
| `compact` | `Promise<Section \| undefined>` | Folds the oldest `count - keep` live messages into a summarized `Section` through the `ConversationSummaryHandler`, removes them from the live tail, regenerates the rollup, and emits `summary` then `compact` — resolving `undefined` when nothing folds (`count <= keep`). Throws a `ConversationError` when no summarizer was supplied. |
|
|
921
|
+
| `rehydrate` | `readonly Message[]` | Returns a section's full original messages — a pure read that emits `rehydrate`, empty for an unknown id and never reinserting. |
|
|
922
|
+
| `search` | `readonly Message[]` | Searches `content` for a case-insensitive substring across every message — each section's retained originals, then the live tail. |
|
|
923
|
+
| `reference` | `string` | Renders this conversation as a self-labeled, fenced provenance block to pull into another conversation — a pure string with no model call: a leading `[Reference — conversation "<label>" — NOT part of this conversation]` marker, the rollup `Summary:` when `summary` is not `false` and a rollup exists, and the cherry-picked excerpts (`- role: content`) when `messages` is supplied. `label` defaults to the `id`. |
|
|
924
|
+
| `snapshot` | `ConversationSnapshot` | Serializes this conversation to a plain, JSON-serializable `ConversationSnapshot` — its `id`, the rollup `summary`, the compacted `sections`, and the live tail; the live `summarize` / `keep` are configuration re-supplied on hydrate rather than serialized. |
|
|
925
|
+
|
|
926
|
+
#### `ConversationManagerInterface`
|
|
927
|
+
|
|
928
|
+
The id-keyed registry of `Conversation`s with an active pointer. `add(input?)` mints a conversation (flowing the manager's default `summarize` / `keep` in unless the input overrides them) and auto-activates the first one; a later `add` leaves `active` unchanged. `switch(id)` re-points `active` (an unknown `id` returns `undefined`, leaving `active` unchanged — lenient, never throws); `remove` carries batch overloads (the array overload first) and clears `active` when the removed conversation was active. The `count` and `active` data members stay Surface rows (`active` is a `readonly` accessor — a property, not a method); the manager is event-free (each conversation owns its `emitter`).
|
|
929
|
+
|
|
930
|
+
| Method | Returns | Summary |
|
|
931
|
+
| --------------- | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
932
|
+
| `conversation` | `ConversationInterface \| undefined` | Looks up one conversation by id (`undefined` when absent). |
|
|
933
|
+
| `conversations` | `readonly ConversationInterface[]` | Lists every conversation, in insertion order. |
|
|
934
|
+
| `add` | `ConversationInterface` | Mints a conversation, taking its `id` from the input or a fresh UUID and flowing the manager's default `summarize` / `keep` in unless the input overrides them — auto-activates the first, and an already-present `id` overwrites, last write wins. |
|
|
935
|
+
| `switch` | `ConversationInterface \| undefined` | Re-points `active` at the conversation with `id` and returns it; an unknown `id` returns `undefined` and leaves `active` unchanged, never throwing. |
|
|
936
|
+
| `open` | `Promise<ConversationInterface \| undefined>` | Resolves a conversation by id and activates it — from the registry when present, else hydrated from the optional `ConversationStoreInterface` (`store`); `undefined` when it is neither registered nor stored. |
|
|
937
|
+
| `save` | `Promise<boolean>` | Persists a registered conversation's `ConversationInterface.snapshot` to the optional `ConversationStoreInterface` (`store`) — `true` when persisted, `false` when there is no store or the id is unknown, and never throwing. |
|
|
938
|
+
| `remove` | `boolean` | Removes one conversation by id, or a batch — `true` only when every supplied id was removed; clears `active` when a removed conversation was the active one. |
|
|
939
|
+
| `clear` | `void` | Removes every conversation and clears `active`. |
|
|
940
|
+
|
|
941
|
+
#### `ConversationStoreInterface`
|
|
942
|
+
|
|
943
|
+
The persistence contract stores a `ConversationSnapshot` under its own identity.
|
|
944
|
+
|
|
945
|
+
| Method | Returns | Summary |
|
|
946
|
+
| -------- | -------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
|
|
947
|
+
| `get` | `Promise<ConversationSnapshot \| undefined>` | Resolves the persisted snapshot for `id`, or `undefined` if none is stored. |
|
|
948
|
+
| `set` | `Promise<void>` | Inserts or replaces a snapshot under its own `snapshot.id` (no separate id param — mirroring `WorkspaceStoreInterface`'s `set`). |
|
|
949
|
+
| `delete` | `Promise<void>` | Drops a snapshot by id; an absent id is a no-op (no throw). |
|
|
950
|
+
|
|
951
|
+
#### `MemoryConversationStore`
|
|
952
|
+
|
|
953
|
+
The in-memory implementation keeps an explicit table because its class name has no same-name interface contract.
|
|
954
|
+
|
|
955
|
+
| Method | Returns | Summary |
|
|
956
|
+
| -------- | -------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
|
|
957
|
+
| `get` | `Promise<ConversationSnapshot \| undefined>` | Resolves the persisted snapshot for `id`, or `undefined` if none is stored. |
|
|
958
|
+
| `set` | `Promise<void>` | Inserts or replaces a snapshot under its own `snapshot.id` (no separate id param — mirroring `WorkspaceStoreInterface`'s `set`). |
|
|
959
|
+
| `delete` | `Promise<void>` | Drops a snapshot by id; an absent id is a no-op (no throw). |
|
|
960
|
+
|
|
961
|
+
#### `DatabaseConversationStore`
|
|
962
|
+
|
|
963
|
+
The driver-backed implementation keeps an explicit table because its class name has no same-name interface contract.
|
|
964
|
+
|
|
965
|
+
| Method | Returns | Summary |
|
|
966
|
+
| -------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
|
|
967
|
+
| `get` | `Promise<ConversationSnapshot \| undefined>` | Resolves the persisted snapshot for `id`, narrowing the opaque JSON column back to a `ConversationSnapshot`. |
|
|
968
|
+
| `set` | `Promise<void>` | Inserts or replaces under the snapshot's own `id` (no separate id param) — the row is `{ id, snapshot }`. |
|
|
969
|
+
| `delete` | `Promise<void>` | Drops a snapshot by id; an absent id is a no-op (no throw). |
|
|
970
|
+
|
|
971
|
+
## Contract
|
|
972
|
+
|
|
973
|
+
These invariants hold across `src/core` ↔ `agent.md`:
|
|
974
|
+
|
|
975
|
+
1. **Doc ↔ source bijection.** Every `function` / `class` / `const` / `interface` / `type` row in the `## Surface` tables is a real export of `src/core`, and every export appears as a Surface row — exhaustive, both directions.
|
|
976
|
+
2. **`ProviderInterface` is the inference boundary; `AgentProvider` is the engine behind it.** A provider turns a conversation (plus optional `tools`, a non-empty `ToolDefinition[]`) into a turn: `generate` resolves the assembled `ProviderResult`, `stream` yields channel-tagged `ProviderDelta`s and returns the assembled result. Both methods accept optional `ProviderStreamOptions`; `think` is the per-call reasoning override. It carries an `id` (a per-instance trace label) and `name` (the backend identifier). This module defines that contract and the host-independent HTTP engine every provider over it would otherwise repeat — the deadline, the transport, the authorization hook, the bounded error read, the decode loop, reasoning separation, and result assembly (the provider-engine clause). One vendor's wire stays the concrete provider's own: it reaches the engine through `frame` / `body` / `read` / `finish`, and nothing in `src/core` names a vendor.
|
|
977
|
+
3. **`stream` yields deltas + returns the assembled result.** A `ProviderInterface.stream` yields each non-empty answer delta as `{ channel: 'content', text }` and each native live reasoning delta as `{ channel: 'thinking', text }`; its return value is the assembled `ProviderResult` whose `content` is the provider's authoritative clean answer, plus any tool calls and usage the turn reported. So a caller can render tokens and reasoning live while still recovering the complete outcome from the generator's return value. A thinking model's reasoning is separated once, at the wire that carried it. A provider reading a raw wire routes its content through a `ThinkSplitter` (`createThinkSplitter` — one per stream, armed by the base's `split` switch), yields the clean content, assembles `content` from the splitter's authoritative accumulation (an implicit pre-seeded open — the qwen3-template bare `</think>` — reclassifies the already-yielded prefix into thinking, the one shape where live content deltas can transiently over-report), and surfaces the accumulated reasoning as `ProviderResult.thinking` — which never re-enters the conversation (the `Agent` joins it across a run's calls onto `AgentResult.thinking` as display/audit metadata). A relay's deltas arrive already separated by the upstream provider that did that work, so `RelayProvider` constructs with `split: false` and preserves each delta verbatim, a literal `<think>` tag in the answer included: splitting a second time would re-classify text the upstream already ruled on (the relay-protocol clause).
|
|
978
|
+
4. **Usage reuses `TokenUsage`.** `ProviderResult.usage` is the [budgets](budget.md) `TokenUsage` shape (`{ prompt, completion, total }`), imported not redefined, present only when the turn reported it — so a caller folds it straight into a token budget. A provider surfaces it when the wire carries it (for example a stream's `done` line / a non-stream body) and omits it otherwise.
|
|
979
|
+
5. **The caller's signal and the engine's own deadline both bound the call.** Both `generate` and `stream` take an `AbortSignal`, so a caller bounds the request — a cancel, a [timeout](timeout.md), and a token [budget](budget.md) folded into one signal through `AbortSignal.any`. An already-aborted signal rejects the call before any content streams. `AgentProvider` arms a second bound the caller does not have to supply: a per-call `Timeout` of `AgentProviderInput.timeout` milliseconds (`DEFAULT_PROVIDER_TIMEOUT` when omitted), folded with the caller's signal through `AbortSignal.any`. Whichever trips first cancels the call, the deadline covers the `headers` hook and the error-body read as well as the stream, and it is cleared in a `finally` on every exit — a completion, a transport rejection, a non-OK response, a decoder failure, an early generator return, and a remote abort alike.
|
|
980
|
+
6. **A local cancel and a remotely reported one are different failures.** A `stream` cancelled mid-flight throws a `ProviderAbortError` whose `partial` is the `ProviderResult` assembled from whatever streamed so far (content + any tool calls + any usage); `isProviderAbortError` narrows a caught value so the loop can recover the partial content. A non-abort error propagates unchanged. `ProviderAbortError` is the boundary's error — it stays in this module so the agent loop catches it regardless of which backend is in use. The origins stay distinguishable. A **local** cancel is the combined bound firing: the engine assembles the partial from what it decoded locally, and when a throw raced that cancel it rides along as the error's `cause` rather than replacing it. A **remote** abort is a report, not a cancel: a relay `abort` frame makes `RelayProvider.read` throw a reconstructed `ProviderAbortError` carrying the upstream's own partial while the local signal stays unaborted, so the engine propagates that instance unchanged instead of folding it into a local partial, and the `Agent` runtime treats it as a genuine error unless its own bound signal aborted. A caller that cancels locally therefore always outranks the wire: the local partial is what it gets.
|
|
981
|
+
7. **The message-store contract (`MessageManagerInterface`).** `MessageManagerInterface` is the immutable message-store contract `AgentContextInterface.messages` is typed to — `count` + `add` / `message` / `messages` / `remove` / `clear`. It has no concrete class in this module: the active `Conversation` (which owns its live tail directly) satisfies it structurally, so `context.messages` is the active conversation (the conversation-layer clause). `add` takes one `MessageInput` or a batch and mints each message's `id` (`crypto.randomUUID()`), carrying the input's `role` / `content` and `calls` only when supplied (an absent `calls` is omitted, never present-but-undefined) — returning the created message(s). A stored message is immutable: assembled once from its input and never mutated, and the object `add` returns is the same one `message(id)` later resolves. `count` is the live tail size, `message(id)` looks one up (`undefined` when absent), `messages()` lists them in insertion order, `remove` (one or a batch) returns `true` only when every supplied id was removed, and `clear` empties it.
|
|
982
|
+
8. **The richer turn context (`AgentContext`).** `AgentContext` composes the optional `system` prompt; the agent-owned `instructions` / `conversations` managers; a `WorkspaceManagerInterface` consumed from `@orkestrel/workspace`; `messages` (the active conversation's live tail); a `ToolManagerInterface` consumed from `@orkestrel/tool` solely for provider advertising and call dispatch; and a readonly active `scope`. Each omitted manager is created fresh, and the context ensures the conversation registry has an active default, so `messages` is always defined. `build()` folds the system prompt, scope-filtered instructions, and the active workspace's scope-filtered text files (the active-workspace clause) into one leading `system` message, then appends the active conversation's `view()`. The tool registry is the one member `build()` never reads: it is loop machinery for advertising and dispatch, not a prompt-context manager, so no tool renders into the prompt. The result is computed fresh on every call and never mutates a manager or stored message.
|
|
983
|
+
9. **Scope filtering + image attachment.** A `Scope` is a named allow-list filter with one list per category (`instructions` by `name`, `tools` by `name`, `files` by the active workspace file's `path`), each three-way through `filterAllowList`: `undefined` ⇒ all pass, `[]` ⇒ none, a non-empty list ⇒ only-listed. `build()` applies the active `scope` to the instruction + workspace-file categories before rendering them (`scope.files` filters the active workspace's `files()` before the carrier split — the active-workspace clause). Conversation messages are deliberately not a scope category: the active conversation owns message inclusion through compaction (the active-conversation clause), so its `view()` is authoritative and a second, competing message filter would only let them disagree. `narrow(config)` composes a tighter child by set-intersection over every category's list (an `undefined` side imposes no constraint — `undefined ∩ list = list`, `undefined ∩ undefined = undefined`), so narrowing only tightens. The active workspace's scoped-in image files' `base64` payload is attached to the last user message (rebuilt as a copy carrying the merged `images`, never mutating the stored message; skipped when no user message exists — the active workspace is the sole image source); `build()` still returns `readonly Message[]`. **Tools stay structural, never in the prompt.** The loop hands the model its tools through `tools.definitions()` (the `tools` argument to `provider.generate` / `.stream`), not by serializing them — and it filters those definitions by the active `scope.tools` first (through `filterAllowList`), so a scoped-out tool is neither advertised nor callable (the model never sees it). `AgentContext.build()`'s output therefore never contains a tool's `name`, `description`, `parameters`, or definition, scoped or not. A `ScopeManager` (`createScopeManager`) is the optional reuse registry of named scopes, keyed by a minted `id` (two scopes may share a `name`), observable like the other managers.
|
|
984
|
+
10. **The agent loop (`Agent` / `createAgent`).** `Agent` composes a `ProviderInterface`, an `AgentContext`, and its `@orkestrel/tool` registry into the bounded context → provider → tools → repeat turn. It builds the provider input, then iterates up to `limit`: stream a provider turn, accumulate content/thinking/usage, append any assistant tool calls, and execute them. Each discriminated `ToolResult` becomes a tool message whose content is `JSON.stringify(result.value)` when `result.success` is true and `result.error` when false, so failure text is never JSON-quoted. A final assistant turn stops the loop.
|
|
985
|
+
11. **One `#run` shared by `generate` + `stream`.** A single private async generator drives the whole turn. `stream()` exposes it as an `AgentStreamInterface` (`events` + `result` + `abort`); `generate()` drains that same stream (iterating `events`, discarding chunks) and returns its `result` — it has zero loop logic of its own, so `generate` and `stream` can never diverge (a `generate()` result deep-equals draining `stream()` on the same input).
|
|
986
|
+
12. **The `AgentChunk` stream.** `stream().events` yields, per turn, each content delta as `{ category: 'token', content }`, each live reasoning delta as `{ category: 'think', content }`, then optional usage, then one `{ category: 'tool', call, result }` per dispatched call. The call and discriminated result use the contracts imported from `@orkestrel/tool`. `result` resolves the settled `AgentResult`: final or partial content, joined thinking, summed optional usage, and `partial`.
|
|
987
|
+
13. **Bounded, paced, capped.** Each run arms one cancel through `createAbort({ signal: AbortSignal.any([…]) })` folding whichever of the external `signal`, the `timeout` deadline (a started `Timeout`), and the `budget` signal (a started `Budget`) are present; `agent.abort(reason)` / `stream.abort(reason)` fires it, and the timeout is always cleared in a `finally`. A cancel — external, deadline, budget, or `abort()` — stops the loop and the `result` promise resolves `{ partial: true, content: <accumulated> }`; it is not an error. The accumulated `content` already holds whatever streamed before the cancel — the loop accumulates each delta as it yields the `token` chunk, and a `ProviderAbortError.partial.content` is exactly those same yielded deltas (the `stream` contract), so the loop never re-adds it. A genuine provider / tool error (the signal is not aborted) rejects the `result` (and `status` → `error`). The optional `scheduler.yield({ signal })`s between turns (before turns 2…N, never after the last); tool iteration is capped at `limit` so the loop always terminates. `status` transitions `idle` → `running` → `done` / `error`.
|
|
988
|
+
14. **Pull and push observation surfaces on the `Agent`; the rest event-free.** The provider contract, the tool registry, the conversation store, and the `AgentContext` itself carry no Emitter, no `EventMap`, no `on` hook — they stay purely functional (of the managers the context composes, the `InstructionManager` and `ScopeManager` carry their own `emitter`s and each `Conversation` owns one, while the `ToolManager`, `WorkspaceManager`, and `ConversationManager` are event-free — as is the context that composes them). The `Agent` itself carries both a pull and a push surface. Pull: the `AgentChunk` stream (`stream().events`) yields per-token answer deltas, per-think reasoning deltas, usage chunks, and tool chunks for a live consumer. Push: the `emitter` (`AgentEventMap`) — `start` (run begins), `turn` (each iteration), `tool` (a dispatched call + result), `usage` (a turn's usage), `deny` (an authority denial — not in the chunk stream), `finish` (the settled result), `error` (a genuine failure), `abort` (a cancel), `exhaust` (a limit exhausted with unresolved tool intent — fires instead of `abort`), and `fault` (a non-fatal automatic-compaction summarizer throw — the run continues; the automatic-compaction clause) — wired through the emitter pattern (`AgentOptions.on` hooks, the `AgentOptions.error` listener-error handler, a `readonly emitter`, `new Emitter({ on: options?.on, error: options?.error })`). Per-token / per-thinking deltas are the stream's job exclusively — there is deliberately no `token` or `think` event. The emitter isolates a listener throw (it can never escape into the settle-once / wake-park loop) and routes it to its own `error` handler (the `error` option, surfaced as `(error, event)`, not a domain event; itself guarded against re-entrancy), so observation is provably side-effect-free on the 3×-hardened loop — and every emit sits after the relevant state transition / settle, so it cannot reorder control flow. A cancelled run emits `abort` (the cancel reason) then `finish` (the settled partial); a genuine error emits `error` instead of `finish`. The loop's deterministic logic is pinned in the `src:core` mirror with a scripted `ProviderInterface`.
|
|
989
|
+
15. **Doc ↔ source method bijection.** The `## Methods` tables list exactly the public methods of every agent-owned interface named there, exhaustive in both directions, and each agent-owned concrete class exposes the same public methods as its interface. `MessageManagerInterface` is implemented structurally by the active `Conversation`; `ProviderInterface` is implemented by the host. Dependency method surfaces belong to [`tool.md`](tool.md) and [`workspace.md`](workspace.md).
|
|
990
|
+
16. **The authority gate (`Authority` / `createAuthority`).** An optional `Authority` is the synchronous policy gate consulted before each tool call runs — passed through `AgentOptions.authority`. `evaluate({ call })` walks its ordered `rules` first-match-wins: the first rule whose `match(context)` is true decides as `{ zone: rule.zone, allowed: rule.allowed ?? true, reason: rule.reason }` (a matched rule allows unless its `allowed` is explicitly `false`); when none match it returns the `fallback`, which defaults to `{ zone: DEFAULT_AUTHORITY_ZONE, allowed: true }` (allow-unmatched — a rules list of denials acts as a denylist; pass an `allowed: false` `fallback` to flip the gate to deny-by-default, an allowlist). The matcher receives the `AuthorityContext` (`{ call }`), so a rule can branch on the call's `name` and its `arguments`. Synchronous — `evaluate` returns the verdict directly. Event-free.
|
|
991
|
+
17. **A denied call is fed back, never executed (the gate's effect on the loop).** With no `authority` set, the loop sends calls straight through the tool registry. With one set, allowed calls run as a batch and each denied call becomes the `@orkestrel/tool` failure union arm `{ success: false, id, name, error }` without executing the tool. Executed results and denials merge back into original call order, so every denial still produces a tool chunk and unquoted failure-text tool message the model can react to. The gate lives in the `Agent`, not the dependency registry.
|
|
992
|
+
18. **Durable, serializable agent jobs (`AgentJobInput` + `AgentRegistry`).** An `AgentJobInput` is a JSON-serializable descriptor — not a live agent: its non-serializable pieces (the `provider`, `tools`, `authority`, `scheduler`) are referenced by name and its data (`messages`, `system`, `limit`, `timeout`, a token `budget` ceiling, nested `children`) carries directly. An `AgentRegistry` (`createAgentRegistry`) holds the named pools (`providers` required; `tools` / `authorities` / `schedulers` optional) and `build(input, signal?)` rehydrates a live, seeded `Agent`: it resolves the `provider`, assembles a fresh `ToolManager` from the `tools` names, rebuilds the token `budget` from its ceiling (`createTokenBudget({ max })`), resolves the `authority` / `scheduler` names, constructs the agent with `system` / `limit` / `timeout` and the threaded `signal`, and seeds its context with the `messages`. Because the descriptor is serializable, a job survives a crash through a Queue's `store` + `restore()` — the registry rehydrates the live pieces from the names on the way back in. An unknown name throws an `AgentError` carrying `code: 'REGISTRY'` and the message `unknown <category>: <name>`, never a silent `undefined`.
|
|
993
|
+
19. **Partial = configurable failure (`partial`).** An `Agent.generate()` resolves `partial: true` on a cancel (abort / budget / timeout) — a cancel is not an error. For a durable job that is, by default, a failure: the `createAgentQueue` / `createAgentRunner` handler throws an `AgentJobError` (carrying the partial `AgentResult`), so the Queue's retries re-run the job and a Runner's fail-fast aborts its siblings. Pass `partial: true` to treat a partial as success instead — the handler resolves the partial result rather than throwing. `isAgentJobError` narrows a caught value to recover its `partial`.
|
|
994
|
+
20. **Bounded concurrency + retries + persistence by composing the substrate (no new engine).** `createAgentQueue` returns a `QueueInterface<AgentJobInput, AgentResult>` built by `createQueue`: its handler `(input, context) => registry.build(input, context.signal).generate()` + the partial policy is the only new logic — bounded `concurrency`, `retries`, the per-attempt `timeout`, and the durable `store` (+ `restore()`) all belong to the `@orkestrel/queue` `Queue`. `createAgentRunner` returns a `RunnerInterface<AgentJobInput, AgentResult>` built by `@orkestrel/workflow`'s `createRunner` (one-shot, ordered, fail-fast). No second concurrency / orchestration engine is written; the verified `Agent` loop + `Authority` gate are untouched. Layering is `agent → (queue, workflow)` with no cycle.
|
|
995
|
+
21. **Sub-agent fan-out through `controller.spawn`; cancellation threaded.** On a `createAgentRunner`, each unit's handler receives a `ControllerInterface`; before running the (parent) job it `controller.spawn`s each of the job's declared `children` (fire-and-track, never inline-awaited — a slot-holding bounded handler awaiting its own spawn can deadlock), so each child is a real sub-agent run through the same bounded queue whose result joins the run after the declared jobs (in spawn order). `createAgentQueue` ignores `children` (a queue has no fan-out). Cancellation threads through in both: the handler passes `context.signal` (queue) / `controller.signal` (runner) into `registry.build`, so a queue / runner `abort()` or a per-attempt timeout fires the rehydrated agent's signal (which commits a partial — the `partial` policy then decides). All event-free.
|
|
996
|
+
22. **The conversation layer (`Conversation` + `ConversationManager`).** A `Conversation` groups messages above the flat `MessageManager`: `messages` is the live uncompacted tail (a real `MessageManagerInterface` a caller appends turns to), `sections` are the compacted history (oldest → newest), and `summary` is the rollup (`undefined` until the first compaction). Compaction folds older messages into summarized `Section`s through a provider-agnostic `ConversationSummaryHandler` seam — `compact(options?)` determines `keep` (`options.keep ?? the conversation's keep ?? DEFAULT_CONVERSATION_KEEP` = `0`), folds the oldest `count - keep` live messages (a no-op resolving `undefined` when `count <= keep`), summarizes that slice into a section (`id` minted, `summary` the seam's output, `messages` the retained originals), removes those messages from the live tail by id, regenerates the rollup (a second seam call over all section summaries), and emits `summary` then `compact` — a summarizer call for the folded slice and another for the rollup. `compact()` throws a `ConversationError` (`code: 'SUMMARIZER'`) when no summarizer was supplied (a conversation can still store + `view()` without one). `summarizable` is `true` exactly when a summarizer was supplied — the clean signal the agent loop's automatic compaction gates on (the automatic-compaction clause), so a non-summarizable conversation is never auto-compacted and the auto path never throws this error; a manual `compact()` still throws. `view()` is the model input — each section as one synthetic recap message (role `'assistant'`, keyed by the section's stable `id`), then the live messages verbatim (the rollup `summary` is not injected). Each recap's content is the section summary prefixed with `CONVERSATION_RECAP_PREFIX` (`'[Summary of earlier messages] '`) so a small model reads it as a condensed recap of earlier turns, not a literal assistant turn it must echo or treat as the live answer — a deliberately lean label (a fixed handful of tokens; the rollup regeneration in `compact()` re-reads the UNframed section summaries, the label being a `view()`-only presentation concern). `rehydrate(id)` returns a section's full original messages (`[]` for an unknown id) and emits `rehydrate` — a pure read (the caller decides whether to re-add them; `rehydrate` never reinserts). `search(query)` is a case-insensitive substring scan of `content` across all messages (every section's originals, then the live tail). `reference(options?)` renders this conversation as a self-labeled, fenced cross-conversation provenance block — a pure string (no model call) to pull into another conversation by writing it to the active context's active workspace: a leading `[Reference — conversation "<label>" — NOT part of this conversation]` marker (`label` defaults to the `id`), the rollup `Summary:` line when `options.summary !== false` and a rollup exists, and the cherry-picked `Relevant messages:` excerpts (each `- role: content`) when `options.messages` is supplied (default none). It frames foreign content so a small model attributes it to its source rather than reading it as part of the live thread; the cherry-pick comes from this conversation's own `search` / `rehydrate`, never its whole history. Observable: the owned `emitter` (`ConversationEventMap` — `compact` / `summary` / `rehydrate`) isolates a listener throw, routing it to its `error` handler (the `error` option), never corrupting a compaction. `ConversationManager` (`createConversationManager`) is the id-keyed registry with an active pointer (the `active` / `switch` seam the `AgentContext` renders): `add(input?)` mints a conversation flowing the manager's default `summarize` / `keep` in unless the input overrides them (an already-present `id` overwrites) and auto-activates the first one (a later `add` leaves `active`); `switch(id)` re-points `active` (an unknown `id` is a lenient `undefined`, leaving `active` unchanged); `conversation(id)` / `conversations()` look up; `remove` (one or a batch) reports `true` only when every supplied id was removed and clears `active` when the removed one was active; `clear` empties it and clears `active`; `count`. It carries the durable `open(id)` / `save(id)` seam over an optional `ConversationManagerOptions.store` (the durable-store clause) and `Conversation.snapshot()` serializes a conversation to a JSON `ConversationSnapshot`. It is event-free (each conversation owns its `emitter`). `estimateTokens(text)` is the deterministic `ceil(length / 4)` char heuristic (it never calls a provider) that `estimateMessages(messages)` sums over a message batch — the default `consumer` estimator for an agent's context `window` budget (the automatic-compaction clause), not a conversation member. The layer's deterministic logic is pinned in the `src:core` mirror with a data-stub summarizer.
|
|
997
|
+
23. **`AgentContext` folds the active conversation's view; the message source is the readonly `conversations` registry.** `messages` is the `conversations` registry's active conversation — always defined: at construction the context ensures the registry has an active conversation, `add`ing a default one when it has none. The dynamic `context.messages` getter returns that active conversation itself (the same reference it exposes — no duplication; computed on every read, so it follows `conversations.switch(id)`, with no captured copy), and `build()` folds that conversation's `view()` (the per-section summaries + the live tail) as the authoritative message inclusion — the conversation owns inclusion through compaction, so there is no competing scope category for messages; scope still filters instructions / tools / workspace files, and the active workspace's scoped-in image-data attachment to the last user message still applies to the view output. With the default (uncompacted) conversation, the message path is exactly the lean `[systemMessage?, ...messages]`. The registry is structural: supply it through `AgentContextOptions.conversations` and change its active conversation through `manager.switch(id)`. **Multi-conversation.** Because `context.messages` / `context.conversations` are read dynamically and the `Agent` reads them fresh on each run, one agent switches its active conversation between runs to serve many conversations from its `conversations` registry (the real app pattern — `manager.switch(id)` per request, creating through `manager.add({ id })` when absent, not an agent per thread): each conversation accumulates its own history and (with `window`, the automatic-compaction clause) compacts independently, one thread's sections never leaking into another. Switch between runs, never during one (the loop drives the run-entry active conversation to completion); for concurrent threads use a separate `Agent` per thread — the framework ships the switch mechanism, the app owns concurrency policy. The `Agent` forwards its `AgentOptions.conversations` straight into this context as the message source; the auto-compaction trigger over the active conversation lives in the loop (the automatic-compaction clause).
|
|
998
|
+
24. **Automatic compaction — the context `window` budget (opt-in, additive, production-hardened).** `AgentOptions.window` is a context [`Budget`](budget.md) (`BudgetInterface<readonly Message[]>`) for automatic conversation compaction, enabled only when both a `window` budget is set and the active conversation is `summarizable` (it has a summarizer — the conversation-layer clause). There is always an active conversation (the active-conversation clause), but the default one has no summarizer, so this `summarizable` gate preserves the shipped behavior: a non-summarizable conversation is never auto-compacted, and the auto path never throws the `compact()` `SUMMARIZER` error. Its `consumer` is a pluggable token estimator (for example the exported `estimateMessages`) and its `max` is the context window — the same consume-to-a-ceiling primitive as the cost `budget` (the bounded-paced-capped clause), but its ceiling action is compact instead of abort. The loop's private `#trim` runs the check at these points: (1) before the first provider request (so a resumed / already-long conversation whose initial prompt already exceeds the window compacts at once, not only after a tool turn), and (2) between turns on the tool-iteration `continue` path (after the prior turn's `usage` was folded and its assistant + tool messages were appended, before the next provider request; never after the final assistant turn that ends the loop). The `window` budget is reset (`clear()`) at run entry, so no stale `consumed` carries across runs or a conversation switch. Each check measures the absolute current prompt: it `clear()`s the budget then `consume`s the working message array — the exact next prompt (the system block + the active conversation's `view()` + this turn's appended messages, that is, what the next `provider.stream` will receive) — so `window.consumed` is the current full prompt's estimated footprint and `window.exhausted` means that prompt has reached the context window `max`. When the prompt `exhausted`s the window, `#trim` `await`s `conversation.compact()` (folding the older live tail into a summarized `Section` through the conversation's own `ConversationSummaryHandler`), then rebuilds the working message array from `context.build(provider.format)` — the same projection the loop opened with — so the run continues on the (now smaller) compacted context. No post-compact `clear()` is needed: the next check's `clear()` + `consume` re-measures the now-shrunken prompt. This is distinct from the hard `budget` ceiling (the bounded-paced-capped clause), which aborts the run with a partial. **Production hardening:** (a) non-fatal summarizer failure — the automatic `compact()` is wrapped in try/catch: a thrown summarizer error does not crash the run; the loop skips compaction that turn, surfaces the error as a `fault` event (so it is observable, never silently lost), and continues (the over-window prompt proceeds to the provider). Only the auto path is resilient — a manual `conversation.compact()` still propagates its error. (b) futile-compaction guard (the single-level limit) — when `compact()` resolves `undefined` (nothing left to fold) while the prompt is still over the window (the section summaries alone exceed it), a per-run flag latches so auto-compaction stops for the rest of that run (no per-turn churn); the over-window prompt then proceeds to the provider, which surfaces a genuine context-length error if it truly cannot fit (the real limit) — the loop does not loop futilely. The whole path is opt-in: with no `window` budget, or a non-summarizable active conversation, `#trim` (the run-entry reset, the pre-first-turn check, the between-turns check) is skipped entirely, adding no `await` before the first provider request, so a cost-budget-only agent's eager-pump and abort timing are untouched. Observability is the conversation's own `compact` / `summary` events (the conversation-layer clause) plus the agent's `fault` (the `strict` clause). These limits are deliberate: the automatic summarizer call is the conversation's configured one and is not separately bound to the run's abort signal; and compaction is single-level, so a conversation whose section summaries alone exceed `window` cannot shrink further — the futile guard stops the churn rather than pretending otherwise. The deterministic behavior (the absolute prompt crossing `max`, the fold, the rebuilt-smaller prompt, the pre-first-turn fold, the non-fatal `fault` path, the futile guard, no-fire below the ceiling) is pinned in the `src:core` mirror with a scripted provider, a data-stub (and a throwing) summarizer, and the real `estimateMessages` estimator, including a run forced through two or more folds that stays coherent.
|
|
999
|
+
|
|
1000
|
+
25. **`context.workspaces` — active workspace rendering by carrier.** `AgentContextInterface.workspaces` is the readonly `WorkspaceManagerInterface` supplied from `@orkestrel/workspace`, or a fresh dependency manager when omitted. `build()` reads only its active workspace, fresh each call, and filters that workspace's files through `scope.files`. The dependency's `isText` narrows text files, which render under `WORKSPACE_SECTION_HEADER` through the agent-owned `renderFencedFile` helper; the image carrier is `isBinary(file.content) && file.content.mime.startsWith('image/')`, and matching base64 data attaches to the last user message. With no active workspace, nothing renders. Both halves of that split are this package's own decision and belong here: `@orkestrel/workspace` holds files without knowing what a prompt is, and only the assembly layer knows that a model reads text as quoted material and images off a user turn. What agent does not own is the workspace domain itself — creation, editing, events, snapshots, and stores all stay in the originating package.
|
|
1001
|
+
|
|
1002
|
+
26. **The durable `ConversationStore` + the manager's `open` / `save` seam.** A `ConversationSnapshot` is the plain JSON payload `{ id, summary?, sections, messages }`; live summarizer/configuration functions are re-supplied when hydrating. `Conversation.snapshot()` creates it, and `ConversationInput.snapshot` restores identity, summary, sections, and live tail without emitting edit events. `isConversationSnapshot` is the total read-boundary guard for unknown storage values. `ConversationStoreInterface` persists that one payload through `get(id)`, `set(snapshot)`, and `delete(id)`, with no TTL. `MemoryConversationStore` keeps snapshots in a process-lifetime map; `DatabaseConversationStore` stores each snapshot in one opaque JSON column and narrows it on read. `ConversationManager.open(id)` activates a registry hit or hydrates a store hit; `save(id)` persists a registered snapshot and returns `false` when no store or conversation exists. The real stores, driver, guard, snapshot hydration, and manager semantics are pinned in the core tests.
|
|
1003
|
+
|
|
1004
|
+
27. **Limit exhaustion, mid-stream budget metering, per-run bounds, and per-run `schema`.** `RunOutcome.exhausted` (and the settled outcome's `partial`) flips `true` when the turn loop exhausts the effective `limit` while the most recently completed turn still held unresolved tool intent (the model requested tools on the very last allowed turn) — a cause distinct from a cancel: it fires the `exhaust` event (carrying the effective `limit`) instead of `abort`, still followed by `finish` carrying the partial result. A natural final answer on the last allowed turn, or `limit: 0` (which never enters the loop), stays `partial: false` with no `exhaust`. **Mid-stream budget charging.** During each provider turn, `#provide`'s `onDelta` re-estimates the turn's accumulated content through `estimateTokens` (`ceil(length / 4)`) and `budget.consume`s only the increment over what was already charged this turn (`{ prompt: 0, completion: increment, total: increment }`) — so the budget trip can land mid-stream, before the turn's final `usage` is known; the tripped signal folds into the run's bound abort exactly like any other cancel (the run resolves `partial: true` with an `abort` event, the provider genuinely cancelled). Once the turn's `usage` is known, a residual reconcile charges the remainder (`{ prompt: usage.prompt, completion: max(0, usage.completion - charged), total: max(0, usage.total - charged) }`), so the turn's total budget draw nets to exactly the authoritative usage — never double-charged, never lost. The reported `AgentResult.usage` / `usage` chunks are always the full authoritative usage, unaffected by how the budget was charged. **Per-run overrides (`AgentRunOptions`).** `limit` / `timeout` / `budget` / `signal` each override their `AgentOptions` construction default for this run only (`??` semantics — an omitted key keeps the constructed default); a per-run `signal` composes with (never replaces) a constructed `signal` through `AbortSignal.any` — either aborting cancels the run; a per-run `budget` is `start()`ed for that run and is the one the loop charges, leaving a constructed `budget` untouched for that run. **Per-run `schema`.** `AgentRunOptions.schema` (and `ProviderStreamOptions.schema`) mirrors `think` — a per-run structured-output constraint forwarded to `provider.stream`. `#provide` composes `think` and `schema` into one options object, omitting whichever key is `undefined`, and passes no options object at all when both are absent.
|
|
1005
|
+
|
|
1006
|
+
28. **`AgentOptions.instructions` / `.workspaces` / `.scope` — construction-time context wiring.** These mirror the identically-named `AgentContextOptions` fields (the richer-turn-context clause) and forward straight into the `AgentContext` the constructor builds: `instructions` a pre-built `InstructionManagerInterface` (an empty one created when omitted), `workspaces` a pre-built `WorkspaceManagerInterface` (a fresh empty one when omitted), and `scope` the initial active `ScopeInterface` (`undefined` ⇒ no filtering). They are construction sugar: the same result is reachable by building an `AgentContext` first and passing it in, and these fields spare that indirection when a caller only needs `createAgent`.
|
|
1007
|
+
29. **`strict` — automatic-compaction failure escalation.** `AgentOptions.strict` (default `false`) governs what happens when the automatic compaction path's `conversation.compact()` throws (the automatic-compaction clause): lenient (the default) surfaces the caught error as the `fault` event and continues over-window; `strict: true` still fires `fault` first (the failure stays observable either way), then rethrows the caught error so it propagates out of `#trim` through `#run`, rejecting the run's `result` with a genuine `error` settle (`status` → `error`) instead of a `partial: true` resolve. A manual `conversation.compact()` is unaffected by `strict` — it always propagates its own error regardless.
|
|
1008
|
+
30. **Bounded `sections` — cap the compacted history, `collapse` on overflow.** `ConversationOptions.sections` / `ConversationManagerOptions.sections` (a manager default, overridden per-`add` by `ConversationInput.sections`) / `CompactOptions.sections` (a per-compaction override) each set a cap (`>= 1`) on `Conversation.sections`'s length; omitted at every level ⇒ unlimited. A sub-1 cap — at construction (`ConversationOptions.sections`) or at a `compact()` call (the effective `options.sections ?? the conversation's own cap`) — throws a `ConversationError` with `code: 'SECTIONS'`. When a `compact()` fold pushes a new section past the effective cap, the oldest overflow sections are immediately folded into one merged section (a third summarizer call over the folded sections' summaries) so `sections.length` never exceeds the cap afterward, and a `collapse` event fires carrying the merged `Section`. This is orthogonal to `keep` (which bounds the live tail folded per compaction) — `sections` bounds the compacted history's length instead.
|
|
1009
|
+
31. **`sanitizeUsage` — normalizing a provider's abort-time partial usage.** A `ProviderAbortError.partial.usage` a provider surfaces mid-abort can be malformed (non-finite, negative, or fractional fields) since it was assembled from a cut-off stream rather than a clean settle. `sanitizeToken(value)` is the shared per-field primitive: it floors non-finite or non-positive values to `0` and positive fractional values to their integer part. `sanitizeUsage(usage)` applies it to `prompt` / `completion` / `total`, so a caller charging a token `Budget` never consumes a negative or fractional amount. The `Agent` loop applies it automatically to an abort's partial usage before folding it into the run's accounted usage / budget charge; both helpers are also exported standalone.
|
|
1010
|
+
32. **`AgentError` — the synchronous shared-accounting concurrency guard.** `Agent.stream()` throws `AgentError('CONCURRENCY', …)` synchronously — before any state mutation or emit — when a run is already in flight on the same agent (`this.#runs.size > 0`) and the agent carries shared per-agent accounting for the new run: either a construction-level `window` (a shared context budget, always shared since it has no per-run override) or a construction-level `budget` with no per-run `AgentRunOptions.budget` override (a shared cost budget). A concurrent run that supplies its own per-run `budget` override (with no `window` set) is still allowed — it charges a separate instance. `isAgentError` narrows a caught value; branch on `error.code` — `'CONCURRENCY'` here, `'REGISTRY'` for an `AgentRegistry` accessor whose name is absent from its pool. A sequential/awaited caller is never affected — this guards only genuinely concurrent `stream()` calls on one agent.
|
|
1011
|
+
|
|
1012
|
+
33. **`agentResultToJSON` — the canonical portable AgentResult projection.** The helper accepts `unknown` and never throws, including for throwing getters, hostile nested usage, and revoked proxies. It captures `content`, `thinking`, `usage`, and `partial` exactly once through Contract's sanctioned `attempt` boundary, so conforming accessors and inherited structural properties are supported without a second read. `content` must be a string and `partial` a boolean; absent/`undefined` `thinking` and `usage` are omitted, while present `thinking` must be a string and present `usage` an object whose `prompt` / `completion` / `total` fields are finite numbers. Finite negative and fractional counts are preserved rather than normalized through `sanitizeUsage`, because the authoritative `TokenUsage` contract is numeric. Extra input fields are dropped. The helper rebuilds a fresh exact plain `{ content, thinking?, usage?: { prompt, completion, total }, partial }` object, deep-gates it through Contract's `parseJSONValue`, and returns Contract's imported `JSONValue`; invalid input returns `undefined`.
|
|
1013
|
+
|
|
1014
|
+
34. **The provider engine and its seams (`AgentProvider`).** The instance owns a minted `id` (`crypto.randomUUID()`), taken in the constructor and returned by every call it serves, and `format` exposed exactly as supplied. The base owns, for each call: the deadline folded with the caller's signal (the bounding clause); `AgentProviderInput.fetch`, or the global transport bound to its global receiver; an awaited `headers` hook merged over a `Content-Type: application/json` default and raced against that combined bound, so an unresolved hook cannot outlive the call and its abort listener is released on every exit; one `POST` of `JSON.stringify(body(request))` to `url + (path ?? '')`; the response body decoded through `readChunks`; one `frame()` parser armed for that call; every framed record through `read`; the `finish(parser)` tail fed through `read` as well; the splitter flushed as a final content delta; and the outcome assembled by `buildProviderResult`. A subclass owns `name`, `frame`, `body`, `read`, and `finish`, and nothing else — an implementation that needs a second transport, a second deadline, or a second error taxonomy has left the seam rather than extended it. `split` (default `true`) arms one `ThinkSplitter` for the call; `strict` (default `false`) decides how a stream may end — with `strict: true` a stream that reaches end of input carrying no `ProviderIncrement.result` throws `ProviderError('PROTOCOL', 'provider error: missing settled result')`, and with `strict: false` the engine assembles the result from its accumulation instead. A record whose `read` returns a `result` is authoritative: the engine returns it at once without folding that record's other fields, leaves any later record undecoded, and cancels the body. **The bounded error read.** A non-OK response's body is decoded through `readText` to at most `MAX_ERROR_BODY_LENGTH` bytes and its remainder is cancelled, so a stalled error body is refused inside the deadline rather than after it. The bound counts bytes handed to the decoder. A source may deliver one chunk larger than the remaining budget; the read admits that chunk's leading bytes up to the budget and cancels the remainder, so the excerpt is decoded from at most `MAX_ERROR_BODY_LENGTH` source bytes and the read may have pulled one whole source chunk from the network. A multibyte character cut at that bound decodes to a replacement character, so the excerpt's own encoded length can exceed the bound by that character. **Reader-owned cancellation.** `readText` and `readChunks` each own their reader: each registers its abort listener on its own controller, cancels the source when the bound fires, and releases the lock in a `finally`, so a cancellation that itself fails never escapes and a body is never left locked. The parser clear and the deadline clear sit in nested `finally` blocks, so one failing cleanup never skips the other.
|
|
1015
|
+
|
|
1016
|
+
35. **The wire shapes and the compiled contracts.** `toolCallShape`, `messageShape`, `providerRequestShape`, `providerResultShape`, and `relayFrameShape` declare each domain type's JSON projection, and `messageContract`, `providerRequestContract`, `providerResultContract`, and `relayFrameContract` compile them into guards and parsers. Every projection is strictly narrower than the domain type it mirrors, and the narrowing is the point: `ToolCall.arguments` is `Record<string, unknown>` in the domain and a JSON record on the wire, so a function-valued argument is refused rather than dropped silently, and a tool's `parameters` and a call's `schema` are refused the same way. `ToolCall.caller` is local context and appears in no shape — the guard refuses a record carrying it and the parser drops it, so it cannot cross a hop. `relayFrameShape` is a channel-discriminated union that admits no member outside the arm it matched, so a decorative field on an `error` frame is a refusal rather than a tolerated extra. The round trip is closed in both directions: the request `RelayProvider.body` projects parses back to the value the relay hands upstream, and the frame `RelayStream` writes parses back to the delta or the result the browser decodes. **The wire body is an owned snapshot.** `body` clones the projection through Contract's `cloneJSONValue` and validates that clone, so the bytes `JSON.stringify` produces are exactly what the guard saw. The clone reads property descriptors rather than accessors, so a serializer reachable only through a `get` trap or a prototype is never consulted; an own function-valued property on a call's `arguments`, a tool's `parameters`, or the options `schema` — such as a `toJSON` method — is a value outside JSON and is refused before fetching, with the clone's own failure as the refusal's `cause` — the snapshot is the sole mechanism, and the refusal that remains also covers a hostile read and a snapshot the contract rejects. A caller mutating its own objects after the call cannot change what was sent.
|
|
1017
|
+
|
|
1018
|
+
36. **The relay protocol (`createRelay` / `RelayStream` / `RelayProvider`).** The request is a `POST` whose body is one `providerRequestContract` JSON document — `{ messages, tools?, options? }` and nothing beyond it. The response is newline-delimited `RelayFrame` records under `RELAY_CONTENT_TYPE` with `cache-control: no-store`, written under response backpressure: `RelayStream` awaits one `provider.stream` step per pull, so a slow reader parks the upstream turn rather than buffering it. `RelayProvider` constructs on `split: false` and `strict: true`, so the browser end preserves each delta verbatim and treats a stream that ended without a `result` frame as a protocol failure rather than a silently assembled answer. **The frame vocabulary.** `{ channel: 'content' | 'thinking', text }` carries each delta, `{ channel: 'result', result }` carries the settled turn, `{ channel: 'abort', partial }` carries an upstream cancel with its partial, and `{ channel: 'error', message }` carries every other failure — `message` is always `RELAY_PROVIDER_MESSAGE`, so an upstream failure's own text never reaches the browser. **`authorize` and its obligations.** It is mandatory, it runs before the body is read, and it must not consume that body: a body-reading hook locks the stream and the handler answers `400`. The mechanism performs no origin check and no method check, so an application authorizing on an ambient credential such as a cookie must compose origin and CSRF middleware in front of the handler; a bearer header the browser sets explicitly is not reachable cross-site and needs no such composition. **The refusals**, each carrying no body: `401` when `authorize` returns anything but `true` or throws; `413` when the request body reaches `limit` bytes (`DEFAULT_RELAY_LIMIT` when omitted) or the inbound read is aborted — at the limit, not merely above it, because a body that fills the budget without reporting end of input is indistinguishable from one that exceeds it; `400` when the body is missing, unreadable, or rejected by the contract; `502` when the upstream provider call cannot be constructed. Each refusal reaches the browser through the engine's HTTP path as a `ProviderError` with code `'HTTP'`, that `status`, and the message `provider error: <status>` — no excerpt, because a refusal carries no body. **Cancellation runs both ways.** An inbound abort — the client disconnecting, or the request's own signal firing — aborts the upstream controller, returns the generator, and releases the request listener even with a frame still queued unread; a browser-side cancel reaches the same path through its `fetch` signal, and the local cancel wins over any frame still in flight (the local-and-remote-cancel clause). What a disconnected client cannot recover is the turn itself: the upstream call is cancelled, whatever streamed but never reached the socket is gone, nothing is replayed, and no frame is held for a reconnect. A client that needs a resumable turn owns that above the relay. **The honest browser limit.** Host independence is proven here by the core scope's typecheck (no host global is in scope), by the transport defaulting to the global `fetch` bound to its global receiver, and by a recorded Chrome 148 run of the built core entry: that entry and its whole `@orkestrel` import closure load as ES modules with no console error, and a browser-side `RelayProvider` round-trips a turn, cancels mid-stream and keeps its local partial, and receives a refused bearer as a `ProviderError` with code `'HTTP'` and status `401`. It is not proven by a browser test project, because this package has none.
|
|
1019
|
+
|
|
1020
|
+
37. **`ProviderError` and its codes.** `ProviderError` carries a machine-readable `code` and a `status` that is present for an `HTTP` failure and `undefined` under every other code; `ProviderErrorOptions` carries that `status` and the underlying `cause`. `isProviderError` narrows a caught value so a caller can branch on the code. `'HTTP'` reports a non-OK response: the message is `provider error: <status>`, and ` - <excerpt>` is appended only when the bounded error read returned text, so an empty body leaves no separator and no trailing space; when that read fails outright the message is `provider error: <status> - (error body unavailable)` and the read's failure rides as the `cause`. `'PROTOCOL'` reports a successful response with no body, a record the concrete provider refuses (a frame outside `relayFrameContract`, a request the JSON wire cannot carry), or a `strict` stream that ended with no settled result. `'PROVIDER'` reports an upstream failure a relay carried in an `error` frame, and its message is `RELAY_PROVIDER_MESSAGE`. A cancel is in none of them: it is `ProviderAbortError` (the local-and-remote-cancel clause).
|
|
1021
|
+
|
|
1022
|
+
## Patterns
|
|
1023
|
+
|
|
1024
|
+
### Bounding any provider call
|
|
1025
|
+
|
|
1026
|
+
`ProviderInterface.generate` / `.stream` take a plain `AbortSignal`, so fold an [abort](abort.md), a [timeout](timeout.md), and a token [budget](budget.md) into one bound through `AbortSignal.any` — whichever trips first cancels the call. This works for any provider. A provider built on `AgentProvider` arms its own deadline as well (`AgentProviderInput.timeout`, `DEFAULT_PROVIDER_TIMEOUT` when omitted) and folds it with the signal you pass, so the bound here is the caller's ceiling and the engine's deadline is the backend's — pass a tighter signal to shorten a call, and construct the provider with a tighter `timeout` to shorten every call it makes.
|
|
1027
|
+
|
|
1028
|
+
```ts
|
|
1029
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1030
|
+
import { createAbort } from '@orkestrel/abort'
|
|
1031
|
+
import { createTimeout } from '@orkestrel/timeout'
|
|
1032
|
+
import { createTokenBudget } from '@orkestrel/budget'
|
|
1033
|
+
|
|
1034
|
+
declare const provider: ProviderInterface // any concrete implementation supplied by the host app
|
|
1035
|
+
const messages = [{ id: '1', role: 'user', content: 'Say hello.' }] as const
|
|
1036
|
+
const abort = createAbort() // external cancel
|
|
1037
|
+
const timeout = createTimeout({ ms: 30_000 }) // wall-clock deadline
|
|
1038
|
+
const budget = createTokenBudget({ max: 50_000, scope: 'total' }) // cost ceiling
|
|
1039
|
+
timeout.start()
|
|
1040
|
+
budget.start()
|
|
1041
|
+
|
|
1042
|
+
const bound = AbortSignal.any([abort.signal, timeout.signal, budget.signal])
|
|
1043
|
+
const result = await provider.generate(messages, bound)
|
|
1044
|
+
budget.consume(result.usage ?? { prompt: 0, completion: 0, total: 0 })
|
|
1045
|
+
```
|
|
1046
|
+
|
|
1047
|
+
### Writing a provider for a new wire
|
|
1048
|
+
|
|
1049
|
+
Reach a backend this package has never heard of by extending `AgentProvider` rather than implementing `ProviderInterface` from nothing. The base already holds the deadline, the transport, the `headers` hook, the request, the bounded error read, the decode loop, reasoning separation, and result assembly (the provider-engine clause), so a subclass is the wire and nothing else: `name` identifies the backend, `frame` returns fresh framing state for the call, `body` projects a `ProviderRequest` onto the vendor's request shape, `read` turns one framed record into a `ProviderIncrement`, and `finish` hands back whatever the parser still held at end of input. Pass the vendor's endpoint through `super`, and set `split` and `strict` to match what the wire actually delivers. The fence writes `AgentProvider<string>` because this wire frames nothing: each decoded chunk is one raw-text record, so `TRecord` is `string`. A wire that frames JSON objects extends bare `AgentProvider` and takes the default record type, and a wire with its own record type names that type instead.
|
|
1050
|
+
|
|
1051
|
+
```ts
|
|
1052
|
+
import type {
|
|
1053
|
+
ProviderIncrement,
|
|
1054
|
+
ProviderOptions,
|
|
1055
|
+
ProviderParserInterface,
|
|
1056
|
+
ProviderRequest,
|
|
1057
|
+
} from '@orkestrel/agent'
|
|
1058
|
+
import { AgentProvider } from '@orkestrel/agent'
|
|
1059
|
+
|
|
1060
|
+
class TextFrame implements ProviderParserInterface<string> {
|
|
1061
|
+
parse(chunk: string): readonly string[] {
|
|
1062
|
+
return [chunk]
|
|
1063
|
+
}
|
|
1064
|
+
clear(): void {} // Raw text retains no framing state.
|
|
1065
|
+
}
|
|
1066
|
+
|
|
1067
|
+
interface TextOptions extends ProviderOptions {
|
|
1068
|
+
readonly url: string
|
|
1069
|
+
}
|
|
1070
|
+
|
|
1071
|
+
class TextProvider extends AgentProvider<string> {
|
|
1072
|
+
readonly name = 'text'
|
|
1073
|
+
constructor(options: TextOptions) {
|
|
1074
|
+
super({ ...options, path: '/generate' })
|
|
1075
|
+
}
|
|
1076
|
+
frame(): ProviderParserInterface<string> {
|
|
1077
|
+
return new TextFrame()
|
|
1078
|
+
}
|
|
1079
|
+
body(request: ProviderRequest): object {
|
|
1080
|
+
return { messages: request.messages }
|
|
1081
|
+
}
|
|
1082
|
+
read(record: string): ProviderIncrement {
|
|
1083
|
+
return { content: record, thinking: '', tools: [] }
|
|
1084
|
+
}
|
|
1085
|
+
finish(_parser: ProviderParserInterface<string>): readonly string[] {
|
|
1086
|
+
return [] // Raw text retains no records at end of input.
|
|
1087
|
+
}
|
|
1088
|
+
}
|
|
1089
|
+
```
|
|
1090
|
+
|
|
1091
|
+
That provider drives the whole runtime unchanged — `createAgent`, the tool loop, the authority gate, and durable jobs all read it through `ProviderInterface`. Take the switches deliberately. Leave `split` at its default when the backend inlines its reasoning as `<think>` spans in the answer, and set `split: false` when the backend already separates reasoning onto its own field, because splitting twice re-classifies text the backend has already ruled on. Set `strict: true` when the protocol always ends with a settled record and a stream that stops short is a protocol failure worth reporting; leave it `false` when the answer is assembled from the deltas themselves.
|
|
1092
|
+
|
|
1093
|
+
### Relaying a browser provider through your own server
|
|
1094
|
+
|
|
1095
|
+
A browser must never hold a model credential. Mount `createRelay` on your own server, where the credential already lives, and give the browser `createRelayProvider` pointed at that route: the browser drives a `ProviderInterface` like any other, and the credential never leaves the server. The handler is a plain `(request: Request) => Promise<Response>`, so any fetch-standard router mounts it — this composition mounts it on an `@orkestrel/router` dispatcher. Each half that follows runs in its own process: copy the server half into your server and the browser half into your browser bundle.
|
|
1096
|
+
|
|
1097
|
+
The `@orkestrel/agent` package declares no dependency on a newline-delimited JSON parser, and no runtime dependency on a router or on a server adapter, so the browser application supplies the parser — `createNDJSONParser` from `@orkestrel/ndjson` here — and the server application supplies the router and the adapter. The `@orkestrel/router` and `@orkestrel/server` development dependencies this package declares serve the executed transcription of these fences in [`tests/guides.test.ts`](../tests/guides.test.ts). That transcription runs the server half and the browser half for real: it mounts the relay handler on the dispatcher route the server half declares, starts `@orkestrel/server` on a loopback listener, and drives the browser half against that listener over an HTTP hop in Node. The hop proves the round trip returning the upstream's settled result, the `405` with its `Allow: POST` header a `GET` to `/relay` answers, the `404` a `POST` to another path answers, the `401` a wrong bearer answers with the upstream provider unentered, and the upstream turn a disconnected reader cancels through the adapter's abort of the inbound request.
|
|
1098
|
+
|
|
1099
|
+
The transcription substitutes where a test process differs from a deployment. The fence's `declare` placeholders become a scripted upstream provider and a fictional bearer, and the `messages` binding is typed as `readonly Message[]` rather than narrowed with `as const`. It frames the relay response with a parser from this repository's own test infrastructure, which throws on a malformed line where the published parser skips it. It adds `host: '127.0.0.1'` to the `createServer` call, because a test listener binds loopback where a deployed server binds every interface. It points the browser half's `url` option at that address and the port `await server.start()` resolved, in place of the fence's fixed public URL. It closes each case with `await server.stop()` — the fence's own call — rather than the `SIGTERM` listener, which a test process outlives. The byte-limit refusal stays a direct handler call. The `limit` option of the `createRelay` factory caps the body the relay reads, and the `limit` option of the `createServer` factory caps the `body()` read a middleware context makes. This composition registers no middleware, so the server's cap never sees these bytes.
|
|
1100
|
+
|
|
1101
|
+
#### Mounting the relay on your server
|
|
1102
|
+
|
|
1103
|
+
```ts
|
|
1104
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1105
|
+
import { createRelay } from '@orkestrel/agent'
|
|
1106
|
+
import { createDispatcher } from '@orkestrel/router'
|
|
1107
|
+
import { createServer } from '@orkestrel/server'
|
|
1108
|
+
|
|
1109
|
+
declare const upstream: ProviderInterface // the server-side provider holding the credential
|
|
1110
|
+
declare const bearer: string
|
|
1111
|
+
|
|
1112
|
+
const handler = createRelay({
|
|
1113
|
+
provider: upstream,
|
|
1114
|
+
authorize: (request) => request.headers.get('authorization') === `Bearer ${bearer}`,
|
|
1115
|
+
})
|
|
1116
|
+
const dispatcher = createDispatcher({
|
|
1117
|
+
routes: [{ method: 'POST', path: '/relay', handler }],
|
|
1118
|
+
})
|
|
1119
|
+
|
|
1120
|
+
export function serve(request: Request): Promise<Response> {
|
|
1121
|
+
return dispatcher.handle(request, undefined)
|
|
1122
|
+
}
|
|
1123
|
+
|
|
1124
|
+
const server = createServer({ dispatcher, state: () => undefined })
|
|
1125
|
+
await server.start()
|
|
1126
|
+
process.on('SIGTERM', () => server.stop()) // signal cancellation, drain, then close the listener
|
|
1127
|
+
```
|
|
1128
|
+
|
|
1129
|
+
The `serve` function is the entry a server runtime's adapter calls, and the adapter owns what the relay deliberately does not: turning the runtime's inbound request into a `Request`, handing the returned `Response` back to the runtime, and aborting that request's signal when the client disconnects, so a reader that goes away cancels the upstream turn instead of leaving it running. `@orkestrel/server` supplies that adapter — `createServer` binds the same `dispatcher` and starts listening, as the fence shows — and a runtime whose own handler is already `(request: Request) => Promise<Response>` needs none.
|
|
1130
|
+
|
|
1131
|
+
#### Reaching the relay from the browser
|
|
1132
|
+
|
|
1133
|
+
```ts
|
|
1134
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1135
|
+
import { createRelayProvider } from '@orkestrel/agent'
|
|
1136
|
+
import { createAbort } from '@orkestrel/abort'
|
|
1137
|
+
// The browser application supplies this parser dependency.
|
|
1138
|
+
import { createNDJSONParser } from '@orkestrel/ndjson'
|
|
1139
|
+
|
|
1140
|
+
declare const bearer: string
|
|
1141
|
+
const abort = createAbort()
|
|
1142
|
+
const messages = [{ id: '1', role: 'user', content: 'Say hello.' }] as const
|
|
1143
|
+
|
|
1144
|
+
const browser: ProviderInterface = createRelayProvider({
|
|
1145
|
+
url: 'https://app.example/relay',
|
|
1146
|
+
parser: createNDJSONParser,
|
|
1147
|
+
headers: () => ({ authorization: `Bearer ${bearer}` }),
|
|
1148
|
+
})
|
|
1149
|
+
const result = await browser.generate(messages, abort.signal) // a ProviderResult like a local provider's
|
|
1150
|
+
```
|
|
1151
|
+
|
|
1152
|
+
`authorize` is mandatory and carries obligations the mechanism deliberately does not assume for you: it must not read the request body, and it performs no origin or method check of its own, so an application that authorizes on an ambient credential such as a cookie composes origin and CSRF middleware in front of the handler. A bearer header the browser sets explicitly, as here, is not reachable cross-site. Every refusal — `401` for a failed authorization, `413` for a body that reaches the byte limit, `400` for a missing, unreadable, or rejected body, `502` for an upstream call that cannot be constructed — arrives at the browser as a `ProviderError` with code `'HTTP'` and that status. The `401`, `400`, and `413` refusals leave the upstream provider unentered; the `502` is answered after `provider.stream` was entered and threw before returning its iterator, so that refusal has already reached the provider. An upstream failure that does happen mid-stream reaches the browser as the fixed `RELAY_PROVIDER_MESSAGE` text, never the upstream's own message (the relay-protocol clause).
|
|
1153
|
+
|
|
1154
|
+
### Dispatching the model's tool calls
|
|
1155
|
+
|
|
1156
|
+
Advertising and dispatch are the halves of one exchange: hand `definitions()` to the provider, and feed the `ToolCall`s that come back through `execute`. Results are correlated by `id` and discriminated on `success`, so a handler throw arrives as a `ToolResult` the model can read rather than an exception the caller must catch.
|
|
1157
|
+
|
|
1158
|
+
```ts
|
|
1159
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1160
|
+
import { createToolManager, createTool } from '@orkestrel/tool'
|
|
1161
|
+
|
|
1162
|
+
declare const provider: ProviderInterface
|
|
1163
|
+
const tools = createToolManager()
|
|
1164
|
+
tools.add(createTool({ name: 'add', execute: (args) => Number(args.a) + Number(args.b) }))
|
|
1165
|
+
|
|
1166
|
+
const turn = await provider.generate(messages, signal, tools.definitions())
|
|
1167
|
+
if (turn.tools) {
|
|
1168
|
+
const results = await tools.execute(turn.tools) // each correlated by id; one bad call never fails the batch
|
|
1169
|
+
// feed `results` back as the next turn's tool messages
|
|
1170
|
+
}
|
|
1171
|
+
```
|
|
1172
|
+
|
|
1173
|
+
### Running the loop (instead of driving the provider by hand)
|
|
1174
|
+
|
|
1175
|
+
The preceding patterns are what an `Agent` does for you turn after turn — bounding the call, dispatching the model's tools, feeding the results back, and repeating until the model stops (or `limit` is hit). Reach for `createAgent` rather than hand-rolling the loop; bound and pace it through `AgentOptions`, and recover a cancel's partial from `result` (which resolves, never rejects, on a cancel).
|
|
1176
|
+
|
|
1177
|
+
```ts
|
|
1178
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1179
|
+
import { createAgent } from '@orkestrel/agent'
|
|
1180
|
+
|
|
1181
|
+
declare const provider: ProviderInterface
|
|
1182
|
+
const agent = createAgent(provider, { timeout: 30_000, limit: 6 })
|
|
1183
|
+
agent.context.messages.add({ role: 'user', content: 'Summarize the news.' })
|
|
1184
|
+
|
|
1185
|
+
// Cancel from elsewhere — the turn commits whatever streamed so far.
|
|
1186
|
+
setTimeout(() => agent.abort('user navigated away'), 5_000)
|
|
1187
|
+
|
|
1188
|
+
const result = await agent.generate()
|
|
1189
|
+
if (result.partial) keep(result.content) // a cancel RESOLVED partial, not an error
|
|
1190
|
+
```
|
|
1191
|
+
|
|
1192
|
+
### Bounding cost mid-stream (the token `budget`)
|
|
1193
|
+
|
|
1194
|
+
An `AgentOptions.budget` (or a per-run override) is not only charged from each turn's final reported `usage` — the loop also charges it incrementally, mid-stream, from an estimated token count as content deltas arrive (the `estimateTokens` `ceil(length / 4)` heuristic), so a runaway completion trips the ceiling without waiting for the turn to finish. When the mid-stream estimate crosses the budget, the budget's `signal` fires, folding into the run's bound abort exactly like an external cancel or a `timeout` — the provider is cancelled, and the run resolves `partial: true` with an `abort` event (the same funnel as any other cancel).
|
|
1195
|
+
|
|
1196
|
+
Once a turn does complete, the loop reconciles: it charges the budget the remainder of the turn's authoritative `usage` (`completion - alreadyCharged`, `total - alreadyCharged`, plus the full `prompt` — which is never estimated mid-stream, having no live delta channel) — so the turn's total budget draw always nets to exactly the reported usage, never double-charged and never under-charged. The `AgentResult.usage` / the `usage` chunks you observe stay the full authoritative usage regardless — this reconcile affects only what the `budget` itself was charged, never what you're told the turn cost. This mid-stream enforcement is bounded, not exact — the estimate can under- or over-shoot the eventual real usage by a turn's tail, so treat the `budget.max` as a firm ceiling with some slack, not a byte-exact cutoff.
|
|
1197
|
+
|
|
1198
|
+
```ts
|
|
1199
|
+
import { createAgent } from '@orkestrel/agent'
|
|
1200
|
+
import { createTokenBudget } from '@orkestrel/budget'
|
|
1201
|
+
|
|
1202
|
+
const budget = createTokenBudget({ max: 2_000, scope: 'completion' })
|
|
1203
|
+
const agent = createAgent(provider, { budget })
|
|
1204
|
+
agent.emitter.on('abort', (reason) => log('budget tripped mid-stream', reason))
|
|
1205
|
+
agent.context.messages.add({ role: 'user', content: 'Write a very long story.' })
|
|
1206
|
+
const result = await agent.generate() // partial: true if the story ran the budget out mid-stream
|
|
1207
|
+
```
|
|
1208
|
+
|
|
1209
|
+
**A `think: true` run needs headroom for reasoning.** Live reasoning deltas (`ProviderDelta` `'thinking'`) are not metered mid-stream (only `'content'` deltas are — thinking is charged, like content, solely through the post-turn reconcile), so a thinking model can spend a large share of a tight budget's ceiling on its reasoning before any answer content streams — the mid-stream trip can land while the model is still reasoning, committing an empty (or near-empty) `content` alongside `partial: true`. Give a `think: true` run enough budget headroom to cover its reasoning, not only its expected answer length.
|
|
1210
|
+
|
|
1211
|
+
### Observing an agent (push vs. pull)
|
|
1212
|
+
|
|
1213
|
+
An `Agent` exposes a pull and a push observation surface. Pull — the `AgentChunk` stream (`stream().events`) — is for a live consumer rendering per-token answer deltas and per-think reasoning deltas as they arrive. Push — the `emitter` (`AgentEventMap`) — is for fire-and-forget observers (logging, metrics, tracing) that want the loop's lifecycle moments without draining the stream: `start` (a run begins), `turn` (each iteration), `tool` (a dispatched call + its result), `usage` (a turn's token usage), `deny` (an authority denial — which never reaches the chunk stream), `finish` (the settled result), `error` (a genuine failure), `abort` (a cancel), and `exhaust` (the limit was reached while the model still held unresolved tool intent — fires instead of `abort`, still followed by `finish`). Per-token / per-thinking deltas stay the stream's job exclusively — there is deliberately no `token` or `think` event on the emitter; reach for the stream when you need live output, the emitter when you need lifecycle.
|
|
1214
|
+
|
|
1215
|
+
```ts
|
|
1216
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1217
|
+
import { createAgent } from '@orkestrel/agent'
|
|
1218
|
+
|
|
1219
|
+
declare const provider: ProviderInterface
|
|
1220
|
+
// Wire fire-and-forget observers at construction through the reserved `on` option …
|
|
1221
|
+
const agent = createAgent(provider, {
|
|
1222
|
+
on: {
|
|
1223
|
+
start: (id) => trace.begin(id),
|
|
1224
|
+
usage: (usage) => meter.add(usage.total),
|
|
1225
|
+
deny: (call, reason) => audit(call.name, reason), // not visible on the chunk stream
|
|
1226
|
+
finish: (result) => trace.end(result),
|
|
1227
|
+
},
|
|
1228
|
+
})
|
|
1229
|
+
// … or subscribe later through `agent.emitter`.
|
|
1230
|
+
agent.emitter.on('abort', (reason) => log('cancelled', reason))
|
|
1231
|
+
```
|
|
1232
|
+
|
|
1233
|
+
**Observation can never corrupt the loop.** The emitter isolates a listener that throws (the throw can never escape into the settle-once / wake-park engine — `generate()` / `stream()` still settle the exact same result), routing the caught error to its own `error` handler (the `error` option, surfaced as `(error, event)`, not a domain event) so the observer bug is not silently lost. Every throwing listener surfaces (not only the first); a throwing `error` handler is swallowed too (it can neither recurse nor escape); with no handler the throw is dropped silently. So a buggy observer degrades to a routed error — it never reorders, throws into, or corrupts the run.
|
|
1234
|
+
|
|
1235
|
+
**A cancelled run emits `abort` then `finish`.** A cancel (an external `signal`, the `timeout` deadline, an exhausted `budget`, or `abort()`) still resolves a partial result — so the emitter fires `abort` (carrying the cancel reason) and then `finish` (carrying the settled partial), letting an observer see both that the run was cancelled and the partial outcome it committed. A natural / cap-bounded finish fires `finish` only; a genuine provider / tool error fires `error` instead of `finish`. `generate()` and `stream()` drive the same events (they share one `#run`).
|
|
1236
|
+
|
|
1237
|
+
### Pulling context from another conversation (with provenance)
|
|
1238
|
+
|
|
1239
|
+
One agent serves many conversations by switching the active conversation between runs (`conversations.switch(id)` — the active-conversation clause); each thread keeps its own history. When the active conversation A needs something decided in another conversation B, don'T merge B's turns into A's live tail — that pollutes A's thread and (for a small model) blurs which conversation said what. Instead pull a provenance-labeled reference of B into A's active workspace (the fenced reference channel `build()` folds into the system block), so the model reads it as clearly-foreign material and attributes it to B.
|
|
1240
|
+
|
|
1241
|
+
The flow is **summary → search / rehydrate → reference → write-to-workspace** — and cherry-pick, never dump:
|
|
1242
|
+
|
|
1243
|
+
```ts
|
|
1244
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1245
|
+
import { createAgent, createConversationManager } from '@orkestrel/agent'
|
|
1246
|
+
|
|
1247
|
+
declare const provider: ProviderInterface
|
|
1248
|
+
const conversations = createConversationManager({
|
|
1249
|
+
summarize: /* a ConversationSummaryHandler */ undefined,
|
|
1250
|
+
})
|
|
1251
|
+
const a = conversations.add({ id: 'auth' }) // the ACTIVE thread (the first add auto-activates it)
|
|
1252
|
+
const b = conversations.add({ id: 'planning' }) // the OTHER thread to pull from
|
|
1253
|
+
|
|
1254
|
+
const agent = createAgent(provider, { conversations }) // its active conversation (a) is the message source
|
|
1255
|
+
|
|
1256
|
+
// 1. summary → decide B is relevant (its rollup is a cheap digest of the whole thread)
|
|
1257
|
+
b.summary // for example "the team evaluated databases and chose Postgres"
|
|
1258
|
+
// 2. search / rehydrate → SELECT the few right turns (never B's whole history)
|
|
1259
|
+
const picked = b.search('database') // or b.rehydrate(sectionId) for a compacted slice
|
|
1260
|
+
// 3. reference → FRAME them as a self-labeled provenance block (a pure string, no model call)
|
|
1261
|
+
const block = b.reference({ label: 'planning', messages: picked })
|
|
1262
|
+
// 4. write-to-workspace → into the ACTIVE conversation's context active workspace, keyed by the source id
|
|
1263
|
+
agent.context.workspaces.add().write(`conversation:${b.id}.md`, block)
|
|
1264
|
+
|
|
1265
|
+
// Now the model can use B's decision AND attribute it: "Postgres, decided in the planning conversation."
|
|
1266
|
+
```
|
|
1267
|
+
|
|
1268
|
+
Why this shape: `reference()` leads with `[Reference — conversation "<label>" — NOT part of this conversation]`, so the model treats the rollup + excerpts as a quoted foreign source (it answers "decided in the planning conversation", not "we decided here"). Keep the excerpts cherry-picked — this content enters another context window a small model must read, and a full dump re-bloats it. `reference()` is event-free and never calls a model; provenance lives in the `label` (default the conversation's `id`).
|
|
1269
|
+
|
|
1270
|
+
**Within a conversation, the same provenance instinct applies to recaps.** `view()` folds each compacted section into a synthetic `assistant` recap — and prefixes it with `CONVERSATION_RECAP_PREFIX` (`[Summary of earlier messages] …`) so a small model reads it as a condensed recap of earlier turns rather than a literal turn to echo or treat as the live answer. The label is deliberately lean (a fixed handful of tokens — no per-section blow-up) and is a `view()`-only presentation concern (the rollup regeneration re-reads the unframed summaries). Empirically, on a 2B model this tightening is the difference between the model correctly attributing a recapped fact and mis-attributing it — at temperature 0 the recap label reliably steers correct attribution where the bare assistant turn does not.
|
|
1271
|
+
|
|
1272
|
+
### Persisting a conversation through either store
|
|
1273
|
+
|
|
1274
|
+
A conversation snapshot carries its generated identity, so the same `set` / `get` / `delete` workflow works through the in-memory store and the driver-backed store without hard-coding an id:
|
|
1275
|
+
|
|
1276
|
+
```ts
|
|
1277
|
+
import {
|
|
1278
|
+
createConversation,
|
|
1279
|
+
createDatabaseConversationStore,
|
|
1280
|
+
createMemoryConversationStore,
|
|
1281
|
+
} from '@orkestrel/agent'
|
|
1282
|
+
import { createMemoryDriver } from '@orkestrel/database'
|
|
1283
|
+
|
|
1284
|
+
const conversation = createConversation()
|
|
1285
|
+
conversation.add({ role: 'user', content: 'hello' })
|
|
1286
|
+
const snapshot = conversation.snapshot()
|
|
1287
|
+
const stores = [
|
|
1288
|
+
createMemoryConversationStore(),
|
|
1289
|
+
createDatabaseConversationStore(createMemoryDriver()),
|
|
1290
|
+
]
|
|
1291
|
+
|
|
1292
|
+
for (const store of stores) {
|
|
1293
|
+
await store.set(snapshot)
|
|
1294
|
+
const stored = await store.get(conversation.id)
|
|
1295
|
+
JSON.stringify(stored) === JSON.stringify(snapshot) // true
|
|
1296
|
+
await store.delete(conversation.id)
|
|
1297
|
+
await store.get(conversation.id) // undefined
|
|
1298
|
+
}
|
|
1299
|
+
```
|
|
1300
|
+
|
|
1301
|
+
### Running many durable agents as jobs
|
|
1302
|
+
|
|
1303
|
+
When you need many agents — bounded, retried, surviving a crash — describe each as a serializable `AgentJobInput` (names for the live pieces, data for the rest), register the live pieces once, and run them through a `createAgentQueue` (durable, bounded) or a `createAgentRunner` (one-shot, ordered, fail-fast, with sub-agent fan-out). The layer composes the `@orkestrel/queue` `Queue` and the `@orkestrel/workflow` `Runner` — it adds only rehydration and the partial policy, no new engine.
|
|
1304
|
+
|
|
1305
|
+
```ts
|
|
1306
|
+
import type { AgentJobInput } from '@orkestrel/agent'
|
|
1307
|
+
import { createAgentQueue, createAgentRegistry } from '@orkestrel/agent'
|
|
1308
|
+
import { createMemoryQueueStore } from '@orkestrel/queue'
|
|
1309
|
+
|
|
1310
|
+
declare const store: ReturnType<typeof createMemoryQueueStore> // or a server JSON / SQLite store
|
|
1311
|
+
|
|
1312
|
+
// Register the live, non-serializable pieces ONCE; jobs reference them by name.
|
|
1313
|
+
const registry = createAgentRegistry({ providers: { main: provider } })
|
|
1314
|
+
const queue = createAgentQueue({ registry, concurrency: 4, retries: 1, store })
|
|
1315
|
+
|
|
1316
|
+
const jobs: readonly AgentJobInput[] = [
|
|
1317
|
+
{ provider: 'main', messages: [{ role: 'user', content: 'Summarize doc A.' }] },
|
|
1318
|
+
{ provider: 'main', messages: [{ role: 'user', content: 'Summarize doc B.' }], budget: 50_000 },
|
|
1319
|
+
]
|
|
1320
|
+
const results = await Promise.all(jobs.map((job) => queue.enqueue(job)))
|
|
1321
|
+
|
|
1322
|
+
// After a crash, re-run whatever was still outstanding — the registry rehydrates them.
|
|
1323
|
+
await queue.restore()
|
|
1324
|
+
```
|
|
1325
|
+
|
|
1326
|
+
Fan out sub-agents by declaring `children` on a parent job; `createAgentRunner` `controller.spawn`s each through the same bounded queue, so the children run as sibling sub-agents:
|
|
1327
|
+
|
|
1328
|
+
```ts
|
|
1329
|
+
import { createAgentRunner } from '@orkestrel/agent'
|
|
1330
|
+
|
|
1331
|
+
const runner = createAgentRunner({ registry, concurrency: 4 })
|
|
1332
|
+
const parent: AgentJobInput = {
|
|
1333
|
+
provider: 'main',
|
|
1334
|
+
messages: [{ role: 'user', content: 'Plan the trip.' }],
|
|
1335
|
+
children: [{ provider: 'main', messages: [{ role: 'user', content: 'Find flights.' }] }],
|
|
1336
|
+
}
|
|
1337
|
+
const results = await runner.execute([parent]) // [parent result, …then spawned child results]
|
|
1338
|
+
```
|
|
1339
|
+
|
|
1340
|
+
### Giving the model documents to read
|
|
1341
|
+
|
|
1342
|
+
A workspace reaches the model through `context.workspaces`, and this is the only channel documents have. `build()` renders the active workspace by carrier on every turn — active-only and scope-filtered — so the workspace the agent is working in is always what the prompt reflects, with nothing to re-mount after an edit. Register one (the first `add` auto-activates it) and its text files fold into the `## Workspace` system section as fenced reference blocks, while its image files' base64 rides the last user message.
|
|
1343
|
+
|
|
1344
|
+
```ts
|
|
1345
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1346
|
+
import { createAgent } from '@orkestrel/agent'
|
|
1347
|
+
import { createToolManager } from '@orkestrel/tool'
|
|
1348
|
+
|
|
1349
|
+
declare const provider: ProviderInterface
|
|
1350
|
+
const agent = createAgent(provider, { tools: createToolManager() })
|
|
1351
|
+
|
|
1352
|
+
// The first add auto-activates — agent.context.workspaces.active is this workspace, so build()
|
|
1353
|
+
// renders its text files into the `## Workspace` section on every turn.
|
|
1354
|
+
const workspace = agent.context.workspaces.add()
|
|
1355
|
+
workspace.write('briefing.txt', 'The vault code is 7731.')
|
|
1356
|
+
```
|
|
1357
|
+
|
|
1358
|
+
Reading is one half. To let the model edit what it reads, register the `createWorkspaceTool` published by `@orkestrel/toolbox` on `agent.context.tools` over this same `context.workspaces` registry: an `operation`-keyed `ToolInterface` whose dispatch and error semantics are that package's to document. The surfaces then close a loop — the model reads the workspace from the prompt, edits it through a tool call, and reads the edited version on the next turn.
|
|
1359
|
+
|
|
1360
|
+
### Switching which workspace the model sees
|
|
1361
|
+
|
|
1362
|
+
Only the active workspace renders; the other registered workspaces never reach the model at all. `switch` changes which one the model sees between runs:
|
|
1363
|
+
|
|
1364
|
+
```ts
|
|
1365
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1366
|
+
import { createAgent, createScope } from '@orkestrel/agent'
|
|
1367
|
+
|
|
1368
|
+
declare const provider: ProviderInterface
|
|
1369
|
+
const agent = createAgent(provider)
|
|
1370
|
+
|
|
1371
|
+
const project = agent.context.workspaces.add() // auto-activates
|
|
1372
|
+
project.write('src/config.ts', 'export const PORT = 8123')
|
|
1373
|
+
|
|
1374
|
+
agent.context.messages.add({ role: 'user', content: 'What port is configured?' })
|
|
1375
|
+
await agent.generate() // the model reads the active workspace's file from the prompt
|
|
1376
|
+
|
|
1377
|
+
// Serve a different workspace next run — switch the active pointer:
|
|
1378
|
+
const other = agent.context.workspaces.add() // NOT active (a later add leaves active unchanged)
|
|
1379
|
+
other.write('notes.txt', 'different context')
|
|
1380
|
+
agent.context.workspaces.switch(other.id) // now build() renders other's files instead
|
|
1381
|
+
// Narrow which files render with scope.files (by path):
|
|
1382
|
+
agent.context.apply(createScope({ name: 'cfg', files: ['src/config.ts'] }))
|
|
1383
|
+
```
|
|
1384
|
+
|
|
1385
|
+
Because the editing tool drives the same registry, a switch moves both surfaces together: the model's next prompt and its next edit land in the same workspace whether the host called `context.workspaces.switch` or the model asked the tool to.
|
|
1386
|
+
|
|
1387
|
+
### Removing / clearing entries, and the less-common accessors
|
|
1388
|
+
|
|
1389
|
+
Agent-owned registries expose their less-common removal, clearing, persistence, and lookup methods here. Tool and workspace registry operations are documented in their dependency guides.
|
|
1390
|
+
|
|
1391
|
+
```ts
|
|
1392
|
+
import type { ProviderInterface } from '@orkestrel/agent'
|
|
1393
|
+
import {
|
|
1394
|
+
createAgent,
|
|
1395
|
+
createAgentContext,
|
|
1396
|
+
createAgentRegistry,
|
|
1397
|
+
createAuthority,
|
|
1398
|
+
createConversationManager,
|
|
1399
|
+
createInstructionManager,
|
|
1400
|
+
createScopeManager,
|
|
1401
|
+
createThinkSplitter,
|
|
1402
|
+
} from '@orkestrel/agent'
|
|
1403
|
+
|
|
1404
|
+
declare const provider: ProviderInterface
|
|
1405
|
+
|
|
1406
|
+
// ThinkSplitter — one per stream; `split` yields clean content, `flush` settles the end.
|
|
1407
|
+
const splitter = createThinkSplitter()
|
|
1408
|
+
splitter.split('hello') // clean content for this raw wire delta
|
|
1409
|
+
splitter.flush() // any held partial tag / unclosed span resolved at stream end
|
|
1410
|
+
|
|
1411
|
+
// The message store (context.messages, a MessageManagerInterface) — same remove / clear.
|
|
1412
|
+
const context = createAgentContext()
|
|
1413
|
+
const added = context.messages.add({ role: 'user', content: 'hi' })
|
|
1414
|
+
context.messages.remove(added.id)
|
|
1415
|
+
context.messages.clear()
|
|
1416
|
+
|
|
1417
|
+
// InstructionManager — same remove / clear.
|
|
1418
|
+
context.instructions.remove('tone')
|
|
1419
|
+
context.instructions.clear()
|
|
1420
|
+
|
|
1421
|
+
// ScopeManager — `create` mints + stores, `scopes` lists, `remove` / `clear` drop.
|
|
1422
|
+
const scopes = createScopeManager()
|
|
1423
|
+
const scope = scopes.create({ name: 'read-only' })
|
|
1424
|
+
scopes.scopes() // every stored scope, in insertion order
|
|
1425
|
+
scopes.remove(scope.id)
|
|
1426
|
+
scopes.clear()
|
|
1427
|
+
|
|
1428
|
+
// Authority — `evaluate` is its one method (also reached through the agent loop internally).
|
|
1429
|
+
const authority = createAuthority()
|
|
1430
|
+
authority.evaluate({ call: { id: '1', name: 'add', arguments: {} } })
|
|
1431
|
+
|
|
1432
|
+
// AgentRegistry — `scheduler(name)` resolves a registered scheduler by name (throws when absent).
|
|
1433
|
+
const registry = createAgentRegistry({ providers: { main: provider } })
|
|
1434
|
+
try {
|
|
1435
|
+
registry.scheduler('paced') // throws 'unknown scheduler: paced' — none registered here
|
|
1436
|
+
} catch {
|
|
1437
|
+
// expected — this registry has no `schedulers` pool
|
|
1438
|
+
}
|
|
1439
|
+
|
|
1440
|
+
// ConversationManager — `save` persists a registered conversation, `remove` / `clear` drop it.
|
|
1441
|
+
const conversations = createConversationManager({ summarize: undefined })
|
|
1442
|
+
const thread = conversations.add({ id: 'thread-1' })
|
|
1443
|
+
await conversations.save(thread.id)
|
|
1444
|
+
conversations.remove(thread.id)
|
|
1445
|
+
conversations.clear()
|
|
1446
|
+
|
|
1447
|
+
// Conversation — `remove` one live message, `clear` the live tail, `snapshot` for durability.
|
|
1448
|
+
const message = thread.add({ role: 'user', content: 'hi' })
|
|
1449
|
+
thread.remove(message.id)
|
|
1450
|
+
thread.clear()
|
|
1451
|
+
thread.snapshot() // { id, summary?, sections, messages } — the durable payload
|
|
1452
|
+
|
|
1453
|
+
const agent = createAgent(provider)
|
|
1454
|
+
void agent
|
|
1455
|
+
```
|
|
1456
|
+
|
|
1457
|
+
### Practices
|
|
1458
|
+
|
|
1459
|
+
- **Bound every call** — pass an `AbortSignal` (an [abort](abort.md), or an `AbortSignal.any` over abort + [timeout](timeout.md) + [budget](budget.md)) so a request can be cancelled, deadlined, or capped.
|
|
1460
|
+
- **Recover the stream's partial** — wrap a driven `stream` in `try`/`catch` and narrow with `isProviderAbortError` to keep the content that arrived before a cancel.
|
|
1461
|
+
- **Fold usage into a budget** — `result.usage` is the [budgets](budget.md) `TokenUsage`; `consume` it per turn to enforce a token ceiling.
|
|
1462
|
+
- **Register tools in a `ToolManager`** — `add` your `Tool`s, hand `definitions()` to the provider, and dispatch the model's `ToolCall`s through `execute`; read the outcome by narrowing on `success` rather than catching, since a handler's throw already arrives as the failure arm. When in-process code wants a typed error instead, call `tools.tool(name)` and `execute` it directly.
|
|
1463
|
+
- **Narrow tool `args`** — a `ToolCall.arguments` is model-supplied `unknown`; narrow it inside `execute` with a guard.
|
|
1464
|
+
- **Collect turns in an `AgentContext`** — `add` to `context.messages` (the `id` is minted for you), then `build()` the provider input each turn; tools travel as the provider's `tools` argument, so never fold a tool's schema into a message yourself.
|
|
1465
|
+
- **Pull cross-conversation context with provenance, never by merging turns** — to use something from another conversation B in the active one, follow `B.summary` (decide relevance) → `B.search` / `B.rehydrate` (cherry-pick the few right turns) → `B.reference({ label, messages })` (frame it) → `context.workspaces.active?.write(...)` (write it into the active workspace). The provenance label keeps a small model from reading B's content (or a recap) as part of the live thread; cherry-pick — don't dump B's whole history into another context window.
|
|
1466
|
+
- **Run the loop with `createAgent`** — don't hand-roll the context → provider → tools cycle; `createAgent` does it, bounded by `AbortSignal.any([signal, timeout, budget])`, paced by `scheduler`, capped at `limit`. Drain `stream().events` to render `token` / `think` / `tool` / `usage` chunks live, or `generate()` for the settled result; either way `result` resolves partial on a cancel (read `result.partial`), rejecting only on a real error.
|
|
1467
|
+
- **Run many durable agents with `createAgentQueue`** — describe each agent as a serializable `AgentJobInput` (names for the provider / tools / authority / scheduler, data for the rest), register the live pieces once in a `createAgentRegistry`, and enqueue the jobs. Persist them with a `store` so `restore()` re-runs outstanding work after a crash. Decide the partial policy up front — `partial: false` (default) retries a cancelled job, `true` accepts the partial.
|
|
1468
|
+
- **Fan out sub-agents with `createAgentRunner`** — declare a parent job's sub-agents in its `children`; the runner `controller.spawn`s each through the same bounded queue. Don't reach for the controller yourself — express fan-out as data on the job (it stays serializable) and let the runner spawn it.
|
|
1469
|
+
- **Pull and push surfaces on the `Agent`, none elsewhere** — observe the `Agent` each way: pull the `AgentChunk` stream for per-token / per-thinking deltas + usage/tool chunks, or push `agent.emitter.on(...)` (`AgentEventMap`) for lifecycle + usage/tool/deny moments a fire-and-forget observer wants. A listener throw can never corrupt the loop (the emitter isolates it, routing it to the `error` option). Do not reach for an Emitter on the provider contract, the tool registry, the conversation store, the context, or the job layer — those stay event-free; and do not expect per-token `token` or `think` events (those stay the stream's job).
|
|
1470
|
+
- **Create and edit files through `@orkestrel/workspace`** — the file domain is that package's, and `AgentContext` borrows only its `isText` / `isBinary` guards to decide a file's carrier. The rendering is agent's: the `## Workspace` fencing and the binary-plus-`image/` attachment are prompt policy that belongs here, not workspace helpers that belong there.
|
|
1471
|
+
- **Give the model both halves of a workspace** — `context.workspaces` is what it reads, rendered by carrier every turn, active-only and filtered by `scope.files`; the `createWorkspaceTool` published by `@orkestrel/toolbox` over that same registry is what it writes. Workspace editing, errors, and persistence live in [`workspace.md`](workspace.md); tool dispatch semantics live in [`tool.md`](tool.md).
|
|
1472
|
+
|
|
1473
|
+
## Tests
|
|
1474
|
+
|
|
1475
|
+
- [`tests/guides.test.ts`](../tests/guides.test.ts) — the `## Surface` ↔ `src/core` bijection for agent-owned values and types, exhaustive method parity for every agent-owned interface and class pair documented here, and the equality gate: every `Summary` cell against its declaration's description paragraph, the titled `Conversations & compaction` fence against the `@example` block of that title (pinned so the titled pair cannot be retired silently), and the README pitch against this guide's tagline. It also runs the flagship fences and asserts the values their comments claim, including the relay halves driven over a started `@orkestrel/server` listener on loopback — the declared route's own `405` and `404`, the `401` a wrong bearer answers, the round trip, and the upstream turn a disconnected reader cancels.
|
|
1476
|
+
- [`tests/src/core/conversations/Conversation.test.ts`](../tests/src/core/conversations/Conversation.test.ts) — `add` single mints a fresh `id` + carries `role` / `content` (and `calls` only when given, the field omitted otherwise); `add` batch returns the created messages in order with unique ids; the returned message reflects its input and is the same object `message(id)` resolves (immutable); `message` lookup / miss; `messages()` insertion order; `remove` single + batch (`true` only when every supplied id was removed) + `clear` + `count`; and hydration through the `ConversationOptions.snapshot` seam — the one way to restore (a `createConversation` restore of the stored id, rollup, sections, and live tail that re-snapshots identically, the snapshot id winning over an options id, a restored conversation's `view` / `search` / `count` and continued compaction, and no event emitted while restoring).
|
|
1477
|
+
- [`tests/src/core/AgentContext.test.ts`](../tests/src/core/AgentContext.test.ts) — `build()` with a system prompt prepends `{ role: 'system', content }` then the conversation in order; without one (and empty managers) returns only the conversation (no system turn, `system === undefined`, empty → `[]`, and an explicit `''` / whitespace system still prepended); `context.tools` is the passed registry (or a fresh empty `ToolManager`) and `build()` never includes a tool's name / description; `build()` is fresh each call (reflects messages added between builds, mints a new system `id`, snapshot-independent). And the richer assembly: the `instructions` manager is fresh empty (or reused by identity); an empty manager contributes nothing (lean behavior preserved); it folds into the system block under its `description` + per-item `override`; the readonly `scope` getter (default `undefined`, initial through options) + `apply(scope)` / `apply(undefined)` through an `AgentContextInterface` binding; and scope filtering per category — `undefined` ⇒ all, a named allow-list ⇒ only-listed, `[]` ⇒ none over the instructions, with the conversation passing through unfiltered, recomputed fresh when the scope is changed between builds. And the active-workspace render by carrier (the only document/image channel): `context.workspaces` is always present (fresh empty `WorkspaceManager`) or supplied structurally through options; the active workspace's text files fold into a `## Workspace` system section (fenced through `renderFencedFile`, placed after instructions) and its image files' base64 attaches to the last user message (own-images-first merge; multiple in insertion order; skipped when no user message; the stored message is never mutated); `scope.files` filters both carriers; the render is active-only (a non-active workspace's files never render, reflected through a `switch`); no active workspace ⇒ nothing rendered. And the injected-conversation message source (a data-stub summarizer): with one, `context.messages` is its live tail and `build()` folds its `view()` (the compacted view after a `compact()`), no scope filtering over messages (the conversation is authoritative) while scope still filters instructions; without one, the plain message-store path. And the structural conversation registry (the multi-conversation switch): supplied through options, its active conversation changes through `conversations.switch(id)`, re-pointing `messages` to the new live tail by identity (no duplication), and `build()` follows the switch; constructing with an empty supplied registry adds a default active conversation; a `compact()` on the active conversation is reflected through the switch.
|
|
1478
|
+
- [`tests/src/core/scopes/Scope.test.ts`](../tests/src/core/scopes/Scope.test.ts) — `Scope` construction (a minted `id` + the per-category lists including `files`, copied in so a later mutation of the caller array can't leak in; `[]` distinct from `undefined`); and `narrow`'s set-intersection semantics — `list ∩ list` = the keys in both, `undefined ∩ list` = the list (no parent constraint), `list ∩ undefined` = the list, `undefined ∩ undefined` = `undefined`, `[]` ⇒ none either side, narrowing only tightens (a parent-excluded key never returns), per-category independence (incl. `files`), name preserved, immutable (a new scope, parent untouched), and chained narrows compose.
|
|
1479
|
+
- [`tests/src/core/scopes/ScopeManager.test.ts`](../tests/src/core/scopes/ScopeManager.test.ts) — the id-keyed registry: `create` mints an `id` + stores (always adds — two scopes sharing a `name` coexist), `scope` / `scopes` insertion-order lookup, `remove` single + batch (`true` only when every supplied id was removed) + `clear` + `count`; the `create` / `remove` / `clear` event emissions (one `remove` per actually-removed id, the reserved `on` option); and emit-safety (a throwing `create` listener can't corrupt the registry + routes to the emitter's `error` handler; a throwing `error` handler neither escapes nor recurses) mirroring the Table / InstructionManager emitter convention.
|
|
1480
|
+
- [`tests/src/core/helpers.test.ts`](../tests/src/core/helpers.test.ts) — agent-owned pure helpers: `agentResultToJSON` full/minimal fresh exact projections + extras dropped + compile-time exhaustive `AgentResult` field-map review, conforming accessor/inherited structural values, and total rejection of wrong/missing fields, malformed present optionals, non-finite/missing usage counts, throwing getters, non-object inputs, and revoked root/nested proxies; `filterAllowList`'s semantics (`undefined` ⇒ all, `[]` ⇒ none, a list ⇒ only-listed) preserving item order (not allow-list order), ignoring unknown keys, matching through the key extractor (not identity), and returning `[]` (not throwing) for an empty item list; `estimateTokens` / `estimateMessages` (the per-message `MESSAGE_TOKEN_OVERHEAD`, an empty batch ⇒ 0, empty content ⇒ overhead only, the JSON-stringified `calls` contribution with its documented fixed fallback on a circular argument and no contribution for an empty `calls` array, and `images.length * IMAGE_TOKEN_ESTIMATE` per image); `renderFencedFile` (the `File:` label + language-tagged fence, the body verbatim across lines, and a workspace text file framed from its own text arm); `sanitizeToken` / `sanitizeUsage` (identity on a well-formed non-negative integer usage; `NaN` / negative / `±Infinity` floored to 0 and a fractional field to its integer part, each field independently); and `settleAgentJob`'s partial policy (a natural finish resolves `partial: false`, a disallowed partial throws an `AgentJobError` carrying the partial, an allowed partial resolves as success). And the extracted loop / cascade / conversation / scope leaves on their own contracts: `joinThinking` / `sumUsage` seeding then accumulating, `assembleResult` omitting an absent `thinking` / `usage` and keeping the loop-internal `exhausted` out of the public result, `denyCall`'s denial texts, `renderSection` rendering nothing for an empty item list, `resolveOpen` / `resolveClose` / `resolveItem` at each cascade level (item override beating every other), `attachImages` merging own-images-first without mutating the source, `attachUserImages` replacing only the last user turn (and returning the conversation unchanged for no data or no user turn), `collectImageData` skipping a text file, `buildSummaryMessage` / `buildRecapMessage` (raw vs. `CONVERSATION_RECAP_PREFIX`-framed), and `intersectKeys` treating `undefined` as the universal set and returning a copy.
|
|
1481
|
+
- [`tests/src/core/factories.test.ts`](../tests/src/core/factories.test.ts) — agent-owned factories over real declared tool/workspace dependencies: `createAgentContext` (the `[system?, ...messages]` assembly; a pre-built tool registry surfacing through `context.tools`); `createAgent` (one turn to its result; passed instructions / workspaces managers surface through `agent.context` — an added text file appears in `build()`; a no-tools scope empties the advertised definitions and filters instructions; omitted managers still yield working empty managers); `createChannel` (pushed values drain in write order then `close` ends it; a buffered value is delivered before a `fail` surfaces); `createAgentRegistry` (a serializable job round-trips build→run; a job naming a provider or tool missing from the registry rejects loudly on enqueue); `createAgentQueue` (each enqueue resolving its own result; the `concurrency` bound incl. a 12-job batch at 3; the partial policy — throws by default carrying the partial, re-runs for the full retry budget then rejects, `partial: true` resolves, a `budget: 0` partial settles the same way, a non-partial sibling resolves beside a throwing partial, a per-attempt timeout rejects as the substrate fault (not an `AgentJobError`) and retries, a pre-aborted entry signal hard-cancels without running; pause parks a job until resume, stop rejects a pending job, and `abort()` threads into the agent's signal); durability (an `AgentJobInput` JSON round-trips unchanged; a queue store round-trips the row; `restore()` re-runs an outstanding job to its real result then removes the row); `createAgentRunner` (an ordered batch, fail-fast on a partial, a parent spawning a child sub-agent through `controller.spawn`, cancel threading); and `AgentJobError` / `isAgentJobError` (carries the partial `AgentResult`; the guard narrows the real error and rejects everything else).
|
|
1482
|
+
- [`tests/src/core/conversations/stores/MemoryConversationStore.test.ts`](../tests/src/core/conversations/stores/MemoryConversationStore.test.ts) — the in-memory `ConversationStoreInterface` (`get` / `set` / `delete`, async, keyed by a snapshot's own id) over real `ConversationSnapshot`s carrying compacted sections (from a genuine `compact()`) + a live tail + a rollup `summary`: set→get round-trip + JSON-portability parity, upsert under the same id, delete + absent, two distinct ids coexist; and the `isToolCall` per-call guard (the fail-closed element check) — accepts the real `ToolCall` shape (string `id` / `name` + a record `arguments`), rejects every hostile shape (non-record, missing / wrong-typed `id` / `name` / `arguments`) without throwing.
|
|
1483
|
+
- [`tests/src/core/validators.test.ts`](../tests/src/core/validators.test.ts) — `isMessage`, `isSection`, and `isConversationSnapshot` (the per-message, per-section, and total read-boundary guards): each accepts the real shape with and without its optionals, rejects a non-record / nullish / primitive without throwing, rejects a missing or wrong-typed required field, and rejects a malformed nested element (`calls`, `messages`, `sections`); `isConversationSnapshot` also accepts a JSON-revived snapshot (the storage-read shape the database store narrows) and rejects a snapshot whose assistant `calls[]` carries a tampered element.
|
|
1484
|
+
- [`tests/src/core/conversations/stores/DatabaseConversationStore.test.ts`](../tests/src/core/conversations/stores/DatabaseConversationStore.test.ts) — the driver-pluggable twin over a real `createMemoryDriver` (no mocks): the same set→get round-trip through the one opaque JSON column (sections + tail + rollup summary survive), the default-driver factory overload (no arg) works the same, cross-instance durability (a second store over the same driver reads the snapshot back), upsert under the same id, delete + absent, two distinct ids coexisting, and a conversation store + a workspace store over separate drivers not colliding.
|
|
1485
|
+
- [`tests/src/core/AgentRegistry.test.ts`](../tests/src/core/AgentRegistry.test.ts) — the registry in isolation (a scripted provider): the accessors resolve a registered `provider` / `tool` / `authority` / `scheduler` and throw a category-specific `unknown <category>: <name>` on a miss; `build` rehydrates a seeded, signal-wired agent — the seed `messages` + `system` reach the agent and its `build()`, the `tools` names resolve into a fresh per-build manager (an unknown name throws), a threaded pre-aborted `signal` commits a partial, a `budget` ceiling becomes a token budget that bounds the loop, and a resolved `limit` / `scheduler` / `authority` reach the rehydrated loop (the cap bounds it, the scheduler paces between turns, the authority denies a call without executing it).
|
|
1486
|
+
- [`tests/src/core/Agent.test.ts`](../tests/src/core/Agent.test.ts) — the loop's deterministic logic over a local scripted `ProviderInterface`: a single no-tools turn → `generate` returns the content; the system prompt prepended + tools advertised structurally; tool iteration (turn 1's `ToolCall` → `execute` → the result fed back → turn 2's final content); a tool throw fed back as the tool message (the loop never throws); `generate` ↔ `stream` parity (same script → deep-equal); the `AgentChunk` sequence (`token` / `think` / `usage` / `tool` chunks, with the tool chunk carrying the executed call + result and usage summed); per-run `think` forwarding into the provider; the iteration cap (an always-tool script stops at `limit`); abort (a pre-aborted `signal` commits a partial without calling the provider; `abort()` mid-stream resolves `partial: true` with the accumulated content; a genuine provider error rejects); the token `budget` bound (exhausted usage stops the turn, `partial`); the `scheduler` yielding between turns (not after the last); `status` transitions; the authority gate wired into the loop (no authority → unchanged; an allowed call executes; a denied call is not executed [a counter tool proves it] yet a `tool` chunk + tool message carry the denial and the next provider call sees it; a mixed batch merges allowed + denied in original call order; an all-denied turn feeds back denials and the cap still bounds it); and the scope-filtered tool advertisement (no scope → all tools advertised; a `tools` allow-list → only the listed definitions reach the provider, a scoped-out tool absent on every turn so its handler never runs — neither described nor callable; an empty `tools` list → the provider is handed `undefined`); and automatic compaction (the context `window` budget): the between-turns trigger (the absolute prompt crossing `max` fires `compact()` + rebuilds smaller, exactly twice over the script, the run answering through the compacted view), the no-fire-below-the-window guard, and the additive regressions (no `window` ⇒ never folds; no conversation ⇒ the trigger is skipped + the budget untouched); and the production hardening (a long-conversation initial prompt compacted pre-first-turn so the first provider call sees the compacted view; a throwing auto-summarizer caught + surfaced as `fault` with the run continuing to a valid answer while a manual `compact()` still throws; the futile guard — a `compact()` that folds nothing while over the window latches per-run so auto-compaction stops, no churn); and the multi-conversation pattern (one agent + a `ConversationManager`, switching the active conversation with `agent.context.conversations.switch(id)` per request: independent accumulated histories with no cross-talk, and independent per-thread compaction whose sections retain only their own thread's originals).
|
|
1487
|
+
- [`tests/src/core/Authority.test.ts`](../tests/src/core/Authority.test.ts) — the `Authority` gate in isolation: ordered first-match-wins; a matched rule allows by default and denies on `allowed: false` (carrying `zone` / `reason`); no-match → the fallback; the default fallback is allow-`'default'`; an empty rules list always returns the fallback; a deny-by-default `fallback` makes unmatched calls denied (an allowlist); the matcher receives the `{ call }` context (branching on `call.name` and `call.arguments`).
|
|
1488
|
+
- [`tests/src/core/AgentProvider.test.ts`](../tests/src/core/AgentProvider.test.ts) — the shared HTTP engine over a scripted wire subclass and real `Response` bodies (no transport mock): identity and exact `format` exposure, the default transport bound to its global receiver, the posted body and case-insensitive header overrides, and a `headers` hook that adds only authentication leaving the JSON content type intact. Stream assembly — content, native reasoning, tool calls, and replacing usage; the qwen3 implicit-open reclassification in the authoritative result; a held content tail flushed as the final delta; verbatim content when `split` is disabled; a buffered record fed through `finish` with the parser cleared; a multibyte character split across byte chunks; a settled `result` record returned unchanged, accepted from `finish`, and leaving a following poison record undecoded with the body cancelled; and `strict` end of input with no settled result rejected. HTTP failures — the bounded error-body read with its remainder cancelled, the documented single-chunk overshoot, an exact-bound stalled body rejected before the deadline, the omitted separator for an empty excerpt, the retained status and cause when the body cannot be read at all, a successful response with no body as a protocol failure, and code / status / cause / `instanceof` narrowing. Failures that reach the caller unchanged — a hostile record decoder error and a remotely reported abort, the remote abort's identity preserved and its open body cancelled with the local signal unaborted. The `headers` hook inside the bound — raced against the deadline, a rejection preserved with no request issued, cancellation through the caller signal, and every abort listener removed after success, rejection, caller cancellation, and deadline expiry. Cancellation and partials — an already-aborted call rejected before framing or fetching, held content flushed into the partial, a stop between channels retaining the complete increment, a stalled body cancelled on the deadline, buffered `finish` records refused after a cancel, a transport `AbortError` normalized, a caller `ProviderAbortError` reason replaced by the local partial, a decoder failure that raced the cancel carried as the abort's `cause`, and that cause left `undefined` when the cancel was the only failure. Deadline clearing and reader release on every exit, and isolation between concurrent calls on one instance (including `generate` deep-equal to a drained `stream`).
|
|
1489
|
+
- [`tests/src/core/RelayStream.test.ts`](../tests/src/core/RelayStream.test.ts) — the relay's response half over real `Response` bodies: validated deltas and the authoritative result written with `RELAY_CONTENT_TYPE` and `cache-control: no-store`; an upstream failure reduced to the fixed `RELAY_PROVIDER_MESSAGE` with no secret text; a remote abort's partial written through the same compiled frame contract; a result outside the JSON frame contract refused; serialized `next` calls that stop pulling while the response queue is full (the backpressure proof); upstream aborted before the iterator returns with a late pending pull suppressed; an already-aborted inbound signal linked before the upstream turn starts; and an inbound abort with a frame still queued unread returning and finalizing the generator.
|
|
1490
|
+
- [`tests/src/core/providers/RelayProvider.test.ts`](../tests/src/core/providers/RelayProvider.test.ts) — the browser end: a synthetic `parameters`, `schema`, or call-`arguments` serializer ignored with the snapshot sent in its place, a projection failure carried as the refusal's `cause`, a valid request owned as a snapshot whose JSON values survive caller mutation, an `error` frame carrying a decorative `code` refused, the declared request fields projected with `caller` omitted across a JSON round trip, non-JSON arguments / parameters / schemas rejected, validated `content` and `thinking` frames mapped with literal `<think>` tags preserved, an `abort` frame reconstructed with its complete partial while the local signal stays unaborted, malformed frames and a missing terminal result refused, and a fragmented unterminated final result recovered with fresh framing for each call.
|
|
1491
|
+
- [`tests/src/core/integration.test.ts`](../tests/src/core/integration.test.ts) — the in-process hop, browser end to server end with no network: each refusal (`401`, `413`, `400`) carried to the browser as a `ProviderError` with its status while the upstream provider is never entered; a full round trip of identified messages, tools, options, deltas, and the authoritative result, deep-equal to the same provider driven directly; browser cancellation propagated through the request and upstream signals; a server-side abort reconstructed while the browser signal stays unaborted; and a secret upstream failure translated to the fixed public provider error. Beside it, provider-agnosticism over a minimal provider driving the full loop, a drop-in swap between differently named providers, and each provider's `format` reaching `build()`.
|
|
1492
|
+
- [`tests/src/core/contracts.test.ts`](../tests/src/core/contracts.test.ts) — the compiled wire contracts: a message, a request, a result, and every relay channel round-tripped, each reporting the path of its malformed field; function-valued tool arguments accepted in the domain and refused on the wire; non-JSON parameters and schemas refused before serialization can drop them; and every accepted message fixture still inside the domain guard.
|
|
1493
|
+
- [`tests/src/core/shapers.test.ts`](../tests/src/core/shapers.test.ts) — the wire shapes at the type level: each projection inferred assignable to the domain type it mirrors, and `ToolCall.caller` excluded from the tool wire shape.
|
|
1494
|
+
|
|
1495
|
+
## See also
|
|
1496
|
+
|
|
1497
|
+
- [`budget.md`](budget.md) — the cost primitive; `ProviderResult.usage` and `AgentResult.usage` reuse its `TokenUsage`, and a token budget bounds a provider call / an agent turn.
|
|
1498
|
+
- [`abort.md`](abort.md) / [`timeout.md`](timeout.md) — the bounding signals folded into a call's `AbortSignal` through `AbortSignal.any`; `@orkestrel/timeout` is also the deadline `AgentProvider` arms for every call it makes.
|
|
1499
|
+
- `@orkestrel/ndjson` — the newline-delimited JSON parser a relay browser application hands `createRelayProvider` as its `parser`. This package declares no dependency on it, so the application chooses the parser.
|
|
1500
|
+
- `@orkestrel/router` — one dispatcher a `RelayHandler` mounts on. The handler is fetch-standard, so any router that routes a `Request` to a `Response` works.
|
|
1501
|
+
- [`queue.md`](queue.md) — the bounded-concurrency, retrying, durable `Queue` `createAgentQueue` composes for many agent jobs.
|
|
1502
|
+
- [`workflow.md`](workflow.md) — the `SchedulerInterface` the loop yields to between turns, and the fail-fast `Runner` `createAgentRunner` composes for sub-agent fan-out.
|
|
1503
|
+
- [`emitter.md`](emitter.md) — the foundational observable primitive the `Agent` owns as its push `emitter`; `AgentEventMap` is its event map, wired through the reserved `on` option.
|
|
1504
|
+
- [`tool.md`](tool.md) — the tool runtime this loop advertises from and dispatches through: definitions, calls, and the success-discriminated `ToolResult`.
|
|
1505
|
+
- [`workspace.md`](workspace.md) — the file domain whose active workspace `AgentContext` renders into a turn: files, editing, events, and persistence.
|
|
1506
|
+
- [`contract.md`](contract.md) — the shape DSL other tools (for example `@orkestrel/toolbox`'s `createWorkspaceTool`) compile against; the shared `describedLiteral` (a discriminant's description-carrier) and `schemaToParameters` (the tool-parameters narrowing) live there.
|
|
1507
|
+
- [`database.md`](database.md) — the `DriverInterface` / `TableInterface` seam `createDatabaseConversationStore` persists conversation snapshots through.
|
|
1508
|
+
- [`AGENTS.md`](../AGENTS.md) — the repository's authority pointer; the coding rules it resolves to live in `@orkestrel/scaffold`.
|
|
1509
|
+
- [`README.md`](README.md) — the guides index.
|