agent-accelerator 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +1220 -0
- package/SYSTEM_PROMPT.md +14 -0
- package/SYSTEM_PROMPT_AGENT.md +34 -0
- package/SYSTEM_PROMPT_TOOLS.md +9 -0
- package/bunfig.toml +2 -0
- package/package.json +59 -0
- package/src/agent/agent.ts +615 -0
- package/src/agent/context.ts +161 -0
- package/src/agent/delegation.ts +481 -0
- package/src/agent/loop.ts +569 -0
- package/src/agent/subagent.ts +83 -0
- package/src/ai-sdk/converters.ts +342 -0
- package/src/ai-sdk/errors.ts +122 -0
- package/src/ai-sdk/executor.ts +454 -0
- package/src/ai-sdk/index.ts +55 -0
- package/src/ai-sdk/model-provider.ts +303 -0
- package/src/ai-sdk/options.ts +306 -0
- package/src/ai-sdk/provider.ts +415 -0
- package/src/ai-sdk/registry.ts +416 -0
- package/src/data/README.md +84 -0
- package/src/index.ts +190 -0
- package/src/models/catalog-cache.ts +273 -0
- package/src/models/catalog.ts +503 -0
- package/src/streaming/event-stream.ts +211 -0
- package/src/streaming/sse-parser.ts +97 -0
- package/src/tokens/counter.ts +136 -0
- package/src/tools/executor.ts +365 -0
- package/src/tools/schema.ts +221 -0
- package/src/tools/tool.ts +101 -0
- package/src/types/agent.ts +87 -0
- package/src/types/core.ts +86 -0
- package/src/types/message.ts +115 -0
- package/src/types/model.ts +212 -0
- package/src/types/provider-payloads.ts +434 -0
- package/src/types/response.ts +158 -0
- package/src/types/tool.ts +61 -0
- package/src/utils/base64.ts +27 -0
- package/src/utils/cache.ts +146 -0
- package/src/utils/env.ts +78 -0
- package/src/utils/headers.ts +110 -0
- package/src/utils/media.ts +137 -0
- package/src/utils/serialization.ts +91 -0
- package/src/utils/session.ts +26 -0
- package/src/utils/thought-signature.ts +27 -0
- package/tsconfig.json +31 -0
package/README.md
ADDED
|
@@ -0,0 +1,1220 @@
|
|
|
1
|
+
# Agent Accelerator
|
|
2
|
+
|
|
3
|
+
[](https://github.com/sashvat-bharat/agent-accelerator)
|
|
4
|
+
[](https://opensource.org/licenses/MIT)
|
|
5
|
+
|
|
6
|
+
A thin, typed transport SDK for calling LLMs through a single `Agent` interface.
|
|
7
|
+
|
|
8
|
+
GitHub: [https://github.com/sashvat-bharat/agent-accelerator](https://github.com/sashvat-bharat/agent-accelerator)
|
|
9
|
+
|
|
10
|
+
Agent Accelerator supports `google`, `opencode`, `openrouter`, `openai`, and any OpenAI-compatible endpoint using `{PREFIX}_API_KEY` and `{PREFIX}_BASE_URL`.
|
|
11
|
+
|
|
12
|
+
|
|
13
|
+
```ts
|
|
14
|
+
import { Agent } from "agent-accelerator";
|
|
15
|
+
|
|
16
|
+
const agent = new Agent({
|
|
17
|
+
model: "google/gemini-3.5-flash-lite",
|
|
18
|
+
instructions: "You are a concise research engineer.",
|
|
19
|
+
cache: { retention: "short" },
|
|
20
|
+
});
|
|
21
|
+
|
|
22
|
+
const res = await agent.run("Explain HBM pricing in 3 bullets");
|
|
23
|
+
|
|
24
|
+
console.log(res.text);
|
|
25
|
+
console.log(res.usage);
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
---
|
|
29
|
+
|
|
30
|
+
## Install
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
# Bun (recommended)
|
|
34
|
+
bun add agent-accelerator
|
|
35
|
+
|
|
36
|
+
# npm
|
|
37
|
+
npm install agent-accelerator
|
|
38
|
+
|
|
39
|
+
# pnpm
|
|
40
|
+
pnpm add agent-accelerator
|
|
41
|
+
|
|
42
|
+
# Yarn
|
|
43
|
+
yarn add agent-accelerator
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
Requires `bun` or `node 22+`.
|
|
47
|
+
|
|
48
|
+
Core dependencies are `zod`, `ai`, and selected `@ai-sdk/*` provider packages used as the backend transport.
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## Setup
|
|
53
|
+
|
|
54
|
+
```bash
|
|
55
|
+
GEMINI_API_KEY=...
|
|
56
|
+
OPENCODE_API_KEY=...
|
|
57
|
+
OPENROUTER_API_KEY=...
|
|
58
|
+
OPENAI_API_KEY=...
|
|
59
|
+
# or OPENAI_BASE_API_KEY=...
|
|
60
|
+
|
|
61
|
+
MODEL="google/gemini-3.5-flash-lite"
|
|
62
|
+
SUB_AGENT_MODEL="google/gemini-3.5-flash-lite"
|
|
63
|
+
|
|
64
|
+
# Any OpenAI-compatible endpoint, no code change:
|
|
65
|
+
# MODEL="groq/llama-3.3-70b-versatile"
|
|
66
|
+
# GROQ_API_KEY="..."
|
|
67
|
+
# GROQ_BASE_URL="https://api.groq.com/openai/v1"
|
|
68
|
+
|
|
69
|
+
# Local, no key needed:
|
|
70
|
+
# MODEL="ollama/qwen2.5-coder"
|
|
71
|
+
# OLLAMA_BASE_URL="http://localhost:11434/v1"
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
`MODEL` and `SUB_AGENT_MODEL` are used when `model` or the dynamic worker model are omitted.
|
|
75
|
+
|
|
76
|
+
Explicit configuration always takes precedence over environment variables.
|
|
77
|
+
|
|
78
|
+
---
|
|
79
|
+
|
|
80
|
+
## Quick Start
|
|
81
|
+
|
|
82
|
+
```ts
|
|
83
|
+
import { Agent } from "agent-accelerator";
|
|
84
|
+
|
|
85
|
+
const agent = new Agent({
|
|
86
|
+
model: "google/gemini-3.5-flash-lite",
|
|
87
|
+
instructions: "Be direct and technical.",
|
|
88
|
+
});
|
|
89
|
+
|
|
90
|
+
const res = await agent.run("Explain lock contention.");
|
|
91
|
+
|
|
92
|
+
console.log(res.text); // final answer
|
|
93
|
+
console.log(res.thinking); // reasoning trace, if any
|
|
94
|
+
console.log(res.usage); // input / output / cached / thinking / cost
|
|
95
|
+
console.log(res.raw.request.headers); // wire audit
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
Use `instructions` for the system prompt.
|
|
99
|
+
|
|
100
|
+
Keep `instructions` stable across turns to improve prefix-cache reuse.
|
|
101
|
+
|
|
102
|
+
---
|
|
103
|
+
|
|
104
|
+
## Agent
|
|
105
|
+
|
|
106
|
+
`new Agent(config: AgentConfig)` is the main entry point.
|
|
107
|
+
|
|
108
|
+
An `Agent` holds :
|
|
109
|
+
|
|
110
|
+
* `instructions`
|
|
111
|
+
* `tools`
|
|
112
|
+
* resolved model
|
|
113
|
+
* thinking configuration
|
|
114
|
+
* cache configuration
|
|
115
|
+
* service tier
|
|
116
|
+
* `sessionId`
|
|
117
|
+
* conversation `context`
|
|
118
|
+
|
|
119
|
+
### AgentConfig
|
|
120
|
+
|
|
121
|
+
| Field | Type | Meaning / Use Case |
|
|
122
|
+
| ----------------- | ---------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
123
|
+
| `model` | `string \| ModelSpec \| ModelProviderInstance` | Model identifier, such as `"google/gemini-3.5-flash-lite"`. Use `ModelProvider.*` or `ModelSpec` to pin a provider and credentials. Falls back to `MODEL`. |
|
|
124
|
+
| `instructions` | `string` | System prompt. Keep it stable; place per-turn additions in `additionalContext`. |
|
|
125
|
+
| `tools` | `Record<string, ToolDefinition> \| ToolDefinition[]` | Deterministic functions the model can call. See [Tools](#tools). |
|
|
126
|
+
| `subagents` | `SubAgent[] \| Agent[]` | Pre-defined workers. Each worker becomes a callable tool. Useful for fixed roles such as researcher or critic. |
|
|
127
|
+
| `subagentModel` | `string \| ModelSpec \| ModelProviderInstance` | Developer-only model for dynamically spawned workers. Never choosable by the Main Agent. Overridden by `dynamicSubagents.model`. Falls back to `SUB_AGENT_MODEL`. |
|
|
128
|
+
| `dynamicSubagents` | `DynamicSubagentsConfig` | Enables and constrains LLM-spawned stateless workers: `{ enabled, model?, maxSpawn?, thinkingLevel?, tools?, timeout? }`. See [Dynamic delegation](#dynamic-delegation). |
|
|
129
|
+
| `thinkingLevel` | `ThinkingLevel` | `none \| dynamic \| minimal \| low \| medium \| high \| xhigh`. Validated against the model catalog before a request. |
|
|
130
|
+
| `cache` | `CacheConfig` | `{ retention, sessionId, cachedContentId, ttlSeconds }`. Controls cache reuse. |
|
|
131
|
+
| `serviceTier` | `"flex" \| "priority"` | Cost / priority routing where supported. Omit for standard routing. |
|
|
132
|
+
| `maxTurns` | `number` | Maximum model → tool → model loops per `run`. Defaults to `10`. |
|
|
133
|
+
| `sessionId` | `string` | Stable identifier used for cache affinity. Auto-generated when omitted. |
|
|
134
|
+
| `headers` | `Record<string,string>` | Additional headers merged into every request. |
|
|
135
|
+
| `apiKey` | `string` | Overrides environment-based API-key lookup for this agent. |
|
|
136
|
+
| `baseUrl` | `string` | Overrides the default endpoint for this agent. |
|
|
137
|
+
| `stateless` | `boolean` | When `true`, history is cleared before and after each `run`. Useful for one-shot evaluators. Defaults to `false`. |
|
|
138
|
+
|
|
139
|
+
When `dynamicSubagents.enabled` is set without a resolvable worker model, construction throws `SubAgentModelError`.
|
|
140
|
+
|
|
141
|
+
### Agent Methods
|
|
142
|
+
|
|
143
|
+
```ts
|
|
144
|
+
await agent.run(prompt, opts?) // non-streaming, or streaming when opts.stream is true
|
|
145
|
+
agent.ask(prompt, optsOrBool?) // string deltas when streaming, otherwise same as run
|
|
146
|
+
agent.stream(prompt, opts?) // AssistantMessageEventStream
|
|
147
|
+
agent.reset() // clears messages + signatures, keeps config
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
`prompt` accepts `string | ContentPart[]`.
|
|
151
|
+
|
|
152
|
+
Use `ContentPart[]` when sending images, audio, or video alongside text.
|
|
153
|
+
|
|
154
|
+
### AgentRunOptions
|
|
155
|
+
|
|
156
|
+
| Field | Meaning / Use Case |
|
|
157
|
+
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
158
|
+
| `stream` | When `true`, returns a stream instead of a regular response promise. |
|
|
159
|
+
| `onDelta(delta, event)` | Called for each text chunk. Use to render answers live. |
|
|
160
|
+
| `onThinkingDelta(delta, event)` | Called for each reasoning chunk. Use to render reasoning separately. |
|
|
161
|
+
| `onEvent(event)` | Called for every event: `text_delta`, `thinking_delta`, `tool_call_complete`, `tool_result`, `subagent_complete`, `usage`, `done`, and `error`. |
|
|
162
|
+
| `wrapThinking` | Wraps reasoning as `<think>\n...\n</think>\n\n`, allowing UIs to render it without maintaining separate state. |
|
|
163
|
+
| `additionalContext` | Added only to the current user turn. Keeps the system prompt stable for better caching. |
|
|
164
|
+
| `sessionId` | Per-run session override. |
|
|
165
|
+
| `headers` | Per-run header merge. |
|
|
166
|
+
| `signal` | `AbortSignal` used to cancel generation, tools, and subagents. |
|
|
167
|
+
|
|
168
|
+
### Streaming
|
|
169
|
+
|
|
170
|
+
```ts
|
|
171
|
+
const res = await agent.run("Design an append-only log", {
|
|
172
|
+
stream: true,
|
|
173
|
+
wrapThinking: true,
|
|
174
|
+
onThinkingDelta: (d) => process.stdout.write(d),
|
|
175
|
+
onDelta: (d) => process.stdout.write(d),
|
|
176
|
+
onEvent: (e) => {
|
|
177
|
+
if (e.type === "subagent_complete") {
|
|
178
|
+
console.log(`done: ${e.subagent!.name}`);
|
|
179
|
+
}
|
|
180
|
+
|
|
181
|
+
if (e.type === "tool_result") {
|
|
182
|
+
console.log(`tool: ${e.toolResult!.name}`);
|
|
183
|
+
}
|
|
184
|
+
},
|
|
185
|
+
});
|
|
186
|
+
```
|
|
187
|
+
|
|
188
|
+
### Manual Iteration
|
|
189
|
+
|
|
190
|
+
```ts
|
|
191
|
+
const stream = agent.stream("Hello");
|
|
192
|
+
|
|
193
|
+
for await (const e of stream) {
|
|
194
|
+
if (e.type === "text_delta") {
|
|
195
|
+
process.stdout.write(e.delta!);
|
|
196
|
+
}
|
|
197
|
+
|
|
198
|
+
if (e.type === "thinking_delta") {
|
|
199
|
+
process.stdout.write(e.thinkingDelta!);
|
|
200
|
+
}
|
|
201
|
+
}
|
|
202
|
+
|
|
203
|
+
const final = await stream.result();
|
|
204
|
+
|
|
205
|
+
stream.on("text_delta", (e) => console.log(e.delta));
|
|
206
|
+
stream.off("text_delta", handler);
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+
---
|
|
210
|
+
|
|
211
|
+
## SubAgent
|
|
212
|
+
|
|
213
|
+
`SubAgent extends Agent` and is designed for fixed worker roles.
|
|
214
|
+
|
|
215
|
+
```ts
|
|
216
|
+
import { Agent, SubAgent } from "agent-accelerator";
|
|
217
|
+
|
|
218
|
+
const researcher = new SubAgent({
|
|
219
|
+
name: "researcher",
|
|
220
|
+
instructions: "You collect quantifiable metrics.",
|
|
221
|
+
model: "google/gemini-3.5-flash-lite",
|
|
222
|
+
});
|
|
223
|
+
|
|
224
|
+
const lead = new Agent({
|
|
225
|
+
model: "google/gemini-3.7-flash",
|
|
226
|
+
instructions: "Delegate research, then synthesize.",
|
|
227
|
+
subagents: [researcher],
|
|
228
|
+
});
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
### Differences from Agent
|
|
232
|
+
|
|
233
|
+
* `model` resolves as `config.model ?? SUB_AGENT_MODEL ?? MODEL` and throws if empty.
|
|
234
|
+
* `cache` defaults to `{ retention: "short" }`.
|
|
235
|
+
* `.asTool(name?, desc?)` / `.toTool(name?, desc?)` converts the subagent into a `ToolDefinition` for manual registration.
|
|
236
|
+
* `stateless: true` is supported for one-shot judges that must not retain history.
|
|
237
|
+
|
|
238
|
+
`SubAgentConfig` is identical to `AgentConfig`.
|
|
239
|
+
|
|
240
|
+
### Pre-defined vs Dynamic Delegation
|
|
241
|
+
|
|
242
|
+
**Pre-defined delegation**
|
|
243
|
+
|
|
244
|
+
Pass `subagents: [a, b]`.
|
|
245
|
+
|
|
246
|
+
Each subagent is automatically registered as an internal tool on the parent that accepts:
|
|
247
|
+
|
|
248
|
+
```ts
|
|
249
|
+
{ task: string }
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
**Dynamic delegation**
|
|
253
|
+
|
|
254
|
+
`subagents` lists pre-defined workers (each becomes a callable tool). For
|
|
255
|
+
LLM-spawned workers, configure `dynamicSubagents`:
|
|
256
|
+
|
|
257
|
+
```ts
|
|
258
|
+
const agent = new Agent({
|
|
259
|
+
name: "Main Agent",
|
|
260
|
+
model: process.env.MODEL,
|
|
261
|
+
tools: { get_topic_brief },
|
|
262
|
+
instructions: "You are the Main Agent",
|
|
263
|
+
dynamicSubagents: {
|
|
264
|
+
enabled: true,
|
|
265
|
+
maxSpawn: 4,
|
|
266
|
+
thinkingLevel: "low",
|
|
267
|
+
tools: { get_weather, recent_news },
|
|
268
|
+
timeout: 60000,
|
|
269
|
+
},
|
|
270
|
+
});
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
A `spawn_subagents` tool is injected with the following task shape:
|
|
274
|
+
|
|
275
|
+
```ts
|
|
276
|
+
{
|
|
277
|
+
tasks: {
|
|
278
|
+
name,
|
|
279
|
+
role?,
|
|
280
|
+
instructions,
|
|
281
|
+
task,
|
|
282
|
+
tools?: string[],
|
|
283
|
+
timeoutMs?: number
|
|
284
|
+
}[]
|
|
285
|
+
}
|
|
286
|
+
```
|
|
287
|
+
|
|
288
|
+
- `maxSpawn` — max workers per call. Extras are trimmed safely, and the limit is stated in the tool description plus the Main Agent's system prompt so the model knows it.
|
|
289
|
+
- `tools` — developer-owned pool. Workers get zero tools unless the Main Agent grants a per-task `tools` subset by name (unknown names are ignored), keeping worker context small.
|
|
290
|
+
- `timeout` (ms) — `0` = no limit, `-1` = the Main Agent sets a per-task `timeoutMs`, `>0` = fixed limit for every worker. Timed-out workers report an error entry; the rest of the batch still completes.
|
|
291
|
+
- Workers are stateless: one task in, one result out, then shut down. No history, no recursion. The Main Agent cannot choose worker models or reasoning levels.
|
|
292
|
+
|
|
293
|
+
```ts
|
|
294
|
+
const res = await agent.run("Audit auth pipeline and write a threat model");
|
|
295
|
+
|
|
296
|
+
console.log(res.subagents.map((s) => s.name));
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
### Delegation Helpers
|
|
300
|
+
|
|
301
|
+
* `buildAgentTools(list)` — wraps multiple agents and deduplicates names. Used internally by `subagents`.
|
|
302
|
+
* `createSubagentSpawnTool(parentAgent)` — builds the dynamic subagent spawner. Rarely needed directly.
|
|
303
|
+
* `DynamicSubagentTask` — `{ name, role?, instructions, task, tools?, timeoutMs? }`, the shape produced for dynamic delegation. No `model` — workers run on the developer-configured model.
|
|
304
|
+
|
|
305
|
+
---
|
|
306
|
+
## Tools
|
|
307
|
+
|
|
308
|
+
```ts
|
|
309
|
+
import { Agent, tool, z } from "agent-accelerator";
|
|
310
|
+
|
|
311
|
+
const agent = new Agent({
|
|
312
|
+
model: "google/gemini-3.5-flash-lite",
|
|
313
|
+
tools: {
|
|
314
|
+
fetch_metrics: tool({
|
|
315
|
+
description: "Fetch cluster metrics",
|
|
316
|
+
input: z.object({
|
|
317
|
+
clusterId: z.string(),
|
|
318
|
+
}),
|
|
319
|
+
execute: async ({ clusterId }) => ({
|
|
320
|
+
clusterId,
|
|
321
|
+
load: 0.42,
|
|
322
|
+
}),
|
|
323
|
+
}),
|
|
324
|
+
},
|
|
325
|
+
});
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
### Tool Helpers
|
|
329
|
+
|
|
330
|
+
* `tool({ name?, description, input?, parameters?, strict?, timeoutMs?, maxTries?, maxConcurrency?, execute })` — creates a `ToolDefinition`. Provide either a Zod `input` schema or raw JSON `parameters`. `execute(input, ctx)` may return arbitrary values; results are converted safely for model context. `timeoutMs` defaults to `0` (no time limit) and is a best-effort event-loop deadline; `ctx.signal` enables cooperative cancellation. `maxTries` accepts a number or numeric string; a positive value is the total attempt limit, while `0`/omitted means no configured limit for transient retries. `maxConcurrency` limits simultaneous calls for that tool (default pool: 8).
|
|
331
|
+
* `toStandardToolDeclarations(record | array)` — converts tools to `{ name, description, parameters }` for providers.
|
|
332
|
+
* `zodToJsonSchema(schema)` — converts Zod schemas using native conversion with a fallback extractor.
|
|
333
|
+
* `cleanJsonSchema(schema)` — removes `$schema`, `$defs`, and `definitions`, resolves `$ref`, and preserves explicit `additionalProperties`.
|
|
334
|
+
* `executeToolCalls({ tools, toolCalls, agentName?, parallel?, signal?, sessionId? })` — executes tool calls. Calls validate Zod input, use bounded parallelism, enforce per-tool timeouts, retry transient failures, recover namespace/camel-case tool-name aliases, and return `ToolResultRecord[]`. Missing tools and aborts are represented as error results rather than thrown.
|
|
335
|
+
|
|
336
|
+
The agent loop also blocks an identical tool name and argument set when the model requests it in the immediately following turn. The synthetic error is returned to the model so it can reuse the prior result or change its arguments.
|
|
337
|
+
|
|
338
|
+
### Tool Timing and Deadlines
|
|
339
|
+
|
|
340
|
+
Every `ToolResultRecord` includes `durationMs`, measured in whole milliseconds with a minimum displayed value of `1ms`. The value covers the tool execution attempt (including retry/backoff time), but excludes queue wait and input-schema validation.
|
|
341
|
+
|
|
342
|
+
`timeoutMs` is specified in milliseconds:
|
|
343
|
+
|
|
344
|
+
```ts
|
|
345
|
+
const get_status = tool({
|
|
346
|
+
name: "get_status",
|
|
347
|
+
description: "Check user authentication status.",
|
|
348
|
+
timeoutMs: 1,
|
|
349
|
+
input: z.object({ username: z.string() }),
|
|
350
|
+
execute: async ({ username }, ctx) => {
|
|
351
|
+
// Pass ctx.signal to cancellable APIs such as fetch.
|
|
352
|
+
return username === "Akshat Dwivedi" ? "Valid" : "Invalid";
|
|
353
|
+
},
|
|
354
|
+
});
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
`timeoutMs: 0` or an omitted value disables the deadline. Positive values are best-effort event-loop deadlines: JavaScript cannot forcibly stop arbitrary synchronous code, so asynchronous implementations should honor `ctx.signal`. A timeout is returned as an error tool result and is not retried indefinitely unless a positive `maxTries` is configured.
|
|
358
|
+
|
|
359
|
+
Tool timing is available after a run:
|
|
360
|
+
|
|
361
|
+
```ts
|
|
362
|
+
const response = await agent.run("Check status");
|
|
363
|
+
for (const result of response.toolResults) {
|
|
364
|
+
console.log(`${result.name}: ${result.durationMs}ms`);
|
|
365
|
+
}
|
|
366
|
+
```
|
|
367
|
+
|
|
368
|
+
### Tool Types
|
|
369
|
+
|
|
370
|
+
`ToolExecutionContext`:
|
|
371
|
+
|
|
372
|
+
```ts
|
|
373
|
+
{
|
|
374
|
+
toolCallId,
|
|
375
|
+
agentName?,
|
|
376
|
+
signal?,
|
|
377
|
+
sessionId?
|
|
378
|
+
}
|
|
379
|
+
```
|
|
380
|
+
|
|
381
|
+
`ToolCallRecord`:
|
|
382
|
+
|
|
383
|
+
```ts
|
|
384
|
+
{
|
|
385
|
+
id,
|
|
386
|
+
name,
|
|
387
|
+
arguments,
|
|
388
|
+
rawArguments?,
|
|
389
|
+
thoughtSignature?
|
|
390
|
+
}
|
|
391
|
+
```
|
|
392
|
+
|
|
393
|
+
`ToolResultRecord`:
|
|
394
|
+
|
|
395
|
+
```ts
|
|
396
|
+
{
|
|
397
|
+
id,
|
|
398
|
+
name,
|
|
399
|
+
result,
|
|
400
|
+
isError?,
|
|
401
|
+
durationMs?
|
|
402
|
+
}
|
|
403
|
+
```
|
|
404
|
+
|
|
405
|
+
---
|
|
406
|
+
## Providers
|
|
407
|
+
|
|
408
|
+
First-class provider IDs are:
|
|
409
|
+
|
|
410
|
+
* `google`
|
|
411
|
+
* `opencode`
|
|
412
|
+
* `openrouter`
|
|
413
|
+
* `openai`
|
|
414
|
+
|
|
415
|
+
Any other provider prefix is treated as an OpenAI-compatible custom provider.
|
|
416
|
+
|
|
417
|
+
### Model Strings
|
|
418
|
+
|
|
419
|
+
* `"google/<id>"` → Google AI Studio at `generativelanguage.googleapis.com`.
|
|
420
|
+
* `"opencode/<id>"` → OpenCode Zen at `opencode.ai/zen/v1`.
|
|
421
|
+
* `"opencode-go/<id>"` → OpenCode Go endpoint.
|
|
422
|
+
* OpenCode automatically selects Chat vs Responses API. Claude models and models using `api=openai-responses` use Responses.
|
|
423
|
+
* `"openrouter/<scope>/<model>"` or `"scope/model:variant"` → OpenRouter.
|
|
424
|
+
* `"openai/<id>"` → OpenAI.
|
|
425
|
+
* `"groq/<id>"`, `"ollama/<id>"`, etc. → custom providers using `{PREFIX}_API_KEY` and `{PREFIX}_BASE_URL`. The API key is optional for local endpoints.
|
|
426
|
+
* A bare `"some-model"` uses `OPENAI_BASE_URL` when configured. Otherwise, catalog lookup is attempted, followed by an `openrouter` fallback.
|
|
427
|
+
|
|
428
|
+
Model resolution never silently selects a default model.
|
|
429
|
+
|
|
430
|
+
An empty model throws.
|
|
431
|
+
|
|
432
|
+
An unknown prefix with no catalog match still routes to a custom provider, allowing private model IDs and endpoints.
|
|
433
|
+
|
|
434
|
+
### Registry
|
|
435
|
+
|
|
436
|
+
* `resolveModel("google/x" | ModelSpec)` → `{ provider, modelId, modelSpec? }`. Use to inspect routing before execution.
|
|
437
|
+
* `getProvider("groq")` → returns a cached provider and automatically creates a custom provider for unknown prefixes.
|
|
438
|
+
* `ensureCustomProvider(prefix, { baseUrl?, apiKey?, name? })` → gets or creates a custom provider. Useful for multiple endpoints in one process.
|
|
439
|
+
* `normalizeProviderPrefix(s)` → lowercases and trims a provider prefix.
|
|
440
|
+
* `ModelProvider.GoogleGenAI(model, apiKey?, { thinkingLevel?, baseUrl? })` — same shape is available for `.OpenCode`, `.OpenRouter`, `.OpenAI`, `.Custom`, `.OpenAICompatible`, and `.Generic`.
|
|
441
|
+
|
|
442
|
+
Each returns a `ModelProviderInstance`:
|
|
443
|
+
|
|
444
|
+
```ts
|
|
445
|
+
{
|
|
446
|
+
model,
|
|
447
|
+
apiKey?,
|
|
448
|
+
baseUrl?,
|
|
449
|
+
thinkingLevel?
|
|
450
|
+
}
|
|
451
|
+
```
|
|
452
|
+
|
|
453
|
+
The result can be passed directly to `Agent.model`.
|
|
454
|
+
|
|
455
|
+
Use `Custom` for Groq, Together, Ollama, vLLM, and similar endpoints.
|
|
456
|
+
|
|
457
|
+
```ts
|
|
458
|
+
import { Agent, ModelProvider } from "agent-accelerator";
|
|
459
|
+
|
|
460
|
+
new Agent({
|
|
461
|
+
model: ModelProvider.GoogleGenAI("gemini-3.5-flash-lite"),
|
|
462
|
+
});
|
|
463
|
+
|
|
464
|
+
new Agent({
|
|
465
|
+
model: ModelProvider.Custom(
|
|
466
|
+
"groq/llama-3.3-70b-versatile",
|
|
467
|
+
process.env.GROQ_API_KEY,
|
|
468
|
+
{
|
|
469
|
+
baseUrl: process.env.GROQ_BASE_URL,
|
|
470
|
+
},
|
|
471
|
+
),
|
|
472
|
+
});
|
|
473
|
+
```
|
|
474
|
+
|
|
475
|
+
`BaseProvider` is the abstract provider base containing `id`, `name`, `models`, `getModel`, `generate`, `stream`, and `countTokens`.
|
|
476
|
+
|
|
477
|
+
Its catalog-first `getModel` behavior includes a hardcoded fallback.
|
|
478
|
+
|
|
479
|
+
Extend `BaseProvider` when implementing a fully custom transport.
|
|
480
|
+
|
|
481
|
+
`ProviderRequestOptions { apiKey?, baseUrl?, headers?, thinking?, cache?, serviceTier?, tools?, toolChoice?, signal?, sessionId?, env? }` is the per-call options object passed to `generate`/`stream`. Agent builds it automatically.
|
|
482
|
+
|
|
483
|
+
### Google
|
|
484
|
+
|
|
485
|
+
`GoogleAIStudioProvider` and `GOOGLE_MODELS` provide Google AI Studio support.
|
|
486
|
+
|
|
487
|
+
The catalog is filtered with fallback support for lite and flash models.
|
|
488
|
+
|
|
489
|
+
Google models matching `gemini-1.x` or `gemini-2.x` are rejected.
|
|
490
|
+
|
|
491
|
+
Thinking levels map to:
|
|
492
|
+
|
|
493
|
+
```text
|
|
494
|
+
OFF | MINIMAL | LOW | MEDIUM | HIGH
|
|
495
|
+
```
|
|
496
|
+
|
|
497
|
+
Google-specific handling includes:
|
|
498
|
+
|
|
499
|
+
* Isolating `thoughtSignature` per part for prefix stability.
|
|
500
|
+
* Converting JSON Schema to OpenAPI 3.0 through `stripSchemaForGoogle`.
|
|
501
|
+
* Supporting explicit `cachedContents` when `retention != implicit` and the prompt is large.
|
|
502
|
+
|
|
503
|
+
Helpers:
|
|
504
|
+
|
|
505
|
+
* `createExplicitCache({ model, systemInstruction?, contents?, tools?, displayName?, ttlSeconds?, expireTime?, apiKey?, baseUrl? })`
|
|
506
|
+
* `isValidThoughtSignature(sig)`
|
|
507
|
+
* `retainThoughtSignature(existing, incoming)`
|
|
508
|
+
* `stripSchemaForGoogle(schema)`
|
|
509
|
+
|
|
510
|
+
### OpenAI
|
|
511
|
+
|
|
512
|
+
`OpenAIProvider` and `OPENAI_MODELS` provide OpenAI support.
|
|
513
|
+
|
|
514
|
+
`OpenAIProvider` is also the base class for OpenAI-wire transports.
|
|
515
|
+
|
|
516
|
+
Thinking levels map to:
|
|
517
|
+
|
|
518
|
+
```text
|
|
519
|
+
reasoning_effort: low | medium | high
|
|
520
|
+
```
|
|
521
|
+
|
|
522
|
+
It also:
|
|
523
|
+
|
|
524
|
+
* passes `service_tier`;
|
|
525
|
+
* derives `max_completion_tokens` from the configured budget;
|
|
526
|
+
* preserves Google thought signatures through `extra_content.google.thought_signature`;
|
|
527
|
+
* exposes `extractGoogleThoughtSignature(obj)`.
|
|
528
|
+
|
|
529
|
+
### OpenCode
|
|
530
|
+
|
|
531
|
+
`OpenCodeProvider` and `OPENCODE_MODELS` provide OpenCode support.
|
|
532
|
+
|
|
533
|
+
Requests include:
|
|
534
|
+
|
|
535
|
+
* `session_id`
|
|
536
|
+
* `prompt_cache_key`
|
|
537
|
+
* headers used for sticky routing
|
|
538
|
+
|
|
539
|
+
The provider handles both Chat and Responses payloads.
|
|
540
|
+
|
|
541
|
+
Transient `500`, `502`, `503`, and `529` errors are retried with thinking disabled, followed by an attempt using the alternate endpoint.
|
|
542
|
+
|
|
543
|
+
### OpenRouter
|
|
544
|
+
|
|
545
|
+
`OpenRouterProvider` and `OPENROUTER_MODELS` provide OpenRouter support.
|
|
546
|
+
|
|
547
|
+
Requests include:
|
|
548
|
+
|
|
549
|
+
* `session_id`
|
|
550
|
+
* `prompt_cache_key`
|
|
551
|
+
|
|
552
|
+
Thinking is mapped to:
|
|
553
|
+
|
|
554
|
+
```text
|
|
555
|
+
reasoning { effort | enabled }
|
|
556
|
+
include_reasoning
|
|
557
|
+
```
|
|
558
|
+
|
|
559
|
+
based on catalog capabilities.
|
|
560
|
+
|
|
561
|
+
`serviceTier` is mapped to provider routing.
|
|
562
|
+
|
|
563
|
+
### Custom
|
|
564
|
+
|
|
565
|
+
`OpenAICompatibleProvider` is also exposed through the `CustomProvider`, `createCustomProvider`, and `createOpenAICompatibleProvider(prefix, opts?)` aliases.
|
|
566
|
+
|
|
567
|
+
`CustomProviderOptions`:
|
|
568
|
+
|
|
569
|
+
```ts
|
|
570
|
+
{
|
|
571
|
+
name?,
|
|
572
|
+
baseUrl?,
|
|
573
|
+
apiKey?,
|
|
574
|
+
defaultBaseUrl?
|
|
575
|
+
}
|
|
576
|
+
```
|
|
577
|
+
|
|
578
|
+
`createGenericModelSpec(provider, modelId)` creates a permissive placeholder so private model IDs can pass validation and window checks.
|
|
579
|
+
|
|
580
|
+
Provider wire types are exported for inspecting raw requests and responses through:
|
|
581
|
+
|
|
582
|
+
```ts
|
|
583
|
+
res.raw.request.body
|
|
584
|
+
res.raw.response.body
|
|
585
|
+
```
|
|
586
|
+
|
|
587
|
+
Examples include:
|
|
588
|
+
|
|
589
|
+
* `OpenAIChatCompletionRequest`
|
|
590
|
+
* `GoogleGenerateContentRequest`
|
|
591
|
+
* `OpenCodeChatRequest`
|
|
592
|
+
* `OpenRouterChatRequest`
|
|
593
|
+
|
|
594
|
+
---
|
|
595
|
+
## Thinking
|
|
596
|
+
|
|
597
|
+
Thinking uses a single `thinkingLevel` flag:
|
|
598
|
+
|
|
599
|
+
```ts
|
|
600
|
+
ThinkingLevel =
|
|
601
|
+
| "none"
|
|
602
|
+
| "dynamic"
|
|
603
|
+
| "minimal"
|
|
604
|
+
| "low"
|
|
605
|
+
| "medium"
|
|
606
|
+
| "high"
|
|
607
|
+
| "xhigh";
|
|
608
|
+
```
|
|
609
|
+
|
|
610
|
+
### Thinking Levels
|
|
611
|
+
|
|
612
|
+
* `none` — disables thinking where supported.
|
|
613
|
+
* `dynamic` — allows the model to decide.
|
|
614
|
+
* `minimal`, `low`, `medium`, `high`, `xhigh` — request increasing levels of reasoning where supported.
|
|
615
|
+
|
|
616
|
+
Thinking can be configured on the `Agent` or through `ModelProvider.*(..., { thinkingLevel })`.
|
|
617
|
+
|
|
618
|
+
The per-run `thinkingLevel` option can override the agent default for one request.
|
|
619
|
+
|
|
620
|
+
Internally, thinking is normalized to:
|
|
621
|
+
|
|
622
|
+
```ts
|
|
623
|
+
ThinkingConfig {
|
|
624
|
+
enabled?,
|
|
625
|
+
level?,
|
|
626
|
+
budgetTokens?,
|
|
627
|
+
includeThoughts?
|
|
628
|
+
}
|
|
629
|
+
```
|
|
630
|
+
|
|
631
|
+
### Catalog Helpers
|
|
632
|
+
|
|
633
|
+
* `getModelThinkingInfo(provider, modelId)` → `{ supportsThinking, reasoningOptions?, allowedLevels, supportsDisable, description }`. Useful for building UI selectors.
|
|
634
|
+
* `validateModelThinking(provider, modelId, level)` → throws `ThinkingLevelError { provider, modelId, requestedLevel, allowedLevels, supportsThinking }` with a fix hint. Automatically called by `run` and `stream`. Catch it and use `err.allowedLevels` to offer valid options.
|
|
635
|
+
* `getModelsForProvider("google")` → returns known model specifications.
|
|
636
|
+
|
|
637
|
+
Unknown models allow all thinking levels.
|
|
638
|
+
|
|
639
|
+
Fixed-reasoning models reject `none`.
|
|
640
|
+
|
|
641
|
+
---
|
|
642
|
+
## Cache
|
|
643
|
+
|
|
644
|
+
```ts
|
|
645
|
+
CacheConfig {
|
|
646
|
+
retention?,
|
|
647
|
+
cachedContentId?,
|
|
648
|
+
sessionId?,
|
|
649
|
+
ttlSeconds?
|
|
650
|
+
}
|
|
651
|
+
```
|
|
652
|
+
|
|
653
|
+
```ts
|
|
654
|
+
CacheRetention =
|
|
655
|
+
| "implicit"
|
|
656
|
+
| "short"
|
|
657
|
+
| "medium"
|
|
658
|
+
| "long";
|
|
659
|
+
```
|
|
660
|
+
|
|
661
|
+
### Retention
|
|
662
|
+
|
|
663
|
+
* `implicit` — automatic prefix reuse without an explicit cache object or storage fee.
|
|
664
|
+
* `short` — approximately 5 minutes.
|
|
665
|
+
* `medium` — approximately 1 hour.
|
|
666
|
+
* `long` — approximately 12 hours for Google explicit caches and approximately 24 hours for prompt caching where supported.
|
|
667
|
+
|
|
668
|
+
### Session Affinity
|
|
669
|
+
|
|
670
|
+
`sessionId` pins provider affinity using mechanisms such as:
|
|
671
|
+
|
|
672
|
+
* `x-session-id`
|
|
673
|
+
* `x-opencode-session`
|
|
674
|
+
* `prompt_cache_key` (openai/openrouter/opencode only — strict endpoints such as groq reject it, so custom providers get headers only)
|
|
675
|
+
|
|
676
|
+
Reuse the same session ID across turns when cache affinity is desired.
|
|
677
|
+
|
|
678
|
+
In browser runtimes only CORS-safe attribution headers are sent; affinity
|
|
679
|
+
still flows via `promptCacheKey` provider options, never via headers.
|
|
680
|
+
|
|
681
|
+
### Explicit Caches
|
|
682
|
+
|
|
683
|
+
`cachedContentId` reuses a previously created Google `cachedContents/...` resource.
|
|
684
|
+
|
|
685
|
+
`ttlSeconds` overrides the retention mapping.
|
|
686
|
+
|
|
687
|
+
### Provider Cache Mechanisms
|
|
688
|
+
|
|
689
|
+
* **Google** keeps system prompts and tools stable and isolates signatures.
|
|
690
|
+
* **OpenCode / OpenRouter** use sticky session keys.
|
|
691
|
+
|
|
692
|
+
The following cache helpers are exported:
|
|
693
|
+
|
|
694
|
+
* `applyAnthropicCacheControl`
|
|
695
|
+
* `getCacheControlForRetention`
|
|
696
|
+
* `getPromptCacheRetention`
|
|
697
|
+
* `clampCacheKey`
|
|
698
|
+
|
|
699
|
+
For the best cache reuse, keep `instructions` and the tool set stable. Put changing data in user messages or `additionalContext`.
|
|
700
|
+
|
|
701
|
+
---
|
|
702
|
+
## serviceTier
|
|
703
|
+
|
|
704
|
+
```ts
|
|
705
|
+
serviceTier = "flex" | "priority";
|
|
706
|
+
```
|
|
707
|
+
|
|
708
|
+
Omit `serviceTier` for standard routing.
|
|
709
|
+
|
|
710
|
+
Where supported, it is passed as `service_tier` or translated into provider-specific routing.
|
|
711
|
+
|
|
712
|
+
* `flex` — suitable for batch-tolerant workloads.
|
|
713
|
+
* `priority` — suitable for latency-sensitive workloads.
|
|
714
|
+
|
|
715
|
+
---
|
|
716
|
+
## Streaming and Events
|
|
717
|
+
|
|
718
|
+
`AssistantMessageEventStream` implements `AsyncIterable<StreamEvent>`.
|
|
719
|
+
|
|
720
|
+
It provides:
|
|
721
|
+
|
|
722
|
+
```ts
|
|
723
|
+
.result(): Promise<AgentResponse>
|
|
724
|
+
.push(...)
|
|
725
|
+
.end(...)
|
|
726
|
+
.fail(...)
|
|
727
|
+
.on(...)
|
|
728
|
+
.off(...)
|
|
729
|
+
.cancel()
|
|
730
|
+
.onCancel(...)
|
|
731
|
+
.isCancelled()
|
|
732
|
+
```
|
|
733
|
+
|
|
734
|
+
### StreamEvent
|
|
735
|
+
|
|
736
|
+
```ts
|
|
737
|
+
StreamEvent {
|
|
738
|
+
type,
|
|
739
|
+
delta?,
|
|
740
|
+
thinkingDelta?,
|
|
741
|
+
toolCall?,
|
|
742
|
+
toolResult?,
|
|
743
|
+
subagent?,
|
|
744
|
+
usage?,
|
|
745
|
+
responseId?,
|
|
746
|
+
finishReason?,
|
|
747
|
+
error?,
|
|
748
|
+
partialText?,
|
|
749
|
+
partialThinking?,
|
|
750
|
+
raw?
|
|
751
|
+
}
|
|
752
|
+
```
|
|
753
|
+
|
|
754
|
+
Event types:
|
|
755
|
+
|
|
756
|
+
```text
|
|
757
|
+
start
|
|
758
|
+
text_start
|
|
759
|
+
text_delta
|
|
760
|
+
text_end
|
|
761
|
+
thinking_start
|
|
762
|
+
thinking_delta
|
|
763
|
+
thinking_end
|
|
764
|
+
tool_call_start
|
|
765
|
+
tool_call_delta
|
|
766
|
+
tool_call_complete
|
|
767
|
+
tool_result
|
|
768
|
+
subagent_complete
|
|
769
|
+
usage
|
|
770
|
+
done
|
|
771
|
+
error
|
|
772
|
+
```
|
|
773
|
+
|
|
774
|
+
Emitted in practice: `start`, `text_delta`, `thinking_delta`, `tool_call_complete`, `tool_result`, `subagent_complete`, `usage`, `done`, `error`. The rest exist in the type for forward compatibility.
|
|
775
|
+
|
|
776
|
+
Use:
|
|
777
|
+
|
|
778
|
+
* `text_delta` for answer output.
|
|
779
|
+
* `thinking_delta` for reasoning output.
|
|
780
|
+
* `tool_call_complete` / `tool_result` for tool progress.
|
|
781
|
+
* `subagent_complete` for each completed worker during streaming multi-agent runs.
|
|
782
|
+
* `usage` for interim usage counts.
|
|
783
|
+
* `done` for final usage and completion information.
|
|
784
|
+
|
|
785
|
+
Breaking out of `for await` auto-cancels. Cancellation joins your
|
|
786
|
+
`AbortSignal` and stream cancellation into one linked controller, so
|
|
787
|
+
`controller.abort()` and `stream.cancel()` both stop the HTTP request.
|
|
788
|
+
Abort rejects with `AbortError` (never retried) and never resolves partial
|
|
789
|
+
results. `agent.run(prompt, { stream: true, signal })` exposes the same
|
|
790
|
+
`cancel` handle.
|
|
791
|
+
|
|
792
|
+
`SSEParser` provides:
|
|
793
|
+
|
|
794
|
+
```ts
|
|
795
|
+
feed(chunk)
|
|
796
|
+
flush()
|
|
797
|
+
```
|
|
798
|
+
|
|
799
|
+
It parses provider SSE streams and is only needed when implementing custom transports.
|
|
800
|
+
|
|
801
|
+
---
|
|
802
|
+
## Response, Usage, and Raw Data
|
|
803
|
+
|
|
804
|
+
### AgentResponse
|
|
805
|
+
|
|
806
|
+
```ts
|
|
807
|
+
AgentResponse {
|
|
808
|
+
text,
|
|
809
|
+
thinking?,
|
|
810
|
+
thoughtSignature?,
|
|
811
|
+
thinkingSignature?,
|
|
812
|
+
textSignature?,
|
|
813
|
+
toolCalls,
|
|
814
|
+
toolResults,
|
|
815
|
+
subagents,
|
|
816
|
+
usage,
|
|
817
|
+
responseId?,
|
|
818
|
+
model,
|
|
819
|
+
provider,
|
|
820
|
+
finishReason?,
|
|
821
|
+
durationMs,
|
|
822
|
+
raw,
|
|
823
|
+
turns
|
|
824
|
+
}
|
|
825
|
+
```
|
|
826
|
+
|
|
827
|
+
`AgentResponse` also provides:
|
|
828
|
+
|
|
829
|
+
```ts
|
|
830
|
+
.toString()
|
|
831
|
+
.toJSON()
|
|
832
|
+
```
|
|
833
|
+
|
|
834
|
+
### SubAgentExecutionMetadata
|
|
835
|
+
|
|
836
|
+
```ts
|
|
837
|
+
SubAgentExecutionMetadata {
|
|
838
|
+
name,
|
|
839
|
+
role?,
|
|
840
|
+
task,
|
|
841
|
+
model,
|
|
842
|
+
provider,
|
|
843
|
+
durationMs,
|
|
844
|
+
usage,
|
|
845
|
+
turns,
|
|
846
|
+
finishReason?,
|
|
847
|
+
responseId?,
|
|
848
|
+
text,
|
|
849
|
+
thinking?,
|
|
850
|
+
toolCalls?,
|
|
851
|
+
raw?,
|
|
852
|
+
isError?,
|
|
853
|
+
error?
|
|
854
|
+
}
|
|
855
|
+
```
|
|
856
|
+
|
|
857
|
+
### TokenUsage
|
|
858
|
+
|
|
859
|
+
```ts
|
|
860
|
+
TokenUsage {
|
|
861
|
+
inputTokens,
|
|
862
|
+
outputTokens,
|
|
863
|
+
totalTokens,
|
|
864
|
+
cachedTokens?,
|
|
865
|
+
cacheReadTokens?,
|
|
866
|
+
cacheWriteTokens?,
|
|
867
|
+
thinkingTokens?,
|
|
868
|
+
cost?: {
|
|
869
|
+
inputCost?,
|
|
870
|
+
outputCost?,
|
|
871
|
+
cacheReadCost?,
|
|
872
|
+
cacheWriteCost?,
|
|
873
|
+
totalCost?
|
|
874
|
+
}
|
|
875
|
+
}
|
|
876
|
+
```
|
|
877
|
+
|
|
878
|
+
Cost calculation prefers provider-reported totals. When unavailable, costs are computed using catalog pricing.
|
|
879
|
+
|
|
880
|
+
Subagent usage is rolled into the parent totals.
|
|
881
|
+
### ProviderRawData
|
|
882
|
+
|
|
883
|
+
```ts
|
|
884
|
+
ProviderRawData {
|
|
885
|
+
request: {
|
|
886
|
+
url,
|
|
887
|
+
method,
|
|
888
|
+
headers,
|
|
889
|
+
body
|
|
890
|
+
},
|
|
891
|
+
response?: {
|
|
892
|
+
status,
|
|
893
|
+
statusText,
|
|
894
|
+
headers,
|
|
895
|
+
body
|
|
896
|
+
}
|
|
897
|
+
}
|
|
898
|
+
```
|
|
899
|
+
|
|
900
|
+
Sensitive keys are redacted.
|
|
901
|
+
|
|
902
|
+
---
|
|
903
|
+
## Errors
|
|
904
|
+
|
|
905
|
+
Provider failures throw `AgentAccelProviderError` — one actionable line,
|
|
906
|
+
never a wire dump:
|
|
907
|
+
|
|
908
|
+
```text
|
|
909
|
+
[openrouter/nvidia/nemotron-3.5-lightning:free] request failed (404): No endpoints found that support input video
|
|
910
|
+
```
|
|
911
|
+
|
|
912
|
+
Raw details (`statusCode`, `url`, truncated body) stay attached as
|
|
913
|
+
non-enumerable properties: available programmatically, invisible in runtime
|
|
914
|
+
dumps. URL query strings are stripped (keys sometimes live there); headers
|
|
915
|
+
and cookies are never attached. Helpers:
|
|
916
|
+
|
|
917
|
+
* `toConciseProviderError(err, providerId, modelId)` — collapse any provider error.
|
|
918
|
+
* `assertModalitiesSupported(context, providerId, modelId)` — pre-request modality gate.
|
|
919
|
+
|
|
920
|
+
Examples print failures as exactly one line via `examples/_shared.ts`:
|
|
921
|
+
|
|
922
|
+
```ts
|
|
923
|
+
const res = await agent.run([...]).catch(fail);
|
|
924
|
+
// ✖ [provider/model] request failed (404): ...
|
|
925
|
+
// exit 1, no stack dump
|
|
926
|
+
```
|
|
927
|
+
|
|
928
|
+
---
|
|
929
|
+
## Model Catalog & Dynamic 12-Hour TTL Sync
|
|
930
|
+
|
|
931
|
+
Agent Accelerator features an adaptive, dynamically synchronized model catalog powered by [models.dev](https://models.dev). To avoid shipping a bloated 4.4 MB static JSON file with production builds, the SDK employs a high-performance **12-hour TTL local caching architecture**:
|
|
932
|
+
|
|
933
|
+
- **Automated 12-Hour Cache Validation**: When `.run()`, `.ask()`, or `.stream()` executes, the runtime checks `src/data/models-cache.json`. If the cache timestamp is within 12 hours, it reads from disk with zero network delay. When the TTL expires, it transparently synchronizes with `https://models.dev/api.json`.
|
|
934
|
+
- **Git & Package Safety**: The dynamic cache file (`src/data/models-cache.json`) is gitignored and excluded from production packages.
|
|
935
|
+
|
|
936
|
+
### Developer Catalog Controls
|
|
937
|
+
|
|
938
|
+
```ts
|
|
939
|
+
import {
|
|
940
|
+
refreshModelCatalog,
|
|
941
|
+
getCatalogStatus,
|
|
942
|
+
setCatalogTTL,
|
|
943
|
+
getModelFromCatalog,
|
|
944
|
+
getModelThinkingInfo,
|
|
945
|
+
} from "agent-accelerator";
|
|
946
|
+
|
|
947
|
+
// Force an immediate refresh from models.dev:
|
|
948
|
+
await refreshModelCatalog({ force: true });
|
|
949
|
+
|
|
950
|
+
// Customize default TTL (e.g. 24 hours):
|
|
951
|
+
setCatalogTTL(24 * 60 * 60 * 1000);
|
|
952
|
+
|
|
953
|
+
// Inspect cache health and metadata:
|
|
954
|
+
const status = getCatalogStatus();
|
|
955
|
+
console.log(status.modelCount, status.providerCount, status.isExpired);
|
|
956
|
+
```
|
|
957
|
+
|
|
958
|
+
### CLI Refresh
|
|
959
|
+
|
|
960
|
+
```bash
|
|
961
|
+
bun run update-models # Refresh model catalog (force or expired)
|
|
962
|
+
bun scripts/update-models.ts --force # Force re-download
|
|
963
|
+
bun scripts/update-models.ts --ttl=24h # Refresh with custom TTL
|
|
964
|
+
```
|
|
965
|
+
|
|
966
|
+
Upstream data lags on some inputs (e.g. `gpt-4o` accepts audio). Record
|
|
967
|
+
verified corrections in `MODALITY_OVERRIDES` in `src/models/catalog.ts`
|
|
968
|
+
(keyed `provider/model`, merged over the snapshot) — never edit the JSON
|
|
969
|
+
directly, a refresh would wipe it.
|
|
970
|
+
|
|
971
|
+
### Catalog Helpers
|
|
972
|
+
|
|
973
|
+
* `getModelFromCatalog(provider, modelId)` → `ModelSpec | undefined`, including alias and global fallback handling. Use it for context-window, pricing, and modality checks.
|
|
974
|
+
* `getModelsForProvider(provider)` → returns `ModelSpec[]`.
|
|
975
|
+
* `getModelThinkingInfo(provider, modelId)` → inspect permitted thinking levels.
|
|
976
|
+
* `refreshModelCatalog(options?)` → programmatically sync with upstream models.dev.
|
|
977
|
+
* `getCatalogStatus()` → diagnostic info on active catalog, provider count, and TTL expiry.
|
|
978
|
+
|
|
979
|
+
### ModelSpec
|
|
980
|
+
|
|
981
|
+
```ts
|
|
982
|
+
ModelSpec {
|
|
983
|
+
id,
|
|
984
|
+
provider,
|
|
985
|
+
name,
|
|
986
|
+
description?,
|
|
987
|
+
family?,
|
|
988
|
+
api?,
|
|
989
|
+
contextWindow,
|
|
990
|
+
maxOutputTokens,
|
|
991
|
+
limit?,
|
|
992
|
+
cost?,
|
|
993
|
+
modalities?,
|
|
994
|
+
reasoning?,
|
|
995
|
+
reasoning_options?,
|
|
996
|
+
tool_call?,
|
|
997
|
+
capabilities,
|
|
998
|
+
pricing?,
|
|
999
|
+
raw?
|
|
1000
|
+
}
|
|
1001
|
+
```
|
|
1002
|
+
|
|
1003
|
+
`contextWindow` and `maxOutputTokens` mirror:
|
|
1004
|
+
|
|
1005
|
+
```ts
|
|
1006
|
+
limit.context
|
|
1007
|
+
limit.output
|
|
1008
|
+
```
|
|
1009
|
+
|
|
1010
|
+
### ModelCapabilities
|
|
1011
|
+
|
|
1012
|
+
```ts
|
|
1013
|
+
ModelCapabilities {
|
|
1014
|
+
supportsThinking?,
|
|
1015
|
+
supportsThinkingBudget?,
|
|
1016
|
+
supportsThinkingLevel?,
|
|
1017
|
+
supportsImplicitCaching?,
|
|
1018
|
+
supportsExplicitCaching?,
|
|
1019
|
+
supportsLongCacheRetention?,
|
|
1020
|
+
supportsParallelToolCalls?,
|
|
1021
|
+
supportsStreaming?,
|
|
1022
|
+
modalities?,
|
|
1023
|
+
supportsReasoningToggle?,
|
|
1024
|
+
supportsReasoningEffort?
|
|
1025
|
+
}
|
|
1026
|
+
```
|
|
1027
|
+
|
|
1028
|
+
When estimated input exceeds the available context budget, `Agent` trims the oldest middle history using:
|
|
1029
|
+
|
|
1030
|
+
```text
|
|
1031
|
+
context * 0.9 - maxOutput
|
|
1032
|
+
```
|
|
1033
|
+
|
|
1034
|
+
The beginning of the conversation and the most recent tail are preserved.
|
|
1035
|
+
|
|
1036
|
+
---
|
|
1037
|
+
## Messages and Media
|
|
1038
|
+
|
|
1039
|
+
### Message
|
|
1040
|
+
|
|
1041
|
+
```ts
|
|
1042
|
+
Message {
|
|
1043
|
+
role: "system" | "user" | "assistant" | "tool",
|
|
1044
|
+
content: string | ContentPart[],
|
|
1045
|
+
name?,
|
|
1046
|
+
thoughtSignature?
|
|
1047
|
+
}
|
|
1048
|
+
```
|
|
1049
|
+
|
|
1050
|
+
### ProviderContext
|
|
1051
|
+
|
|
1052
|
+
```ts
|
|
1053
|
+
ProviderContext {
|
|
1054
|
+
systemPrompt?,
|
|
1055
|
+
messages,
|
|
1056
|
+
tools?,
|
|
1057
|
+
cachedContentId?
|
|
1058
|
+
}
|
|
1059
|
+
```
|
|
1060
|
+
|
|
1061
|
+
### ContentPart
|
|
1062
|
+
|
|
1063
|
+
Supported content parts include:
|
|
1064
|
+
|
|
1065
|
+
* `TextPart { type:"text", text, thoughtSignature? }` — plain text.
|
|
1066
|
+
* `ThinkingPart { type:"thinking", thinking, thoughtSignature? }` — preserved reasoning.
|
|
1067
|
+
* `ToolCallPart { type:"tool_call", id, name, arguments, rawArguments?, thoughtSignature? }`
|
|
1068
|
+
* `ToolResultPart { type:"tool_result", id, name, result, isError? }`
|
|
1069
|
+
* `ImagePart { type:"image", image, mimeType? }`
|
|
1070
|
+
* `AudioPart { type:"audio", audio, mimeType? }`
|
|
1071
|
+
* `VideoPart { type:"video", video, mimeType? }`
|
|
1072
|
+
* `FilePart { type:"file", file, mimeType?, filename? }` — PDFs/documents.
|
|
1073
|
+
|
|
1074
|
+
`image`, `audio`, `video`, and `file` inputs accept:
|
|
1075
|
+
|
|
1076
|
+
* data URLs
|
|
1077
|
+
* remote `http(s)` URLs
|
|
1078
|
+
* local paths
|
|
1079
|
+
* raw base64
|
|
1080
|
+
* `Uint8Array`
|
|
1081
|
+
* `ArrayBuffer`
|
|
1082
|
+
|
|
1083
|
+
`normalizeMediaInput(input, mime?)` returns:
|
|
1084
|
+
|
|
1085
|
+
```ts
|
|
1086
|
+
{
|
|
1087
|
+
mimeType,
|
|
1088
|
+
base64Data,
|
|
1089
|
+
dataUrl
|
|
1090
|
+
}
|
|
1091
|
+
```
|
|
1092
|
+
|
|
1093
|
+
`inferMimeType(path)` infers the MIME type from a file extension.
|
|
1094
|
+
|
|
1095
|
+
Media normalization is handled automatically by providers (mapped to Vercel V4 `file` parts). Thinking traces are echoed in follow-up turns only where accepted — strict endpoints receive tool calls without `reasoning_content`.
|
|
1096
|
+
|
|
1097
|
+
Provider matrix: text + image + wav/mp3 audio + PDF work on all supported
|
|
1098
|
+
providers. Video works on Gemini only.
|
|
1099
|
+
|
|
1100
|
+
Before any network call, the executor checks `image` / `audio` / `video` /
|
|
1101
|
+
`file`-as-PDF parts against the catalog's `modalities.input` for that
|
|
1102
|
+
provider/model and fails fast with a one-line error naming the gap
|
|
1103
|
+
(`assertModalitiesSupported`). Models absent from the catalog (custom
|
|
1104
|
+
providers, dynamic routers) are skipped — the provider endpoint decides and
|
|
1105
|
+
its verdict surfaces as a concise error (see [Errors](#errors)). Verified
|
|
1106
|
+
catalog corrections live in `MODALITY_OVERRIDES` (`src/models/catalog.ts`),
|
|
1107
|
+
never in the gitignored snapshot.
|
|
1108
|
+
|
|
1109
|
+
---
|
|
1110
|
+
## Tokens
|
|
1111
|
+
|
|
1112
|
+
Token counting is heuristic only.
|
|
1113
|
+
|
|
1114
|
+
Actual billing information comes from `res.usage`.
|
|
1115
|
+
|
|
1116
|
+
Available helpers:
|
|
1117
|
+
|
|
1118
|
+
* `countTokens(string | Message[] | ProviderContext)` — useful for preflight sizing and history trimming.
|
|
1119
|
+
* `estimateTokensFromText(text)`
|
|
1120
|
+
* `estimateTokensFromMessage(msg)`
|
|
1121
|
+
* `estimateTokensFromPart(part)`
|
|
1122
|
+
|
|
1123
|
+
---
|
|
1124
|
+
## Utils
|
|
1125
|
+
|
|
1126
|
+
* `createSessionId(prefix="accel")` — creates a UUID-based session ID clamped to 64 characters for affinity.
|
|
1127
|
+
* `getEnv(key, fallback?)` — resolves an environment variable.
|
|
1128
|
+
* `getApiKey(provider, explicit?, env?)` — resolves API keys using the configured environment lookup order.
|
|
1129
|
+
* `buildSessionHeaders(provider, cache?, custom?, sessionId?)` — builds provider-specific affinity headers. Normally handled automatically.
|
|
1130
|
+
* `z` — re-exported from Zod so tools do not require a separate Zod import.
|
|
1131
|
+
|
|
1132
|
+
---
|
|
1133
|
+
## Examples
|
|
1134
|
+
|
|
1135
|
+
```bash
|
|
1136
|
+
bun run examples/06-chat.ts
|
|
1137
|
+
# persistent CLI, /model "..." /level <lvl> /help /exit
|
|
1138
|
+
|
|
1139
|
+
bun run examples/05-sub-agents.ts
|
|
1140
|
+
# fixed researcher + critic pipeline
|
|
1141
|
+
|
|
1142
|
+
bun run examples/04-multi_agent.ts
|
|
1143
|
+
# dynamic spawn_subagents demo
|
|
1144
|
+
|
|
1145
|
+
bun run examples/01-metadata.ts
|
|
1146
|
+
# usage + raw inspection
|
|
1147
|
+
|
|
1148
|
+
bun run examples/02-function_calling.ts
|
|
1149
|
+
# single tool call
|
|
1150
|
+
|
|
1151
|
+
bun run examples/03-multimodal_image.ts [./photo.png]
|
|
1152
|
+
# image input (remote URL default, local path optional)
|
|
1153
|
+
|
|
1154
|
+
bun run examples/07-multimodal_audio.ts
|
|
1155
|
+
# audio input
|
|
1156
|
+
|
|
1157
|
+
bun run examples/08-multimodal_document.ts
|
|
1158
|
+
# PDF/file input
|
|
1159
|
+
|
|
1160
|
+
bun run examples/09-multimodal_video.ts
|
|
1161
|
+
# video input (video-capable model required)
|
|
1162
|
+
```
|
|
1163
|
+
|
|
1164
|
+
Every example ends its `run()` with `.catch(fail)` (`examples/_shared.ts`),
|
|
1165
|
+
so failures print one line and exit `1` — no stack dumps.
|
|
1166
|
+
|
|
1167
|
+
The chat example persists conversations to:
|
|
1168
|
+
|
|
1169
|
+
```text
|
|
1170
|
+
.session.jsonl
|
|
1171
|
+
```
|
|
1172
|
+
|
|
1173
|
+
It resumes from that file on next launch.
|
|
1174
|
+
|
|
1175
|
+
It also displays per-turn:
|
|
1176
|
+
|
|
1177
|
+
```text
|
|
1178
|
+
input / output / cached / cost / context %
|
|
1179
|
+
```
|
|
1180
|
+
|
|
1181
|
+
---
|
|
1182
|
+
## Scripts and Structure
|
|
1183
|
+
|
|
1184
|
+
### Scripts
|
|
1185
|
+
|
|
1186
|
+
```bash
|
|
1187
|
+
bun run typecheck # tsc --noEmit
|
|
1188
|
+
bun test # bun test test/
|
|
1189
|
+
bun run update-models # refresh model catalog cache (supports --force, --ttl=24h)
|
|
1190
|
+
```
|
|
1191
|
+
### Project Structure
|
|
1192
|
+
|
|
1193
|
+
```text
|
|
1194
|
+
src/
|
|
1195
|
+
├── agent/ # Agent, context, loop, delegation, subagent
|
|
1196
|
+
├── ai-sdk/ # Provider backend: provider factories, converters, executor, registry
|
|
1197
|
+
├── models/ # Dynamic catalog cache, parser, verified overrides
|
|
1198
|
+
├── data/ # Dynamic model catalog cache (gitignored, excluded from bundle)
|
|
1199
|
+
├── tools/ # tool(), schema, executor
|
|
1200
|
+
├── streaming/ # event stream, SSE parser
|
|
1201
|
+
├── tokens/ # estimator
|
|
1202
|
+
├── types/ # agent, core, message, model, response, tool
|
|
1203
|
+
└── utils/ # base64, cache, env, headers, media, serialization, session
|
|
1204
|
+
examples/
|
|
1205
|
+
├── 01-metadata.ts 02-function_calling.ts 03-multimodal_image.ts
|
|
1206
|
+
├── 04-multi_agent.ts 05-sub-agents.ts 06-chat.ts
|
|
1207
|
+
├── 07-multimodal_audio.ts 08-multimodal_document.ts
|
|
1208
|
+
├── 09-multimodal_video.ts _shared.ts
|
|
1209
|
+
test/ # 105 tests mirroring the above
|
|
1210
|
+
```
|
|
1211
|
+
|
|
1212
|
+
---
|
|
1213
|
+
## License
|
|
1214
|
+
|
|
1215
|
+
MIT License — see [LICENSE](file:///data/projects/Agent-Accelerator/LICENSE) for details.
|
|
1216
|
+
|
|
1217
|
+
Copyright (c) 2026 Sashvat Bharat.
|
|
1218
|
+
|
|
1219
|
+
---
|
|
1220
|
+
|