agent-accelerator 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/README.md +83 -112
  2. package/{SYSTEM_PROMPT_AGENT.md → examples/prompts/SYSTEM_PROMPT_AGENT.md} +1 -1
  3. package/package.json +14 -9
  4. package/src/agent/agent.ts +191 -71
  5. package/src/agent/context.ts +1 -0
  6. package/src/agent/delegation.ts +61 -12
  7. package/src/agent/loop.ts +71 -48
  8. package/src/data/README.md +6 -6
  9. package/src/index.ts +130 -45
  10. package/src/models/catalog-cache.ts +60 -9
  11. package/src/models/catalog.ts +53 -7
  12. package/src/providers/google.ts +926 -0
  13. package/src/providers/openai-compat.ts +1147 -0
  14. package/src/providers/openai.ts +959 -0
  15. package/src/providers/openrouter-responses.ts +949 -0
  16. package/src/providers/openrouter.ts +1037 -0
  17. package/src/{ai-sdk → providers}/registry.ts +43 -55
  18. package/src/providers.ts +490 -0
  19. package/src/streaming/sse-parser.ts +6 -4
  20. package/src/tools/executor.ts +19 -6
  21. package/src/tools/schema.ts +21 -11
  22. package/src/types/agent.ts +8 -1
  23. package/src/types/core.ts +1 -1
  24. package/src/types/message.ts +5 -0
  25. package/src/types/model.ts +3 -7
  26. package/src/types/provider-payloads.ts +2 -84
  27. package/src/types/tool.ts +6 -0
  28. package/src/update-models.ts +56 -0
  29. package/src/utils/cache.ts +1 -1
  30. package/src/utils/documents.ts +517 -0
  31. package/src/utils/env.ts +0 -7
  32. package/src/{ai-sdk → utils}/errors.ts +61 -2
  33. package/src/utils/headers.ts +10 -20
  34. package/src/utils/media.ts +5 -2
  35. package/src/utils/retry.ts +89 -0
  36. package/src/utils/serialization.ts +15 -0
  37. package/src/ai-sdk/converters.ts +0 -342
  38. package/src/ai-sdk/executor.ts +0 -454
  39. package/src/ai-sdk/index.ts +0 -55
  40. package/src/ai-sdk/model-provider.ts +0 -303
  41. package/src/ai-sdk/options.ts +0 -306
  42. package/src/ai-sdk/provider.ts +0 -415
  43. package/src/tokens/counter.ts +0 -136
  44. /package/{SYSTEM_PROMPT.md → examples/prompts/SYSTEM_PROMPT.md} +0 -0
  45. /package/{SYSTEM_PROMPT_TOOLS.md → examples/prompts/SYSTEM_PROMPT_TOOLS.md} +0 -0
package/README.md CHANGED
@@ -1,13 +1,14 @@
1
1
  # Agent Accelerator
2
2
 
3
+ [![npm version](https://img.shields.io/npm/v/agent-accelerator.svg?color=cb3837&logo=npm)](https://www.npmjs.com/package/agent-accelerator)
3
4
  [![GitHub](https://img.shields.io/badge/GitHub-sashvat--bharat%2Fagent--accelerator-blue?logo=github)](https://github.com/sashvat-bharat/agent-accelerator)
4
5
  [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
5
6
 
6
- A thin, typed transport SDK for calling LLMs through a single `Agent` interface.
7
+ > **Stability notice:** Agent Accelerator is pre-1.0 and under active development. Public APIs, provider transports, and configuration options may still change in breaking ways between releases. For production use, please pin your dependency to an exact version.
7
8
 
8
- GitHub: [https://github.com/sashvat-bharat/agent-accelerator](https://github.com/sashvat-bharat/agent-accelerator)
9
+ A thin, typed transport SDK for calling LLMs through a single `Agent` interface.
9
10
 
10
- Agent Accelerator supports `google`, `opencode`, `openrouter`, `openai`, and any OpenAI-compatible endpoint using `{PREFIX}_API_KEY` and `{PREFIX}_BASE_URL`.
11
+ Agent Accelerator supports `google`, `openrouter`, `openai`, and any OpenAI-compatible endpoint using `{PREFIX}_API_KEY` and `{PREFIX}_BASE_URL`.
11
12
 
12
13
 
13
14
  ```ts
@@ -45,7 +46,7 @@ yarn add agent-accelerator
45
46
 
46
47
  Requires `bun` or `node 22+`.
47
48
 
48
- Core dependencies are `zod`, `ai`, and selected `@ai-sdk/*` provider packages used as the backend transport.
49
+ Core dependencies are `zod` only. Document conversion additionally uses the optional `@firecrawl/anydoc` peer. All provider transports are native REST (`fetch`, no SDK).
49
50
 
50
51
  ---
51
52
 
@@ -53,7 +54,6 @@ Core dependencies are `zod`, `ai`, and selected `@ai-sdk/*` provider packages us
53
54
 
54
55
  ```bash
55
56
  GEMINI_API_KEY=...
56
- OPENCODE_API_KEY=...
57
57
  OPENROUTER_API_KEY=...
58
58
  OPENAI_API_KEY=...
59
59
  # or OPENAI_BASE_API_KEY=...
@@ -129,6 +129,7 @@ An `Agent` holds :
129
129
  | `thinkingLevel` | `ThinkingLevel` | `none \| dynamic \| minimal \| low \| medium \| high \| xhigh`. Validated against the model catalog before a request. |
130
130
  | `cache` | `CacheConfig` | `{ retention, sessionId, cachedContentId, ttlSeconds }`. Controls cache reuse. |
131
131
  | `serviceTier` | `"flex" \| "priority"` | Cost / priority routing where supported. Omit for standard routing. |
132
+ | `bypassInputFileModality` | `boolean` | When `true`, `file` parts are converted client-side to Markdown for models lacking native support. Capable models still receive files natively. Defaults to `false`. |
132
133
  | `maxTurns` | `number` | Maximum model → tool → model loops per `run`. Defaults to `10`. |
133
134
  | `sessionId` | `string` | Stable identifier used for cache affinity. Auto-generated when omitted. |
134
135
  | `headers` | `Record<string,string>` | Additional headers merged into every request. |
@@ -327,11 +328,12 @@ const agent = new Agent({
327
328
 
328
329
  ### Tool Helpers
329
330
 
330
- * `tool({ name?, description, input?, parameters?, strict?, timeoutMs?, maxTries?, maxConcurrency?, execute })` — creates a `ToolDefinition`. Provide either a Zod `input` schema or raw JSON `parameters`. `execute(input, ctx)` may return arbitrary values; results are converted safely for model context. `timeoutMs` defaults to `0` (no time limit) and is a best-effort event-loop deadline; `ctx.signal` enables cooperative cancellation. `maxTries` accepts a number or numeric string; a positive value is the total attempt limit, while `0`/omitted means no configured limit for transient retries. `maxConcurrency` limits simultaneous calls for that tool (default pool: 8).
331
+ * `tool({ name?, description, input?, parameters?, strict?, timeoutMs?, maxTries?, maxConcurrency?, execute })` — creates a `ToolDefinition`. Provide either a Zod `input` schema or raw JSON `parameters`. `execute(input, ctx)` may return arbitrary values; results are converted safely for model context. `timeoutMs` defaults to `0` (no time limit) and is a best-effort event-loop deadline; `ctx.signal` enables cooperative cancellation. `maxTries` accepts a number or numeric string; a positive value is the total attempt limit, while `0`/omitted falls back to 3 attempts. `maxConcurrency` limits simultaneous calls for that tool (default pool: 8).
331
332
  * `toStandardToolDeclarations(record | array)` — converts tools to `{ name, description, parameters }` for providers.
332
333
  * `zodToJsonSchema(schema)` — converts Zod schemas using native conversion with a fallback extractor.
333
334
  * `cleanJsonSchema(schema)` — removes `$schema`, `$defs`, and `definitions`, resolves `$ref`, and preserves explicit `additionalProperties`.
334
335
  * `executeToolCalls({ tools, toolCalls, agentName?, parallel?, signal?, sessionId? })` — executes tool calls. Calls validate Zod input, use bounded parallelism, enforce per-tool timeouts, retry transient failures, recover namespace/camel-case tool-name aliases, and return `ToolResultRecord[]`. Missing tools and aborts are represented as error results rather than thrown.
336
+ * `convert_document_to_markdown` — built-in tool that converts document files (PDF/Word/PowerPoint/Excel/OpenDocument/RTF/EPUB/CSV, local path or URL) to Markdown for models without native document parsing. Requires the optional `@firecrawl/anydoc` peer.
335
337
 
336
338
  The agent loop also blocks an identical tool name and argument set when the model requests it in the immediately following turn. The synthetic error is returned to the model so it can reuse the prior result or change its arguments.
337
339
 
@@ -386,7 +388,8 @@ for (const result of response.toolResults) {
386
388
  name,
387
389
  arguments,
388
390
  rawArguments?,
389
- thoughtSignature?
391
+ thoughtSignature?,
392
+ callId?,
390
393
  }
391
394
  ```
392
395
 
@@ -408,7 +411,6 @@ for (const result of response.toolResults) {
408
411
  First-class provider IDs are:
409
412
 
410
413
  * `google`
411
- * `opencode`
412
414
  * `openrouter`
413
415
  * `openai`
414
416
 
@@ -417,9 +419,6 @@ Any other provider prefix is treated as an OpenAI-compatible custom provider.
417
419
  ### Model Strings
418
420
 
419
421
  * `"google/<id>"` → Google AI Studio at `generativelanguage.googleapis.com`.
420
- * `"opencode/<id>"` → OpenCode Zen at `opencode.ai/zen/v1`.
421
- * `"opencode-go/<id>"` → OpenCode Go endpoint.
422
- * OpenCode automatically selects Chat vs Responses API. Claude models and models using `api=openai-responses` use Responses.
423
422
  * `"openrouter/<scope>/<model>"` or `"scope/model:variant"` → OpenRouter.
424
423
  * `"openai/<id>"` → OpenAI.
425
424
  * `"groq/<id>"`, `"ollama/<id>"`, etc. → custom providers using `{PREFIX}_API_KEY` and `{PREFIX}_BASE_URL`. The API key is optional for local endpoints.
@@ -437,7 +436,7 @@ An unknown prefix with no catalog match still routes to a custom provider, allow
437
436
  * `getProvider("groq")` → returns a cached provider and automatically creates a custom provider for unknown prefixes.
438
437
  * `ensureCustomProvider(prefix, { baseUrl?, apiKey?, name? })` → gets or creates a custom provider. Useful for multiple endpoints in one process.
439
438
  * `normalizeProviderPrefix(s)` → lowercases and trims a provider prefix.
440
- * `ModelProvider.GoogleGenAI(model, apiKey?, { thinkingLevel?, baseUrl? })` — same shape is available for `.OpenCode`, `.OpenRouter`, `.OpenAI`, `.Custom`, `.OpenAICompatible`, and `.Generic`.
439
+ * `ModelProvider.GoogleGenAI(model, apiKey?, { thinkingLevel?, baseUrl? })` — same shape is available for `.OpenRouter`, `.OpenAI`, `.Custom`, `.OpenAICompatible`, and `.Generic`.
441
440
 
442
441
  Each returns a `ModelProviderInstance`:
443
442
 
@@ -472,93 +471,82 @@ new Agent({
472
471
  });
473
472
  ```
474
473
 
475
- `BaseProvider` is the abstract provider base containing `id`, `name`, `models`, `getModel`, `generate`, `stream`, and `countTokens`.
474
+ `Provider` is the provider contract containing `id`, `name`, `models`, `getModel`, `generate`, and `stream`.
476
475
 
477
- Its catalog-first `getModel` behavior includes a hardcoded fallback.
476
+ Its catalog-first `getModel` behavior includes a permissive fallback (`createGenericModelSpec`) so private model IDs never fail preflight.
478
477
 
479
- Extend `BaseProvider` when implementing a fully custom transport.
478
+ Implement the `Provider` interface when adding a fully custom transport.
480
479
 
481
480
  `ProviderRequestOptions { apiKey?, baseUrl?, headers?, thinking?, cache?, serviceTier?, tools?, toolChoice?, signal?, sessionId?, env? }` is the per-call options object passed to `generate`/`stream`. Agent builds it automatically.
482
481
 
483
482
  ### Google
484
483
 
485
- `GoogleAIStudioProvider` and `GOOGLE_MODELS` provide Google AI Studio support.
484
+ `GoogleInteractionsProvider` (aliased as `GoogleAIStudioProvider`) and `GOOGLE_MODELS` provide Google support over the Interactions API (`POST {baseUrl}/interactions`, streaming via `?alt=sse`).
486
485
 
487
- The catalog is filtered with fallback support for lite and flash models.
486
+ Turns chain statefully through `previous_interaction_id` per session, falling back to stateless full-history sends when the session switches providers or models.
488
487
 
489
- Google models matching `gemini-1.x` or `gemini-2.x` are rejected.
490
-
491
- Thinking levels map to:
488
+ Thinking levels map to `thinking_level`, with `thinking_summaries` enabled whenever thinking is active:
492
489
 
493
490
  ```text
494
- OFF | MINIMAL | LOW | MEDIUM | HIGH
491
+ minimal | low | medium | high (xhigh clamps to high; none is unsupported and omitted; dynamic omits)
495
492
  ```
496
493
 
497
494
  Google-specific handling includes:
498
495
 
499
496
  * Isolating `thoughtSignature` per part for prefix stability.
500
497
  * Converting JSON Schema to OpenAPI 3.0 through `stripSchemaForGoogle`.
501
- * Supporting explicit `cachedContents` when `retention != implicit` and the prompt is large.
498
+ * Implicit/automatic caching only — explicit retention and `cachedContentId` are warned about and dropped (the Interactions API defines no explicit cache primitives).
502
499
 
503
500
  Helpers:
504
501
 
505
502
  * `createExplicitCache({ model, systemInstruction?, contents?, tools?, displayName?, ttlSeconds?, expireTime?, apiKey?, baseUrl? })`
506
503
  * `isValidThoughtSignature(sig)`
507
504
  * `retainThoughtSignature(existing, incoming)`
505
+ * `extractGoogleThoughtSignature(obj)`
508
506
  * `stripSchemaForGoogle(schema)`
507
+ * `clearInteractionChains(sessionId?)`
509
508
 
510
509
  ### OpenAI
511
510
 
512
- `OpenAIProvider` and `OPENAI_MODELS` provide OpenAI support.
511
+ `OpenAIResponsesProvider` (aliased as `OpenAIProvider`) and `OPENAI_MODELS` provide OpenAI support over the Responses API (`POST {baseUrl}/responses`).
513
512
 
514
- `OpenAIProvider` is also the base class for OpenAI-wire transports.
513
+ Every turn is stateless: the full history is sent explicitly with `store: false`, so no `previous_response_id`, `background`, or conversation chaining is ever used.
515
514
 
516
- Thinking levels map to:
515
+ Thinking levels map to `reasoning.effort` verbatim:
517
516
 
518
517
  ```text
519
- reasoning_effort: low | medium | high
518
+ none | minimal | low | medium | high | xhigh (dynamic omits — server default)
520
519
  ```
521
520
 
522
521
  It also:
523
522
 
524
- * passes `service_tier`;
525
- * derives `max_completion_tokens` from the configured budget;
526
- * preserves Google thought signatures through `extra_content.google.thought_signature`;
527
- * exposes `extractGoogleThoughtSignature(obj)`.
528
-
529
- ### OpenCode
530
-
531
- `OpenCodeProvider` and `OPENCODE_MODELS` provide OpenCode support.
532
-
533
- Requests include:
534
-
535
- * `session_id`
536
- * `prompt_cache_key`
537
- * headers used for sticky routing
538
-
539
- The provider handles both Chat and Responses payloads.
540
-
541
- Transient `500`, `502`, `503`, and `529` errors are retried with thinking disabled, followed by an attempt using the alternate endpoint.
523
+ * passes `service_tier` (`flex` | `priority`);
524
+ * pins cache affinity with `prompt_cache_key` (clamped to 64 chars);
525
+ * rejects `video` parts up front with a one-line error (the Responses wire carries text/image/file only);
526
+ * never sends `temperature`, token caps, or background/conversation fields.
542
527
 
543
528
  ### OpenRouter
544
529
 
545
- `OpenRouterProvider` and `OPENROUTER_MODELS` provide OpenRouter support.
530
+ `OpenRouterChatCompletionsProvider` (aliased as `OpenRouterProvider` / `OpenRouter`) and `OPENROUTER_MODELS` provide OpenRouter support over Chat Completions (`POST {baseUrl}/chat/completions`).
546
531
 
547
- Requests include:
532
+ Every turn is stateless: the full `messages[]` array is sent explicitly — no chaining primitives exist on this endpoint.
548
533
 
549
- * `session_id`
550
- * `prompt_cache_key`
534
+ Session affinity is a top-level body `session_id` plus the `x-session-id` header fallback (body takes precedence per OpenRouter docs).
551
535
 
552
- Thinking is mapped to:
536
+ Thinking levels map to `reasoning.effort` verbatim:
553
537
 
554
538
  ```text
555
- reasoning { effort | enabled }
556
- include_reasoning
539
+ none | minimal | low | medium | high | xhigh (dynamic omits — server default)
557
540
  ```
558
541
 
559
- based on catalog capabilities.
542
+ It also:
543
+
544
+ * passes `service_tier` (`flex` | `priority`);
545
+ * enables the `file-parser` plugin when file parts are present (remote file URLs are additionally surfaced in text);
546
+ * maps `tool_choice` (`auto` omitted; `none`, `required`, and function pins supported);
547
+ * reports `cached_tokens` / `cache_write_tokens` / reasoning tokens, preferring provider-reported `cost` when present.
560
548
 
561
- `serviceTier` is mapped to provider routing.
549
+ `serviceTier` values other than `flex` / `priority` are omitted (standard routing).
562
550
 
563
551
  ### Custom
564
552
 
@@ -588,7 +576,6 @@ Examples include:
588
576
 
589
577
  * `OpenAIChatCompletionRequest`
590
578
  * `GoogleGenerateContentRequest`
591
- * `OpenCodeChatRequest`
592
579
  * `OpenRouterChatRequest`
593
580
 
594
581
  ---
@@ -632,6 +619,7 @@ ThinkingConfig {
632
619
 
633
620
  * `getModelThinkingInfo(provider, modelId)` → `{ supportsThinking, reasoningOptions?, allowedLevels, supportsDisable, description }`. Useful for building UI selectors.
634
621
  * `validateModelThinking(provider, modelId, level)` → throws `ThinkingLevelError { provider, modelId, requestedLevel, allowedLevels, supportsThinking }` with a fix hint. Automatically called by `run` and `stream`. Catch it and use `err.allowedLevels` to offer valid options.
622
+ * `resolveEffectiveThinking(base, override?)` → resolves a per-run `thinkingLevel` override into a `ThinkingConfig` without mutating agent config.
635
623
  * `getModelsForProvider("google")` → returns known model specifications.
636
624
 
637
625
  Unknown models allow all thinking levels.
@@ -660,40 +648,40 @@ CacheRetention =
660
648
 
661
649
  ### Retention
662
650
 
663
- * `implicit` — automatic prefix reuse without an explicit cache object or storage fee.
664
- * `short` — approximately 5 minutes.
665
- * `medium` — approximately 1 hour.
666
- * `long` — approximately 12 hours for Google explicit caches and approximately 24 hours for prompt caching where supported.
651
+ * `implicit` — automatic prefix reuse without an explicit cache object or storage fee. This is the only retention the native adapters act on (by doing nothing special — stable prompts plus session affinity do the work).
652
+ * `short` / `medium` / `long` — TTL hints (`5m` / `1h` / `12h`) consumed only by the opt-in `createExplicitCache` helper. Provider adapters warn about and drop any non-`implicit` retention: none of the native transports expose retention control.
667
653
 
668
654
  ### Session Affinity
669
655
 
670
656
  `sessionId` pins provider affinity using mechanisms such as:
671
657
 
672
- * `x-session-id`
673
- * `x-opencode-session`
674
- * `prompt_cache_key` (openai/openrouter/opencode only — strict endpoints such as groq reject it, so custom providers get headers only)
658
+ * `x-session-id` (+ `x-client-request-id`) headers — all providers, best effort
659
+ * top-level body `session_id` — OpenRouter (takes precedence over the header)
660
+ * `prompt_cache_key` — OpenAI only (custom endpoints get headers only; strict ones reject unknown body fields)
675
661
 
676
662
  Reuse the same session ID across turns when cache affinity is desired.
677
663
 
678
- In browser runtimes only CORS-safe attribution headers are sent; affinity
679
- still flows via `promptCacheKey` provider options, never via headers.
664
+ In browser runtimes custom `x-*` headers are stripped to avoid CORS preflights: OpenAI keeps affinity through its body key, while custom endpoints lose affinity entirely on browsers.
680
665
 
681
666
  ### Explicit Caches
682
667
 
683
- `cachedContentId` reuses a previously created Google `cachedContents/...` resource.
668
+ `createExplicitCache` mints a Google `cachedContents/...` resource directly (TTL via `retention`/`ttlSeconds`).
669
+
670
+ `cachedContentId` is accepted and carried in context, but no native adapter currently sends it — explicit references are warned about and dropped, so stable prompts plus session affinity remain the cache-reuse path.
684
671
 
685
672
  `ttlSeconds` overrides the retention mapping.
686
673
 
687
674
  ### Provider Cache Mechanisms
688
675
 
689
- * **Google** keeps system prompts and tools stable and isolates signatures.
690
- * **OpenCode / OpenRouter** use sticky session keys.
676
+ * **Google** keeps system prompts and tools stable and isolates signatures (implicit caching only).
677
+ * **OpenAI** pins affinity with `prompt_cache_key`.
678
+ * **OpenRouter** pins affinity with body `session_id` plus headers (sticky routing from the first request).
679
+ * **Custom endpoints** get headers-only affinity; anything else cache-related is warned about and dropped.
691
680
 
692
681
  The following cache helpers are exported:
693
682
 
694
- * `applyAnthropicCacheControl`
695
- * `getCacheControlForRetention`
696
683
  * `getPromptCacheRetention`
684
+ * `retentionToTtlSeconds`
697
685
  * `clampCacheKey`
698
686
 
699
687
  For the best cache reuse, keep `instructions` and the tool set stable. Put changing data in user messages or `additionalContext`.
@@ -808,8 +796,6 @@ AgentResponse {
808
796
  text,
809
797
  thinking?,
810
798
  thoughtSignature?,
811
- thinkingSignature?,
812
- textSignature?,
813
799
  toolCalls,
814
800
  toolResults,
815
801
  subagents,
@@ -930,8 +916,8 @@ const res = await agent.run([...]).catch(fail);
930
916
 
931
917
  Agent Accelerator features an adaptive, dynamically synchronized model catalog powered by [models.dev](https://models.dev). To avoid shipping a bloated 4.4 MB static JSON file with production builds, the SDK employs a high-performance **12-hour TTL local caching architecture**:
932
918
 
933
- - **Automated 12-Hour Cache Validation**: When `.run()`, `.ask()`, or `.stream()` executes, the runtime checks `src/data/models-cache.json`. If the cache timestamp is within 12 hours, it reads from disk with zero network delay. When the TTL expires, it transparently synchronizes with `https://models.dev/api.json`.
934
- - **Git & Package Safety**: The dynamic cache file (`src/data/models-cache.json`) is gitignored and excluded from production packages.
919
+ - **Automated 12-Hour Cache Validation**: When `.run()`, `.ask()`, or `.stream()` executes, the runtime checks `src/data/models.dev.json`. If the cache timestamp is within 12 hours, it reads from disk with zero network delay. When the TTL expires, it transparently synchronizes with `https://models.dev/api.json`.
920
+ - **Git & Package Safety**: The dynamic cache file (`src/data/models.dev.json`) is gitignored and excluded from production packages.
935
921
 
936
922
  ### Developer Catalog Controls
937
923
 
@@ -958,9 +944,9 @@ console.log(status.modelCount, status.providerCount, status.isExpired);
958
944
  ### CLI Refresh
959
945
 
960
946
  ```bash
961
- bun run update-models # Refresh model catalog (force or expired)
962
- bun scripts/update-models.ts --force # Force re-download
963
- bun scripts/update-models.ts --ttl=24h # Refresh with custom TTL
947
+ bun run update-models # Refresh model catalog (force or expired)
948
+ bun src/update-models.ts --force # Force re-download
949
+ bun src/update-models.ts --ttl=24h # Refresh with custom TTL
964
950
  ```
965
951
 
966
952
  Upstream data lags on some inputs (e.g. `gpt-4o` accepts audio). Record
@@ -1025,15 +1011,8 @@ ModelCapabilities {
1025
1011
  }
1026
1012
  ```
1027
1013
 
1028
- When estimated input exceeds the available context budget, `Agent` trims the oldest middle history using:
1029
-
1030
- ```text
1031
- context * 0.9 - maxOutput
1032
- ```
1033
-
1034
- The beginning of the conversation and the most recent tail are preserved.
1035
-
1036
1014
  ---
1015
+
1037
1016
  ## Messages and Media
1038
1017
 
1039
1018
  ### Message
@@ -1092,7 +1071,7 @@ Supported content parts include:
1092
1071
 
1093
1072
  `inferMimeType(path)` infers the MIME type from a file extension.
1094
1073
 
1095
- Media normalization is handled automatically by providers (mapped to Vercel V4 `file` parts). Thinking traces are echoed in follow-up turns only where accepted — strict endpoints receive tool calls without `reasoning_content`.
1074
+ Media normalization is handled automatically by providers (mapped to provider-native media shapes). Thinking traces are echoed in follow-up turns only where accepted — strict endpoints receive tool calls without `reasoning_content`.
1096
1075
 
1097
1076
  Provider matrix: text + image + wav/mp3 audio + PDF work on all supported
1098
1077
  providers. Video works on Gemini only.
@@ -1106,20 +1085,6 @@ its verdict surfaces as a concise error (see [Errors](#errors)). Verified
1106
1085
  catalog corrections live in `MODALITY_OVERRIDES` (`src/models/catalog.ts`),
1107
1086
  never in the gitignored snapshot.
1108
1087
 
1109
- ---
1110
- ## Tokens
1111
-
1112
- Token counting is heuristic only.
1113
-
1114
- Actual billing information comes from `res.usage`.
1115
-
1116
- Available helpers:
1117
-
1118
- * `countTokens(string | Message[] | ProviderContext)` — useful for preflight sizing and history trimming.
1119
- * `estimateTokensFromText(text)`
1120
- * `estimateTokensFromMessage(msg)`
1121
- * `estimateTokensFromPart(part)`
1122
-
1123
1088
  ---
1124
1089
  ## Utils
1125
1090
 
@@ -1127,19 +1092,22 @@ Available helpers:
1127
1092
  * `getEnv(key, fallback?)` — resolves an environment variable.
1128
1093
  * `getApiKey(provider, explicit?, env?)` — resolves API keys using the configured environment lookup order.
1129
1094
  * `buildSessionHeaders(provider, cache?, custom?, sessionId?)` — builds provider-specific affinity headers. Normally handled automatically.
1095
+ * `withRetries(fn, { maxRetries?, maxRetryDelayMs?, signal?, label? })` / `isTransientError(err)` — bounded retries for transient provider failures (defaults: 2 retries, 5s cap, aborts never retried).
1096
+ * `convertDocumentToMarkdown(input, { filename?, format?, mimeType?, maxChars? })` — converts PDF/Word/PowerPoint/Excel/OpenDocument/RTF/EPUB/CSV to Markdown via the optional `@firecrawl/anydoc` peer. Throws `DocumentConversionError` on failure.
1097
+ * `convert_document_to_markdown` — built-in model-callable version of the above (returns a `<Document>` envelope); auto-registered with `bypassInputFileModality: true`.
1130
1098
  * `z` — re-exported from Zod so tools do not require a separate Zod import.
1131
1099
 
1132
1100
  ---
1133
1101
  ## Examples
1134
1102
 
1135
1103
  ```bash
1136
- bun run examples/06-chat.ts
1104
+ bun run examples/05-chat.ts
1137
1105
  # persistent CLI, /model "..." /level <lvl> /help /exit
1138
1106
 
1139
- bun run examples/05-sub-agents.ts
1107
+ bun run examples/04-sub-agents.ts
1140
1108
  # fixed researcher + critic pipeline
1141
1109
 
1142
- bun run examples/04-multi_agent.ts
1110
+ bun run examples/03-multi_agent.ts
1143
1111
  # dynamic spawn_subagents demo
1144
1112
 
1145
1113
  bun run examples/01-metadata.ts
@@ -1148,17 +1116,20 @@ bun run examples/01-metadata.ts
1148
1116
  bun run examples/02-function_calling.ts
1149
1117
  # single tool call
1150
1118
 
1151
- bun run examples/03-multimodal_image.ts [./photo.png]
1119
+ bun run examples/06-multimodal_image.ts [./photo.png]
1152
1120
  # image input (remote URL default, local path optional)
1153
1121
 
1154
1122
  bun run examples/07-multimodal_audio.ts
1155
1123
  # audio input
1156
1124
 
1157
- bun run examples/08-multimodal_document.ts
1125
+ bun run examples/09-multimodal_document.ts
1158
1126
  # PDF/file input
1159
1127
 
1160
- bun run examples/09-multimodal_video.ts
1128
+ bun run examples/08-multimodal_video.ts
1161
1129
  # video input (video-capable model required)
1130
+
1131
+ bun run examples/10-document_markdown.ts
1132
+ # document → Markdown preprocessing (any model, optional @firecrawl/anydoc peer)
1162
1133
  ```
1163
1134
 
1164
1135
  Every example ends its `run()` with `.catch(fail)` (`examples/_shared.ts`),
@@ -1193,20 +1164,20 @@ bun run update-models # refresh model catalog cache (supports --force, --ttl=24h
1193
1164
  ```text
1194
1165
  src/
1195
1166
  ├── agent/ # Agent, context, loop, delegation, subagent
1196
- ├── ai-sdk/ # Provider backend: provider factories, converters, executor, registry
1167
+ ├── providers/ # native REST adapters (google/openai/openrouter/openai-compat) + registry + canonical contract
1197
1168
  ├── models/ # Dynamic catalog cache, parser, verified overrides
1198
1169
  ├── data/ # Dynamic model catalog cache (gitignored, excluded from bundle)
1199
1170
  ├── tools/ # tool(), schema, executor
1200
1171
  ├── streaming/ # event stream, SSE parser
1201
- ├── tokens/ # estimator
1202
1172
  ├── types/ # agent, core, message, model, response, tool
1203
- └── utils/ # base64, cache, env, headers, media, serialization, session
1173
+ └── utils/ # base64, cache, env, headers, media, serialization, session, thought-signature, documents, retry, errors
1204
1174
  examples/
1205
- ├── 01-metadata.ts 02-function_calling.ts 03-multimodal_image.ts
1206
- ├── 04-multi_agent.ts 05-sub-agents.ts 06-chat.ts
1207
- ├── 07-multimodal_audio.ts 08-multimodal_document.ts
1208
- ├── 09-multimodal_video.ts _shared.ts
1209
- test/ # 105 tests mirroring the above
1175
+ ├── 01-metadata.ts 02-function_calling.ts 03-multi_agent.ts
1176
+ ├── 04-sub-agents.ts 05-chat.ts 06-multimodal_image.ts
1177
+ ├── 07-multimodal_audio.ts 08-multimodal_video.ts
1178
+ ├── 09-multimodal_document.ts 10-document_markdown.ts
1179
+ ├── files/ prompts/ research-agent.ts _shared.ts
1180
+ test/ # unit + mocked-provider tests mirroring src/
1210
1181
  ```
1211
1182
 
1212
1183
  ---
@@ -5,7 +5,7 @@ You are an Agent — you evaluate user objectives and decide whether to solve th
5
5
  - Single-focus, factual, code-generation, or conversational prompts → Answer directly, concisely, and completely.
6
6
  - Do NOT spawn sub-agents for trivial, single-step tasks.
7
7
  2. **Delegate**:
8
- - Complex research, multi-angle analysis, deep technical trade-offs, or parallel explorations → Invoke `spawn_subagents` with 2 to 5 focused sub-agents in a single batched call.
8
+ - Complex research, multi-angle analysis, deep technical trade-offs, or parallel explorations → Invoke `spawn_subagents` with 2 to 4 focused sub-agents in a single batched call.
9
9
  - For each sub-agent define:
10
10
  - `name`: `UPPER_SNAKE_CASE` identifier describing domain focus (e.g., `MARKET_ANALYST`, `SYSTEMS_ARCHITECT`).
11
11
  - `role`: Distinct specialized persona.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "agent-accelerator",
3
- "version": "0.1.0",
3
+ "version": "0.2.0",
4
4
  "description": "High-performance SDK for AI agents with a unified API across AI providers. Optimized for high cache hit rates, observability, and seamless agent orchestration. Built to power production-grade agentic systems and harnesses for teams at any scale.",
5
5
  "license": "MIT",
6
6
  "author": "Sashvat Bharat",
@@ -23,15 +23,14 @@
23
23
  ],
24
24
  "files": [
25
25
  "src",
26
- "SYSTEM_PROMPT.md",
27
- "SYSTEM_PROMPT_AGENT.md",
28
- "SYSTEM_PROMPT_TOOLS.md",
26
+ "examples/prompts/SYSTEM_PROMPT.md",
27
+ "examples/prompts/SYSTEM_PROMPT_AGENT.md",
28
+ "examples/prompts/SYSTEM_PROMPT_TOOLS.md",
29
29
  "tsconfig.json",
30
30
  "bunfig.toml",
31
31
  "bun.lock",
32
32
  "README.md",
33
33
  "LICENSE",
34
- "!src/data/models-cache.json",
35
34
  "!src/data/models.dev.json"
36
35
  ],
37
36
  "main": "./src/index.ts",
@@ -44,16 +43,22 @@
44
43
  "scripts": {
45
44
  "typecheck": "tsc --noEmit",
46
45
  "test": "bun test test/",
47
- "update-models": "bun scripts/update-models.ts"
46
+ "update-models": "bun src/update-models.ts"
48
47
  },
49
48
  "devDependencies": {
49
+ "@firecrawl/anydoc": "^0.2.4",
50
50
  "@types/bun": "latest",
51
51
  "typescript": "^7.0.0"
52
52
  },
53
+ "peerDependencies": {
54
+ "@firecrawl/anydoc": "^0.2.4"
55
+ },
56
+ "peerDependenciesMeta": {
57
+ "@firecrawl/anydoc": {
58
+ "optional": true
59
+ }
60
+ },
53
61
  "dependencies": {
54
- "@ai-sdk/google": "^4.0.64",
55
- "@ai-sdk/openai": "^4.0.60",
56
- "@ai-sdk/openai-compatible": "^3.0.44",
57
62
  "zod": "^4.4.3"
58
63
  }
59
64
  }