@x12i/ai-dispatcher 1.4.0 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +160 -31
- package/dist/index.cjs +2118 -198
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +157 -7
- package/dist/index.d.ts +157 -7
- package/dist/index.js +2127 -200
- package/dist/index.js.map +1 -1
- package/package.json +4 -4
package/README.md
CHANGED
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
# @x12i/ai-dispatcher
|
|
2
2
|
|
|
3
|
-
One request shape for OpenRouter, AWS Bedrock,
|
|
3
|
+
One request shape for OpenRouter, AWS Bedrock, the OpenAI Responses API, and Cloudflare AI. You send the same request shape. The package picks the provider, and `@x12i/ai-profiles@5.0.0` turns `reasoningEffort` into that provider's wire fields. Callers do not map effort levels onto provider fields or parse provider reasoning channels.
|
|
4
4
|
|
|
5
|
-
**Implemented providers:** `openrouter`, `bedrock`, `openai`. Any other provider throws `PROVIDER_NOT_IMPLEMENTED`.
|
|
5
|
+
**Implemented providers:** `openrouter`, `bedrock`, `openai`, `cloudflare`. Any other provider throws `PROVIDER_NOT_IMPLEMENTED`.
|
|
6
6
|
|
|
7
|
-
The package is `@x12i/ai-dispatcher`
|
|
7
|
+
The package is `@x12i/ai-dispatcher` 2.0.0. It runs on Node 20 or newer and publishes ESM and CommonJS from the same entry.
|
|
8
8
|
|
|
9
9
|
Connected MCP servers can be passed to `createAiDispatcher`. `run()` and `compile()` expose those tools to the model. `executeStreamingChat` does not. Details are in [MCP tools](#mcp-tools).
|
|
10
10
|
|
|
@@ -14,7 +14,7 @@ Connected MCP servers can be passed to `createAiDispatcher`. `run()` and `compil
|
|
|
14
14
|
npm install @x12i/ai-dispatcher
|
|
15
15
|
```
|
|
16
16
|
|
|
17
|
-
The package depends on `@x12i/ai-profiles@5.0.0`, `@x12i/openrouter-runtime@^
|
|
17
|
+
The package depends on `@x12i/ai-profiles@5.0.0`, `@x12i/openrouter-runtime@^2.0.0`, and `@x12i/bedrock-runtime@^2.0.0`.
|
|
18
18
|
|
|
19
19
|
## One request
|
|
20
20
|
|
|
@@ -24,7 +24,11 @@ import { createAiDispatcher } from "@x12i/ai-dispatcher";
|
|
|
24
24
|
const dispatcher = createAiDispatcher({
|
|
25
25
|
openrouter: { apiKey: process.env.OPENROUTER_API_KEY },
|
|
26
26
|
bedrock: { region: process.env.AWS_REGION },
|
|
27
|
-
openai: { apiKey: process.env.OPENAI_API_KEY }
|
|
27
|
+
openai: { apiKey: process.env.OPENAI_API_KEY },
|
|
28
|
+
cloudflare: {
|
|
29
|
+
apiToken: process.env.CLOUDFLARE_API_TOKEN,
|
|
30
|
+
accountId: process.env.CLOUDFLARE_ACCOUNT_ID
|
|
31
|
+
}
|
|
28
32
|
});
|
|
29
33
|
|
|
30
34
|
const response = await dispatcher.run({
|
|
@@ -50,6 +54,7 @@ Use model ids that belong to the selected catalog target. An OpenRouter id is no
|
|
|
50
54
|
| `openrouter` | `openai/gpt-5.2` | `reasoning.effort: "low"` on the chat-completions body |
|
|
51
55
|
| `bedrock` | `anthropic.claude-fable-5-1` | `additionalModelRequestFields.thinking.type: "adaptive"` and `output_config.effort: "low"` |
|
|
52
56
|
| `openai` | `gpt-5.4` | `reasoning.effort: "low"` on the Responses body |
|
|
57
|
+
| `cloudflare` | `openai/gpt-5.4` | `reasoning.effort: "low"` on the Cloudflare Responses body |
|
|
53
58
|
|
|
54
59
|
Representative catalog results for `@x12i/ai-profiles@5.0.0`:
|
|
55
60
|
|
|
@@ -85,15 +90,24 @@ createAiDispatcher({
|
|
|
85
90
|
project: process.env.OPENAI_PROJECT_ID,
|
|
86
91
|
timeoutMs: 60_000,
|
|
87
92
|
maxAttempts: 2
|
|
93
|
+
},
|
|
94
|
+
cloudflare: {
|
|
95
|
+
apiToken: process.env.CLOUDFLARE_API_TOKEN,
|
|
96
|
+
accountId: process.env.CLOUDFLARE_ACCOUNT_ID,
|
|
97
|
+
gatewayId: "default",
|
|
98
|
+
defaultModel: "@cf/meta/llama-3.1-8b-instruct",
|
|
99
|
+
endpoint: "auto",
|
|
100
|
+
timeoutMs: 120_000,
|
|
101
|
+
maxAttempts: 2
|
|
88
102
|
}
|
|
89
103
|
});
|
|
90
104
|
```
|
|
91
105
|
|
|
92
106
|
`dispatcher.provider` is the default provider id. Adapters are created on first use and reused. The default provider's adapter is created immediately.
|
|
93
107
|
|
|
94
|
-
The `logger` callback is `{ debug, info, warn, error }`. A provider-specific `logger` on `openrouter`, `bedrock`, or `
|
|
108
|
+
The `logger` callback is `{ debug, info, warn, error }`. A provider-specific `logger` on `openrouter`, `bedrock`, `openai`, or `cloudflare` wins over the dispatcher logger for that adapter. You can inject `fetch` on OpenAI, OpenRouter, and Cloudflare, and a `client` on Bedrock.
|
|
95
109
|
|
|
96
|
-
Missing OpenAI credentials throw `OPENAI_API_KEY_MISSING`. Missing Bedrock region throws the runtime's `REGION_REQUIRED`. Missing OpenRouter key throws `OPENROUTER_API_KEY_MISSING`.
|
|
110
|
+
Missing OpenAI credentials throw `OPENAI_API_KEY_MISSING`. Missing Bedrock region throws the runtime's `REGION_REQUIRED`. Missing OpenRouter key throws `OPENROUTER_API_KEY_MISSING`. Missing Cloudflare token throws `CLOUDFLARE_API_TOKEN_MISSING`. Missing Cloudflare account id throws `CLOUDFLARE_ACCOUNT_ID_MISSING`.
|
|
97
111
|
|
|
98
112
|
### Credentials and transport
|
|
99
113
|
|
|
@@ -102,35 +116,76 @@ Missing OpenAI credentials throw `OPENAI_API_KEY_MISSING`. Missing Bedrock regio
|
|
|
102
116
|
| OpenRouter | `openrouter.apiKey` | OpenRouter chat completions |
|
|
103
117
|
| Bedrock | `bedrock.region` | Converse / ConverseStream. Credentials are the AWS default chain, or `bedrock.credentials` |
|
|
104
118
|
| OpenAI | `openai.apiKey` or `OPENAI_API_KEY` | `POST {baseUrl}/responses`. Default base URL is `https://api.openai.com/v1`. Default timeout is 120 seconds. Default attempts is 2 |
|
|
119
|
+
| Cloudflare | `cloudflare.apiToken` or `CLOUDFLARE_API_TOKEN`, and `cloudflare.accountId` or `CLOUDFLARE_ACCOUNT_ID` | `POST https://api.cloudflare.com/client/v4/accounts/{accountId}/ai/...`. Default timeout is 120 seconds. Default attempts is 2 |
|
|
105
120
|
|
|
106
121
|
Optional OpenAI headers are `OpenAI-Organization` and `OpenAI-Project`.
|
|
107
122
|
|
|
123
|
+
Cloudflare uses one account token for Workers AI (`@cf/...`) and third-party models (`openai/...`, `anthropic/...`, and the rest). `auto` sends `anthropic/` models to `/ai/v1/messages`, `openai/` and bare `gpt` or `o1`/`o3`/`o4` ids to `/ai/v1/responses`, and every other id, including `@cf/`, to `/ai/v1/chat/completions`. Set `cloudflare.endpoint` to `chat`, `responses`, `messages`, or `run` to force a path. That field wins over `apiMode`. `apiMode: "chat"` and `"responses"` still select those two paths when the endpoint is `auto`. Workers AI always sends `cf-aig-gateway-id`, using `default` when `gatewayId` is omitted. Third-party models omit that header unless `gatewayId` is set. `/ai/run` wraps the chat payload as `{ model, input }`. `run.webhookUrl` requires `run.background: true`. `executeStreamingChat` rejects `run` because that endpoint is synchronous or webhook, not SSE.
|
|
124
|
+
|
|
125
|
+
Cloudflare retries `429`, `408`, and `5xx` with backoff capped at 5 seconds. After the last retryable failure, `run()` returns a failed response and streaming yields `stream.error` and then throws. `401` and `403` become `PROVIDER_AUTH_FAILED`. `404` becomes `PROVIDER_MODEL_NOT_FOUND`. A last `429` becomes `PROVIDER_RATE_LIMITED`. Other HTTP failures, including a Cloudflare `{ success: false, errors: [...] }` body, become `PROVIDER_REQUEST_FAILED`. Timeouts and network failures are `PROVIDER_GATEWAY_UNREACHABLE`. Config mistakes (`CLOUDFLARE_API_TOKEN_MISSING`, `CLOUDFLARE_ACCOUNT_ID_MISSING`, `MODEL_REQUIRED`, `INPUT_REQUIRED`, `UNSUPPORTED_REQUEST_FIELD`) throw from `run()` instead of becoming a failed response.
|
|
126
|
+
|
|
108
127
|
OpenAI retries `429`, `408`, and `5xx` with backoff capped at 5 seconds. After the last retryable failure, `run()` returns a failed response and streaming yields `stream.error` and then throws. `401` and `403` become `PROVIDER_AUTH_FAILED`. `404` becomes `PROVIDER_MODEL_NOT_FOUND`. Other HTTP failures become `PROVIDER_REQUEST_FAILED`. Timeouts and network failures are `PROVIDER_GATEWAY_UNREACHABLE`. Config mistakes (`OPENAI_API_KEY_MISSING`, `MODEL_REQUIRED`, `INPUT_REQUIRED`, `UNSUPPORTED_REQUEST_FIELD`) throw from `run()` instead of becoming a failed response.
|
|
109
128
|
|
|
110
129
|
## The request
|
|
111
130
|
|
|
112
|
-
`AiDispatcherRequest` is the OpenRouter runtime request plus `provider`, `reasoningEffort`, `bedrockControls`, and `mcp`.
|
|
131
|
+
`AiDispatcherRequest` is the OpenRouter runtime request plus `provider`, `orgId`, `stepId`, `skillId`, `reasoningEffort`, `bedrockControls`, `cache`, `cloudflare`, and `mcp`. `agentId` is already on that request.
|
|
113
132
|
|
|
114
133
|
Text can arrive as `prompt`, `messages`, `input`, `system`, or `instructions`. A message role is `system`, `user`, `assistant`, or `tool`. Content is a string or parts: `text`, `image_url`, or `file`.
|
|
115
134
|
|
|
116
135
|
| Field | Meaning |
|
|
117
136
|
| --- | --- |
|
|
118
137
|
| `id` | Caller request id. Logged when set. OpenAI uses it as a fallback response id. |
|
|
138
|
+
| `agentId` | Optional caller agent id. Stored in the one metadata record. On OpenRouter it is also the activity-log App name. |
|
|
139
|
+
| `orgId` | Optional caller organization id. Stored in the one metadata record as `orgId`. |
|
|
140
|
+
| `stepId` | Optional caller step id. Stored in the one metadata record as `stepId`. |
|
|
141
|
+
| `skillId` | Optional caller skill id. Stored in the one metadata record as `skillId`. |
|
|
119
142
|
| `temperature`, `maxTokens` | Sampling and output cap. OpenAI sends `maxTokens` as `max_output_tokens`. |
|
|
120
|
-
| `reasoning` | Explicit wire object for OpenRouter and
|
|
143
|
+
| `reasoning` | Explicit wire object for OpenRouter, OpenAI, and Cloudflare. |
|
|
121
144
|
| `functionTools` | Local function tools: `name`, `description`, `parameters`, optional `executor`. |
|
|
122
145
|
| `toolChoice` | `"auto"`, `"none"`, `"required"`, `{ type: "function", functionName }`, or an OpenRouter `{ type: "server_tool", serverTool }`. |
|
|
123
146
|
| `execution` | `maxToolIterations`, `maxServerToolCalls`, `maxFunctionToolCalls`, `timeoutMs`. Bedrock keeps the iteration, function-call, and timeout limits. |
|
|
124
|
-
| `metadata` |
|
|
125
|
-
| `serverTools` | OpenRouter server tools only. Bedrock
|
|
126
|
-
| `apiMode` | `"chat"`, `"responses"`, or `"auto"`. Direct OpenAI rejects `"chat"`. Bedrock rejects any `apiMode`. |
|
|
127
|
-
| `responseFormat` | OpenRouter forwards it. Direct OpenAI
|
|
128
|
-
| `rawOpenRouterOverrides` | Merged onto an OpenRouter body.
|
|
147
|
+
| `metadata` | The one metadata record. String, number, and boolean values, including `xtivix.execution`, `recordId`, and `objectType`. |
|
|
148
|
+
| `serverTools` | OpenRouter server tools only. Bedrock, direct OpenAI, and Cloudflare reject an enabled policy. |
|
|
149
|
+
| `apiMode` | `"chat"`, `"responses"`, or `"auto"`. Direct OpenAI rejects `"chat"`. Bedrock rejects any `apiMode`. On Cloudflare, `"chat"` and `"responses"` select those endpoints when `cloudflare.endpoint` is omitted or `auto`. |
|
|
150
|
+
| `responseFormat` | OpenRouter forwards it. Direct OpenAI and Cloudflare Responses place it on `text`. Cloudflare chat sends `response_format`. Cloudflare messages and run reject it. Bedrock rejects it. |
|
|
151
|
+
| `rawOpenRouterOverrides` | Merged onto an OpenRouter body. A `metadata` object inside it is folded into `metadata` on every provider. Any other override key is rejected on Bedrock, direct OpenAI, and Cloudflare. |
|
|
152
|
+
| `cache` | Optional prompt-cache intent: `mode` (`auto`, `explicit`, or `disabled`), `key`, and `ttl` (`5m`, `30m`, or `1h`). Omit it and the compiled body is unchanged. `explicit` places one breakpoint at the end of `system` or `instructions` when that model supports it. On Cloudflare it also sets AI Gateway cache headers. |
|
|
153
|
+
| `cloudflare` | Per-call overrides: `endpoint`, `gatewayId`, `run`, and `cf-aig-*` controls. Stripped before compilation. |
|
|
129
154
|
| `mcp` | `false` or `[]` omits MCP tools registered at initialization. A list of exposed names attaches only those tools. Omit the field to attach every registered MCP tool. |
|
|
130
155
|
|
|
156
|
+
Set `metadata` for the pairs you want back from a log. `agentId`, `orgId`, `stepId`, and `skillId` are fields on the same record. A dedicated field wins over the same key inside `metadata`. Omit an id you do not have. The dispatcher does not invent one.
|
|
157
|
+
|
|
158
|
+
```ts
|
|
159
|
+
import { decodeProviderMetadata } from "@x12i/ai-dispatcher";
|
|
160
|
+
|
|
161
|
+
await dispatcher.run({
|
|
162
|
+
model: "openai/gpt-5.2",
|
|
163
|
+
prompt: "Explain the tradeoff in one paragraph.",
|
|
164
|
+
agentId: "agent-42",
|
|
165
|
+
orgId: "org-7",
|
|
166
|
+
stepId: "step-3",
|
|
167
|
+
skillId: "skill-9",
|
|
168
|
+
metadata: {
|
|
169
|
+
"xtivix.execution": executionJson,
|
|
170
|
+
recordId: "rec-1",
|
|
171
|
+
objectType: "note"
|
|
172
|
+
}
|
|
173
|
+
});
|
|
174
|
+
|
|
175
|
+
const record = decodeProviderMetadata(logMetadata);
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
`metadata` is the one record. When `rawOpenRouterOverrides` only repeats `metadata`, the dispatcher folds it in. `metadata` wins when the same key is in both places. String, number, and boolean values are stored, including a JSON string. Objects and arrays stay on the normalized response and the tool context.
|
|
179
|
+
|
|
180
|
+
Every provider writes that record as `m.0`, `m.1`, and so on: one base64 string, 256 characters per piece, at most 16 pieces. The base64 is the UTF-8 JSON object of the string map, with keys sorted. `decodeProviderMetadata` reads those pieces from OpenRouter body `metadata`, OpenAI Responses `metadata`, Bedrock `requestMetadata`, or Cloudflare body `metadata` and `cf-aig-metadata`. The pieces are the same on every provider. A larger record throws `METADATA_RECORD_TOO_LARGE`.
|
|
181
|
+
|
|
182
|
+
OpenRouter also puts `agentId` in `X-Title` and `X-OpenRouter-Title`, ahead of `appAttribution.appName` and `defaultHeaders`. That header is the activity-log App name. `appAttribution.siteUrl` is still `HTTP-Referer` when set. The App header is not a second metadata record.
|
|
183
|
+
|
|
131
184
|
OpenAI compilation order: `instructions` wins, then `system`, then system messages joined with newlines. `input` is sent as written. Otherwise non-system messages are converted, and a string `prompt` is appended as a user item. Tool messages become `function_call_output`. Image parts become `input_image`. A file part with `url` becomes `input_file` `file_url`. A file part with `data` becomes `input_file` `file_data`. `fileName` becomes `filename`. A file part with neither `url` nor `data` throws `UNSUPPORTED_REQUEST_FIELD`.
|
|
132
185
|
|
|
133
|
-
|
|
186
|
+
Cloudflare Responses uses that same compilation. Chat and run use OpenAI chat messages. Messages uses Anthropic content blocks, with `system` kept as its own field. Chat, messages, and run reject a request that sets both `messages` and `input`. `toolChoice: "none"` omits tools. A file part with neither `url` nor `data` throws `UNSUPPORTED_REQUEST_FIELD`.
|
|
187
|
+
|
|
188
|
+
Bedrock receives `id`, `model`, `messages`, `system`, `prompt`, `temperature`, `maxTokens`, `functionTools`, a Bedrock-compatible `toolChoice`, the execution limits above, `agentId`, and `metadata` (sent as `requestMetadata`). `instructions` is used as `system` when `system` is empty. A string `input` is used as `prompt` when `prompt` and `messages` are empty. `input` items with a role and content become `messages` when `messages` is empty. Sending both `messages` and `input` items throws `UNSUPPORTED_REQUEST_FIELD`. An enabled server tool, `responseFormat`, a `rawOpenRouterOverrides` entry other than `metadata`, `apiMode`, or a `server_tool` tool choice also throws. `toolChoice: "none"` omits tools. A `data:image/(png|jpeg|jpg|gif|webp);base64,...` part is sent as Converse image bytes (`jpg` is sent as `jpeg`). Any other image URL, and any other non-text part, throws `UNSUPPORTED_CONTENT`.
|
|
134
189
|
|
|
135
190
|
## Entrypoints
|
|
136
191
|
|
|
@@ -146,6 +201,7 @@ Bedrock receives `id`, `model`, `messages`, `system`, `prompt`, `temperature`, `
|
|
|
146
201
|
- OpenRouter: `{ provider, url, headers, body, warnings }`
|
|
147
202
|
- Bedrock: `{ provider, modelId, input }`
|
|
148
203
|
- OpenAI: `{ provider, url, headers, body, warnings }`
|
|
204
|
+
- Cloudflare: `{ provider, endpoint, url, headers, body, warnings }`
|
|
149
205
|
|
|
150
206
|
`Authorization`, `x-api-key`, and `api-key` are replaced with `[redacted]` in the compiled result. Execution sends the real headers. A streaming call adds only the transport stream flag (`stream: true`, or Bedrock `ConverseStream`).
|
|
151
207
|
|
|
@@ -177,7 +233,7 @@ Each tool is exposed as `${serverId}__${toolName}`. `callTool` still receives th
|
|
|
177
233
|
|
|
178
234
|
`run()` and `compile()` attach every registered tool. `executeStreamingChat` does not attach them and does not run them. Set `mcp: false` or `mcp: []` on a request to attach none. Set `mcp: ["fs__read_file"]` to attach that tool only. An unknown name in that list throws `MCP_TOOL_NOT_FOUND`.
|
|
179
235
|
|
|
180
|
-
A `functionTools` entry with the same exposed name replaces the MCP tool for that call. If that entry has no executor, the MCP `callTool` does not run. An executor already stored on `openrouter.tools`, `bedrock.tools`, or `
|
|
236
|
+
A `functionTools` entry with the same exposed name replaces the MCP tool for that call. If that entry has no executor, the MCP `callTool` does not run. An executor already stored on `openrouter.tools`, `bedrock.tools`, `openai.tools`, or `cloudflare.tools` under the exposed name runs instead of `callTool`.
|
|
181
237
|
|
|
182
238
|
Text blocks in an MCP result are joined and returned to the model. `isError: true` fails that tool call. Any other result is JSON. Bedrock `toolChoice: "none"` still omits tools. OpenRouter advisor, subagent, and fusion cannot call these sessions.
|
|
183
239
|
|
|
@@ -196,7 +252,7 @@ Text blocks in an MCP result are joined and returned to the model. `isError: tru
|
|
|
196
252
|
|
|
197
253
|
Explicit controls are transport fields:
|
|
198
254
|
|
|
199
|
-
- OpenRouter and
|
|
255
|
+
- OpenRouter, OpenAI, and Cloudflare: `request.reasoning`.
|
|
200
256
|
- Bedrock: `request.bedrockControls.additionalModelRequestFields`.
|
|
201
257
|
|
|
202
258
|
An OpenRouter `reasoning` object on a Bedrock call is not a Bedrock control. With `reasoningEffort`, the catalog still runs and that object is removed. Without `reasoningEffort`, the call fails with `UNSUPPORTED_REQUEST_FIELD`.
|
|
@@ -209,6 +265,8 @@ Outcomes:
|
|
|
209
265
|
|
|
210
266
|
Unknown models and missing mappings throw. The error keeps the catalog `code` and `details`, including `UNKNOWN_MODEL` and `REASONING_MAPPING_MISSING`. `reasoningEffort` without any model id throws `REASONING_MODEL_REQUIRED`.
|
|
211
267
|
|
|
268
|
+
Cloudflare has no catalog target of its own. `reasoningResolution.target` stays `"cloudflare"`. The catalog is queried as `openai` after stripping an `openai/` prefix (and bare `gpt` or `o1`/`o3`/`o4` ids are prefixed back to `openai/` when the catalog rewrites the model). Every other id, including `anthropic/` and `@cf/`, is queried as `openrouter`. `UNKNOWN_MODEL` and `REASONING_MAPPING_MISSING` become outcome `ignored` with an empty body, so a Workers AI model still runs. The other providers still throw those codes.
|
|
269
|
+
|
|
212
270
|
The dispatcher does not rank efforts, invent token budgets, or rewrite `maxTokens`. Profile notes about output-token headroom stay on the catalog profile. The directive body already contains companion fields the wire needs, and those fields are merged onto the compiled body.
|
|
213
271
|
|
|
214
272
|
## Answers and reasoning text
|
|
@@ -232,7 +290,7 @@ The dispatcher does not rank efforts, invent token budgets, or rewrite `maxToken
|
|
|
232
290
|
|
|
233
291
|
The same record is on `compile()` output, `stream.start`, and `stream.done`.
|
|
234
292
|
|
|
235
|
-
The rest of the response is the runtime shape: `id`, `status` (`completed`, `failed`, `requires_action`, `policy_violation`), `apiMode`, `model`, `messages`, `citations`, `images`, `patches`, `toolUsage`, `finishReason`, `warnings`, `errors`, `raw`, and `metadata`. Direct OpenAI sets `apiMode` to `"responses"` and returns function calls as `requires_action` without running them.
|
|
293
|
+
The rest of the response is the runtime shape: `id`, `status` (`completed`, `failed`, `requires_action`, `policy_violation`), `apiMode`, `model`, `messages`, `citations`, `images`, `patches`, `toolUsage`, `finishReason`, `warnings`, `errors`, `raw`, and `metadata`. Direct OpenAI sets `apiMode` to `"responses"` and returns function calls as `requires_action` without running them. Cloudflare Responses does the same. Cloudflare chat, messages, and run report `apiMode: "chat"`.
|
|
236
294
|
|
|
237
295
|
Streaming emits `stream.reasoning.delta` for thought text. That text is not copied into `stream.text.delta`, tool-argument deltas, or structured-output text. If a future catalog asks for think-tag stripping, tags are split out of answer deltas before those deltas are yielded.
|
|
238
296
|
|
|
@@ -262,9 +320,9 @@ Events, each with `type`, `requestId`, and `data`:
|
|
|
262
320
|
|
|
263
321
|
## Tools
|
|
264
322
|
|
|
265
|
-
Local function tools work on all
|
|
323
|
+
Local function tools work on all four providers. OpenRouter, Bedrock, direct OpenAI, and Cloudflare chat, responses, and messages run the tool loop when every returned call has an `executor` on the tool or in that provider's `tools` registry. The loop defaults are 8 iterations and 20 function calls. On direct OpenAI and Cloudflare, a call with no executor stops the loop: `run()` returns `requires_action` and does not send another request. Send that tool result back as a `role: "tool"` message with `toolCallId`. Streaming stays one request and does not run executors. Exceeding `maxFunctionToolCalls` fails the run with `FUNCTION_TOOL_LIMIT`. Cloudflare `/ai/run` forwards tools inside `input` and continues the loop only when the result is a chat completion with `tool_calls`.
|
|
266
324
|
|
|
267
|
-
OpenRouter server tools are `webSearch`, `webFetch`, `datetime`, `imageGeneration`, `applyPatch`, `fusion`, `advisor`, and `subagent`. Each policy has a `mode` of `disabled`, `allowed`, or `required`. Bedrock
|
|
325
|
+
OpenRouter server tools are `webSearch`, `webFetch`, `datetime`, `imageGeneration`, `applyPatch`, `fusion`, `advisor`, and `subagent`. Each policy has a `mode` of `disabled`, `allowed`, or `required`. Bedrock, direct OpenAI, and Cloudflare throw `UNSUPPORTED_REQUEST_FIELD` when any of them is enabled, and also reject a `server_tool` tool choice. A policy whose mode is `disabled` is not enabled.
|
|
268
326
|
|
|
269
327
|
### Nested OpenRouter tools
|
|
270
328
|
|
|
@@ -278,6 +336,70 @@ Fusion, advisor, and subagent policies can set `reasoning` or `reasoningEffort`.
|
|
|
278
336
|
|
|
279
337
|
Each nested model is resolved on its own. A parent Bedrock or OpenAI body is never reused as another target's thinking config. A parent `reasoningEffort` also stops the OpenRouter runtime from copying the parent `reasoning` object onto those nested tools.
|
|
280
338
|
|
|
339
|
+
## Prompt caching
|
|
340
|
+
|
|
341
|
+
`cache` is optional. A request that omits it compiles to the same provider body as before. Caching is resolved in the dispatcher from a local model profile. `@x12i/ai-profiles` is not called, and an unknown model still runs: the outcome is `ignored` and no marker is written.
|
|
342
|
+
|
|
343
|
+
`explicit` places one breakpoint at the end of the stable prefix. That prefix is `system` or `instructions`. The alert, user prompt, or other trailing turn stays after the breakpoint, so it can change without invalidating the prefix. There is no breakpoint when that prefix is missing. `auto` does not place a breakpoint. On GPT-5.6, GPT-6, Sol, and Luna it sends implicit cache options and `prompt_cache_key`. On Claude, Nova, and Qwen it does not cache unless you set `mode` to `explicit`. `disabled` removes cache markers that arrived through `rawOpenRouterOverrides` and writes none of its own.
|
|
344
|
+
|
|
345
|
+
`key` is sent as `prompt_cache_key` only for the OpenAI dialect. On other dialects it stays off the wire. `cacheResolution.keyPresent` is still `true` when you passed one. `ttl` is folded to the nearest value that model accepts. Folding sets `cacheResolution.outcome` to `degraded` and adds a `CACHE_TTL_DEGRADED` warning. The call still runs.
|
|
346
|
+
|
|
347
|
+
On Cloudflare, `cache` is also AI Gateway cache. `disabled` sends `cf-aig-skip-cache: true` and the outcome is `applied`. `key` is `cf-aig-cache-key`. `5m`, `30m`, and `1h` are `300`, `1800`, and `3600` on `cf-aig-cache-ttl`. A model with no native prompt-cache profile still gets those headers and outcome `applied`. GPT on the Responses endpoint and Claude on the Messages endpoint also get the native breakpoint for that profile. Gateway headers are still sent. A per-request `cloudflare.cacheTtlSeconds`, `cacheKey`, or `skipCache` wins over the policy headers.
|
|
348
|
+
|
|
349
|
+
```ts
|
|
350
|
+
await dispatcher.run({
|
|
351
|
+
model: "gpt-6-sol",
|
|
352
|
+
system: playbook,
|
|
353
|
+
prompt: alertPayload,
|
|
354
|
+
cache: { mode: "explicit", key: "triage-playbook-v3", ttl: "30m" }
|
|
355
|
+
});
|
|
356
|
+
```
|
|
357
|
+
|
|
358
|
+
`compile()` returns `cacheResolution` and the warnings without sending the request. `run()` copies the same record onto the response. When the provider reports them, `usage.cachedTokens` and `usage.cacheWriteTokens` are set from OpenRouter `prompt_tokens_details`, OpenAI `input_tokens_details`, or Bedrock `cacheReadInputTokens` and `cacheWriteInputTokens`. Those fields are also filled when the provider cached implicitly and the request had no `cache` object.
|
|
359
|
+
|
|
360
|
+
| Target | Models | Explicit marker | Default TTL | Allowed TTL |
|
|
361
|
+
| --- | --- | --- | --- | --- |
|
|
362
|
+
| OpenAI | `gpt-5.6`, `gpt-6`, Sol, Luna | `prompt_cache_breakpoint` on the prefix, plus `prompt_cache_options` | `30m` | `30m` |
|
|
363
|
+
| OpenAI | earlier GPT models | none; caching is automatic | n/a | n/a |
|
|
364
|
+
| OpenRouter | `openai/gpt-5.6`, `openai/gpt-6`, Sol, Luna | same OpenAI marker on the chat or responses prefix | `30m` | `30m` |
|
|
365
|
+
| OpenRouter | Claude | `cache_control: { type: "ephemeral", ttl }` on the last system block | `5m` | `5m`, `1h` |
|
|
366
|
+
| OpenRouter | Qwen | `cache_control` on the last system block | `5m` | `5m` |
|
|
367
|
+
| OpenRouter | Gemini, DeepSeek, Grok, Moonshot, earlier GPT | none; caching is automatic | n/a | n/a |
|
|
368
|
+
| Bedrock Converse | Claude, Nova | `{ cachePoint: { type: "default", ttl } }` after `system` | `5m` | `5m`, `1h` |
|
|
369
|
+
| Bedrock Converse | OpenAI model ids | none. Those models cache through the Responses API, which this Converse runtime does not send | n/a | n/a |
|
|
370
|
+
|
|
371
|
+
TTL folding:
|
|
372
|
+
|
|
373
|
+
| Requested | GPT-5.6 / GPT-6 / Sol / Luna | Claude and Nova | Qwen |
|
|
374
|
+
| --- | --- | --- | --- |
|
|
375
|
+
| omitted | `30m` | `5m` | `5m` |
|
|
376
|
+
| `5m` | `30m`, `degraded` | `5m` | `5m` |
|
|
377
|
+
| `30m` | `30m` | `5m`, `degraded` | `5m`, `degraded` |
|
|
378
|
+
| `1h` | `30m`, `degraded` | `1h` | `5m`, `degraded` |
|
|
379
|
+
|
|
380
|
+
`30m` on Bedrock becomes `5m` because 30 minutes is closer to 5 minutes than to 1 hour.
|
|
381
|
+
|
|
382
|
+
Warnings, not thrown errors:
|
|
383
|
+
|
|
384
|
+
| Code | When |
|
|
385
|
+
| --- | --- |
|
|
386
|
+
| `CACHE_TTL_DEGRADED` | The requested TTL was folded to a supported value and that value was sent. |
|
|
387
|
+
| `CACHE_UNSUPPORTED` | The model has no explicit marker, `auto` was set on a model that only caches with a breakpoint, or the request has no stable prefix. |
|
|
388
|
+
| `CACHE_DIALECT_UNSUPPORTED` | An OpenAI model id was sent to Bedrock Converse. |
|
|
389
|
+
|
|
390
|
+
Providers ignore a prefix under their minimum size (often 1,024 tokens) and cap how many breakpoints a request may carry. This dispatcher writes one breakpoint and does not count tokens.
|
|
391
|
+
|
|
392
|
+
A cache write can cost more than an uncached input. The first hit pays for the write when `(write multiplier - 1) / (1 - read multiplier)` is under 1. Later hits are the savings.
|
|
393
|
+
|
|
394
|
+
| Profile | Cache read | Cache write | Reads that pay for one write |
|
|
395
|
+
| --- | --- | --- | --- |
|
|
396
|
+
| Claude or Qwen, 5-minute TTL | 0.1x input | 1.25x input | 1 |
|
|
397
|
+
| Claude, 1-hour TTL | 0.1x input | 2x input | 2 |
|
|
398
|
+
| GPT-5.6, GPT-6, Sol, Luna | 0.25x to 0.5x input | 1.25x input | 1 |
|
|
399
|
+
| Gemini, Grok, Moonshot, OpenAI before GPT-5.6 | about 0.25x to 0.5x input | no write premium | 1 |
|
|
400
|
+
|
|
401
|
+
Reuse the same `key` and the same prefix bytes. A changed playbook needs a new key, or the next call pays for another write.
|
|
402
|
+
|
|
281
403
|
## Escape hatches
|
|
282
404
|
|
|
283
405
|
Set the provider control and leave `reasoningEffort` unset when you already have a wire object:
|
|
@@ -303,7 +425,7 @@ await dispatcher.run({
|
|
|
303
425
|
});
|
|
304
426
|
```
|
|
305
427
|
|
|
306
|
-
`rawOpenRouterOverrides` still merges onto an OpenRouter body.
|
|
428
|
+
`rawOpenRouterOverrides` still merges onto an OpenRouter body. Its `metadata` object is folded into `metadata` on every provider. Any other override key is rejected on Bedrock, direct OpenAI, and Cloudflare.
|
|
307
429
|
|
|
308
430
|
## Supported features
|
|
309
431
|
|
|
@@ -311,6 +433,7 @@ await dispatcher.run({
|
|
|
311
433
|
| --- | --- | --- | --- |
|
|
312
434
|
| `run`, `compile`, `executeStreamingChat` | yes | yes | yes |
|
|
313
435
|
| `reasoningEffort` via the catalog | yes | yes | yes |
|
|
436
|
+
| `cache` explicit breakpoint | GPT-5.6 / GPT-6 family; Claude and Qwen via `cache_control`. Other models stay implicit | Claude and Nova via `cachePoint`. OpenAI ids are not marked | GPT-5.6 / GPT-6 family, including Sol and Luna. Earlier models stay implicit |
|
|
314
437
|
| Explicit reasoning control | `reasoning` | `bedrockControls.additionalModelRequestFields` | `reasoning` |
|
|
315
438
|
| Local function tools | yes, including the tool loop | yes, including the tool loop | yes, when every returned call has an executor; otherwise `requires_action` |
|
|
316
439
|
| MCP tools from `mcp.servers` | `run` and `compile` | `run` and `compile` | `run` and `compile`. Streaming does not receive them |
|
|
@@ -319,26 +442,31 @@ await dispatcher.run({
|
|
|
319
442
|
| `apiMode: "chat"`, `rawOpenRouterOverrides` | OpenRouter fields | `UNSUPPORTED_REQUEST_FIELD` | `UNSUPPORTED_REQUEST_FIELD` |
|
|
320
443
|
| File content parts | forwarded | non-text parts throw `UNSUPPORTED_CONTENT`, except data-URL images | `input_file` |
|
|
321
444
|
|
|
445
|
+
Cloudflare supports the same `run`, `compile`, and `executeStreamingChat` entrypoints. `executeStreamingChat` rejects `/ai/run`. `reasoningEffort` uses the OpenAI or OpenRouter catalog, as described above. `cache` sets AI Gateway headers and, for GPT on Responses and Claude on Messages, the native breakpoint. Local function tools and MCP tools follow the OpenAI loop. OpenRouter server tools are rejected. `rawOpenRouterOverrides` is rejected except for a `metadata` object, which is folded into `metadata`. `responseFormat` is `text` on Responses, `response_format` on chat, and rejected on messages and run.
|
|
446
|
+
|
|
322
447
|
## Errors
|
|
323
448
|
|
|
324
449
|
Catalog failures throw `AiDispatcherError` with `code`, `message`, and the catalog `details`. Dispatcher codes include:
|
|
325
450
|
|
|
326
451
|
| Code | When |
|
|
327
452
|
| --- | --- |
|
|
328
|
-
| `PROVIDER_NOT_IMPLEMENTED` | Provider is not `openrouter`, `bedrock`, or `
|
|
453
|
+
| `PROVIDER_NOT_IMPLEMENTED` | Provider is not `openrouter`, `bedrock`, `openai`, or `cloudflare` |
|
|
329
454
|
| `PROVIDER_REQUIRED` | No provider string is available |
|
|
330
455
|
| `PROVIDER_ADAPTER_MISSING` | An implemented provider has no adapter |
|
|
331
456
|
| `REASONING_MODEL_REQUIRED` | `reasoningEffort` is set and no model id is available |
|
|
332
457
|
| `REASONING_PREFIX_UNSUPPORTED` | A catalog prompt prefix has no user text to attach to |
|
|
333
458
|
| `UNSUPPORTED_REQUEST_FIELD` | A field cannot be sent on the selected provider |
|
|
334
|
-
| `MODEL_REQUIRED` | Direct OpenAI has no model id |
|
|
335
|
-
| `INPUT_REQUIRED` | Direct OpenAI has no messages, input, or prompt |
|
|
459
|
+
| `MODEL_REQUIRED` | Direct OpenAI or Cloudflare has no model id |
|
|
460
|
+
| `INPUT_REQUIRED` | Direct OpenAI or Cloudflare has no messages, input, or prompt |
|
|
336
461
|
| `OPENAI_API_KEY_MISSING` | Direct OpenAI has no key |
|
|
337
|
-
| `
|
|
338
|
-
| `
|
|
339
|
-
| `
|
|
340
|
-
| `
|
|
341
|
-
| `
|
|
462
|
+
| `CLOUDFLARE_API_TOKEN_MISSING` | Cloudflare has no API token |
|
|
463
|
+
| `CLOUDFLARE_ACCOUNT_ID_MISSING` | Cloudflare has no account id |
|
|
464
|
+
| `PROVIDER_AUTH_FAILED` | OpenAI or Cloudflare HTTP `401` or `403` |
|
|
465
|
+
| `PROVIDER_MODEL_NOT_FOUND` | OpenAI or Cloudflare HTTP `404` |
|
|
466
|
+
| `PROVIDER_REQUEST_FAILED` | Other OpenAI or Cloudflare HTTP failures |
|
|
467
|
+
| `PROVIDER_RATE_LIMITED` | Cloudflare HTTP `429` after retries |
|
|
468
|
+
| `PROVIDER_RETRYABLE` / `PROVIDER_GATEWAY_UNREACHABLE` | OpenAI or Cloudflare transport failures after retries, including timeout |
|
|
469
|
+
| `FUNCTION_TOOL_LIMIT` | Function-tool calls exceed `maxFunctionToolCalls` |
|
|
342
470
|
| `UNSUPPORTED_CONTENT` | Bedrock received an image URL or content part it cannot send |
|
|
343
471
|
| `MCP_SERVER_INVALID` | An MCP server is missing an id, a tools array, or `callTool`, or a server id is repeated |
|
|
344
472
|
| `MCP_TOOL_NAME_INVALID` | An MCP server id or tool name uses characters other than letters, numbers, underscores, and hyphens, or the exposed name is longer than 64 characters |
|
|
@@ -355,9 +483,10 @@ Pass `logger` on the dispatcher. `run` and `executeStreamingChat` emit `ai-dispa
|
|
|
355
483
|
## Limitations
|
|
356
484
|
|
|
357
485
|
- Catalog 5.0.0 does not currently emit a prompt prefix, a replacement model, think-tag stripping, history deletions, or a reasoning token budget. Those paths exist for the directive type and are covered by synthetic tests.
|
|
486
|
+
- Prompt caching is a dispatcher profile, not an `@x12i/ai-profiles` field. It writes one breakpoint after `system` or `instructions`. It does not place extra breakpoints on tools or on a prefix that changes in the middle of the call.
|
|
358
487
|
- Provenance notes about output-token headroom are not enforced, because the resolver does not return them as fields.
|
|
359
|
-
- OpenRouter server tools, `apiMode: "chat"`, and `rawOpenRouterOverrides` stay on OpenRouter. Bedrock
|
|
360
|
-
- Direct OpenAI streaming
|
|
488
|
+
- OpenRouter server tools, `apiMode: "chat"`, and `rawOpenRouterOverrides` stay on OpenRouter. Bedrock, direct OpenAI, and Cloudflare reject the server-tool fields and `rawOpenRouterOverrides` keys other than `metadata`. Cloudflare accepts `apiMode: "chat"` as the chat-completions endpoint.
|
|
489
|
+
- Direct OpenAI and Cloudflare streaming do not run the function-tool loop. MCP tools are not attached to `executeStreamingChat` on any provider. Cloudflare `/ai/run` does not stream.
|
|
361
490
|
- OpenRouter advisor, subagent, and fusion cannot call MCP sessions registered on the dispatcher.
|
|
362
491
|
- Bedrock `run()` returns a failed response for `UNSUPPORTED_CONTENT`. `compile()` throws that error.
|
|
363
492
|
- Live provider calls are not part of the package test suite. Tests mock HTTP and the Bedrock client and use the real 5.0.0 catalog for contract cases.
|