@mastra/mcp-docs-server 1.2.17 → 1.2.18-alpha.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,297 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # Context engineering
4
+
5
+ A model can only work with the information available in its context window. That might include the current conversation, remembered details, application data, tool results, or relevant passages from a knowledge base.
6
+
7
+ Context engineering is the practice of deciding what information the model should see and when. The goal isn't to provide as much as possible, but to keep the context relevant and current. Too little context leaves the model without information it needs; too much can make important details harder to find, increase cost, and reduce the quality of the response long before the model reaches its context limit.
8
+
9
+ Mastra provides different ways to bring information into context, keep it available over time, retrieve it when needed, and reduce or isolate it as a task grows. This guide explains when to use each mechanism and how they fit together.
10
+
11
+ | Need | Start with | What the model sees |
12
+ | ----------------------------------------- | --------------------------------------------- | -------------------------------------------------------- |
13
+ | Stable identity, rules, or constraints | [Instructions](#instructions) | System context on each model call |
14
+ | Data from a database or API | [Tools](#tools) | Tool definitions followed by selected results |
15
+ | Current customer or application data | [Inline context](#inline-context) | Data interpolated into the user message |
16
+ | A large, stable knowledge base | [RAG](#rag) | Semantically relevant chunks from an index |
17
+ | User- or organization-managed documents | [Filesystems](#filesystems) | Files selected through read or search tools |
18
+ | Recent conversation or durable facts | [Memory](#memory) | History, observations, or retrieved memories |
19
+ | A long-running conversation | [Observational Memory](#observational-memory) | Dense observations plus recent unobserved messages |
20
+ | New events or changing state during a run | [Signals](#signals) | User, reactive, notification, or state messages |
21
+ | Instructions needed only for some tasks | [Dynamic skills](#dynamic-skills) | Skill metadata followed by instructions loaded on demand |
22
+
23
+ ## Instructions
24
+
25
+ An agent's [`instructions`](https://mastra.ai/reference/agents/agent) define its stable identity, behavior, and constraints. They're system messages and appear before conversation messages in the model request.
26
+
27
+ ```typescript
28
+ import { Agent } from '@mastra/core/agent'
29
+
30
+ export const supportAgent = new Agent({
31
+ id: 'support-agent',
32
+ name: 'Support Agent',
33
+ instructions: `You help customers understand their account.
34
+ Today is ${new Date().toDateString()}.
35
+ Use plain language and don't invent account details.`,
36
+ model: 'openai/gpt-5.6-sol',
37
+ })
38
+ ```
39
+
40
+ Keep instructions focused on behavior that applies to most calls. Adding current account data, retrieved documents, or task-specific details makes the base prompt larger and harder to reuse.
41
+
42
+ Instructions can also be resolved at runtime from [`RequestContext`](https://mastra.ai/docs/server/request-context):
43
+
44
+ ```typescript
45
+ instructions: ({ requestContext }) => {
46
+ const name = requestContext.get('name')
47
+
48
+ return `You help ${name} understand their account.`
49
+ }
50
+ ```
51
+
52
+ Use `RequestContext` when instructions depend on data that changes with each request, such as the current user, tenant, locale, role, or feature flags. Values that don't come from the request, such as the current date, can be interpolated directly as shown in the first example.
53
+
54
+ > **Tip:** If the resolved instructions change often, the model provider may not be able to reuse the same prompt cache prefix. Keep the stable part first, and pass frequently changing background through messages or signals instead.
55
+ >
56
+ > Watch [this short video on prompt caching](https://youtu.be/eBB0dBqfvuQ) to learn how cacheable prompt prefixes reduce latency and cost.
57
+
58
+ ## Inline context
59
+
60
+ Most applications pass runtime context by interpolating relevant values into the current message. This works well when your code has already loaded customer or application data:
61
+
62
+ ```typescript
63
+ const customer = await db.customer.findById(customerId)
64
+
65
+ await supportAgent.generate(`
66
+ Customer: ${customer.name}
67
+ Plan: ${customer.plan}
68
+ Question: ${question}
69
+ `)
70
+ ```
71
+
72
+ Select and label the fields the model needs instead of serializing an entire database record. This keeps the prompt smaller and makes the meaning of each value clear.
73
+
74
+ When memory is enabled, Mastra saves the current user message. Don't interpolate sensitive or temporary data that shouldn't appear in conversation history.
75
+
76
+ For the less common case where background should affect one response without being saved as conversation history, pass a [`context`](https://mastra.ai/reference/agents/generate) message:
77
+
78
+ ```typescript
79
+ await supportAgent.generate('Recommend the next action.', {
80
+ context: [{ role: 'user', content: 'The customer has an unresolved billing dispute.' }],
81
+ })
82
+ ```
83
+
84
+ The model sees this background for the current execution, but Mastra doesn't save it to memory. Use `context` when persisting the background would pollute the conversation or expose temporary application state on later turns.
85
+
86
+ ## Tools
87
+
88
+ [Tools](https://mastra.ai/docs/agents/tools) are the recommended way to fetch current data from a database, API, or service. The model decides when it needs the data and supplies the tool arguments, while your application controls the query and returned fields.
89
+
90
+ ```typescript
91
+ import { createTool } from '@mastra/core/tools'
92
+ import { z } from 'zod'
93
+
94
+ export const getCustomer = createTool({
95
+ id: 'get-customer',
96
+ description: 'Gets the current profile and plan for a customer',
97
+ inputSchema: z.object({ customerId: z.string() }),
98
+ execute: async ({ customerId }) => {
99
+ const customer = await db.customer.findById(customerId)
100
+ return { name: customer.name, plan: customer.plan, status: customer.status }
101
+ },
102
+ })
103
+ ```
104
+
105
+ Use [`toModelOutput`](https://mastra.ai/docs/agents/tools) when application code needs the full result but the model needs a smaller representation.
106
+
107
+ ## RAG
108
+
109
+ [Retrieval-Augmented Generation (RAG)](https://mastra.ai/reference/rag/overview) retrieves semantically relevant chunks from an indexed corpus. It still fits large, stable knowledge bases where users ask open-ended questions that don't map cleanly to structured database queries.
110
+
111
+ ```typescript
112
+ import { ModelRouterEmbeddingModel } from '@mastra/core/llm'
113
+ import { createVectorQueryTool } from '@mastra/rag'
114
+
115
+ const knowledgeBase = createVectorQueryTool({
116
+ vectorStoreName: 'knowledgeBase',
117
+ indexName: 'support-docs',
118
+ model: new ModelRouterEmbeddingModel('openai/text-embedding-3-small'),
119
+ })
120
+ ```
121
+
122
+ Register the vector store referenced by `vectorStoreName` on the same Mastra instance as the agent. Mastra supports [multiple vector databases](https://mastra.ai/reference/rag/vector-databases). RAG is often exposed through a tool, as in this example. The design choice is whether the agent needs semantic retrieval or can query the source directly.
123
+
124
+ Many applications now start with direct, source-specific tools. Models have become better at selecting them, and a direct query is often simpler and cheaper because it doesn't require a chunking, embedding, and vector-index pipeline. Choose RAG when semantic search over unstructured content is the actual requirement, then constrain the returned context with metadata filters, reranking, and a conservative `topK`.
125
+
126
+ ## Filesystems
127
+
128
+ A [filesystem](https://mastra.ai/docs/sandbox/filesystem) gives an agent persistent access to documents and other files. Files can live in a local directory or in providers such as Amazon S3, AgentFS, or Google Drive. The agent receives built-in tools to list, read, and search them.
129
+
130
+ ```typescript
131
+ import { LocalFilesystem, Workspace } from '@mastra/core/workspace'
132
+
133
+ export const workspace = new Workspace({
134
+ filesystem: new LocalFilesystem({ basePath: './knowledge-base' }),
135
+ bm25: true,
136
+ autoIndexPaths: ['**/*.md'],
137
+ })
138
+
139
+ // Agents receive tools including read_file, list_files, grep,
140
+ // mastra_workspace_search, and mastra_workspace_index.
141
+ await workspace.init()
142
+ ```
143
+
144
+ [Workspace search](https://mastra.ai/docs/sandbox/search) supports BM25 keyword search, vector semantic search, or a hybrid of both. Use a filesystem for a personal assistant that works with a user's files or an organization knowledge base that teammates update in a service such as Google Drive. Search runs against the workspace index, so changed files must be indexed before the agent can retrieve their latest contents.
145
+
146
+ > **Tip:** Filesystem search can also use vectors, so it overlaps with RAG. Choose a filesystem when the source of truth is a set of files that the agent may need to list, read, or update. Choose a standalone RAG pipeline when retrieval is the main requirement and the source content doesn't need to behave like files.
147
+
148
+ ## Memory
149
+
150
+ [Memory](https://mastra.ai/docs/memory/overview) gives an agent conversational coherence across turns. It brings recent messages and remembered details into context without requiring the application to resend the full transcript on every turn.
151
+
152
+ Memory requires a storage provider. Each call also identifies a `resource` that owns the memory and a `thread` that identifies the conversation. Reuse both values to continue the same conversation:
153
+
154
+ ```typescript
155
+ import { Agent } from '@mastra/core/agent'
156
+ import { Memory } from '@mastra/memory'
157
+
158
+ export const assistant = new Agent({
159
+ id: 'assistant',
160
+ name: 'Assistant',
161
+ model: 'openai/gpt-5.6-sol',
162
+ memory: new Memory({
163
+ options: {
164
+ lastMessages: 20,
165
+ },
166
+ }),
167
+ })
168
+
169
+ await assistant.generate('Help me plan the next project milestone.', {
170
+ memory: {
171
+ resource: 'user-123',
172
+ thread: 'project-456',
173
+ },
174
+ })
175
+ ```
176
+
177
+ The example assumes storage is configured on the registered Mastra instance or directly on `Memory`. `lastMessages` controls how many recent messages Mastra loads from the thread. The default is 10.
178
+
179
+ Message history works well for shorter conversations where recent turns contain the context the agent needs. For long-running conversations, Mastra recommends [Observational Memory](https://mastra.ai/docs/memory/observational-memory), which keeps recent conversation available and turns older history into a dense observation log.
180
+
181
+ ## Observational Memory
182
+
183
+ Conversation history grows with every user message, response, and tool call. Even before it reaches the model's hard context limit, a long transcript can increase cost and make relevant details harder for the model to find. Compression replaces old, verbose history with a smaller representation.
184
+
185
+ [Observational Memory](https://mastra.ai/docs/memory/observational-memory) handles this continuously. An Observer turns older messages and tool interactions into dense observations, while periodic reflection reorganizes and compresses those observations.
186
+
187
+ You don't need to configure `lastMessages` when Observational Memory is enabled. Observational Memory manages history itself, keeping recent unobserved messages in context and replacing older messages with observations.
188
+
189
+ ```typescript
190
+ import { Agent } from '@mastra/core/agent'
191
+ import { Memory } from '@mastra/memory'
192
+
193
+ export const assistant = new Agent({
194
+ id: 'assistant',
195
+ name: 'Assistant',
196
+ model: 'openai/gpt-5.6-sol',
197
+ memory: new Memory({
198
+ options: {
199
+ observationalMemory: true,
200
+ },
201
+ }),
202
+ })
203
+ ```
204
+
205
+ After messages are observed, the model receives the observation log, recent messages that haven't been observed, and a continuation reminder. The raw messages remain stored but no longer occupy the active model context.
206
+
207
+ Observations are added in stable chunks, which helps providers reuse the existing prompt prefix. Observational Memory can also activate buffered observations after a prompt cache is likely to expire or before the agent changes providers.
208
+
209
+ ## Signals
210
+
211
+ > **Beta:** Signals may change without a major version bump until the API is stable.
212
+
213
+ [Signals](https://mastra.ai/docs/harness/signals) add messages or system-generated context to a memory-backed thread. Delivery depends on the thread's state: a signal can wake an idle agent or enter an active loop. It can also wait for the next turn or persist without waking the agent.
214
+
215
+ State signals require memory and an existing thread. Notification inbox signals require a storage adapter with notification support.
216
+
217
+ | API | Use | Context behavior |
218
+ | ------------------- | ------------------------------------------------------------------ | ----------------------------------------------------------------------- |
219
+ | `sendMessage()` | User input that the active agent should see now | Enters the active loop or wakes an idle thread |
220
+ | `queueMessage()` | User input that should wait for the next turn | Starts after the current run finishes |
221
+ | `sendSignal()` | Background results, policy reminders, or external events | Adds reactive or notification context according to its delivery options |
222
+ | `sendStateSignal()` | Browser state, editor state, task state, or another changing value | Maintains a thread-scoped state lane with snapshots and deltas |
223
+
224
+ Use `sendSignal()` for context produced by the system rather than the user:
225
+
226
+ ```typescript
227
+ const result = agent.sendSignal(
228
+ {
229
+ type: 'notification',
230
+ contents: 'CI failed on pull request 123: three tests failed.',
231
+ attributes: { source: 'github', pullRequest: 123 },
232
+ },
233
+ {
234
+ resourceId: 'user-123',
235
+ threadId: 'project-456',
236
+ },
237
+ )
238
+
239
+ await result.accepted
240
+ ```
241
+
242
+ A processor can send a reactive signal during `processInputStep()`. This is useful for guidance that depends on the current step or a recent tool result. Set `transient: true` when the signal should reach only the current model call. Re-send it when needed instead of storing repeated reminders in conversation history.
243
+
244
+ State signals represent context that changes over time. Mastra tracks snapshots and deltas for each state lane and can reinsert a fresh snapshot after the previous one leaves the active context window. Use `computeStateSignal()` when a processor owns the state. Working memory, browser context, and task lists can use this lane to stay available even after history or Observational Memory removes older messages.
245
+
246
+ Signals append changing context near the current turn instead of rewriting the agent's base instructions. Transient and state signals can therefore preserve a more stable prompt prefix while keeping current guidance and state visible to the model.
247
+
248
+ ## Dynamic skills
249
+
250
+ [Agent skills](https://mastra.ai/docs/skills) let an agent load task-specific instructions only when needed instead of carrying every procedure in its base instructions. Use them for specialized guidance that applies to some requests and keep the default context smaller.
251
+
252
+ ## Context control
253
+
254
+ Context control limits what the model sees as a task grows. Compaction and processors reduce context within one agent, while subagent boundaries control what moves between agents.
255
+
256
+ ### Compaction
257
+
258
+ If you've used Claude Code, you may have seen compaction happen during a long session. The Mastra team likes to joke, "Friends don't let friends do compaction."
259
+
260
+ Compaction waits until a conversation reaches a token threshold. It then summarizes the transcript and replaces earlier messages. It's a blunt fallback. The compaction turn adds latency, and a single summary has to represent everything that came before. Repeated summaries can flatten chronology or lose details that later become important.
261
+
262
+ Prefer [Observational Memory](#observational-memory) for long-running conversations. It can process history asynchronously in the background while preserving temporal context. Reflection revisits accumulated memories and naturally prunes details that no longer matter. Mastra doesn't provide compaction out of the box, though you could implement it with a custom [processor](https://mastra.ai/docs/agents/processors).
263
+
264
+ ### Processors
265
+
266
+ [Processors](https://mastra.ai/docs/agents/processors) control what enters model context and can rewrite content before a model call. Use them when information should remain in stored history or application output but doesn't need to be sent back to the model on every step.
267
+
268
+ For example, `ToolCallFilter` removes old tool arguments and results from the next model request without deleting those messages from storage. The model gets a smaller prompt, while your application can still display or inspect the complete interaction.
269
+
270
+ Use `processInput()` or `processInputStep()` to change the active message list. Those changes may later be saved to memory. Use `processLLMRequest()` when a rewrite should apply only to the current provider call and leave memory untouched.
271
+
272
+ Mastra includes several controls for common sources of context bloat:
273
+
274
+ - [`toModelOutput`](https://mastra.ai/docs/agents/tools): Replace a verbose tool result with a smaller model-facing representation.
275
+ - [`ToolCallFilter`](https://mastra.ai/reference/processors/tool-call-filter): Remove old tool calls and results from model input while retaining them in memory and the UI.
276
+ - [`ToolSearchProcessor`](https://mastra.ai/reference/processors/tool-search-processor): Replace a large tool catalog with search and load tools.
277
+ - [`TokenLimiter`](https://mastra.ai/reference/processors/token-limiter-processor): Prune non-system messages until the prompt fits a token budget.
278
+
279
+ Start with [`toModelOutput`](https://mastra.ai/docs/agents/tools) for verbose tool results and `ToolCallFilter` for old tool interactions. Add `TokenLimiter` as a final budget guard rather than relying on the model's maximum context window.
280
+
281
+ ### Subagents
282
+
283
+ A common source of context bloat is passing too much information to and from [subagents](https://mastra.ai/docs/subagents). A subagent gets a separate model context for its delegated task, but the boundary still needs deliberate controls.
284
+
285
+ By default, Mastra forwards the parent's conversation to the subagent. Use `messageFilter` to pass only the messages that specialist needs:
286
+
287
+ ```typescript
288
+ await supervisor.generate('Investigate the failed deployment.', {
289
+ delegation: {
290
+ messageFilter: ({ messages }) => messages.slice(-10),
291
+ },
292
+ })
293
+ ```
294
+
295
+ In the other direction, Mastra returns the subagent's text to the parent model by default while keeping nested tool calls and metadata available to application code. Leave `includeSubAgentToolResultsInModelContext` disabled unless the parent must reason over those details. Use `onDelegationStart` to refine the child prompt and `onDelegationComplete` to reduce or replace the text returned to the parent.
296
+
297
+ See [Subagents](https://mastra.ai/docs/subagents) for delegation hooks, memory isolation, iteration monitoring, and result controls.
@@ -50,6 +50,7 @@ List of required environment variables for each model provider and gateway suppo
50
50
  | [DigitalOcean](https://mastra.ai/models/providers/digitalocean) | `digitalocean/*` | `DIGITALOCEAN_ACCESS_TOKEN` |
51
51
  | [DInference](https://mastra.ai/models/providers/dinference) | `dinference/*` | `DINFERENCE_API_KEY` |
52
52
  | [EBCloud](https://mastra.ai/models/providers/ebcloud) | `ebcloud/*` | `EBCLOUD_API_KEY` |
53
+ | [Echo](https://mastra.ai/models/providers/echo) | `echo/*` | `ECHO_API_KEY` |
53
54
  | [Eden AI](https://mastra.ai/models/providers/edenai) | `edenai/*` | `EDENAI_API_KEY` |
54
55
  | [EmpirioLabs AI](https://mastra.ai/models/providers/empiriolabs) | `empiriolabs/*` | `EMPIRIOLABS_API_KEY` |
55
56
  | [evroc](https://mastra.ai/models/providers/evroc) | `evroc/*` | `EVROC_API_KEY` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # Model Providers
4
4
 
5
- Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 6106 models from 178 providers through a single API.
5
+ Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 6202 models from 179 providers through a single API.
6
6
 
7
7
  ## Features
8
8
 
@@ -46,7 +46,7 @@ for await (const chunk of stream) {
46
46
  | `chutes/Qwen/Qwen3-32B-TEE` | 41K | | | | | | $0.10 | $0.42 |
47
47
  | `chutes/Qwen/Qwen3.5-397B-A17B-TEE` | 262K | | | | | | $0.45 | $3 |
48
48
  | `chutes/Qwen/Qwen3.6-27B-TEE` | 262K | | | | | | $0.30 | $2 |
49
- | `chutes/Qwen/Qwen3.8-27B-TEE` | 262K | | | | | | $0.45 | $3 |
49
+ | `chutes/Qwen/Qwen3.8-27B-TEE` | 262K | | | | | | $0.40 | $3 |
50
50
  | `chutes/unsloth/Mistral-Nemo-Instruct-2407-TEE` | 131K | | | | | | $0.02 | $0.10 |
51
51
  | `chutes/zai-org/GLM-5.1-TEE` | 203K | | | | | | $0.98 | $3 |
52
52
  | `chutes/zai-org/GLM-5.2-TEE` | 1.0M | | | | | | $1 | $4 |
@@ -0,0 +1,73 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # ![Echo logo](https://models.dev/logos/echo.svg)Echo
4
+
5
+ Access 1 Echo model through Mastra's model router. Authentication is handled automatically using the `ECHO_API_KEY` environment variable.
6
+
7
+ Learn more in the [Echo documentation](https://echo.tracerml.ai).
8
+
9
+ ```bash
10
+ ECHO_API_KEY=your-api-key
11
+ ```
12
+
13
+ ```typescript
14
+ import { Agent } from "@mastra/core/agent";
15
+
16
+ const agent = new Agent({
17
+ id: "my-agent",
18
+ name: "My Agent",
19
+ instructions: "You are a helpful assistant",
20
+ model: "echo/echo"
21
+ });
22
+
23
+ // Generate a response
24
+ const response = await agent.generate("Hello!");
25
+
26
+ // Stream a response
27
+ const stream = await agent.stream("Tell me a story");
28
+ for await (const chunk of stream) {
29
+ console.log(chunk);
30
+ }
31
+ ```
32
+
33
+ > **Note:** Mastra uses the OpenAI-compatible `/chat/completions` endpoint. Some provider-specific features may not be available. Check the [Echo documentation](https://echo.tracerml.ai) for details.
34
+
35
+ ## Models
36
+
37
+ | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
+ | ----------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
+ | `echo/echo` | 262K | | | | | | $10 | $50 |
40
+
41
+ ## Advanced configuration
42
+
43
+ ### Custom headers
44
+
45
+ ```typescript
46
+ const agent = new Agent({
47
+ id: "custom-agent",
48
+ name: "custom-agent",
49
+ model: {
50
+ url: "https://echo.tracerml.ai/v1",
51
+ id: "echo/echo",
52
+ apiKey: process.env.ECHO_API_KEY,
53
+ headers: {
54
+ "X-Custom-Header": "value"
55
+ }
56
+ }
57
+ });
58
+ ```
59
+
60
+ ### Dynamic model selection
61
+
62
+ ```typescript
63
+ const agent = new Agent({
64
+ id: "dynamic-agent",
65
+ name: "Dynamic Agent",
66
+ model: ({ requestContext }) => {
67
+ const useAdvanced = requestContext.task === "complex";
68
+ return useAdvanced
69
+ ? "echo/echo"
70
+ : "echo/echo";
71
+ }
72
+ });
73
+ ```
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![Eden AI logo](https://models.dev/logos/edenai.svg)Eden AI
4
4
 
5
- Access 234 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
5
+ Access 235 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [Eden AI documentation](https://docs.edenai.co).
8
8
 
@@ -129,8 +129,8 @@ for await (const chunk of stream) {
129
129
  | `edenai/google/gemini-3.5-flash` | 1.0M | | | | | | $2 | $9 |
130
130
  | `edenai/google/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
131
131
  | `edenai/google/gemini-3.6-flash` | 1.0M | | | | | | $0.75 | $4 |
132
- | `edenai/google/gemini-3.7-flash` | 1.0M | | | | | | $0.75 | $4 |
133
- | `edenai/google/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
132
+ | `edenai/google/gemini-3.7-flash` | 1.0M | | | | | | $2 | $8 |
133
+ | `edenai/google/gemini-flash-latest` | 1.0M | | | | | | $2 | $8 |
134
134
  | `edenai/google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
135
135
  | `edenai/google/lyria-3-clip-preview` | 1.0M | | | | | | — | — |
136
136
  | `edenai/groq/openai/gpt-oss-120b` | 131K | | | | | | $0.15 | $0.60 |
@@ -203,7 +203,7 @@ for await (const chunk of stream) {
203
203
  | `edenai/perplexityai/sonar-pro` | 200K | | | | | | $3 | $15 |
204
204
  | `edenai/perplexityai/sonar-reasoning-pro` | 128K | | | | | | $2 | $8 |
205
205
  | `edenai/qwen/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.20 | $0.40 |
206
- | `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $0.66 | $2 |
206
+ | `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $4 |
207
207
  | `edenai/qwen/qwen-max` | 33K | | | | | | $2 | $6 |
208
208
  | `edenai/qwen/qwen-vl-max` | 131K | | | | | | $0.80 | $3 |
209
209
  | `edenai/qwen/qwen-vl-plus` | 131K | | | | | | $0.21 | $0.63 |
@@ -250,10 +250,10 @@ for await (const chunk of stream) {
250
250
  | `edenai/vertex/gemini-3.6-flash` | 1.0M | | | | | | $0.75 | $4 |
251
251
  | `edenai/vertex/gemini-3.6-flash@eu` | 1.0M | | | | | | $0.75 | $4 |
252
252
  | `edenai/vertex/gemini-3.6-flash@us` | 1.0M | | | | | | $0.75 | $4 |
253
- | `edenai/vertex/gemini-3.7-flash` | 1.0M | | | | | | $0.75 | $4 |
254
- | `edenai/vertex/gemini-3.7-flash@eu` | 1.0M | | | | | | $0.75 | $4 |
255
- | `edenai/vertex/gemini-3.7-flash@us` | 1.0M | | | | | | $0.75 | $4 |
256
- | `edenai/vertex/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
253
+ | `edenai/vertex/gemini-3.7-flash` | 1.0M | | | | | | $2 | $8 |
254
+ | `edenai/vertex/gemini-3.7-flash@eu` | 1.0M | | | | | | $2 | $8 |
255
+ | `edenai/vertex/gemini-3.7-flash@us` | 1.0M | | | | | | $2 | $8 |
256
+ | `edenai/vertex/gemini-flash-latest` | 1.0M | | | | | | $2 | $8 |
257
257
  | `edenai/vertex/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
258
258
  | `edenai/xai/grok-4.20-0309-non-reasoning` | 1.0M | | | | | | $1 | $3 |
259
259
  | `edenai/xai/grok-4.20-0309-reasoning` | 1.0M | | | | | | $1 | $3 |
@@ -269,6 +269,7 @@ for await (const chunk of stream) {
269
269
  | `edenai/zai/glm-5-turbo` | 203K | | | | | | $1 | $4 |
270
270
  | `edenai/zai/glm-5.1` | 203K | | | | | | $1 | $4 |
271
271
  | `edenai/zai/glm-5.2` | 1.0M | | | | | | $1 | $4 |
272
+ | `edenai/zai/glm-5.3` | 1.0M | | | | | | $1 | $4 |
272
273
  | `edenai/zai/glm-5v-turbo` | 203K | | | | | | $1 | $4 |
273
274
 
274
275
  ## Advanced configuration
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![EmpirioLabs AI logo](https://models.dev/logos/empiriolabs.svg)EmpirioLabs AI
4
4
 
5
- Access 53 EmpirioLabs AI models through Mastra's model router. Authentication is handled automatically using the `EMPIRIOLABS_API_KEY` environment variable.
5
+ Access 54 EmpirioLabs AI models through Mastra's model router. Authentication is handled automatically using the `EMPIRIOLABS_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [EmpirioLabs AI documentation](https://docs.empiriolabs.ai).
8
8
 
@@ -48,6 +48,7 @@ for await (const chunk of stream) {
48
48
  | `empiriolabs/glm-4-7-flash` | 200K | | | | | | — | — |
49
49
  | `empiriolabs/glm-5-1` | 202K | | | | | | $0.82 | $3 |
50
50
  | `empiriolabs/glm-5-2` | 1.0M | | | | | | $1 | $4 |
51
+ | `empiriolabs/glm-5-3` | 1.0M | | | | | | $1 | $4 |
51
52
  | `empiriolabs/kimi-k2-6` | 256K | | | | | | $0.89 | $4 |
52
53
  | `empiriolabs/kimi-k2-7-code` | 256K | | | | | | $0.95 | $4 |
53
54
  | `empiriolabs/kimi-k2-7-code-highspeed` | 256K | | | | | | $2 | $8 |
@@ -48,7 +48,7 @@ for await (const chunk of stream) {
48
48
  | `hyper/kimi-k2.5` | 262K | | | | | | $0.57 | $3 |
49
49
  | `hyper/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
50
50
  | `hyper/kimi-k2.7-code` | 256K | | | | | | $0.95 | $4 |
51
- | `hyper/kimi-k3` | 1.0M | | | | | | $3 | $15 |
51
+ | `hyper/kimi-k3` | 1.0M | | | | | | $3 | $16 |
52
52
  | `hyper/llama-3.3-70b-instruct` | 128K | | | | | | $0.64 | $0.77 |
53
53
  | `hyper/llama-4-maverick-17b-128e-instruct-fp8` | 430K | | | | | | $0.27 | $0.90 |
54
54
  | `hyper/minimax-m2.7` | 262K | | | | | | $0.47 | $2 |
@@ -332,7 +332,7 @@ for await (const chunk of stream) {
332
332
  | `kilo/qwen/qwen3.5-flash-02-23` | 1.0M | | | | | | $0.07 | $0.26 |
333
333
  | `kilo/qwen/qwen3.5-plus-02-15` | 1.0M | | | | | | $0.26 | $2 |
334
334
  | `kilo/qwen/qwen3.5-plus-20260420` | 1.0M | | | | | | $0.30 | $2 |
335
- | `kilo/qwen/qwen3.6-27b` | 131K | | | | | | $0.45 | $3 |
335
+ | `kilo/qwen/qwen3.6-27b` | 262K | | | | | | $0.45 | $3 |
336
336
  | `kilo/qwen/qwen3.6-35b-a3b` | 262K | | | | | | $0.14 | $1 |
337
337
  | `kilo/qwen/qwen3.6-flash` | 1.0M | | | | | | $0.19 | $1 |
338
338
  | `kilo/qwen/qwen3.6-max-preview` | 262K | | | | | | $1 | $6 |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![NanoGPT logo](https://models.dev/logos/nano-gpt.svg)NanoGPT
4
4
 
5
- Access 595 NanoGPT models through Mastra's model router. Authentication is handled automatically using the `NANO_GPT_API_KEY` environment variable.
5
+ Access 593 NanoGPT models through Mastra's model router. Authentication is handled automatically using the `NANO_GPT_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [NanoGPT documentation](https://docs.nano-gpt.com).
8
8
 
@@ -523,8 +523,6 @@ for await (const chunk of stream) {
523
523
  | `nano-gpt/stepfun/step-3.7-flash:thinking` | 262K | | | | | | $0.20 | $1 |
524
524
  | `nano-gpt/TEE/deepseek-v3.2` | 164K | | | | | | $0.50 | $1 |
525
525
  | `nano-gpt/TEE/deepseek-v4-flash` | 1.0M | | | | | | $0.20 | $0.40 |
526
- | `nano-gpt/TEE/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $4 |
527
- | `nano-gpt/TEE/deepseek-v4-pro-0813:thinking` | 1.0M | | | | | | $1 | $4 |
528
526
  | `nano-gpt/TEE/gemma-3-27b-it` | 131K | | | | | | $0.20 | $0.80 |
529
527
  | `nano-gpt/TEE/gemma-4-26b-a4b-uncensored` | 66K | | | | | | $0.15 | $0.70 |
530
528
  | `nano-gpt/TEE/gemma-4-31b-it` | 262K | | | | | | $0.15 | $0.46 |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![Ofox logo](https://models.dev/logos/ofox.svg)Ofox
4
4
 
5
- Access 103 Ofox models through Mastra's model router. Authentication is handled automatically using the `OFOX_API_KEY` environment variable.
5
+ Access 107 Ofox models through Mastra's model router. Authentication is handled automatically using the `OFOX_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [Ofox documentation](https://ofox.ai/docs).
8
8
 
@@ -81,6 +81,7 @@ for await (const chunk of stream) {
81
81
  | `ofox/google/gemini-3.5-flash` | 1.0M | | | | | | $2 | $9 |
82
82
  | `ofox/google/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
83
83
  | `ofox/google/gemini-3.6-flash` | 1.0M | | | | | | $2 | $8 |
84
+ | `ofox/google/gemini-3.7-flash` | 1.0M | | | | | | $2 | $8 |
84
85
  | `ofox/minimax/m2-her` | 200K | | | | | | $0.30 | $1 |
85
86
  | `ofox/minimax/minimax-m2` | 197K | | | | | | $0.30 | $1 |
86
87
  | `ofox/minimax/minimax-m2.1` | 205K | | | | | | $0.30 | $1 |
@@ -131,6 +132,8 @@ for await (const chunk of stream) {
131
132
  | `ofox/x-ai/grok-4.1-fast` | 2.0M | | | | | | $0.20 | $0.50 |
132
133
  | `ofox/x-ai/grok-4.20` | 2.0M | | | | | | $4 | $12 |
133
134
  | `ofox/x-ai/grok-4.3` | 1.0M | | | | | | $1 | $3 |
135
+ | `ofox/x-ai/grok-4.5` | 500K | | | | | | $2 | $6 |
136
+ | `ofox/x-ai/grok-4.6` | 500K | | | | | | $2 | $6 |
134
137
  | `ofox/z-ai/glm-4.6` | 205K | | | | | | $0.40 | $2 |
135
138
  | `ofox/z-ai/glm-4.7` | 205K | | | | | | $0.40 | $2 |
136
139
  | `ofox/z-ai/glm-4.7-flashx` | 200K | | | | | | $0.07 | $0.43 |
@@ -138,6 +141,7 @@ for await (const chunk of stream) {
138
141
  | `ofox/z-ai/glm-5-turbo` | 200K | | | | | | $1 | $4 |
139
142
  | `ofox/z-ai/glm-5.1` | 200K | | | | | | $1 | $4 |
140
143
  | `ofox/z-ai/glm-5.2` | 1.0M | | | | | | $1 | $4 |
144
+ | `ofox/z-ai/glm-5.3` | 1.0M | | | | | | $1 | $4 |
141
145
  | `ofox/z-ai/glm-5v-turbo` | 200K | | | | | | $1 | $4 |
142
146
 
143
147
  ## Advanced configuration
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![OpenCode Go logo](https://models.dev/logos/opencode-go.svg)OpenCode Go
4
4
 
5
- Access 25 OpenCode Go models through Mastra's model router. Authentication is handled automatically using the `OPENCODE_API_KEY` environment variable.
5
+ Access 26 OpenCode Go models through Mastra's model router. Authentication is handled automatically using the `OPENCODE_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [OpenCode Go documentation](https://opencode.ai/docs/zen).
8
8
 
@@ -34,27 +34,28 @@ for await (const chunk of stream) {
34
34
 
35
35
  ## Models
36
36
 
37
- | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
- | ------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
- | `opencode-go/deepseek-v4-flash` | 1.0M | | | | | | $0.22 | $0.66 |
40
- | `opencode-go/deepseek-v4-pro` | 1.0M | | | | | | $0.66 | $2 |
41
- | `opencode-go/glm-5.1` | 203K | | | | | | $1 | $4 |
42
- | `opencode-go/glm-5.2` | 1.0M | | | | | | $1 | $4 |
43
- | `opencode-go/glm-5.3` | 1.0M | | | | | | $1 | $4 |
44
- | `opencode-go/gpt-5.6-luna` | 1.1M | | | | | | $0.10 | $0.60 |
45
- | `opencode-go/grok-4.5` | 500K | | | | | | $2 | $6 |
46
- | `opencode-go/hy3` | 256K | | | | | | $0.14 | $0.58 |
47
- | `opencode-go/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
48
- | `opencode-go/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
49
- | `opencode-go/kimi-k3` | 1.0M | | | | | | $3 | $15 |
50
- | `opencode-go/mimo-v2.5` | 1.0M | | | | | | $0.14 | $0.28 |
51
- | `opencode-go/mimo-v2.5-pro` | 1.0M | | | | | | $0.43 | $0.87 |
52
- | `opencode-go/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
53
- | `opencode-go/minimax-m3` | 1.0M | | | | | | $0.30 | $1 |
54
- | `opencode-go/qwen3.6-plus` | 1.0M | | | | | | $0.50 | $3 |
55
- | `opencode-go/qwen3.7-max` | 1.0M | | | | | | $3 | $8 |
56
- | `opencode-go/qwen3.7-plus` | 1.0M | | | | | | $0.40 | $2 |
57
- | `opencode-go/qwen3.8-max` | 1.0M | | | | | | $2 | $6 |
37
+ | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
+ | ---------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
+ | `opencode-go/deepseek-v4-flash` | 1.0M | | | | | | $0.22 | $0.66 |
40
+ | `opencode-go/deepseek-v4-pro` | 1.0M | | | | | | $0.66 | $2 |
41
+ | `opencode-go/glm-5.1` | 203K | | | | | | $1 | $4 |
42
+ | `opencode-go/glm-5.2` | 1.0M | | | | | | $1 | $4 |
43
+ | `opencode-go/glm-5.3` | 1.0M | | | | | | $1 | $4 |
44
+ | `opencode-go/gpt-5.6-luna` | 1.1M | | | | | | $0.20 | $1 |
45
+ | `opencode-go/grok-4.5` | 500K | | | | | | $2 | $6 |
46
+ | `opencode-go/hy3` | 256K | | | | | | $0.14 | $0.58 |
47
+ | `opencode-go/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
48
+ | `opencode-go/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
49
+ | `opencode-go/kimi-k3` | 1.0M | | | | | | $3 | $15 |
50
+ | `opencode-go/mimo-v2.5` | 1.0M | | | | | | $0.14 | $0.28 |
51
+ | `opencode-go/mimo-v2.5-pro` | 1.0M | | | | | | $0.43 | $0.87 |
52
+ | `opencode-go/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
53
+ | `opencode-go/minimax-m3` | 1.0M | | | | | | $0.30 | $1 |
54
+ | `opencode-go/muse-spark-1.2-contributor` | 1.0M | | | | | | $0.10 | $0.20 |
55
+ | `opencode-go/qwen3.6-plus` | 1.0M | | | | | | $0.50 | $3 |
56
+ | `opencode-go/qwen3.7-max` | 1.0M | | | | | | $3 | $8 |
57
+ | `opencode-go/qwen3.7-plus` | 1.0M | | | | | | $0.40 | $2 |
58
+ | `opencode-go/qwen3.8-max` | 1.0M | | | | | | $2 | $6 |
58
59
 
59
60
  ## Advanced configuration
60
61
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![Requesty logo](https://models.dev/logos/requesty.svg)Requesty
4
4
 
5
- Access 49 Requesty models through Mastra's model router. Authentication is handled automatically using the `REQUESTY_API_KEY` environment variable.
5
+ Access 139 Requesty models through Mastra's model router. Authentication is handled automatically using the `REQUESTY_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [Requesty documentation](https://requesty.ai/solution/llm-routing/models).
8
8
 
@@ -34,57 +34,147 @@ for await (const chunk of stream) {
34
34
 
35
35
  ## Models
36
36
 
37
- | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
- | ------------------------------------ | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
- | `requesty/claude-fable-5` | 1.0M | | | | | | $10 | $50 |
40
- | `requesty/claude-haiku-4-5` | 200K | | | | | | $1 | $5 |
41
- | `requesty/claude-haiku-4-5@eu` | 200K | | | | | | $1 | $6 |
42
- | `requesty/claude-opus-4-1` | 200K | | | | | | $15 | $75 |
43
- | `requesty/claude-opus-4-5` | 200K | | | | | | $5 | $25 |
44
- | `requesty/claude-opus-4-5@eu` | 200K | | | | | | $6 | $28 |
45
- | `requesty/claude-opus-4-6` | 1.0M | | | | | | $5 | $25 |
46
- | `requesty/claude-opus-4-6@eu` | 1.0M | | | | | | $6 | $28 |
47
- | `requesty/claude-opus-4-7` | 1.0M | | | | | | $5 | $25 |
48
- | `requesty/claude-opus-4-7@eu` | 1.0M | | | | | | $6 | $28 |
49
- | `requesty/claude-opus-4-8` | 1.0M | | | | | | $5 | $25 |
50
- | `requesty/claude-opus-4-8@eu` | 1.0M | | | | | | $6 | $28 |
51
- | `requesty/claude-opus-5` | 1.0M | | | | | | $5 | $25 |
52
- | `requesty/claude-opus-5@eu` | 1.0M | | | | | | $6 | $28 |
53
- | `requesty/claude-sonnet-4-5` | 1.0M | | | | | | $3 | $15 |
54
- | `requesty/claude-sonnet-4-5@eu` | 1.0M | | | | | | $3 | $17 |
55
- | `requesty/claude-sonnet-4-6` | 1.0M | | | | | | $3 | $15 |
56
- | `requesty/claude-sonnet-4-6@eu` | 1.0M | | | | | | $3 | $17 |
57
- | `requesty/claude-sonnet-4@eu` | 1.0M | | | | | | $3 | $15 |
58
- | `requesty/claude-sonnet-5` | 1.0M | | | | | | $2 | $10 |
59
- | `requesty/claude-sonnet-5@eu` | 1.0M | | | | | | $2 | $11 |
60
- | `requesty/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.14 | $0.28 |
61
- | `requesty/deepseek-v4-flash-0731@eu` | 1.0M | | | | | | $0.14 | $0.28 |
62
- | `requesty/gemini-2.5-flash@eu` | 1.0M | | | | | | $0.30 | $3 |
63
- | `requesty/gemini-3.1-flash-lite@eu` | 1.0M | | | | | | $0.28 | $2 |
64
- | `requesty/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
65
- | `requesty/gemini-3.5-flash@eu` | 1.0M | | | | | | $2 | $10 |
66
- | `requesty/gemini-3.6-flash` | 1.0M | | | | | | $2 | $7 |
67
- | `requesty/gemini-3.7-flash` | 1.0M | | | | | | $0.75 | $4 |
68
- | `requesty/gemini-3.7-flash@eu` | 1.0M | | | | | | $0.82 | $4 |
69
- | `requesty/glm-5.2` | 1.0M | | | | | | $1 | $4 |
70
- | `requesty/glm-5.2@eu` | 1.0M | | | | | | $1 | $4 |
71
- | `requesty/gpt-4.1-mini@eu` | 1.0M | | | | | | $0.44 | $2 |
72
- | `requesty/gpt-4.1-nano@eu` | 1.0M | | | | | | $0.11 | $0.44 |
73
- | `requesty/gpt-4.1@eu` | 1.0M | | | | | | $2 | $9 |
74
- | `requesty/gpt-4o-mini@eu` | 128K | | | | | | $0.17 | $0.66 |
75
- | `requesty/gpt-5-mini@eu` | 200K | | | | | | $0.28 | $2 |
76
- | `requesty/gpt-5-nano@eu` | 200K | | | | | | $0.06 | $0.44 |
77
- | `requesty/gpt-5.1@eu` | 400K | | | | | | $1 | $11 |
78
- | `requesty/gpt-5.4@eu` | 1.1M | | | | | | $3 | $15 |
79
- | `requesty/gpt-5.5@eu` | 1.1M | | | | | | $5 | $30 |
80
- | `requesty/gpt-5.6-luna@eu` | 1.1M | | | | | | $0.22 | $1 |
81
- | `requesty/gpt-5.6-sol@eu` | 1.1M | | | | | | $6 | $33 |
82
- | `requesty/gpt-5.6-terra@eu` | 1.1M | | | | | | $2 | $13 |
83
- | `requesty/gpt-5@eu` | 400K | | | | | | $1 | $11 |
84
- | `requesty/grok-4.5` | 500K | | | | | | $2 | $6 |
85
- | `requesty/kimi-k3` | 1.0M | | | | | | $2 | $11 |
86
- | `requesty/kimi-k3@eu` | 1.0M | | | | | | $2 | $11 |
87
- | `requesty/o4-mini@eu` | 200K | | | | | | $1 | $5 |
37
+ | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
+ | ------------------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
+ | `requesty/claude-fable-5` | 1.0M | | | | | | $10 | $50 |
40
+ | `requesty/claude-fable-5@eu` | 1.0M | | | | | | $11 | $55 |
41
+ | `requesty/claude-haiku-4-5` | 200K | | | | | | $1 | $5 |
42
+ | `requesty/claude-haiku-4-5@eu` | 200K | | | | | | $1 | $6 |
43
+ | `requesty/claude-opus-4-1` | 200K | | | | | | $15 | $75 |
44
+ | `requesty/claude-opus-4-5` | 200K | | | | | | $5 | $25 |
45
+ | `requesty/claude-opus-4-5@eu` | 200K | | | | | | $6 | $28 |
46
+ | `requesty/claude-opus-4-6` | 1.0M | | | | | | $5 | $25 |
47
+ | `requesty/claude-opus-4-6@eu` | 1.0M | | | | | | $6 | $28 |
48
+ | `requesty/claude-opus-4-7` | 1.0M | | | | | | $5 | $25 |
49
+ | `requesty/claude-opus-4-7@eu` | 1.0M | | | | | | $6 | $28 |
50
+ | `requesty/claude-opus-4-8` | 1.0M | | | | | | $5 | $25 |
51
+ | `requesty/claude-opus-4-8@eu` | 1.0M | | | | | | $6 | $28 |
52
+ | `requesty/claude-opus-5` | 1.0M | | | | | | $5 | $25 |
53
+ | `requesty/claude-opus-5@eu` | 1.0M | | | | | | $6 | $28 |
54
+ | `requesty/claude-sonnet-4-5` | 1.0M | | | | | | $3 | $15 |
55
+ | `requesty/claude-sonnet-4-5@eu` | 1.0M | | | | | | $3 | $17 |
56
+ | `requesty/claude-sonnet-4-6` | 1.0M | | | | | | $3 | $15 |
57
+ | `requesty/claude-sonnet-4-6@eu` | 1.0M | | | | | | $3 | $17 |
58
+ | `requesty/claude-sonnet-4@eu` | 1.0M | | | | | | $3 | $15 |
59
+ | `requesty/claude-sonnet-5` | 1.0M | | | | | | $2 | $10 |
60
+ | `requesty/claude-sonnet-5@eu` | 1.0M | | | | | | $2 | $11 |
61
+ | `requesty/deepseek-v4-flash` | 1.0M | | | | | | $0.44 | $1 |
62
+ | `requesty/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.14 | $0.28 |
63
+ | `requesty/deepseek-v4-flash-0731@eu` | 1.0M | | | | | | $0.14 | $0.28 |
64
+ | `requesty/deepseek-v4-pro` | 1.0M | | | | | | $1 | $4 |
65
+ | `requesty/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $4 |
66
+ | `requesty/deepseek-v4-pro@eu` | 1.0M | | | | | | $2 | $4 |
67
+ | `requesty/devstral-latest` | 256K | | | | | | $0.44 | $2 |
68
+ | `requesty/devstral-latest@eu` | 256K | | | | | | $0.44 | $2 |
69
+ | `requesty/fugu-ultra` | 1.0M | | | | | | $5 | $30 |
70
+ | `requesty/gemini-2.5-flash@eu` | 1.0M | | | | | | $0.30 | $3 |
71
+ | `requesty/gemini-3-pro-image` | 1.0M | | | | | | $2 | $12 |
72
+ | `requesty/gemini-3.1-flash-image` | 131K | | | | | | $0.50 | $2 |
73
+ | `requesty/gemini-3.1-flash-lite` | 1.0M | | | | | | $0.25 | $2 |
74
+ | `requesty/gemini-3.1-flash-lite@eu` | 1.0M | | | | | | $0.28 | $2 |
75
+ | `requesty/gemini-3.1-pro-preview` | 1.0M | | | | | | $2 | $12 |
76
+ | `requesty/gemini-3.5-flash` | 1.0M | | | | | | $2 | $9 |
77
+ | `requesty/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
78
+ | `requesty/gemini-3.5-flash-lite@eu` | 1.0M | | | | | | $0.33 | $3 |
79
+ | `requesty/gemini-3.5-flash@eu` | 1.0M | | | | | | $2 | $10 |
80
+ | `requesty/gemini-3.6-flash` | 1.0M | | | | | | $2 | $7 |
81
+ | `requesty/gemini-3.7-flash` | 1.0M | | | | | | $0.75 | $4 |
82
+ | `requesty/gemini-3.7-flash@eu` | 1.0M | | | | | | $0.82 | $4 |
83
+ | `requesty/gemma-4-26b-a4b-it` | 262K | | | | | | $0.07 | $0.34 |
84
+ | `requesty/gemma-4-31b-it` | 262K | | | | | | — | — |
85
+ | `requesty/glm-5.1` | 200K | | | | | | $1 | $4 |
86
+ | `requesty/glm-5.1@eu` | 200K | | | | | | $1 | $4 |
87
+ | `requesty/glm-5.2` | 1.0M | | | | | | $1 | $4 |
88
+ | `requesty/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
89
+ | `requesty/glm-5.2@eu` | 1.0M | | | | | | $1 | $4 |
90
+ | `requesty/glm-5.3` | 1.0M | | | | | | $1 | $4 |
91
+ | `requesty/gpt-4.1-mini@eu` | 1.0M | | | | | | $0.44 | $2 |
92
+ | `requesty/gpt-4.1-nano@eu` | 1.0M | | | | | | $0.11 | $0.44 |
93
+ | `requesty/gpt-4.1@eu` | 1.0M | | | | | | $2 | $9 |
94
+ | `requesty/gpt-4o-mini@eu` | 128K | | | | | | $0.17 | $0.66 |
95
+ | `requesty/gpt-5-mini@eu` | 200K | | | | | | $0.28 | $2 |
96
+ | `requesty/gpt-5-nano@eu` | 200K | | | | | | $0.06 | $0.44 |
97
+ | `requesty/gpt-5.1@eu` | 400K | | | | | | $1 | $11 |
98
+ | `requesty/gpt-5.3-chat` | 128K | | | | | | $2 | $14 |
99
+ | `requesty/gpt-5.3-codex` | 400K | | | | | | $2 | $14 |
100
+ | `requesty/gpt-5.4` | 1.1M | | | | | | $3 | $17 |
101
+ | `requesty/gpt-5.4-mini` | 400K | | | | | | $0.75 | $5 |
102
+ | `requesty/gpt-5.4-nano` | 400K | | | | | | $0.20 | $1 |
103
+ | `requesty/gpt-5.4-pro` | 1.1M | | | | | | $30 | $180 |
104
+ | `requesty/gpt-5.4@eu` | 1.1M | | | | | | $3 | $15 |
105
+ | `requesty/gpt-5.5` | 1.1M | | | | | | $6 | $33 |
106
+ | `requesty/gpt-5.5-pro` | 1.1M | | | | | | $30 | $180 |
107
+ | `requesty/gpt-5.5@eu` | 1.1M | | | | | | $5 | $30 |
108
+ | `requesty/gpt-5.6-luna` | 1.1M | | | | | | $0.20 | $1 |
109
+ | `requesty/gpt-5.6-luna@eu` | 1.1M | | | | | | $0.22 | $1 |
110
+ | `requesty/gpt-5.6-sol` | 1.1M | | | | | | $5 | $30 |
111
+ | `requesty/gpt-5.6-sol@eu` | 1.1M | | | | | | $6 | $33 |
112
+ | `requesty/gpt-5.6-terra` | 1.1M | | | | | | $2 | $12 |
113
+ | `requesty/gpt-5.6-terra@eu` | 1.1M | | | | | | $2 | $13 |
114
+ | `requesty/gpt-5@eu` | 400K | | | | | | $1 | $11 |
115
+ | `requesty/grok-4.2-beta` | 2.0M | | | | | | $2 | $6 |
116
+ | `requesty/grok-4.3` | 1.0M | | | | | | $1 | $3 |
117
+ | `requesty/grok-4.5` | 500K | | | | | | $2 | $6 |
118
+ | `requesty/grok-4.6` | 500K | | | | | | $2 | $6 |
119
+ | `requesty/grok-build-0.1` | 256K | | | | | | $1 | $2 |
120
+ | `requesty/hy3` | 262K | | | | | | $0.14 | $0.58 |
121
+ | `requesty/inkling` | 66K | | | | | | $2 | $5 |
122
+ | `requesty/inkling-256k` | 262K | | | | | | $2 | $5 |
123
+ | `requesty/kat-coder-pro` | 256K | | | | | | $0.30 | $1 |
124
+ | `requesty/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
125
+ | `requesty/kimi-k2.6@eu` | 256K | | | | | | $0.95 | $4 |
126
+ | `requesty/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
127
+ | `requesty/kimi-k2.7-code@eu` | 262K | | | | | | $1 | $5 |
128
+ | `requesty/kimi-k3` | 1.0M | | | | | | $2 | $11 |
129
+ | `requesty/kimi-k3@eu` | 1.0M | | | | | | $2 | $11 |
130
+ | `requesty/laguna-m.1` | 33K | | | | | | — | — |
131
+ | `requesty/laguna-xs.2` | 33K | | | | | | — | — |
132
+ | `requesty/leanstral-1-5` | 262K | | | | | | — | — |
133
+ | `requesty/leanstral-1-5@eu` | 262K | | | | | | — | — |
134
+ | `requesty/ling-2.6-1t` | 262K | | | | | | $0.30 | $3 |
135
+ | `requesty/ling-2.6-flash` | 262K | | | | | | $0.10 | $0.30 |
136
+ | `requesty/ling-3.0-tiny` | 262K | | | | | | — | — |
137
+ | `requesty/mimo-v2.5` | 1.0M | | | | | | $0.14 | $0.28 |
138
+ | `requesty/mimo-v2.5-pro` | 1.0M | | | | | | $0.43 | $0.87 |
139
+ | `requesty/minimax-m2.7` | 200K | | | | | | $0.30 | $1 |
140
+ | `requesty/minimax-m2.7-highspeed` | 200K | | | | | | $0.60 | $2 |
141
+ | `requesty/minimax-m3` | 1.0M | | | | | | $0.30 | $1 |
142
+ | `requesty/minimax-m3@eu` | 1.0M | | | | | | $0.40 | $2 |
143
+ | `requesty/mistral-medium-3-5` | 262K | | | | | | $2 | $8 |
144
+ | `requesty/mistral-medium-3-5@eu` | 262K | | | | | | $2 | $8 |
145
+ | `requesty/mistral-medium-latest` | 131K | | | | | | $0.44 | $2 |
146
+ | `requesty/mistral-medium-latest@eu` | 131K | | | | | | $0.44 | $2 |
147
+ | `requesty/mistral-small-2603` | 256K | | | | | | $0.17 | $0.66 |
148
+ | `requesty/mistral-small-2603@eu` | 256K | | | | | | $0.17 | $0.66 |
149
+ | `requesty/muse-glimmer-30b` | 131K | | | | | | — | — |
150
+ | `requesty/nemotron-3-nano-omni` | 300K | | | | | | $0.06 | $0.24 |
151
+ | `requesty/nemotron-3-nano-omni-30b-a3b-reasoning` | 131K | | | | | | — | — |
152
+ | `requesty/nemotron-3-nano-omni@eu` | 300K | | | | | | $0.06 | $0.24 |
153
+ | `requesty/nemotron-3-super-120b-a12b` | 1.0M | | | | | | — | — |
154
+ | `requesty/nemotron-3-ultra-550b-a55b` | 1.0M | | | | | | — | — |
155
+ | `requesty/nemotron-3-ultra-nvfp4` | 262K | | | | | | $0.60 | $2 |
156
+ | `requesty/nemotron-3.5-content-safety` | 131K | | | | | | — | — |
157
+ | `requesty/nemotron-3.5-lightning-30b-a3b` | 1.0M | | | | | | — | — |
158
+ | `requesty/nemotron-lightning-3.5-30b-a3b` | 262K | | | | | | $0.05 | $0.20 |
159
+ | `requesty/nvidia-nemotron-3-super-120b-a12b` | 262K | | | | | | $0.10 | $0.50 |
160
+ | `requesty/nvidia-nemotron-3-ultra` | 262K | | | | | | $0.50 | $3 |
161
+ | `requesty/o4-mini@eu` | 200K | | | | | | $1 | $5 |
162
+ | `requesty/qwen3.5-27b` | 262K | | | | | | $0.26 | $3 |
163
+ | `requesty/qwen3.5-2b` | 262K | | | | | | $0.02 | $0.10 |
164
+ | `requesty/qwen3.5-35b-a3b` | 262K | | | | | | $0.14 | $1 |
165
+ | `requesty/qwen3.6-plus` | 1.0M | | | | | | $0.50 | $3 |
166
+ | `requesty/qwen3.7-max` | 1.0M | | | | | | $3 | $8 |
167
+ | `requesty/qwen3.7-plus` | 1.0M | | | | | | $0.32 | $1 |
168
+ | `requesty/qwen3.8-2.4T-A95B` | 262K | | | | | | $2 | $6 |
169
+ | `requesty/qwen3.8-max` | 1.0M | | | | | | $2 | $6 |
170
+ | `requesty/ring-2.6-1t` | 262K | | | | | | $0.30 | $3 |
171
+ | `requesty/seed-1.8` | 256K | | | | | | $0.25 | $2 |
172
+ | `requesty/seed-2.0-code` | 256K | | | | | | $0.50 | $3 |
173
+ | `requesty/seed-2.0-mini` | 256K | | | | | | $0.10 | $0.40 |
174
+ | `requesty/seed-2.0-pro` | 256K | | | | | | $0.50 | $3 |
175
+ | `requesty/step-3.7-flash` | 262K | | | | | | $0.20 | $1 |
176
+ | `requesty/thinkingcap-qwen3.6-27b` | 262K | | | | | | $0.40 | $3 |
177
+ | `requesty/thinkingcap-qwen3.6-27b@eu` | 262K | | | | | | $0.40 | $3 |
88
178
 
89
179
  ## Advanced configuration
90
180
 
@@ -114,7 +204,7 @@ const agent = new Agent({
114
204
  model: ({ requestContext }) => {
115
205
  const useAdvanced = requestContext.task === "complex";
116
206
  return useAdvanced
117
- ? "requesty/o4-mini@eu"
207
+ ? "requesty/thinkingcap-qwen3.6-27b@eu"
118
208
  : "requesty/claude-fable-5";
119
209
  }
120
210
  });
@@ -53,6 +53,7 @@ Direct access to individual AI model providers. Each provider offers unique mode
53
53
  - [DigitalOcean](https://mastra.ai/models/providers/digitalocean)
54
54
  - [DInference](https://mastra.ai/models/providers/dinference)
55
55
  - [EBCloud](https://mastra.ai/models/providers/ebcloud)
56
+ - [Echo](https://mastra.ai/models/providers/echo)
56
57
  - [Eden AI](https://mastra.ai/models/providers/edenai)
57
58
  - [EmpirioLabs AI](https://mastra.ai/models/providers/empiriolabs)
58
59
  - [evroc](https://mastra.ai/models/providers/evroc)
@@ -93,6 +93,27 @@ await run.resume({
93
93
 
94
94
  If `forEachIndex` is omitted, every suspended iteration of the step is resumed with the same `resumeData`.
95
95
 
96
+ ## Concurrent resume calls
97
+
98
+ Only one `resume()` call can continue a suspension. Before it runs anything, `resume()` atomically claims the run by moving its stored status from `suspended` to `running`. If another caller already claimed it, this call throws `WORKFLOW_RESUME_ALREADY_CLAIMED` and executes no steps, so downstream steps and their side effects run once per suspension.
99
+
100
+ This matters whenever a resume can be triggered more than once, such as an approval button pressed twice or a retried webhook. It also applies when several server instances react to the same event.
101
+
102
+ ```typescript
103
+ try {
104
+ await run.resume({ step: 'approval', resumeData: { approved: true } })
105
+ } catch (error) {
106
+ if (error.id === 'WORKFLOW_RESUME_ALREADY_CLAIMED') {
107
+ // Another caller is already continuing this run. Re-read the run state
108
+ // instead of resuming again.
109
+ }
110
+ }
111
+ ```
112
+
113
+ Over HTTP, a losing resume returns `409 Conflict`.
114
+
115
+ > **Note:** The claim is enforced atomically by storage adapters that report `supportsConcurrentUpdates()`. Adapters without atomic read-modify-write support (such as ClickHouse, Cloudflare D1, Cloudflare KV, Cloudflare Durable Objects, LanceDB, and Redis) can't enforce it, and the claim is also skipped when `shouldPersistSnapshot` excludes the `running` status.
116
+
96
117
  ## Related
97
118
 
98
119
  - [Workflows overview](https://mastra.ai/docs/workflows/overview)
package/CHANGELOG.md CHANGED
@@ -1,5 +1,13 @@
1
1
  # @mastra/mcp-docs-server
2
2
 
3
+ ## 1.2.18-alpha.0
4
+
5
+ ### Patch Changes
6
+
7
+ - Updated dependencies [[`88d14ca`](https://github.com/mastra-ai/mastra/commit/88d14cac008582a618fecc3d5c7fd3bdf4f6ddc3), [`84a5b69`](https://github.com/mastra-ai/mastra/commit/84a5b699f84d6bae0a34efe5a970d891090b9f41), [`84a5b69`](https://github.com/mastra-ai/mastra/commit/84a5b699f84d6bae0a34efe5a970d891090b9f41), [`64cd7ac`](https://github.com/mastra-ai/mastra/commit/64cd7ac22c2c7a6e6b533a4b3a9ede432700f1fb), [`84a5b69`](https://github.com/mastra-ai/mastra/commit/84a5b699f84d6bae0a34efe5a970d891090b9f41), [`038b7b4`](https://github.com/mastra-ai/mastra/commit/038b7b405cb4ac25ab3f3031334111b1f87ac112), [`4132d61`](https://github.com/mastra-ai/mastra/commit/4132d61f8367077120ee9e6420d3224dffd93c93)]:
8
+ - @mastra/core@1.60.1-alpha.0
9
+ - @mastra/mcp@1.17.1-alpha.0
10
+
3
11
  ## 1.2.17
4
12
 
5
13
  ### Patch Changes
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mastra/mcp-docs-server",
3
- "version": "1.2.17",
3
+ "version": "1.2.18-alpha.1",
4
4
  "description": "MCP server for accessing Mastra.ai documentation, changelogs, and news.",
5
5
  "type": "module",
6
6
  "main": "dist/index.js",
@@ -28,8 +28,8 @@
28
28
  "jsdom": "^26.1.0",
29
29
  "local-pkg": "^1.1.2",
30
30
  "zod": "^4.4.3",
31
- "@mastra/mcp": "^1.17.0",
32
- "@mastra/core": "1.60.0"
31
+ "@mastra/mcp": "^1.17.1-alpha.0",
32
+ "@mastra/core": "1.60.1-alpha.0"
33
33
  },
34
34
  "devDependencies": {
35
35
  "@hono/node-server": "^2.0.0",
@@ -46,7 +46,7 @@
46
46
  "typescript": "^6.0.3",
47
47
  "vitest": "4.1.10",
48
48
  "@internal/lint": "0.0.124",
49
- "@mastra/core": "1.60.0",
49
+ "@mastra/core": "1.60.1-alpha.0",
50
50
  "@internal/types-builder": "0.0.99"
51
51
  },
52
52
  "homepage": "https://mastra.ai",