@mastra/mcp-docs-server 1.2.24 → 1.2.25-alpha.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/.docs/docs/agents/processors.md +2 -2
  2. package/.docs/docs/auth/fga.md +1 -1
  3. package/.docs/docs/channels.md +1 -1
  4. package/.docs/docs/evals/custom-scorers.md +2 -2
  5. package/.docs/docs/evals/experiments.md +3 -3
  6. package/.docs/docs/evals/multi-turn.md +1 -1
  7. package/.docs/docs/guides/streaming.md +1 -1
  8. package/.docs/docs/harness/agent-controller.md +1 -1
  9. package/.docs/docs/mastra-platform/database.md +3 -1
  10. package/.docs/docs/memory/message-history.md +1 -1
  11. package/.docs/docs/memory/multi-user-threads.md +1 -1
  12. package/.docs/docs/memory/observational-memory.md +41 -1
  13. package/.docs/docs/memory/overview.md +4 -4
  14. package/.docs/docs/observability/tracing/overview.md +2 -2
  15. package/.docs/docs/server/pubsub.md +1 -1
  16. package/.docs/docs/studio/overview.md +1 -1
  17. package/.docs/models/gateways/merge-gateway.md +2 -1
  18. package/.docs/models/gateways/netlify.md +1 -1
  19. package/.docs/models/gateways/openrouter.md +1 -2
  20. package/.docs/models/index.md +1 -1
  21. package/.docs/models/providers/cloudflare-workers-ai.md +1 -2
  22. package/.docs/models/providers/crossmodel.md +1 -1
  23. package/.docs/models/providers/edenai.md +15 -7
  24. package/.docs/models/providers/hyper.md +1 -1
  25. package/.docs/models/providers/kilo.md +3 -4
  26. package/.docs/models/providers/llmgateway-providers.md +6 -6
  27. package/.docs/models/providers/llmgateway.md +6 -6
  28. package/.docs/models/providers/nano-gpt.md +5 -5
  29. package/.docs/models/providers/opencode-go.md +1 -1
  30. package/.docs/reference/agents/generate.md +1 -1
  31. package/.docs/reference/agents/network.md +1 -1
  32. package/.docs/reference/coding-agent/build-base-prompt.md +4 -4
  33. package/.docs/reference/evals/completeness.md +1 -1
  34. package/.docs/reference/evals/faithfulness.md +1 -1
  35. package/.docs/reference/evals/noise-sensitivity.md +1 -1
  36. package/.docs/reference/evals/prompt-alignment.md +2 -4
  37. package/.docs/reference/memory/observational-memory.md +1 -1
  38. package/.docs/reference/migrations/upgrade-to-v1/workflows.md +1 -1
  39. package/.docs/reference/observability/tracing/trace-query.md +169 -9
  40. package/.docs/reference/processors/regex-filter-processor.md +1 -1
  41. package/.docs/reference/rag/vector-databases.md +1 -1
  42. package/.docs/reference/streaming/agents/stream.md +7 -3
  43. package/.docs/reference/tools/mcp-server.md +1 -1
  44. package/.docs/reference/workspace/sandbox.md +2 -2
  45. package/package.json +4 -4
@@ -442,7 +442,7 @@ Caching is implemented as the [`ResponseCache`](https://mastra.ai/reference/proc
442
442
 
443
443
  ### When to use response caching
444
444
 
445
- Reach for it when the same request shape repeats across users or sessions, for example prompt templates, suggested-prompt buttons, agentic search re-asks, or guardrail LLMs that classify the same input over and over. Skip it when calls trigger external side effects through tools, since cache hits replay tool calls without re-executing them.
445
+ Use caching when identical requests recur across users or sessions. Examples include suggested-prompt buttons and repeated searches, or guardrail LLMs that classify the same input. Skip it when calls trigger external side effects through tools, since cache hits replay tool calls without re-executing them.
446
446
 
447
447
  ### Quickstart
448
448
 
@@ -546,7 +546,7 @@ The cache key is derived from the resolved `LanguageModelV2Prompt` Mastra is abo
546
546
 
547
547
  When you don't supply `key`, the processor derives one deterministically from the inputs that change the LLM's response at this step: `agentId`, `stepNumber` (so each step in a tool loop has its own cache entry), `scope`, model identity (`provider`, `modelId`, spec version), and the resolved `prompt` (post-memory + post-processors). Any change to these inputs automatically invalidates the cache.
548
548
 
549
- Multimodal prompts are included too. Image and file parts reach the key by value: the key includes a URL's full href, and a digest of the bytes for inline binary data (`Uint8Array`, `ArrayBuffer`). Requests that differ only in which image they reference therefore get different cache entries.
549
+ Multimodal prompts are included too. For image and file parts, the key includes the full URL or a digest of the bytes for inline binary data (`Uint8Array`, `ArrayBuffer`). Requests that differ only in which image they reference therefore get different cache entries.
550
550
 
551
551
  #### Customize the cache key
552
552
 
@@ -321,7 +321,7 @@ The actor signal is trusted input, so construct it server-side:
321
321
  - Establish tenant scope server-side. Built-in agent HTTP routes ignore a client-supplied `organizationId` in the request context, and the trusted-actor path requires an `organizationId` to be set.
322
322
  - Durable resume keeps its existing request-context recovery and merge behavior. This doesn't make a persisted actor trusted for a later workflow segment.
323
323
  - The tenant-scope check confirms that a trusted `organizationId` exists. It doesn't verify that `actor.agentId` belongs to that organization. When that relationship matters, verify it in `requireActor` using authoritative provider data.
324
- - Treat `actor.permissions` as an unverified claim. Resolve authoritative grants from a trusted source. A provider that enforces least privilege resolves the agent's authoritative permissions from a trusted source, for example a manifest or your FGA backend keyed by `agentId`, rather than trusting the inline values.
324
+ - Treat `actor.permissions` as an unverified claim. To enforce least privilege, resolve the agent's permissions from a trusted source, such as a manifest or your FGA backend keyed by `agentId`. Don't trust the inline values.
325
325
  - Once a provider implements `requireActor`, errors from that method stop execution. Mastra doesn't fall back to organization-only authorization.
326
326
 
327
327
  ## Related
@@ -70,7 +70,7 @@ export const mastra = new Mastra({
70
70
 
71
71
  ## Webhook routes
72
72
 
73
- Platforms send channel activity to Mastra through webhooks. A webhook is an HTTP endpoint that the platform calls when something happens, such as a new message, a mention, or a user selecting "Approve" on an interactive tool approval card. This is how your agent receives a new message and starts processing it, plus responds in the same channel.
73
+ Platforms send channel activity to Mastra through webhooks. A webhook is an HTTP endpoint that the platform calls when something happens, such as a new message, a mention, or a user selecting "Approve" on an interactive tool approval card. The webhook delivers the message to your agent for processing, and the agent responds in the same channel.
74
74
 
75
75
  Mastra registers a webhook route for each configured adapter and handles the request for you:
76
76
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Custom scorers
6
6
 
7
- Mastra provides a unified `createScorer` factory that allows you to build custom evaluation logic using either JavaScript functions or LLM-based prompt objects for each step. This flexibility lets you choose the best approach for each part of your evaluation pipeline.
7
+ Use `createScorer` to build custom evaluation logic. Each step can use a JavaScript function or an LLM-based prompt object, depending on what you're evaluating.
8
8
 
9
9
  ## The four-step pipeline
10
10
 
@@ -15,7 +15,7 @@ All scorers in Mastra follow a consistent four-step evaluation pipeline:
15
15
  3. **generateScore** (required): Convert analysis into a numerical score
16
16
  4. **generateReason** (optional): Generate human-readable explanations
17
17
 
18
- Each step can use either **functions** or **prompt objects** (LLM-based evaluation), giving you the flexibility to combine deterministic algorithms with AI judgment as needed.
18
+ Combine **functions** and **prompt objects** within a scorer to use deterministic checks for some steps and LLM-based evaluation for others.
19
19
 
20
20
  ## Functions vs prompt objects
21
21
 
@@ -315,7 +315,7 @@ Teardown failures are logged rather than propagated. By the time `afterEach` run
315
315
 
316
316
  When an experiment runs an agent that calls side-effecting tools, attach static tool mocks to individual dataset items to make the run deterministic. During the experiment, a mocked tool returns its declared output instead of executing. Tools without a mock on the item run live by default.
317
317
 
318
- Mocks live on the dataset item, so they version with the row and travel with the test case. Each mock declares a tool name, the arguments it expects, and the output to return:
318
+ Mocks are stored and versioned with the dataset item. Each mock specifies the tool name and expected arguments, along with the output to return:
319
319
 
320
320
  ```typescript
321
321
  await dataset.addItem({
@@ -357,7 +357,7 @@ The item value takes precedence over the experiment value. A denied call fails w
357
357
 
358
358
  ### Matching and consumption
359
359
 
360
- Arguments are matched strictly: object key order is ignored and array order is substantial, plus there is no type coercion. A mock is served only when the agent calls the tool with arguments that deep-equal the mock's `args`.
360
+ Arguments are matched strictly: object key order is ignored, but array order matters. Values aren't coerced to other types. A mock is served only when the agent calls the tool with arguments that deep-equal the mock's `args`.
361
361
 
362
362
  When an item declares several mocks for the same tool and arguments, they're consumed in order, the first call gets the first mock, the next call gets the second, and so on. Ordering is tracked per `(toolName, args)` group and is independent across different arguments.
363
363
 
@@ -390,7 +390,7 @@ While mock interception is active, the agent's tools execute sequentially so rep
390
390
 
391
391
  ### Diagnostics
392
392
 
393
- Each item result carries a `toolMockReport` describing what the run did with the item's mocks:
393
+ Each item result includes a `toolMockReport` listing which mocks were used and which tool calls ran without a mock:
394
394
 
395
395
  ```typescript
396
396
  for (const item of summary.results) {
@@ -191,7 +191,7 @@ result.turnResults // per-turn gate/threshold/scorer outcomes
191
191
 
192
192
  Semantics:
193
193
 
194
- - A per-turn gate or scorer sees **only that turn's** `run.input` and `run.output`: never the accumulated conversation. This fixes both blind spots of `inputs`: the wrong turn can't satisfy a check, and `run.input` is correct for each turn.
194
+ - A per-turn gate or scorer receives **only that turn's** `run.input` and `run.output`, rather than the accumulated conversation. A different turn's output can't satisfy the check, and `run.input` contains the input for the turn being evaluated.
195
195
  - Per-turn outcomes fold into the [verdict](https://mastra.ai/docs/evals/gates-and-verdicts): a failing turn gate makes the verdict `failed`. A missed turn threshold (with gates passing) makes it `scored`.
196
196
  - `result.turnResults[i]` reports each turn's `gateResults`, `thresholdResults`, and `scores`, so a failure points at the exact turn. Across multiple conversations, turn results are averaged by turn index.
197
197
  - A turn with no `gates` or `scorers` advances the conversation.
@@ -108,7 +108,7 @@ Visit [Run.stream()](https://mastra.ai/reference/streaming/workflows/stream) for
108
108
 
109
109
  ### Output from `Run.stream()`
110
110
 
111
- The event structure includes `runId` and `from` at the top level, making it easier to identify and track workflow runs without digging into the payload.
111
+ Events include `runId` and `from` at the top level, so you can identify the workflow run without inspecting the payload.
112
112
 
113
113
  ```typescript
114
114
  {
@@ -22,7 +22,7 @@ Use the Agent Controller when your application needs:
22
22
  - Subagent orchestration to delegate focused subtasks with constrained tools
23
23
  - Persistent threads and selected thread settings across restarts, with isolated live state for each Session
24
24
 
25
- You could assemble all of this yourself on top of the [Agent class](https://mastra.ai/docs/agents/overview), which exposes the full agent loop, tools, and memory. The AgentController provides opinionated defaults for an ongoing session where the agent acts as a collaborator rather than a one-shot endpoint. Reach for the Agent class directly when you want full control or a request-response call. Reach for the AgentController when you want the collaborative-session model without building the runtime around it.
25
+ You could assemble all of this yourself on top of the [Agent class](https://mastra.ai/docs/agents/overview), which exposes the full agent loop, tools, and memory. The AgentController provides opinionated defaults for an ongoing session where the agent acts as a collaborator rather than a one-shot endpoint. Use the Agent class directly for full control or request-response calls. Choose AgentController for ongoing sessions without building your own session runtime.
26
26
 
27
27
  ## Quickstart
28
28
 
@@ -97,7 +97,9 @@ Databases attached from project settings are project-scoped. Use the [CLI](#atta
97
97
 
98
98
  ## Connect from your code
99
99
 
100
- When a database is `ready`, the provider has finished provisioning and the platform has injected connection details as managed environment variables. Check status in **Project Settings → Database**, each attached database shows `provisioning` while setup runs in the background, then `ready` when you can connect. Open a `ready` database to view its environment variables and a copy-pasteable code snippet. Wire those variables into a Mastra storage adapter, with no manual configuration required.
100
+ Check the database status in **Project Settings → Database**. An attached database shows `provisioning` while setup runs in the background. Once it's `ready`, the platform has injected its connection details as managed environment variables.
101
+
102
+ Open a `ready` database to view those variables and a code snippet. Use the variables to configure a Mastra storage adapter, as shown below.
101
103
 
102
104
  ### Turso (LibSQL)
103
105
 
@@ -178,7 +178,7 @@ const agent = mastra.getAgentById('test-agent')
178
178
  const memory = await agent.getMemory()
179
179
  ```
180
180
 
181
- The `Memory` instance gives you access to functions for listing threads and recalling messages, plus cloning conversations, and more.
181
+ Use the `Memory` instance to query stored threads and messages or clone a conversation.
182
182
 
183
183
  ## Querying
184
184
 
@@ -160,7 +160,7 @@ const memory = new Memory({
160
160
 
161
161
  OM requires a storage adapter that supports it: `@mastra/libsql`, `@mastra/pg`, `@mastra/mongodb`, or `@mastra/oracledb`.
162
162
 
163
- > **Note:** If you switch the Observer to a weaker model and see facts collapse to a generic `User`, use [`observation.instruction`](https://mastra.ai/reference/memory/observational-memory) to teach the Observer how to read the `<turn>` tag.
163
+ > **Note:** If you switch the Observer to a less capable model and see facts attributed to a generic `User` instead of individual participants, use [`observation.instruction`](https://mastra.ai/reference/memory/observational-memory) to explain how to interpret the `<turn>` tag.
164
164
 
165
165
  ### With working memory
166
166
 
@@ -718,7 +718,7 @@ Buffered observations also include continuation hints, a suggested next response
718
718
 
719
719
  When message production outpaces the Observer, the `blockAfter` safety threshold allows activation to overshoot the retention target instead of using fewer chunks. Activation still uses no more chunks than needed to reach the target, and the default settings remain unaffected. A synchronous observation runs when the `messageTokens` threshold is reached and buffered activation didn't happen. Buffered activation usually preserves a minimum remaining context (the smaller of \~1k tokens or the configured retention floor), but a single buffered chunk that covers the whole pending window still activates and can leave less.
720
720
 
721
- Reflection works similarly, the Reflector runs in the background when observations reach a fraction of the reflection threshold.
721
+ Reflection works similarly: the Reflector runs in the background when observations reach a fraction of the reflection threshold.
722
722
 
723
723
  ### Settings
724
724
 
@@ -809,6 +809,46 @@ const memory = new Memory({
809
809
  - `previousObserverTokens: 0` → omit previous observations completely.
810
810
  - `previousObserverTokens: false` → disable truncation and keep full previous observations.
811
811
 
812
+ ## Hooks
813
+
814
+ OM exposes two kinds of config-level hooks on `observationalMemory.hooks`:
815
+
816
+ - **Lifecycle hooks** (`onObservationStart`, `onObservationEnd`, `onReflectionStart`, `onReflectionEnd`) are telemetry callbacks. They receive `threadId`, `resourceId`, and `trigger`, and the end hooks also receive the model call's `usage`, `providerMetadata`, and any `error`. They never change what OM stores.
817
+ - **Transform hooks** (`beforeObservation`, `afterObservation`, `beforeReflection`, `afterReflection`) intercept the data flowing through a cycle. Return `void` to pass the input through unchanged, or return a replacement to change what the Observer/Reflector sees or what gets persisted.
818
+
819
+ ```typescript
820
+ const memory = new Memory({
821
+ options: {
822
+ observationalMemory: {
823
+ model: 'google/gemini-2.5-flash',
824
+ hooks: {
825
+ // Drop or redact messages before the Observer sees them.
826
+ beforeObservation: ({ messages }) => ({
827
+ messages: messages.filter(m => !isSensitive(m)),
828
+ }),
829
+ // Rewrite observations before they are persisted.
830
+ afterObservation: ({ observations, threadId }) => ({
831
+ observations: redact(observations),
832
+ }),
833
+ // Rewrite the text the Reflector condenses, or its output.
834
+ beforeReflection: ({ observations }) => ({
835
+ observations: stripInternalNotes(observations),
836
+ }),
837
+ afterReflection: async ({ observations, resourceId }) => {
838
+ await syncToExternalStore(resourceId, observations)
839
+ },
840
+ },
841
+ },
842
+ },
843
+ })
844
+ ```
845
+
846
+ Transform hooks are always awaited, on every path (manual `observe()`/`reflect()`, turn-synchronous observation, and async buffering). If `beforeObservation` returns an empty `messages` array, the Observer model call is skipped and the filtered messages are still marked as observed. If a transform hook throws, the cycle fails before committing the transformed observation or reflection text. This doesn't roll back extractor callbacks or other side effects that have already run.
847
+
848
+ `afterObservation` and `afterReflection` replace only the observation or reflection text. They don't recompute or redact the separate structured extractor results stored in thread metadata. Reflection extraction and its callbacks run before `afterReflection`, so rewriting the reflection doesn't rerun those callbacks. Use extractor configuration and callbacks to control structured values. Don't treat an after hook as a redaction boundary for all cycle data.
849
+
850
+ Because hooks receive `threadId` and `resourceId`, you can also use them to update [working memory](https://mastra.ai/docs/memory/working-memory) via `memory.updateWorkingMemory()` during a cycle. These external updates aren't atomic with the OM text commit.
851
+
812
852
  ## Migrating existing threads
813
853
 
814
854
  No manual migration needed. OM reads existing messages and observes them lazily when thresholds are exceeded.
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Memory
6
6
 
7
- Memory enables your agent to remember user messages and agent replies, and tool results across interactions, giving it the context it needs to stay consistent, maintain conversation flow, plus produce better answers over time.
7
+ Memory gives your agent access to earlier messages and tool results. The agent can use this context to answer follow-up questions and recall information from previous interactions.
8
8
 
9
9
  Mastra agents can be configured to store [message history](https://mastra.ai/docs/memory/message-history). Additionally, you can enable:
10
10
 
@@ -153,7 +153,7 @@ To list all threads for a resource, or retrieve a specific thread, [use the memo
153
153
 
154
154
  ## Observational Memory
155
155
 
156
- For long-running conversations, raw message history grows until it fills the context window, degrading agent performance. [Observational Memory](https://mastra.ai/docs/memory/observational-memory) solves this by running background agents that compress old messages into dense observations, keeping the context window small while preserving long-term memory.
156
+ Long conversations can fill the context window with raw message history and reduce agent performance. [Observational Memory](https://mastra.ai/docs/memory/observational-memory) uses background agents to compress older messages into observations. This reduces the context used by message history while retaining information for later turns.
157
157
 
158
158
  **For AI agents:** Using Observational Memory requires a storage provider! You either need to set it on the Mastra instance at `src/mastra/index.ts` or pass it to the Agent constructor.
159
159
 
@@ -176,7 +176,7 @@ See [Observational Memory](https://mastra.ai/docs/memory/observational-memory) f
176
176
 
177
177
  ## What the model sees
178
178
 
179
- Each memory feature is added to either the system messages or the conversation messages in the request sent to the model. The layers depend on the features you've enabled. Working memory and semantic recall only appear when configured. The same applies to Observational Memory, while message history is on by default. The diagram shows where each enabled layer is placed in the request. The list below describes what each layer contributes:
179
+ Memory adds context to the system messages or conversation messages sent to the model. Message history is enabled by default. Other memory features contribute context only when configured. The diagram shows where each feature adds its context, and the list below explains what it contributes:
180
180
 
181
181
  ![Diagram showing how Mastra assembles the model context: system messages containing agent instructions, call-time system messages, working memory, cross-thread semantic recall, and Observational Memory, followed by conversation messages where message history and same-thread semantic recall interleave by timestamp, then call-time context messages, and finally the new user message](/img/memory/memory-context-window-light.svg)
182
182
 
@@ -202,7 +202,7 @@ Each delegation creates a fresh `threadId` and a deterministic `resourceId` for
202
202
 
203
203
  > **Note:** Title generation (`generateTitle`) is a top-level thread concern and **isn't** applied to inherited subagent threads. Because each delegation creates an ephemeral thread that no one sees, running title generation for it would waste an LLM call per delegation. To generate titles for a subagent's own threads, give that subagent its own memory configuration.
204
204
 
205
- The supervisor forwards its conversation context to the subagent so it has enough background to complete the task. Only the delegation prompt and the subagent's response are saved, the full parent conversation isn't stored. You can control which messages reach the subagent with the [`messageFilter`](https://mastra.ai/docs/subagents) callback.
205
+ The supervisor forwards its conversation context to the subagent so it has enough background to complete the task. Only the delegation prompt and the subagent's response are saved; the full parent conversation isn't stored. You can control which messages reach the subagent with the [`messageFilter`](https://mastra.ai/docs/subagents) callback.
206
206
 
207
207
  > **Note:** Subagent resource IDs are always suffixed with the agent name (`{parentResourceId}-{agentName}`). Different subagents under the same supervisor never share a resource ID through delegation.
208
208
 
@@ -98,7 +98,7 @@ The `sampling` option allows you to control which traces are collected, helping
98
98
 
99
99
  ## Adding custom metadata
100
100
 
101
- Custom metadata allows you to attach additional context to your traces, making it easier to debug issues and understand system behavior in production.
101
+ Add custom metadata to record application-specific context in your traces for debugging production issues.
102
102
 
103
103
  Metadata can include business logic and performance metrics. It can also carry user context or any other information that explains what happened during execution.
104
104
 
@@ -764,7 +764,7 @@ The trace ID is only available when tracing is enabled. If tracing is disabled o
764
764
 
765
765
  ## Integrating with external tracing systems
766
766
 
767
- When running Mastra agents or workflows within applications that have existing distributed tracing (OpenTelemetry, Datadog, etc.), you can connect Mastra traces to your parent trace context. This creates a unified view of your entire request flow, making it easier to understand how Mastra operations fit into the broader system.
767
+ If your application already uses distributed tracing, such as OpenTelemetry or Datadog, you can connect Mastra traces to the parent trace context. This lets you follow a request through your application and its Mastra agent or workflow calls.
768
768
 
769
769
  ### Passing external trace IDs
770
770
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # PubSub
6
6
 
7
- Mastra uses a publish/subscribe (pub/sub) system as its internal event bus. Components publish events to topics, and other components subscribe to those topics to react. The backend you configure decides how far those events travel: within one process or across processes on one host, or alternatively across separate instances.
7
+ Mastra uses a publish/subscribe (pub/sub) system as its internal event bus. Components publish events to topics, and other components subscribe to those topics to react. The backend determines whether events are delivered within a single process, between processes on one host, or across separate instances.
8
8
 
9
9
  Set the backend once on the `Mastra` instance for use throughout the system. The default is an in-process backend that requires no setup.
10
10
 
@@ -67,7 +67,7 @@ Use [Agent Builder](https://agent-builder.mastra.ai) to create and manage fully
67
67
 
68
68
  Visualize your workflow as a graph and run it step by step with a custom input. During execution, the interface updates in real time to show the active step and the path taken.
69
69
 
70
- When running a workflow, you can also view detailed traces showing tool calls and raw JSON outputs, plus any errors that might have occurred along the way.
70
+ Workflow traces show tool calls and raw JSON outputs, along with any execution errors.
71
71
 
72
72
  ### Processors
73
73
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Merge Gateway logo](https://models.dev/logos/merge-gateway.svg)Merge Gateway
6
6
 
7
- Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 180 models through Mastra's model router.
7
+ Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 181 models through Mastra's model router.
8
8
 
9
9
  Learn more in the [Merge Gateway documentation](https://docs.merge.dev/merge-gateway).
10
10
 
@@ -95,6 +95,7 @@ ANTHROPIC_API_KEY=ant-...
95
95
  | `google/gemini-flash-lite-latest` |
96
96
  | `google/gemma-4-26b-a4b-it` |
97
97
  | `google/gemma-4-31b-it` |
98
+ | `meta/llama-3.1-70b-instruct` |
98
99
  | `meta/llama-3.1-8b-instruct` |
99
100
  | `meta/llama-3.3-70b-instruct` |
100
101
  | `meta/muse-spark-1.1` |
@@ -197,7 +197,6 @@ ANTHROPIC_API_KEY=ant-...
197
197
  | `openrouter/nousresearch/hermes-3-llama-3.1-405b` |
198
198
  | `openrouter/nousresearch/hermes-3-llama-3.1-70b` |
199
199
  | `openrouter/nousresearch/hermes-4-405b` |
200
- | `openrouter/nousresearch/hermes-4-70b` |
201
200
  | `openrouter/nvidia/nemotron-3-nano-30b-a3b` |
202
201
  | `openrouter/nvidia/nemotron-3-super-120b-a12b` |
203
202
  | `openrouter/nvidia/nemotron-3-ultra-550b-a55b` |
@@ -242,6 +241,7 @@ ANTHROPIC_API_KEY=ant-...
242
241
  | `openrouter/qwen/qwen3.6-35b-a3b` |
243
242
  | `openrouter/qwen/qwen3.8-2.4t-a95b` |
244
243
  | `openrouter/qwen/qwen3.8-27b` |
244
+ | `openrouter/qwen/qwen3.8-flash` |
245
245
  | `openrouter/rekaai/reka-edge` |
246
246
  | `openrouter/rekaai/reka-flash-3` |
247
247
  | `openrouter/relace/relace-apply-3` |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![OpenRouter logo](https://models.dev/logos/openrouter.svg)OpenRouter
6
6
 
7
- OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 358 models through Mastra's model router.
7
+ OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 357 models through Mastra's model router.
8
8
 
9
9
  Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
10
10
 
@@ -208,7 +208,6 @@ ANTHROPIC_API_KEY=ant-...
208
208
  | `nousresearch/hermes-3-llama-3.1-405b` |
209
209
  | `nousresearch/hermes-3-llama-3.1-70b` |
210
210
  | `nousresearch/hermes-4-405b` |
211
- | `nousresearch/hermes-4-70b` |
212
211
  | `nvidia/nemotron-3-nano-30b-a3b` |
213
212
  | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` |
214
213
  | `nvidia/nemotron-3-super-120b-a12b` |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Model Providers
6
6
 
7
- Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7118 models from 200 providers through a single API.
7
+ Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7124 models from 200 providers through a single API.
8
8
 
9
9
  ## Features
10
10
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Cloudflare Workers AI logo](https://models.dev/logos/cloudflare-workers-ai.svg)Cloudflare Workers AI
6
6
 
7
- Access 27 Cloudflare Workers AI models through Mastra's model router. Authentication is handled automatically using the `CLOUDFLARE_API_KEY` environment variable. Configure `CLOUDFLARE_ACCOUNT_ID` as well.
7
+ Access 26 Cloudflare Workers AI models through Mastra's model router. Authentication is handled automatically using the `CLOUDFLARE_API_KEY` environment variable. Configure `CLOUDFLARE_ACCOUNT_ID` as well.
8
8
 
9
9
  Learn more in the [Cloudflare Workers AI documentation](https://developers.cloudflare.com/workers-ai/models/).
10
10
 
@@ -64,7 +64,6 @@ for await (const chunk of stream) {
64
64
  | `cloudflare-workers-ai/@cf/qwen/qwq-32b` | 24K | | | | | | $0.66 | $1 |
65
65
  | `cloudflare-workers-ai/@cf/zai-org/glm-4.7-flash` | 131K | | | | | | $0.06 | $0.40 |
66
66
  | `cloudflare-workers-ai/@cf/zai-org/glm-5.2` | 262K | | | | | | $1 | $4 |
67
- | `cloudflare-workers-ai/@cf/zai-org/glm-5.3` | 1.3M | | | | | | $1 | $4 |
68
67
  | `cloudflare-workers-ai/@cf/zai-org/glm-5.3-flash` | 1.3M | | | | | | $0.15 | $0.50 |
69
68
 
70
69
  Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
@@ -95,7 +95,7 @@ for await (const chunk of stream) {
95
95
  | `crossmodel/z-ai/glm-5.1` | 200K | | | | | | $1 | $4 |
96
96
  | `crossmodel/z-ai/glm-5.2` | 1.0M | | | | | | $1 | $4 |
97
97
  | `crossmodel/z-ai/glm-5.3` | 1.0M | | | | | | $1 | $4 |
98
- | `crossmodel/z-ai/glm-5.3-flash` | 1.0M | | | | | | $0.07 | $0.25 |
98
+ | `crossmodel/z-ai/glm-5.3-flash` | 1.0M | | | | | | $0.15 | $0.50 |
99
99
 
100
100
  Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
101
101
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Eden AI logo](https://models.dev/logos/edenai.svg)Eden AI
6
6
 
7
- Access 250 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
7
+ Access 258 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Eden AI documentation](https://docs.edenai.co).
10
10
 
@@ -19,7 +19,7 @@ const agent = new Agent({
19
19
  id: "my-agent",
20
20
  name: "My Agent",
21
21
  instructions: "You are a helpful assistant",
22
- model: "edenai/amazon/moonshot.kimi-k2-thinking"
22
+ model: "edenai/amazon/amazon.nova-lite-v1:0"
23
23
  });
24
24
 
25
25
  // Generate a response
@@ -38,6 +38,14 @@ for await (const chunk of stream) {
38
38
 
39
39
  | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
40
40
  | ---------------------------------------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
41
+ | `edenai/amazon/amazon.nova-lite-v1:0` | 300K | | | | | | $0.06 | $0.24 |
42
+ | `edenai/amazon/amazon.nova-lite-v1:0@us` | 300K | | | | | | $0.06 | $0.24 |
43
+ | `edenai/amazon/amazon.nova-micro-v1:0` | 128K | | | | | | $0.04 | $0.14 |
44
+ | `edenai/amazon/amazon.nova-micro-v1:0@us` | 128K | | | | | | $0.04 | $0.14 |
45
+ | `edenai/amazon/amazon.nova-pro-v1:0` | 300K | | | | | | $0.80 | $3 |
46
+ | `edenai/amazon/amazon.nova-pro-v1:0@us` | 300K | | | | | | $0.80 | $3 |
47
+ | `edenai/amazon/mistral.pixtral-large-2502-v1:0` | 128K | | | | | | $2 | $6 |
48
+ | `edenai/amazon/mistral.pixtral-large-2502-v1:0@us` | 128K | | | | | | $2 | $6 |
41
49
  | `edenai/amazon/moonshot.kimi-k2-thinking` | 128K | | | | | | $0.60 | $3 |
42
50
  | `edenai/amazon/moonshotai.kimi-k2.5` | 262K | | | | | | $0.60 | $3 |
43
51
  | `edenai/amazon/zai.glm-4.7-flash` | 200K | | | | | | $0.07 | $0.40 |
@@ -137,8 +145,8 @@ for await (const chunk of stream) {
137
145
  | `edenai/google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
138
146
  | `edenai/groq/openai/gpt-oss-120b` | 131K | | | | | | $0.15 | $0.60 |
139
147
  | `edenai/groq/openai/gpt-oss-20b` | 131K | | | | | | $0.07 | $0.30 |
140
- | `edenai/ionos/meta-llama/Llama-3.3-70B-Instruct` | 128K | | | | | | $0.75 | $0.75 |
141
- | `edenai/ionos/openai/gpt-oss-120b` | 131K | | | | | | $0.17 | $0.75 |
148
+ | `edenai/ionos/meta-llama/Llama-3.3-70B-Instruct` | 128K | | | | | | $0.76 | $0.76 |
149
+ | `edenai/ionos/openai/gpt-oss-120b` | 131K | | | | | | $0.17 | $0.76 |
142
150
  | `edenai/minimax/MiniMax-M2` | 205K | | | | | | $0.30 | $1 |
143
151
  | `edenai/minimax/MiniMax-M2.1` | 205K | | | | | | $0.30 | $1 |
144
152
  | `edenai/minimax/MiniMax-M2.5` | 205K | | | | | | $0.30 | $1 |
@@ -232,7 +240,7 @@ for await (const chunk of stream) {
232
240
  | `edenai/qwen/qwen3.8-max` | 1.0M | | | | | | $2 | $6 |
233
241
  | `edenai/qwen/qwen3.8-max-0902` | 1.0M | | | | | | $2 | $6 |
234
242
  | `edenai/qwen/qwq-plus` | 131K | | | | | | $0.80 | $2 |
235
- | `edenai/scaleway/deepseek-v4-flash-0731` | 256K | | | | | | $0.46 | $0.93 |
243
+ | `edenai/scaleway/deepseek-v4-flash-0731` | 256K | | | | | | $0.47 | $0.93 |
236
244
  | `edenai/scaleway/gpt-oss-120b` | 128K | | | | | | $0.17 | $0.70 |
237
245
  | `edenai/scaleway/llama-3.3-70b-instruct` | 128K | | | | | | $1 | $1 |
238
246
  | `edenai/tensorx/deepseek/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.25 | $0.30 |
@@ -301,7 +309,7 @@ const agent = new Agent({
301
309
  name: "custom-agent",
302
310
  model: {
303
311
  url: "https://api.edenai.run/v3",
304
- id: "edenai/amazon/moonshot.kimi-k2-thinking",
312
+ id: "edenai/amazon/amazon.nova-lite-v1:0",
305
313
  apiKey: process.env.EDENAI_API_KEY,
306
314
  headers: {
307
315
  "X-Custom-Header": "value"
@@ -320,7 +328,7 @@ const agent = new Agent({
320
328
  const useAdvanced = requestContext.task === "complex";
321
329
  return useAdvanced
322
330
  ? "edenai/zai/glm-5v-turbo"
323
- : "edenai/amazon/moonshot.kimi-k2-thinking";
331
+ : "edenai/amazon/amazon.nova-lite-v1:0";
324
332
  }
325
333
  });
326
334
  ```
@@ -48,7 +48,7 @@ for await (const chunk of stream) {
48
48
  | `hyper/glm-5.2` | 1.0M | | | | | | $2 | $5 |
49
49
  | `hyper/glm-5.3` | 1.0M | | | | | | $2 | $5 |
50
50
  | `hyper/glm-5.3-flash` | 1.0M | | | | | | $0.16 | $0.54 |
51
- | `hyper/gpt-oss-120b` | 128K | | | | | | $0.18 | $0.68 |
51
+ | `hyper/gpt-oss-120b` | 128K | | | | | | $0.18 | $0.61 |
52
52
  | `hyper/inkling` | 1.0M | | | | | | $1 | $4 |
53
53
  | `hyper/kimi-k2-thinking` | 262K | | | | | | $0.60 | $3 |
54
54
  | `hyper/kimi-k2.5` | 262K | | | | | | $0.56 | $3 |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Kilo Gateway logo](https://models.dev/logos/kilo.svg)Kilo Gateway
6
6
 
7
- Access 367 Kilo Gateway models through Mastra's model router. Authentication is handled automatically using the `KILO_API_KEY` environment variable.
7
+ Access 366 Kilo Gateway models through Mastra's model router. Authentication is handled automatically using the `KILO_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Kilo Gateway documentation](https://kilo.ai).
10
10
 
@@ -45,7 +45,7 @@ for await (const chunk of stream) {
45
45
  | `kilo/~deepseek/deepseek-v4-flash-latest` | 1.0M | | | | | | $0.05 | $0.16 |
46
46
  | `kilo/~google/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
47
47
  | `kilo/~google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
48
- | `kilo/~moonshotai/kimi-latest` | 1.0M | | | | | | $3 | $13 |
48
+ | `kilo/~moonshotai/kimi-latest` | 1.0M | | | | | | $2 | $12 |
49
49
  | `kilo/~openai/gpt-latest` | 1.1M | | | | | | $2 | $10 |
50
50
  | `kilo/~openai/gpt-mini-latest` | 400K | | | | | | $0.75 | $5 |
51
51
  | `kilo/~x-ai/grok-latest` | 500K | | | | | | $2 | $6 |
@@ -211,7 +211,6 @@ for await (const chunk of stream) {
211
211
  | `kilo/nousresearch/hermes-3-llama-3.1-405b` | 131K | | | | | | $1 | $1 |
212
212
  | `kilo/nousresearch/hermes-3-llama-3.1-70b` | 131K | | | | | | $0.70 | $0.70 |
213
213
  | `kilo/nousresearch/hermes-4-405b` | 131K | | | | | | $1 | $3 |
214
- | `kilo/nousresearch/hermes-4-70b` | 131K | | | | | | $0.13 | $0.40 |
215
214
  | `kilo/nvidia/nemotron-3-nano-30b-a3b` | 262K | | | | | | $0.05 | $0.20 |
216
215
  | `kilo/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` | 256K | | | | | | — | — |
217
216
  | `kilo/nvidia/nemotron-3-super-120b-a12b` | 262K | | | | | | $0.09 | $0.40 |
@@ -394,7 +393,7 @@ for await (const chunk of stream) {
394
393
  | `kilo/z-ai/glm-4.5` | 131K | | | | | | $0.60 | $2 |
395
394
  | `kilo/z-ai/glm-4.5-air` | 131K | | | | | | $0.13 | $0.85 |
396
395
  | `kilo/z-ai/glm-4.5v` | 66K | | | | | | $0.60 | $2 |
397
- | `kilo/z-ai/glm-4.6` | 205K | | | | | | $0.55 | $2 |
396
+ | `kilo/z-ai/glm-4.6` | 198K | | | | | | $0.43 | $2 |
398
397
  | `kilo/z-ai/glm-4.6v` | 131K | | | | | | $0.30 | $0.90 |
399
398
  | `kilo/z-ai/glm-4.7` | 203K | | | | | | $0.40 | $2 |
400
399
  | `kilo/z-ai/glm-4.7-flash` | 131K | | | | | | $0.06 | $0.40 |
@@ -347,13 +347,13 @@ for await (const chunk of stream) {
347
347
  | `llmgateway-providers/runware/kimi-k2.6` | 262K | | | | | | $0.60 | $3 |
348
348
  | `llmgateway-providers/runware/kimi-k3` | 1.0M | | | | | | $3 | $15 |
349
349
  | `llmgateway-providers/sakana/fugu-ultra` | 1.0M | | | | | | $5 | $30 |
350
- | `llmgateway-providers/scx-ai-gp/glm-5.2` | 1.0M | | | | | | $0.55 | $2 |
351
- | `llmgateway-providers/scx-ai-gp/glm-5.2-fast` | 1.0M | | | | | | $2 | $6 |
350
+ | `llmgateway-providers/scx-ai-gp/glm-5.2` | 1.0M | | | | | | $0.80 | $3 |
351
+ | `llmgateway-providers/scx-ai-gp/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
352
352
  | `llmgateway-providers/scx-ai-gp/glm-5.3` | 1.0M | | | | | | $1 | $4 |
353
- | `llmgateway-providers/scx-ai-gp/glm-5.3-flash` | 1.0M | | | | | | $0.13 | $0.40 |
354
- | `llmgateway-providers/scx-ai-gp/kimi-k2.7-code` | 262K | | | | | | $0.89 | $4 |
355
- | `llmgateway-providers/scx-ai-gp/kimi-k3` | 1.0M | | | | | | $3 | $14 |
356
- | `llmgateway-providers/scx-ai-gp/qwen3.8-max` | 1.0M | | | | | | $2 | $5 |
353
+ | `llmgateway-providers/scx-ai-gp/glm-5.3-flash` | 1.0M | | | | | | $0.09 | $0.25 |
354
+ | `llmgateway-providers/scx-ai-gp/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
355
+ | `llmgateway-providers/scx-ai-gp/kimi-k3` | 1.0M | | | | | | $4 | $18 |
356
+ | `llmgateway-providers/scx-ai-gp/qwen3.8-max` | 1.0M | | | | | | $2 | $6 |
357
357
  | `llmgateway-providers/scx-ai/gemma-4-31b-it` | 131K | | | | | | $0.30 | $0.91 |
358
358
  | `llmgateway-providers/scx-ai/gpt-oss-120b` | 131K | | | | | | $0.17 | $0.55 |
359
359
  | `llmgateway-providers/scx-ai/llama-4-maverick-17b-instruct` | 131K | | | | | | $0.53 | $2 |
@@ -89,10 +89,10 @@ for await (const chunk of stream) {
89
89
  | `llmgateway/glm-4.7-flashx` | 200K | | | | | | $0.07 | $0.40 |
90
90
  | `llmgateway/glm-5` | 203K | | | | | | $0.72 | $2 |
91
91
  | `llmgateway/glm-5.1` | 205K | | | | | | $0.93 | $3 |
92
- | `llmgateway/glm-5.2` | 1.0M | | | | | | $0.55 | $2 |
93
- | `llmgateway/glm-5.2-fast` | 1.0M | | | | | | $2 | $6 |
92
+ | `llmgateway/glm-5.2` | 1.0M | | | | | | $0.80 | $3 |
93
+ | `llmgateway/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
94
94
  | `llmgateway/glm-5.3` | 1.0M | | | | | | $1 | $4 |
95
- | `llmgateway/glm-5.3-flash` | 1.0M | | | | | | $0.10 | $0.25 |
95
+ | `llmgateway/glm-5.3-flash` | 1.0M | | | | | | $0.09 | $0.25 |
96
96
  | `llmgateway/gpt-3.5-turbo` | 16K | | | | | | $0.50 | $2 |
97
97
  | `llmgateway/gpt-4` | 8K | | | | | | $30 | $60 |
98
98
  | `llmgateway/gpt-4-turbo` | 128K | | | | | | $10 | $30 |
@@ -142,9 +142,9 @@ for await (const chunk of stream) {
142
142
  | `llmgateway/kimi-k2-thinking` | 262K | | | | | | $0.60 | $3 |
143
143
  | `llmgateway/kimi-k2.5` | 262K | | | | | | $0.41 | $2 |
144
144
  | `llmgateway/kimi-k2.6` | 262K | | | | | | $0.60 | $3 |
145
- | `llmgateway/kimi-k2.7-code` | 262K | | | | | | $0.89 | $4 |
145
+ | `llmgateway/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
146
146
  | `llmgateway/kimi-k2.7-code-highspeed` | 262K | | | | | | $2 | $8 |
147
- | `llmgateway/kimi-k3` | 1.0M | | | | | | $3 | $14 |
147
+ | `llmgateway/kimi-k3` | 1.0M | | | | | | $3 | $15 |
148
148
  | `llmgateway/kimi-k3-fast` | 1.0M | | | | | | $5 | $23 |
149
149
  | `llmgateway/ling-3.0-flash` | 262K | | | | | | $0.06 | $0.18 |
150
150
  | `llmgateway/llama-3-70b-instruct` | 8K | | | | | | $0.51 | $0.74 |
@@ -214,7 +214,7 @@ for await (const chunk of stream) {
214
214
  | `llmgateway/qwen3.7-plus` | 1.0M | | | | | | $0.40 | $2 |
215
215
  | `llmgateway/Qwen3.8-27B` | 33K | | | | | | $0.20 | $2 |
216
216
  | `llmgateway/qwen3.8-flash` | 1.0M | | | | | | $0.15 | $0.47 |
217
- | `llmgateway/qwen3.8-max` | 1.0M | | | | | | $2 | $5 |
217
+ | `llmgateway/qwen3.8-max` | 1.0M | | | | | | $2 | $6 |
218
218
  | `llmgateway/qwen35-397b-a17b` | 262K | | | | | | $0.60 | $4 |
219
219
  | `llmgateway/seed-1-6-250615` | 256K | | | | | | $0.25 | $2 |
220
220
  | `llmgateway/seed-1-6-250915` | 256K | | | | | | $0.25 | $2 |
@@ -42,6 +42,7 @@ for await (const chunk of stream) {
42
42
  | `nano-gpt/abliteration-ai/abliterated-model` | 262K | | | | | | $3 | $3 |
43
43
  | `nano-gpt/abliteration-ai/abliterated-model-large` | 1.0M | | | | | | $5 | $5 |
44
44
  | `nano-gpt/abliteration-ai/abliterated-model-large-v2` | 1.0M | | | | | | $5 | $5 |
45
+ | `nano-gpt/agnes-3.0-flash` | 524K | | | | | | $0.05 | $0.15 |
45
46
  | `nano-gpt/aion-labs/aion-2.0` | 131K | | | | | | $0.80 | $2 |
46
47
  | `nano-gpt/aion-labs/aion-3.0` | 131K | | | | | | $3 | $6 |
47
48
  | `nano-gpt/aion-labs/aion-3.0-mini` | 131K | | | | | | $0.70 | $1 |
@@ -82,11 +83,6 @@ for await (const chunk of stream) {
82
83
  | `nano-gpt/auto-model-basic` | 1.0M | | | | | | $10 | $20 |
83
84
  | `nano-gpt/auto-model-premium` | 1.0M | | | | | | $10 | $20 |
84
85
  | `nano-gpt/auto-model-standard` | 1.0M | | | | | | $10 | $20 |
85
- | `nano-gpt/azure-gpt-4-turbo` | 128K | | | | | | $10 | $30 |
86
- | `nano-gpt/azure-gpt-4o` | 128K | | | | | | $3 | $10 |
87
- | `nano-gpt/azure-gpt-4o-mini` | 128K | | | | | | $0.15 | $0.60 |
88
- | `nano-gpt/azure-o1` | 200K | | | | | | $15 | $60 |
89
- | `nano-gpt/azure-o3-mini` | 200K | | | | | | $1 | $4 |
90
86
  | `nano-gpt/baseten/Kimi-K2-Instruct-FP4` | 131K | | | | | | $0.40 | $2 |
91
87
  | `nano-gpt/brave` | 8K | | | | | | $5 | $5 |
92
88
  | `nano-gpt/brave-pro` | 8K | | | | | | $5 | $5 |
@@ -159,6 +155,7 @@ for await (const chunk of stream) {
159
155
  | `nano-gpt/deepseek/deepseek-v4-pro-0813:thinking` | 1.0M | | | | | | $1 | $3 |
160
156
  | `nano-gpt/deepseek/deepseek-v4-pro:thinking` | 1.0M | | | | | | $1 | $2 |
161
157
  | `nano-gpt/deepseek/deepseek-v4.1-flash` | 1.0M | | | | | | $0.16 | $0.31 |
158
+ | `nano-gpt/deepseek/deepseek-v4.1-flash:thinking` | 1.0M | | | | | | $0.16 | $0.31 |
162
159
  | `nano-gpt/Doctor-Shotgun/MS3.2-24B-Magnum-Diamond` | 33K | | | | | | $0.49 | $0.49 |
163
160
  | `nano-gpt/doubao-1.5-pro-256k` | 256K | | | | | | $0.80 | $1 |
164
161
  | `nano-gpt/doubao-1.5-pro-32k` | 32K | | | | | | $0.13 | $0.33 |
@@ -206,6 +203,8 @@ for await (const chunk of stream) {
206
203
  | `nano-gpt/gemini-3-pro-image-preview` | 66K | | | | | | $2 | $12 |
207
204
  | `nano-gpt/gemini-exp-1206` | 2.1M | | | | | | $1 | $5 |
208
205
  | `nano-gpt/gemma-4-12b-it` | 262K | | | | | | $0.05 | $0.25 |
206
+ | `nano-gpt/gemma-4-12b-it-semancer` | 131K | | | | | | $0.05 | $0.25 |
207
+ | `nano-gpt/gemma-4-12b-it-station-keeper` | 131K | | | | | | $0.05 | $0.25 |
209
208
  | `nano-gpt/gemma-4-26b-a4b-it-chimerax` | 262K | | | | | | $0.12 | $0.38 |
210
209
  | `nano-gpt/gemma-4-26b-a4b-it-darksoul` | 262K | | | | | | $0.12 | $0.38 |
211
210
  | `nano-gpt/gemma-4-26b-a4b-it-luminous` | 262K | | | | | | $0.12 | $0.38 |
@@ -461,6 +460,7 @@ for await (const chunk of stream) {
461
460
  | `nano-gpt/qwen/qwen3.8-27b-fable` | 262K | | | | | | $0.25 | $2 |
462
461
  | `nano-gpt/qwen/qwen3.8-27b-obliterated` | 262K | | | | | | $0.25 | $2 |
463
462
  | `nano-gpt/qwen/qwen3.8-27b-obliterated:thinking` | 262K | | | | | | $0.25 | $2 |
463
+ | `nano-gpt/qwen/qwen3.8-27b-queen` | 262K | | | | | | $0.25 | $2 |
464
464
  | `nano-gpt/qwen/qwen3.8-27b-uncensored` | 262K | | | | | | $0.25 | $2 |
465
465
  | `nano-gpt/qwen/qwen3.8-27b-uncensored:thinking` | 262K | | | | | | $0.25 | $2 |
466
466
  | `nano-gpt/qwen25-vl-72b-instruct` | 32K | | | | | | $0.70 | $0.70 |
@@ -44,7 +44,7 @@ for await (const chunk of stream) {
44
44
  | `opencode-go/glm-5.1` | 203K | | | | | | $1 | $4 |
45
45
  | `opencode-go/glm-5.2` | 1.0M | | | | | | $1 | $4 |
46
46
  | `opencode-go/glm-5.3` | 1.0M | | | | | | $1 | $4 |
47
- | `opencode-go/glm-5.3-flash` | 1.0M | | | | | | $0.07 | $0.25 |
47
+ | `opencode-go/glm-5.3-flash` | 1.0M | | | | | | $0.15 | $0.50 |
48
48
  | `opencode-go/gpt-5.6-luna` | 1.1M | | | | | | $0.20 | $1 |
49
49
  | `opencode-go/grok-4.6` | 500K | | | | | | $2 | $6 |
50
50
  | `opencode-go/hy3` | 256K | | | | | | $0.14 | $0.58 |
@@ -166,7 +166,7 @@ const result = await agent.generate('message for agent')
166
166
 
167
167
  **options.modelSettings.frequencyPenalty** (`number`): Penalty for token frequency (-2 to 2). Reduces repetition of frequent tokens.
168
168
 
169
- **options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured.
169
+ **options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured. Also accepts firstChunkMs, which only applies to streaming calls and is the maximum time the model may take to emit its first content-bearing chunk (text, reasoning, tool call, file or source; stream-start and metadata chunks do not count). A firstChunkMs timeout fails with timeoutType: 'firstChunk', behaves like stepMs for fallback, and is reset for each provider retry attempt. Nested timeout keys are merged across call-time and per-model settings.
170
170
 
171
171
  **options.modelSettings.stopSequences** (`string[]`): Stop sequences. If set, the model will stop generating text when one of the stop sequences is generated.
172
172
 
@@ -102,7 +102,7 @@ await agent.network(`
102
102
 
103
103
  **options.modelSettings.frequencyPenalty** (`number`): Penalty for token frequency (-2 to 2). Reduces repetition of frequent tokens.
104
104
 
105
- **options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured.
105
+ **options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured. Also accepts firstChunkMs, which only applies to streaming calls and is the maximum time the model may take to emit its first content-bearing chunk (text, reasoning, tool call, file or source; stream-start and metadata chunks do not count). A firstChunkMs timeout fails with timeoutType: 'firstChunk', behaves like stepMs for fallback, and is reset for each provider retry attempt. Nested timeout keys are merged across call-time and per-model settings.
106
106
 
107
107
  **options.modelSettings.stopSequences** (`string[]`): Stop sequences. If set, the model will stop generating text when one of the stop sequences is generated.
108
108
 
@@ -6,7 +6,7 @@
6
6
 
7
7
  `buildBasePrompt()` builds the shared base system prompt for a coding agent: the behavioral instructions that make the agent a good coding assistant. It takes a `PromptContext` describing the current environment (project, platform, git branch, mode, model) and returns the prompt as a string.
8
8
 
9
- Product-specific strings are parameterized through `productName`, `coAuthorName`, and `coAuthorEmail`, so you can rebrand the prompt and commit trailer without forking it. They default to `"Mastra Code"` / `"noreply@mastra.ai"`, so existing callers keep identical output.
9
+ Product-specific strings are parameterized through `productName`, `coAuthorName`, and `coAuthorEmail`, so you can rebrand the prompt and commit trailer without forking it. The co-author defaults to `"mastra-platform[bot]"` and `"284800079+mastra-platform[bot]@users.noreply.github.com"`.
10
10
 
11
11
  Use this with [`createCodingAgent()`](https://mastra.ai/reference/coding-agent/create-coding-agent) when you build runtime-defined instructions for the agent.
12
12
 
@@ -58,7 +58,7 @@ const prompt = buildBasePrompt({
58
58
 
59
59
  **mode** (`string`): Active agent mode (for example, "build" or "plan").
60
60
 
61
- **modelId** (`string`): Identifier of the active model, included in the commit Co-Authored-By line when provided.
61
+ **modelId** (`string`): Identifier of the active model.
62
62
 
63
63
  **activePlan** (`{ title: string; plan: string; approvedAt: string } | null`): The currently approved plan, if any.
64
64
 
@@ -66,9 +66,9 @@ const prompt = buildBasePrompt({
66
66
 
67
67
  **productName** (`string`): Display name used in the prompt header. (Default: `Mastra Code`)
68
68
 
69
- **coAuthorName** (`string`): Name used in the commit Co-Authored-By line. (Default: `Mastra Code`)
69
+ **coAuthorName** (`string`): Name used in the commit Co-Authored-By line. (Default: `mastra-platform[bot]`)
70
70
 
71
- **coAuthorEmail** (`string`): Email used in the commit Co-Authored-By line. (Default: `noreply@mastra.ai`)
71
+ **coAuthorEmail** (`string`): Email used in the commit Co-Authored-By line. (Default: `284800079+mastra-platform[bot]@users.noreply.github.com`)
72
72
 
73
73
  ## Returns
74
74
 
@@ -85,7 +85,7 @@ Final score: `(covered_elements / total_input_elements) * scale`
85
85
 
86
86
  A completeness score between 0 and 1:
87
87
 
88
- - **1.0**: Thoroughly addresses all aspects of the query with detailed detail.
88
+ - **1.0**: Addresses all aspects of the query in detail.
89
89
  - **0.7 to 0.9**: Covers most important aspects with good detail, minor gaps.
90
90
  - **0.4 to 0.6**: Addresses some key points but missing important aspects or lacking detail.
91
91
  - **0.1 to 0.3**: Only partially addresses the query with substantial gaps.
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Faithfulness scorer
6
6
 
7
- The `createFaithfulnessScorer()` function evaluates how factually accurate an LLM's output is compared to the provided context. It extracts claims from the output and verifies them against the context, making it essential to measure RAG pipeline responses' reliability.
7
+ The `createFaithfulnessScorer()` function evaluates how factually accurate an LLM's output is compared to the provided context. It extracts claims from the output and verifies them against the context. Use it to check whether a RAG response is supported by the retrieved information.
8
8
 
9
9
  ## Parameters
10
10
 
@@ -136,7 +136,7 @@ Each dimension receives an impact level with corresponding weights:
136
136
  - **Minimal (0.85)**: Slight phrasing changes but maintains correctness
137
137
  - **Moderate (0.6)**: Noticeable changes affecting quality but core info correct
138
138
  - **Significant (0.3)**: Major degradation in quality or accuracy
139
- - **Severe (0.1)**: Response substantially worse or completely derailed
139
+ - **Severe (0.1)**: Response much worse than the baseline or unrelated to the original query
140
140
 
141
141
  ### Conservative Scoring
142
142
 
@@ -180,10 +180,8 @@ Final Score = Weighted Score × scale
180
180
 
181
181
  **Both Mode (`'both'`)** - Use when (default, recommended):
182
182
 
183
- - Comprehensive evaluation of AI agent performance
184
- - Balancing user satisfaction with system compliance
183
+ - Evaluating whether responses satisfy user requests while following system instructions
185
184
  - Production monitoring where both user and system requirements matter
186
- - Holistic assessment of prompt-response alignment
187
185
 
188
186
  ## Common use cases
189
187
 
@@ -423,7 +421,7 @@ console.log(result)
423
421
 
424
422
  ### Excellent alignment output
425
423
 
426
- The output receives a high score because it perfectly addresses the intent and fulfills all requirements. It also uses the appropriate format.
424
+ In this example, the response receives a high score because it provides the requested Python factorial function with error handling for negative numbers.
427
425
 
428
426
  ```typescript
429
427
  {
@@ -51,7 +51,7 @@ OM performs thresholding with fast local token estimation. Text uses `tokenx`, a
51
51
 
52
52
  **retrieval** (`boolean | { vector?: boolean; scope?: 'thread' | 'resource'; instructions?: string }`): Let the agent look up the raw message history behind its observations. Observation groups keep durable pointers to the original messages, and a recall tool is registered so the agent can browse them. true enables cross-thread browsing by default. { vector: true } also enables semantic search using Memory's vector store and embedder. { scope: 'thread' } restricts the recall tool to the current thread only. Default scope is 'resource'. { instructions: '...' } appends application-specific recall guidance after Mastra's built-in retrieval instructions. (Default: `false`)
53
53
 
54
- **hooks** (`ObserveHooks`): Lifecycle hooks fired for every observation/reflection cycle — the manual observe()/reflect() APIs, turn-driven synchronous observation, and fire-and-forget async buffering. Callbacks receive threadId/resourceId/trigger call context ('manual' | 'turn-sync' | 'async-buffer'), and the end hooks (onObservationEnd/onReflectionEnd) additionally receive the OM model call's token usage and providerMetadata (where providers such as the AI Gateway report per-call cost), so apps can account for OM model spend without wrapping the observer/reflector models in middleware. Failed async-buffered cycles never throw; they report through the end hook's error field. Errors thrown by these hooks are caught and logged — they never fail the cycle.
54
+ **hooks** (`ObserveHooks`): Lifecycle hooks fired for every observation/reflection cycle — the manual observe()/reflect() APIs, turn-driven synchronous observation, and fire-and-forget async buffering. Callbacks receive threadId/resourceId/trigger call context ('manual' | 'turn-sync' | 'async-buffer'), and the end hooks (onObservationEnd/onReflectionEnd) additionally receive the OM model call's token usage and providerMetadata (where providers such as the AI Gateway report per-call cost), so apps can account for OM model spend without wrapping the observer/reflector models in middleware. Config-level lifecycle hooks are non-blocking by default: their errors are logged without failing the cycle. With hookExecution: 'await', config-level lifecycle errors can reject observe() or reflect(). Per-call observe({ hooks }) lifecycle hooks are also awaited and can reject observe(). A failing awaited start hook skips the model call; an end-hook failure can reject the call after its work has completed. Async-buffered cycles always use non-blocking config-level lifecycle hooks, regardless of hookExecution. Their cycle failures are reported through onObservationEnd.error or onReflectionEnd.error, rather than thrown to the caller. ObserveHooks also accepts transform hooks that intercept and can replace cycle data: beforeObservation({ messages, ...context }) runs on the messages about to be observed (return { messages } to filter/redact; an empty array skips the Observer call), afterObservation({ observations, ...context }) runs on the Observer's output before it is persisted, beforeReflection({ observations, ...context }) runs on the text sent to the Reflector, and afterReflection({ observations, ...context }) runs on the Reflector's output before it is persisted. Transform hooks return void to pass data through unchanged and are always awaited on every path. A thrown error fails the cycle before committing its transformed observation or reflection text, but doesn't roll back extractor callbacks or other side effects that already ran. The after hooks replace only text, not the separate structured extractor results stored in thread metadata. Reflection extraction and its callbacks run before afterReflection; this hook neither recomputes extracted values nor reruns callbacks. After hooks aren't a redaction boundary for all cycle data.
55
55
 
56
56
  **observation** (`ObservationalMemoryObservationConfig`): Configuration for the observation step. Controls when the Observer agent runs and how it behaves.
57
57
 
@@ -217,7 +217,7 @@ const run = await workflow.createRun();
217
217
 
218
218
  ### `setState()` is now async and the data passed is validated
219
219
 
220
- The `setState()` function is now async. The data passed is now validated against the `stateSchema` defined in the step. The state data validation also uses the `validateInputs` flag to determine whether to validate the state data or not. Also, when calling `setState()`, you can now pass only the state data being updated, instead of adding the previous state spread `(...state)`.
220
+ The `setState()` function is now async and validates data against the step's `stateSchema` when `validateInputs` is enabled. Pass only the fields you want to update, without spreading the previous state (`...state`).
221
221
 
222
222
  To migrate, update the `setState()` function to be async.
223
223
 
@@ -24,6 +24,7 @@ const result = await mastraClient.queryTraces({
24
24
  op: 'and',
25
25
  args: [
26
26
  { op: 'eq', left: { path: 'scorerId' }, right: { literal: 'factuality' } },
27
+ { op: 'eq', left: { path: 'scorerVersion' }, right: { literal: '2.1.0' } },
27
28
  { op: 'lt', left: { path: 'score' }, right: { literal: 0.6 } },
28
29
  ],
29
30
  },
@@ -34,7 +35,7 @@ const result = await mastraClient.queryTraces({
34
35
  })
35
36
  ```
36
37
 
37
- A `some` clause matches when one related record satisfies its complete nested predicate. In this example, `scorerId` and `score` must match on the same score record. A `none` clause matches when no related record satisfies its complete nested predicate.
38
+ A `some` clause matches when one related record satisfies its complete nested predicate. In this example, `scorerId`, `scorerVersion`, and `score` must match on the same score record. A `none` clause uses anti-existence semantics: it matches when no related record satisfies its complete nested predicate. A trace with no related scores therefore matches `scores.none`.
38
39
 
39
40
  Span clauses examine the current root span and current child spans. The root `timeRange` applies only to the selected current root's `startedAt`. Related spans and scores can participate even when their own timestamps are outside that range. Related records correlate only through a matching non-null `traceId`.
40
41
 
@@ -99,17 +100,176 @@ Hono and Fastify enforce the request-body limit before JSON parsing. Express and
99
100
 
100
101
  ## Fields and operators
101
102
 
102
- | Predicate context | Fields | Operators |
103
- | ----------------- | ---------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- |
104
- | Trace | `traceId`, `threadId`, `resourceId`, `entityName`, `entityType`, `environment`, `status` | `eq`, `ne`, `in`, `notIn`, `exists`, `notExists` |
105
- | Trace | `startedAt`, `endedAt` | `eq`, `ne`, `in`, `notIn`, `lt`, `lte`, `gt`, `gte`, `exists`, `notExists` |
106
- | Span | `spanType` | `eq`, `ne`, `in`, `notIn`, `exists`, `notExists` |
107
- | Span | `error` | `exists`, `notExists` |
108
- | Score | `scorerId` | `eq`, `ne`, `in`, `notIn`, `exists`, `notExists` |
109
- | Score | `score` | `eq`, `ne`, `in`, `notIn`, `lt`, `lte`, `gt`, `gte`, `exists`, `notExists` |
103
+ | Predicate context | Fields | Operators |
104
+ | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------- |
105
+ | Trace | `traceId`, `threadId`, `resourceId`, `entityName`, `entityType`, `environment`, `status` | `eq`, `ne`, `in`, `notIn`, `exists`, `notExists` |
106
+ | Trace metadata | `metadata.<key>` | `eq`, `ne`, `in`, `notIn`, `exists`, `notExists` |
107
+ | Trace | `startedAt`, `endedAt` | `eq`, `ne`, `in`, `notIn`, `lt`, `lte`, `gt`, `gte`, `exists`, `notExists` |
108
+ | Span | `name`, `spanType`, `model`, `provider`, `status`, `entityType`, `entityId`, `entityName`, `entityVersionId`, `parentEntityVersionId`, `rootEntityVersionId` | `eq`, `ne`, `in`, `notIn`, `exists`, `notExists` |
109
+ | Span | `startedAt`, `endedAt`, `durationMs` | `eq`, `ne`, `in`, `notIn`, `lt`, `lte`, `gt`, `gte`, `exists`, `notExists` |
110
+ | Span | `error` | `exists`, `notExists` |
111
+ | Score | `scorerId`, `scorerVersion`, `scoreSource`, `entityVersionId`, `parentEntityVersionId`, `rootEntityVersionId` | `eq`, `ne`, `in`, `notIn`, `exists`, `notExists` |
112
+ | Score | `score`, `timestamp` | `eq`, `ne`, `in`, `notIn`, `lt`, `lte`, `gt`, `gte`, `exists`, `notExists` |
113
+ | Score | `spanId` | `exists`, `notExists` |
110
114
 
111
115
  Compose predicates with `{ op: 'and', args: [...] }`, `{ op: 'or', args: [...] }`, and `{ op: 'not', arg: ... }`. Comparison predicates place a field reference on the left and a literal on the right. Membership predicates use a field reference in `value` and a homogeneous literal array in `set`.
112
116
 
117
+ String comparisons are case-sensitive, and literals are never coerced. Canonical string fields compare exact stored values. Metadata string values are trimmed before comparison, as described below. Numeric predicates, including `score` and `durationMs`, require numbers, while timestamp predicates require ISO timestamp strings. A missing value satisfies neither positive nor ordered predicates, although it does satisfy the negative operators `ne` and `notIn`. Combine a negative predicate with `exists` when the field must also be present.
118
+
119
+ ### Filter by span properties
120
+
121
+ Every condition inside one `spans.some` or `spans.none` clause applies to the same current span. For example, this predicate finds a failed tool span whose name is `medication_lookup`; a matching name on one span and an error on another don't satisfy it:
122
+
123
+ ```typescript
124
+ const failedMedicationLookup = {
125
+ spans: {
126
+ some: {
127
+ op: 'and',
128
+ args: [
129
+ { op: 'eq', left: { path: 'name' }, right: { literal: 'medication_lookup' } },
130
+ { op: 'eq', left: { path: 'spanType' }, right: { literal: 'tool_call' } },
131
+ { op: 'exists', path: 'error' },
132
+ ],
133
+ },
134
+ },
135
+ }
136
+ ```
137
+
138
+ Use `status` for the derived `success` or `error` outcome. Use `error` when you only need to test whether error details are present.
139
+
140
+ `durationMs` is the span's elapsed time in milliseconds. Span `startedAt` and `endedAt` conditions filter related spans independently of the top-level root `timeRange`:
141
+
142
+ ```typescript
143
+ const slowModelCalls = {
144
+ spans: {
145
+ some: {
146
+ op: 'and',
147
+ args: [
148
+ { op: 'eq', left: { path: 'spanType' }, right: { literal: 'model_generation' } },
149
+ { op: 'gt', left: { path: 'durationMs' }, right: { literal: 5000 } },
150
+ {
151
+ op: 'gte',
152
+ left: { path: 'startedAt' },
153
+ right: { literal: '2026-07-15T00:00:00.000Z' },
154
+ },
155
+ ],
156
+ },
157
+ },
158
+ }
159
+ ```
160
+
161
+ `model` and `provider` read the canonical `attributes.model` and `attributes.provider` span attributes. Only stored string values are queryable; missing, null, and non-string values count as missing.
162
+
163
+ ```typescript
164
+ const selectedModel = {
165
+ spans: {
166
+ some: {
167
+ op: 'and',
168
+ args: [
169
+ {
170
+ op: 'eq',
171
+ left: { path: 'model' },
172
+ right: { literal: 'claude-sonnet-4-6' },
173
+ },
174
+ { op: 'eq', left: { path: 'provider' }, right: { literal: 'anthropic' } },
175
+ ],
176
+ },
177
+ },
178
+ }
179
+ ```
180
+
181
+ Use generic entity fields for tool and retrieval identity. The version fields let one clause qualify the same span by its entity lineage:
182
+
183
+ ```typescript
184
+ const versionedRetrieval = {
185
+ spans: {
186
+ some: {
187
+ op: 'and',
188
+ args: [
189
+ { op: 'eq', left: { path: 'entityType' }, right: { literal: 'rag_ingestion' } },
190
+ { op: 'eq', left: { path: 'entityId' }, right: { literal: 'medication-index' } },
191
+ { op: 'eq', left: { path: 'entityVersionId' }, right: { literal: 'index-v3' } },
192
+ { op: 'eq', left: { path: 'rootEntityVersionId' }, right: { literal: 'agent-v2' } },
193
+ ],
194
+ },
195
+ },
196
+ }
197
+ ```
198
+
199
+ ### Filter by score time
200
+
201
+ Score timestamps are independent of the required top-level trace `timeRange`. The top-level range selects trace candidates by trace start time. A nested `timestamp` predicate limits related score records:
202
+
203
+ ```typescript
204
+ const query = {
205
+ timeRange: {
206
+ from: '2026-08-01T00:00:00.000Z',
207
+ to: '2026-08-08T00:00:00.000Z',
208
+ },
209
+ where: {
210
+ scores: {
211
+ some: {
212
+ op: 'and',
213
+ args: [
214
+ { op: 'eq', left: { path: 'scoreSource' }, right: { literal: 'automated' } },
215
+ {
216
+ op: 'gte',
217
+ left: { path: 'timestamp' },
218
+ right: { literal: '2026-07-15T00:00:00.000Z' },
219
+ },
220
+ {
221
+ op: 'lt',
222
+ left: { path: 'timestamp' },
223
+ right: { literal: '2026-08-01T00:00:00.000Z' },
224
+ },
225
+ ],
226
+ },
227
+ },
228
+ },
229
+ }
230
+ ```
231
+
232
+ ### Filter by span anchoring
233
+
234
+ Use `spanId` only to test whether a score targets a span. The predicate doesn't join the score to a matching span predicate:
235
+
236
+ ```typescript
237
+ // At least one score targets a span.
238
+ const anchoredToSpanWhere = { scores: { some: { op: 'exists', path: 'spanId' } } }
239
+
240
+ // At least one score applies to the trace rather than a span.
241
+ const traceLevelScoreWhere = { scores: { some: { op: 'notExists', path: 'spanId' } } }
242
+ ```
243
+
244
+ The deprecated `source` field and `scorerName` aren't available as score predicate fields. Use the canonical `scoreSource` and `scorerId` fields.
245
+
246
+ ### Filter by metadata
247
+
248
+ Use `metadata.<key>` paths for custom dimensions stored on the current root span. Metadata predicates compare trimmed, non-empty string values exactly and case-sensitively. Leading and trailing whitespace in a stored string value is ignored. Values empty after trimming count as missing. Null, numeric, boolean, object, and array values also count as missing.
249
+
250
+ ```typescript
251
+ const messageTrace = {
252
+ op: 'and',
253
+ args: [
254
+ { op: 'eq', left: { path: 'metadata.messageId' }, right: { literal: 'message-123' } },
255
+ { op: 'in', value: { path: 'metadata.actorRole' }, set: ['assistant', 'tool'] },
256
+ { op: 'exists', path: 'metadata.protocolVersion' },
257
+ {
258
+ op: 'or',
259
+ args: [
260
+ { op: 'eq', left: { path: 'metadata.temporalRunId' }, right: { literal: 'run-456' } },
261
+ { op: 'eq', left: { path: 'metadata.externalTraceId' }, right: { literal: 'trace-789' } },
262
+ ],
263
+ },
264
+ { op: 'not', arg: { op: 'exists', path: 'metadata.parentMessageId' } },
265
+ ],
266
+ }
267
+ ```
268
+
269
+ The key must name one top-level property. Empty keys and nested paths are rejected. Metadata keys aren't trimmed, so leading and trailing whitespace remains part of the exact key identity. When metadata duplicates a canonical trace field, such as `resourceId`, `threadId`, or `environment`, prefer the canonical field because it uses the dedicated storage column. The `metadata.<key>` form remains available when you specifically need the value from the metadata object. Metadata and canonical values may differ.
270
+
271
+ Metadata fields aren't available for grouping or field discovery.
272
+
113
273
  ## Responses
114
274
 
115
275
  An ungrouped query returns only lightweight completed traces:
@@ -4,7 +4,7 @@
4
4
 
5
5
  # RegexFilterProcessor
6
6
 
7
- The `RegexFilterProcessor` applies zero-cost regex pattern matching to filter, redact, or block content in agent messages. No LLM calls are made. All detection is regex-based.
7
+ The `RegexFilterProcessor` uses regex pattern matching to filter, redact, or block content in agent messages. No LLM calls are made.
8
8
 
9
9
  Supports built-in presets for common patterns (PII, secrets, URLs) and custom regex rules. Can be applied to input, output, or both phases.
10
10
 
@@ -650,7 +650,7 @@ Key metadata considerations:
650
650
 
651
651
  ## Deleting vectors
652
652
 
653
- When building RAG applications, you often need to clean up stale vectors when documents are deleted or updated. Mastra provides the `deleteVectors` method that supports deleting vectors by metadata filters, making it straightforward to remove all embeddings associated with a specific document.
653
+ Use `deleteVectors` with a metadata filter to remove embeddings associated with a document. This is useful for cleaning up stale vectors after a document is deleted or updated.
654
654
 
655
655
  ### Delete by Metadata Filter
656
656
 
@@ -160,7 +160,7 @@ const stream = await agent.stream('message for agent')
160
160
 
161
161
  **options.modelSettings.frequencyPenalty** (`number`): Penalty for token frequency (-2 to 2). Reduces repetition of frequent tokens.
162
162
 
163
- **options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured.
163
+ **options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured. Also accepts firstChunkMs, which only applies to streaming calls and is the maximum time the model may take to emit its first content-bearing chunk (text, reasoning, tool call, file or source; stream-start and metadata chunks do not count). A firstChunkMs timeout fails with timeoutType: 'firstChunk', behaves like stepMs for fallback, and is reset for each provider retry attempt. Nested timeout keys are merged across call-time and per-model settings.
164
164
 
165
165
  **options.modelSettings.stopSequences** (`string[]`): Stop sequences. If set, the model will stop generating text when one of the stop sequences is generated.
166
166
 
@@ -268,7 +268,7 @@ const fullText = await stream.text
268
268
 
269
269
  ### Limiting execution time
270
270
 
271
- Use `modelSettings.timeout` to bound how long a run may take. `totalMs` limits the entire run, including every loop iteration, tool call and retry. `stepMs` limits a single model call, covering both establishing the stream and consuming it.
271
+ Use `modelSettings.timeout` to bound how long a run may take. `totalMs` limits the entire run, including every loop iteration, tool call and retry. `stepMs` limits a single model call, covering both establishing the stream and consuming it. `firstChunkMs` only applies to streaming calls and limits how long the model may take to emit its first content-bearing chunk.
272
272
 
273
273
  ```ts
274
274
  const stream = await agent.stream('Tell me a story', {
@@ -276,15 +276,19 @@ const stream = await agent.stream('Tell me a story', {
276
276
  timeout: {
277
277
  totalMs: 30000, // fail the run if it takes longer than 30s
278
278
  stepMs: 10000, // fail an individual model call after 10s
279
+ firstChunkMs: 3000, // fail a model call that produces no content within 3s
279
280
  },
280
281
  },
281
282
  })
282
283
  ```
283
284
 
284
- Exceeding either budget fails with a `MastraTimeoutError`, which carries a `timeoutType` of `'total'` or `'step'`. Each budget behaves differently when the agent is configured with fallback [`models`](https://mastra.ai/reference/agents/agent):
285
+ Exceeding a budget fails with a `MastraTimeoutError`, which carries a `timeoutType` of `'total'`, `'step'` or `'firstChunk'`. Each budget behaves differently when the agent is configured with fallback [`models`](https://mastra.ai/reference/agents/agent):
285
286
 
286
287
  - A `totalMs` timeout ends the run immediately and doesn't try the next model, because it's a hard deadline for the run as a whole.
287
288
  - A `stepMs` timeout isn't retried against the same model, but does advance to the next model, which makes it a way to fail over from a slow provider.
289
+ - A `firstChunkMs` timeout behaves like `stepMs`: it isn't retried against the same model, but does advance to the next model. Stream-start and metadata chunks don't satisfy the budget; only the first text, reasoning, tool call, file or source chunk does. Once content begins, only `stepMs` and `totalMs` remain active. Each provider retry attempt gets a fresh `firstChunkMs` budget.
290
+
291
+ Nested `timeout` settings are merged per key across call-time and per-model `modelSettings`, so a per-model `stepMs` override keeps a call-time `firstChunkMs`.
288
292
 
289
293
  ### AI SDK v5+ Format
290
294
 
@@ -646,7 +646,7 @@ async executeTool(
646
646
 
647
647
  ### What are MCP Resources?
648
648
 
649
- Resources are a core primitive in the Model Context Protocol (MCP) that allow servers to expose data and content that can be read by clients and used as context for LLM interactions. They represent any kind of data that an MCP server wants to make available, such as:
649
+ MCP resources expose server data that clients can read and use as context for LLM interactions. Examples include:
650
650
 
651
651
  - File contents
652
652
  - Database records
@@ -35,7 +35,7 @@ A sandbox constructed with a known `id` resolves that id on `start()`: reconnect
35
35
 
36
36
  Providers that predate the contract return `void`, which the base class treats as unknown.
37
37
 
38
- Provider implementations plug into the start lifecycle at one of three rungs; the best available wins, and the base class always owns coalescing, status management, the `onStart` hook, and mount processing:
38
+ Providers implement the start lifecycle using one of the following approaches, listed in order of preference. The base class handles concurrent start calls and status management, as well as the `onStart` hook and mount processing:
39
39
 
40
40
  1. **Acquisition primitives**: implement protected `find()` (side-effect-free lookup by logical id, returning a provider-native handle or `undefined`), `connect(handle)` (wake/resume/adopt), and `create()` (provision fresh) without overriding `start()`. The base orchestrates find → connect → `{ outcome: 'connected' }`, else create → `{ outcome: 'created' }`. The outcome is derived structurally from which branch ran. Used by `E2BSandbox`, `DaytonaSandbox`, and `LocalSandbox`.
41
41
  2. **`start()` override returning `SandboxStartResult`**: for providers whose API is a fused get-or-create where decomposition would add round-trips (`PlatformSandbox`, `RailwaySandbox`).
@@ -47,7 +47,7 @@ The result is also forwarded to the `onStart` lifecycle hook as `{ sandbox, outc
47
47
 
48
48
  ### `onStart` (constructor option)
49
49
 
50
- `onStart` runs inside the start lifecycle, after the sandbox reaches `running` status and before pending mounts are processed. It fires on every start regardless of trigger, whether an explicit call, a lazy `ensureRunning()` from a command, or a revival after the provider replaced the VM. That makes it the seam for once-per-VM setup: branch on `outcome` and probe or run whatever the environment needs.
50
+ `onStart` runs inside the start lifecycle, after the sandbox reaches `running` status and before pending mounts are processed. It fires on every start regardless of trigger, whether an explicit call, a lazy `ensureRunning()` from a command, or a revival after the provider replaced the VM. Use this hook for once-per-VM setup: check `outcome` to decide whether to run setup or check that it already completed.
51
51
 
52
52
  ```typescript
53
53
  new E2BSandbox({
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mastra/mcp-docs-server",
3
- "version": "1.2.24",
3
+ "version": "1.2.25-alpha.3",
4
4
  "description": "MCP server for accessing Mastra.ai documentation, changelogs, and news.",
5
5
  "type": "module",
6
6
  "main": "dist/index.js",
@@ -27,7 +27,7 @@
27
27
  "jsdom": "^26.1.0",
28
28
  "local-pkg": "^1.1.2",
29
29
  "zod": "^4.4.3",
30
- "@mastra/core": "1.65.0",
30
+ "@mastra/core": "1.66.0-alpha.1",
31
31
  "@mastra/mcp": "^1.17.3"
32
32
  },
33
33
  "devDependencies": {
@@ -45,8 +45,8 @@
45
45
  "typescript": "^7.0.2",
46
46
  "vitest": "4.1.10",
47
47
  "@internal/lint": "0.0.131",
48
- "@mastra/core": "1.65.0",
49
- "@internal/types-builder": "0.0.106"
48
+ "@internal/types-builder": "0.0.106",
49
+ "@mastra/core": "1.66.0-alpha.1"
50
50
  },
51
51
  "homepage": "https://mastra.ai",
52
52
  "repository": {