@mastra/mcp-docs-server 1.3.1 → 1.3.2-alpha.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. package/.docs/docs/agents/processors.md +1 -1
  2. package/.docs/docs/evals/evals-with-memory.md +3 -1
  3. package/.docs/docs/evals/experiments.md +30 -10
  4. package/.docs/docs/harness/durable-agents.md +23 -5
  5. package/.docs/docs/memory/observational-memory.md +59 -12
  6. package/.docs/docs/memory/overview.md +2 -2
  7. package/.docs/integrations/deploy/kubernetes.md +6 -4
  8. package/.docs/integrations/sandboxes/modal.md +40 -0
  9. package/.docs/models/gateways/netlify.md +7 -2
  10. package/.docs/models/gateways/openrouter.md +4 -5
  11. package/.docs/models/index.md +1 -1
  12. package/.docs/models/providers/above.md +2 -2
  13. package/.docs/models/providers/cerebras.md +1 -1
  14. package/.docs/models/providers/cortecs.md +3 -3
  15. package/.docs/models/providers/deepseek.md +9 -3
  16. package/.docs/models/providers/edenai.md +285 -288
  17. package/.docs/models/providers/fireworks-ai.md +4 -7
  18. package/.docs/models/providers/kilo.md +19 -20
  19. package/.docs/models/providers/melious.md +1 -4
  20. package/.docs/models/providers/nano-gpt.md +10 -4
  21. package/.docs/models/providers/opencode-go.md +2 -1
  22. package/.docs/models/providers/opencode.md +3 -2
  23. package/.docs/models/providers/requesty.md +2 -2
  24. package/.docs/models/providers/zenmux.md +8 -1
  25. package/.docs/reference/agents/agent.md +15 -0
  26. package/.docs/reference/agents/durable-agent.md +3 -1
  27. package/.docs/reference/channels/channel-provider.md +20 -1
  28. package/.docs/reference/client-js/agents.md +2 -0
  29. package/.docs/reference/coding-agent/create-coding-agent.md +22 -14
  30. package/.docs/reference/index.md +1 -0
  31. package/.docs/reference/memory/cloneThread.md +2 -2
  32. package/.docs/reference/memory/observational-memory.md +12 -6
  33. package/.docs/reference/processors/agents-md-injector.md +2 -0
  34. package/.docs/reference/processors/cyber-refusal-handler.md +76 -0
  35. package/.docs/reference/pubsub/base.md +11 -0
  36. package/.docs/reference/pubsub/redis-streams.md +8 -0
  37. package/.docs/reference/workspace/local-sandbox.md +2 -0
  38. package/.docs/reference/workspace/workspace-class.md +14 -1
  39. package/package.json +5 -5
@@ -879,7 +879,7 @@ const agent = new Agent({
879
879
  The retry mechanism:
880
880
 
881
881
  - Works in `processOutputStep()` and `processInputStep()` methods
882
- - Replays the step with the abort reason added as context for the LLM
882
+ - Replays the step with the abort reason appended verbatim to the end of the conversation as a system reminder. The retry keeps the previous request as its prefix, which preserves provider prompt caching
883
883
  - Tracks retry count via the `retryCount` parameter
884
884
  - Requires an explicit `maxProcessorRetries` limit on the agent or call
885
885
 
@@ -90,7 +90,9 @@ const average = scores.reduce((a, b) => a + b, 0) / scores.length
90
90
 
91
91
  ## Dataset experiments with an inline task
92
92
 
93
- `dataset.startExperiment({ target: agent })` **doesn't** forward a `memory` option to the agent, only `requestContext`. To run a stored dataset against a memory-enabled agent, use an inline `task` function and stash `{ threadId, resourceId }` in each item's `metadata`. The scorer pipeline still runs as normal.
93
+ `dataset.startExperiment()` accepts `targetType` and `targetId` for a registered target, or an inline `task` function. For a registered memory-enabled agent, the runner creates a fresh thread per item by default. See [memory-enabled experiment targets](https://mastra.ai/docs/evals/experiments) for resource isolation and request context options.
94
+
95
+ To choose each item's thread explicitly through `agent.generate()`, use an inline `task` and store `{ threadId, resourceId }` in each item's `metadata`. The scorer pipeline still runs as normal.
94
96
 
95
97
  ```typescript
96
98
  import { randomUUID } from 'node:crypto'
@@ -53,7 +53,9 @@ You can also delete experiments with the [Core API](https://mastra.ai/reference/
53
53
 
54
54
  ## Experiment targets
55
55
 
56
- You can point an experiment at a registered agent, workflow, or scorer.
56
+ Use `targetType` and `targetId` to select an agent, workflow, or scorer registered on your Mastra instance. The target processes each dataset item. The experiment's `scorers` evaluate the target's output afterward.
57
+
58
+ For custom execution, pass an inline `task` function instead. See [dataset experiments with an inline task](https://mastra.ai/docs/evals/evals-with-memory) for an example that controls each item's memory thread.
57
59
 
58
60
  ### Registered agent
59
61
 
@@ -68,25 +70,27 @@ const summary = await dataset.startExperiment({
68
70
  })
69
71
  ```
70
72
 
71
- Each item's `input` is passed directly to `agent.generate()`, so it must be a `string`, `string[]`, or `CoreMessage[]`.
73
+ Each item's `input` is passed directly to the agent. Use a supported message format, such as a prompt string, a message object, or an array of messages. See the [`agent.generate()` reference](https://mastra.ai/reference/agents/generate) for accepted inputs.
72
74
 
73
75
  #### Memory-enabled agents
74
76
 
75
- When the target agent has its own memory and the request context carries a resource id (`MASTRA_RESOURCE_ID_KEY`, set by auth middleware, the experiment or item `requestContext`, or the Studio **Run Experiment** form), the experiment runner injects a fresh memory thread for each item. A resource id in the request context means "run as this resource": each item's conversation persists as a thread under that resource, and retried items get a new thread per attempt so earlier failed attempts can't leak into the retry's context.
77
+ When the target agent has its own memory, the experiment runner injects a fresh memory thread for each item and retry. Without a resource id in the request context, each thread gets an isolated, experiment-owned resource.
78
+
79
+ To run against an existing resource, set `MASTRA_RESOURCE_ID_KEY` in the experiment's `requestContext` or in an item's `requestContext`. Item values override experiment values. The runner still creates a fresh thread for each item and retry, but all threads using the same resource id share that resource's memory.
76
80
 
77
81
  Injected threads are tagged so you can map them back to the run: thread metadata carries the `experimentId` and the dataset item's id as `experimentItemId`. No thread title is generated for them.
78
82
 
79
- Because the threads belong to the caller's resource, resource-scoped memory features both read and write that resource's state during the run:
83
+ When items share a resource, resource-scoped memory features can read and write the same state:
80
84
 
81
- - Resource-scoped working memory updates persist to the resource, and later items in the run see updates made by earlier items.
85
+ - Resource-scoped working memory updates persist to the resource and can affect other items. Items run concurrently by default, so don't rely on an execution order.
82
86
  - Resource-scoped semantic recall can surface the resource's prior conversations to the experiment, and experiment transcripts become recallable in that resource's later conversations.
83
87
 
84
88
  Use this approach to evaluate an agent against a real user's accumulated context. To keep experiment runs from touching real user state, use a dedicated evaluation resource id instead.
85
89
 
86
90
  Thread injection is skipped in the following cases:
87
91
 
88
- - If the request context also sets `MASTRA_THREAD_ID_KEY`, the runner uses that thread as-is, so every item (and retry) shares the same conversation.
89
- - If the agent has no memory, or the request context has no resource id, the run is memoryless and nothing is persisted.
92
+ - If the request context sets `MASTRA_THREAD_ID_KEY`, the runner uses that thread as-is. Items and retries using the same thread id share a conversation.
93
+ - If the agent has no memory, or the request context explicitly sets the resource id to an empty string or `null`, the runner skips thread injection.
90
94
 
91
95
  ### Registered workflow
92
96
 
@@ -101,13 +105,24 @@ const summary = await dataset.startExperiment({
101
105
  })
102
106
  ```
103
107
 
104
- The workflow receives each item's `input` as its trigger data.
108
+ The workflow receives each item's `input` as `inputData` in `run.start()`. Structure the input to match the workflow's input schema.
105
109
 
106
110
  ### Registered scorer
107
111
 
108
- Point to a scorer to evaluate an LLM judge against ground truth:
112
+ Use a scorer as the target to evaluate the judge itself. The runner calls `scorer.run(item.input)`, so the dataset item's `input` must contain the full payload the scorer expects.
113
+
114
+ This example assumes a registered `accuracy` scorer that accepts string `input`, `output`, and `groundTruth` fields. Adapt those fields to your scorer's expected shape:
109
115
 
110
116
  ```typescript
117
+ await dataset.addItem({
118
+ input: {
119
+ input: 'What is the capital of France?',
120
+ output: 'Paris',
121
+ groundTruth: 'Paris',
122
+ },
123
+ groundTruth: { score: 1 },
124
+ })
125
+
111
126
  const summary = await dataset.startExperiment({
112
127
  name: 'judge-accuracy-eval',
113
128
  targetType: 'scorer',
@@ -115,7 +130,12 @@ const summary = await dataset.startExperiment({
115
130
  })
116
131
  ```
117
132
 
118
- The scorer receives each item's `input` and `groundTruth`. LLM-based judges can drift over time as underlying models change, so it's important to periodically realign them against known-good labels. A dataset gives you a stable benchmark to detect that drift.
133
+ The two `groundTruth` fields serve different purposes:
134
+
135
+ - `input.groundTruth` is the reference answer passed to the judge: `'Paris'`.
136
+ - The top-level `groundTruth` is the expected result from the judge: `{ score: 1 }`. It isn't passed to the target scorer.
137
+
138
+ The target's output contains its `score` and `reason`. To evaluate that output against the top-level ground truth, add another scorer through the experiment's `scorers` option. Without an additional scorer, the experiment records the judge's output but doesn't score its agreement with the expected result.
119
139
 
120
140
  ## Scoring results
121
141
 
@@ -163,15 +163,29 @@ Visit the [`createInngestAgent()` reference](https://mastra.ai/reference/agents/
163
163
  Durable agents support resumable streams through PubSub and an event cache. When a client disconnects mid-stream, the cache continues storing events. The same client can reconnect by calling `observe()` with the `runId`:
164
164
 
165
165
  ```typescript
166
- const { output, cleanup } = await durableResearcher.observe(runId)
166
+ const { output } = await durableResearcher.observe(runId)
167
167
 
168
168
  for await (const chunk of output.fullStream) {
169
169
  // Chunks from the run, including any missed while disconnected
170
170
  }
171
+ ```
171
172
 
172
- cleanup()
173
+ Observing doesn't own the run. When the loop ends, or you leave it early with `break`, `return`, or an error, the observer unsubscribes and the run keeps going for anyone else watching it. To stop observing from outside the loop, for example when the client disconnects, call `detach()`:
174
+
175
+ ```typescript
176
+ const { output, detach } = await durableResearcher.observe(runId)
177
+
178
+ request.signal.addEventListener('abort', detach, { once: true })
179
+
180
+ if (request.signal.aborted) detach()
181
+
182
+ for await (const chunk of output.fullStream) {
183
+ // Stops when the run finishes or the client disconnects
184
+ }
173
185
  ```
174
186
 
187
+ If you pass `idleTimeoutMs` and no chunks arrive for that long, Mastra checks `isAlive` when you provide it. If `isAlive` returns `true` or throws, the timer restarts and the stream stays open. Otherwise the stream ends with an error chunk and detaches the observer. The run itself is only cleaned up when `isAlive` returns `false`.
188
+
175
189
  `createDurableAgent()` and `createEventedAgent()` use an in-memory cache by default, which means resumable streams work within a single process. For production, provide a persistent cache backend (e.g., Redis) so cached events survive process restarts:
176
190
 
177
191
  ```typescript
@@ -231,7 +245,11 @@ Visit [Background tasks](https://mastra.ai/docs/harness/background-tasks) for th
231
245
 
232
246
  ## Cleanup
233
247
 
234
- Every `stream()` and `observe()` call returns a `cleanup` function. Calling it unsubscribes from PubSub and removes the run from the internal registry. If you forget to call it, an automatic timer fires after the stream ends, but calling `cleanup()` yourself frees resources immediately.
248
+ Every `stream()` and `observe()` call returns a `cleanup` function that tears the run down: it unsubscribes from PubSub, removes the run from the internal registry, and deletes the run's cached events so nobody can replay it afterward.
249
+
250
+ When you don't call `cleanup()` yourself, Mastra calls it `cleanupTimeoutMs` after the run finishes, errors, or is aborted. Detaching an observer or reaching a plain `idleTimeoutMs` doesn't start this timer. Set `cleanupTimeoutMs: 0` to turn the timer off.
251
+
252
+ Call `cleanup()` from the process that started the run. To stop watching a run from `observe()`, use `detach()` or leave the loop instead.
235
253
 
236
254
  ## Tool approval
237
255
 
@@ -260,6 +278,8 @@ Under the default persistence policy, these `running` checkpoints are only writt
260
278
 
261
279
  Durable agent runs are excluded from the generic boot-time restart of active workflow runs. The only automatic recovery path for durable agent runs is `recovery.durableAgents: 'auto'`, which holds a recovery lease and registers thread runtimes before re-driving each run.
262
280
 
281
+ > **Warning:** Automatic and manual recovery re-run the agentic loop from the last persisted snapshot, which can reissue LLM calls (with real cost) and re-execute tool calls. A sub-agent delegation is one generated `agent-<name>` tool call, and recovery doesn't checkpoint the LLM or tool calls inside that delegation independently. If recovery replays the delegation step, the entire sub-agent run starts again and can repeat inner LLM calls and already-completed tool side effects. Make tools, including tools used by sub-agents, idempotent, or deduplicate their external side effects.
282
+
263
283
  ### Snapshot persistence
264
284
 
265
285
  The `shouldPersistSnapshot` option on `createDurableAgent()` (also accepted in the agent-level `durable` config) controls which workflow snapshots a durable agent writes. The default policy:
@@ -295,8 +315,6 @@ export const mastra = new Mastra({
295
315
 
296
316
  On startup, this discovers every registered durable agent with runs stuck in `running` status and re-drives them from the last persisted snapshot.
297
317
 
298
- > **Warning:** Recovery re-runs the agentic loop from the last snapshot, which re-issues LLM calls (real cost) and re-executes tool calls. Make sure your tools are idempotent before enabling automatic recovery.
299
-
300
318
  ### Manual recovery
301
319
 
302
320
  If you need finer control, such as gating recovery behind a leader election or running it on a schedule, call the methods directly:
@@ -191,6 +191,40 @@ const memory = new Memory({
191
191
 
192
192
  `bufferOnIdle` is off by default. It's separate from `bufferTokens`: `bufferTokens` controls step-time async buffering, while `bufferOnIdle` controls end-of-turn buffering for idle turns.
193
193
 
194
+ ### Retries and failure policy
195
+
196
+ Observer and Reflector model calls are retried on transient provider errors, and a turn aborts if they still fail. Each stage is controlled independently by:
197
+
198
+ - `maxRetries`: retries after the initial model call (default `8`).
199
+ - `failurePolicy`: what happens once retries are exhausted, either `'abort'` (default) or `'continue'`.
200
+
201
+ ```typescript
202
+ const memory = new Memory({
203
+ options: {
204
+ observationalMemory: {
205
+ observation: {
206
+ maxRetries: 8,
207
+ failurePolicy: 'continue',
208
+ },
209
+ reflection: {
210
+ maxRetries: 8,
211
+ failurePolicy: 'continue',
212
+ },
213
+ },
214
+ },
215
+ })
216
+ ```
217
+
218
+ `maxRetries` governs OM's own retry ladder; the underlying model call is configured with no provider-level retries, so the option is the single retry knob for the stage.
219
+
220
+ With `failurePolicy: 'continue'`, Mastra retries as configured, reports the failure through the existing OM diagnostics, keeps the failed input pending for a later cycle, and lets the main agent turn continue. It doesn't advance observation boundaries or discard unobserved messages. A Reflector failure under `'continue'` leaves any already-persisted observations committed and defers reflection to the next threshold crossing.
221
+
222
+ A provider outage usually takes out both stages, so set the policy on both if you want the turn to survive one.
223
+
224
+ This policy applies only to Observer and Reflector model and provider failures in synchronous, resource-scoped, and buffered observation. Persistence, indexing, transform, locking, invariant, and explicit abort failures remain fatal. It doesn't prevent the underlying provider or network error, change `blockAfter`, or change attachment handling.
225
+
226
+ > **Warning:** `'continue'` has no backstop for a sustained outage. Unobserved messages stay pending and keep accruing in the main agent's context, so a long outage moves the failure from the memory layer to the model's context limit.
227
+
194
228
  See [the API reference](https://mastra.ai/reference/memory/observational-memory) for the full configuration shape.
195
229
 
196
230
  ## Benefits
@@ -615,29 +649,42 @@ const memory = new Memory({
615
649
 
616
650
  Thread scope requires a valid `threadId` to be provided when calling the agent. If `threadId` is missing, Observational Memory throws an error. This prevents multiple threads from silently sharing a single observation record, which can cause database deadlocks.
617
651
 
618
- ### Resource scope (experimental)
652
+ ### Resource scope (deprecated)
653
+
654
+ > **Deprecated:** `scope: 'resource'` is deprecated and will be removed in a future release. Mastra logs a warning the first time Observational Memory is created with resource scope. Remove the `scope` option to use the default thread scope, and use the [alternatives below](#cross-thread-continuity-without-resource-scope) for cross-conversation continuity.
655
+
656
+ In resource scope, observations are shared across all threads for a resource (typically a user). Resource scope works much worse than thread scope for prompt caching and for the agent's understanding of the conversation:
657
+
658
+ - **Prompt caching:** Observations from every thread share one record. When any thread is observed, the observation context changes for all other threads, which invalidates their cached prompt prefix.
659
+ - **Agent understanding:** Each thread sees a mix of observations from every thread for the resource. Every new thread starts with context from earlier threads, and the agent may continue work that another thread started but didn't finish.
660
+ - **Performance:** Unobserved messages across _all_ threads are processed together, which is slow for users with many existing threads.
661
+
662
+ A new knowledge and subconscious memory primitive for cross-thread memory is coming soon and will replace resource scope. Until then, use the alternatives below.
663
+
664
+ #### Cross-thread continuity without resource scope
665
+
666
+ Switching from resource scope to thread scope doesn't migrate resource-scoped observations. Each thread uses its existing thread-scoped record if it has one, or starts a new, empty one, and messages already observed under resource scope aren't observed again. Keep facts that must carry over in resource-scoped working memory.
619
667
 
620
- Observations are shared across all threads for a resource (typically a user). Enables cross-conversation memory.
668
+ Use thread scope with one or both of these features instead:
669
+
670
+ - [Retrieval mode](#retrieval-mode): The agent gets a `recall` tool that can list and browse other threads for the same resource, and optionally search them semantically. The agent looks up earlier conversations only when it needs them.
671
+ - [Resource-scoped working memory](https://mastra.ai/docs/memory/working-memory): Stores small, durable facts about the user that every thread can read.
621
672
 
622
673
  ```typescript
623
674
  const memory = new Memory({
624
675
  options: {
625
676
  observationalMemory: {
626
677
  model: 'google/gemini-2.5-flash',
678
+ retrieval: true,
679
+ },
680
+ workingMemory: {
681
+ enabled: true,
627
682
  scope: 'resource',
628
683
  },
629
684
  },
630
685
  })
631
686
  ```
632
687
 
633
- Resource scope works, however it's marked as experimental for now until we prove task adherence/continuity across multiple ongoing simultaneous threads. As of today, you may need to tweak your system prompt to prevent one thread from continuing the work that another had already started (but hadn't finished).
634
-
635
- This is because in resource scope, each thread is a perspective on _all_ threads for the resource.
636
-
637
- For your use-case this may not be a problem, so your mileage may vary.
638
-
639
- > **Warning:** In resource scope, unobserved messages across _all_ threads are processed together. For users with many existing threads, this can be slow. Use thread scope for existing apps.
640
-
641
688
  ## Token budgets
642
689
 
643
690
  OM uses token thresholds to decide when to observe and reflect. See [token budget configuration](https://mastra.ai/reference/memory/observational-memory) for details.
@@ -786,7 +833,7 @@ const memory = new Memory({
786
833
 
787
834
  Setting `bufferTokens: false` disables both observation and reflection async buffering. See [async buffering configuration](https://mastra.ai/reference/memory/observational-memory) for the full API.
788
835
 
789
- > **Note:** Resource scope automatically disables async buffering.
836
+ > **Note:** Resource scope (deprecated) automatically disables async buffering.
790
837
 
791
838
  ## Observer Context Optimization
792
839
 
@@ -889,7 +936,7 @@ hooks: {
889
936
  No manual migration needed. OM reads existing messages and observes them lazily when thresholds are exceeded.
890
937
 
891
938
  - **Thread scope**: The first time a thread exceeds `observation.messageTokens`, the Observer processes the backlog.
892
- - **Resource scope**: All unobserved messages across all threads for a resource are processed together. For users with many existing threads, this could take substantial time.
939
+ - **Resource scope (deprecated)**: All unobserved messages across all threads for a resource are processed together. For users with many existing threads, this could take substantial time.
893
940
 
894
941
  ## Comparing OM with other memory features
895
942
 
@@ -212,7 +212,7 @@ To go beyond this default isolation, you can share memory between agents by pass
212
212
 
213
213
  When you call agents directly (outside the delegation flow), memory sharing is controlled by two identifiers: `resourceId` and `threadId`. Agents that use the same values read and write to the same data. This is useful when agents collaborate on a shared context, for example, a researcher that saves notes and a writer that reads them.
214
214
 
215
- **Resource-scoped sharing** is the most common pattern. [Working memory](https://mastra.ai/docs/memory/working-memory) and [semantic recall](https://mastra.ai/docs/memory/semantic-recall) default to `scope: 'resource'`. If two agents share a `resourceId`, they share observations, working memory, and embeddings, even across different threads:
215
+ **Resource-scoped sharing** is the most common pattern. [Working memory](https://mastra.ai/docs/memory/working-memory) and [semantic recall](https://mastra.ai/docs/memory/semantic-recall) default to `scope: 'resource'`. If two agents share a `resourceId`, they share working memory and embeddings, even across different threads:
216
216
 
217
217
  ```typescript
218
218
  // Both agents share the same resource-scoped memory
@@ -225,7 +225,7 @@ await writer.generate('Write a summary from the research notes.', {
225
225
  })
226
226
  ```
227
227
 
228
- Because both calls use `resource: 'project-42'`, the writer can access the researcher's observations and working memory. Semantic embeddings are also shared through the resource. Each agent still has its own thread, so message histories stay separate.
228
+ Because both calls use `resource: 'project-42'`, the writer can access the researcher's working memory. Semantic embeddings are also shared through the resource. Each agent still has its own thread, so message histories stay separate.
229
229
 
230
230
  **Thread-scoped sharing** gives tighter coupling. [Observational Memory](https://mastra.ai/docs/memory/observational-memory) uses `scope: 'thread'` by default. If two agents use the same `resource` and `thread`, they share the full message history. Each agent sees every message the other has written. This is useful when agents need to build on each other's exact outputs.
231
231
 
@@ -251,18 +251,20 @@ Register the durable agent with the `Mastra` instance above. Its run state now l
251
251
  A durable agent publishes stream chunks to a per-run topic through the shared pub/sub. A client that disconnects reconnects by calling `observe()` with the run's ID, and replays the chunks it missed from the shared cache:
252
252
 
253
253
  ```typescript
254
- const { output, cleanup } = await durableAssistant.observe(runId)
254
+ const { output, detach } = await durableAssistant.observe(runId)
255
+
256
+ request.signal.addEventListener('abort', detach, { once: true })
257
+
258
+ if (request.signal.aborted) detach()
255
259
 
256
260
  for await (const chunk of output.fullStream) {
257
261
  // Chunks from the run, including any missed while disconnected
258
262
  }
259
-
260
- cleanup()
261
263
  ```
262
264
 
263
265
  Because the run state is in Postgres and the events are in Redis, the reconnecting request can be served by any pod, not only the one that started the run. See [Resumable streams](https://mastra.ai/docs/harness/durable-agents).
264
266
 
265
- Multiple clients can observe the same run at once. Each `observe()` call receives the full stream, so a user watching from two devices, or two people following the same run, stay in sync.
267
+ Multiple clients can observe the same run at once. Each `observe()` call receives the full stream, so a user watching from two devices, or two people following the same run, stay in sync. When one client leaves the loop or calls `detach()`, only that client stops watching. Don't call `cleanup()` from an observer, because it clears the cached events the other clients need to reconnect.
266
268
 
267
269
  ## Tool approval across pods
268
270
 
@@ -100,6 +100,8 @@ Get your credentials from the [Modal dashboard](https://modal.com/settings/token
100
100
 
101
101
  **onStart** (`function`): Lifecycle hook called after the sandbox reaches running status.
102
102
 
103
+ **volumes** (`Record<string, Volume>`): Modal Volumes to mount into the sandbox, keyed by absolute mount path. The working directory must not be at or inside a mount path.
104
+
103
105
  **onStop** (`function`): Lifecycle hook called before the sandbox stops.
104
106
 
105
107
  **onDestroy** (`function`): Lifecycle hook called before the sandbox is destroyed.
@@ -137,6 +139,44 @@ await sandbox._stop()
137
139
  await sandbox._start()
138
140
  ```
139
141
 
142
+ ## Volumes and persistent files
143
+
144
+ Mount [Modal Volumes](https://modal.com/docs/guide/volumes) with the `volumes` option, then use `ModalFilesystem` to expose a mounted path to the agent's workspace file tools. Files written under the mount persist across sandbox restarts.
145
+
146
+ ```typescript
147
+ import { Workspace } from '@mastra/core/workspace'
148
+ import { ModalFilesystem, ModalSandbox } from '@mastra/modal'
149
+ import { ModalClient } from 'modal'
150
+
151
+ const modal = new ModalClient()
152
+ const privateVolume = await modal.volumes.fromName('agent-user-123', { createIfMissing: true })
153
+ const teamVolume = await modal.volumes.fromName('agent-team', { createIfMissing: true })
154
+
155
+ const sandbox = new ModalSandbox({
156
+ id: 'agent-sandbox',
157
+ workingDirectory: '/workspace',
158
+ volumes: {
159
+ '/mnt/agent': privateVolume,
160
+ '/mnt/team': teamVolume,
161
+ },
162
+ })
163
+
164
+ const workspace = new Workspace({
165
+ sandbox,
166
+ filesystem: new ModalFilesystem({ sandbox, basePath: '/mnt/agent' }),
167
+ })
168
+ ```
169
+
170
+ `ModalFilesystem` runs file operations inside the sandbox and confines all paths to `basePath`. Set `readOnly: true` to reject writes.
171
+
172
+ Writes made by other sandboxes to the same Volume are only visible after a reload. Call `reloadVolumes()` on a long-lived sandbox to pick them up:
173
+
174
+ ```typescript
175
+ await sandbox.reloadVolumes()
176
+ ```
177
+
178
+ > **Note:** Keep the working directory outside every Volume mount (for example `/workspace`). Modal can't reload a Volume that is the current working directory, so `ModalSandbox` throws if `workingDirectory` is at or inside a mount path.
179
+
140
180
  ## Background processes
141
181
 
142
182
  `ModalSandbox` includes a built-in process manager for spawning and managing background processes. Each `spawn()` call creates a new `ContainerProcess` via the Modal SDK's `Sandbox.exec()` API.
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Netlify
6
6
 
7
- Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 265 models through Mastra's model router.
7
+ Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 270 models through Mastra's model router.
8
8
 
9
9
  Learn more in the [Netlify documentation](https://docs.netlify.com/build/ai-gateway/overview/).
10
10
 
@@ -188,11 +188,15 @@ ANTHROPIC_API_KEY=ant-...
188
188
  | `openrouter/minimax/minimax-m2.5` |
189
189
  | `openrouter/minimax/minimax-m2.7` |
190
190
  | `openrouter/minimax/minimax-m3` |
191
+ | `openrouter/mistralai/devstral-2512` |
191
192
  | `openrouter/mistralai/ministral-14b-2512` |
192
193
  | `openrouter/mistralai/ministral-3b-2512` |
193
194
  | `openrouter/mistralai/ministral-8b-2512` |
194
195
  | `openrouter/mistralai/mistral-large-2407` |
196
+ | `openrouter/mistralai/mistral-large-2512` |
197
+ | `openrouter/mistralai/mistral-medium-3` |
195
198
  | `openrouter/mistralai/mistral-medium-3-5` |
199
+ | `openrouter/mistralai/mistral-medium-3.1` |
196
200
  | `openrouter/mistralai/mistral-nemo` |
197
201
  | `openrouter/mistralai/mistral-saba` |
198
202
  | `openrouter/mistralai/mistral-small-24b-instruct-2501` |
@@ -225,6 +229,7 @@ ANTHROPIC_API_KEY=ant-...
225
229
  | `openrouter/openrouter/free` |
226
230
  | `openrouter/openrouter/pareto-code` |
227
231
  | `openrouter/perceptron/perceptron-mk1` |
232
+ | `openrouter/perceptron/perceptron-mk1.5` |
228
233
  | `openrouter/perplexity/sonar` |
229
234
  | `openrouter/perplexity/sonar-deep-research` |
230
235
  | `openrouter/perplexity/sonar-pro` |
@@ -275,6 +280,7 @@ ANTHROPIC_API_KEY=ant-...
275
280
  | `openrouter/thedrummer/unslopnemo-12b` |
276
281
  | `openrouter/thinkingmachines/inkling` |
277
282
  | `openrouter/thinkingmachines/inkling-small` |
283
+ | `openrouter/typesafe/jev-router` |
278
284
  | `openrouter/undi95/remm-slerp-l2-13b` |
279
285
  | `openrouter/upstage/solar-mini4` |
280
286
  | `openrouter/upstage/solar-pro4` |
@@ -298,7 +304,6 @@ ANTHROPIC_API_KEY=ant-...
298
304
  | `openrouter/z-ai/glm-5` |
299
305
  | `openrouter/z-ai/glm-5.1` |
300
306
  | `openrouter/z-ai/glm-5.2` |
301
- | `openrouter/z-ai/glm-5.2:free` |
302
307
  | `openrouter/z-ai/glm-5.3` |
303
308
  | `openrouter/z-ai/glm-5.3-flash` |
304
309
  | `typesafe/jev-1.13.0` |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![OpenRouter logo](https://models.dev/logos/openrouter.svg)OpenRouter
6
6
 
7
- OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 385 models through Mastra's model router.
7
+ OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 384 models through Mastra's model router.
8
8
 
9
9
  Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
10
10
 
@@ -68,7 +68,6 @@ ANTHROPIC_API_KEY=ant-...
68
68
  | `amazon/nova-premier-v1` |
69
69
  | `amazon/nova-pro-v1` |
70
70
  | `anthracite-org/magnum-v4-72b` |
71
- | `anthropic/claude-3-haiku` |
72
71
  | `anthropic/claude-fable-5` |
73
72
  | `anthropic/claude-fable-5.1` |
74
73
  | `anthropic/claude-haiku-4.5` |
@@ -187,11 +186,13 @@ ANTHROPIC_API_KEY=ant-...
187
186
  | `minimax/minimax-m2.7` |
188
187
  | `minimax/minimax-m3` |
189
188
  | `mistralai/codestral-2508` |
189
+ | `mistralai/devstral-2512` |
190
190
  | `mistralai/ministral-14b-2512` |
191
191
  | `mistralai/ministral-3b-2512` |
192
192
  | `mistralai/ministral-8b-2512` |
193
193
  | `mistralai/mistral-large` |
194
194
  | `mistralai/mistral-large-2407` |
195
+ | `mistralai/mistral-large-2512` |
195
196
  | `mistralai/mistral-medium-3` |
196
197
  | `mistralai/mistral-medium-3-5` |
197
198
  | `mistralai/mistral-medium-3.1` |
@@ -212,8 +213,6 @@ ANTHROPIC_API_KEY=ant-...
212
213
  | `moonshotai/kimi-k3` |
213
214
  | `morph/morph-v3-fast` |
214
215
  | `morph/morph-v3-large` |
215
- | `nex-agi/nex-n2.5-mini:free` |
216
- | `nex-agi/nex-n2.5-pro:free` |
217
216
  | `nousresearch/hermes-3-llama-3.1-405b` |
218
217
  | `nousresearch/hermes-3-llama-3.1-70b` |
219
218
  | `nousresearch/hermes-4-405b` |
@@ -296,6 +295,7 @@ ANTHROPIC_API_KEY=ant-...
296
295
  | `openrouter/fusion` |
297
296
  | `openrouter/pareto-code` |
298
297
  | `perceptron/perceptron-mk1` |
298
+ | `perceptron/perceptron-mk1.5` |
299
299
  | `perplexity/sonar` |
300
300
  | `perplexity/sonar-deep-research` |
301
301
  | `perplexity/sonar-pro` |
@@ -417,7 +417,6 @@ ANTHROPIC_API_KEY=ant-...
417
417
  | `z-ai/glm-5-turbo` |
418
418
  | `z-ai/glm-5.1` |
419
419
  | `z-ai/glm-5.2` |
420
- | `z-ai/glm-5.2:free` |
421
420
  | `z-ai/glm-5.3` |
422
421
  | `z-ai/glm-5.3-flash` |
423
422
  | `z-ai/glm-5.3-flashx` |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Model Providers
6
6
 
7
- Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7616 models from 210 providers through a single API.
7
+ Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7618 models from 210 providers through a single API.
8
8
 
9
9
  ## Features
10
10
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![above.dev logo](https://models.dev/logos/above.svg)above.dev
6
6
 
7
- Access 9 above.dev models through Mastra's model router. Authentication is handled automatically using the `ABOVE_API_KEY` environment variable.
7
+ Access 10 above.dev models through Mastra's model router. Authentication is handled automatically using the `ABOVE_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [above.dev documentation](https://above.dev/docs).
10
10
 
@@ -38,8 +38,8 @@ for await (const chunk of stream) {
38
38
 
39
39
  | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
40
40
  | -------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
41
- | `above/deepseek-v4-flash` | 1.0M | | | | | | $0.17 | $0.66 |
42
41
  | `above/deepseek-v4-pro` | 1.0M | | | | | | $0.73 | $2 |
42
+ | `above/deepseek-v4.1-flash` | 1.0M | | | | | | $0.17 | $0.66 |
43
43
  | `above/glm-5.2` | 1.0M | | | | | | $2 | $5 |
44
44
  | `above/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
45
45
  | `above/glm-5.3-flash` | 1.0M | | | | | | $0.17 | $0.55 |
@@ -37,7 +37,7 @@ for await (const chunk of stream) {
37
37
  | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
38
  | ----------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
39
  | `cerebras/gpt-oss-120b` | 131K | | | | | | $0.35 | $0.75 |
40
- | `cerebras/qwen-3.8-27b` | 66K | | | | | | $0.99 | $1 |
40
+ | `cerebras/qwen-3.8-27b` | 131K | | | | | | $0.99 | $1 |
41
41
 
42
42
  Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
43
43
 
@@ -56,7 +56,7 @@ for await (const chunk of stream) {
56
56
  | `cortecs/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.09 | $0.17 |
57
57
  | `cortecs/deepseek-v4-pro` | 1.0M | | | | | | $2 | $3 |
58
58
  | `cortecs/deepseek-v4-pro-0813` | 1.0M | | | | | | $2 | $4 |
59
- | `cortecs/deepseek-v4.1-flash` | 1.0M | | | | | | $0.50 | $1 |
59
+ | `cortecs/deepseek-v4.1-flash` | 1.0M | | | | | | $0.30 | $1 |
60
60
  | `cortecs/devstral-2512` | 256K | | | | | | $0.48 | $2 |
61
61
  | `cortecs/gemini-2.5-flash` | 1.0M | | | | | | $0.30 | $2 |
62
62
  | `cortecs/gemini-2.5-pro` | 1.0M | | | | | | $1 | $10 |
@@ -107,7 +107,7 @@ for await (const chunk of stream) {
107
107
  | `cortecs/minimax-m2.1` | 196K | | | | | | $0.36 | $1 |
108
108
  | `cortecs/minimax-m2.5` | 196K | | | | | | $0.30 | $1 |
109
109
  | `cortecs/minimax-m2.7` | 197K | | | | | | $0.67 | $3 |
110
- | `cortecs/minimax-m3` | 1.0M | | | | | | $0.40 | $2 |
110
+ | `cortecs/minimax-m3` | 1.0M | | | | | | $0.30 | $1 |
111
111
  | `cortecs/ministral-14b-2512` | 256K | | | | | | $0.24 | $0.24 |
112
112
  | `cortecs/ministral-3b-2512` | 256K | | | | | | $0.12 | $0.12 |
113
113
  | `cortecs/ministral-8b-2512` | 256K | | | | | | $0.18 | $0.18 |
@@ -142,7 +142,7 @@ for await (const chunk of stream) {
142
142
  | `cortecs/qwen3.6-27b` | 262K | | | | | | $0.45 | $3 |
143
143
  | `cortecs/qwen3.6-35b-a3b` | 262K | | | | | | $0.17 | $0.56 |
144
144
  | `cortecs/qwen3.8-2.4t-a95b` | 262K | | | | | | $3 | $6 |
145
- | `cortecs/qwen3.8-27b` | 262K | | | | | | $0.10 | $0.40 |
145
+ | `cortecs/qwen3.8-27b` | 1.0M | | | | | | $0.10 | $0.40 |
146
146
  | `cortecs/qwen3.8-flash-next` | 262K | | | | | | $0.20 | $0.50 |
147
147
  | `cortecs/qwen3guard-gen-0.6b` | 32K | | | | | | — | — |
148
148
  | `cortecs/qwen3guard-gen-8b` | 32K | | | | | | — | — |
@@ -91,8 +91,14 @@ const response = await agent.generate("Hello!", {
91
91
 
92
92
  ### Available Options
93
93
 
94
- **thinking** (`{ type?: "adaptive" | "enabled" | "disabled" | undefined; } | undefined`)
94
+ **logprobs** (`boolean | undefined`): Whether to return log probabilities for generated tokens.
95
95
 
96
- **reasoningEffort** (`"low" | "medium" | "high" | "xhigh" | "max" | undefined`)
96
+ **topLogprobs** (`number | undefined`): Number of most likely tokens to return at each token position. Setting this option automatically enables logprobs.
97
97
 
98
- **strictJsonSchema** (`boolean | undefined`)
98
+ **userId** (`string | undefined`): An opaque identifier for the end user. DeepSeek uses this identifier for content-safety tracing and request isolation. Must contain only ASCII letters, numbers, underscores, and hyphens, and must be at most 512 characters long.
99
+
100
+ **strictJsonSchema** (`boolean | undefined`): Whether to use strict JSON schema validation for structured outputs. Only applies when the serving endpoint supports JSON schema response formats (e.g. Azure). Defaults to true.
101
+
102
+ **thinking** (`{ type?: "enabled" | "disabled" | "adaptive" | undefined; } | undefined`)
103
+
104
+ **reasoningEffort** (`"low" | "high" | "max" | "medium" | "xhigh" | undefined`)