@mastra/mcp-docs-server 1.2.19-alpha.0 → 1.2.19-alpha.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -22,6 +22,7 @@ These scorers evaluate how correct, truthful, and complete your agent's answers
22
22
  - [`tool-call-accuracy`](https://mastra.ai/reference/evals/tool-call-accuracy): Evaluates whether the LLM selects the correct tool from available options (`0-1`, higher is better)
23
23
  - [`trajectory-accuracy`](https://mastra.ai/reference/evals/trajectory-accuracy): Evaluates the expected action sequence for all span types. Covered spans include tool and model activity plus workflow steps (`0-1`, higher is better)
24
24
  - [`prompt-alignment`](https://mastra.ai/reference/evals/prompt-alignment): Measures how well agent responses align with user prompt intent, requirements, completeness, and format (`0-1`, higher is better)
25
+ - [`multi-turn-judge`](https://mastra.ai/reference/evals/multi-turn-judge): Grades every assistant turn of a [multi-turn conversation](https://mastra.ai/docs/evals/multi-turn) against a plain-English criterion (`0` or `1`)
25
26
 
26
27
  ### Context quality
27
28
 
@@ -35,6 +35,8 @@ const result = await runEvals({
35
35
 
36
36
  Each turn runs `agent.generate()` with the same thread ID, so the agent sees the full conversation history. Scorers receive the accumulated output messages from all turns.
37
37
 
38
+ > **Prebuilt LLM judges grade one turn:** Prebuilt LLM-judge scorers (rubric, answer relevancy, faithfulness, and the others listed in [Scorer compatibility](#scorer-compatibility)) read a single assistant message, so with `inputs` they grade only the last turn's response. For semantic grading across a conversation, use [`createMultiTurnJudgeScorer()`](https://mastra.ai/reference/evals/multi-turn-judge), which grades every assistant turn against one criterion, or per-turn [`turns[].scorers`](#per-turn-assertions-with-turns), where each turn grades its own output.
39
+
38
40
  ## Memory is required for cross-turn recall
39
41
 
40
42
  Multi-turn recall depends on the agent having a **memory store configured**. The shared thread ID is what lets each turn see the earlier ones, but a thread only persists history when the agent has memory. If the agent has no memory configured, the turns still run sequentially and their outputs still accumulate for scoring, but the agent won't recall earlier turns (each input runs in isolation). `runEvals` logs a warning when you use `inputs` on an agent without memory.
@@ -70,6 +72,85 @@ These details matter when writing scorers for multi-turn items:
70
72
  - **`run.output` is the accumulated output from every turn.** Output-based scorers: `checks.includes`, `checks.calledTool`, `checks.similarity`, and similar: evaluate the whole conversation. For example, `checks.calledTool('get_weather', { times: 2 })` counts calls across all turns.
71
73
  - **`run.input` is only the first turn's input.** Scorers that compare input against output (faithfulness, answer relevancy, and other input-relative LLM scorers) only see the first user message, not the full conversation. Prefer output-based checks for multi-turn, or build scorers that read the accumulated `run.output` directly.
72
74
 
75
+ ### Scorer compatibility
76
+
77
+ Accumulated output only helps if the scorer reads all of it:
78
+
79
+ | Scorer | What it sees with `inputs` |
80
+ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- |
81
+ | [Quick Checks](https://mastra.ai/docs/evals/quick-checks): `checks.calledTool`, `checks.includes`, `checks.similarity`, `checks.noToolErrors`, and the rest | The whole conversation. Tool-call checks count calls across every turn, and text checks search all assistant text |
82
+ | [Multi-turn Judge scorer](https://mastra.ai/reference/evals/multi-turn-judge) | Every assistant turn, graded together against one plain-English criterion |
83
+ | Prebuilt LLM-judge scorers: rubric, answer relevancy, answer similarity, faithfulness, hallucination, bias, toxicity, context precision, context recall, context relevance, noise sensitivity, prompt alignment, summarization | Only the last assistant message that carries text. Input-relative judges also see only the first turn's input |
84
+ | Trajectory scorers (`AgentScorerConfig.trajectory`) | The last turn's span when traces are available, otherwise tool calls from all accumulated messages |
85
+ | Custom scorers | Whatever they read from `run.output` |
86
+
87
+ ### Grading a whole conversation
88
+
89
+ To judge every turn together, use [`createMultiTurnJudgeScorer()`](https://mastra.ai/reference/evals/multi-turn-judge). It builds a transcript of every assistant turn and asks a judge model whether the conversation satisfies one plain-English criterion, scoring `1` or `0`:
90
+
91
+ ```typescript
92
+ import { runEvals } from '@mastra/core/evals'
93
+ import { createMultiTurnJudgeScorer } from '@mastra/evals/scorers/prebuilt'
94
+ import { weatherAgent } from '../agents'
95
+
96
+ const result = await runEvals({
97
+ data: [
98
+ {
99
+ inputs: [
100
+ "How's the weather in London?",
101
+ 'And Paris?',
102
+ 'Should I pack an umbrella for London?',
103
+ ],
104
+ },
105
+ ],
106
+ target: weatherAgent,
107
+ scorers: [
108
+ {
109
+ scorer: createMultiTurnJudgeScorer({
110
+ model: 'anthropic/claude-haiku-4-5',
111
+ criterion:
112
+ 'The agent provided forecasts for London and Paris, and gave weather-appropriate packing advice.',
113
+ }),
114
+ threshold: 1,
115
+ },
116
+ ],
117
+ })
118
+ ```
119
+
120
+ To grade something the criterion can't express, write your own scorer that reads the assistant messages out of `run.output`. [`extractAgentResponseMessages()`](https://mastra.ai/reference/evals/scorer-utils) returns the text of each assistant message, in order:
121
+
122
+ ```typescript
123
+ import { createScorer } from '@mastra/core/evals'
124
+ import { extractAgentResponseMessages } from '@mastra/evals/scorers/utils'
125
+ import { z } from 'zod'
126
+
127
+ export const conversationJudge = createScorer({
128
+ id: 'conversation-judge',
129
+ name: 'Conversation Judge',
130
+ description: 'Grades every assistant turn in a conversation against one criterion',
131
+ type: 'agent',
132
+ judge: {
133
+ model: 'anthropic/claude-haiku-4-5',
134
+ instructions: 'You grade multi-turn assistant transcripts against a single criterion.',
135
+ },
136
+ })
137
+ .analyze({
138
+ description: 'Judge the transcript as a whole',
139
+ outputSchema: z.object({ satisfied: z.boolean(), reason: z.string() }),
140
+ createPrompt: ({ run }) => {
141
+ const transcript = extractAgentResponseMessages(run.output)
142
+ .map((text, index) => `Turn ${index + 1}: ${text}`)
143
+ .join('\n\n')
144
+
145
+ return `Grade this conversation:\n\n${transcript}\n\nCriterion: the assistant keeps the forecast consistent across turns.`
146
+ },
147
+ })
148
+ .generateScore(({ results }) => (results.analyzeStepResult?.satisfied ? 1 : 0))
149
+ .generateReason(({ results }) => results.analyzeStepResult?.reason ?? '')
150
+ ```
151
+
152
+ Pass it to `runEvals` like any other scorer. See [Custom scorers](https://mastra.ai/docs/evals/custom-scorers) for the full `createScorer()` pipeline.
153
+
73
154
  ## Per-turn assertions with `turns`
74
155
 
75
156
  The `inputs` form scores the **accumulated** output as a whole, a single score over every turn's output. That can hide per-turn failures: an output-based check like `checks.includes('Brooklyn')` passes if _any_ turn mentions Brooklyn, even when the follow-up turn is broken.
@@ -175,4 +256,6 @@ await runEvals({
175
256
 
176
257
  - [`runEvals()` reference](https://mastra.ai/reference/evals/run-evals): Full API for `runEvals` parameters and returns
177
258
  - [Gates and verdicts](https://mastra.ai/docs/evals/gates-and-verdicts): Enforce hard requirements and quality thresholds
178
- - [Quick Checks](https://mastra.ai/docs/evals/quick-checks): Zero-LLM composable micro-scorers
259
+ - [Quick Checks](https://mastra.ai/docs/evals/quick-checks): Zero-LLM composable micro-scorers
260
+ - [Scorer utilities](https://mastra.ai/reference/evals/scorer-utils): Helpers for reading messages out of a scorer run
261
+ - [Multi-turn Judge scorer](https://mastra.ai/reference/evals/multi-turn-judge): LLM judge that grades every assistant turn together
@@ -115,15 +115,77 @@ For the step-level `scorers` API, see the [Step class reference](https://mastra.
115
115
 
116
116
  **Asynchronous execution**: Live evaluations run in the background without blocking your agent responses or workflow execution. Your AI systems remain responsive while live evaluations monitor them.
117
117
 
118
- **Sampling control**: The `sampling.rate` parameter (0-1) controls what percentage of outputs get scored:
118
+ **Sampling control**: The `sampling.rate` parameter (0-1) controls what fraction of outputs get scored:
119
119
 
120
120
  - `1.0`: Score every single response (100%)
121
121
  - `0.5`: Score half of all responses (50%)
122
122
  - `0.1`: Score 10% of responses
123
123
  - `0.0`: Disable scoring
124
124
 
125
+ Sampling is deterministic per trace: the decision is derived from the trace ID, not drawn at random. In practice:
126
+
127
+ - Scorers configured at the same rate score the same traces, so their scores are comparable on shared traffic.
128
+ - Re-running the same trace produces the same sampling decision, so sampled coverage is reproducible.
129
+
130
+ When a run has no trace (observability not configured), the decision is derived from the run ID instead. If [trace sampling](https://mastra.ai/docs/observability/tracing/overview) declined the trace, scorers skip that run entirely, so scores aren't created for traces that were never stored.
131
+
132
+ **Eligibility filters**: The optional `filter` parameter restricts which runs a scorer is eligible for, using a declarative predicate over the run's context. Filters are evaluated before sampling, so `sampling.rate` applies only to runs that match the filter:
133
+
134
+ ```typescript
135
+ export const myAgent = new Agent({
136
+ // ...
137
+ scorers: {
138
+ relevancy: {
139
+ scorer: createAnswerRelevancyScorer({ model: 'openai/gpt-5-mini' }),
140
+ filter: {
141
+ op: 'eq',
142
+ left: { path: 'requestContext.plan' },
143
+ right: { literal: 'enterprise' },
144
+ },
145
+ sampling: { type: 'ratio', rate: 0.1 },
146
+ },
147
+ },
148
+ })
149
+ ```
150
+
151
+ This scores 10% of enterprise-plan traffic and none of the rest. To score different segments at different rates, bind the same scorer twice with complementary filters.
152
+
153
+ Predicates can reference `requestContext.*`, `entity.*`, `entityType`, `source`, `threadId`, `resourceId`, and `projectId`. They support comparisons (`eq`, `ne`, `lt`, `lte`, `gt`, `gte`), membership (`in`, `notIn`), existence (`exists`, `notExists`), truthiness (`truthy`, `falsy`), and boolean composition (`and`, `or`, `not`). A filter that references an unknown root fails at agent construction rather than silently skipping scoring at runtime. Filters are plain JSON, so they're unaffected by durable agent state serialization.
154
+
155
+ Eligibility filters decide _whether a scorer runs_; to filter _which messages a scorer sees_ once it runs, use [`filterRun()`](https://mastra.ai/reference/evals/filter-run).
156
+
125
157
  **Automatic storage**: All scoring results are automatically stored in the `mastra_scorers` table in your configured database, allowing you to analyze performance trends over time.
126
158
 
159
+ ## Score persistence
160
+
161
+ Scores are persisted when the Mastra instance has `storage` configured, and when the scorer is registered on that instance. Registration is what lets Mastra resolve the scorer's metadata (name, description, type) through [`getScorerById()`](https://mastra.ai/reference/core/getScorerById) before writing the score.
162
+
163
+ Scorers you attach to an agent or a workflow step register themselves. Scorers you pass directly to [`runEvals()`](https://mastra.ai/reference/evals/run-evals), including [Quick Checks](https://mastra.ai/docs/evals/quick-checks), need the `scorers` option on the [`Mastra` class](https://mastra.ai/reference/core/mastra-class):
164
+
165
+ ```typescript
166
+ import { Mastra } from '@mastra/core'
167
+ import { LibSQLStore } from '@mastra/libsql'
168
+ import { checks } from '@mastra/evals/checks'
169
+ import { createAnswerRelevancyScorer } from '@mastra/evals/scorers/prebuilt'
170
+ import { myAgent } from './agents/my-agent'
171
+
172
+ export const mastra = new Mastra({
173
+ agents: { myAgent },
174
+ storage: new LibSQLStore({ url: 'file:./mastra.db' }),
175
+ scorers: {
176
+ calledTool: checks.calledTool('get_weather'),
177
+ includes: checks.includes('Brooklyn'),
178
+ relevancy: createAnswerRelevancyScorer({ model: 'openai/gpt-5-mini' }),
179
+ },
180
+ })
181
+ ```
182
+
183
+ The lookup only compares scorer IDs, so a registered instance can be configured differently from the one you evaluate with. Every `checks.calledTool()` instance has the id `check-called-tool`, so registering one covers all of them, whatever tool name you evaluate with.
184
+
185
+ The arguments still shape the `description` stored alongside each score, which comes from the registered instance, not the one you evaluate with. `checks.calledTool('')` is accepted and persists scores fine, but stores `Checks that "" was called`, so pass a representative value.
186
+
187
+ Skipping registration doesn't change scoring results, but each save fails with a `Scorer with id <id> not found` warning and the scores never reach the store.
188
+
127
189
  ## Trace evaluations
128
190
 
129
191
  In addition to live evaluations, you can use scorers to evaluate historical traces from your agent interactions and workflows.
@@ -7,7 +7,7 @@ A filesystem gives an agent tools for reading, writing, listing, and [searching]
7
7
  Configure files in two ways:
8
8
 
9
9
  - [Direct filesystem access](#direct-filesystem-access) uses `filesystem` with one filesystem provider or a [`CompositeFilesystem`](#manual-composition) that you create yourself.
10
- - [Mounts](#mounts) uses `mounts` to create a `CompositeFilesystem` from path-prefixed providers. When the workspace also has a static sandbox, Mastra attempts to mount each provider at its configured path inside the sandbox so commands can access the same files. Remote sandbox mounts typically use Filesystem in Userspace (FUSE).
10
+ - [Mounts](#mounts) uses `mounts` to create a `CompositeFilesystem` from path-prefixed providers. When a static sandbox and filesystem provider support mounting, Mastra automatically mounts the provider at its configured path. Remote sandbox mounts typically use Filesystem in Userspace (FUSE).
11
11
 
12
12
  Configure either `filesystem` or `mounts`, not both. Configuring both throws a `WorkspaceError` with the code `INVALID_CONFIG`.
13
13
 
@@ -71,7 +71,7 @@ See [Search](https://mastra.ai/docs/sandbox/search) to index the files for keywo
71
71
 
72
72
  ## Mounts
73
73
 
74
- Use `mounts` when programs inside a sandbox need to access persistent files by path. The agent still receives file tools, while command tools can run commands such as `ls`, `cat`, or `python` against the same files.
74
+ Use `mounts` when programs inside a sandbox need to access persistent files by path. Mastra creates the mount automatically when the sandbox and filesystem provider support it. The agent still receives file tools, while command tools can run commands such as `ls`, `cat`, or `python` against the same files.
75
75
 
76
76
  For example, mount an S3 bucket at `/workspace` inside a Daytona sandbox:
77
77
 
@@ -180,7 +180,7 @@ Use manual composition when another part of your application needs the composite
180
180
 
181
181
  ## Mount availability
182
182
 
183
- File tools and composite routing work without FUSE. Remote sandboxes can use FUSE to make cloud storage visible to commands. `LocalSandbox` uses symlinks instead.
183
+ File tools and composite routing work without a sandbox mount. When a remote sandbox and filesystem provider support mounting, `mounts` automatically uses FUSE to make the files visible to commands. `LocalSandbox` uses symlinks instead.
184
184
 
185
185
  Built-in sandbox mounting currently includes:
186
186
 
@@ -195,27 +195,27 @@ Remote mounts may require `s3fs`, `gcsfuse`, or `blobfuse2` inside the sandbox.
195
195
 
196
196
  If a sandbox mount is unavailable or fails, the workspace remains usable. File tools continue to access the provider through its SDK, but commands can't see that path. Mastra describes these providers to the agent as available through file tools only.
197
197
 
198
- ## Multi-tenant filesystems
198
+ ## Filesystems per user or thread
199
199
 
200
200
  The `filesystem` option accepts a resolver when storage should vary by request, user, role, or tenant:
201
201
 
202
202
  ```typescript
203
203
  const workspace = new Workspace({
204
204
  filesystem: ({ requestContext }) => {
205
- const tenantId = requestContext.get('tenant-id') as string
205
+ const userId = requestContext.get('user-id') as string
206
206
 
207
207
  return new S3Filesystem({
208
208
  bucket: process.env.S3_BUCKET!,
209
209
  region: process.env.S3_REGION!,
210
- prefix: `tenants/${tenantId}`,
210
+ prefix: `users/${userId}`,
211
211
  })
212
212
  },
213
213
  })
214
214
  ```
215
215
 
216
- Each tenant gets its own filesystem view, and file tools resolve the provider from the request context automatically.
216
+ Each user gets a separate filesystem view, and file tools resolve the provider from the request context automatically.
217
217
 
218
- `mounts` doesn't accept a resolver and can't be combined with a sandbox resolver. When each user or thread needs separate storage inside a separate sandbox, create and mount the provider inside the sandbox resolver. See [Multi-tenant sandboxes](https://mastra.ai/docs/sandbox/overview).
218
+ `mounts` doesn't accept a resolver and can't be combined with a sandbox resolver. When each user or thread needs separate storage inside a separate sandbox, create and mount the provider inside the sandbox resolver. See [Sandboxes per user or thread](https://mastra.ai/docs/sandbox/overview).
219
219
 
220
220
  ## Policies and containment
221
221
 
@@ -2,27 +2,27 @@
2
2
 
3
3
  # Sandboxes and filesystems
4
4
 
5
- A sandbox gives your agent an environment where it can run commands, execute code, install dependencies, and manage processes. Sandboxes can isolate potentially risky agent-run code from your host, and they can also provide a place to run longer-lived or more resource-intensive work outside your application process.
5
+ A sandbox gives your agent an isolated environment where it can run commands, execute code, install dependencies, and manage processes. This lets agents perform work that would be risky, resource-intensive, or impractical to run directly inside your application.
6
6
 
7
- [Filesystems](https://mastra.ai/docs/sandbox/filesystem) give the agent files it can read, write, and [search](https://mastra.ai/docs/sandbox/search). They can provide a knowledge base for your agent, or be shared with a sandbox so commands can work with the same files and keep data beyond the sandbox's lifetime.
7
+ Sandboxes are often temporary, so files created inside them may disappear when the environment stops. A [filesystem](https://mastra.ai/docs/sandbox/filesystem) gives the agent a place to read, write, and [search](https://mastra.ai/docs/sandbox/search) files that can outlive the sandbox. You can use one to keep outputs between runs, seed a new sandbox with existing files, or give the agent documents it can search while working. Filesystems also work without a sandbox, for example when an agent only needs a knowledge base or access to files in a service such as Google Drive.
8
8
 
9
9
  ## When to use sandboxes
10
10
 
11
- Sandboxes work well for:
11
+ Use a sandbox when an agent needs to:
12
12
 
13
- - **Coding agents and software factories**: Clone repositories and run shell, Git, build, or test workflows without giving autonomous agents access to the host system.
14
- - **Deep research and analysis**: Use specialized libraries to process downloaded PDFs or presentations and produce new artifacts.
15
- - **Long-running and parallel tasks**: Run work in separate environments without tying it to one request or sharing files and processes.
13
+ - Clone repositories and run shell, Git, build, or test workflows in a separate environment.
14
+ - Process PDFs, presentations, or other files with specialized libraries and produce new artifacts.
15
+ - Run long-lived or parallel work in separate environments with isolated files and processes.
16
16
 
17
17
  If an agent only needs to read, write, or search files, configure [direct filesystem access](https://mastra.ai/docs/sandbox/filesystem).
18
18
 
19
19
  ## Quickstart
20
20
 
21
- Give an agent a local sandbox and filesystem:
21
+ Give an agent a local sandbox:
22
22
 
23
23
  ```typescript
24
24
  import { Agent } from '@mastra/core/agent'
25
- import { LocalFilesystem, LocalSandbox, Workspace } from '@mastra/core/workspace'
25
+ import { LocalSandbox, Workspace } from '@mastra/core/workspace'
26
26
 
27
27
  export const codingAgent = new Agent({
28
28
  id: 'coding-agent',
@@ -33,20 +33,15 @@ export const codingAgent = new Agent({
33
33
  sandbox: new LocalSandbox({
34
34
  workingDirectory: './workspace',
35
35
  }),
36
- filesystem: new LocalFilesystem({
37
- basePath: './workspace',
38
- }),
39
36
  }),
40
37
  })
41
38
 
42
39
  await codingAgent.generate('List the files in the sandbox directory')
43
40
  ```
44
41
 
45
- > **Warning:** `LocalSandbox` is the quickest sandbox to set up, but commands run on the host by default. Enable [native isolation](#localsandbox) whenever possible. For applications exposed to untrusted users, use a remote or container backend with a stronger isolation boundary instead.
46
-
47
- To use a different sandbox for each user, tenant, or thread, configure a [sandbox resolver](#multi-tenant-sandboxes) instead.
42
+ > **Warning:** `LocalSandbox` runs commands on the application host by default and isn't isolated or secure. Enable [native isolation](#localsandbox), or use a remote or container sandbox when running untrusted code.
48
43
 
49
- The `LocalSandbox` above is one static instance. Every request and memory thread that uses `codingAgent` shares its files, processes, and lifecycle state. A memory thread doesn't create an isolated sandbox by itself.
44
+ A static sandbox is shared across every request and memory thread that uses the agent. Use a [resolver](#sandboxes-per-user-or-thread) when each user or thread needs a separate environment. See [Lifecycle and persistence](#lifecycle-and-persistence) for sharing and cleanup details.
50
45
 
51
46
  ## Using the sandbox
52
47
 
@@ -74,7 +69,7 @@ const workspace = new Workspace({
74
69
  })
75
70
  ```
76
71
 
77
- Omitting a tool entry keeps its capability-driven default. Set `{ enabled: false }` on one tool to remove it, or set the top-level `enabled: false` to disable generated tools by default. A per-tool `{ enabled: true }` overrides that global setting. The top-level `requireApproval` policy applies to every generated tool unless a per-tool entry overrides it.
72
+ Set `{ enabled: false }` on one tool to remove it, or set the top-level `enabled: false` to disable generated tools by default. A per-tool `{ enabled: true }` overrides that global setting. The top-level `requireApproval` policy applies to every generated tool unless a per-tool entry overrides it.
78
73
 
79
74
  See the [sandbox tools reference](https://mastra.ai/reference/workspace/workspace-class) for all generated tools and the [tool configuration reference](https://mastra.ai/reference/workspace/workspace-class) for approvals, output limits, and hooks.
80
75
 
@@ -109,17 +104,12 @@ export const startServerTool = createTool({
109
104
  })
110
105
  ```
111
106
 
112
- ## Supported backends
107
+ ## Sandboxes
113
108
 
114
109
  ### `LocalSandbox`
115
110
 
116
111
  [`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) executes commands on the same machine as your Mastra application. By default, commands run directly on the host with the permissions of the application process.
117
112
 
118
- Enable native isolation to restrict filesystem and network access at the operating-system level:
119
-
120
- - **macOS**: Seatbelt (`sandbox-exec`)
121
- - **Linux**: Bubblewrap (`bwrap`)
122
-
123
113
  ```typescript
124
114
  const sandbox = new LocalSandbox({
125
115
  workingDirectory: './workspace',
@@ -131,11 +121,13 @@ const sandbox = new LocalSandbox({
131
121
  })
132
122
  ```
133
123
 
134
- Use `LocalSandbox.detectIsolation()` to check whether Seatbelt or Bubblewrap is available on the current operating system. The [`nativeSandbox` options](https://mastra.ai/reference/workspace/local-sandbox) control network access, read-only or writable paths, write access to the working directory, and system binaries. You can also provide a custom Seatbelt profile or Bubblewrap arguments. See [`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) for the full configuration.
124
+ Enable native isolation to restrict filesystem and network access at the operating-system level. On macOS, native isolation uses Seatbelt (`sandbox-exec`). On Linux, it uses Bubblewrap (`bwrap`). Use `LocalSandbox.detectIsolation()` to check whether Seatbelt or Bubblewrap is available on the current operating system.
135
125
 
136
- ### Other backends
126
+ The [`nativeSandbox` options](https://mastra.ai/reference/workspace/local-sandbox) control network access, read-only or writable paths, write access to the working directory, and system binaries. You can also provide a custom Seatbelt profile or Bubblewrap arguments. See [`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) for the full configuration.
137
127
 
138
- Use a remote or container backend when commands need a stronger boundary from the host application or when workloads need to scale beyond the resources of your application server. Each backend has its own isolation, persistence, networking, and mount behavior:
128
+ ### Remote sandboxes
129
+
130
+ Use a remote or container sandbox when commands need a stronger boundary from the host application or when workloads need to scale beyond the resources of your application server. Each sandbox has its own isolation, persistence, networking, and mount behavior:
139
131
 
140
132
  - [AgentCore](https://mastra.ai/integrations/sandboxes/agentcore)
141
133
  - [Apple Container](https://mastra.ai/integrations/sandboxes/apple-container)
@@ -149,47 +141,40 @@ Use a remote or container backend when commands need a stronger boundary from th
149
141
  - [Railway](https://mastra.ai/integrations/sandboxes/railway)
150
142
  - [Vercel](https://mastra.ai/integrations/sandboxes/vercel)
151
143
 
152
- If Mastra doesn't support your execution backend, implement the [sandbox provider interface](https://mastra.ai/reference/workspace/sandbox) to add it.
144
+ If Mastra doesn't support your sandbox provider, implement the [sandbox provider interface](https://mastra.ai/reference/workspace/sandbox) to add it.
153
145
 
154
146
  ## Filesystem
155
147
 
156
- Each sandbox has a native filesystem that commands can use. Treat it as ephemeral because its contents follow the sandbox's lifecycle.
148
+ Every sandbox has a filesystem for commands. `LocalSandbox` uses the host filesystem, while remote sandboxes have isolated filesystems that are often temporary. For files that need to outlive a sandbox, use external storage.
149
+
150
+ Mastra filesystems connect agents and sandboxes to provider-backed storage such as Amazon S3, Google Cloud Storage, or Google Drive. Configuring one gives the agent tools to read, write, and search files. When a remote sandbox and provider support mounting, Mastra automatically mounts the storage through Filesystem in Userspace (FUSE). Commands use normal file paths while the provider stores changes outside the sandbox.
157
151
 
158
- To seed a sandbox with files or keep files between runs, give the agent a [filesystem](https://mastra.ai/docs/sandbox/filesystem). Use mounts when sandbox commands need access to the same files.
152
+ See [Filesystem](https://mastra.ai/docs/sandbox/filesystem) for providers, file tools, and mounts.
159
153
 
160
- ## Multi-tenant sandboxes
154
+ ## Sandboxes per user or thread
161
155
 
162
156
  Use a **resolver** when each user, tenant, or thread needs a separate sandbox. Set `sandboxCacheKey` to the identity that owns the sandbox so later requests reuse the same live environment.
163
157
 
164
158
  This example creates and caches one Daytona sandbox per user:
165
159
 
166
160
  ```typescript
167
- import type { RequestContext } from '@mastra/core/request-context'
168
161
  import { Workspace } from '@mastra/core/workspace'
169
162
  import { DaytonaSandbox } from '@mastra/daytona'
170
163
 
171
- const getUserId = (requestContext: RequestContext) => {
172
- const userId = requestContext.get('user-id')
173
- if (typeof userId !== 'string' || !userId) {
174
- throw new Error('A user ID is required to resolve this sandbox')
175
- }
176
-
177
- return userId
178
- }
179
-
180
164
  const workspace = new Workspace({
181
165
  sandbox: async ({ requestContext }) => {
182
- const sandbox = new DaytonaSandbox({ id: `user-${getUserId(requestContext)}` })
166
+ const userId = requestContext.get('user-id') as string
167
+ const sandbox = new DaytonaSandbox({ id: `user-${userId}` })
183
168
  await sandbox.start()
184
169
  return sandbox
185
170
  },
186
- sandboxCacheKey: ({ requestContext }) => getUserId(requestContext),
171
+ sandboxCacheKey: ({ requestContext }) => requestContext.get('user-id') as string,
187
172
  })
188
173
  ```
189
174
 
190
175
  The first request from a user creates the sandbox. Later requests with the same user ID reuse it.
191
176
 
192
- For a sandbox with persistent storage, create the filesystem and mount it inside the resolver. This example creates one Daytona sandbox and one S3 storage prefix per memory thread, then mounts the storage at `/workspace`:
177
+ For a sandbox with persistent storage, create a [filesystem](https://mastra.ai/docs/sandbox/filesystem) and mount it inside the resolver. This example creates one Daytona sandbox and one S3 storage prefix per memory thread, then mounts the storage at `/workspace`:
193
178
 
194
179
  ```typescript
195
180
  import { MASTRA_THREAD_ID_KEY, type RequestContext } from '@mastra/core/request-context'
@@ -197,14 +182,8 @@ import { Workspace } from '@mastra/core/workspace'
197
182
  import { DaytonaSandbox } from '@mastra/daytona'
198
183
  import { S3Filesystem } from '@mastra/s3'
199
184
 
200
- const getThreadId = (requestContext: RequestContext) => {
201
- const threadId = requestContext.get(MASTRA_THREAD_ID_KEY)
202
- if (typeof threadId !== 'string' || !threadId) {
203
- throw new Error('A memory thread is required to resolve this sandbox')
204
- }
205
-
206
- return threadId
207
- }
185
+ const getThreadId = (requestContext: RequestContext) =>
186
+ requestContext.get(MASTRA_THREAD_ID_KEY) as string
208
187
 
209
188
  const createThreadFilesystem = (threadId: string) =>
210
189
  new S3Filesystem({
@@ -234,25 +213,13 @@ The first request in a thread runs the resolver and starts the sandbox. Later re
234
213
 
235
214
  The example mounts the filesystem inside the resolver because static `mounts` can't be combined with a sandbox resolver. Generated tools resolve the filesystem and sandbox from the request context automatically.
236
215
 
237
- ### Resolver ownership
238
-
239
- You are responsible for every sandbox returned by a resolver. It must be ready to use when returned. Keep track of it and destroy it through your application lifecycle code when it's no longer needed. `workspace.destroy()` doesn't destroy resolver-returned sandboxes.
240
-
241
- Resolvers are incompatible with `mounts` and [`lsp: true`](https://mastra.ai/docs/sandbox/lsp), because both require a static sandbox during configuration. Using a resolver with `mounts` throws an `INVALID_CONFIG` error. With `lsp: true`, Mastra disables LSP and logs a warning.
216
+ > **Warning:** Your application owns sandboxes returned by a resolver. Destroy them and call `workspace.clearSandboxCache(cacheKey)` when the user, thread, or session ends. `workspace.destroy()` doesn't destroy resolver-returned sandboxes.
242
217
 
243
218
  ### Tool availability
244
219
 
245
- With a static sandbox, Mastra knows which capabilities the backend supports and only gives the agent the corresponding tools. Application code should also check that an optional capability is available before using it:
246
-
247
- ```typescript
248
- const processManager = workspace.sandbox?.processes
249
-
250
- if (processManager) {
251
- await processManager.spawn('pnpm dev')
252
- }
253
- ```
220
+ With a static sandbox, Mastra knows which capabilities the backend supports and only gives the agent the corresponding tools. Application code should also check that an optional capability is available before using it.
254
221
 
255
- With a [resolver-backed sandbox](#multi-tenant-sandboxes), the backend isn't known until a request runs, so Mastra initially makes all sandbox tools available. If the resolved backend doesn't support the tool the agent calls, the call fails with `SandboxFeatureNotSupportedError`.
222
+ With a [resolver-backed sandbox](#sandboxes-per-user-or-thread), the backend isn't known until a request runs, so Mastra initially makes all sandbox tools available. If the resolved backend doesn't support the tool the agent calls, the call fails with `SandboxFeatureNotSupportedError`.
256
223
 
257
224
  ## Background processes
258
225
 
@@ -292,7 +259,7 @@ Who shares a sandbox depends on where you configure it and whether you use a res
292
259
  | Resource-scoped resolver | The resolver caches one sandbox for each resource ID. |
293
260
  | Thread-scoped resolver | A memory thread keeps its sandbox across requests in that thread. |
294
261
 
295
- A static sandbox isn't automatically scoped to the current resource or memory thread. For resource or thread scope, use a resolver and set `sandboxCacheKey` to the corresponding ID. See [Multi-tenant sandboxes](#multi-tenant-sandboxes).
262
+ A static sandbox isn't automatically scoped to the current resource or memory thread. For resource or thread scope, use a resolver and set `sandboxCacheKey` to the corresponding ID. See [Sandboxes per user or thread](#sandboxes-per-user-or-thread).
296
263
 
297
264
  ### Start
298
265
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![Merge Gateway logo](https://models.dev/logos/merge-gateway.svg)Merge Gateway
4
4
 
5
- Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 174 models through Mastra's model router.
5
+ Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 175 models through Mastra's model router.
6
6
 
7
7
  Learn more in the [Merge Gateway documentation](https://docs.merge.dev/merge-gateway).
8
8
 
@@ -66,6 +66,7 @@ ANTHROPIC_API_KEY=ant-...
66
66
  | `deepseek/deepseek-v4-flash` |
67
67
  | `deepseek/deepseek-v4-flash-0731` |
68
68
  | `deepseek/deepseek-v4-pro` |
69
+ | `deepseek/deepseek-v4-pro-0423` |
69
70
  | `deepseek/deepseek-v4-pro-0813` |
70
71
  | `google/gemini-2.5-computer-use-preview-10-2025` |
71
72
  | `google/gemini-2.5-flash` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # Netlify
4
4
 
5
- Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 228 models through Mastra's model router.
5
+ Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 227 models through Mastra's model router.
6
6
 
7
7
  Learn more in the [Netlify documentation](https://docs.netlify.com/build/ai-gateway/overview/).
8
8
 
@@ -120,7 +120,6 @@ ANTHROPIC_API_KEY=ant-...
120
120
  | `openrouter/bytedance-seed/seed-2.0-mini` |
121
121
  | `openrouter/bytedance/ui-tars-1.5-7b` |
122
122
  | `openrouter/cognitivecomputations/dolphin-mistral-24b-venice-edition` |
123
- | `openrouter/deepcogito/cogito-v2.1-671b` |
124
123
  | `openrouter/deepseek/deepseek-chat` |
125
124
  | `openrouter/deepseek/deepseek-chat-v3-0324` |
126
125
  | `openrouter/deepseek/deepseek-chat-v3.1` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![OpenRouter logo](https://models.dev/logos/openrouter.svg)OpenRouter
4
4
 
5
- OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 356 models through Mastra's model router.
5
+ OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 360 models through Mastra's model router.
6
6
 
7
7
  Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
8
8
 
@@ -92,7 +92,6 @@ ANTHROPIC_API_KEY=ant-...
92
92
  | `cohere/command-r-plus-08-2024` |
93
93
  | `cohere/command-r7b-12-2024` |
94
94
  | `cohere/north-mini-code:free` |
95
- | `deepcogito/cogito-v2.1-671b` |
96
95
  | `deepseek/deepseek-chat` |
97
96
  | `deepseek/deepseek-chat-v3-0324` |
98
97
  | `deepseek/deepseek-chat-v3.1` |
@@ -104,6 +103,7 @@ ANTHROPIC_API_KEY=ant-...
104
103
  | `deepseek/deepseek-v3.2-exp` |
105
104
  | `deepseek/deepseek-v4-flash` |
106
105
  | `deepseek/deepseek-v4-flash-0731` |
106
+ | `deepseek/deepseek-v4-flash-vision-exp` |
107
107
  | `deepseek/deepseek-v4-pro` |
108
108
  | `deepseek/deepseek-v4-pro-0813` |
109
109
  | `dots-studio/dots-3-note-preview:free` |
@@ -150,6 +150,7 @@ ANTHROPIC_API_KEY=ant-...
150
150
  | `kwaipilot/kat-coder-pro-v2` |
151
151
  | `kwaipilot/kat-coder-pro-v2.5` |
152
152
  | `liquid/lfm-2.5-2.6b:free` |
153
+ | `mancer/weaver` |
153
154
  | `meituan/longcat-2.0` |
154
155
  | `meta-llama/llama-3.1-70b-instruct` |
155
156
  | `meta-llama/llama-3.1-8b-instruct` |
@@ -162,6 +163,7 @@ ANTHROPIC_API_KEY=ant-...
162
163
  | `meta/muse-glimmer-30b` |
163
164
  | `meta/muse-spark-1.1` |
164
165
  | `meta/muse-spark-1.2` |
166
+ | `meta/muse-spark-1.2-contributor` |
165
167
  | `microsoft/phi-4` |
166
168
  | `microsoft/wizardlm-2-8x22b` |
167
169
  | `minimax/minimax-01` |
@@ -366,6 +368,8 @@ ANTHROPIC_API_KEY=ant-...
366
368
  | `thedrummer/unslopnemo-12b` |
367
369
  | `thinkingmachines/inkling` |
368
370
  | `thinkingmachines/inkling-small` |
371
+ | `thinkingmachines/inkling-small:free` |
372
+ | `thinkingmachines/inkling:free` |
369
373
  | `undi95/remm-slerp-l2-13b` |
370
374
  | `upstage/solar-pro-3` |
371
375
  | `upstage/solar-pro4` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![Vercel logo](https://models.dev/logos/vercel.svg)Vercel
4
4
 
5
- Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 347 models through Mastra's model router.
5
+ Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 351 models through Mastra's model router.
6
6
 
7
7
  Learn more in the [Vercel documentation](https://ai-sdk.dev/providers/ai-sdk-providers).
8
8
 
@@ -133,6 +133,7 @@ ANTHROPIC_API_KEY=ant-...
133
133
  | `deepseek/deepseek-v3.2-thinking` |
134
134
  | `deepseek/deepseek-v4-flash` |
135
135
  | `deepseek/deepseek-v4-flash-0731` |
136
+ | `deepseek/deepseek-v4-flash-vision-exp` |
136
137
  | `deepseek/deepseek-v4-pro` |
137
138
  | `deepseek/deepseek-v4-pro-0813` |
138
139
  | `fish-audio/s1` |
@@ -352,6 +353,9 @@ ANTHROPIC_API_KEY=ant-...
352
353
  | `spacexai/grok-voice-think-fast-2.0` |
353
354
  | `stepfun/step-3.5-flash` |
354
355
  | `stepfun/step-3.7-flash` |
356
+ | `tencent/hy-mt2-lite` |
357
+ | `tencent/hy-mt2-plus` |
358
+ | `tencent/hy-mt2-pro` |
355
359
  | `tencent/hy3` |
356
360
  | `thinkingmachines/inkling` |
357
361
  | `thinkingmachines/inkling-small` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # Model Providers
4
4
 
5
- Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 6740 models from 180 providers through a single API.
5
+ Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 6767 models from 180 providers through a single API.
6
6
 
7
7
  ## Features
8
8