@mastra/mcp-docs-server 1.2.18 → 1.2.19-alpha.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -22,6 +22,7 @@ These scorers evaluate how correct, truthful, and complete your agent's answers
22
22
  - [`tool-call-accuracy`](https://mastra.ai/reference/evals/tool-call-accuracy): Evaluates whether the LLM selects the correct tool from available options (`0-1`, higher is better)
23
23
  - [`trajectory-accuracy`](https://mastra.ai/reference/evals/trajectory-accuracy): Evaluates the expected action sequence for all span types. Covered spans include tool and model activity plus workflow steps (`0-1`, higher is better)
24
24
  - [`prompt-alignment`](https://mastra.ai/reference/evals/prompt-alignment): Measures how well agent responses align with user prompt intent, requirements, completeness, and format (`0-1`, higher is better)
25
+ - [`multi-turn-judge`](https://mastra.ai/reference/evals/multi-turn-judge): Grades every assistant turn of a [multi-turn conversation](https://mastra.ai/docs/evals/multi-turn) against a plain-English criterion (`0` or `1`)
25
26
 
26
27
  ### Context quality
27
28
 
@@ -35,6 +35,8 @@ const result = await runEvals({
35
35
 
36
36
  Each turn runs `agent.generate()` with the same thread ID, so the agent sees the full conversation history. Scorers receive the accumulated output messages from all turns.
37
37
 
38
+ > **Prebuilt LLM judges grade one turn:** Prebuilt LLM-judge scorers (rubric, answer relevancy, faithfulness, and the others listed in [Scorer compatibility](#scorer-compatibility)) read a single assistant message, so with `inputs` they grade only the last turn's response. For semantic grading across a conversation, use [`createMultiTurnJudgeScorer()`](https://mastra.ai/reference/evals/multi-turn-judge), which grades every assistant turn against one criterion, or per-turn [`turns[].scorers`](#per-turn-assertions-with-turns), where each turn grades its own output.
39
+
38
40
  ## Memory is required for cross-turn recall
39
41
 
40
42
  Multi-turn recall depends on the agent having a **memory store configured**. The shared thread ID is what lets each turn see the earlier ones, but a thread only persists history when the agent has memory. If the agent has no memory configured, the turns still run sequentially and their outputs still accumulate for scoring, but the agent won't recall earlier turns (each input runs in isolation). `runEvals` logs a warning when you use `inputs` on an agent without memory.
@@ -70,6 +72,85 @@ These details matter when writing scorers for multi-turn items:
70
72
  - **`run.output` is the accumulated output from every turn.** Output-based scorers: `checks.includes`, `checks.calledTool`, `checks.similarity`, and similar: evaluate the whole conversation. For example, `checks.calledTool('get_weather', { times: 2 })` counts calls across all turns.
71
73
  - **`run.input` is only the first turn's input.** Scorers that compare input against output (faithfulness, answer relevancy, and other input-relative LLM scorers) only see the first user message, not the full conversation. Prefer output-based checks for multi-turn, or build scorers that read the accumulated `run.output` directly.
72
74
 
75
+ ### Scorer compatibility
76
+
77
+ Accumulated output only helps if the scorer reads all of it:
78
+
79
+ | Scorer | What it sees with `inputs` |
80
+ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- |
81
+ | [Quick Checks](https://mastra.ai/docs/evals/quick-checks): `checks.calledTool`, `checks.includes`, `checks.similarity`, `checks.noToolErrors`, and the rest | The whole conversation. Tool-call checks count calls across every turn, and text checks search all assistant text |
82
+ | [Multi-turn Judge scorer](https://mastra.ai/reference/evals/multi-turn-judge) | Every assistant turn, graded together against one plain-English criterion |
83
+ | Prebuilt LLM-judge scorers: rubric, answer relevancy, answer similarity, faithfulness, hallucination, bias, toxicity, context precision, context recall, context relevance, noise sensitivity, prompt alignment, summarization | Only the last assistant message that carries text. Input-relative judges also see only the first turn's input |
84
+ | Trajectory scorers (`AgentScorerConfig.trajectory`) | The last turn's span when traces are available, otherwise tool calls from all accumulated messages |
85
+ | Custom scorers | Whatever they read from `run.output` |
86
+
87
+ ### Grading a whole conversation
88
+
89
+ To judge every turn together, use [`createMultiTurnJudgeScorer()`](https://mastra.ai/reference/evals/multi-turn-judge). It builds a transcript of every assistant turn and asks a judge model whether the conversation satisfies one plain-English criterion, scoring `1` or `0`:
90
+
91
+ ```typescript
92
+ import { runEvals } from '@mastra/core/evals'
93
+ import { createMultiTurnJudgeScorer } from '@mastra/evals/scorers/prebuilt'
94
+ import { weatherAgent } from '../agents'
95
+
96
+ const result = await runEvals({
97
+ data: [
98
+ {
99
+ inputs: [
100
+ "How's the weather in London?",
101
+ 'And Paris?',
102
+ 'Should I pack an umbrella for London?',
103
+ ],
104
+ },
105
+ ],
106
+ target: weatherAgent,
107
+ scorers: [
108
+ {
109
+ scorer: createMultiTurnJudgeScorer({
110
+ model: 'anthropic/claude-haiku-4-5',
111
+ criterion:
112
+ 'The agent provided forecasts for London and Paris, and gave weather-appropriate packing advice.',
113
+ }),
114
+ threshold: 1,
115
+ },
116
+ ],
117
+ })
118
+ ```
119
+
120
+ To grade something the criterion can't express, write your own scorer that reads the assistant messages out of `run.output`. [`extractAgentResponseMessages()`](https://mastra.ai/reference/evals/scorer-utils) returns the text of each assistant message, in order:
121
+
122
+ ```typescript
123
+ import { createScorer } from '@mastra/core/evals'
124
+ import { extractAgentResponseMessages } from '@mastra/evals/scorers/utils'
125
+ import { z } from 'zod'
126
+
127
+ export const conversationJudge = createScorer({
128
+ id: 'conversation-judge',
129
+ name: 'Conversation Judge',
130
+ description: 'Grades every assistant turn in a conversation against one criterion',
131
+ type: 'agent',
132
+ judge: {
133
+ model: 'anthropic/claude-haiku-4-5',
134
+ instructions: 'You grade multi-turn assistant transcripts against a single criterion.',
135
+ },
136
+ })
137
+ .analyze({
138
+ description: 'Judge the transcript as a whole',
139
+ outputSchema: z.object({ satisfied: z.boolean(), reason: z.string() }),
140
+ createPrompt: ({ run }) => {
141
+ const transcript = extractAgentResponseMessages(run.output)
142
+ .map((text, index) => `Turn ${index + 1}: ${text}`)
143
+ .join('\n\n')
144
+
145
+ return `Grade this conversation:\n\n${transcript}\n\nCriterion: the assistant keeps the forecast consistent across turns.`
146
+ },
147
+ })
148
+ .generateScore(({ results }) => (results.analyzeStepResult?.satisfied ? 1 : 0))
149
+ .generateReason(({ results }) => results.analyzeStepResult?.reason ?? '')
150
+ ```
151
+
152
+ Pass it to `runEvals` like any other scorer. See [Custom scorers](https://mastra.ai/docs/evals/custom-scorers) for the full `createScorer()` pipeline.
153
+
73
154
  ## Per-turn assertions with `turns`
74
155
 
75
156
  The `inputs` form scores the **accumulated** output as a whole, a single score over every turn's output. That can hide per-turn failures: an output-based check like `checks.includes('Brooklyn')` passes if _any_ turn mentions Brooklyn, even when the follow-up turn is broken.
@@ -175,4 +256,6 @@ await runEvals({
175
256
 
176
257
  - [`runEvals()` reference](https://mastra.ai/reference/evals/run-evals): Full API for `runEvals` parameters and returns
177
258
  - [Gates and verdicts](https://mastra.ai/docs/evals/gates-and-verdicts): Enforce hard requirements and quality thresholds
178
- - [Quick Checks](https://mastra.ai/docs/evals/quick-checks): Zero-LLM composable micro-scorers
259
+ - [Quick Checks](https://mastra.ai/docs/evals/quick-checks): Zero-LLM composable micro-scorers
260
+ - [Scorer utilities](https://mastra.ai/reference/evals/scorer-utils): Helpers for reading messages out of a scorer run
261
+ - [Multi-turn Judge scorer](https://mastra.ai/reference/evals/multi-turn-judge): LLM judge that grades every assistant turn together
@@ -124,6 +124,36 @@ For the step-level `scorers` API, see the [Step class reference](https://mastra.
124
124
 
125
125
  **Automatic storage**: All scoring results are automatically stored in the `mastra_scorers` table in your configured database, allowing you to analyze performance trends over time.
126
126
 
127
+ ## Score persistence
128
+
129
+ Scores are persisted when the Mastra instance has `storage` configured, and when the scorer is registered on that instance. Registration is what lets Mastra resolve the scorer's metadata (name, description, type) through [`getScorerById()`](https://mastra.ai/reference/core/getScorerById) before writing the score.
130
+
131
+ Scorers you attach to an agent or a workflow step register themselves. Scorers you pass directly to [`runEvals()`](https://mastra.ai/reference/evals/run-evals), including [Quick Checks](https://mastra.ai/docs/evals/quick-checks), need the `scorers` option on the [`Mastra` class](https://mastra.ai/reference/core/mastra-class):
132
+
133
+ ```typescript
134
+ import { Mastra } from '@mastra/core'
135
+ import { LibSQLStore } from '@mastra/libsql'
136
+ import { checks } from '@mastra/evals/checks'
137
+ import { createAnswerRelevancyScorer } from '@mastra/evals/scorers/prebuilt'
138
+ import { myAgent } from './agents/my-agent'
139
+
140
+ export const mastra = new Mastra({
141
+ agents: { myAgent },
142
+ storage: new LibSQLStore({ url: 'file:./mastra.db' }),
143
+ scorers: {
144
+ calledTool: checks.calledTool('get_weather'),
145
+ includes: checks.includes('Brooklyn'),
146
+ relevancy: createAnswerRelevancyScorer({ model: 'openai/gpt-5-mini' }),
147
+ },
148
+ })
149
+ ```
150
+
151
+ The lookup only compares scorer IDs, so a registered instance can be configured differently from the one you evaluate with. Every `checks.calledTool()` instance has the id `check-called-tool`, so registering one covers all of them, whatever tool name you evaluate with.
152
+
153
+ The arguments still shape the `description` stored alongside each score, which comes from the registered instance, not the one you evaluate with. `checks.calledTool('')` is accepted and persists scores fine, but stores `Checks that "" was called`, so pass a representative value.
154
+
155
+ Skipping registration doesn't change scoring results, but each save fails with a `Scorer with id <id> not found` warning and the scores never reach the store.
156
+
127
157
  ## Trace evaluations
128
158
 
129
159
  In addition to live evaluations, you can use scorers to evaluate historical traces from your agent interactions and workflows.
@@ -7,7 +7,7 @@ A filesystem gives an agent tools for reading, writing, listing, and [searching]
7
7
  Configure files in two ways:
8
8
 
9
9
  - [Direct filesystem access](#direct-filesystem-access) uses `filesystem` with one filesystem provider or a [`CompositeFilesystem`](#manual-composition) that you create yourself.
10
- - [Mounts](#mounts) uses `mounts` to create a `CompositeFilesystem` from path-prefixed providers. When the workspace also has a static sandbox, Mastra attempts to mount each provider at its configured path inside the sandbox so commands can access the same files. Remote sandbox mounts typically use Filesystem in Userspace (FUSE).
10
+ - [Mounts](#mounts) uses `mounts` to create a `CompositeFilesystem` from path-prefixed providers. When a static sandbox and filesystem provider support mounting, Mastra automatically mounts the provider at its configured path. Remote sandbox mounts typically use Filesystem in Userspace (FUSE).
11
11
 
12
12
  Configure either `filesystem` or `mounts`, not both. Configuring both throws a `WorkspaceError` with the code `INVALID_CONFIG`.
13
13
 
@@ -71,7 +71,7 @@ See [Search](https://mastra.ai/docs/sandbox/search) to index the files for keywo
71
71
 
72
72
  ## Mounts
73
73
 
74
- Use `mounts` when programs inside a sandbox need to access persistent files by path. The agent still receives file tools, while command tools can run commands such as `ls`, `cat`, or `python` against the same files.
74
+ Use `mounts` when programs inside a sandbox need to access persistent files by path. Mastra creates the mount automatically when the sandbox and filesystem provider support it. The agent still receives file tools, while command tools can run commands such as `ls`, `cat`, or `python` against the same files.
75
75
 
76
76
  For example, mount an S3 bucket at `/workspace` inside a Daytona sandbox:
77
77
 
@@ -180,7 +180,7 @@ Use manual composition when another part of your application needs the composite
180
180
 
181
181
  ## Mount availability
182
182
 
183
- File tools and composite routing work without FUSE. Remote sandboxes can use FUSE to make cloud storage visible to commands. `LocalSandbox` uses symlinks instead.
183
+ File tools and composite routing work without a sandbox mount. When a remote sandbox and filesystem provider support mounting, `mounts` automatically uses FUSE to make the files visible to commands. `LocalSandbox` uses symlinks instead.
184
184
 
185
185
  Built-in sandbox mounting currently includes:
186
186
 
@@ -195,27 +195,27 @@ Remote mounts may require `s3fs`, `gcsfuse`, or `blobfuse2` inside the sandbox.
195
195
 
196
196
  If a sandbox mount is unavailable or fails, the workspace remains usable. File tools continue to access the provider through its SDK, but commands can't see that path. Mastra describes these providers to the agent as available through file tools only.
197
197
 
198
- ## Multi-tenant filesystems
198
+ ## Filesystems per user or thread
199
199
 
200
200
  The `filesystem` option accepts a resolver when storage should vary by request, user, role, or tenant:
201
201
 
202
202
  ```typescript
203
203
  const workspace = new Workspace({
204
204
  filesystem: ({ requestContext }) => {
205
- const tenantId = requestContext.get('tenant-id') as string
205
+ const userId = requestContext.get('user-id') as string
206
206
 
207
207
  return new S3Filesystem({
208
208
  bucket: process.env.S3_BUCKET!,
209
209
  region: process.env.S3_REGION!,
210
- prefix: `tenants/${tenantId}`,
210
+ prefix: `users/${userId}`,
211
211
  })
212
212
  },
213
213
  })
214
214
  ```
215
215
 
216
- Each tenant gets its own filesystem view, and file tools resolve the provider from the request context automatically.
216
+ Each user gets a separate filesystem view, and file tools resolve the provider from the request context automatically.
217
217
 
218
- `mounts` doesn't accept a resolver and can't be combined with a sandbox resolver. When each user or thread needs separate storage inside a separate sandbox, create and mount the provider inside the sandbox resolver. See [Multi-tenant sandboxes](https://mastra.ai/docs/sandbox/overview).
218
+ `mounts` doesn't accept a resolver and can't be combined with a sandbox resolver. When each user or thread needs separate storage inside a separate sandbox, create and mount the provider inside the sandbox resolver. See [Sandboxes per user or thread](https://mastra.ai/docs/sandbox/overview).
219
219
 
220
220
  ## Policies and containment
221
221
 
@@ -2,27 +2,27 @@
2
2
 
3
3
  # Sandboxes and filesystems
4
4
 
5
- A sandbox gives your agent an environment where it can run commands, execute code, install dependencies, and manage processes. Sandboxes can isolate potentially risky agent-run code from your host, and they can also provide a place to run longer-lived or more resource-intensive work outside your application process.
5
+ A sandbox gives your agent an isolated environment where it can run commands, execute code, install dependencies, and manage processes. This lets agents perform work that would be risky, resource-intensive, or impractical to run directly inside your application.
6
6
 
7
- [Filesystems](https://mastra.ai/docs/sandbox/filesystem) give the agent files it can read, write, and [search](https://mastra.ai/docs/sandbox/search). They can provide a knowledge base for your agent, or be shared with a sandbox so commands can work with the same files and keep data beyond the sandbox's lifetime.
7
+ Sandboxes are often temporary, so files created inside them may disappear when the environment stops. A [filesystem](https://mastra.ai/docs/sandbox/filesystem) gives the agent a place to read, write, and [search](https://mastra.ai/docs/sandbox/search) files that can outlive the sandbox. You can use one to keep outputs between runs, seed a new sandbox with existing files, or give the agent documents it can search while working. Filesystems also work without a sandbox, for example when an agent only needs a knowledge base or access to files in a service such as Google Drive.
8
8
 
9
9
  ## When to use sandboxes
10
10
 
11
- Sandboxes work well for:
11
+ Use a sandbox when an agent needs to:
12
12
 
13
- - **Coding agents and software factories**: Clone repositories and run shell, Git, build, or test workflows without giving autonomous agents access to the host system.
14
- - **Deep research and analysis**: Use specialized libraries to process downloaded PDFs or presentations and produce new artifacts.
15
- - **Long-running and parallel tasks**: Run work in separate environments without tying it to one request or sharing files and processes.
13
+ - Clone repositories and run shell, Git, build, or test workflows in a separate environment.
14
+ - Process PDFs, presentations, or other files with specialized libraries and produce new artifacts.
15
+ - Run long-lived or parallel work in separate environments with isolated files and processes.
16
16
 
17
17
  If an agent only needs to read, write, or search files, configure [direct filesystem access](https://mastra.ai/docs/sandbox/filesystem).
18
18
 
19
19
  ## Quickstart
20
20
 
21
- Give an agent a local sandbox and filesystem:
21
+ Give an agent a local sandbox:
22
22
 
23
23
  ```typescript
24
24
  import { Agent } from '@mastra/core/agent'
25
- import { LocalFilesystem, LocalSandbox, Workspace } from '@mastra/core/workspace'
25
+ import { LocalSandbox, Workspace } from '@mastra/core/workspace'
26
26
 
27
27
  export const codingAgent = new Agent({
28
28
  id: 'coding-agent',
@@ -33,20 +33,15 @@ export const codingAgent = new Agent({
33
33
  sandbox: new LocalSandbox({
34
34
  workingDirectory: './workspace',
35
35
  }),
36
- filesystem: new LocalFilesystem({
37
- basePath: './workspace',
38
- }),
39
36
  }),
40
37
  })
41
38
 
42
39
  await codingAgent.generate('List the files in the sandbox directory')
43
40
  ```
44
41
 
45
- > **Warning:** `LocalSandbox` is the quickest sandbox to set up, but commands run on the host by default. Enable [native isolation](#localsandbox) whenever possible. For applications exposed to untrusted users, use a remote or container backend with a stronger isolation boundary instead.
46
-
47
- To use a different sandbox for each user, tenant, or thread, configure a [sandbox resolver](#multi-tenant-sandboxes) instead.
42
+ > **Warning:** `LocalSandbox` runs commands on the application host by default and isn't isolated or secure. Enable [native isolation](#localsandbox), or use a remote or container sandbox when running untrusted code.
48
43
 
49
- The `LocalSandbox` above is one static instance. Every request and memory thread that uses `codingAgent` shares its files, processes, and lifecycle state. A memory thread doesn't create an isolated sandbox by itself.
44
+ A static sandbox is shared across every request and memory thread that uses the agent. Use a [resolver](#sandboxes-per-user-or-thread) when each user or thread needs a separate environment. See [Lifecycle and persistence](#lifecycle-and-persistence) for sharing and cleanup details.
50
45
 
51
46
  ## Using the sandbox
52
47
 
@@ -74,7 +69,7 @@ const workspace = new Workspace({
74
69
  })
75
70
  ```
76
71
 
77
- Omitting a tool entry keeps its capability-driven default. Set `{ enabled: false }` on one tool to remove it, or set the top-level `enabled: false` to disable generated tools by default. A per-tool `{ enabled: true }` overrides that global setting. The top-level `requireApproval` policy applies to every generated tool unless a per-tool entry overrides it.
72
+ Set `{ enabled: false }` on one tool to remove it, or set the top-level `enabled: false` to disable generated tools by default. A per-tool `{ enabled: true }` overrides that global setting. The top-level `requireApproval` policy applies to every generated tool unless a per-tool entry overrides it.
78
73
 
79
74
  See the [sandbox tools reference](https://mastra.ai/reference/workspace/workspace-class) for all generated tools and the [tool configuration reference](https://mastra.ai/reference/workspace/workspace-class) for approvals, output limits, and hooks.
80
75
 
@@ -109,17 +104,12 @@ export const startServerTool = createTool({
109
104
  })
110
105
  ```
111
106
 
112
- ## Supported backends
107
+ ## Sandboxes
113
108
 
114
109
  ### `LocalSandbox`
115
110
 
116
111
  [`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) executes commands on the same machine as your Mastra application. By default, commands run directly on the host with the permissions of the application process.
117
112
 
118
- Enable native isolation to restrict filesystem and network access at the operating-system level:
119
-
120
- - **macOS**: Seatbelt (`sandbox-exec`)
121
- - **Linux**: Bubblewrap (`bwrap`)
122
-
123
113
  ```typescript
124
114
  const sandbox = new LocalSandbox({
125
115
  workingDirectory: './workspace',
@@ -131,11 +121,13 @@ const sandbox = new LocalSandbox({
131
121
  })
132
122
  ```
133
123
 
134
- Use `LocalSandbox.detectIsolation()` to check whether Seatbelt or Bubblewrap is available on the current operating system. The [`nativeSandbox` options](https://mastra.ai/reference/workspace/local-sandbox) control network access, read-only or writable paths, write access to the working directory, and system binaries. You can also provide a custom Seatbelt profile or Bubblewrap arguments. See [`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) for the full configuration.
124
+ Enable native isolation to restrict filesystem and network access at the operating-system level. On macOS, native isolation uses Seatbelt (`sandbox-exec`). On Linux, it uses Bubblewrap (`bwrap`). Use `LocalSandbox.detectIsolation()` to check whether Seatbelt or Bubblewrap is available on the current operating system.
135
125
 
136
- ### Other backends
126
+ The [`nativeSandbox` options](https://mastra.ai/reference/workspace/local-sandbox) control network access, read-only or writable paths, write access to the working directory, and system binaries. You can also provide a custom Seatbelt profile or Bubblewrap arguments. See [`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) for the full configuration.
137
127
 
138
- Use a remote or container backend when commands need a stronger boundary from the host application or when workloads need to scale beyond the resources of your application server. Each backend has its own isolation, persistence, networking, and mount behavior:
128
+ ### Remote sandboxes
129
+
130
+ Use a remote or container sandbox when commands need a stronger boundary from the host application or when workloads need to scale beyond the resources of your application server. Each sandbox has its own isolation, persistence, networking, and mount behavior:
139
131
 
140
132
  - [AgentCore](https://mastra.ai/integrations/sandboxes/agentcore)
141
133
  - [Apple Container](https://mastra.ai/integrations/sandboxes/apple-container)
@@ -149,47 +141,40 @@ Use a remote or container backend when commands need a stronger boundary from th
149
141
  - [Railway](https://mastra.ai/integrations/sandboxes/railway)
150
142
  - [Vercel](https://mastra.ai/integrations/sandboxes/vercel)
151
143
 
152
- If Mastra doesn't support your execution backend, implement the [sandbox provider interface](https://mastra.ai/reference/workspace/sandbox) to add it.
144
+ If Mastra doesn't support your sandbox provider, implement the [sandbox provider interface](https://mastra.ai/reference/workspace/sandbox) to add it.
153
145
 
154
146
  ## Filesystem
155
147
 
156
- Each sandbox has a native filesystem that commands can use. Treat it as ephemeral because its contents follow the sandbox's lifecycle.
148
+ Every sandbox has a filesystem for commands. `LocalSandbox` uses the host filesystem, while remote sandboxes have isolated filesystems that are often temporary. For files that need to outlive a sandbox, use external storage.
149
+
150
+ Mastra filesystems connect agents and sandboxes to provider-backed storage such as Amazon S3, Google Cloud Storage, or Google Drive. Configuring one gives the agent tools to read, write, and search files. When a remote sandbox and provider support mounting, Mastra automatically mounts the storage through Filesystem in Userspace (FUSE). Commands use normal file paths while the provider stores changes outside the sandbox.
157
151
 
158
- To seed a sandbox with files or keep files between runs, give the agent a [filesystem](https://mastra.ai/docs/sandbox/filesystem). Use mounts when sandbox commands need access to the same files.
152
+ See [Filesystem](https://mastra.ai/docs/sandbox/filesystem) for providers, file tools, and mounts.
159
153
 
160
- ## Multi-tenant sandboxes
154
+ ## Sandboxes per user or thread
161
155
 
162
156
  Use a **resolver** when each user, tenant, or thread needs a separate sandbox. Set `sandboxCacheKey` to the identity that owns the sandbox so later requests reuse the same live environment.
163
157
 
164
158
  This example creates and caches one Daytona sandbox per user:
165
159
 
166
160
  ```typescript
167
- import type { RequestContext } from '@mastra/core/request-context'
168
161
  import { Workspace } from '@mastra/core/workspace'
169
162
  import { DaytonaSandbox } from '@mastra/daytona'
170
163
 
171
- const getUserId = (requestContext: RequestContext) => {
172
- const userId = requestContext.get('user-id')
173
- if (typeof userId !== 'string' || !userId) {
174
- throw new Error('A user ID is required to resolve this sandbox')
175
- }
176
-
177
- return userId
178
- }
179
-
180
164
  const workspace = new Workspace({
181
165
  sandbox: async ({ requestContext }) => {
182
- const sandbox = new DaytonaSandbox({ id: `user-${getUserId(requestContext)}` })
166
+ const userId = requestContext.get('user-id') as string
167
+ const sandbox = new DaytonaSandbox({ id: `user-${userId}` })
183
168
  await sandbox.start()
184
169
  return sandbox
185
170
  },
186
- sandboxCacheKey: ({ requestContext }) => getUserId(requestContext),
171
+ sandboxCacheKey: ({ requestContext }) => requestContext.get('user-id') as string,
187
172
  })
188
173
  ```
189
174
 
190
175
  The first request from a user creates the sandbox. Later requests with the same user ID reuse it.
191
176
 
192
- For a sandbox with persistent storage, create the filesystem and mount it inside the resolver. This example creates one Daytona sandbox and one S3 storage prefix per memory thread, then mounts the storage at `/workspace`:
177
+ For a sandbox with persistent storage, create a [filesystem](https://mastra.ai/docs/sandbox/filesystem) and mount it inside the resolver. This example creates one Daytona sandbox and one S3 storage prefix per memory thread, then mounts the storage at `/workspace`:
193
178
 
194
179
  ```typescript
195
180
  import { MASTRA_THREAD_ID_KEY, type RequestContext } from '@mastra/core/request-context'
@@ -197,14 +182,8 @@ import { Workspace } from '@mastra/core/workspace'
197
182
  import { DaytonaSandbox } from '@mastra/daytona'
198
183
  import { S3Filesystem } from '@mastra/s3'
199
184
 
200
- const getThreadId = (requestContext: RequestContext) => {
201
- const threadId = requestContext.get(MASTRA_THREAD_ID_KEY)
202
- if (typeof threadId !== 'string' || !threadId) {
203
- throw new Error('A memory thread is required to resolve this sandbox')
204
- }
205
-
206
- return threadId
207
- }
185
+ const getThreadId = (requestContext: RequestContext) =>
186
+ requestContext.get(MASTRA_THREAD_ID_KEY) as string
208
187
 
209
188
  const createThreadFilesystem = (threadId: string) =>
210
189
  new S3Filesystem({
@@ -234,25 +213,13 @@ The first request in a thread runs the resolver and starts the sandbox. Later re
234
213
 
235
214
  The example mounts the filesystem inside the resolver because static `mounts` can't be combined with a sandbox resolver. Generated tools resolve the filesystem and sandbox from the request context automatically.
236
215
 
237
- ### Resolver ownership
238
-
239
- You are responsible for every sandbox returned by a resolver. It must be ready to use when returned. Keep track of it and destroy it through your application lifecycle code when it's no longer needed. `workspace.destroy()` doesn't destroy resolver-returned sandboxes.
240
-
241
- Resolvers are incompatible with `mounts` and [`lsp: true`](https://mastra.ai/docs/sandbox/lsp), because both require a static sandbox during configuration. Using a resolver with `mounts` throws an `INVALID_CONFIG` error. With `lsp: true`, Mastra disables LSP and logs a warning.
216
+ > **Warning:** Your application owns sandboxes returned by a resolver. Destroy them and call `workspace.clearSandboxCache(cacheKey)` when the user, thread, or session ends. `workspace.destroy()` doesn't destroy resolver-returned sandboxes.
242
217
 
243
218
  ### Tool availability
244
219
 
245
- With a static sandbox, Mastra knows which capabilities the backend supports and only gives the agent the corresponding tools. Application code should also check that an optional capability is available before using it:
246
-
247
- ```typescript
248
- const processManager = workspace.sandbox?.processes
249
-
250
- if (processManager) {
251
- await processManager.spawn('pnpm dev')
252
- }
253
- ```
220
+ With a static sandbox, Mastra knows which capabilities the backend supports and only gives the agent the corresponding tools. Application code should also check that an optional capability is available before using it.
254
221
 
255
- With a [resolver-backed sandbox](#multi-tenant-sandboxes), the backend isn't known until a request runs, so Mastra initially makes all sandbox tools available. If the resolved backend doesn't support the tool the agent calls, the call fails with `SandboxFeatureNotSupportedError`.
222
+ With a [resolver-backed sandbox](#sandboxes-per-user-or-thread), the backend isn't known until a request runs, so Mastra initially makes all sandbox tools available. If the resolved backend doesn't support the tool the agent calls, the call fails with `SandboxFeatureNotSupportedError`.
256
223
 
257
224
  ## Background processes
258
225
 
@@ -292,7 +259,7 @@ Who shares a sandbox depends on where you configure it and whether you use a res
292
259
  | Resource-scoped resolver | The resolver caches one sandbox for each resource ID. |
293
260
  | Thread-scoped resolver | A memory thread keeps its sandbox across requests in that thread. |
294
261
 
295
- A static sandbox isn't automatically scoped to the current resource or memory thread. For resource or thread scope, use a resolver and set `sandboxCacheKey` to the corresponding ID. See [Multi-tenant sandboxes](#multi-tenant-sandboxes).
262
+ A static sandbox isn't automatically scoped to the current resource or memory thread. For resource or thread scope, use a resolver and set `sandboxCacheKey` to the corresponding ID. See [Sandboxes per user or thread](#sandboxes-per-user-or-thread).
296
263
 
297
264
  ### Start
298
265
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![OpenRouter logo](https://models.dev/logos/openrouter.svg)OpenRouter
4
4
 
5
- OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 356 models through Mastra's model router.
5
+ OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 357 models through Mastra's model router.
6
6
 
7
7
  Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
8
8
 
@@ -150,6 +150,7 @@ ANTHROPIC_API_KEY=ant-...
150
150
  | `kwaipilot/kat-coder-pro-v2` |
151
151
  | `kwaipilot/kat-coder-pro-v2.5` |
152
152
  | `liquid/lfm-2.5-2.6b:free` |
153
+ | `mancer/weaver` |
153
154
  | `meituan/longcat-2.0` |
154
155
  | `meta-llama/llama-3.1-70b-instruct` |
155
156
  | `meta-llama/llama-3.1-8b-instruct` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![Vercel logo](https://models.dev/logos/vercel.svg)Vercel
4
4
 
5
- Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 347 models through Mastra's model router.
5
+ Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 350 models through Mastra's model router.
6
6
 
7
7
  Learn more in the [Vercel documentation](https://ai-sdk.dev/providers/ai-sdk-providers).
8
8
 
@@ -352,6 +352,9 @@ ANTHROPIC_API_KEY=ant-...
352
352
  | `spacexai/grok-voice-think-fast-2.0` |
353
353
  | `stepfun/step-3.5-flash` |
354
354
  | `stepfun/step-3.7-flash` |
355
+ | `tencent/hy-mt2-lite` |
356
+ | `tencent/hy-mt2-plus` |
357
+ | `tencent/hy-mt2-pro` |
355
358
  | `tencent/hy3` |
356
359
  | `thinkingmachines/inkling` |
357
360
  | `thinkingmachines/inkling-small` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # Model Providers
4
4
 
5
- Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 6740 models from 180 providers through a single API.
5
+ Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 6747 models from 180 providers through a single API.
6
6
 
7
7
  ## Features
8
8
 
@@ -68,18 +68,18 @@ for await (const chunk of stream) {
68
68
  | `digitalocean/fal-ai/flux/schnell` | — | | | | | | — | — |
69
69
  | `digitalocean/fal-ai/stable-audio-25/text-to-audio` | — | | | | | | — | — |
70
70
  | `digitalocean/gemma-4-31B-it` | 256K | | | | | | $0.18 | $0.50 |
71
- | `digitalocean/glm-5` | 64K | | | | | | $0.75 | $2 |
72
- | `digitalocean/glm-5.1` | 164K | | | | | | $0.97 | $4 |
71
+ | `digitalocean/glm-5` | 64K | | | | | | $1 | $3 |
72
+ | `digitalocean/glm-5.1` | 164K | | | | | | $1 | $4 |
73
73
  | `digitalocean/glm-5.2` | 262K | | | | | | $0.70 | $2 |
74
74
  | `digitalocean/gte-large-en-v1.5` | 8K | | | | | | $0.09 | — |
75
- | `digitalocean/kimi-k2.5` | 262K | | | | | | $0.38 | $2 |
76
- | `digitalocean/kimi-k2.6` | 262K | | | | | | $0.76 | $3 |
75
+ | `digitalocean/kimi-k2.5` | 262K | | | | | | $0.50 | $3 |
76
+ | `digitalocean/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
77
77
  | `digitalocean/kimi-k3` | 1.0M | | | | | | $3 | $14 |
78
78
  | `digitalocean/llama-4-maverick` | 128K | | | | | | $0.20 | $0.70 |
79
79
  | `digitalocean/llama3-8b-instruct` | 131K | | | | | | $0.20 | $0.20 |
80
80
  | `digitalocean/llama3.3-70b-instruct` | 128K | | | | | | $0.65 | $0.65 |
81
81
  | `digitalocean/mimo-v2.5-pro` | 262K | | | | | | $0.40 | $2 |
82
- | `digitalocean/minimax-m2.5` | 66K | | | | | | $0.23 | $0.90 |
82
+ | `digitalocean/minimax-m2.5` | 66K | | | | | | $0.30 | $1 |
83
83
  | `digitalocean/ministral-3-8b-instruct-2512` | 262K | | | | | | — | — |
84
84
  | `digitalocean/mistral-3-14B` | 262K | | | | | | $0.20 | $0.20 |
85
85
  | `digitalocean/mistral-7b-instruct-v0.3` | 33K | | | | | | — | — |
@@ -88,7 +88,7 @@ for await (const chunk of stream) {
88
88
  | `digitalocean/nemotron-3-nano-omni` | 66K | | | | | | $0.50 | $0.90 |
89
89
  | `digitalocean/nemotron-3-ultra-550b` | 131K | | | | | | $0.90 | $2 |
90
90
  | `digitalocean/nemotron-nano-12b-v2-vl` | 128K | | | | | | $0.20 | $0.60 |
91
- | `digitalocean/nvidia-nemotron-3-super-120b` | 1.0M | | | | | | $0.17 | $0.36 |
91
+ | `digitalocean/nvidia-nemotron-3-super-120b` | 1.0M | | | | | | $0.30 | $0.65 |
92
92
  | `digitalocean/openai-gpt-4.1` | 1.0M | | | | | | $2 | $8 |
93
93
  | `digitalocean/openai-gpt-4o` | 128K | | | | | | $3 | $10 |
94
94
  | `digitalocean/openai-gpt-4o-mini` | 128K | | | | | | $0.15 | $0.60 |
@@ -119,7 +119,7 @@ for await (const chunk of stream) {
119
119
  | `digitalocean/qwen3-coder-flash` | 262K | | | | | | $0.45 | $2 |
120
120
  | `digitalocean/qwen3-embedding-0.6b` | 8K | | | | | | $0.04 | — |
121
121
  | `digitalocean/qwen3-tts-voicedesign` | 33K | | | | | | — | — |
122
- | `digitalocean/qwen3.5-397b-a17b` | 131K | | | | | | $0.30 | $2 |
122
+ | `digitalocean/qwen3.5-397b-a17b` | 131K | | | | | | $0.55 | $4 |
123
123
  | `digitalocean/qwen3.8-max` | 262K | | | | | | $2 | $6 |
124
124
  | `digitalocean/stable-diffusion-3.5-large` | 256 | | | | | | $0.08 | — |
125
125
  | `digitalocean/wan2-2-t2v-a14b` | 100 | | | | | | $0.60 | — |
@@ -97,8 +97,8 @@ for await (const chunk of stream) {
97
97
  | `edenai/deepinfra/zai-org/GLM-4.7-Flash` | 203K | | | | | | $0.06 | $0.40 |
98
98
  | `edenai/deepseek/deepseek-chat` | 131K | | | | | | $0.28 | $0.42 |
99
99
  | `edenai/deepseek/deepseek-reasoner` | 131K | | | | | | $0.28 | $0.42 |
100
- | `edenai/deepseek/deepseek-v4-flash` | 1.0M | | | | | | $0.22 | $0.66 |
101
- | `edenai/deepseek/deepseek-v4-pro` | 1.0M | | | | | | $0.66 | $2 |
100
+ | `edenai/deepseek/deepseek-v4-flash` | 1.0M | | | | | | $0.44 | $1 |
101
+ | `edenai/deepseek/deepseek-v4-pro` | 1.0M | | | | | | $1 | $4 |
102
102
  | `edenai/fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.14 | $0.28 |
103
103
  | `edenai/fireworks_ai/accounts/fireworks/models/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $4 |
104
104
  | `edenai/fireworks_ai/accounts/fireworks/models/gpt-oss-120b` | 131K | | | | | | $0.15 | $0.60 |
@@ -139,7 +139,7 @@ for await (const chunk of stream) {
139
139
  | `edenai/minimax/MiniMax-M2.5` | 205K | | | | | | $0.30 | $1 |
140
140
  | `edenai/minimax/MiniMax-M2.7` | 205K | | | | | | $0.30 | $1 |
141
141
  | `edenai/minimax/MiniMax-M3` | 524K | | | | | | $0.30 | $1 |
142
- | `edenai/mistral/codestral-latest` | 256K | | | | | | $1 | $3 |
142
+ | `edenai/mistral/codestral-latest` | 256K | | | | | | $0.30 | $0.90 |
143
143
  | `edenai/mistral/devstral-2512` | 262K | | | | | | $0.40 | $2 |
144
144
  | `edenai/mistral/devstral-medium-latest` | 262K | | | | | | $0.40 | $2 |
145
145
  | `edenai/mistral/magistral-medium-latest` | 262K | | | | | | $2 | $5 |
@@ -149,7 +149,7 @@ for await (const chunk of stream) {
149
149
  | `edenai/mistral/mistral-medium-2604` | 262K | | | | | | $2 | $8 |
150
150
  | `edenai/mistral/mistral-medium-latest` | 262K | | | | | | $2 | $8 |
151
151
  | `edenai/mistral/mistral-small-2603` | 262K | | | | | | $0.15 | $0.60 |
152
- | `edenai/mistral/mistral-small-latest` | 262K | | | | | | $0.06 | $0.18 |
152
+ | `edenai/mistral/mistral-small-latest` | 262K | | | | | | $0.15 | $0.60 |
153
153
  | `edenai/moonshot/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
154
154
  | `edenai/moonshot/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
155
155
  | `edenai/moonshot/kimi-k3` | 1.0M | | | | | | $3 | $15 |
@@ -199,8 +199,8 @@ for await (const chunk of stream) {
199
199
  | `edenai/perplexityai/sonar-deep-research` | 128K | | | | | | $2 | $8 |
200
200
  | `edenai/perplexityai/sonar-pro` | 200K | | | | | | $3 | $15 |
201
201
  | `edenai/perplexityai/sonar-reasoning-pro` | 128K | | | | | | $2 | $8 |
202
- | `edenai/qwen/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.24 | $0.73 |
203
- | `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $0.73 | $2 |
202
+ | `edenai/qwen/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.40 | $1 |
203
+ | `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $4 |
204
204
  | `edenai/qwen/qwen-max` | 33K | | | | | | $2 | $6 |
205
205
  | `edenai/qwen/qwen-vl-max` | 131K | | | | | | $0.80 | $3 |
206
206
  | `edenai/qwen/qwen-vl-plus` | 131K | | | | | | $0.21 | $0.63 |
@@ -58,12 +58,12 @@ for await (const chunk of stream) {
58
58
  | `google/gemini-3.5-flash` | 1.0M | | | | | | $2 | $9 |
59
59
  | `google/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
60
60
  | `google/gemini-3.5-live-translate-preview` | 16K | | | | | | $4 | $21 |
61
- | `google/gemini-3.6-flash` | 1.0M | | | | | | $2 | $8 |
61
+ | `google/gemini-3.6-flash` | 1.0M | | | | | | $0.75 | $4 |
62
62
  | `google/gemini-3.7-flash` | 1.0M | | | | | | $0.75 | $4 |
63
63
  | `google/gemini-embedding-001` | 2K | | | | | | $0.15 | — |
64
64
  | `google/gemini-embedding-2` | 8K | | | | | | $0.20 | — |
65
- | `google/gemini-flash-latest` | 1.0M | | | | | | $2 | $9 |
66
- | `google/gemini-flash-lite-latest` | 1.0M | | | | | | $0.25 | $2 |
65
+ | `google/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
66
+ | `google/gemini-flash-lite-latest` | 1.0M | | | | | | $0.30 | $3 |
67
67
  | `google/gemini-omni-flash-preview` | 131K | | | | | | $2 | $18 |
68
68
  | `google/gemini-robotics-er-1.6-preview` | 131K | | | | | | $1 | $5 |
69
69
  | `google/gemma-4-26b-a4b-it` | 262K | | | | | | — | — |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![Kilo Gateway logo](https://models.dev/logos/kilo.svg)Kilo Gateway
4
4
 
5
- Access 361 Kilo Gateway models through Mastra's model router. Authentication is handled automatically using the `KILO_API_KEY` environment variable.
5
+ Access 362 Kilo Gateway models through Mastra's model router. Authentication is handled automatically using the `KILO_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [Kilo Gateway documentation](https://kilo.ai).
8
8
 
@@ -40,7 +40,7 @@ for await (const chunk of stream) {
40
40
  | `kilo/~anthropic/claude-haiku-latest` | 200K | | | | | | $1 | $5 |
41
41
  | `kilo/~anthropic/claude-opus-latest` | 1.0M | | | | | | $5 | $25 |
42
42
  | `kilo/~anthropic/claude-sonnet-latest` | 1.0M | | | | | | $2 | $10 |
43
- | `kilo/~deepseek/deepseek-v4-flash-latest` | 262K | | | | | | $0.07 | $0.14 |
43
+ | `kilo/~deepseek/deepseek-v4-flash-latest` | 1.0M | | | | | | $0.07 | $0.18 |
44
44
  | `kilo/~google/gemini-flash-latest` | 1.0M | | | | | | $0.38 | $2 |
45
45
  | `kilo/~google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
46
46
  | `kilo/~moonshotai/kimi-latest` | 975K | | | | | | $3 | $13 |
@@ -152,7 +152,8 @@ for await (const chunk of stream) {
152
152
  | `kilo/kwaipilot/kat-coder-air-v2.5` | 256K | | | | | | $0.15 | $0.60 |
153
153
  | `kilo/kwaipilot/kat-coder-pro-v2` | 256K | | | | | | $0.30 | $1 |
154
154
  | `kilo/kwaipilot/kat-coder-pro-v2.5` | 256K | | | | | | $0.74 | $3 |
155
- | `kilo/liquid/lfm-2.5-2.6b:free` | 128K | | | | | | — | — |
155
+ | `kilo/liquid/lfm-2.5-2.6b:free` | 66K | | | | | | — | — |
156
+ | `kilo/mancer/weaver` | 8K | | | | | | $0.50 | $0.75 |
156
157
  | `kilo/meituan/longcat-2.0` | 1.0M | | | | | | $0.75 | $3 |
157
158
  | `kilo/meta-llama/llama-3.1-70b-instruct` | 131K | | | | | | $0.40 | $0.40 |
158
159
  | `kilo/meta-llama/llama-3.1-8b-instruct` | 131K | | | | | | $0.02 | $0.04 |
@@ -172,7 +173,7 @@ for await (const chunk of stream) {
172
173
  | `kilo/minimax/minimax-m2` | 205K | | | | | | $0.30 | $1 |
173
174
  | `kilo/minimax/minimax-m2-her` | 66K | | | | | | $0.30 | $1 |
174
175
  | `kilo/minimax/minimax-m2.1` | 205K | | | | | | $0.30 | $1 |
175
- | `kilo/minimax/minimax-m2.5` | 66K | | | | | | $0.30 | $1 |
176
+ | `kilo/minimax/minimax-m2.5` | 198K | | | | | | $0.30 | $1 |
176
177
  | `kilo/minimax/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
177
178
  | `kilo/minimax/minimax-m3` | 524K | | | | | | $0.30 | $1 |
178
179
  | `kilo/mistralai/codestral-2508` | 256K | | | | | | $0.30 | $0.90 |
@@ -363,7 +364,7 @@ for await (const chunk of stream) {
363
364
  | `kilo/tencent/hunyuan-a13b-instruct` | 131K | | | | | | $0.14 | $0.57 |
364
365
  | `kilo/tencent/hy-mt2-1.8b` | 8K | | | | | | $0.04 | $0.18 |
365
366
  | `kilo/tencent/hy-mt2-30b-a3b` | 8K | | | | | | $0.07 | $0.29 |
366
- | `kilo/tencent/hy3` | 262K | | | | | | $0.13 | $0.53 |
367
+ | `kilo/tencent/hy3` | 262K | | | | | | $0.14 | $0.58 |
367
368
  | `kilo/tencent/hy3-preview` | 262K | | | | | | $0.18 | $0.60 |
368
369
  | `kilo/tencent/hy3:free` | 262K | | | | | | — | — |
369
370
  | `kilo/thedrummer/cydonia-24b-v4.1` | 131K | | | | | | $0.30 | $0.50 |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![LLM Gateway logo](https://models.dev/logos/llmgateway-providers.svg)LLM Gateway
4
4
 
5
- Access 368 LLM Gateway models through Mastra's model router. Authentication is handled automatically using the `LLMGATEWAY_API_KEY` environment variable.
5
+ Access 367 LLM Gateway models through Mastra's model router. Authentication is handled automatically using the `LLMGATEWAY_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [LLM Gateway documentation](https://llmgateway.io/docs).
8
8
 
@@ -205,7 +205,6 @@ for await (const chunk of stream) {
205
205
  | `llmgateway-providers/google-vertex/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
206
206
  | `llmgateway-providers/google-vertex/gemini-3.6-flash` | 1.0M | | | | | | $0.75 | $4 |
207
207
  | `llmgateway-providers/google-vertex/gemini-3.7-flash` | 1.0M | | | | | | $0.75 | $4 |
208
- | `llmgateway-providers/granite/qwen3.7-max` | 1.0M | | | | | | $3 | $8 |
209
208
  | `llmgateway-providers/groq/gpt-oss-120b` | 131K | | | | | | $0.15 | $0.75 |
210
209
  | `llmgateway-providers/groq/gpt-oss-20b` | 131K | | | | | | $0.10 | $0.50 |
211
210
  | `llmgateway-providers/iceberg/gemini-3-flash-preview` | 1.0M | | | | | | $0.50 | $3 |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![NanoGPT logo](https://models.dev/logos/nano-gpt.svg)NanoGPT
4
4
 
5
- Access 595 NanoGPT models through Mastra's model router. Authentication is handled automatically using the `NANO_GPT_API_KEY` environment variable.
5
+ Access 597 NanoGPT models through Mastra's model router. Authentication is handled automatically using the `NANO_GPT_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [NanoGPT documentation](https://docs.nano-gpt.com).
8
8
 
@@ -242,9 +242,9 @@ for await (const chunk of stream) {
242
242
  | `nano-gpt/google/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
243
243
  | `nano-gpt/google/gemini-flash-lite-latest` | 1.0M | | | | | | $0.30 | $3 |
244
244
  | `nano-gpt/google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
245
- | `nano-gpt/google/gemma-4-26b-a4b-it` | 262K | | | | | | $0.13 | $0.40 |
245
+ | `nano-gpt/google/gemma-4-26b-a4b-it` | 262K | | | | | | $0.08 | $0.33 |
246
246
  | `nano-gpt/google/gemma-4-26b-a4b-it:thinking` | 262K | | | | | | $0.13 | $0.40 |
247
- | `nano-gpt/google/gemma-4-31b-it` | 262K | | | | | | $0.10 | $0.35 |
247
+ | `nano-gpt/google/gemma-4-31b-it` | 262K | | | | | | $0.08 | $0.33 |
248
248
  | `nano-gpt/google/gemma-4-31b-it:thinking` | 262K | | | | | | $0.10 | $0.35 |
249
249
  | `nano-gpt/Gryphe/MythoMax-L2-13b` | 4K | | | | | | $0.10 | $0.10 |
250
250
  | `nano-gpt/hermes-high` | 1.0M | | | | | | $5 | $25 |
@@ -462,7 +462,9 @@ for await (const chunk of stream) {
462
462
  | `nano-gpt/qwen/qwen3.5-plus` | 984K | | | | | | $0.40 | $2 |
463
463
  | `nano-gpt/qwen/qwen3.5-plus-thinking` | 984K | | | | | | $0.40 | $2 |
464
464
  | `nano-gpt/qwen/Qwen3.6-35B-A3B` | 262K | | | | | | $0.11 | $0.80 |
465
+ | `nano-gpt/qwen/qwen3.6-35b-a3b-uncensored` | 66K | | | | | | $0.15 | $0.50 |
465
466
  | `nano-gpt/qwen/Qwen3.6-35B-A3B:thinking` | 262K | | | | | | $0.11 | $0.80 |
467
+ | `nano-gpt/qwen/qwen3.8-27b-uncensored` | 131K | | | | | | $0.18 | $0.50 |
466
468
  | `nano-gpt/qwen25-vl-72b-instruct` | 32K | | | | | | $0.70 | $0.70 |
467
469
  | `nano-gpt/qwen3-30b-a3b-instruct-2507` | 256K | | | | | | $0.20 | $0.50 |
468
470
  | `nano-gpt/qwen3-coder-30b-a3b-instruct` | 128K | | | | | | $0.10 | $0.40 |
@@ -80,8 +80,8 @@ for await (const chunk of stream) {
80
80
  | `ofox/google/gemini-3.1-pro-preview` | 1.0M | | | | | | $2 | $12 |
81
81
  | `ofox/google/gemini-3.5-flash` | 1.0M | | | | | | $2 | $9 |
82
82
  | `ofox/google/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
83
- | `ofox/google/gemini-3.6-flash` | 1.0M | | | | | | $2 | $8 |
84
- | `ofox/google/gemini-3.7-flash` | 1.0M | | | | | | $2 | $8 |
83
+ | `ofox/google/gemini-3.6-flash` | 1.0M | | | | | | $0.75 | $4 |
84
+ | `ofox/google/gemini-3.7-flash` | 1.0M | | | | | | $0.75 | $4 |
85
85
  | `ofox/minimax/m2-her` | 200K | | | | | | $0.30 | $1 |
86
86
  | `ofox/minimax/minimax-m2` | 197K | | | | | | $0.30 | $1 |
87
87
  | `ofox/minimax/minimax-m2.1` | 205K | | | | | | $0.30 | $1 |
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![OpenCode Go logo](https://models.dev/logos/opencode-go.svg)OpenCode Go
4
4
 
5
- Access 26 OpenCode Go models through Mastra's model router. Authentication is handled automatically using the `OPENCODE_API_KEY` environment variable.
5
+ Access 27 OpenCode Go models through Mastra's model router. Authentication is handled automatically using the `OPENCODE_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [OpenCode Go documentation](https://opencode.ai/docs/zen).
8
8
 
@@ -52,6 +52,7 @@ for await (const chunk of stream) {
52
52
  | `opencode-go/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
53
53
  | `opencode-go/minimax-m3` | 1.0M | | | | | | $0.30 | $1 |
54
54
  | `opencode-go/muse-spark-1.2-contributor` | 1.0M | | | | | | $0.10 | $0.20 |
55
+ | `opencode-go/ox-alpha-free` | 1.0M | | | | | | — | — |
55
56
  | `opencode-go/qwen3.6-plus` | 1.0M | | | | | | $0.50 | $3 |
56
57
  | `opencode-go/qwen3.7-max` | 1.0M | | | | | | $3 | $8 |
57
58
  | `opencode-go/qwen3.7-plus` | 1.0M | | | | | | $0.40 | $2 |
@@ -79,7 +79,7 @@ Visit the [Configuration reference](https://mastra.ai/reference/configuration) f
79
79
 
80
80
  **bundler** (`BundlerConfig`): Configuration for the asset bundler with options for externals, sourcemap, transpilePackages, and dynamicPackages. (Default: `{ externals: [], sourcemap: false, transpilePackages: [], dynamicPackages: [] }`)
81
81
 
82
- **scorers** (`Record<string, Scorer>`): Scorers for evaluating agent responses and workflow outputs (Default: `{}`)
82
+ **scorers** (`Record<string, Scorer>`): Scorers for evaluating agent responses and workflow outputs. Registration also makes a scorer resolvable by ID, which is required to persist its scores. See Score persistence (Default: `{}`)
83
83
 
84
84
  **processors** (`Record<string, Processor>`): Input/output processors for transforming agent inputs and outputs (Default: `{}`)
85
85
 
@@ -6,6 +6,8 @@ Quick Checks are zero-LLM, composable micro-scorers for common assertions. They
6
6
 
7
7
  Internally they're standard `createScorer()` instances, so they have the same observability, storage, and pipeline integration as any other scorer.
8
8
 
9
+ > **Register checks to persist their scores:** A check's score is only written to the scores store when the check is registered on the Mastra instance, alongside `storage`. Each check has a fixed id (`checks.includes()` is `check-includes`, `checks.calledTool()` is `check-called-tool`, and so on), so one registered instance covers every use of that check regardless of its arguments. See [Score persistence](https://mastra.ai/docs/evals/overview).
10
+
9
11
  ## Usage example
10
12
 
11
13
  ```typescript
@@ -203,6 +205,10 @@ const result = await runEvals({
203
205
  })
204
206
  ```
205
207
 
208
+ ## Multi-turn behavior
209
+
210
+ Checks read the accumulated `run.output`, so in a [multi-turn eval](https://mastra.ai/docs/evals/multi-turn) they see every turn. `checks.calledTool('get_weather', { times: 2 })` counts calls across the whole conversation, and `checks.includes()` searches all assistant text. Use per-turn `turns[].scorers` when a specific turn has to satisfy the check.
211
+
206
212
  ## Related
207
213
 
208
214
  - [Quick Checks overview](https://mastra.ai/docs/evals/quick-checks)
@@ -0,0 +1,101 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # Multi-turn Judge scorer
4
+
5
+ **Added in:** `@mastra/evals@1.9.0`
6
+
7
+ The `createMultiTurnJudgeScorer()` function creates an LLM-as-judge scorer that grades a whole conversation against a single plain-English criterion. It returns a **binary** score: `1` when the criterion is satisfied, otherwise `0`, and the `reason` echoes the criterion with the judge's explanation.
8
+
9
+ Unlike the other prebuilt LLM judges, which read a single assistant message, this scorer reads every assistant turn accumulated in `run.output`, so it works with the [multi-turn `inputs`](https://mastra.ai/docs/evals/multi-turn) form of [`runEvals()`](https://mastra.ai/reference/evals/run-evals).
10
+
11
+ ## Parameters
12
+
13
+ **model** (`MastraModelConfig`): The language model used to grade the conversation. A smaller, cheaper model is usually sufficient for grading.
14
+
15
+ **criterion** (`string`): What the conversation must satisfy, in plain English, e.g. "The agent gave forecasts for London and Paris, and weather-appropriate packing advice".
16
+
17
+ **options** (`MultiTurnJudgeScorerOptions`): Configuration options for the scorer
18
+
19
+ ## `.run()` returns
20
+
21
+ **score** (`number`): 1 when the judge considers the criterion satisfied, otherwise 0 (multiplied by scale).
22
+
23
+ **reason** (`string`): The verdict, the criterion it graded, and the judge's explanation of why the criterion is or is not satisfied.
24
+
25
+ ## Usage with multi-turn evals
26
+
27
+ Pass the scorer to `runEvals` alongside an `inputs` array. Every assistant turn is included in the prompt sent to the judge:
28
+
29
+ ```typescript
30
+ import { runEvals } from '@mastra/core/evals'
31
+ import { createMultiTurnJudgeScorer } from '@mastra/evals/scorers/prebuilt'
32
+ import { weatherAgent } from '../agents'
33
+
34
+ const result = await runEvals({
35
+ data: [
36
+ {
37
+ inputs: [
38
+ "I'm planning a trip to London, Paris, and Tokyo next week.",
39
+ "How's the weather looking in London?",
40
+ 'And Paris?',
41
+ 'Tokyo too?',
42
+ 'Should I pack an umbrella for the London leg?',
43
+ ],
44
+ },
45
+ ],
46
+ target: weatherAgent,
47
+ scorers: [
48
+ {
49
+ scorer: createMultiTurnJudgeScorer({
50
+ model: 'anthropic/claude-haiku-4-5',
51
+ criterion:
52
+ 'The agent provided weather forecasts for London, Paris, and Tokyo, and gave weather-appropriate packing or clothing advice.',
53
+ }),
54
+ threshold: 1,
55
+ },
56
+ ],
57
+ })
58
+ ```
59
+
60
+ Use `threshold: 1` to turn the verdict into a pass or fail: the score is binary, so any lower threshold always passes.
61
+
62
+ ## Persisting scores
63
+
64
+ Scores are only written to the scores store when a scorer with the same ID is registered on the Mastra instance, because persistence resolves scorer metadata through `Mastra.getScorerById()`. Only the ID is looked up, so the registered instance's `criterion` can be a placeholder:
65
+
66
+ ```typescript
67
+ import { Mastra } from '@mastra/core'
68
+ import { LibSQLStore } from '@mastra/libsql'
69
+ import { createMultiTurnJudgeScorer } from '@mastra/evals/scorers/prebuilt'
70
+
71
+ export const mastra = new Mastra({
72
+ agents: { weatherAgent },
73
+ storage: new LibSQLStore({ url: 'file:./mastra.db' }),
74
+ scorers: {
75
+ 'multi-turn-judge-scorer': createMultiTurnJudgeScorer({
76
+ model: 'anthropic/claude-haiku-4-5',
77
+ criterion: 'placeholder',
78
+ }),
79
+ },
80
+ })
81
+ ```
82
+
83
+ See [Score persistence](https://mastra.ai/docs/evals/overview) for the full requirement and the warning you get when a scorer isn't registered.
84
+
85
+ ## Scoring details
86
+
87
+ The scorer runs in two phases:
88
+
89
+ 1. **Grade**: Every assistant message in `run.output` is collected in order and rendered as a numbered transcript, then the judge decides whether the conversation as a whole satisfies the criterion. Assistant messages with no text (a turn that only carried tool calls, for example) are skipped.
90
+ 2. **Score**: A `satisfied` verdict scores `1` and anything else scores `0`, multiplied by `scale`.
91
+
92
+ The judge only sees what the assistant said. The user's turns and any tool results aren't included, so write criteria in terms of the agent's responses. This keeps the graded text limited to the agent's own output, but it also means a reply that only makes sense next to the question that prompted it ("Yes, bring one.") can't be judged on its own. For criteria that depend on the user's turns, grade each turn with `turns[].scorers` or write a [custom scorer](https://mastra.ai/docs/evals/multi-turn) that renders both roles.
93
+
94
+ The transcript is passed to the judge as untrusted data, fenced with explicit delimiters and an instruction to ignore anything inside it that reads as an instruction, so an agent response can't talk its way into a passing verdict.
95
+
96
+ ## Related
97
+
98
+ - [Multi-turn evals](https://mastra.ai/docs/evals/multi-turn)
99
+ - [`runEvals()`](https://mastra.ai/reference/evals/run-evals)
100
+ - [Rubric scorer](https://mastra.ai/reference/evals/rubric)
101
+ - [createScorer](https://mastra.ai/reference/evals/create-scorer)
@@ -152,6 +152,7 @@ The Reference section provides documentation of Mastra's API, including paramete
152
152
  - [Faithfulness](https://mastra.ai/reference/evals/faithfulness)
153
153
  - [Hallucination](https://mastra.ai/reference/evals/hallucination)
154
154
  - [Keyword Coverage Scorer](https://mastra.ai/reference/evals/keyword-coverage)
155
+ - [Multi-turn Judge Scorer](https://mastra.ai/reference/evals/multi-turn-judge)
155
156
  - [Noise Sensitivity Scorer](https://mastra.ai/reference/evals/noise-sensitivity)
156
157
  - [Prompt Alignment Scorer](https://mastra.ai/reference/evals/prompt-alignment)
157
158
  - [Rubric Scorer](https://mastra.ai/reference/evals/rubric)
@@ -31,9 +31,9 @@ const workspace = new Workspace({
31
31
 
32
32
  **name** (`string`): Human-readable name (Default: `workspace-{id}`)
33
33
 
34
- **filesystem** (`WorkspaceFilesystem | WorkspaceFilesystemResolver`): Filesystem provider instance, or a resolver function that receives requestContext and returns a filesystem per request. See multi-tenant filesystems.
34
+ **filesystem** (`WorkspaceFilesystem | WorkspaceFilesystemResolver`): Filesystem provider instance, or a resolver function that receives requestContext and returns a filesystem per request. See filesystems per user or thread.
35
35
 
36
- **sandbox** (`WorkspaceSandbox | WorkspaceSandboxResolver`): Sandbox provider instance, or a resolver function that receives requestContext and returns a sandbox per request. See multi-tenant sandboxes.
36
+ **sandbox** (`WorkspaceSandbox | WorkspaceSandboxResolver`): Sandbox provider instance, or a resolver function that receives requestContext and returns a sandbox per request. See sandboxes per user or thread.
37
37
 
38
38
  **instructions.dynamicSandbox** (`'placeholder' | 'resolve' | (({ requestContext }) => string)`): Controls how a resolver-backed sandbox contributes to workspace instructions. 'placeholder' (default) emits stable text without calling the resolver. 'resolve' calls the resolver and uses the sandbox's own instructions. A function returns custom text without resolving. Has no effect on a static sandbox. (Default: `'placeholder'`)
39
39
 
package/CHANGELOG.md CHANGED
@@ -1,5 +1,19 @@
1
1
  # @mastra/mcp-docs-server
2
2
 
3
+ ## 1.2.19-alpha.2
4
+
5
+ ### Patch Changes
6
+
7
+ - Updated dependencies [[`f95f468`](https://github.com/mastra-ai/mastra/commit/f95f468cf1e7c2b924a13826494f98b8f2ccd581)]:
8
+ - @mastra/core@1.61.1-alpha.1
9
+
10
+ ## 1.2.19-alpha.1
11
+
12
+ ### Patch Changes
13
+
14
+ - Updated dependencies [[`1e47b75`](https://github.com/mastra-ai/mastra/commit/1e47b7520cab4cfaa8daed52f17e2e6d14ff7539)]:
15
+ - @mastra/core@1.61.1-alpha.0
16
+
3
17
  ## 1.2.18
4
18
 
5
19
  ### Patch Changes
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mastra/mcp-docs-server",
3
- "version": "1.2.18",
3
+ "version": "1.2.19-alpha.3",
4
4
  "description": "MCP server for accessing Mastra.ai documentation, changelogs, and news.",
5
5
  "type": "module",
6
6
  "main": "dist/index.js",
@@ -28,7 +28,7 @@
28
28
  "jsdom": "^26.1.0",
29
29
  "local-pkg": "^1.1.2",
30
30
  "zod": "^4.4.3",
31
- "@mastra/core": "1.61.0",
31
+ "@mastra/core": "1.61.1-alpha.1",
32
32
  "@mastra/mcp": "^1.17.1"
33
33
  },
34
34
  "devDependencies": {
@@ -45,9 +45,9 @@
45
45
  "tsx": "^4.23.1",
46
46
  "typescript": "^6.0.3",
47
47
  "vitest": "4.1.10",
48
- "@internal/lint": "0.0.125",
49
48
  "@internal/types-builder": "0.0.100",
50
- "@mastra/core": "1.61.0"
49
+ "@mastra/core": "1.61.1-alpha.1",
50
+ "@internal/lint": "0.0.125"
51
51
  },
52
52
  "homepage": "https://mastra.ai",
53
53
  "repository": {