@mastra/mcp-docs-server 1.2.19-alpha.0 → 1.2.19-alpha.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.docs/docs/evals/built-in-scorers.md +1 -0
- package/.docs/docs/evals/multi-turn.md +84 -1
- package/.docs/docs/evals/overview.md +63 -1
- package/.docs/docs/sandbox/filesystem.md +8 -8
- package/.docs/docs/sandbox/overview.md +33 -66
- package/.docs/models/gateways/merge-gateway.md +2 -1
- package/.docs/models/gateways/netlify.md +1 -2
- package/.docs/models/gateways/openrouter.md +6 -2
- package/.docs/models/gateways/vercel.md +5 -1
- package/.docs/models/index.md +1 -1
- package/.docs/models/providers/crossmodel.md +56 -55
- package/.docs/models/providers/deepseek.md +8 -7
- package/.docs/models/providers/digitalocean.md +8 -8
- package/.docs/models/providers/edenai.md +8 -7
- package/.docs/models/providers/google.md +3 -3
- package/.docs/models/providers/hyper.md +3 -3
- package/.docs/models/providers/kilo.md +14 -9
- package/.docs/models/providers/llmgateway-providers.md +3 -2
- package/.docs/models/providers/nano-gpt.md +7 -3
- package/.docs/models/providers/ofox.md +113 -110
- package/.docs/models/providers/opencode-go.md +25 -23
- package/.docs/models/providers/opencode.md +1 -1
- package/.docs/models/providers/scaleway.md +2 -1
- package/.docs/reference/core/mastra-class.md +1 -1
- package/.docs/reference/evals/checks.md +6 -0
- package/.docs/reference/evals/multi-turn-judge.md +101 -0
- package/.docs/reference/index.md +1 -0
- package/.docs/reference/workspace/workspace-class.md +15 -3
- package/CHANGELOG.md +21 -0
- package/package.json +5 -5
|
@@ -22,6 +22,7 @@ These scorers evaluate how correct, truthful, and complete your agent's answers
|
|
|
22
22
|
- [`tool-call-accuracy`](https://mastra.ai/reference/evals/tool-call-accuracy): Evaluates whether the LLM selects the correct tool from available options (`0-1`, higher is better)
|
|
23
23
|
- [`trajectory-accuracy`](https://mastra.ai/reference/evals/trajectory-accuracy): Evaluates the expected action sequence for all span types. Covered spans include tool and model activity plus workflow steps (`0-1`, higher is better)
|
|
24
24
|
- [`prompt-alignment`](https://mastra.ai/reference/evals/prompt-alignment): Measures how well agent responses align with user prompt intent, requirements, completeness, and format (`0-1`, higher is better)
|
|
25
|
+
- [`multi-turn-judge`](https://mastra.ai/reference/evals/multi-turn-judge): Grades every assistant turn of a [multi-turn conversation](https://mastra.ai/docs/evals/multi-turn) against a plain-English criterion (`0` or `1`)
|
|
25
26
|
|
|
26
27
|
### Context quality
|
|
27
28
|
|
|
@@ -35,6 +35,8 @@ const result = await runEvals({
|
|
|
35
35
|
|
|
36
36
|
Each turn runs `agent.generate()` with the same thread ID, so the agent sees the full conversation history. Scorers receive the accumulated output messages from all turns.
|
|
37
37
|
|
|
38
|
+
> **Prebuilt LLM judges grade one turn:** Prebuilt LLM-judge scorers (rubric, answer relevancy, faithfulness, and the others listed in [Scorer compatibility](#scorer-compatibility)) read a single assistant message, so with `inputs` they grade only the last turn's response. For semantic grading across a conversation, use [`createMultiTurnJudgeScorer()`](https://mastra.ai/reference/evals/multi-turn-judge), which grades every assistant turn against one criterion, or per-turn [`turns[].scorers`](#per-turn-assertions-with-turns), where each turn grades its own output.
|
|
39
|
+
|
|
38
40
|
## Memory is required for cross-turn recall
|
|
39
41
|
|
|
40
42
|
Multi-turn recall depends on the agent having a **memory store configured**. The shared thread ID is what lets each turn see the earlier ones, but a thread only persists history when the agent has memory. If the agent has no memory configured, the turns still run sequentially and their outputs still accumulate for scoring, but the agent won't recall earlier turns (each input runs in isolation). `runEvals` logs a warning when you use `inputs` on an agent without memory.
|
|
@@ -70,6 +72,85 @@ These details matter when writing scorers for multi-turn items:
|
|
|
70
72
|
- **`run.output` is the accumulated output from every turn.** Output-based scorers: `checks.includes`, `checks.calledTool`, `checks.similarity`, and similar: evaluate the whole conversation. For example, `checks.calledTool('get_weather', { times: 2 })` counts calls across all turns.
|
|
71
73
|
- **`run.input` is only the first turn's input.** Scorers that compare input against output (faithfulness, answer relevancy, and other input-relative LLM scorers) only see the first user message, not the full conversation. Prefer output-based checks for multi-turn, or build scorers that read the accumulated `run.output` directly.
|
|
72
74
|
|
|
75
|
+
### Scorer compatibility
|
|
76
|
+
|
|
77
|
+
Accumulated output only helps if the scorer reads all of it:
|
|
78
|
+
|
|
79
|
+
| Scorer | What it sees with `inputs` |
|
|
80
|
+
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- |
|
|
81
|
+
| [Quick Checks](https://mastra.ai/docs/evals/quick-checks): `checks.calledTool`, `checks.includes`, `checks.similarity`, `checks.noToolErrors`, and the rest | The whole conversation. Tool-call checks count calls across every turn, and text checks search all assistant text |
|
|
82
|
+
| [Multi-turn Judge scorer](https://mastra.ai/reference/evals/multi-turn-judge) | Every assistant turn, graded together against one plain-English criterion |
|
|
83
|
+
| Prebuilt LLM-judge scorers: rubric, answer relevancy, answer similarity, faithfulness, hallucination, bias, toxicity, context precision, context recall, context relevance, noise sensitivity, prompt alignment, summarization | Only the last assistant message that carries text. Input-relative judges also see only the first turn's input |
|
|
84
|
+
| Trajectory scorers (`AgentScorerConfig.trajectory`) | The last turn's span when traces are available, otherwise tool calls from all accumulated messages |
|
|
85
|
+
| Custom scorers | Whatever they read from `run.output` |
|
|
86
|
+
|
|
87
|
+
### Grading a whole conversation
|
|
88
|
+
|
|
89
|
+
To judge every turn together, use [`createMultiTurnJudgeScorer()`](https://mastra.ai/reference/evals/multi-turn-judge). It builds a transcript of every assistant turn and asks a judge model whether the conversation satisfies one plain-English criterion, scoring `1` or `0`:
|
|
90
|
+
|
|
91
|
+
```typescript
|
|
92
|
+
import { runEvals } from '@mastra/core/evals'
|
|
93
|
+
import { createMultiTurnJudgeScorer } from '@mastra/evals/scorers/prebuilt'
|
|
94
|
+
import { weatherAgent } from '../agents'
|
|
95
|
+
|
|
96
|
+
const result = await runEvals({
|
|
97
|
+
data: [
|
|
98
|
+
{
|
|
99
|
+
inputs: [
|
|
100
|
+
"How's the weather in London?",
|
|
101
|
+
'And Paris?',
|
|
102
|
+
'Should I pack an umbrella for London?',
|
|
103
|
+
],
|
|
104
|
+
},
|
|
105
|
+
],
|
|
106
|
+
target: weatherAgent,
|
|
107
|
+
scorers: [
|
|
108
|
+
{
|
|
109
|
+
scorer: createMultiTurnJudgeScorer({
|
|
110
|
+
model: 'anthropic/claude-haiku-4-5',
|
|
111
|
+
criterion:
|
|
112
|
+
'The agent provided forecasts for London and Paris, and gave weather-appropriate packing advice.',
|
|
113
|
+
}),
|
|
114
|
+
threshold: 1,
|
|
115
|
+
},
|
|
116
|
+
],
|
|
117
|
+
})
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
To grade something the criterion can't express, write your own scorer that reads the assistant messages out of `run.output`. [`extractAgentResponseMessages()`](https://mastra.ai/reference/evals/scorer-utils) returns the text of each assistant message, in order:
|
|
121
|
+
|
|
122
|
+
```typescript
|
|
123
|
+
import { createScorer } from '@mastra/core/evals'
|
|
124
|
+
import { extractAgentResponseMessages } from '@mastra/evals/scorers/utils'
|
|
125
|
+
import { z } from 'zod'
|
|
126
|
+
|
|
127
|
+
export const conversationJudge = createScorer({
|
|
128
|
+
id: 'conversation-judge',
|
|
129
|
+
name: 'Conversation Judge',
|
|
130
|
+
description: 'Grades every assistant turn in a conversation against one criterion',
|
|
131
|
+
type: 'agent',
|
|
132
|
+
judge: {
|
|
133
|
+
model: 'anthropic/claude-haiku-4-5',
|
|
134
|
+
instructions: 'You grade multi-turn assistant transcripts against a single criterion.',
|
|
135
|
+
},
|
|
136
|
+
})
|
|
137
|
+
.analyze({
|
|
138
|
+
description: 'Judge the transcript as a whole',
|
|
139
|
+
outputSchema: z.object({ satisfied: z.boolean(), reason: z.string() }),
|
|
140
|
+
createPrompt: ({ run }) => {
|
|
141
|
+
const transcript = extractAgentResponseMessages(run.output)
|
|
142
|
+
.map((text, index) => `Turn ${index + 1}: ${text}`)
|
|
143
|
+
.join('\n\n')
|
|
144
|
+
|
|
145
|
+
return `Grade this conversation:\n\n${transcript}\n\nCriterion: the assistant keeps the forecast consistent across turns.`
|
|
146
|
+
},
|
|
147
|
+
})
|
|
148
|
+
.generateScore(({ results }) => (results.analyzeStepResult?.satisfied ? 1 : 0))
|
|
149
|
+
.generateReason(({ results }) => results.analyzeStepResult?.reason ?? '')
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
Pass it to `runEvals` like any other scorer. See [Custom scorers](https://mastra.ai/docs/evals/custom-scorers) for the full `createScorer()` pipeline.
|
|
153
|
+
|
|
73
154
|
## Per-turn assertions with `turns`
|
|
74
155
|
|
|
75
156
|
The `inputs` form scores the **accumulated** output as a whole, a single score over every turn's output. That can hide per-turn failures: an output-based check like `checks.includes('Brooklyn')` passes if _any_ turn mentions Brooklyn, even when the follow-up turn is broken.
|
|
@@ -175,4 +256,6 @@ await runEvals({
|
|
|
175
256
|
|
|
176
257
|
- [`runEvals()` reference](https://mastra.ai/reference/evals/run-evals): Full API for `runEvals` parameters and returns
|
|
177
258
|
- [Gates and verdicts](https://mastra.ai/docs/evals/gates-and-verdicts): Enforce hard requirements and quality thresholds
|
|
178
|
-
- [Quick Checks](https://mastra.ai/docs/evals/quick-checks): Zero-LLM composable micro-scorers
|
|
259
|
+
- [Quick Checks](https://mastra.ai/docs/evals/quick-checks): Zero-LLM composable micro-scorers
|
|
260
|
+
- [Scorer utilities](https://mastra.ai/reference/evals/scorer-utils): Helpers for reading messages out of a scorer run
|
|
261
|
+
- [Multi-turn Judge scorer](https://mastra.ai/reference/evals/multi-turn-judge): LLM judge that grades every assistant turn together
|
|
@@ -115,15 +115,77 @@ For the step-level `scorers` API, see the [Step class reference](https://mastra.
|
|
|
115
115
|
|
|
116
116
|
**Asynchronous execution**: Live evaluations run in the background without blocking your agent responses or workflow execution. Your AI systems remain responsive while live evaluations monitor them.
|
|
117
117
|
|
|
118
|
-
**Sampling control**: The `sampling.rate` parameter (0-1) controls what
|
|
118
|
+
**Sampling control**: The `sampling.rate` parameter (0-1) controls what fraction of outputs get scored:
|
|
119
119
|
|
|
120
120
|
- `1.0`: Score every single response (100%)
|
|
121
121
|
- `0.5`: Score half of all responses (50%)
|
|
122
122
|
- `0.1`: Score 10% of responses
|
|
123
123
|
- `0.0`: Disable scoring
|
|
124
124
|
|
|
125
|
+
Sampling is deterministic per trace: the decision is derived from the trace ID, not drawn at random. In practice:
|
|
126
|
+
|
|
127
|
+
- Scorers configured at the same rate score the same traces, so their scores are comparable on shared traffic.
|
|
128
|
+
- Re-running the same trace produces the same sampling decision, so sampled coverage is reproducible.
|
|
129
|
+
|
|
130
|
+
When a run has no trace (observability not configured), the decision is derived from the run ID instead. If [trace sampling](https://mastra.ai/docs/observability/tracing/overview) declined the trace, scorers skip that run entirely, so scores aren't created for traces that were never stored.
|
|
131
|
+
|
|
132
|
+
**Eligibility filters**: The optional `filter` parameter restricts which runs a scorer is eligible for, using a declarative predicate over the run's context. Filters are evaluated before sampling, so `sampling.rate` applies only to runs that match the filter:
|
|
133
|
+
|
|
134
|
+
```typescript
|
|
135
|
+
export const myAgent = new Agent({
|
|
136
|
+
// ...
|
|
137
|
+
scorers: {
|
|
138
|
+
relevancy: {
|
|
139
|
+
scorer: createAnswerRelevancyScorer({ model: 'openai/gpt-5-mini' }),
|
|
140
|
+
filter: {
|
|
141
|
+
op: 'eq',
|
|
142
|
+
left: { path: 'requestContext.plan' },
|
|
143
|
+
right: { literal: 'enterprise' },
|
|
144
|
+
},
|
|
145
|
+
sampling: { type: 'ratio', rate: 0.1 },
|
|
146
|
+
},
|
|
147
|
+
},
|
|
148
|
+
})
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
This scores 10% of enterprise-plan traffic and none of the rest. To score different segments at different rates, bind the same scorer twice with complementary filters.
|
|
152
|
+
|
|
153
|
+
Predicates can reference `requestContext.*`, `entity.*`, `entityType`, `source`, `threadId`, `resourceId`, and `projectId`. They support comparisons (`eq`, `ne`, `lt`, `lte`, `gt`, `gte`), membership (`in`, `notIn`), existence (`exists`, `notExists`), truthiness (`truthy`, `falsy`), and boolean composition (`and`, `or`, `not`). A filter that references an unknown root fails at agent construction rather than silently skipping scoring at runtime. Filters are plain JSON, so they're unaffected by durable agent state serialization.
|
|
154
|
+
|
|
155
|
+
Eligibility filters decide _whether a scorer runs_; to filter _which messages a scorer sees_ once it runs, use [`filterRun()`](https://mastra.ai/reference/evals/filter-run).
|
|
156
|
+
|
|
125
157
|
**Automatic storage**: All scoring results are automatically stored in the `mastra_scorers` table in your configured database, allowing you to analyze performance trends over time.
|
|
126
158
|
|
|
159
|
+
## Score persistence
|
|
160
|
+
|
|
161
|
+
Scores are persisted when the Mastra instance has `storage` configured, and when the scorer is registered on that instance. Registration is what lets Mastra resolve the scorer's metadata (name, description, type) through [`getScorerById()`](https://mastra.ai/reference/core/getScorerById) before writing the score.
|
|
162
|
+
|
|
163
|
+
Scorers you attach to an agent or a workflow step register themselves. Scorers you pass directly to [`runEvals()`](https://mastra.ai/reference/evals/run-evals), including [Quick Checks](https://mastra.ai/docs/evals/quick-checks), need the `scorers` option on the [`Mastra` class](https://mastra.ai/reference/core/mastra-class):
|
|
164
|
+
|
|
165
|
+
```typescript
|
|
166
|
+
import { Mastra } from '@mastra/core'
|
|
167
|
+
import { LibSQLStore } from '@mastra/libsql'
|
|
168
|
+
import { checks } from '@mastra/evals/checks'
|
|
169
|
+
import { createAnswerRelevancyScorer } from '@mastra/evals/scorers/prebuilt'
|
|
170
|
+
import { myAgent } from './agents/my-agent'
|
|
171
|
+
|
|
172
|
+
export const mastra = new Mastra({
|
|
173
|
+
agents: { myAgent },
|
|
174
|
+
storage: new LibSQLStore({ url: 'file:./mastra.db' }),
|
|
175
|
+
scorers: {
|
|
176
|
+
calledTool: checks.calledTool('get_weather'),
|
|
177
|
+
includes: checks.includes('Brooklyn'),
|
|
178
|
+
relevancy: createAnswerRelevancyScorer({ model: 'openai/gpt-5-mini' }),
|
|
179
|
+
},
|
|
180
|
+
})
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
The lookup only compares scorer IDs, so a registered instance can be configured differently from the one you evaluate with. Every `checks.calledTool()` instance has the id `check-called-tool`, so registering one covers all of them, whatever tool name you evaluate with.
|
|
184
|
+
|
|
185
|
+
The arguments still shape the `description` stored alongside each score, which comes from the registered instance, not the one you evaluate with. `checks.calledTool('')` is accepted and persists scores fine, but stores `Checks that "" was called`, so pass a representative value.
|
|
186
|
+
|
|
187
|
+
Skipping registration doesn't change scoring results, but each save fails with a `Scorer with id <id> not found` warning and the scores never reach the store.
|
|
188
|
+
|
|
127
189
|
## Trace evaluations
|
|
128
190
|
|
|
129
191
|
In addition to live evaluations, you can use scorers to evaluate historical traces from your agent interactions and workflows.
|
|
@@ -7,7 +7,7 @@ A filesystem gives an agent tools for reading, writing, listing, and [searching]
|
|
|
7
7
|
Configure files in two ways:
|
|
8
8
|
|
|
9
9
|
- [Direct filesystem access](#direct-filesystem-access) uses `filesystem` with one filesystem provider or a [`CompositeFilesystem`](#manual-composition) that you create yourself.
|
|
10
|
-
- [Mounts](#mounts) uses `mounts` to create a `CompositeFilesystem` from path-prefixed providers. When
|
|
10
|
+
- [Mounts](#mounts) uses `mounts` to create a `CompositeFilesystem` from path-prefixed providers. When a static sandbox and filesystem provider support mounting, Mastra automatically mounts the provider at its configured path. Remote sandbox mounts typically use Filesystem in Userspace (FUSE).
|
|
11
11
|
|
|
12
12
|
Configure either `filesystem` or `mounts`, not both. Configuring both throws a `WorkspaceError` with the code `INVALID_CONFIG`.
|
|
13
13
|
|
|
@@ -71,7 +71,7 @@ See [Search](https://mastra.ai/docs/sandbox/search) to index the files for keywo
|
|
|
71
71
|
|
|
72
72
|
## Mounts
|
|
73
73
|
|
|
74
|
-
Use `mounts` when programs inside a sandbox need to access persistent files by path. The agent still receives file tools, while command tools can run commands such as `ls`, `cat`, or `python` against the same files.
|
|
74
|
+
Use `mounts` when programs inside a sandbox need to access persistent files by path. Mastra creates the mount automatically when the sandbox and filesystem provider support it. The agent still receives file tools, while command tools can run commands such as `ls`, `cat`, or `python` against the same files.
|
|
75
75
|
|
|
76
76
|
For example, mount an S3 bucket at `/workspace` inside a Daytona sandbox:
|
|
77
77
|
|
|
@@ -180,7 +180,7 @@ Use manual composition when another part of your application needs the composite
|
|
|
180
180
|
|
|
181
181
|
## Mount availability
|
|
182
182
|
|
|
183
|
-
File tools and composite routing work without
|
|
183
|
+
File tools and composite routing work without a sandbox mount. When a remote sandbox and filesystem provider support mounting, `mounts` automatically uses FUSE to make the files visible to commands. `LocalSandbox` uses symlinks instead.
|
|
184
184
|
|
|
185
185
|
Built-in sandbox mounting currently includes:
|
|
186
186
|
|
|
@@ -195,27 +195,27 @@ Remote mounts may require `s3fs`, `gcsfuse`, or `blobfuse2` inside the sandbox.
|
|
|
195
195
|
|
|
196
196
|
If a sandbox mount is unavailable or fails, the workspace remains usable. File tools continue to access the provider through its SDK, but commands can't see that path. Mastra describes these providers to the agent as available through file tools only.
|
|
197
197
|
|
|
198
|
-
##
|
|
198
|
+
## Filesystems per user or thread
|
|
199
199
|
|
|
200
200
|
The `filesystem` option accepts a resolver when storage should vary by request, user, role, or tenant:
|
|
201
201
|
|
|
202
202
|
```typescript
|
|
203
203
|
const workspace = new Workspace({
|
|
204
204
|
filesystem: ({ requestContext }) => {
|
|
205
|
-
const
|
|
205
|
+
const userId = requestContext.get('user-id') as string
|
|
206
206
|
|
|
207
207
|
return new S3Filesystem({
|
|
208
208
|
bucket: process.env.S3_BUCKET!,
|
|
209
209
|
region: process.env.S3_REGION!,
|
|
210
|
-
prefix: `
|
|
210
|
+
prefix: `users/${userId}`,
|
|
211
211
|
})
|
|
212
212
|
},
|
|
213
213
|
})
|
|
214
214
|
```
|
|
215
215
|
|
|
216
|
-
Each
|
|
216
|
+
Each user gets a separate filesystem view, and file tools resolve the provider from the request context automatically.
|
|
217
217
|
|
|
218
|
-
`mounts` doesn't accept a resolver and can't be combined with a sandbox resolver. When each user or thread needs separate storage inside a separate sandbox, create and mount the provider inside the sandbox resolver. See [
|
|
218
|
+
`mounts` doesn't accept a resolver and can't be combined with a sandbox resolver. When each user or thread needs separate storage inside a separate sandbox, create and mount the provider inside the sandbox resolver. See [Sandboxes per user or thread](https://mastra.ai/docs/sandbox/overview).
|
|
219
219
|
|
|
220
220
|
## Policies and containment
|
|
221
221
|
|
|
@@ -2,27 +2,27 @@
|
|
|
2
2
|
|
|
3
3
|
# Sandboxes and filesystems
|
|
4
4
|
|
|
5
|
-
A sandbox gives your agent an environment where it can run commands, execute code, install dependencies, and manage processes.
|
|
5
|
+
A sandbox gives your agent an isolated environment where it can run commands, execute code, install dependencies, and manage processes. This lets agents perform work that would be risky, resource-intensive, or impractical to run directly inside your application.
|
|
6
6
|
|
|
7
|
-
[
|
|
7
|
+
Sandboxes are often temporary, so files created inside them may disappear when the environment stops. A [filesystem](https://mastra.ai/docs/sandbox/filesystem) gives the agent a place to read, write, and [search](https://mastra.ai/docs/sandbox/search) files that can outlive the sandbox. You can use one to keep outputs between runs, seed a new sandbox with existing files, or give the agent documents it can search while working. Filesystems also work without a sandbox, for example when an agent only needs a knowledge base or access to files in a service such as Google Drive.
|
|
8
8
|
|
|
9
9
|
## When to use sandboxes
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
Use a sandbox when an agent needs to:
|
|
12
12
|
|
|
13
|
-
-
|
|
14
|
-
-
|
|
15
|
-
-
|
|
13
|
+
- Clone repositories and run shell, Git, build, or test workflows in a separate environment.
|
|
14
|
+
- Process PDFs, presentations, or other files with specialized libraries and produce new artifacts.
|
|
15
|
+
- Run long-lived or parallel work in separate environments with isolated files and processes.
|
|
16
16
|
|
|
17
17
|
If an agent only needs to read, write, or search files, configure [direct filesystem access](https://mastra.ai/docs/sandbox/filesystem).
|
|
18
18
|
|
|
19
19
|
## Quickstart
|
|
20
20
|
|
|
21
|
-
Give an agent a local sandbox
|
|
21
|
+
Give an agent a local sandbox:
|
|
22
22
|
|
|
23
23
|
```typescript
|
|
24
24
|
import { Agent } from '@mastra/core/agent'
|
|
25
|
-
import {
|
|
25
|
+
import { LocalSandbox, Workspace } from '@mastra/core/workspace'
|
|
26
26
|
|
|
27
27
|
export const codingAgent = new Agent({
|
|
28
28
|
id: 'coding-agent',
|
|
@@ -33,20 +33,15 @@ export const codingAgent = new Agent({
|
|
|
33
33
|
sandbox: new LocalSandbox({
|
|
34
34
|
workingDirectory: './workspace',
|
|
35
35
|
}),
|
|
36
|
-
filesystem: new LocalFilesystem({
|
|
37
|
-
basePath: './workspace',
|
|
38
|
-
}),
|
|
39
36
|
}),
|
|
40
37
|
})
|
|
41
38
|
|
|
42
39
|
await codingAgent.generate('List the files in the sandbox directory')
|
|
43
40
|
```
|
|
44
41
|
|
|
45
|
-
> **Warning:** `LocalSandbox`
|
|
46
|
-
|
|
47
|
-
To use a different sandbox for each user, tenant, or thread, configure a [sandbox resolver](#multi-tenant-sandboxes) instead.
|
|
42
|
+
> **Warning:** `LocalSandbox` runs commands on the application host by default and isn't isolated or secure. Enable [native isolation](#localsandbox), or use a remote or container sandbox when running untrusted code.
|
|
48
43
|
|
|
49
|
-
|
|
44
|
+
A static sandbox is shared across every request and memory thread that uses the agent. Use a [resolver](#sandboxes-per-user-or-thread) when each user or thread needs a separate environment. See [Lifecycle and persistence](#lifecycle-and-persistence) for sharing and cleanup details.
|
|
50
45
|
|
|
51
46
|
## Using the sandbox
|
|
52
47
|
|
|
@@ -74,7 +69,7 @@ const workspace = new Workspace({
|
|
|
74
69
|
})
|
|
75
70
|
```
|
|
76
71
|
|
|
77
|
-
|
|
72
|
+
Set `{ enabled: false }` on one tool to remove it, or set the top-level `enabled: false` to disable generated tools by default. A per-tool `{ enabled: true }` overrides that global setting. The top-level `requireApproval` policy applies to every generated tool unless a per-tool entry overrides it.
|
|
78
73
|
|
|
79
74
|
See the [sandbox tools reference](https://mastra.ai/reference/workspace/workspace-class) for all generated tools and the [tool configuration reference](https://mastra.ai/reference/workspace/workspace-class) for approvals, output limits, and hooks.
|
|
80
75
|
|
|
@@ -109,17 +104,12 @@ export const startServerTool = createTool({
|
|
|
109
104
|
})
|
|
110
105
|
```
|
|
111
106
|
|
|
112
|
-
##
|
|
107
|
+
## Sandboxes
|
|
113
108
|
|
|
114
109
|
### `LocalSandbox`
|
|
115
110
|
|
|
116
111
|
[`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) executes commands on the same machine as your Mastra application. By default, commands run directly on the host with the permissions of the application process.
|
|
117
112
|
|
|
118
|
-
Enable native isolation to restrict filesystem and network access at the operating-system level:
|
|
119
|
-
|
|
120
|
-
- **macOS**: Seatbelt (`sandbox-exec`)
|
|
121
|
-
- **Linux**: Bubblewrap (`bwrap`)
|
|
122
|
-
|
|
123
113
|
```typescript
|
|
124
114
|
const sandbox = new LocalSandbox({
|
|
125
115
|
workingDirectory: './workspace',
|
|
@@ -131,11 +121,13 @@ const sandbox = new LocalSandbox({
|
|
|
131
121
|
})
|
|
132
122
|
```
|
|
133
123
|
|
|
134
|
-
|
|
124
|
+
Enable native isolation to restrict filesystem and network access at the operating-system level. On macOS, native isolation uses Seatbelt (`sandbox-exec`). On Linux, it uses Bubblewrap (`bwrap`). Use `LocalSandbox.detectIsolation()` to check whether Seatbelt or Bubblewrap is available on the current operating system.
|
|
135
125
|
|
|
136
|
-
|
|
126
|
+
The [`nativeSandbox` options](https://mastra.ai/reference/workspace/local-sandbox) control network access, read-only or writable paths, write access to the working directory, and system binaries. You can also provide a custom Seatbelt profile or Bubblewrap arguments. See [`LocalSandbox`](https://mastra.ai/reference/workspace/local-sandbox) for the full configuration.
|
|
137
127
|
|
|
138
|
-
|
|
128
|
+
### Remote sandboxes
|
|
129
|
+
|
|
130
|
+
Use a remote or container sandbox when commands need a stronger boundary from the host application or when workloads need to scale beyond the resources of your application server. Each sandbox has its own isolation, persistence, networking, and mount behavior:
|
|
139
131
|
|
|
140
132
|
- [AgentCore](https://mastra.ai/integrations/sandboxes/agentcore)
|
|
141
133
|
- [Apple Container](https://mastra.ai/integrations/sandboxes/apple-container)
|
|
@@ -149,47 +141,40 @@ Use a remote or container backend when commands need a stronger boundary from th
|
|
|
149
141
|
- [Railway](https://mastra.ai/integrations/sandboxes/railway)
|
|
150
142
|
- [Vercel](https://mastra.ai/integrations/sandboxes/vercel)
|
|
151
143
|
|
|
152
|
-
If Mastra doesn't support your
|
|
144
|
+
If Mastra doesn't support your sandbox provider, implement the [sandbox provider interface](https://mastra.ai/reference/workspace/sandbox) to add it.
|
|
153
145
|
|
|
154
146
|
## Filesystem
|
|
155
147
|
|
|
156
|
-
|
|
148
|
+
Every sandbox has a filesystem for commands. `LocalSandbox` uses the host filesystem, while remote sandboxes have isolated filesystems that are often temporary. For files that need to outlive a sandbox, use external storage.
|
|
149
|
+
|
|
150
|
+
Mastra filesystems connect agents and sandboxes to provider-backed storage such as Amazon S3, Google Cloud Storage, or Google Drive. Configuring one gives the agent tools to read, write, and search files. When a remote sandbox and provider support mounting, Mastra automatically mounts the storage through Filesystem in Userspace (FUSE). Commands use normal file paths while the provider stores changes outside the sandbox.
|
|
157
151
|
|
|
158
|
-
|
|
152
|
+
See [Filesystem](https://mastra.ai/docs/sandbox/filesystem) for providers, file tools, and mounts.
|
|
159
153
|
|
|
160
|
-
##
|
|
154
|
+
## Sandboxes per user or thread
|
|
161
155
|
|
|
162
156
|
Use a **resolver** when each user, tenant, or thread needs a separate sandbox. Set `sandboxCacheKey` to the identity that owns the sandbox so later requests reuse the same live environment.
|
|
163
157
|
|
|
164
158
|
This example creates and caches one Daytona sandbox per user:
|
|
165
159
|
|
|
166
160
|
```typescript
|
|
167
|
-
import type { RequestContext } from '@mastra/core/request-context'
|
|
168
161
|
import { Workspace } from '@mastra/core/workspace'
|
|
169
162
|
import { DaytonaSandbox } from '@mastra/daytona'
|
|
170
163
|
|
|
171
|
-
const getUserId = (requestContext: RequestContext) => {
|
|
172
|
-
const userId = requestContext.get('user-id')
|
|
173
|
-
if (typeof userId !== 'string' || !userId) {
|
|
174
|
-
throw new Error('A user ID is required to resolve this sandbox')
|
|
175
|
-
}
|
|
176
|
-
|
|
177
|
-
return userId
|
|
178
|
-
}
|
|
179
|
-
|
|
180
164
|
const workspace = new Workspace({
|
|
181
165
|
sandbox: async ({ requestContext }) => {
|
|
182
|
-
const
|
|
166
|
+
const userId = requestContext.get('user-id') as string
|
|
167
|
+
const sandbox = new DaytonaSandbox({ id: `user-${userId}` })
|
|
183
168
|
await sandbox.start()
|
|
184
169
|
return sandbox
|
|
185
170
|
},
|
|
186
|
-
sandboxCacheKey: ({ requestContext }) =>
|
|
171
|
+
sandboxCacheKey: ({ requestContext }) => requestContext.get('user-id') as string,
|
|
187
172
|
})
|
|
188
173
|
```
|
|
189
174
|
|
|
190
175
|
The first request from a user creates the sandbox. Later requests with the same user ID reuse it.
|
|
191
176
|
|
|
192
|
-
For a sandbox with persistent storage, create
|
|
177
|
+
For a sandbox with persistent storage, create a [filesystem](https://mastra.ai/docs/sandbox/filesystem) and mount it inside the resolver. This example creates one Daytona sandbox and one S3 storage prefix per memory thread, then mounts the storage at `/workspace`:
|
|
193
178
|
|
|
194
179
|
```typescript
|
|
195
180
|
import { MASTRA_THREAD_ID_KEY, type RequestContext } from '@mastra/core/request-context'
|
|
@@ -197,14 +182,8 @@ import { Workspace } from '@mastra/core/workspace'
|
|
|
197
182
|
import { DaytonaSandbox } from '@mastra/daytona'
|
|
198
183
|
import { S3Filesystem } from '@mastra/s3'
|
|
199
184
|
|
|
200
|
-
const getThreadId = (requestContext: RequestContext) =>
|
|
201
|
-
|
|
202
|
-
if (typeof threadId !== 'string' || !threadId) {
|
|
203
|
-
throw new Error('A memory thread is required to resolve this sandbox')
|
|
204
|
-
}
|
|
205
|
-
|
|
206
|
-
return threadId
|
|
207
|
-
}
|
|
185
|
+
const getThreadId = (requestContext: RequestContext) =>
|
|
186
|
+
requestContext.get(MASTRA_THREAD_ID_KEY) as string
|
|
208
187
|
|
|
209
188
|
const createThreadFilesystem = (threadId: string) =>
|
|
210
189
|
new S3Filesystem({
|
|
@@ -234,25 +213,13 @@ The first request in a thread runs the resolver and starts the sandbox. Later re
|
|
|
234
213
|
|
|
235
214
|
The example mounts the filesystem inside the resolver because static `mounts` can't be combined with a sandbox resolver. Generated tools resolve the filesystem and sandbox from the request context automatically.
|
|
236
215
|
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
You are responsible for every sandbox returned by a resolver. It must be ready to use when returned. Keep track of it and destroy it through your application lifecycle code when it's no longer needed. `workspace.destroy()` doesn't destroy resolver-returned sandboxes.
|
|
240
|
-
|
|
241
|
-
Resolvers are incompatible with `mounts` and [`lsp: true`](https://mastra.ai/docs/sandbox/lsp), because both require a static sandbox during configuration. Using a resolver with `mounts` throws an `INVALID_CONFIG` error. With `lsp: true`, Mastra disables LSP and logs a warning.
|
|
216
|
+
> **Warning:** Your application owns sandboxes returned by a resolver. Destroy them and call `workspace.clearSandboxCache(cacheKey)` when the user, thread, or session ends. `workspace.destroy()` doesn't destroy resolver-returned sandboxes.
|
|
242
217
|
|
|
243
218
|
### Tool availability
|
|
244
219
|
|
|
245
|
-
With a static sandbox, Mastra knows which capabilities the backend supports and only gives the agent the corresponding tools. Application code should also check that an optional capability is available before using it
|
|
246
|
-
|
|
247
|
-
```typescript
|
|
248
|
-
const processManager = workspace.sandbox?.processes
|
|
249
|
-
|
|
250
|
-
if (processManager) {
|
|
251
|
-
await processManager.spawn('pnpm dev')
|
|
252
|
-
}
|
|
253
|
-
```
|
|
220
|
+
With a static sandbox, Mastra knows which capabilities the backend supports and only gives the agent the corresponding tools. Application code should also check that an optional capability is available before using it.
|
|
254
221
|
|
|
255
|
-
With a [resolver-backed sandbox](#
|
|
222
|
+
With a [resolver-backed sandbox](#sandboxes-per-user-or-thread), the backend isn't known until a request runs, so Mastra initially makes all sandbox tools available. If the resolved backend doesn't support the tool the agent calls, the call fails with `SandboxFeatureNotSupportedError`.
|
|
256
223
|
|
|
257
224
|
## Background processes
|
|
258
225
|
|
|
@@ -292,7 +259,7 @@ Who shares a sandbox depends on where you configure it and whether you use a res
|
|
|
292
259
|
| Resource-scoped resolver | The resolver caches one sandbox for each resource ID. |
|
|
293
260
|
| Thread-scoped resolver | A memory thread keeps its sandbox across requests in that thread. |
|
|
294
261
|
|
|
295
|
-
A static sandbox isn't automatically scoped to the current resource or memory thread. For resource or thread scope, use a resolver and set `sandboxCacheKey` to the corresponding ID. See [
|
|
262
|
+
A static sandbox isn't automatically scoped to the current resource or memory thread. For resource or thread scope, use a resolver and set `sandboxCacheKey` to the corresponding ID. See [Sandboxes per user or thread](#sandboxes-per-user-or-thread).
|
|
296
263
|
|
|
297
264
|
### Start
|
|
298
265
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# Merge Gateway
|
|
4
4
|
|
|
5
|
-
Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
5
|
+
Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 175 models through Mastra's model router.
|
|
6
6
|
|
|
7
7
|
Learn more in the [Merge Gateway documentation](https://docs.merge.dev/merge-gateway).
|
|
8
8
|
|
|
@@ -66,6 +66,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
66
66
|
| `deepseek/deepseek-v4-flash` |
|
|
67
67
|
| `deepseek/deepseek-v4-flash-0731` |
|
|
68
68
|
| `deepseek/deepseek-v4-pro` |
|
|
69
|
+
| `deepseek/deepseek-v4-pro-0423` |
|
|
69
70
|
| `deepseek/deepseek-v4-pro-0813` |
|
|
70
71
|
| `google/gemini-2.5-computer-use-preview-10-2025` |
|
|
71
72
|
| `google/gemini-2.5-flash` |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# Netlify
|
|
4
4
|
|
|
5
|
-
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access
|
|
5
|
+
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 227 models through Mastra's model router.
|
|
6
6
|
|
|
7
7
|
Learn more in the [Netlify documentation](https://docs.netlify.com/build/ai-gateway/overview/).
|
|
8
8
|
|
|
@@ -120,7 +120,6 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
120
120
|
| `openrouter/bytedance-seed/seed-2.0-mini` |
|
|
121
121
|
| `openrouter/bytedance/ui-tars-1.5-7b` |
|
|
122
122
|
| `openrouter/cognitivecomputations/dolphin-mistral-24b-venice-edition` |
|
|
123
|
-
| `openrouter/deepcogito/cogito-v2.1-671b` |
|
|
124
123
|
| `openrouter/deepseek/deepseek-chat` |
|
|
125
124
|
| `openrouter/deepseek/deepseek-chat-v3-0324` |
|
|
126
125
|
| `openrouter/deepseek/deepseek-chat-v3.1` |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# OpenRouter
|
|
4
4
|
|
|
5
|
-
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
5
|
+
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 360 models through Mastra's model router.
|
|
6
6
|
|
|
7
7
|
Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
|
|
8
8
|
|
|
@@ -92,7 +92,6 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
92
92
|
| `cohere/command-r-plus-08-2024` |
|
|
93
93
|
| `cohere/command-r7b-12-2024` |
|
|
94
94
|
| `cohere/north-mini-code:free` |
|
|
95
|
-
| `deepcogito/cogito-v2.1-671b` |
|
|
96
95
|
| `deepseek/deepseek-chat` |
|
|
97
96
|
| `deepseek/deepseek-chat-v3-0324` |
|
|
98
97
|
| `deepseek/deepseek-chat-v3.1` |
|
|
@@ -104,6 +103,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
104
103
|
| `deepseek/deepseek-v3.2-exp` |
|
|
105
104
|
| `deepseek/deepseek-v4-flash` |
|
|
106
105
|
| `deepseek/deepseek-v4-flash-0731` |
|
|
106
|
+
| `deepseek/deepseek-v4-flash-vision-exp` |
|
|
107
107
|
| `deepseek/deepseek-v4-pro` |
|
|
108
108
|
| `deepseek/deepseek-v4-pro-0813` |
|
|
109
109
|
| `dots-studio/dots-3-note-preview:free` |
|
|
@@ -150,6 +150,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
150
150
|
| `kwaipilot/kat-coder-pro-v2` |
|
|
151
151
|
| `kwaipilot/kat-coder-pro-v2.5` |
|
|
152
152
|
| `liquid/lfm-2.5-2.6b:free` |
|
|
153
|
+
| `mancer/weaver` |
|
|
153
154
|
| `meituan/longcat-2.0` |
|
|
154
155
|
| `meta-llama/llama-3.1-70b-instruct` |
|
|
155
156
|
| `meta-llama/llama-3.1-8b-instruct` |
|
|
@@ -162,6 +163,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
162
163
|
| `meta/muse-glimmer-30b` |
|
|
163
164
|
| `meta/muse-spark-1.1` |
|
|
164
165
|
| `meta/muse-spark-1.2` |
|
|
166
|
+
| `meta/muse-spark-1.2-contributor` |
|
|
165
167
|
| `microsoft/phi-4` |
|
|
166
168
|
| `microsoft/wizardlm-2-8x22b` |
|
|
167
169
|
| `minimax/minimax-01` |
|
|
@@ -366,6 +368,8 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
366
368
|
| `thedrummer/unslopnemo-12b` |
|
|
367
369
|
| `thinkingmachines/inkling` |
|
|
368
370
|
| `thinkingmachines/inkling-small` |
|
|
371
|
+
| `thinkingmachines/inkling-small:free` |
|
|
372
|
+
| `thinkingmachines/inkling:free` |
|
|
369
373
|
| `undi95/remm-slerp-l2-13b` |
|
|
370
374
|
| `upstage/solar-pro-3` |
|
|
371
375
|
| `upstage/solar-pro4` |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# Vercel
|
|
4
4
|
|
|
5
|
-
Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
5
|
+
Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 351 models through Mastra's model router.
|
|
6
6
|
|
|
7
7
|
Learn more in the [Vercel documentation](https://ai-sdk.dev/providers/ai-sdk-providers).
|
|
8
8
|
|
|
@@ -133,6 +133,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
133
133
|
| `deepseek/deepseek-v3.2-thinking` |
|
|
134
134
|
| `deepseek/deepseek-v4-flash` |
|
|
135
135
|
| `deepseek/deepseek-v4-flash-0731` |
|
|
136
|
+
| `deepseek/deepseek-v4-flash-vision-exp` |
|
|
136
137
|
| `deepseek/deepseek-v4-pro` |
|
|
137
138
|
| `deepseek/deepseek-v4-pro-0813` |
|
|
138
139
|
| `fish-audio/s1` |
|
|
@@ -352,6 +353,9 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
352
353
|
| `spacexai/grok-voice-think-fast-2.0` |
|
|
353
354
|
| `stepfun/step-3.5-flash` |
|
|
354
355
|
| `stepfun/step-3.7-flash` |
|
|
356
|
+
| `tencent/hy-mt2-lite` |
|
|
357
|
+
| `tencent/hy-mt2-plus` |
|
|
358
|
+
| `tencent/hy-mt2-pro` |
|
|
355
359
|
| `tencent/hy3` |
|
|
356
360
|
| `thinkingmachines/inkling` |
|
|
357
361
|
| `thinkingmachines/inkling-small` |
|
package/.docs/models/index.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# Model Providers
|
|
4
4
|
|
|
5
|
-
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to
|
|
5
|
+
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 6767 models from 180 providers through a single API.
|
|
6
6
|
|
|
7
7
|
## Features
|
|
8
8
|
|