@mastra/mcp-docs-server 1.3.1 → 1.3.2-alpha.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.docs/docs/evals/evals-with-memory.md +3 -1
- package/.docs/docs/evals/experiments.md +30 -10
- package/.docs/docs/harness/durable-agents.md +21 -3
- package/.docs/docs/memory/observational-memory.md +25 -12
- package/.docs/docs/memory/overview.md +2 -2
- package/.docs/integrations/deploy/kubernetes.md +6 -4
- package/.docs/integrations/sandboxes/modal.md +40 -0
- package/.docs/models/gateways/netlify.md +5 -1
- package/.docs/models/gateways/openrouter.md +3 -1
- package/.docs/models/index.md +1 -1
- package/.docs/models/providers/cerebras.md +1 -1
- package/.docs/models/providers/cortecs.md +3 -3
- package/.docs/models/providers/kilo.md +12 -10
- package/.docs/models/providers/nano-gpt.md +5 -4
- package/.docs/models/providers/opencode.md +2 -2
- package/.docs/reference/agents/agent.md +15 -0
- package/.docs/reference/agents/durable-agent.md +1 -1
- package/.docs/reference/client-js/agents.md +2 -0
- package/.docs/reference/memory/cloneThread.md +2 -2
- package/.docs/reference/memory/observational-memory.md +4 -6
- package/.docs/reference/processors/agents-md-injector.md +2 -0
- package/.docs/reference/pubsub/base.md +11 -0
- package/.docs/reference/pubsub/redis-streams.md +8 -0
- package/.docs/reference/workspace/workspace-class.md +14 -1
- package/package.json +4 -4
|
@@ -90,7 +90,9 @@ const average = scores.reduce((a, b) => a + b, 0) / scores.length
|
|
|
90
90
|
|
|
91
91
|
## Dataset experiments with an inline task
|
|
92
92
|
|
|
93
|
-
`dataset.startExperiment(
|
|
93
|
+
`dataset.startExperiment()` accepts `targetType` and `targetId` for a registered target, or an inline `task` function. For a registered memory-enabled agent, the runner creates a fresh thread per item by default. See [memory-enabled experiment targets](https://mastra.ai/docs/evals/experiments) for resource isolation and request context options.
|
|
94
|
+
|
|
95
|
+
To choose each item's thread explicitly through `agent.generate()`, use an inline `task` and store `{ threadId, resourceId }` in each item's `metadata`. The scorer pipeline still runs as normal.
|
|
94
96
|
|
|
95
97
|
```typescript
|
|
96
98
|
import { randomUUID } from 'node:crypto'
|
|
@@ -53,7 +53,9 @@ You can also delete experiments with the [Core API](https://mastra.ai/reference/
|
|
|
53
53
|
|
|
54
54
|
## Experiment targets
|
|
55
55
|
|
|
56
|
-
|
|
56
|
+
Use `targetType` and `targetId` to select an agent, workflow, or scorer registered on your Mastra instance. The target processes each dataset item. The experiment's `scorers` evaluate the target's output afterward.
|
|
57
|
+
|
|
58
|
+
For custom execution, pass an inline `task` function instead. See [dataset experiments with an inline task](https://mastra.ai/docs/evals/evals-with-memory) for an example that controls each item's memory thread.
|
|
57
59
|
|
|
58
60
|
### Registered agent
|
|
59
61
|
|
|
@@ -68,25 +70,27 @@ const summary = await dataset.startExperiment({
|
|
|
68
70
|
})
|
|
69
71
|
```
|
|
70
72
|
|
|
71
|
-
Each item's `input` is passed directly to
|
|
73
|
+
Each item's `input` is passed directly to the agent. Use a supported message format, such as a prompt string, a message object, or an array of messages. See the [`agent.generate()` reference](https://mastra.ai/reference/agents/generate) for accepted inputs.
|
|
72
74
|
|
|
73
75
|
#### Memory-enabled agents
|
|
74
76
|
|
|
75
|
-
When the target agent has its own memory
|
|
77
|
+
When the target agent has its own memory, the experiment runner injects a fresh memory thread for each item and retry. Without a resource id in the request context, each thread gets an isolated, experiment-owned resource.
|
|
78
|
+
|
|
79
|
+
To run against an existing resource, set `MASTRA_RESOURCE_ID_KEY` in the experiment's `requestContext` or in an item's `requestContext`. Item values override experiment values. The runner still creates a fresh thread for each item and retry, but all threads using the same resource id share that resource's memory.
|
|
76
80
|
|
|
77
81
|
Injected threads are tagged so you can map them back to the run: thread metadata carries the `experimentId` and the dataset item's id as `experimentItemId`. No thread title is generated for them.
|
|
78
82
|
|
|
79
|
-
|
|
83
|
+
When items share a resource, resource-scoped memory features can read and write the same state:
|
|
80
84
|
|
|
81
|
-
- Resource-scoped working memory updates persist to the resource
|
|
85
|
+
- Resource-scoped working memory updates persist to the resource and can affect other items. Items run concurrently by default, so don't rely on an execution order.
|
|
82
86
|
- Resource-scoped semantic recall can surface the resource's prior conversations to the experiment, and experiment transcripts become recallable in that resource's later conversations.
|
|
83
87
|
|
|
84
88
|
Use this approach to evaluate an agent against a real user's accumulated context. To keep experiment runs from touching real user state, use a dedicated evaluation resource id instead.
|
|
85
89
|
|
|
86
90
|
Thread injection is skipped in the following cases:
|
|
87
91
|
|
|
88
|
-
- If the request context
|
|
89
|
-
- If the agent has no memory, or the request context
|
|
92
|
+
- If the request context sets `MASTRA_THREAD_ID_KEY`, the runner uses that thread as-is. Items and retries using the same thread id share a conversation.
|
|
93
|
+
- If the agent has no memory, or the request context explicitly sets the resource id to an empty string or `null`, the runner skips thread injection.
|
|
90
94
|
|
|
91
95
|
### Registered workflow
|
|
92
96
|
|
|
@@ -101,13 +105,24 @@ const summary = await dataset.startExperiment({
|
|
|
101
105
|
})
|
|
102
106
|
```
|
|
103
107
|
|
|
104
|
-
The workflow receives each item's `input` as
|
|
108
|
+
The workflow receives each item's `input` as `inputData` in `run.start()`. Structure the input to match the workflow's input schema.
|
|
105
109
|
|
|
106
110
|
### Registered scorer
|
|
107
111
|
|
|
108
|
-
|
|
112
|
+
Use a scorer as the target to evaluate the judge itself. The runner calls `scorer.run(item.input)`, so the dataset item's `input` must contain the full payload the scorer expects.
|
|
113
|
+
|
|
114
|
+
This example assumes a registered `accuracy` scorer that accepts string `input`, `output`, and `groundTruth` fields. Adapt those fields to your scorer's expected shape:
|
|
109
115
|
|
|
110
116
|
```typescript
|
|
117
|
+
await dataset.addItem({
|
|
118
|
+
input: {
|
|
119
|
+
input: 'What is the capital of France?',
|
|
120
|
+
output: 'Paris',
|
|
121
|
+
groundTruth: 'Paris',
|
|
122
|
+
},
|
|
123
|
+
groundTruth: { score: 1 },
|
|
124
|
+
})
|
|
125
|
+
|
|
111
126
|
const summary = await dataset.startExperiment({
|
|
112
127
|
name: 'judge-accuracy-eval',
|
|
113
128
|
targetType: 'scorer',
|
|
@@ -115,7 +130,12 @@ const summary = await dataset.startExperiment({
|
|
|
115
130
|
})
|
|
116
131
|
```
|
|
117
132
|
|
|
118
|
-
The
|
|
133
|
+
The two `groundTruth` fields serve different purposes:
|
|
134
|
+
|
|
135
|
+
- `input.groundTruth` is the reference answer passed to the judge: `'Paris'`.
|
|
136
|
+
- The top-level `groundTruth` is the expected result from the judge: `{ score: 1 }`. It isn't passed to the target scorer.
|
|
137
|
+
|
|
138
|
+
The target's output contains its `score` and `reason`. To evaluate that output against the top-level ground truth, add another scorer through the experiment's `scorers` option. Without an additional scorer, the experiment records the judge's output but doesn't score its agreement with the expected result.
|
|
119
139
|
|
|
120
140
|
## Scoring results
|
|
121
141
|
|
|
@@ -163,15 +163,29 @@ Visit the [`createInngestAgent()` reference](https://mastra.ai/reference/agents/
|
|
|
163
163
|
Durable agents support resumable streams through PubSub and an event cache. When a client disconnects mid-stream, the cache continues storing events. The same client can reconnect by calling `observe()` with the `runId`:
|
|
164
164
|
|
|
165
165
|
```typescript
|
|
166
|
-
const { output
|
|
166
|
+
const { output } = await durableResearcher.observe(runId)
|
|
167
167
|
|
|
168
168
|
for await (const chunk of output.fullStream) {
|
|
169
169
|
// Chunks from the run, including any missed while disconnected
|
|
170
170
|
}
|
|
171
|
+
```
|
|
171
172
|
|
|
172
|
-
|
|
173
|
+
Observing doesn't own the run. When the loop ends, or you leave it early with `break`, `return`, or an error, the observer unsubscribes and the run keeps going for anyone else watching it. To stop observing from outside the loop, for example when the client disconnects, call `detach()`:
|
|
174
|
+
|
|
175
|
+
```typescript
|
|
176
|
+
const { output, detach } = await durableResearcher.observe(runId)
|
|
177
|
+
|
|
178
|
+
request.signal.addEventListener('abort', detach, { once: true })
|
|
179
|
+
|
|
180
|
+
if (request.signal.aborted) detach()
|
|
181
|
+
|
|
182
|
+
for await (const chunk of output.fullStream) {
|
|
183
|
+
// Stops when the run finishes or the client disconnects
|
|
184
|
+
}
|
|
173
185
|
```
|
|
174
186
|
|
|
187
|
+
If you pass `idleTimeoutMs` and no chunks arrive for that long, Mastra checks `isAlive` when you provide it. If `isAlive` returns `true` or throws, the timer restarts and the stream stays open. Otherwise the stream ends with an error chunk and detaches the observer. The run itself is only cleaned up when `isAlive` returns `false`.
|
|
188
|
+
|
|
175
189
|
`createDurableAgent()` and `createEventedAgent()` use an in-memory cache by default, which means resumable streams work within a single process. For production, provide a persistent cache backend (e.g., Redis) so cached events survive process restarts:
|
|
176
190
|
|
|
177
191
|
```typescript
|
|
@@ -231,7 +245,11 @@ Visit [Background tasks](https://mastra.ai/docs/harness/background-tasks) for th
|
|
|
231
245
|
|
|
232
246
|
## Cleanup
|
|
233
247
|
|
|
234
|
-
Every `stream()` and `observe()` call returns a `cleanup` function
|
|
248
|
+
Every `stream()` and `observe()` call returns a `cleanup` function that tears the run down: it unsubscribes from PubSub, removes the run from the internal registry, and deletes the run's cached events so nobody can replay it afterward.
|
|
249
|
+
|
|
250
|
+
When you don't call `cleanup()` yourself, Mastra calls it `cleanupTimeoutMs` after the run finishes, errors, or is aborted. Detaching an observer or reaching a plain `idleTimeoutMs` doesn't start this timer. Set `cleanupTimeoutMs: 0` to turn the timer off.
|
|
251
|
+
|
|
252
|
+
Call `cleanup()` from the process that started the run. To stop watching a run from `observe()`, use `detach()` or leave the loop instead.
|
|
235
253
|
|
|
236
254
|
## Tool approval
|
|
237
255
|
|
|
@@ -615,29 +615,42 @@ const memory = new Memory({
|
|
|
615
615
|
|
|
616
616
|
Thread scope requires a valid `threadId` to be provided when calling the agent. If `threadId` is missing, Observational Memory throws an error. This prevents multiple threads from silently sharing a single observation record, which can cause database deadlocks.
|
|
617
617
|
|
|
618
|
-
### Resource scope (
|
|
618
|
+
### Resource scope (deprecated)
|
|
619
619
|
|
|
620
|
-
|
|
620
|
+
> **Deprecated:** `scope: 'resource'` is deprecated and will be removed in a future release. Mastra logs a warning the first time Observational Memory is created with resource scope. Remove the `scope` option to use the default thread scope, and use the [alternatives below](#cross-thread-continuity-without-resource-scope) for cross-conversation continuity.
|
|
621
|
+
|
|
622
|
+
In resource scope, observations are shared across all threads for a resource (typically a user). Resource scope works much worse than thread scope for prompt caching and for the agent's understanding of the conversation:
|
|
623
|
+
|
|
624
|
+
- **Prompt caching:** Observations from every thread share one record. When any thread is observed, the observation context changes for all other threads, which invalidates their cached prompt prefix.
|
|
625
|
+
- **Agent understanding:** Each thread sees a mix of observations from every thread for the resource. Every new thread starts with context from earlier threads, and the agent may continue work that another thread started but didn't finish.
|
|
626
|
+
- **Performance:** Unobserved messages across _all_ threads are processed together, which is slow for users with many existing threads.
|
|
627
|
+
|
|
628
|
+
A new knowledge and subconscious memory primitive for cross-thread memory is coming soon and will replace resource scope. Until then, use the alternatives below.
|
|
629
|
+
|
|
630
|
+
#### Cross-thread continuity without resource scope
|
|
631
|
+
|
|
632
|
+
Switching from resource scope to thread scope doesn't migrate resource-scoped observations. Each thread uses its existing thread-scoped record if it has one, or starts a new, empty one, and messages already observed under resource scope aren't observed again. Keep facts that must carry over in resource-scoped working memory.
|
|
633
|
+
|
|
634
|
+
Use thread scope with one or both of these features instead:
|
|
635
|
+
|
|
636
|
+
- [Retrieval mode](#retrieval-mode): The agent gets a `recall` tool that can list and browse other threads for the same resource, and optionally search them semantically. The agent looks up earlier conversations only when it needs them.
|
|
637
|
+
- [Resource-scoped working memory](https://mastra.ai/docs/memory/working-memory): Stores small, durable facts about the user that every thread can read.
|
|
621
638
|
|
|
622
639
|
```typescript
|
|
623
640
|
const memory = new Memory({
|
|
624
641
|
options: {
|
|
625
642
|
observationalMemory: {
|
|
626
643
|
model: 'google/gemini-2.5-flash',
|
|
644
|
+
retrieval: true,
|
|
645
|
+
},
|
|
646
|
+
workingMemory: {
|
|
647
|
+
enabled: true,
|
|
627
648
|
scope: 'resource',
|
|
628
649
|
},
|
|
629
650
|
},
|
|
630
651
|
})
|
|
631
652
|
```
|
|
632
653
|
|
|
633
|
-
Resource scope works, however it's marked as experimental for now until we prove task adherence/continuity across multiple ongoing simultaneous threads. As of today, you may need to tweak your system prompt to prevent one thread from continuing the work that another had already started (but hadn't finished).
|
|
634
|
-
|
|
635
|
-
This is because in resource scope, each thread is a perspective on _all_ threads for the resource.
|
|
636
|
-
|
|
637
|
-
For your use-case this may not be a problem, so your mileage may vary.
|
|
638
|
-
|
|
639
|
-
> **Warning:** In resource scope, unobserved messages across _all_ threads are processed together. For users with many existing threads, this can be slow. Use thread scope for existing apps.
|
|
640
|
-
|
|
641
654
|
## Token budgets
|
|
642
655
|
|
|
643
656
|
OM uses token thresholds to decide when to observe and reflect. See [token budget configuration](https://mastra.ai/reference/memory/observational-memory) for details.
|
|
@@ -786,7 +799,7 @@ const memory = new Memory({
|
|
|
786
799
|
|
|
787
800
|
Setting `bufferTokens: false` disables both observation and reflection async buffering. See [async buffering configuration](https://mastra.ai/reference/memory/observational-memory) for the full API.
|
|
788
801
|
|
|
789
|
-
> **Note:** Resource scope automatically disables async buffering.
|
|
802
|
+
> **Note:** Resource scope (deprecated) automatically disables async buffering.
|
|
790
803
|
|
|
791
804
|
## Observer Context Optimization
|
|
792
805
|
|
|
@@ -889,7 +902,7 @@ hooks: {
|
|
|
889
902
|
No manual migration needed. OM reads existing messages and observes them lazily when thresholds are exceeded.
|
|
890
903
|
|
|
891
904
|
- **Thread scope**: The first time a thread exceeds `observation.messageTokens`, the Observer processes the backlog.
|
|
892
|
-
- **Resource scope**: All unobserved messages across all threads for a resource are processed together. For users with many existing threads, this could take substantial time.
|
|
905
|
+
- **Resource scope (deprecated)**: All unobserved messages across all threads for a resource are processed together. For users with many existing threads, this could take substantial time.
|
|
893
906
|
|
|
894
907
|
## Comparing OM with other memory features
|
|
895
908
|
|
|
@@ -212,7 +212,7 @@ To go beyond this default isolation, you can share memory between agents by pass
|
|
|
212
212
|
|
|
213
213
|
When you call agents directly (outside the delegation flow), memory sharing is controlled by two identifiers: `resourceId` and `threadId`. Agents that use the same values read and write to the same data. This is useful when agents collaborate on a shared context, for example, a researcher that saves notes and a writer that reads them.
|
|
214
214
|
|
|
215
|
-
**Resource-scoped sharing** is the most common pattern. [Working memory](https://mastra.ai/docs/memory/working-memory) and [semantic recall](https://mastra.ai/docs/memory/semantic-recall) default to `scope: 'resource'`. If two agents share a `resourceId`, they share
|
|
215
|
+
**Resource-scoped sharing** is the most common pattern. [Working memory](https://mastra.ai/docs/memory/working-memory) and [semantic recall](https://mastra.ai/docs/memory/semantic-recall) default to `scope: 'resource'`. If two agents share a `resourceId`, they share working memory and embeddings, even across different threads:
|
|
216
216
|
|
|
217
217
|
```typescript
|
|
218
218
|
// Both agents share the same resource-scoped memory
|
|
@@ -225,7 +225,7 @@ await writer.generate('Write a summary from the research notes.', {
|
|
|
225
225
|
})
|
|
226
226
|
```
|
|
227
227
|
|
|
228
|
-
Because both calls use `resource: 'project-42'`, the writer can access the researcher's
|
|
228
|
+
Because both calls use `resource: 'project-42'`, the writer can access the researcher's working memory. Semantic embeddings are also shared through the resource. Each agent still has its own thread, so message histories stay separate.
|
|
229
229
|
|
|
230
230
|
**Thread-scoped sharing** gives tighter coupling. [Observational Memory](https://mastra.ai/docs/memory/observational-memory) uses `scope: 'thread'` by default. If two agents use the same `resource` and `thread`, they share the full message history. Each agent sees every message the other has written. This is useful when agents need to build on each other's exact outputs.
|
|
231
231
|
|
|
@@ -251,18 +251,20 @@ Register the durable agent with the `Mastra` instance above. Its run state now l
|
|
|
251
251
|
A durable agent publishes stream chunks to a per-run topic through the shared pub/sub. A client that disconnects reconnects by calling `observe()` with the run's ID, and replays the chunks it missed from the shared cache:
|
|
252
252
|
|
|
253
253
|
```typescript
|
|
254
|
-
const { output,
|
|
254
|
+
const { output, detach } = await durableAssistant.observe(runId)
|
|
255
|
+
|
|
256
|
+
request.signal.addEventListener('abort', detach, { once: true })
|
|
257
|
+
|
|
258
|
+
if (request.signal.aborted) detach()
|
|
255
259
|
|
|
256
260
|
for await (const chunk of output.fullStream) {
|
|
257
261
|
// Chunks from the run, including any missed while disconnected
|
|
258
262
|
}
|
|
259
|
-
|
|
260
|
-
cleanup()
|
|
261
263
|
```
|
|
262
264
|
|
|
263
265
|
Because the run state is in Postgres and the events are in Redis, the reconnecting request can be served by any pod, not only the one that started the run. See [Resumable streams](https://mastra.ai/docs/harness/durable-agents).
|
|
264
266
|
|
|
265
|
-
Multiple clients can observe the same run at once. Each `observe()` call receives the full stream, so a user watching from two devices, or two people following the same run, stay in sync.
|
|
267
|
+
Multiple clients can observe the same run at once. Each `observe()` call receives the full stream, so a user watching from two devices, or two people following the same run, stay in sync. When one client leaves the loop or calls `detach()`, only that client stops watching. Don't call `cleanup()` from an observer, because it clears the cached events the other clients need to reconnect.
|
|
266
268
|
|
|
267
269
|
## Tool approval across pods
|
|
268
270
|
|
|
@@ -100,6 +100,8 @@ Get your credentials from the [Modal dashboard](https://modal.com/settings/token
|
|
|
100
100
|
|
|
101
101
|
**onStart** (`function`): Lifecycle hook called after the sandbox reaches running status.
|
|
102
102
|
|
|
103
|
+
**volumes** (`Record<string, Volume>`): Modal Volumes to mount into the sandbox, keyed by absolute mount path. The working directory must not be at or inside a mount path.
|
|
104
|
+
|
|
103
105
|
**onStop** (`function`): Lifecycle hook called before the sandbox stops.
|
|
104
106
|
|
|
105
107
|
**onDestroy** (`function`): Lifecycle hook called before the sandbox is destroyed.
|
|
@@ -137,6 +139,44 @@ await sandbox._stop()
|
|
|
137
139
|
await sandbox._start()
|
|
138
140
|
```
|
|
139
141
|
|
|
142
|
+
## Volumes and persistent files
|
|
143
|
+
|
|
144
|
+
Mount [Modal Volumes](https://modal.com/docs/guide/volumes) with the `volumes` option, then use `ModalFilesystem` to expose a mounted path to the agent's workspace file tools. Files written under the mount persist across sandbox restarts.
|
|
145
|
+
|
|
146
|
+
```typescript
|
|
147
|
+
import { Workspace } from '@mastra/core/workspace'
|
|
148
|
+
import { ModalFilesystem, ModalSandbox } from '@mastra/modal'
|
|
149
|
+
import { ModalClient } from 'modal'
|
|
150
|
+
|
|
151
|
+
const modal = new ModalClient()
|
|
152
|
+
const privateVolume = await modal.volumes.fromName('agent-user-123', { createIfMissing: true })
|
|
153
|
+
const teamVolume = await modal.volumes.fromName('agent-team', { createIfMissing: true })
|
|
154
|
+
|
|
155
|
+
const sandbox = new ModalSandbox({
|
|
156
|
+
id: 'agent-sandbox',
|
|
157
|
+
workingDirectory: '/workspace',
|
|
158
|
+
volumes: {
|
|
159
|
+
'/mnt/agent': privateVolume,
|
|
160
|
+
'/mnt/team': teamVolume,
|
|
161
|
+
},
|
|
162
|
+
})
|
|
163
|
+
|
|
164
|
+
const workspace = new Workspace({
|
|
165
|
+
sandbox,
|
|
166
|
+
filesystem: new ModalFilesystem({ sandbox, basePath: '/mnt/agent' }),
|
|
167
|
+
})
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
`ModalFilesystem` runs file operations inside the sandbox and confines all paths to `basePath`. Set `readOnly: true` to reject writes.
|
|
171
|
+
|
|
172
|
+
Writes made by other sandboxes to the same Volume are only visible after a reload. Call `reloadVolumes()` on a long-lived sandbox to pick them up:
|
|
173
|
+
|
|
174
|
+
```typescript
|
|
175
|
+
await sandbox.reloadVolumes()
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
> **Note:** Keep the working directory outside every Volume mount (for example `/workspace`). Modal can't reload a Volume that is the current working directory, so `ModalSandbox` throws if `workingDirectory` is at or inside a mount path.
|
|
179
|
+
|
|
140
180
|
## Background processes
|
|
141
181
|
|
|
142
182
|
`ModalSandbox` includes a built-in process manager for spawning and managing background processes. Each `spawn()` call creates a new `ContainerProcess` via the Modal SDK's `Sandbox.exec()` API.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Netlify
|
|
6
6
|
|
|
7
|
-
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access
|
|
7
|
+
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 269 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Netlify documentation](https://docs.netlify.com/build/ai-gateway/overview/).
|
|
10
10
|
|
|
@@ -188,11 +188,15 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
188
188
|
| `openrouter/minimax/minimax-m2.5` |
|
|
189
189
|
| `openrouter/minimax/minimax-m2.7` |
|
|
190
190
|
| `openrouter/minimax/minimax-m3` |
|
|
191
|
+
| `openrouter/mistralai/devstral-2512` |
|
|
191
192
|
| `openrouter/mistralai/ministral-14b-2512` |
|
|
192
193
|
| `openrouter/mistralai/ministral-3b-2512` |
|
|
193
194
|
| `openrouter/mistralai/ministral-8b-2512` |
|
|
194
195
|
| `openrouter/mistralai/mistral-large-2407` |
|
|
196
|
+
| `openrouter/mistralai/mistral-large-2512` |
|
|
197
|
+
| `openrouter/mistralai/mistral-medium-3` |
|
|
195
198
|
| `openrouter/mistralai/mistral-medium-3-5` |
|
|
199
|
+
| `openrouter/mistralai/mistral-medium-3.1` |
|
|
196
200
|
| `openrouter/mistralai/mistral-nemo` |
|
|
197
201
|
| `openrouter/mistralai/mistral-saba` |
|
|
198
202
|
| `openrouter/mistralai/mistral-small-24b-instruct-2501` |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# OpenRouter
|
|
6
6
|
|
|
7
|
-
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
7
|
+
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 387 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
|
|
10
10
|
|
|
@@ -187,11 +187,13 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
187
187
|
| `minimax/minimax-m2.7` |
|
|
188
188
|
| `minimax/minimax-m3` |
|
|
189
189
|
| `mistralai/codestral-2508` |
|
|
190
|
+
| `mistralai/devstral-2512` |
|
|
190
191
|
| `mistralai/ministral-14b-2512` |
|
|
191
192
|
| `mistralai/ministral-3b-2512` |
|
|
192
193
|
| `mistralai/ministral-8b-2512` |
|
|
193
194
|
| `mistralai/mistral-large` |
|
|
194
195
|
| `mistralai/mistral-large-2407` |
|
|
196
|
+
| `mistralai/mistral-large-2512` |
|
|
195
197
|
| `mistralai/mistral-medium-3` |
|
|
196
198
|
| `mistralai/mistral-medium-3-5` |
|
|
197
199
|
| `mistralai/mistral-medium-3.1` |
|
package/.docs/models/index.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Model Providers
|
|
6
6
|
|
|
7
|
-
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to
|
|
7
|
+
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7626 models from 210 providers through a single API.
|
|
8
8
|
|
|
9
9
|
## Features
|
|
10
10
|
|
|
@@ -37,7 +37,7 @@ for await (const chunk of stream) {
|
|
|
37
37
|
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|
|
38
38
|
| ----------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
|
|
39
39
|
| `cerebras/gpt-oss-120b` | 131K | | | | | | $0.35 | $0.75 |
|
|
40
|
-
| `cerebras/qwen-3.8-27b` |
|
|
40
|
+
| `cerebras/qwen-3.8-27b` | 131K | | | | | | $0.99 | $1 |
|
|
41
41
|
|
|
42
42
|
Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
|
|
43
43
|
|
|
@@ -56,7 +56,7 @@ for await (const chunk of stream) {
|
|
|
56
56
|
| `cortecs/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.09 | $0.17 |
|
|
57
57
|
| `cortecs/deepseek-v4-pro` | 1.0M | | | | | | $2 | $3 |
|
|
58
58
|
| `cortecs/deepseek-v4-pro-0813` | 1.0M | | | | | | $2 | $4 |
|
|
59
|
-
| `cortecs/deepseek-v4.1-flash` | 1.0M | | | | | | $0.
|
|
59
|
+
| `cortecs/deepseek-v4.1-flash` | 1.0M | | | | | | $0.30 | $1 |
|
|
60
60
|
| `cortecs/devstral-2512` | 256K | | | | | | $0.48 | $2 |
|
|
61
61
|
| `cortecs/gemini-2.5-flash` | 1.0M | | | | | | $0.30 | $2 |
|
|
62
62
|
| `cortecs/gemini-2.5-pro` | 1.0M | | | | | | $1 | $10 |
|
|
@@ -107,7 +107,7 @@ for await (const chunk of stream) {
|
|
|
107
107
|
| `cortecs/minimax-m2.1` | 196K | | | | | | $0.36 | $1 |
|
|
108
108
|
| `cortecs/minimax-m2.5` | 196K | | | | | | $0.30 | $1 |
|
|
109
109
|
| `cortecs/minimax-m2.7` | 197K | | | | | | $0.67 | $3 |
|
|
110
|
-
| `cortecs/minimax-m3` | 1.0M | | | | | | $0.
|
|
110
|
+
| `cortecs/minimax-m3` | 1.0M | | | | | | $0.30 | $1 |
|
|
111
111
|
| `cortecs/ministral-14b-2512` | 256K | | | | | | $0.24 | $0.24 |
|
|
112
112
|
| `cortecs/ministral-3b-2512` | 256K | | | | | | $0.12 | $0.12 |
|
|
113
113
|
| `cortecs/ministral-8b-2512` | 256K | | | | | | $0.18 | $0.18 |
|
|
@@ -142,7 +142,7 @@ for await (const chunk of stream) {
|
|
|
142
142
|
| `cortecs/qwen3.6-27b` | 262K | | | | | | $0.45 | $3 |
|
|
143
143
|
| `cortecs/qwen3.6-35b-a3b` | 262K | | | | | | $0.17 | $0.56 |
|
|
144
144
|
| `cortecs/qwen3.8-2.4t-a95b` | 262K | | | | | | $3 | $6 |
|
|
145
|
-
| `cortecs/qwen3.8-27b` |
|
|
145
|
+
| `cortecs/qwen3.8-27b` | 1.0M | | | | | | $0.10 | $0.40 |
|
|
146
146
|
| `cortecs/qwen3.8-flash-next` | 262K | | | | | | $0.20 | $0.50 |
|
|
147
147
|
| `cortecs/qwen3guard-gen-0.6b` | 32K | | | | | | — | — |
|
|
148
148
|
| `cortecs/qwen3guard-gen-8b` | 32K | | | | | | — | — |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Kilo Gateway
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 394 Kilo Gateway models through Mastra's model router. Authentication is handled automatically using the `KILO_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Kilo Gateway documentation](https://kilo.ai).
|
|
10
10
|
|
|
@@ -42,12 +42,12 @@ for await (const chunk of stream) {
|
|
|
42
42
|
| `kilo/~anthropic/claude-haiku-latest` | 200K | | | | | | $1 | $5 |
|
|
43
43
|
| `kilo/~anthropic/claude-opus-latest` | 1.0M | | | | | | $4 | $20 |
|
|
44
44
|
| `kilo/~anthropic/claude-sonnet-latest` | 1.0M | | | | | | $2 | $10 |
|
|
45
|
-
| `kilo/~deepseek/deepseek-flash-latest` | 1.0M | | | | | | $0.04 | $
|
|
46
|
-
| `kilo/~deepseek/deepseek-pro-latest` | 1.0M | | | | | | $0.
|
|
45
|
+
| `kilo/~deepseek/deepseek-flash-latest` | 1.0M | | | | | | $0.04 | $0.49 |
|
|
46
|
+
| `kilo/~deepseek/deepseek-pro-latest` | 1.0M | | | | | | $0.25 | $3 |
|
|
47
47
|
| `kilo/~deepseek/deepseek-v4-flash-latest` | 1.0M | | | | | | $0.03 | $0.32 |
|
|
48
48
|
| `kilo/~google/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
|
|
49
49
|
| `kilo/~google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
|
|
50
|
-
| `kilo/~moonshotai/kimi-latest` | 1.0M | | | | | | $
|
|
50
|
+
| `kilo/~moonshotai/kimi-latest` | 1.0M | | | | | | $0.88 | $11 |
|
|
51
51
|
| `kilo/~openai/gpt-astra-latest` | 1.1M | | | | | | $10 | $50 |
|
|
52
52
|
| `kilo/~openai/gpt-luna-latest` | 1.1M | | | | | | $0.10 | $0.50 |
|
|
53
53
|
| `kilo/~openai/gpt-mini-latest` | 400K | | | | | | $0.75 | $5 |
|
|
@@ -190,20 +190,22 @@ for await (const chunk of stream) {
|
|
|
190
190
|
| `kilo/minimax/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
|
|
191
191
|
| `kilo/minimax/minimax-m3` | 524K | | | | | | $0.30 | $1 |
|
|
192
192
|
| `kilo/mistralai/codestral-2508` | 256K | | | | | | $0.30 | $0.90 |
|
|
193
|
+
| `kilo/mistralai/devstral-2512` | 262K | | | | | | $0.40 | $2 |
|
|
193
194
|
| `kilo/mistralai/ministral-14b-2512` | 262K | | | | | | $0.20 | $0.20 |
|
|
194
195
|
| `kilo/mistralai/ministral-3b-2512` | 131K | | | | | | $0.10 | $0.10 |
|
|
195
196
|
| `kilo/mistralai/ministral-8b-2512` | 262K | | | | | | $0.15 | $0.15 |
|
|
196
197
|
| `kilo/mistralai/mistral-large` | 128K | | | | | | $2 | $6 |
|
|
197
198
|
| `kilo/mistralai/mistral-large-2407` | 131K | | | | | | $2 | $6 |
|
|
199
|
+
| `kilo/mistralai/mistral-large-2512` | 262K | | | | | | $0.50 | $2 |
|
|
198
200
|
| `kilo/mistralai/mistral-medium-3` | 131K | | | | | | $0.40 | $2 |
|
|
199
201
|
| `kilo/mistralai/mistral-medium-3-5` | 262K | | | | | | $2 | $8 |
|
|
200
202
|
| `kilo/mistralai/mistral-medium-3.1` | 131K | | | | | | $0.40 | $2 |
|
|
201
|
-
| `kilo/mistralai/mistral-nemo` | 131K | | | | | | $0.
|
|
203
|
+
| `kilo/mistralai/mistral-nemo` | 131K | | | | | | $0.15 | $0.15 |
|
|
202
204
|
| `kilo/mistralai/mistral-saba` | 33K | | | | | | $0.20 | $0.60 |
|
|
203
205
|
| `kilo/mistralai/mistral-small-24b-instruct-2501` | 33K | | | | | | $0.05 | $0.08 |
|
|
204
206
|
| `kilo/mistralai/mistral-small-2603` | 262K | | | | | | $0.15 | $0.60 |
|
|
205
207
|
| `kilo/mistralai/mistral-small-3.1-24b-instruct` | 128K | | | | | | $0.35 | $0.56 |
|
|
206
|
-
| `kilo/mistralai/mistral-small-3.2-24b-instruct` | 256K | | | | | | $0.
|
|
208
|
+
| `kilo/mistralai/mistral-small-3.2-24b-instruct` | 256K | | | | | | $0.10 | $0.30 |
|
|
207
209
|
| `kilo/mistralai/mixtral-8x22b-instruct` | 66K | | | | | | $2 | $6 |
|
|
208
210
|
| `kilo/mistralai/voxtral-small-24b-2507` | 33K | | | | | | $0.10 | $0.30 |
|
|
209
211
|
| `kilo/moonshotai/kimi-k2` | 131K | | | | | | $0.57 | $2 |
|
|
@@ -334,7 +336,7 @@ for await (const chunk of stream) {
|
|
|
334
336
|
| `kilo/qwen/qwen3-next-80b-a3b-thinking` | 262K | | | | | | $0.15 | $1 |
|
|
335
337
|
| `kilo/qwen/qwen3-vl-235b-a22b-instruct` | 131K | | | | | | $0.26 | $1 |
|
|
336
338
|
| `kilo/qwen/qwen3-vl-235b-a22b-thinking` | 131K | | | | | | $0.40 | $4 |
|
|
337
|
-
| `kilo/qwen/qwen3-vl-30b-a3b-instruct` |
|
|
339
|
+
| `kilo/qwen/qwen3-vl-30b-a3b-instruct` | 262K | | | | | | $0.13 | $0.52 |
|
|
338
340
|
| `kilo/qwen/qwen3-vl-30b-a3b-thinking` | 131K | | | | | | $0.20 | $2 |
|
|
339
341
|
| `kilo/qwen/qwen3-vl-32b-instruct` | 131K | | | | | | $0.10 | $0.42 |
|
|
340
342
|
| `kilo/qwen/qwen3-vl-8b-instruct` | 131K | | | | | | $0.12 | $0.46 |
|
|
@@ -386,7 +388,7 @@ for await (const chunk of stream) {
|
|
|
386
388
|
| `kilo/tencent/hy-mt2-1.8b` | 8K | | | | | | $0.04 | $0.18 |
|
|
387
389
|
| `kilo/tencent/hy-mt2-30b-a3b` | 8K | | | | | | $0.07 | $0.29 |
|
|
388
390
|
| `kilo/tencent/hy-mt2-7b` | 8K | | | | | | $0.07 | $0.29 |
|
|
389
|
-
| `kilo/tencent/hy3` | 262K | | | | | | $0.
|
|
391
|
+
| `kilo/tencent/hy3` | 262K | | | | | | $0.13 | $0.53 |
|
|
390
392
|
| `kilo/tencent/hy3-preview` | 262K | | | | | | $0.18 | $0.60 |
|
|
391
393
|
| `kilo/tencent/hy4-preview` | 1.0M | | | | | | $0.83 | $3 |
|
|
392
394
|
| `kilo/thedrummer/cydonia-24b-v4.1` | 131K | | | | | | $0.30 | $0.50 |
|
|
@@ -418,11 +420,11 @@ for await (const chunk of stream) {
|
|
|
418
420
|
| `kilo/z-ai/glm-4.5v` | 66K | | | | | | $0.60 | $2 |
|
|
419
421
|
| `kilo/z-ai/glm-4.6` | 198K | | | | | | $0.43 | $2 |
|
|
420
422
|
| `kilo/z-ai/glm-4.6v` | 131K | | | | | | $0.30 | $0.90 |
|
|
421
|
-
| `kilo/z-ai/glm-4.7` | 203K | | | | | | $0.
|
|
423
|
+
| `kilo/z-ai/glm-4.7` | 203K | | | | | | $0.60 | $2 |
|
|
422
424
|
| `kilo/z-ai/glm-4.7-flash` | 131K | | | | | | $0.06 | $0.40 |
|
|
423
425
|
| `kilo/z-ai/glm-5` | 198K | | | | | | $0.60 | $2 |
|
|
424
426
|
| `kilo/z-ai/glm-5-turbo` | 203K | | | | | | $1 | $4 |
|
|
425
|
-
| `kilo/z-ai/glm-5.1` |
|
|
427
|
+
| `kilo/z-ai/glm-5.1` | 203K | | | | | | $1 | $4 |
|
|
426
428
|
| `kilo/z-ai/glm-5.2` | 1.0M | | | | | | $1 | $4 |
|
|
427
429
|
| `kilo/z-ai/glm-5.2:free` | 33K | | | | | | — | — |
|
|
428
430
|
| `kilo/z-ai/glm-5.3` | 1.0M | | | | | | $1 | $4 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# NanoGPT
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 593 NanoGPT models through Mastra's model router. Authentication is handled automatically using the `NANO_GPT_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [NanoGPT documentation](https://docs.nano-gpt.com).
|
|
10
10
|
|
|
@@ -83,7 +83,7 @@ for await (const chunk of stream) {
|
|
|
83
83
|
| `nano-gpt/anthropic/claude-opus-4.8:thinking` | 1.0M | | | | | | $5 | $25 |
|
|
84
84
|
| `nano-gpt/anthropic/claude-opus-5` | 1.0M | | | | | | $5 | $25 |
|
|
85
85
|
| `nano-gpt/anthropic/claude-opus-5.5` | 1.0M | | | | | | $4 | $20 |
|
|
86
|
-
| `nano-gpt/anthropic/claude-opus-latest` | 1.0M | | | | | | $
|
|
86
|
+
| `nano-gpt/anthropic/claude-opus-latest` | 1.0M | | | | | | $4 | $20 |
|
|
87
87
|
| `nano-gpt/anthropic/claude-sonnet-4` | 200K | | | | | | $3 | $15 |
|
|
88
88
|
| `nano-gpt/anthropic/claude-sonnet-4:thinking` | 1.0M | | | | | | $3 | $15 |
|
|
89
89
|
| `nano-gpt/anthropic/claude-sonnet-4:thinking:1024` | 1.0M | | | | | | $3 | $15 |
|
|
@@ -400,7 +400,7 @@ for await (const chunk of stream) {
|
|
|
400
400
|
| `nano-gpt/openai/gpt-astra-latest` | 1.1M | | | | | | $10 | $50 |
|
|
401
401
|
| `nano-gpt/openai/gpt-chat-latest` | 1.1M | | | | | | $2 | $10 |
|
|
402
402
|
| `nano-gpt/openai/gpt-latest` | 1.1M | | | | | | $10 | $50 |
|
|
403
|
-
| `nano-gpt/openai/gpt-luna-latest` | 1.1M | | | | | | $0.
|
|
403
|
+
| `nano-gpt/openai/gpt-luna-latest` | 1.1M | | | | | | $0.10 | $0.50 |
|
|
404
404
|
| `nano-gpt/openai/gpt-oss-120b` | 128K | | | | | | $0.35 | $0.75 |
|
|
405
405
|
| `nano-gpt/openai/gpt-oss-20b` | 128K | | | | | | $0.20 | $0.30 |
|
|
406
406
|
| `nano-gpt/openai/gpt-oss-safeguard-20b` | 128K | | | | | | $0.07 | $0.30 |
|
|
@@ -586,12 +586,13 @@ for await (const chunk of stream) {
|
|
|
586
586
|
| `nano-gpt/x-ai/grok-4.6` | 500K | | | | | | $2 | $6 |
|
|
587
587
|
| `nano-gpt/x-ai/grok-4.7` | 500K | | | | | | $2 | $5 |
|
|
588
588
|
| `nano-gpt/x-ai/grok-build-0.1` | 256K | | | | | | $1 | $2 |
|
|
589
|
-
| `nano-gpt/x-ai/grok-latest` | 500K | | | | | | $2 | $
|
|
589
|
+
| `nano-gpt/x-ai/grok-latest` | 500K | | | | | | $2 | $5 |
|
|
590
590
|
| `nano-gpt/xiaomi/mimo-v2.5` | 1.0M | | | | | | $0.14 | $0.28 |
|
|
591
591
|
| `nano-gpt/xiaomi/mimo-v2.5-pro` | 1.0M | | | | | | $0.43 | $0.87 |
|
|
592
592
|
| `nano-gpt/xiaomi/mimo-v2.5-pro:thinking` | 1.0M | | | | | | $0.43 | $0.87 |
|
|
593
593
|
| `nano-gpt/xiaomi/mimo-v2.5:thinking` | 1.0M | | | | | | $0.14 | $0.28 |
|
|
594
594
|
| `nano-gpt/xiaomi/mimo-v2.6-flash` | 1.0M | | | | | | $0.14 | $0.28 |
|
|
595
|
+
| `nano-gpt/xiaomi/mimo-v2.6-flash-uncensored` | 1.0M | | | | | | $0.50 | $2 |
|
|
595
596
|
| `nano-gpt/xiaomi/mimo-v2.6-pro` | 1.0M | | | | | | $0.43 | $0.87 |
|
|
596
597
|
| `nano-gpt/xiaomi/mimo-v2.6-pro-ultraspeed` | 1.0M | | | | | | $4 | $9 |
|
|
597
598
|
| `nano-gpt/z-ai/glm-4.5` | 128K | | | | | | $0.30 | $1 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# OpenCode Zen
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 111 OpenCode Zen models through Mastra's model router. Authentication is handled automatically using the `OPENCODE_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [OpenCode Zen documentation](https://opencode.ai/docs/zen).
|
|
10
10
|
|
|
@@ -105,7 +105,6 @@ for await (const chunk of stream) {
|
|
|
105
105
|
| `opencode/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
|
|
106
106
|
| `opencode/minimax-m3` | 512K | | | | | | $0.30 | $1 |
|
|
107
107
|
| `opencode/muse-spark-1.2` | 1.0M | | | | | | $1 | $4 |
|
|
108
|
-
| `opencode/muse-spark-1.2-contributor-free` | 1.0M | | | | | | — | — |
|
|
109
108
|
| `opencode/muse-spark-1.3` | 1.0M | | | | | | $1 | $4 |
|
|
110
109
|
| `opencode/muse-spark-1.3-contributor-free` | 1.0M | | | | | | — | — |
|
|
111
110
|
| `opencode/nemotron-3-ultra-free` | 1.0M | | | | | | — | — |
|
|
@@ -113,6 +112,7 @@ for await (const chunk of stream) {
|
|
|
113
112
|
| `opencode/qwen3.5-plus` | 262K | | | | | | $0.20 | $1 |
|
|
114
113
|
| `opencode/qwen3.6-plus` | 262K | | | | | | $0.50 | $3 |
|
|
115
114
|
| `opencode/qwen3.8-flash` | 1.0M | | | | | | $0.15 | $0.47 |
|
|
115
|
+
| `opencode/qwen3.8-max` | 262K | | | | | | $2 | $6 |
|
|
116
116
|
| `opencode/space-bunny-free` | 1.0M | | | | | | — | — |
|
|
117
117
|
|
|
118
118
|
Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
|
|
@@ -510,6 +510,8 @@ Subscribes to raw stream chunks for a memory thread. Use this before calling `se
|
|
|
510
510
|
|
|
511
511
|
**options.hideSignals** (`boolean | AgentSignalType[]`): Use true to hide all recognized signals, false to show all, or an array to hide selected types from this subscription, including live, idle-persisted, and replayed signals. Other subscribers and the initiating stream keep their own policies.
|
|
512
512
|
|
|
513
|
+
**options.withInitialHistory** (`boolean | { perPage?: number }`): Emit the stored thread messages as one thread-history chunk first, then only retained parts newer than that history, then live parts. Completed runs already in storage aren't replayed; pending tool approvals are still emitted. perPage limits how many recent messages are loaded (default 40). Requires agent memory; without it the subscription is live-only.
|
|
514
|
+
|
|
513
515
|
By default, subscriptions include every signal type, including reactive reminders. To hide reminders for one subscriber without affecting another:
|
|
514
516
|
|
|
515
517
|
```ts
|
|
@@ -520,6 +522,19 @@ const filtered = await agent.subscribeToThread({
|
|
|
520
522
|
})
|
|
521
523
|
```
|
|
522
524
|
|
|
525
|
+
To render a thread without replaying finished runs, request initial history. The first chunk carries the stored messages:
|
|
526
|
+
|
|
527
|
+
```ts
|
|
528
|
+
const subscription = await agent.subscribeToThread({
|
|
529
|
+
threadId: 'thread-abc',
|
|
530
|
+
withInitialHistory: true,
|
|
531
|
+
})
|
|
532
|
+
for await (const chunk of subscription.stream) {
|
|
533
|
+
if (chunk.type === 'thread-history') render(chunk.payload.messages)
|
|
534
|
+
else handle(chunk)
|
|
535
|
+
}
|
|
536
|
+
```
|
|
537
|
+
|
|
523
538
|
`hideSignals: true` hides all recognized signal types. Set it to `false` to show all signals. An array accepts `user`, `state`, `reactive`, `notification`, `user-message`, and `system-reminder`. Matching normalizes `system-reminder` to `reactive` and `user-message` to `user`. An omitted option, `false`, or `[]` excludes nothing. Each subscriber filters its own live and replayed chunks, including remote pubsub events and idle-persisted signals. Filtering preserves non-signal chunks, unknown or malformed signal chunks, ordering, and completion/error events, even when every signal type is excluded.
|
|
524
539
|
|
|
525
540
|
Exclusions don't change model context, storage, shared broadcasts, or another caller's output. They aren't a security boundary and don't replace `ifActive`/`ifIdle` delivery policies or `transient` persistence behavior. This option is supported by the in-process core API, not HTTP or client-js subscription requests. See [stream signal visibility](https://mastra.ai/reference/streaming/agents/stream) for transform ordering and the distinction from recall's exact stored-type matching and reminder-hidden history default.
|
|
@@ -96,7 +96,7 @@ Returns: `DurableAgent`
|
|
|
96
96
|
|
|
97
97
|
## `createEventedAgent(options)`
|
|
98
98
|
|
|
99
|
-
Wraps an `Agent` with fire-and-forget durable execution on the built-in workflow engine. Like `createDurableAgent`, it returns a result you stream from, but the
|
|
99
|
+
Wraps an `Agent` with fire-and-forget durable execution on the built-in evented workflow engine. Like `createDurableAgent`, it returns a result you stream from, but the run is started without being awaited: execution is driven by events on the Mastra instance's PubSub, so the run progresses independently of the caller. Register the agent on a `Mastra` instance with storage to get evented execution; without one, the agent falls back to the default in-process engine and logs a warning. It doesn't accept `id` or `name` overrides.
|
|
100
100
|
|
|
101
101
|
```typescript
|
|
102
102
|
import { createEventedAgent } from '@mastra/core/agent/durable'
|
|
@@ -302,6 +302,8 @@ await subscription.processDataStream({
|
|
|
302
302
|
})
|
|
303
303
|
```
|
|
304
304
|
|
|
305
|
+
Pass `withInitialHistory: true` (or `{ perPage }`) to receive the stored messages as a first `thread-history` chunk instead of fetching history separately. Finished runs aren't replayed after it, so completed answers don't stream again on reconnect.
|
|
306
|
+
|
|
305
307
|
`subscribeToThread()` returns the underlying `Response` plus a `processDataStream()` helper. The helper reads the subscription stream until the connection closes or the request is aborted. Pass `reconnect: true` to resubscribe when the transport closes or a reconnect request fails, such as after a proxy idle timeout.
|
|
306
308
|
|
|
307
309
|
**resourceId** (`string`): Resource ID for the memory thread.
|
|
@@ -171,7 +171,7 @@ When working memory is enabled, `cloneThread()` copies or shares working memory
|
|
|
171
171
|
When [Observational Memory](https://mastra.ai/docs/memory/observational-memory) is enabled, `cloneThread()` automatically clones the OM records associated with the source thread. The behavior depends on the OM scope:
|
|
172
172
|
|
|
173
173
|
- **Thread-scoped OM**: The OM record is cloned to the new thread. All internal message ID references are remapped to point to the cloned messages.
|
|
174
|
-
- **Resource-scoped OM (same `resourceId`)**: The OM record is shared between the source and cloned threads since they belong to the same resource. No duplication occurs.
|
|
175
|
-
- **Resource-scoped OM (different `resourceId`)**: The OM record is cloned to the new resource. Message IDs are remapped and any thread-identifying tags within observations are updated to reference the cloned thread.
|
|
174
|
+
- **Resource-scoped OM (deprecated, same `resourceId`)**: The OM record is shared between the source and cloned threads since they belong to the same resource. No duplication occurs.
|
|
175
|
+
- **Resource-scoped OM (deprecated, different `resourceId`)**: The OM record is cloned to the new resource. Message IDs are remapped and any thread-identifying tags within observations are updated to reference the cloned thread.
|
|
176
176
|
|
|
177
177
|
Only the current (most recent) OM generation is cloned. Older history generations aren't copied. Transient processing state (observation/reflection in-progress flags) is reset on the cloned record.
|
|
@@ -39,7 +39,7 @@ OM performs thresholding with fast local token estimation. Text uses `tokenx`, a
|
|
|
39
39
|
|
|
40
40
|
**model** (`string | LanguageModel | DynamicModel | ModelByInputTokens | ModelWithRetries[]`): Model for both the Observer and Reflector agents. Sets the model for both at once. Cannot be used together with observation.model or reflection.model — an error will be thrown if both are set. When this and observation.model/reflection.model are all omitted, OM falls back to google/gemini-2.5-flash. Use "default" to explicitly use the default model (google/gemini-2.5-flash). (Default: `'google/gemini-2.5-flash'`)
|
|
41
41
|
|
|
42
|
-
**scope** (`'resource' | 'thread'`): Memory scope for observations. 'thread' keeps observations per-thread. 'resource'
|
|
42
|
+
**scope** (`'resource' | 'thread'`): Memory scope for observations. 'thread' keeps observations per-thread. 'resource' is deprecated and will be removed in a future release. It shares observations across all threads for a resource and works much worse than thread scope for prompt caching and agent understanding. For cross-thread continuity, use thread scope with retrieval or resource-scoped working memory. See resource scope (deprecated). (Default: `'thread'`)
|
|
43
43
|
|
|
44
44
|
**activateAfterIdle** (`number | string | false | "auto"`): Time before buffered observations are forced to activate after inactivity, even before observation.messageTokens is reached. Accepts a numeric millisecond value such as 300\_000, duration strings like "5m" or "1hr", "auto" for a provider-aware prompt cache TTL, or false to disable inherited observation idle activation. Reflections do not inherit this setting. Use reflection.activateAfterIdle to opt reflections into idle activation.
|
|
45
45
|
|
|
@@ -71,7 +71,7 @@ OM performs thresholding with fast local token estimation. Text uses `tokenx`, a
|
|
|
71
71
|
|
|
72
72
|
**observation.messageTokens** (`number`): Token count of unobserved messages that triggers observation. When unobserved message tokens exceed this threshold, the Observer agent is called. Text is estimated locally with tokenx. Image parts are included with model-aware heuristics when possible, with deterministic fallbacks when image metadata is incomplete. Image-like file parts are counted the same way when uploads are normalized as files.
|
|
73
73
|
|
|
74
|
-
**observation.maxTokensPerBatch** (`number`): Maximum tokens per batch when observing multiple threads in resource scope. Threads are chunked into batches of this size and processed in parallel. Lower values mean more parallelism but more API calls.
|
|
74
|
+
**observation.maxTokensPerBatch** (`number`): Maximum tokens per batch when observing multiple threads in resource scope (deprecated). Threads are chunked into batches of this size and processed in parallel. Lower values mean more parallelism but more API calls.
|
|
75
75
|
|
|
76
76
|
**observation.modelSettings** (`ObservationalMemoryModelSettings`): Model settings for the Observer agent. The temperature: 0.3 default is only applied when the resolved model is known to support temperature. The maxOutputTokens: 100\_000 default is only applied with default model selection (no model set, "default", or a ModelByInputTokens selector). Custom models get no maxOutputTokens default.
|
|
77
77
|
|
|
@@ -216,7 +216,7 @@ const memory = new Memory({
|
|
|
216
216
|
|
|
217
217
|
Set `workingMemory.agentManaged: true` if the main agent should still receive working memory tool and instruction injection.
|
|
218
218
|
|
|
219
|
-
###
|
|
219
|
+
### Custom thresholds
|
|
220
220
|
|
|
221
221
|
```typescript
|
|
222
222
|
import { Memory } from '@mastra/memory'
|
|
@@ -231,7 +231,6 @@ export const agent = new Agent({
|
|
|
231
231
|
options: {
|
|
232
232
|
observationalMemory: {
|
|
233
233
|
model: 'google/gemini-2.5-flash',
|
|
234
|
-
scope: 'resource',
|
|
235
234
|
observation: {
|
|
236
235
|
messageTokens: 20_000,
|
|
237
236
|
},
|
|
@@ -420,7 +419,7 @@ observationalMemory: {
|
|
|
420
419
|
|
|
421
420
|
Setting `bufferTokens: false` disables both observation and reflection async buffering. Observations and reflections will run synchronously when their thresholds are reached.
|
|
422
421
|
|
|
423
|
-
> **Note:** Async buffering isn't supported with `scope: 'resource'` and is automatically disabled in resource scope.
|
|
422
|
+
> **Note:** Async buffering isn't supported with the deprecated `scope: 'resource'` and is automatically disabled in resource scope.
|
|
424
423
|
|
|
425
424
|
## Streaming data parts
|
|
426
425
|
|
|
@@ -734,7 +733,6 @@ const om = new ObservationalMemory({
|
|
|
734
733
|
storage: storage.stores.memory!,
|
|
735
734
|
memory,
|
|
736
735
|
model: 'google/gemini-2.5-flash',
|
|
737
|
-
scope: 'resource',
|
|
738
736
|
observation: {
|
|
739
737
|
messageTokens: 20_000,
|
|
740
738
|
},
|
|
@@ -32,6 +32,8 @@ Add the processor to an agent's `inputProcessors`. Each invocation injects at mo
|
|
|
32
32
|
|
|
33
33
|
**getReader** (`(args: ProcessInputStepArgs) => ReminderFileReader | undefined`): Select a reader for this request. Returning undefined keeps the instance defaults.
|
|
34
34
|
|
|
35
|
+
**getBasePath** (`(args: ProcessInputStepArgs) => string | undefined`): Return the project root for this request. Relative tool paths resolve against it instead of process.cwd(), and instruction files outside it are never loaded. Returning undefined keeps process.cwd() resolution with no boundary.
|
|
36
|
+
|
|
35
37
|
## `ReminderFileReader`
|
|
36
38
|
|
|
37
39
|
A reader controls both file access and optional path identity. Return one from `getReader` when instruction files live in a virtual filesystem or a trusted git ref rather than the current checkout.
|
|
@@ -119,6 +119,17 @@ The default implementation is a no-op because transports that retain nothing per
|
|
|
119
119
|
await pubsub.clearTopic('workflow.events.v2.run-123')
|
|
120
120
|
```
|
|
121
121
|
|
|
122
|
+
#### `trimTopic(topic, { runId, producedBefore? })`
|
|
123
|
+
|
|
124
|
+
Deletes every retained entry on the topic that was published with `runId`. Other entries are untouched. Mastra calls this once a run completes successfully and its messages are saved to storage, so finished runs leave the topic while runs still in progress or waiting for approval stay. The default implementation is a no-op. [`RedisStreamsPubSub`](https://mastra.ai/reference/pubsub/redis-streams) overrides it. Same best-effort contract as `clearTopic`.
|
|
125
|
+
|
|
126
|
+
With `producedBefore` (epoch ms), only entries whose `data.producedAt` is at or before it are deleted, and entries marked `data.pinned` (approval and suspension prompts) are kept. Mastra uses this while a run continues: each time the run's messages are saved mid-run, the parts up to the last finished step leave the topic.
|
|
127
|
+
|
|
128
|
+
```typescript
|
|
129
|
+
await pubsub.trimTopic('agent.thread.thread-abc', { runId: 'run-123' })
|
|
130
|
+
await pubsub.trimTopic('agent.thread.thread-abc', { runId: 'run-123', producedBefore: Date.now() })
|
|
131
|
+
```
|
|
132
|
+
|
|
122
133
|
### Replay methods
|
|
123
134
|
|
|
124
135
|
These methods support resuming a stream after a disconnect. The default implementations fall back to a regular `subscribe`, so backends without history support behave as live-only. [`CachingPubSub`](https://mastra.ai/reference/pubsub/caching-pubsub) overrides them to replay cached events.
|
|
@@ -159,6 +159,14 @@ Not every topic reaches a `clearTopic` call. Topics without a defined end of lif
|
|
|
159
159
|
await pubsub.clearTopic('workflow.events.run-123')
|
|
160
160
|
```
|
|
161
161
|
|
|
162
|
+
### `trimTopic(topic, { runId, producedBefore? })`
|
|
163
|
+
|
|
164
|
+
Pages through the stream with `XRANGE` and runs `XDEL` for every entry published with `runId`. With `producedBefore`, only entries produced at or before that time are deleted, skipping approval and suspension prompts. Mastra uses this to trim a long run each time its messages are saved. Mastra calls this once a thread run's messages are persisted, so an idle thread's stream empties instead of growing to `maxLen`. Because entries are matched by `runId`, a run resumed after a restart also removes the entries it published before the restart. Entries from other runs, including those from other processes, are never touched. Best-effort. Failures are logged at warn level.
|
|
165
|
+
|
|
166
|
+
```typescript
|
|
167
|
+
await pubsub.trimTopic('agent.thread.thread-abc', { runId: 'run-123' })
|
|
168
|
+
```
|
|
169
|
+
|
|
162
170
|
### `close()`
|
|
163
171
|
|
|
164
172
|
Closes the Redis connections and stops all subscriptions. Call this during graceful shutdown.
|
|
@@ -516,7 +516,7 @@ Added when a sandbox is configured:
|
|
|
516
516
|
|
|
517
517
|
With a static sandbox, capability checks (`executeCommand`, `processes`) decide which tool variants are exposed. With a [runtime-defined sandbox](https://mastra.ai/docs/sandbox/overview), all sandbox tools are registered and the runtime throws a clear error if the resolved sandbox doesn't implement a requested capability.
|
|
518
518
|
|
|
519
|
-
The `execute_command` tool accepts
|
|
519
|
+
The `execute_command` tool accepts these options:
|
|
520
520
|
|
|
521
521
|
**backgroundProcesses** (`BackgroundProcessesConfig`): Configuration for handling background processes. Only applicable if the sandbox supports background execution.
|
|
522
522
|
|
|
@@ -528,6 +528,19 @@ The `execute_command` tool accepts a `backgroundProcesses` option for lifecycle
|
|
|
528
528
|
|
|
529
529
|
**backgroundProcesses.abortSignal** (`AbortSignal | null | false`): Abort signal for background processes. undefined (default) uses the agent's signal. null or false disables abort — processes persist after agent shutdown.
|
|
530
530
|
|
|
531
|
+
**requireDescription** (`boolean`): Adds a required description argument, listed before command, where the model says in a few words what the command does. Use it to show a readable label in place of the raw command. When false, the tool schema has no description argument. (Default: `false`)
|
|
532
|
+
|
|
533
|
+
```typescript
|
|
534
|
+
const workspace = new Workspace({
|
|
535
|
+
sandbox: new LocalSandbox({ workingDirectory: './workspace' }),
|
|
536
|
+
tools: {
|
|
537
|
+
[WORKSPACE_TOOLS.SANDBOX.EXECUTE_COMMAND]: {
|
|
538
|
+
requireDescription: true,
|
|
539
|
+
},
|
|
540
|
+
},
|
|
541
|
+
})
|
|
542
|
+
```
|
|
543
|
+
|
|
531
544
|
See [Background processes](https://mastra.ai/docs/sandbox/overview) for callback examples.
|
|
532
545
|
|
|
533
546
|
### Search tools
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@mastra/mcp-docs-server",
|
|
3
|
-
"version": "1.3.
|
|
3
|
+
"version": "1.3.2-alpha.3",
|
|
4
4
|
"description": "MCP server for accessing Mastra.ai documentation, changelogs, and news.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "dist/index.js",
|
|
@@ -26,7 +26,7 @@
|
|
|
26
26
|
"@mastra/mcp-legacy": "npm:@mastra/mcp@^1.18.0",
|
|
27
27
|
"local-pkg": "^1.1.2",
|
|
28
28
|
"zod": "^4.6.4",
|
|
29
|
-
"@mastra/core": "1.
|
|
29
|
+
"@mastra/core": "1.72.0-alpha.1",
|
|
30
30
|
"@mastra/mcp": "^2.1.0"
|
|
31
31
|
},
|
|
32
32
|
"devDependencies": {
|
|
@@ -44,8 +44,8 @@
|
|
|
44
44
|
"typescript": "^7.0.2",
|
|
45
45
|
"vitest": "4.1.11",
|
|
46
46
|
"@internal/lint": "0.0.137",
|
|
47
|
-
"@
|
|
48
|
-
"@
|
|
47
|
+
"@internal/types-builder": "0.0.112",
|
|
48
|
+
"@mastra/core": "1.72.0-alpha.1"
|
|
49
49
|
},
|
|
50
50
|
"homepage": "https://mastra.ai",
|
|
51
51
|
"repository": {
|