@mastra/mcp-docs-server 1.2.24 → 1.2.25-alpha.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.docs/docs/agents/processors.md +2 -2
- package/.docs/docs/auth/fga.md +1 -1
- package/.docs/docs/channels.md +1 -1
- package/.docs/docs/evals/custom-scorers.md +2 -2
- package/.docs/docs/evals/experiments.md +3 -3
- package/.docs/docs/evals/multi-turn.md +1 -1
- package/.docs/docs/guides/streaming.md +1 -1
- package/.docs/docs/harness/agent-controller.md +1 -1
- package/.docs/docs/mastra-platform/database.md +3 -1
- package/.docs/docs/memory/message-history.md +1 -1
- package/.docs/docs/memory/multi-user-threads.md +1 -1
- package/.docs/docs/memory/observational-memory.md +41 -1
- package/.docs/docs/memory/overview.md +4 -4
- package/.docs/docs/observability/tracing/overview.md +2 -2
- package/.docs/docs/server/pubsub.md +1 -1
- package/.docs/docs/studio/overview.md +1 -1
- package/.docs/models/gateways/merge-gateway.md +2 -1
- package/.docs/models/index.md +1 -1
- package/.docs/models/providers/edenai.md +14 -6
- package/.docs/models/providers/kilo.md +3 -3
- package/.docs/models/providers/nano-gpt.md +1 -6
- package/.docs/reference/agents/generate.md +1 -1
- package/.docs/reference/agents/network.md +1 -1
- package/.docs/reference/coding-agent/build-base-prompt.md +4 -4
- package/.docs/reference/evals/completeness.md +1 -1
- package/.docs/reference/evals/faithfulness.md +1 -1
- package/.docs/reference/evals/noise-sensitivity.md +1 -1
- package/.docs/reference/evals/prompt-alignment.md +2 -4
- package/.docs/reference/memory/observational-memory.md +1 -1
- package/.docs/reference/migrations/upgrade-to-v1/workflows.md +1 -1
- package/.docs/reference/processors/regex-filter-processor.md +1 -1
- package/.docs/reference/rag/vector-databases.md +1 -1
- package/.docs/reference/streaming/agents/stream.md +7 -3
- package/.docs/reference/tools/mcp-server.md +1 -1
- package/.docs/reference/workspace/sandbox.md +2 -2
- package/package.json +3 -3
|
@@ -442,7 +442,7 @@ Caching is implemented as the [`ResponseCache`](https://mastra.ai/reference/proc
|
|
|
442
442
|
|
|
443
443
|
### When to use response caching
|
|
444
444
|
|
|
445
|
-
|
|
445
|
+
Use caching when identical requests recur across users or sessions. Examples include suggested-prompt buttons and repeated searches, or guardrail LLMs that classify the same input. Skip it when calls trigger external side effects through tools, since cache hits replay tool calls without re-executing them.
|
|
446
446
|
|
|
447
447
|
### Quickstart
|
|
448
448
|
|
|
@@ -546,7 +546,7 @@ The cache key is derived from the resolved `LanguageModelV2Prompt` Mastra is abo
|
|
|
546
546
|
|
|
547
547
|
When you don't supply `key`, the processor derives one deterministically from the inputs that change the LLM's response at this step: `agentId`, `stepNumber` (so each step in a tool loop has its own cache entry), `scope`, model identity (`provider`, `modelId`, spec version), and the resolved `prompt` (post-memory + post-processors). Any change to these inputs automatically invalidates the cache.
|
|
548
548
|
|
|
549
|
-
Multimodal prompts are included too.
|
|
549
|
+
Multimodal prompts are included too. For image and file parts, the key includes the full URL or a digest of the bytes for inline binary data (`Uint8Array`, `ArrayBuffer`). Requests that differ only in which image they reference therefore get different cache entries.
|
|
550
550
|
|
|
551
551
|
#### Customize the cache key
|
|
552
552
|
|
package/.docs/docs/auth/fga.md
CHANGED
|
@@ -321,7 +321,7 @@ The actor signal is trusted input, so construct it server-side:
|
|
|
321
321
|
- Establish tenant scope server-side. Built-in agent HTTP routes ignore a client-supplied `organizationId` in the request context, and the trusted-actor path requires an `organizationId` to be set.
|
|
322
322
|
- Durable resume keeps its existing request-context recovery and merge behavior. This doesn't make a persisted actor trusted for a later workflow segment.
|
|
323
323
|
- The tenant-scope check confirms that a trusted `organizationId` exists. It doesn't verify that `actor.agentId` belongs to that organization. When that relationship matters, verify it in `requireActor` using authoritative provider data.
|
|
324
|
-
- Treat `actor.permissions` as an unverified claim.
|
|
324
|
+
- Treat `actor.permissions` as an unverified claim. To enforce least privilege, resolve the agent's permissions from a trusted source, such as a manifest or your FGA backend keyed by `agentId`. Don't trust the inline values.
|
|
325
325
|
- Once a provider implements `requireActor`, errors from that method stop execution. Mastra doesn't fall back to organization-only authorization.
|
|
326
326
|
|
|
327
327
|
## Related
|
package/.docs/docs/channels.md
CHANGED
|
@@ -70,7 +70,7 @@ export const mastra = new Mastra({
|
|
|
70
70
|
|
|
71
71
|
## Webhook routes
|
|
72
72
|
|
|
73
|
-
Platforms send channel activity to Mastra through webhooks. A webhook is an HTTP endpoint that the platform calls when something happens, such as a new message, a mention, or a user selecting "Approve" on an interactive tool approval card.
|
|
73
|
+
Platforms send channel activity to Mastra through webhooks. A webhook is an HTTP endpoint that the platform calls when something happens, such as a new message, a mention, or a user selecting "Approve" on an interactive tool approval card. The webhook delivers the message to your agent for processing, and the agent responds in the same channel.
|
|
74
74
|
|
|
75
75
|
Mastra registers a webhook route for each configured adapter and handles the request for you:
|
|
76
76
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Custom scorers
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
Use `createScorer` to build custom evaluation logic. Each step can use a JavaScript function or an LLM-based prompt object, depending on what you're evaluating.
|
|
8
8
|
|
|
9
9
|
## The four-step pipeline
|
|
10
10
|
|
|
@@ -15,7 +15,7 @@ All scorers in Mastra follow a consistent four-step evaluation pipeline:
|
|
|
15
15
|
3. **generateScore** (required): Convert analysis into a numerical score
|
|
16
16
|
4. **generateReason** (optional): Generate human-readable explanations
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
Combine **functions** and **prompt objects** within a scorer to use deterministic checks for some steps and LLM-based evaluation for others.
|
|
19
19
|
|
|
20
20
|
## Functions vs prompt objects
|
|
21
21
|
|
|
@@ -315,7 +315,7 @@ Teardown failures are logged rather than propagated. By the time `afterEach` run
|
|
|
315
315
|
|
|
316
316
|
When an experiment runs an agent that calls side-effecting tools, attach static tool mocks to individual dataset items to make the run deterministic. During the experiment, a mocked tool returns its declared output instead of executing. Tools without a mock on the item run live by default.
|
|
317
317
|
|
|
318
|
-
Mocks
|
|
318
|
+
Mocks are stored and versioned with the dataset item. Each mock specifies the tool name and expected arguments, along with the output to return:
|
|
319
319
|
|
|
320
320
|
```typescript
|
|
321
321
|
await dataset.addItem({
|
|
@@ -357,7 +357,7 @@ The item value takes precedence over the experiment value. A denied call fails w
|
|
|
357
357
|
|
|
358
358
|
### Matching and consumption
|
|
359
359
|
|
|
360
|
-
Arguments are matched strictly: object key order is ignored
|
|
360
|
+
Arguments are matched strictly: object key order is ignored, but array order matters. Values aren't coerced to other types. A mock is served only when the agent calls the tool with arguments that deep-equal the mock's `args`.
|
|
361
361
|
|
|
362
362
|
When an item declares several mocks for the same tool and arguments, they're consumed in order, the first call gets the first mock, the next call gets the second, and so on. Ordering is tracked per `(toolName, args)` group and is independent across different arguments.
|
|
363
363
|
|
|
@@ -390,7 +390,7 @@ While mock interception is active, the agent's tools execute sequentially so rep
|
|
|
390
390
|
|
|
391
391
|
### Diagnostics
|
|
392
392
|
|
|
393
|
-
Each item result
|
|
393
|
+
Each item result includes a `toolMockReport` listing which mocks were used and which tool calls ran without a mock:
|
|
394
394
|
|
|
395
395
|
```typescript
|
|
396
396
|
for (const item of summary.results) {
|
|
@@ -191,7 +191,7 @@ result.turnResults // per-turn gate/threshold/scorer outcomes
|
|
|
191
191
|
|
|
192
192
|
Semantics:
|
|
193
193
|
|
|
194
|
-
- A per-turn gate or scorer
|
|
194
|
+
- A per-turn gate or scorer receives **only that turn's** `run.input` and `run.output`, rather than the accumulated conversation. A different turn's output can't satisfy the check, and `run.input` contains the input for the turn being evaluated.
|
|
195
195
|
- Per-turn outcomes fold into the [verdict](https://mastra.ai/docs/evals/gates-and-verdicts): a failing turn gate makes the verdict `failed`. A missed turn threshold (with gates passing) makes it `scored`.
|
|
196
196
|
- `result.turnResults[i]` reports each turn's `gateResults`, `thresholdResults`, and `scores`, so a failure points at the exact turn. Across multiple conversations, turn results are averaged by turn index.
|
|
197
197
|
- A turn with no `gates` or `scorers` advances the conversation.
|
|
@@ -108,7 +108,7 @@ Visit [Run.stream()](https://mastra.ai/reference/streaming/workflows/stream) for
|
|
|
108
108
|
|
|
109
109
|
### Output from `Run.stream()`
|
|
110
110
|
|
|
111
|
-
|
|
111
|
+
Events include `runId` and `from` at the top level, so you can identify the workflow run without inspecting the payload.
|
|
112
112
|
|
|
113
113
|
```typescript
|
|
114
114
|
{
|
|
@@ -22,7 +22,7 @@ Use the Agent Controller when your application needs:
|
|
|
22
22
|
- Subagent orchestration to delegate focused subtasks with constrained tools
|
|
23
23
|
- Persistent threads and selected thread settings across restarts, with isolated live state for each Session
|
|
24
24
|
|
|
25
|
-
You could assemble all of this yourself on top of the [Agent class](https://mastra.ai/docs/agents/overview), which exposes the full agent loop, tools, and memory. The AgentController provides opinionated defaults for an ongoing session where the agent acts as a collaborator rather than a one-shot endpoint.
|
|
25
|
+
You could assemble all of this yourself on top of the [Agent class](https://mastra.ai/docs/agents/overview), which exposes the full agent loop, tools, and memory. The AgentController provides opinionated defaults for an ongoing session where the agent acts as a collaborator rather than a one-shot endpoint. Use the Agent class directly for full control or request-response calls. Choose AgentController for ongoing sessions without building your own session runtime.
|
|
26
26
|
|
|
27
27
|
## Quickstart
|
|
28
28
|
|
|
@@ -97,7 +97,9 @@ Databases attached from project settings are project-scoped. Use the [CLI](#atta
|
|
|
97
97
|
|
|
98
98
|
## Connect from your code
|
|
99
99
|
|
|
100
|
-
|
|
100
|
+
Check the database status in **Project Settings → Database**. An attached database shows `provisioning` while setup runs in the background. Once it's `ready`, the platform has injected its connection details as managed environment variables.
|
|
101
|
+
|
|
102
|
+
Open a `ready` database to view those variables and a code snippet. Use the variables to configure a Mastra storage adapter, as shown below.
|
|
101
103
|
|
|
102
104
|
### Turso (LibSQL)
|
|
103
105
|
|
|
@@ -178,7 +178,7 @@ const agent = mastra.getAgentById('test-agent')
|
|
|
178
178
|
const memory = await agent.getMemory()
|
|
179
179
|
```
|
|
180
180
|
|
|
181
|
-
|
|
181
|
+
Use the `Memory` instance to query stored threads and messages or clone a conversation.
|
|
182
182
|
|
|
183
183
|
## Querying
|
|
184
184
|
|
|
@@ -160,7 +160,7 @@ const memory = new Memory({
|
|
|
160
160
|
|
|
161
161
|
OM requires a storage adapter that supports it: `@mastra/libsql`, `@mastra/pg`, `@mastra/mongodb`, or `@mastra/oracledb`.
|
|
162
162
|
|
|
163
|
-
> **Note:** If you switch the Observer to a
|
|
163
|
+
> **Note:** If you switch the Observer to a less capable model and see facts attributed to a generic `User` instead of individual participants, use [`observation.instruction`](https://mastra.ai/reference/memory/observational-memory) to explain how to interpret the `<turn>` tag.
|
|
164
164
|
|
|
165
165
|
### With working memory
|
|
166
166
|
|
|
@@ -718,7 +718,7 @@ Buffered observations also include continuation hints, a suggested next response
|
|
|
718
718
|
|
|
719
719
|
When message production outpaces the Observer, the `blockAfter` safety threshold allows activation to overshoot the retention target instead of using fewer chunks. Activation still uses no more chunks than needed to reach the target, and the default settings remain unaffected. A synchronous observation runs when the `messageTokens` threshold is reached and buffered activation didn't happen. Buffered activation usually preserves a minimum remaining context (the smaller of \~1k tokens or the configured retention floor), but a single buffered chunk that covers the whole pending window still activates and can leave less.
|
|
720
720
|
|
|
721
|
-
Reflection works similarly
|
|
721
|
+
Reflection works similarly: the Reflector runs in the background when observations reach a fraction of the reflection threshold.
|
|
722
722
|
|
|
723
723
|
### Settings
|
|
724
724
|
|
|
@@ -809,6 +809,46 @@ const memory = new Memory({
|
|
|
809
809
|
- `previousObserverTokens: 0` → omit previous observations completely.
|
|
810
810
|
- `previousObserverTokens: false` → disable truncation and keep full previous observations.
|
|
811
811
|
|
|
812
|
+
## Hooks
|
|
813
|
+
|
|
814
|
+
OM exposes two kinds of config-level hooks on `observationalMemory.hooks`:
|
|
815
|
+
|
|
816
|
+
- **Lifecycle hooks** (`onObservationStart`, `onObservationEnd`, `onReflectionStart`, `onReflectionEnd`) are telemetry callbacks. They receive `threadId`, `resourceId`, and `trigger`, and the end hooks also receive the model call's `usage`, `providerMetadata`, and any `error`. They never change what OM stores.
|
|
817
|
+
- **Transform hooks** (`beforeObservation`, `afterObservation`, `beforeReflection`, `afterReflection`) intercept the data flowing through a cycle. Return `void` to pass the input through unchanged, or return a replacement to change what the Observer/Reflector sees or what gets persisted.
|
|
818
|
+
|
|
819
|
+
```typescript
|
|
820
|
+
const memory = new Memory({
|
|
821
|
+
options: {
|
|
822
|
+
observationalMemory: {
|
|
823
|
+
model: 'google/gemini-2.5-flash',
|
|
824
|
+
hooks: {
|
|
825
|
+
// Drop or redact messages before the Observer sees them.
|
|
826
|
+
beforeObservation: ({ messages }) => ({
|
|
827
|
+
messages: messages.filter(m => !isSensitive(m)),
|
|
828
|
+
}),
|
|
829
|
+
// Rewrite observations before they are persisted.
|
|
830
|
+
afterObservation: ({ observations, threadId }) => ({
|
|
831
|
+
observations: redact(observations),
|
|
832
|
+
}),
|
|
833
|
+
// Rewrite the text the Reflector condenses, or its output.
|
|
834
|
+
beforeReflection: ({ observations }) => ({
|
|
835
|
+
observations: stripInternalNotes(observations),
|
|
836
|
+
}),
|
|
837
|
+
afterReflection: async ({ observations, resourceId }) => {
|
|
838
|
+
await syncToExternalStore(resourceId, observations)
|
|
839
|
+
},
|
|
840
|
+
},
|
|
841
|
+
},
|
|
842
|
+
},
|
|
843
|
+
})
|
|
844
|
+
```
|
|
845
|
+
|
|
846
|
+
Transform hooks are always awaited, on every path (manual `observe()`/`reflect()`, turn-synchronous observation, and async buffering). If `beforeObservation` returns an empty `messages` array, the Observer model call is skipped and the filtered messages are still marked as observed. If a transform hook throws, the cycle fails before committing the transformed observation or reflection text. This doesn't roll back extractor callbacks or other side effects that have already run.
|
|
847
|
+
|
|
848
|
+
`afterObservation` and `afterReflection` replace only the observation or reflection text. They don't recompute or redact the separate structured extractor results stored in thread metadata. Reflection extraction and its callbacks run before `afterReflection`, so rewriting the reflection doesn't rerun those callbacks. Use extractor configuration and callbacks to control structured values. Don't treat an after hook as a redaction boundary for all cycle data.
|
|
849
|
+
|
|
850
|
+
Because hooks receive `threadId` and `resourceId`, you can also use them to update [working memory](https://mastra.ai/docs/memory/working-memory) via `memory.updateWorkingMemory()` during a cycle. These external updates aren't atomic with the OM text commit.
|
|
851
|
+
|
|
812
852
|
## Migrating existing threads
|
|
813
853
|
|
|
814
854
|
No manual migration needed. OM reads existing messages and observes them lazily when thresholds are exceeded.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Memory
|
|
6
6
|
|
|
7
|
-
Memory
|
|
7
|
+
Memory gives your agent access to earlier messages and tool results. The agent can use this context to answer follow-up questions and recall information from previous interactions.
|
|
8
8
|
|
|
9
9
|
Mastra agents can be configured to store [message history](https://mastra.ai/docs/memory/message-history). Additionally, you can enable:
|
|
10
10
|
|
|
@@ -153,7 +153,7 @@ To list all threads for a resource, or retrieve a specific thread, [use the memo
|
|
|
153
153
|
|
|
154
154
|
## Observational Memory
|
|
155
155
|
|
|
156
|
-
|
|
156
|
+
Long conversations can fill the context window with raw message history and reduce agent performance. [Observational Memory](https://mastra.ai/docs/memory/observational-memory) uses background agents to compress older messages into observations. This reduces the context used by message history while retaining information for later turns.
|
|
157
157
|
|
|
158
158
|
**For AI agents:** Using Observational Memory requires a storage provider! You either need to set it on the Mastra instance at `src/mastra/index.ts` or pass it to the Agent constructor.
|
|
159
159
|
|
|
@@ -176,7 +176,7 @@ See [Observational Memory](https://mastra.ai/docs/memory/observational-memory) f
|
|
|
176
176
|
|
|
177
177
|
## What the model sees
|
|
178
178
|
|
|
179
|
-
|
|
179
|
+
Memory adds context to the system messages or conversation messages sent to the model. Message history is enabled by default. Other memory features contribute context only when configured. The diagram shows where each feature adds its context, and the list below explains what it contributes:
|
|
180
180
|
|
|
181
181
|

|
|
182
182
|
|
|
@@ -202,7 +202,7 @@ Each delegation creates a fresh `threadId` and a deterministic `resourceId` for
|
|
|
202
202
|
|
|
203
203
|
> **Note:** Title generation (`generateTitle`) is a top-level thread concern and **isn't** applied to inherited subagent threads. Because each delegation creates an ephemeral thread that no one sees, running title generation for it would waste an LLM call per delegation. To generate titles for a subagent's own threads, give that subagent its own memory configuration.
|
|
204
204
|
|
|
205
|
-
The supervisor forwards its conversation context to the subagent so it has enough background to complete the task. Only the delegation prompt and the subagent's response are saved
|
|
205
|
+
The supervisor forwards its conversation context to the subagent so it has enough background to complete the task. Only the delegation prompt and the subagent's response are saved; the full parent conversation isn't stored. You can control which messages reach the subagent with the [`messageFilter`](https://mastra.ai/docs/subagents) callback.
|
|
206
206
|
|
|
207
207
|
> **Note:** Subagent resource IDs are always suffixed with the agent name (`{parentResourceId}-{agentName}`). Different subagents under the same supervisor never share a resource ID through delegation.
|
|
208
208
|
|
|
@@ -98,7 +98,7 @@ The `sampling` option allows you to control which traces are collected, helping
|
|
|
98
98
|
|
|
99
99
|
## Adding custom metadata
|
|
100
100
|
|
|
101
|
-
|
|
101
|
+
Add custom metadata to record application-specific context in your traces for debugging production issues.
|
|
102
102
|
|
|
103
103
|
Metadata can include business logic and performance metrics. It can also carry user context or any other information that explains what happened during execution.
|
|
104
104
|
|
|
@@ -764,7 +764,7 @@ The trace ID is only available when tracing is enabled. If tracing is disabled o
|
|
|
764
764
|
|
|
765
765
|
## Integrating with external tracing systems
|
|
766
766
|
|
|
767
|
-
|
|
767
|
+
If your application already uses distributed tracing, such as OpenTelemetry or Datadog, you can connect Mastra traces to the parent trace context. This lets you follow a request through your application and its Mastra agent or workflow calls.
|
|
768
768
|
|
|
769
769
|
### Passing external trace IDs
|
|
770
770
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# PubSub
|
|
6
6
|
|
|
7
|
-
Mastra uses a publish/subscribe (pub/sub) system as its internal event bus. Components publish events to topics, and other components subscribe to those topics to react. The backend
|
|
7
|
+
Mastra uses a publish/subscribe (pub/sub) system as its internal event bus. Components publish events to topics, and other components subscribe to those topics to react. The backend determines whether events are delivered within a single process, between processes on one host, or across separate instances.
|
|
8
8
|
|
|
9
9
|
Set the backend once on the `Mastra` instance for use throughout the system. The default is an in-process backend that requires no setup.
|
|
10
10
|
|
|
@@ -67,7 +67,7 @@ Use [Agent Builder](https://agent-builder.mastra.ai) to create and manage fully
|
|
|
67
67
|
|
|
68
68
|
Visualize your workflow as a graph and run it step by step with a custom input. During execution, the interface updates in real time to show the active step and the path taken.
|
|
69
69
|
|
|
70
|
-
|
|
70
|
+
Workflow traces show tool calls and raw JSON outputs, along with any execution errors.
|
|
71
71
|
|
|
72
72
|
### Processors
|
|
73
73
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Merge Gateway
|
|
6
6
|
|
|
7
|
-
Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
7
|
+
Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 181 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Merge Gateway documentation](https://docs.merge.dev/merge-gateway).
|
|
10
10
|
|
|
@@ -95,6 +95,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
95
95
|
| `google/gemini-flash-lite-latest` |
|
|
96
96
|
| `google/gemma-4-26b-a4b-it` |
|
|
97
97
|
| `google/gemma-4-31b-it` |
|
|
98
|
+
| `meta/llama-3.1-70b-instruct` |
|
|
98
99
|
| `meta/llama-3.1-8b-instruct` |
|
|
99
100
|
| `meta/llama-3.3-70b-instruct` |
|
|
100
101
|
| `meta/muse-spark-1.1` |
|
package/.docs/models/index.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Model Providers
|
|
6
6
|
|
|
7
|
-
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to
|
|
7
|
+
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7122 models from 200 providers through a single API.
|
|
8
8
|
|
|
9
9
|
## Features
|
|
10
10
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Eden AI
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 258 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Eden AI documentation](https://docs.edenai.co).
|
|
10
10
|
|
|
@@ -19,7 +19,7 @@ const agent = new Agent({
|
|
|
19
19
|
id: "my-agent",
|
|
20
20
|
name: "My Agent",
|
|
21
21
|
instructions: "You are a helpful assistant",
|
|
22
|
-
model: "edenai/amazon/
|
|
22
|
+
model: "edenai/amazon/amazon.nova-lite-v1:0"
|
|
23
23
|
});
|
|
24
24
|
|
|
25
25
|
// Generate a response
|
|
@@ -38,6 +38,14 @@ for await (const chunk of stream) {
|
|
|
38
38
|
|
|
39
39
|
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|
|
40
40
|
| ---------------------------------------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
|
|
41
|
+
| `edenai/amazon/amazon.nova-lite-v1:0` | 300K | | | | | | $0.06 | $0.24 |
|
|
42
|
+
| `edenai/amazon/amazon.nova-lite-v1:0@us` | 300K | | | | | | $0.06 | $0.24 |
|
|
43
|
+
| `edenai/amazon/amazon.nova-micro-v1:0` | 128K | | | | | | $0.04 | $0.14 |
|
|
44
|
+
| `edenai/amazon/amazon.nova-micro-v1:0@us` | 128K | | | | | | $0.04 | $0.14 |
|
|
45
|
+
| `edenai/amazon/amazon.nova-pro-v1:0` | 300K | | | | | | $0.80 | $3 |
|
|
46
|
+
| `edenai/amazon/amazon.nova-pro-v1:0@us` | 300K | | | | | | $0.80 | $3 |
|
|
47
|
+
| `edenai/amazon/mistral.pixtral-large-2502-v1:0` | 128K | | | | | | $2 | $6 |
|
|
48
|
+
| `edenai/amazon/mistral.pixtral-large-2502-v1:0@us` | 128K | | | | | | $2 | $6 |
|
|
41
49
|
| `edenai/amazon/moonshot.kimi-k2-thinking` | 128K | | | | | | $0.60 | $3 |
|
|
42
50
|
| `edenai/amazon/moonshotai.kimi-k2.5` | 262K | | | | | | $0.60 | $3 |
|
|
43
51
|
| `edenai/amazon/zai.glm-4.7-flash` | 200K | | | | | | $0.07 | $0.40 |
|
|
@@ -208,8 +216,8 @@ for await (const chunk of stream) {
|
|
|
208
216
|
| `edenai/perplexityai/sonar-deep-research` | 128K | | | | | | $2 | $8 |
|
|
209
217
|
| `edenai/perplexityai/sonar-pro` | 200K | | | | | | $3 | $15 |
|
|
210
218
|
| `edenai/perplexityai/sonar-reasoning-pro` | 128K | | | | | | $2 | $8 |
|
|
211
|
-
| `edenai/qwen/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.
|
|
212
|
-
| `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $
|
|
219
|
+
| `edenai/qwen/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.35 | $1 |
|
|
220
|
+
| `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $3 |
|
|
213
221
|
| `edenai/qwen/qwen-max` | 33K | | | | | | $2 | $6 |
|
|
214
222
|
| `edenai/qwen/qwen-vl-max` | 131K | | | | | | $0.80 | $3 |
|
|
215
223
|
| `edenai/qwen/qwen-vl-plus` | 131K | | | | | | $0.21 | $0.63 |
|
|
@@ -301,7 +309,7 @@ const agent = new Agent({
|
|
|
301
309
|
name: "custom-agent",
|
|
302
310
|
model: {
|
|
303
311
|
url: "https://api.edenai.run/v3",
|
|
304
|
-
id: "edenai/amazon/
|
|
312
|
+
id: "edenai/amazon/amazon.nova-lite-v1:0",
|
|
305
313
|
apiKey: process.env.EDENAI_API_KEY,
|
|
306
314
|
headers: {
|
|
307
315
|
"X-Custom-Header": "value"
|
|
@@ -320,7 +328,7 @@ const agent = new Agent({
|
|
|
320
328
|
const useAdvanced = requestContext.task === "complex";
|
|
321
329
|
return useAdvanced
|
|
322
330
|
? "edenai/zai/glm-5v-turbo"
|
|
323
|
-
: "edenai/amazon/
|
|
331
|
+
: "edenai/amazon/amazon.nova-lite-v1:0";
|
|
324
332
|
}
|
|
325
333
|
});
|
|
326
334
|
```
|
|
@@ -45,7 +45,7 @@ for await (const chunk of stream) {
|
|
|
45
45
|
| `kilo/~deepseek/deepseek-v4-flash-latest` | 1.0M | | | | | | $0.05 | $0.16 |
|
|
46
46
|
| `kilo/~google/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
|
|
47
47
|
| `kilo/~google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
|
|
48
|
-
| `kilo/~moonshotai/kimi-latest` | 1.0M | | | | | | $
|
|
48
|
+
| `kilo/~moonshotai/kimi-latest` | 1.0M | | | | | | $2 | $12 |
|
|
49
49
|
| `kilo/~openai/gpt-latest` | 1.1M | | | | | | $2 | $10 |
|
|
50
50
|
| `kilo/~openai/gpt-mini-latest` | 400K | | | | | | $0.75 | $5 |
|
|
51
51
|
| `kilo/~x-ai/grok-latest` | 500K | | | | | | $2 | $6 |
|
|
@@ -369,7 +369,7 @@ for await (const chunk of stream) {
|
|
|
369
369
|
| `kilo/tencent/hy-mt2-1.8b` | 8K | | | | | | $0.04 | $0.18 |
|
|
370
370
|
| `kilo/tencent/hy-mt2-30b-a3b` | 8K | | | | | | $0.07 | $0.29 |
|
|
371
371
|
| `kilo/tencent/hy-mt2-7b` | 8K | | | | | | $0.07 | $0.29 |
|
|
372
|
-
| `kilo/tencent/hy3` | 262K | | | | | | $0.
|
|
372
|
+
| `kilo/tencent/hy3` | 262K | | | | | | $0.13 | $0.53 |
|
|
373
373
|
| `kilo/tencent/hy3-preview` | 262K | | | | | | $0.18 | $0.60 |
|
|
374
374
|
| `kilo/tencent/hy4-preview` | 1.0M | | | | | | $0.83 | $3 |
|
|
375
375
|
| `kilo/thedrummer/cydonia-24b-v4.1` | 131K | | | | | | $0.30 | $0.50 |
|
|
@@ -394,7 +394,7 @@ for await (const chunk of stream) {
|
|
|
394
394
|
| `kilo/z-ai/glm-4.5` | 131K | | | | | | $0.60 | $2 |
|
|
395
395
|
| `kilo/z-ai/glm-4.5-air` | 131K | | | | | | $0.13 | $0.85 |
|
|
396
396
|
| `kilo/z-ai/glm-4.5v` | 66K | | | | | | $0.60 | $2 |
|
|
397
|
-
| `kilo/z-ai/glm-4.6` |
|
|
397
|
+
| `kilo/z-ai/glm-4.6` | 198K | | | | | | $0.43 | $2 |
|
|
398
398
|
| `kilo/z-ai/glm-4.6v` | 131K | | | | | | $0.30 | $0.90 |
|
|
399
399
|
| `kilo/z-ai/glm-4.7` | 203K | | | | | | $0.40 | $2 |
|
|
400
400
|
| `kilo/z-ai/glm-4.7-flash` | 131K | | | | | | $0.06 | $0.40 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# NanoGPT
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 584 NanoGPT models through Mastra's model router. Authentication is handled automatically using the `NANO_GPT_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [NanoGPT documentation](https://docs.nano-gpt.com).
|
|
10
10
|
|
|
@@ -82,11 +82,6 @@ for await (const chunk of stream) {
|
|
|
82
82
|
| `nano-gpt/auto-model-basic` | 1.0M | | | | | | $10 | $20 |
|
|
83
83
|
| `nano-gpt/auto-model-premium` | 1.0M | | | | | | $10 | $20 |
|
|
84
84
|
| `nano-gpt/auto-model-standard` | 1.0M | | | | | | $10 | $20 |
|
|
85
|
-
| `nano-gpt/azure-gpt-4-turbo` | 128K | | | | | | $10 | $30 |
|
|
86
|
-
| `nano-gpt/azure-gpt-4o` | 128K | | | | | | $3 | $10 |
|
|
87
|
-
| `nano-gpt/azure-gpt-4o-mini` | 128K | | | | | | $0.15 | $0.60 |
|
|
88
|
-
| `nano-gpt/azure-o1` | 200K | | | | | | $15 | $60 |
|
|
89
|
-
| `nano-gpt/azure-o3-mini` | 200K | | | | | | $1 | $4 |
|
|
90
85
|
| `nano-gpt/baseten/Kimi-K2-Instruct-FP4` | 131K | | | | | | $0.40 | $2 |
|
|
91
86
|
| `nano-gpt/brave` | 8K | | | | | | $5 | $5 |
|
|
92
87
|
| `nano-gpt/brave-pro` | 8K | | | | | | $5 | $5 |
|
|
@@ -166,7 +166,7 @@ const result = await agent.generate('message for agent')
|
|
|
166
166
|
|
|
167
167
|
**options.modelSettings.frequencyPenalty** (`number`): Penalty for token frequency (-2 to 2). Reduces repetition of frequent tokens.
|
|
168
168
|
|
|
169
|
-
**options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured.
|
|
169
|
+
**options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured. Also accepts firstChunkMs, which only applies to streaming calls and is the maximum time the model may take to emit its first content-bearing chunk (text, reasoning, tool call, file or source; stream-start and metadata chunks do not count). A firstChunkMs timeout fails with timeoutType: 'firstChunk', behaves like stepMs for fallback, and is reset for each provider retry attempt. Nested timeout keys are merged across call-time and per-model settings.
|
|
170
170
|
|
|
171
171
|
**options.modelSettings.stopSequences** (`string[]`): Stop sequences. If set, the model will stop generating text when one of the stop sequences is generated.
|
|
172
172
|
|
|
@@ -102,7 +102,7 @@ await agent.network(`
|
|
|
102
102
|
|
|
103
103
|
**options.modelSettings.frequencyPenalty** (`number`): Penalty for token frequency (-2 to 2). Reduces repetition of frequent tokens.
|
|
104
104
|
|
|
105
|
-
**options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured.
|
|
105
|
+
**options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured. Also accepts firstChunkMs, which only applies to streaming calls and is the maximum time the model may take to emit its first content-bearing chunk (text, reasoning, tool call, file or source; stream-start and metadata chunks do not count). A firstChunkMs timeout fails with timeoutType: 'firstChunk', behaves like stepMs for fallback, and is reset for each provider retry attempt. Nested timeout keys are merged across call-time and per-model settings.
|
|
106
106
|
|
|
107
107
|
**options.modelSettings.stopSequences** (`string[]`): Stop sequences. If set, the model will stop generating text when one of the stop sequences is generated.
|
|
108
108
|
|
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
|
|
7
7
|
`buildBasePrompt()` builds the shared base system prompt for a coding agent: the behavioral instructions that make the agent a good coding assistant. It takes a `PromptContext` describing the current environment (project, platform, git branch, mode, model) and returns the prompt as a string.
|
|
8
8
|
|
|
9
|
-
Product-specific strings are parameterized through `productName`, `coAuthorName`, and `coAuthorEmail`, so you can rebrand the prompt and commit trailer without forking it.
|
|
9
|
+
Product-specific strings are parameterized through `productName`, `coAuthorName`, and `coAuthorEmail`, so you can rebrand the prompt and commit trailer without forking it. The co-author defaults to `"mastra-platform[bot]"` and `"284800079+mastra-platform[bot]@users.noreply.github.com"`.
|
|
10
10
|
|
|
11
11
|
Use this with [`createCodingAgent()`](https://mastra.ai/reference/coding-agent/create-coding-agent) when you build runtime-defined instructions for the agent.
|
|
12
12
|
|
|
@@ -58,7 +58,7 @@ const prompt = buildBasePrompt({
|
|
|
58
58
|
|
|
59
59
|
**mode** (`string`): Active agent mode (for example, "build" or "plan").
|
|
60
60
|
|
|
61
|
-
**modelId** (`string`): Identifier of the active model
|
|
61
|
+
**modelId** (`string`): Identifier of the active model.
|
|
62
62
|
|
|
63
63
|
**activePlan** (`{ title: string; plan: string; approvedAt: string } | null`): The currently approved plan, if any.
|
|
64
64
|
|
|
@@ -66,9 +66,9 @@ const prompt = buildBasePrompt({
|
|
|
66
66
|
|
|
67
67
|
**productName** (`string`): Display name used in the prompt header. (Default: `Mastra Code`)
|
|
68
68
|
|
|
69
|
-
**coAuthorName** (`string`): Name used in the commit Co-Authored-By line. (Default: `
|
|
69
|
+
**coAuthorName** (`string`): Name used in the commit Co-Authored-By line. (Default: `mastra-platform[bot]`)
|
|
70
70
|
|
|
71
|
-
**coAuthorEmail** (`string`): Email used in the commit Co-Authored-By line. (Default: `
|
|
71
|
+
**coAuthorEmail** (`string`): Email used in the commit Co-Authored-By line. (Default: `284800079+mastra-platform[bot]@users.noreply.github.com`)
|
|
72
72
|
|
|
73
73
|
## Returns
|
|
74
74
|
|
|
@@ -85,7 +85,7 @@ Final score: `(covered_elements / total_input_elements) * scale`
|
|
|
85
85
|
|
|
86
86
|
A completeness score between 0 and 1:
|
|
87
87
|
|
|
88
|
-
- **1.0**:
|
|
88
|
+
- **1.0**: Addresses all aspects of the query in detail.
|
|
89
89
|
- **0.7 to 0.9**: Covers most important aspects with good detail, minor gaps.
|
|
90
90
|
- **0.4 to 0.6**: Addresses some key points but missing important aspects or lacking detail.
|
|
91
91
|
- **0.1 to 0.3**: Only partially addresses the query with substantial gaps.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Faithfulness scorer
|
|
6
6
|
|
|
7
|
-
The `createFaithfulnessScorer()` function evaluates how factually accurate an LLM's output is compared to the provided context. It extracts claims from the output and verifies them against the context
|
|
7
|
+
The `createFaithfulnessScorer()` function evaluates how factually accurate an LLM's output is compared to the provided context. It extracts claims from the output and verifies them against the context. Use it to check whether a RAG response is supported by the retrieved information.
|
|
8
8
|
|
|
9
9
|
## Parameters
|
|
10
10
|
|
|
@@ -136,7 +136,7 @@ Each dimension receives an impact level with corresponding weights:
|
|
|
136
136
|
- **Minimal (0.85)**: Slight phrasing changes but maintains correctness
|
|
137
137
|
- **Moderate (0.6)**: Noticeable changes affecting quality but core info correct
|
|
138
138
|
- **Significant (0.3)**: Major degradation in quality or accuracy
|
|
139
|
-
- **Severe (0.1)**: Response
|
|
139
|
+
- **Severe (0.1)**: Response much worse than the baseline or unrelated to the original query
|
|
140
140
|
|
|
141
141
|
### Conservative Scoring
|
|
142
142
|
|
|
@@ -180,10 +180,8 @@ Final Score = Weighted Score × scale
|
|
|
180
180
|
|
|
181
181
|
**Both Mode (`'both'`)** - Use when (default, recommended):
|
|
182
182
|
|
|
183
|
-
-
|
|
184
|
-
- Balancing user satisfaction with system compliance
|
|
183
|
+
- Evaluating whether responses satisfy user requests while following system instructions
|
|
185
184
|
- Production monitoring where both user and system requirements matter
|
|
186
|
-
- Holistic assessment of prompt-response alignment
|
|
187
185
|
|
|
188
186
|
## Common use cases
|
|
189
187
|
|
|
@@ -423,7 +421,7 @@ console.log(result)
|
|
|
423
421
|
|
|
424
422
|
### Excellent alignment output
|
|
425
423
|
|
|
426
|
-
|
|
424
|
+
In this example, the response receives a high score because it provides the requested Python factorial function with error handling for negative numbers.
|
|
427
425
|
|
|
428
426
|
```typescript
|
|
429
427
|
{
|
|
@@ -51,7 +51,7 @@ OM performs thresholding with fast local token estimation. Text uses `tokenx`, a
|
|
|
51
51
|
|
|
52
52
|
**retrieval** (`boolean | { vector?: boolean; scope?: 'thread' | 'resource'; instructions?: string }`): Let the agent look up the raw message history behind its observations. Observation groups keep durable pointers to the original messages, and a recall tool is registered so the agent can browse them. true enables cross-thread browsing by default. { vector: true } also enables semantic search using Memory's vector store and embedder. { scope: 'thread' } restricts the recall tool to the current thread only. Default scope is 'resource'. { instructions: '...' } appends application-specific recall guidance after Mastra's built-in retrieval instructions. (Default: `false`)
|
|
53
53
|
|
|
54
|
-
**hooks** (`ObserveHooks`): Lifecycle hooks fired for every observation/reflection cycle — the manual observe()/reflect() APIs, turn-driven synchronous observation, and fire-and-forget async buffering. Callbacks receive threadId/resourceId/trigger call context ('manual' | 'turn-sync' | 'async-buffer'), and the end hooks (onObservationEnd/onReflectionEnd) additionally receive the OM model call's token usage and providerMetadata (where providers such as the AI Gateway report per-call cost), so apps can account for OM model spend without wrapping the observer/reflector models in middleware.
|
|
54
|
+
**hooks** (`ObserveHooks`): Lifecycle hooks fired for every observation/reflection cycle — the manual observe()/reflect() APIs, turn-driven synchronous observation, and fire-and-forget async buffering. Callbacks receive threadId/resourceId/trigger call context ('manual' | 'turn-sync' | 'async-buffer'), and the end hooks (onObservationEnd/onReflectionEnd) additionally receive the OM model call's token usage and providerMetadata (where providers such as the AI Gateway report per-call cost), so apps can account for OM model spend without wrapping the observer/reflector models in middleware. Config-level lifecycle hooks are non-blocking by default: their errors are logged without failing the cycle. With hookExecution: 'await', config-level lifecycle errors can reject observe() or reflect(). Per-call observe({ hooks }) lifecycle hooks are also awaited and can reject observe(). A failing awaited start hook skips the model call; an end-hook failure can reject the call after its work has completed. Async-buffered cycles always use non-blocking config-level lifecycle hooks, regardless of hookExecution. Their cycle failures are reported through onObservationEnd.error or onReflectionEnd.error, rather than thrown to the caller. ObserveHooks also accepts transform hooks that intercept and can replace cycle data: beforeObservation({ messages, ...context }) runs on the messages about to be observed (return { messages } to filter/redact; an empty array skips the Observer call), afterObservation({ observations, ...context }) runs on the Observer's output before it is persisted, beforeReflection({ observations, ...context }) runs on the text sent to the Reflector, and afterReflection({ observations, ...context }) runs on the Reflector's output before it is persisted. Transform hooks return void to pass data through unchanged and are always awaited on every path. A thrown error fails the cycle before committing its transformed observation or reflection text, but doesn't roll back extractor callbacks or other side effects that already ran. The after hooks replace only text, not the separate structured extractor results stored in thread metadata. Reflection extraction and its callbacks run before afterReflection; this hook neither recomputes extracted values nor reruns callbacks. After hooks aren't a redaction boundary for all cycle data.
|
|
55
55
|
|
|
56
56
|
**observation** (`ObservationalMemoryObservationConfig`): Configuration for the observation step. Controls when the Observer agent runs and how it behaves.
|
|
57
57
|
|
|
@@ -217,7 +217,7 @@ const run = await workflow.createRun();
|
|
|
217
217
|
|
|
218
218
|
### `setState()` is now async and the data passed is validated
|
|
219
219
|
|
|
220
|
-
The `setState()` function is now async
|
|
220
|
+
The `setState()` function is now async and validates data against the step's `stateSchema` when `validateInputs` is enabled. Pass only the fields you want to update, without spreading the previous state (`...state`).
|
|
221
221
|
|
|
222
222
|
To migrate, update the `setState()` function to be async.
|
|
223
223
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# RegexFilterProcessor
|
|
6
6
|
|
|
7
|
-
The `RegexFilterProcessor`
|
|
7
|
+
The `RegexFilterProcessor` uses regex pattern matching to filter, redact, or block content in agent messages. No LLM calls are made.
|
|
8
8
|
|
|
9
9
|
Supports built-in presets for common patterns (PII, secrets, URLs) and custom regex rules. Can be applied to input, output, or both phases.
|
|
10
10
|
|
|
@@ -650,7 +650,7 @@ Key metadata considerations:
|
|
|
650
650
|
|
|
651
651
|
## Deleting vectors
|
|
652
652
|
|
|
653
|
-
|
|
653
|
+
Use `deleteVectors` with a metadata filter to remove embeddings associated with a document. This is useful for cleaning up stale vectors after a document is deleted or updated.
|
|
654
654
|
|
|
655
655
|
### Delete by Metadata Filter
|
|
656
656
|
|
|
@@ -160,7 +160,7 @@ const stream = await agent.stream('message for agent')
|
|
|
160
160
|
|
|
161
161
|
**options.modelSettings.frequencyPenalty** (`number`): Penalty for token frequency (-2 to 2). Reduces repetition of frequent tokens.
|
|
162
162
|
|
|
163
|
-
**options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured.
|
|
163
|
+
**options.modelSettings.timeout** (`object`): Time-based execution budget for the run. Accepts totalMs, the maximum duration of the entire agent run across every loop iteration, tool call and retry, and stepMs, the maximum duration of a single model call including the time spent consuming its stream. Exceeding either budget fails with a MastraTimeoutError. A totalMs timeout ends the run and does not try fallback models, because it is a hard deadline for the whole run. A stepMs timeout is not retried against the same model but does advance to the next entry in models when fallback models are configured. Also accepts firstChunkMs, which only applies to streaming calls and is the maximum time the model may take to emit its first content-bearing chunk (text, reasoning, tool call, file or source; stream-start and metadata chunks do not count). A firstChunkMs timeout fails with timeoutType: 'firstChunk', behaves like stepMs for fallback, and is reset for each provider retry attempt. Nested timeout keys are merged across call-time and per-model settings.
|
|
164
164
|
|
|
165
165
|
**options.modelSettings.stopSequences** (`string[]`): Stop sequences. If set, the model will stop generating text when one of the stop sequences is generated.
|
|
166
166
|
|
|
@@ -268,7 +268,7 @@ const fullText = await stream.text
|
|
|
268
268
|
|
|
269
269
|
### Limiting execution time
|
|
270
270
|
|
|
271
|
-
Use `modelSettings.timeout` to bound how long a run may take. `totalMs` limits the entire run, including every loop iteration, tool call and retry. `stepMs` limits a single model call, covering both establishing the stream and consuming it.
|
|
271
|
+
Use `modelSettings.timeout` to bound how long a run may take. `totalMs` limits the entire run, including every loop iteration, tool call and retry. `stepMs` limits a single model call, covering both establishing the stream and consuming it. `firstChunkMs` only applies to streaming calls and limits how long the model may take to emit its first content-bearing chunk.
|
|
272
272
|
|
|
273
273
|
```ts
|
|
274
274
|
const stream = await agent.stream('Tell me a story', {
|
|
@@ -276,15 +276,19 @@ const stream = await agent.stream('Tell me a story', {
|
|
|
276
276
|
timeout: {
|
|
277
277
|
totalMs: 30000, // fail the run if it takes longer than 30s
|
|
278
278
|
stepMs: 10000, // fail an individual model call after 10s
|
|
279
|
+
firstChunkMs: 3000, // fail a model call that produces no content within 3s
|
|
279
280
|
},
|
|
280
281
|
},
|
|
281
282
|
})
|
|
282
283
|
```
|
|
283
284
|
|
|
284
|
-
Exceeding
|
|
285
|
+
Exceeding a budget fails with a `MastraTimeoutError`, which carries a `timeoutType` of `'total'`, `'step'` or `'firstChunk'`. Each budget behaves differently when the agent is configured with fallback [`models`](https://mastra.ai/reference/agents/agent):
|
|
285
286
|
|
|
286
287
|
- A `totalMs` timeout ends the run immediately and doesn't try the next model, because it's a hard deadline for the run as a whole.
|
|
287
288
|
- A `stepMs` timeout isn't retried against the same model, but does advance to the next model, which makes it a way to fail over from a slow provider.
|
|
289
|
+
- A `firstChunkMs` timeout behaves like `stepMs`: it isn't retried against the same model, but does advance to the next model. Stream-start and metadata chunks don't satisfy the budget; only the first text, reasoning, tool call, file or source chunk does. Once content begins, only `stepMs` and `totalMs` remain active. Each provider retry attempt gets a fresh `firstChunkMs` budget.
|
|
290
|
+
|
|
291
|
+
Nested `timeout` settings are merged per key across call-time and per-model `modelSettings`, so a per-model `stepMs` override keeps a call-time `firstChunkMs`.
|
|
288
292
|
|
|
289
293
|
### AI SDK v5+ Format
|
|
290
294
|
|
|
@@ -646,7 +646,7 @@ async executeTool(
|
|
|
646
646
|
|
|
647
647
|
### What are MCP Resources?
|
|
648
648
|
|
|
649
|
-
|
|
649
|
+
MCP resources expose server data that clients can read and use as context for LLM interactions. Examples include:
|
|
650
650
|
|
|
651
651
|
- File contents
|
|
652
652
|
- Database records
|
|
@@ -35,7 +35,7 @@ A sandbox constructed with a known `id` resolves that id on `start()`: reconnect
|
|
|
35
35
|
|
|
36
36
|
Providers that predate the contract return `void`, which the base class treats as unknown.
|
|
37
37
|
|
|
38
|
-
|
|
38
|
+
Providers implement the start lifecycle using one of the following approaches, listed in order of preference. The base class handles concurrent start calls and status management, as well as the `onStart` hook and mount processing:
|
|
39
39
|
|
|
40
40
|
1. **Acquisition primitives**: implement protected `find()` (side-effect-free lookup by logical id, returning a provider-native handle or `undefined`), `connect(handle)` (wake/resume/adopt), and `create()` (provision fresh) without overriding `start()`. The base orchestrates find → connect → `{ outcome: 'connected' }`, else create → `{ outcome: 'created' }`. The outcome is derived structurally from which branch ran. Used by `E2BSandbox`, `DaytonaSandbox`, and `LocalSandbox`.
|
|
41
41
|
2. **`start()` override returning `SandboxStartResult`**: for providers whose API is a fused get-or-create where decomposition would add round-trips (`PlatformSandbox`, `RailwaySandbox`).
|
|
@@ -47,7 +47,7 @@ The result is also forwarded to the `onStart` lifecycle hook as `{ sandbox, outc
|
|
|
47
47
|
|
|
48
48
|
### `onStart` (constructor option)
|
|
49
49
|
|
|
50
|
-
`onStart` runs inside the start lifecycle, after the sandbox reaches `running` status and before pending mounts are processed. It fires on every start regardless of trigger, whether an explicit call, a lazy `ensureRunning()` from a command, or a revival after the provider replaced the VM.
|
|
50
|
+
`onStart` runs inside the start lifecycle, after the sandbox reaches `running` status and before pending mounts are processed. It fires on every start regardless of trigger, whether an explicit call, a lazy `ensureRunning()` from a command, or a revival after the provider replaced the VM. Use this hook for once-per-VM setup: check `outcome` to decide whether to run setup or check that it already completed.
|
|
51
51
|
|
|
52
52
|
```typescript
|
|
53
53
|
new E2BSandbox({
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@mastra/mcp-docs-server",
|
|
3
|
-
"version": "1.2.
|
|
3
|
+
"version": "1.2.25-alpha.1",
|
|
4
4
|
"description": "MCP server for accessing Mastra.ai documentation, changelogs, and news.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "dist/index.js",
|
|
@@ -27,7 +27,7 @@
|
|
|
27
27
|
"jsdom": "^26.1.0",
|
|
28
28
|
"local-pkg": "^1.1.2",
|
|
29
29
|
"zod": "^4.4.3",
|
|
30
|
-
"@mastra/core": "1.
|
|
30
|
+
"@mastra/core": "1.66.0-alpha.0",
|
|
31
31
|
"@mastra/mcp": "^1.17.3"
|
|
32
32
|
},
|
|
33
33
|
"devDependencies": {
|
|
@@ -45,7 +45,7 @@
|
|
|
45
45
|
"typescript": "^7.0.2",
|
|
46
46
|
"vitest": "4.1.10",
|
|
47
47
|
"@internal/lint": "0.0.131",
|
|
48
|
-
"@mastra/core": "1.
|
|
48
|
+
"@mastra/core": "1.66.0-alpha.0",
|
|
49
49
|
"@internal/types-builder": "0.0.106"
|
|
50
50
|
},
|
|
51
51
|
"homepage": "https://mastra.ai",
|