@mastra/mcp-docs-server 1.2.25 → 1.2.26-alpha.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.docs/docs/deployment/monorepo.md +12 -0
- package/.docs/docs/evals/datasets.md +5 -1
- package/.docs/docs/memory/message-history.md +1 -1
- package/.docs/docs/subagents.md +25 -0
- package/.docs/integrations/file-storage/amazon-s3.md +7 -1
- package/.docs/integrations/sandboxes/daytona.md +33 -0
- package/.docs/integrations/voice/livekit.md +26 -2
- package/.docs/models/gateways/merge-gateway.md +2 -1
- package/.docs/models/gateways/netlify.md +2 -1
- package/.docs/models/gateways/openrouter.md +2 -1
- package/.docs/models/gateways/vercel.md +1 -1
- package/.docs/models/index.md +1 -1
- package/.docs/models/providers/above.md +1 -2
- package/.docs/models/providers/alibaba-token-plan-cn.md +0 -1
- package/.docs/models/providers/alibaba-token-plan.md +0 -1
- package/.docs/models/providers/bothub.md +9 -4
- package/.docs/models/providers/deepinfra.md +2 -4
- package/.docs/models/providers/deepseek.md +7 -6
- package/.docs/models/providers/edenai.md +5 -4
- package/.docs/models/providers/fireworks-ai.md +3 -1
- package/.docs/models/providers/huggingface.md +2 -1
- package/.docs/models/providers/hyper.md +5 -4
- package/.docs/models/providers/inception.md +2 -1
- package/.docs/models/providers/kilo.md +7 -6
- package/.docs/models/providers/llmgateway-providers.md +6 -2
- package/.docs/models/providers/llmgateway.md +1 -1
- package/.docs/models/providers/nano-gpt.md +9 -11
- package/.docs/models/providers/ofox.md +23 -1
- package/.docs/models/providers/opencode-go.md +4 -4
- package/.docs/models/providers/requesty.md +2 -1
- package/.docs/models/providers/scnet-token-plan.md +4 -6
- package/.docs/reference/agents/agent.md +16 -0
- package/.docs/reference/agents/generate.md +2 -0
- package/.docs/reference/client-js/datasets.md +1 -1
- package/.docs/reference/datasets/purgeItem.md +3 -3
- package/.docs/reference/index.md +1 -0
- package/.docs/reference/memory/recall.md +51 -0
- package/.docs/reference/processors/agents-md-injector.md +55 -0
- package/.docs/reference/streaming/agents/stream.md +30 -0
- package/package.json +4 -4
|
@@ -118,6 +118,18 @@ Keep dependencies consistent to avoid version conflicts and build errors:
|
|
|
118
118
|
- Use a **single lockfile** at the monorepo root so all packages resolve the same versions
|
|
119
119
|
- Align versions of **shared libraries** (like Mastra or frameworks) to prevent duplicates
|
|
120
120
|
|
|
121
|
+
### Skipping dependency installation
|
|
122
|
+
|
|
123
|
+
By default `mastra build` installs dependencies into the `.mastra/output` directory and generates a `package-lock.json` for the deploy target. Hermetic build systems (such as Bazel) supply `node_modules` externally and run in a network-less sandbox, where this install is both redundant and fatal.
|
|
124
|
+
|
|
125
|
+
Set `MASTRA_BUILD_SKIP_INSTALL` to `true` or `1` to skip the output install and lockfile generation:
|
|
126
|
+
|
|
127
|
+
```bash
|
|
128
|
+
MASTRA_BUILD_SKIP_INSTALL=1 mastra build
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
When enabled, you are responsible for providing `node_modules` in the output directory yourself.
|
|
132
|
+
|
|
121
133
|
## Troubleshooting
|
|
122
134
|
|
|
123
135
|
### Workspace package not found
|
|
@@ -142,7 +142,11 @@ Deleting an item hides it from the current dataset version but retains its conte
|
|
|
142
142
|
await dataset.purgeItem({ itemId: 'item-abc-123' })
|
|
143
143
|
```
|
|
144
144
|
|
|
145
|
-
Purging replaces content in every historical row and deletion tombstone
|
|
145
|
+
Purging replaces content in every historical row and deletion tombstone with redacted values, and scrubs linked experiment-result payloads, tags, and comments. Later experiment-result submissions for the item are also stored with redacted content.
|
|
146
|
+
|
|
147
|
+
Purge serializes or conflicts with concurrent dataset item writers without guaranteeing which operation completes first. If a mutating `updateItem()` call loses the race, it re-reads the purge marker and rejects with `DATASET_ITEM_PURGED`. `deleteItem()` remains idempotent, and any deletion tombstone created during the race stays redacted.
|
|
148
|
+
|
|
149
|
+
Normal item mutations use Slowly Changing Dimension Type 2 (SCD-2) versioning. Permanent purge intentionally overrides historical immutability for erasure while preserving item identity and the dataset version timeline. It doesn't create a dataset version and can't be undone. Experiment counters and review status are also preserved. Avoid storing sensitive data in `externalId`, which remains unchanged as the item's identity key.
|
|
146
150
|
|
|
147
151
|
MongoDB storage requires a replica set or sharded deployment with transaction support for this operation. Purging fails before changing data when MongoDB transactions aren't available.
|
|
148
152
|
|
|
@@ -232,7 +232,7 @@ const thread = await memory.getThreadById({ threadId: 'thread-123' })
|
|
|
232
232
|
|
|
233
233
|
Once you have a thread, use [`recall()`](https://mastra.ai/reference/memory/recall) to retrieve its messages. It supports pagination and [semantic search](https://mastra.ai/docs/memory/semantic-recall), with optional date filtering.
|
|
234
234
|
|
|
235
|
-
|
|
235
|
+
Fetch a thread's history without pagination. Recall hides reminder signals by default; pass `hideSignals: false` to include them, `true` to hide all recognized signals, or an array to omit selected types. See [signal visibility and compatibility](https://mastra.ai/reference/memory/recall) for matching rules and precedence.
|
|
236
236
|
|
|
237
237
|
```typescript
|
|
238
238
|
const { messages } = await memory.recall({
|
package/.docs/docs/subagents.md
CHANGED
|
@@ -240,6 +240,31 @@ await parentAgent.generate('Research AI trends', {
|
|
|
240
240
|
})
|
|
241
241
|
```
|
|
242
242
|
|
|
243
|
+
### Reusing an earlier subagent result
|
|
244
|
+
|
|
245
|
+
By default, subagent results reach the parent agent as tool results, which are stripped from the context forwarded to later subagents. The parent agent must restate an earlier result in the next delegation prompt, which costs tokens and loses detail.
|
|
246
|
+
|
|
247
|
+
Set `enableResultReferences` to let a later delegation reuse an earlier result verbatim:
|
|
248
|
+
|
|
249
|
+
```typescript
|
|
250
|
+
await parentAgent.generate('Find and fix the token refresh bug', {
|
|
251
|
+
delegation: {
|
|
252
|
+
enableResultReferences: true,
|
|
253
|
+
},
|
|
254
|
+
})
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
When enabled:
|
|
258
|
+
|
|
259
|
+
- Each successful, non-empty subagent result gets a reference ID such as `explorer-1`. The parent agent's model sees it as a `[ref: explorer-1]` line after the subagent's text.
|
|
260
|
+
- The delegation tools gain a `contextFromRefs` input. The parent agent can pass earlier IDs, either as strings (`["explorer-1"]`) or as objects with an optional label and note (`[{ ref: "explorer-1", as: "investigation", note: "bug location" }]`).
|
|
261
|
+
- The referenced text is inserted before the delegation prompt, each result in its own labeled block, exactly as the earlier subagent produced it. `onDelegationStart` and `messageFilter` receive the expanded prompt.
|
|
262
|
+
- If `onDelegationComplete` returns `resultText`, the replaced text is what later delegations receive.
|
|
263
|
+
|
|
264
|
+
References are held in memory for a single parent agent run and aren't persisted. Rejected, failed, empty, and background-task delegations don't receive a reference ID. Unknown IDs are skipped with a warning and the delegation continues.
|
|
265
|
+
|
|
266
|
+
Referenced text is output from another agent. Each block uses a fresh, unpredictable tag and tells the receiving subagent to treat the contents as data, but if subagents handle untrusted input, add your own checks in `onDelegationStart` or through processors.
|
|
267
|
+
|
|
243
268
|
## Iteration monitoring
|
|
244
269
|
|
|
245
270
|
`onIterationComplete` is called after each iteration of the parent agent's loop. Use it to monitor execution or guide the next iteration. You can also stop execution early.
|
|
@@ -111,7 +111,13 @@ const filesystem = new S3Filesystem({
|
|
|
111
111
|
})
|
|
112
112
|
```
|
|
113
113
|
|
|
114
|
-
Provider functions only apply to `S3Filesystem` API calls. When mounting the filesystem into an E2B sandbox, mount configuration only supports static `accessKeyId`, `secretAccessKey`, and `sessionToken` values, so credential refresh must be handled outside the mount.
|
|
114
|
+
Provider functions only apply to `S3Filesystem` API calls. When mounting the filesystem into an E2B or Daytona sandbox, mount configuration only supports static `accessKeyId`, `secretAccessKey`, and `sessionToken` values, so credential refresh must be handled outside the mount. See [temporary credentials in Daytona](https://mastra.ai/integrations/sandboxes/daytona) for mount lifetime and isolation requirements.
|
|
115
|
+
|
|
116
|
+
### Prefix-scoped permissions
|
|
117
|
+
|
|
118
|
+
When `prefix` is set, initialization calls `ListObjectsV2` with the normalized prefix, including its trailing `/`, and `MaxKeys: 1`. Credentials must permit listing that prefix. Without a prefix, initialization uses `HeadBucket`. These checks verify access, not whether a directory exists.
|
|
119
|
+
|
|
120
|
+
The prefix limits which keys the filesystem addresses; it isn't an authorization boundary. For isolation, use credentials whose storage-provider policy restricts access to that prefix. Keep parent credentials on your backend. Read and write permissions alone aren't sufficient for prefixed filesystem initialization.
|
|
115
121
|
|
|
116
122
|
### Cloudflare R2
|
|
117
123
|
|
|
@@ -200,6 +200,39 @@ const workspace = new Workspace({
|
|
|
200
200
|
|
|
201
201
|
When the workspace starts, the filesystems are automatically mounted at the specified paths. Code running in the sandbox can then access files at `/s3-data` and `/gcs-data` as if they were local directories.
|
|
202
202
|
|
|
203
|
+
#### Temporary S3 credentials
|
|
204
|
+
|
|
205
|
+
Pass all three credential values to `S3Filesystem` when using temporary credentials. Daytona forwards the session token to s3fs:
|
|
206
|
+
|
|
207
|
+
```typescript
|
|
208
|
+
import { Workspace } from '@mastra/core/workspace'
|
|
209
|
+
import { DaytonaSandbox } from '@mastra/daytona'
|
|
210
|
+
import { S3Filesystem } from '@mastra/s3'
|
|
211
|
+
|
|
212
|
+
const workspace = new Workspace({
|
|
213
|
+
mounts: {
|
|
214
|
+
'/s3-data': new S3Filesystem({
|
|
215
|
+
bucket: process.env.S3_BUCKET!,
|
|
216
|
+
region: process.env.S3_REGION ?? 'us-east-1',
|
|
217
|
+
endpoint: process.env.S3_ENDPOINT,
|
|
218
|
+
prefix: 'resource-123/thread-456/',
|
|
219
|
+
accessKeyId: process.env.SCOPED_S3_ACCESS_KEY_ID!,
|
|
220
|
+
secretAccessKey: process.env.SCOPED_S3_SECRET_ACCESS_KEY!,
|
|
221
|
+
sessionToken: process.env.SCOPED_S3_SESSION_TOKEN!,
|
|
222
|
+
}),
|
|
223
|
+
},
|
|
224
|
+
sandbox: new DaytonaSandbox({ language: 'python', ephemeral: true }),
|
|
225
|
+
})
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
Before starting the workspace, obtain credentials from your storage provider that authorize only the intended prefix. Credential issuance is provider-specific; Mastra doesn't mint or restrict credentials. A `prefix` selects a directory but doesn't enforce authorization. Authenticate each request, authorize its resource and thread, and use separate sandboxes for separate security scopes.
|
|
229
|
+
|
|
230
|
+
Temporary credentials are loaded when the mount starts and aren't automatically refreshed. Host-side credential provider functions don't refresh the mount. Keep runs within the credential lifetime and create a new sandbox with fresh credentials for later runs. Reconnecting to a sandbox doesn't renew its credentials.
|
|
231
|
+
|
|
232
|
+
Credentials are uploaded into owner-only files inside private directories. After launching s3fs, Daytona removes the temporary-credential staging file and directory. The daemon retains the credentials in its environment, which sandbox code running as the same user or root can still read. Never supply broader credentials than the sandbox needs. Mounts using long-lived credentials retain their password files while s3fs needs them. Unmount and reconnect cleanup remove those files after verifying that the daemon has exited. If a daemon is still active, including after a mount is moved aside, or its status can't be checked, the files are retained and a warning is logged. Cleanup on a later unmount of the same path retries removal; deleting the sandbox removes any remaining files.
|
|
233
|
+
|
|
234
|
+
Prefixed filesystems require permission to list their prefix during initialization. With s3fs, a prefixed mount may also require a zero-byte object at the exact `<prefix>/` key. Provision that directory marker before mounting. After launching s3fs, Daytona checks the mounted directory's metadata without listing its contents, with a 15-second timeout and forced termination after another 5 seconds. A failed check reports a mount failure. Cleanup attempts to unmount the failed mount, moving a stuck mount aside if necessary to free the original path for retry. Moved stale mounts may remain until sandbox deletion. This startup check doesn't guarantee read or write access to individual files or ongoing daemon health; verify a read and write through the mounted path.
|
|
235
|
+
|
|
203
236
|
#### Via `sandbox.mount()`
|
|
204
237
|
|
|
205
238
|
Mount manually at any point after the sandbox has started:
|
|
@@ -205,6 +205,28 @@ export default createLiveKitWorker({
|
|
|
205
205
|
|
|
206
206
|
`configuration.stt` works the same way for per-call transcription, for example a different transcription model or language per tenant. The greeting has a matching per-call form: `configuration.greeting.text` accepts a resolver with the same call context, so one worker can open with each tenant's own phrasing.
|
|
207
207
|
|
|
208
|
+
### Per-call turn detection
|
|
209
|
+
|
|
210
|
+
LiveKit's `TurnDetector` classes read the job's inference executor when constructed, so they can only be created inside a LiveKit job, not at module scope where the worker options live. To use one, set the `configuration.turnDetection` resolver. It runs once per call with the same call context as `configuration.stt` and returns anything the top-level `turnDetection` option accepts. Return `undefined` to fall back to the top-level option.
|
|
211
|
+
|
|
212
|
+
```typescript
|
|
213
|
+
import { turnDetector } from '@livekit/agents-plugin-livekit'
|
|
214
|
+
|
|
215
|
+
export default createLiveKitWorker({
|
|
216
|
+
mastra,
|
|
217
|
+
agent: 'support',
|
|
218
|
+
stt: 'deepgram/nova-3',
|
|
219
|
+
tts: 'cartesia/sonic-3',
|
|
220
|
+
turnDetection: 'multilingual',
|
|
221
|
+
configuration: {
|
|
222
|
+
// Constructed inside the job, where the inference executor is available.
|
|
223
|
+
turnDetection: () => new turnDetector.MultilingualModel(0.2),
|
|
224
|
+
},
|
|
225
|
+
})
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
The semantic model's inference runners must be registered before the agent server boots, so when this resolver is set the worker imports `@livekit/agents-plugin-livekit` up front. Keep the top-level `turnDetection` set to `'multilingual'` or `'english'` to pre-register only that model; otherwise both stay available.
|
|
229
|
+
|
|
208
230
|
### Memory and threads
|
|
209
231
|
|
|
210
232
|
When the resolved Mastra agent has memory configured, each call becomes one memory thread:
|
|
@@ -539,7 +561,7 @@ if (process.argv[1] === fileURLToPath(import.meta.url)) {
|
|
|
539
561
|
|
|
540
562
|
**vad** (`VAD | 'silero' | false`): Voice activity detection. 'silero' loads the Silero VAD from @livekit/agents-plugin-silero during prewarm. Pass an instance to bring your own, or false to disable. (Default: `'silero'`)
|
|
541
563
|
|
|
542
|
-
**turnDetection** (`'multilingual' | 'english' | TurnDetectionMode`): End-of-turn detection. 'multilingual' and 'english' load LiveKit's semantic turn detector from @livekit/agents-plugin-livekit. Other values such as 'vad', 'stt', or 'manual' pass through.
|
|
564
|
+
**turnDetection** (`'multilingual' | 'english' | TurnDetectionMode`): End-of-turn detection. 'multilingual' and 'english' load LiveKit's semantic turn detector from @livekit/agents-plugin-livekit. Other values such as 'vad', 'stt', or 'manual' pass through. To construct a TurnDetector instance per call, set the configuration.turnDetection resolver — it takes precedence, with this option as the fallback.
|
|
543
565
|
|
|
544
566
|
**turnHandling** (`Partial<TurnHandlingOptions>`): Turn handling tuning: endpointing delays, interruption sensitivity, preemptive generation. The worker disables preemptiveGeneration unless set here — each preemptive attempt re-runs the Mastra agent and persists partial user and assistant messages unless memory.options.readOnly is set.
|
|
545
567
|
|
|
@@ -551,7 +573,7 @@ if (process.argv[1] === fileURLToPath(import.meta.url)) {
|
|
|
551
573
|
|
|
552
574
|
**onTurnComplete** (`(ctx: VoiceTurnCompleteContext) => void | Promise<void>`): Called once per turn after the reply finished streaming to text-to-speech. Runs off the audio path and is not awaited. The context carries the produced reply (text, toolCalls, interrupted, usage) and the resolved memory mapping.
|
|
553
575
|
|
|
554
|
-
**configuration** (`LiveKitWorkerConfiguration`): Grouped conversation and compliance configuration: the opening greeting and AI disclosure, consent requirements, agent-initiated hang-up, and per-call STT/TTS selection.
|
|
576
|
+
**configuration** (`LiveKitWorkerConfiguration`): Grouped conversation and compliance configuration: the opening greeting and AI disclosure, consent requirements, agent-initiated hang-up, and per-call STT/TTS/turn detection selection.
|
|
555
577
|
|
|
556
578
|
**configuration.greeting** (`GreetingConfiguration`): The opening greeting and AI disclosure: text (a fixed string or a per-call resolver for per-tenant greetings), allowInterruptions, awaitPlayout, persist, and periodic re-disclosure via repeatEvery and repeatText.
|
|
557
579
|
|
|
@@ -563,6 +585,8 @@ if (process.argv[1] === fileURLToPath(import.meta.url)) {
|
|
|
563
585
|
|
|
564
586
|
**configuration.tts** (`(context: VoiceCallContext) => TTS | string | undefined`): Per-call text-to-speech: a resolver invoked once per call (post-connect) with { metadata, requestContext, roomName, ctx }, returning anything the top-level tts option accepts — one voice or language per tenant. Return undefined to fall back to the top-level tts. Cache plugin instances across calls.
|
|
565
587
|
|
|
588
|
+
**configuration.turnDetection** (`(context: VoiceCallContext) => TurnDetectionMode | undefined`): Per-call end-of-turn detection: a resolver invoked once per call (post-connect, inside the LiveKit job) with { metadata, requestContext, roomName, ctx }, returning anything the top-level turnDetection option accepts. Use it to construct LiveKit TurnDetector instances, which need the job's inference executor. Return undefined to fall back to the top-level turnDetection.
|
|
589
|
+
|
|
566
590
|
**greeting** (`string`): Static greeting spoken when the session starts. Deprecated: prefer configuration.greeting.text.
|
|
567
591
|
|
|
568
592
|
**persistGreeting** (`boolean`): Save the spoken greeting to the memory thread as an assistant message, making the saved thread a faithful call transcript. Only applies when a greeting is set and memory is enabled. Deprecated: prefer configuration.greeting.persist. (Default: `true`)
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Merge Gateway
|
|
6
6
|
|
|
7
|
-
Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
7
|
+
Merge Gateway aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 182 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Merge Gateway documentation](https://docs.merge.dev/merge-gateway).
|
|
10
10
|
|
|
@@ -72,6 +72,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
72
72
|
| `deepseek/deepseek-v4-pro` |
|
|
73
73
|
| `deepseek/deepseek-v4-pro-0423` |
|
|
74
74
|
| `deepseek/deepseek-v4-pro-0813` |
|
|
75
|
+
| `deepseek/deepseek-v4.1-flash` |
|
|
75
76
|
| `google/gemini-2.5-computer-use-preview-10-2025` |
|
|
76
77
|
| `google/gemini-2.5-flash` |
|
|
77
78
|
| `google/gemini-2.5-flash-image` |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Netlify
|
|
6
6
|
|
|
7
|
-
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access
|
|
7
|
+
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 252 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Netlify documentation](https://docs.netlify.com/build/ai-gateway/overview/).
|
|
10
10
|
|
|
@@ -161,6 +161,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
161
161
|
| `openrouter/inclusionai/ling-3.0-flash-fin` |
|
|
162
162
|
| `openrouter/inclusionai/ling-3.0-flash-fin:free` |
|
|
163
163
|
| `openrouter/inclusionai/ling-3.0-flash-sante:free` |
|
|
164
|
+
| `openrouter/inclusionai/ling-3.0-flash-vl:free` |
|
|
164
165
|
| `openrouter/mancer/weaver` |
|
|
165
166
|
| `openrouter/meta-llama/llama-3.1-70b-instruct` |
|
|
166
167
|
| `openrouter/meta-llama/llama-3.1-8b-instruct` |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# OpenRouter
|
|
6
6
|
|
|
7
|
-
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
7
|
+
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 359 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
|
|
10
10
|
|
|
@@ -147,6 +147,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
147
147
|
| `inclusionai/ling-3.0-flash-fin` |
|
|
148
148
|
| `inclusionai/ling-3.0-flash-fin:free` |
|
|
149
149
|
| `inclusionai/ling-3.0-flash-sante:free` |
|
|
150
|
+
| `inclusionai/ling-3.0-flash-vl:free` |
|
|
150
151
|
| `kwaipilot/kat-coder-pro-v2` |
|
|
151
152
|
| `kwaipilot/kat-coder-pro-v2.5` |
|
|
152
153
|
| `liquid/lfm-2.5-2.6b:free` |
|
|
@@ -143,7 +143,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
143
143
|
| `deepseek/deepseek-v4-flash-vision-exp` |
|
|
144
144
|
| `deepseek/deepseek-v4-pro` |
|
|
145
145
|
| `deepseek/deepseek-v4-pro-0813` |
|
|
146
|
-
| `deepseek/deepseek-v4.1-flash
|
|
146
|
+
| `deepseek/deepseek-v4.1-flash` |
|
|
147
147
|
| `fish-audio/s1` |
|
|
148
148
|
| `fish-audio/s1-free` |
|
|
149
149
|
| `fish-audio/s2-pro` |
|
package/.docs/models/index.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Model Providers
|
|
6
6
|
|
|
7
|
-
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to
|
|
7
|
+
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7171 models from 200 providers through a single API.
|
|
8
8
|
|
|
9
9
|
## Features
|
|
10
10
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# above.dev
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 8 above.dev models through Mastra's model router. Authentication is handled automatically using the `ABOVE_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [above.dev documentation](https://above.dev/docs).
|
|
10
10
|
|
|
@@ -45,7 +45,6 @@ for await (const chunk of stream) {
|
|
|
45
45
|
| `above/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
|
|
46
46
|
| `above/glm-5.3-flash` | 1.0M | | | | | | $0.17 | $0.55 |
|
|
47
47
|
| `above/mimo-v2.5-pro` | 1.0M | | | | | | $0.51 | $1 |
|
|
48
|
-
| `above/mimo-v2.5-pro-ultraspeed` | 1.0M | | | | | | $2 | $3 |
|
|
49
48
|
| `above/qwen3.8-max` | 1.0M | | | | | | $2 | $7 |
|
|
50
49
|
|
|
51
50
|
Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
|
|
@@ -61,7 +61,6 @@ for await (const chunk of stream) {
|
|
|
61
61
|
| `alibaba-token-plan-cn/qwen3.7-plus` | 1.0M | | | | | | — | — |
|
|
62
62
|
| `alibaba-token-plan-cn/qwen3.8-flash` | 1.0M | | | | | | — | — |
|
|
63
63
|
| `alibaba-token-plan-cn/qwen3.8-max` | 1.0M | | | | | | — | — |
|
|
64
|
-
| `alibaba-token-plan-cn/qwen3.8-max-preview` | 1.0M | | | | | | — | — |
|
|
65
64
|
| `alibaba-token-plan-cn/wan2.7-image` | 8K | | | | | | — | — |
|
|
66
65
|
| `alibaba-token-plan-cn/wan2.7-image-pro` | 8K | | | | | | — | — |
|
|
67
66
|
|
|
@@ -61,7 +61,6 @@ for await (const chunk of stream) {
|
|
|
61
61
|
| `alibaba-token-plan/qwen3.7-plus` | 1.0M | | | | | | — | — |
|
|
62
62
|
| `alibaba-token-plan/qwen3.8-flash` | 1.0M | | | | | | — | — |
|
|
63
63
|
| `alibaba-token-plan/qwen3.8-max` | 1.0M | | | | | | — | — |
|
|
64
|
-
| `alibaba-token-plan/qwen3.8-max-preview` | 1.0M | | | | | | — | — |
|
|
65
64
|
| `alibaba-token-plan/wan2.7-image` | 8K | | | | | | — | — |
|
|
66
65
|
| `alibaba-token-plan/wan2.7-image-pro` | 8K | | | | | | — | — |
|
|
67
66
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Bothub
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 7 Bothub models through Mastra's model router. Authentication is handled automatically using the `BOTHUB_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Bothub documentation](https://bothub.ru/models).
|
|
10
10
|
|
|
@@ -19,7 +19,7 @@ const agent = new Agent({
|
|
|
19
19
|
id: "my-agent",
|
|
20
20
|
name: "My Agent",
|
|
21
21
|
instructions: "You are a helpful assistant",
|
|
22
|
-
model: "bothub/
|
|
22
|
+
model: "bothub/deepseek-v4-flash-0731"
|
|
23
23
|
});
|
|
24
24
|
|
|
25
25
|
// Generate a response
|
|
@@ -38,7 +38,12 @@ for await (const chunk of stream) {
|
|
|
38
38
|
|
|
39
39
|
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|
|
40
40
|
| ---------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
|
|
41
|
+
| `bothub/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.10 | $0.28 |
|
|
42
|
+
| `bothub/deepseek-v4-pro-0813` | 1.0M | | | | | | $2 | $5 |
|
|
41
43
|
| `bothub/gemma-4-31b-it:free` | 262K | | | | | | — | — |
|
|
44
|
+
| `bothub/glm-5.3` | 1.0M | | | | | | $2 | $5 |
|
|
45
|
+
| `bothub/glm-5.3-flash` | 1.0M | | | | | | $0.12 | $0.44 |
|
|
46
|
+
| `bothub/gpt-5.6-luna` | 1.1M | | | | | | $0.06 | $0.37 |
|
|
42
47
|
| `bothub/nemotron-3-ultra-550b-a55b:free` | 1.0M | | | | | | — | — |
|
|
43
48
|
|
|
44
49
|
Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
|
|
@@ -53,7 +58,7 @@ const agent = new Agent({
|
|
|
53
58
|
name: "custom-agent",
|
|
54
59
|
model: {
|
|
55
60
|
url: "https://openai.bothub.ru/v1",
|
|
56
|
-
id: "bothub/
|
|
61
|
+
id: "bothub/deepseek-v4-flash-0731",
|
|
57
62
|
apiKey: process.env.BOTHUB_API_KEY,
|
|
58
63
|
headers: {
|
|
59
64
|
"X-Custom-Header": "value"
|
|
@@ -72,7 +77,7 @@ const agent = new Agent({
|
|
|
72
77
|
const useAdvanced = requestContext.task === "complex";
|
|
73
78
|
return useAdvanced
|
|
74
79
|
? "bothub/nemotron-3-ultra-550b-a55b:free"
|
|
75
|
-
: "bothub/
|
|
80
|
+
: "bothub/deepseek-v4-flash-0731";
|
|
76
81
|
}
|
|
77
82
|
});
|
|
78
83
|
```
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Deep Infra
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 64 Deep Infra models through Mastra's model router. Authentication is handled automatically using the `DEEPINFRA_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Deep Infra documentation](https://deepinfra.com/models).
|
|
10
10
|
|
|
@@ -49,13 +49,13 @@ for await (const chunk of stream) {
|
|
|
49
49
|
| `deepinfra/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp` | 1.0M | | | | | | $0.44 | $1 |
|
|
50
50
|
| `deepinfra/deepseek-ai/DeepSeek-V4-Pro` | 1.0M | | | | | | $1 | $3 |
|
|
51
51
|
| `deepinfra/deepseek-ai/DeepSeek-V4-Pro-0813` | 1.0M | | | | | | $1 | $3 |
|
|
52
|
+
| `deepinfra/deepseek-ai/DeepSeek-V4.1-Flash` | 1.0M | | | | | | $0.30 | $1 |
|
|
52
53
|
| `deepinfra/google/gemma-4-26B-A4B-it` | 262K | | | | | | $0.07 | $0.34 |
|
|
53
54
|
| `deepinfra/google/gemma-4-31B-it` | 262K | | | | | | $0.13 | $0.38 |
|
|
54
55
|
| `deepinfra/google/gemma-4-E4B-it` | 131K | | | | | | $0.02 | $0.10 |
|
|
55
56
|
| `deepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo` | 131K | | | | | | $0.10 | $0.32 |
|
|
56
57
|
| `deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8` | 1.0M | | | | | | $0.20 | $0.80 |
|
|
57
58
|
| `deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct` | 328K | | | | | | $0.10 | $0.30 |
|
|
58
|
-
| `deepinfra/MiniMaxAI/MiniMax-M2.7` | 197K | | | | | | $0.25 | $1 |
|
|
59
59
|
| `deepinfra/MiniMaxAI/MiniMax-M3` | 524K | | | | | | $0.28 | $1 |
|
|
60
60
|
| `deepinfra/moonshotai/Kimi-K2.6` | 262K | | | | | | $0.75 | $4 |
|
|
61
61
|
| `deepinfra/moonshotai/Kimi-K2.7-Code` | 262K | | | | | | $0.68 | $3 |
|
|
@@ -89,8 +89,6 @@ for await (const chunk of stream) {
|
|
|
89
89
|
| `deepinfra/XiaomiMiMo/MiMo-V2.5-Pro` | 1.0M | | | | | | $1 | $3 |
|
|
90
90
|
| `deepinfra/zai-org/GLM-4.6` | 203K | | | | | | $0.50 | $2 |
|
|
91
91
|
| `deepinfra/zai-org/GLM-4.7` | 203K | | | | | | $0.40 | $2 |
|
|
92
|
-
| `deepinfra/zai-org/GLM-4.7-Flash` | 203K | | | | | | $0.06 | $0.40 |
|
|
93
|
-
| `deepinfra/zai-org/GLM-5` | 203K | | | | | | $0.60 | $2 |
|
|
94
92
|
| `deepinfra/zai-org/GLM-5.1` | 203K | | | | | | $1 | $4 |
|
|
95
93
|
| `deepinfra/zai-org/GLM-5.2` | 1.0M | | | | | | $0.75 | $2 |
|
|
96
94
|
| `deepinfra/zai-org/GLM-5.3` | 1.0M | | | | | | $1 | $4 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# DeepSeek
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 4 DeepSeek models through Mastra's model router. Authentication is handled automatically using the `DEEPSEEK_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [DeepSeek documentation](https://api-docs.deepseek.com/quick_start/pricing).
|
|
10
10
|
|
|
@@ -19,7 +19,7 @@ const agent = new Agent({
|
|
|
19
19
|
id: "my-agent",
|
|
20
20
|
name: "My Agent",
|
|
21
21
|
instructions: "You are a helpful assistant",
|
|
22
|
-
model: "deepseek/deepseek-
|
|
22
|
+
model: "deepseek/deepseek-flash"
|
|
23
23
|
});
|
|
24
24
|
|
|
25
25
|
// Generate a response
|
|
@@ -36,8 +36,9 @@ for await (const chunk of stream) {
|
|
|
36
36
|
|
|
37
37
|
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|
|
38
38
|
| --------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
|
|
39
|
-
| `deepseek/deepseek-
|
|
40
|
-
| `deepseek/deepseek-v4-flash
|
|
39
|
+
| `deepseek/deepseek-flash` | 1.0M | | | | | | $0.15 | $0.60 |
|
|
40
|
+
| `deepseek/deepseek-v4-flash` | 1.0M | | | | | | $0.15 | $0.60 |
|
|
41
|
+
| `deepseek/deepseek-v4-flash-vision-exp` | 1.0M | | | | | | $0.15 | $0.60 |
|
|
41
42
|
| `deepseek/deepseek-v4-pro` | 1.0M | | | | | | $0.43 | $0.87 |
|
|
42
43
|
|
|
43
44
|
Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
|
|
@@ -52,7 +53,7 @@ const agent = new Agent({
|
|
|
52
53
|
name: "custom-agent",
|
|
53
54
|
model: {
|
|
54
55
|
url: "https://api.deepseek.com",
|
|
55
|
-
id: "deepseek/deepseek-
|
|
56
|
+
id: "deepseek/deepseek-flash",
|
|
56
57
|
apiKey: process.env.DEEPSEEK_API_KEY,
|
|
57
58
|
headers: {
|
|
58
59
|
"X-Custom-Header": "value"
|
|
@@ -71,7 +72,7 @@ const agent = new Agent({
|
|
|
71
72
|
const useAdvanced = requestContext.task === "complex";
|
|
72
73
|
return useAdvanced
|
|
73
74
|
? "deepseek/deepseek-v4-pro"
|
|
74
|
-
: "deepseek/deepseek-
|
|
75
|
+
: "deepseek/deepseek-flash";
|
|
75
76
|
}
|
|
76
77
|
});
|
|
77
78
|
```
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Eden AI
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 260 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Eden AI documentation](https://docs.edenai.co).
|
|
10
10
|
|
|
@@ -94,6 +94,7 @@ for await (const chunk of stream) {
|
|
|
94
94
|
| `edenai/deepinfra/deepseek-ai/DeepSeek-V3-0324` | 164K | | | | | | $0.24 | $0.90 |
|
|
95
95
|
| `edenai/deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731` | 1.0M | | | | | | $0.06 | $0.18 |
|
|
96
96
|
| `edenai/deepinfra/deepseek-ai/DeepSeek-V4-Pro-0813` | 1.0M | | | | | | $1 | $3 |
|
|
97
|
+
| `edenai/deepinfra/deepseek-ai/DeepSeek-V4.1-Flash` | 1.0M | | | | | | $0.30 | $1 |
|
|
97
98
|
| `edenai/deepinfra/meta-llama/Llama-3.2-11B-Vision-Instruct` | 131K | | | | | | $0.34 | $0.34 |
|
|
98
99
|
| `edenai/deepinfra/meta-llama/Llama-3.3-70B-Instruct` | 131K | | | | | | $0.10 | $0.32 |
|
|
99
100
|
| `edenai/deepinfra/meta-llama/Llama-Guard-3-8B` | 131K | | | | | | $0.06 | $0.06 |
|
|
@@ -217,8 +218,8 @@ for await (const chunk of stream) {
|
|
|
217
218
|
| `edenai/perplexityai/sonar-deep-research` | 128K | | | | | | $2 | $8 |
|
|
218
219
|
| `edenai/perplexityai/sonar-pro` | 200K | | | | | | $3 | $15 |
|
|
219
220
|
| `edenai/perplexityai/sonar-reasoning-pro` | 128K | | | | | | $2 | $8 |
|
|
220
|
-
| `edenai/qwen/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.
|
|
221
|
-
| `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $
|
|
221
|
+
| `edenai/qwen/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.18 | $0.53 |
|
|
222
|
+
| `edenai/qwen/deepseek-v4-pro-0813` | 1.0M | | | | | | $0.58 | $2 |
|
|
222
223
|
| `edenai/qwen/qwen-max` | 33K | | | | | | $2 | $6 |
|
|
223
224
|
| `edenai/qwen/qwen-vl-max` | 131K | | | | | | $0.80 | $3 |
|
|
224
225
|
| `edenai/qwen/qwen-vl-plus` | 131K | | | | | | $0.21 | $0.63 |
|
|
@@ -241,7 +242,7 @@ for await (const chunk of stream) {
|
|
|
241
242
|
| `edenai/qwen/qwen3.8-max` | 1.0M | | | | | | $2 | $6 |
|
|
242
243
|
| `edenai/qwen/qwen3.8-max-0902` | 1.0M | | | | | | $2 | $6 |
|
|
243
244
|
| `edenai/qwen/qwq-plus` | 131K | | | | | | $0.80 | $2 |
|
|
244
|
-
| `edenai/scaleway/deepseek-v4-flash-0731` | 256K | | | | | | $0.
|
|
245
|
+
| `edenai/scaleway/deepseek-v4-flash-0731` | 256K | | | | | | $0.46 | $0.93 |
|
|
245
246
|
| `edenai/scaleway/gpt-oss-120b` | 128K | | | | | | $0.17 | $0.70 |
|
|
246
247
|
| `edenai/scaleway/llama-3.3-70b-instruct` | 128K | | | | | | $1 | $1 |
|
|
247
248
|
| `edenai/tensorx/deepseek/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.25 | $0.30 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Fireworks AI
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 23 Fireworks AI models through Mastra's model router. Authentication is handled automatically using the `FIREWORKS_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Fireworks AI documentation](https://fireworks.ai/docs/).
|
|
10
10
|
|
|
@@ -41,6 +41,7 @@ for await (const chunk of stream) {
|
|
|
41
41
|
| `fireworks-ai/accounts/fireworks/models/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.22 | $0.66 |
|
|
42
42
|
| `fireworks-ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp` | 1.0M | | | | | | $0.22 | $0.66 |
|
|
43
43
|
| `fireworks-ai/accounts/fireworks/models/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $4 |
|
|
44
|
+
| `fireworks-ai/accounts/fireworks/models/deepseek-v4p1-flash` | 1.0M | | | | | | $0.22 | $0.66 |
|
|
44
45
|
| `fireworks-ai/accounts/fireworks/models/glm-5p2` | 1.0M | | | | | | $1 | $4 |
|
|
45
46
|
| `fireworks-ai/accounts/fireworks/models/glm-5p3` | 1.0M | | | | | | $1 | $4 |
|
|
46
47
|
| `fireworks-ai/accounts/fireworks/models/glm-5p3-flash` | 1.0M | | | | | | $0.15 | $0.50 |
|
|
@@ -50,6 +51,7 @@ for await (const chunk of stream) {
|
|
|
50
51
|
| `fireworks-ai/accounts/fireworks/models/kimi-k2p7-code` | 262K | | | | | | $0.95 | $4 |
|
|
51
52
|
| `fireworks-ai/accounts/fireworks/models/kimi-k3` | 1.0M | | | | | | $3 | $15 |
|
|
52
53
|
| `fireworks-ai/accounts/fireworks/models/minimax-m3` | 512K | | | | | | $0.30 | $1 |
|
|
54
|
+
| `fireworks-ai/accounts/fireworks/models/mistral-large-3-fp8` | 262K | | | | | | — | — |
|
|
53
55
|
| `fireworks-ai/accounts/fireworks/models/muse-glimmer-30b` | 131K | | | | | | $0.35 | $2 |
|
|
54
56
|
| `fireworks-ai/accounts/fireworks/models/nemotron-3-ultra-nvfp4` | 262K | | | | | | $0.60 | $2 |
|
|
55
57
|
| `fireworks-ai/accounts/fireworks/models/nemotron-lightning-3p5-30b-a3b` | 262K | | | | | | $0.05 | $0.20 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Hugging Face
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 74 Hugging Face models through Mastra's model router. Authentication is handled automatically using the `HF_TOKEN` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Hugging Face documentation](https://huggingface.co).
|
|
10
10
|
|
|
@@ -49,6 +49,7 @@ for await (const chunk of stream) {
|
|
|
49
49
|
| `huggingface/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp` | 1.0M | | | | | | $0.44 | $1 |
|
|
50
50
|
| `huggingface/deepseek-ai/DeepSeek-V4-Pro` | 1.0M | | | | | | $0.43 | $0.87 |
|
|
51
51
|
| `huggingface/deepseek-ai/DeepSeek-V4-Pro-0813` | 1.0M | | | | | | $1 | $4 |
|
|
52
|
+
| `huggingface/deepseek-ai/DeepSeek-V4.1-Flash` | 1.0M | | | | | | $0.30 | $1 |
|
|
52
53
|
| `huggingface/google/gemma-4-26B-A4B-it` | 262K | | | | | | $0.13 | $0.40 |
|
|
53
54
|
| `huggingface/google/gemma-4-31B-it` | 262K | | | | | | $0.14 | $0.40 |
|
|
54
55
|
| `huggingface/meta-llama/Llama-3.1-8B-Instruct` | 131K | | | | | | $0.06 | $0.06 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Charm Hyper
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 34 Charm Hyper models through Mastra's model router. Authentication is handled automatically using the `HYPER_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Charm Hyper documentation](https://hyper.charm.land).
|
|
10
10
|
|
|
@@ -42,13 +42,14 @@ for await (const chunk of stream) {
|
|
|
42
42
|
| `hyper/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.44 | $1 |
|
|
43
43
|
| `hyper/deepseek-v4-pro` | 1.0M | | | | | | $2 | $5 |
|
|
44
44
|
| `hyper/deepseek-v4-pro-0813` | 1.0M | | | | | | $1 | $4 |
|
|
45
|
-
| `hyper/
|
|
45
|
+
| `hyper/deepseek-v4.1-flash` | 1.0M | | | | | | $0.30 | $1 |
|
|
46
|
+
| `hyper/gemma-4-26b-a4b-it` | 256K | | | | | | $0.12 | $0.42 |
|
|
46
47
|
| `hyper/glm-5` | 203K | | | | | | $0.86 | $3 |
|
|
47
48
|
| `hyper/glm-5.1` | 203K | | | | | | $1 | $4 |
|
|
48
49
|
| `hyper/glm-5.2` | 1.0M | | | | | | $2 | $5 |
|
|
49
50
|
| `hyper/glm-5.3` | 1.0M | | | | | | $2 | $5 |
|
|
50
51
|
| `hyper/glm-5.3-flash` | 1.0M | | | | | | $0.16 | $0.54 |
|
|
51
|
-
| `hyper/gpt-oss-120b` | 128K | | | | | | $0.18 | $0.
|
|
52
|
+
| `hyper/gpt-oss-120b` | 128K | | | | | | $0.18 | $0.68 |
|
|
52
53
|
| `hyper/inkling` | 1.0M | | | | | | $1 | $4 |
|
|
53
54
|
| `hyper/kimi-k2-thinking` | 262K | | | | | | $0.60 | $3 |
|
|
54
55
|
| `hyper/kimi-k2.5` | 262K | | | | | | $0.56 | $3 |
|
|
@@ -57,7 +58,7 @@ for await (const chunk of stream) {
|
|
|
57
58
|
| `hyper/kimi-k3` | 1.0M | | | | | | $3 | $16 |
|
|
58
59
|
| `hyper/llama-3.3-70b-instruct` | 128K | | | | | | $0.61 | $1 |
|
|
59
60
|
| `hyper/llama-4-maverick-17b-128e-instruct-fp8` | 430K | | | | | | $0.27 | $0.90 |
|
|
60
|
-
| `hyper/minimax-m2.7` | 262K | | | | | | $0.
|
|
61
|
+
| `hyper/minimax-m2.7` | 262K | | | | | | $0.40 | $1 |
|
|
61
62
|
| `hyper/minimax-m3` | 512K | | | | | | $0.33 | $1 |
|
|
62
63
|
| `hyper/qwen3-coder-480b-a35b-instruct-int4-mixed-ar` | 106K | | | | | | $0.45 | $2 |
|
|
63
64
|
| `hyper/qwen3-next-80b-a3b-instruct` | 262K | | | | | | $0.12 | $1 |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Inception
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 3 Inception models through Mastra's model router. Authentication is handled automatically using the `INCEPTION_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Inception documentation](https://platform.inceptionlabs.ai/docs).
|
|
10
10
|
|
|
@@ -39,6 +39,7 @@ for await (const chunk of stream) {
|
|
|
39
39
|
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|
|
40
40
|
| -------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
|
|
41
41
|
| `inception/mercury-2` | 128K | | | | | | $0.25 | $0.75 |
|
|
42
|
+
| `inception/mercury-2.5` | 260K | | | | | | $0.04 | $0.15 |
|
|
42
43
|
| `inception/mercury-edit-2` | 128K | | | | | | $0.25 | $0.75 |
|
|
43
44
|
|
|
44
45
|
Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
|