@mastra/mcp-docs-server 1.3.2-alpha.1 → 1.3.2-alpha.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.docs/docs/agents/processors.md +1 -1
- package/.docs/docs/evals/evals-with-memory.md +3 -1
- package/.docs/docs/evals/experiments.md +30 -10
- package/.docs/docs/harness/durable-agents.md +2 -2
- package/.docs/docs/memory/observational-memory.md +34 -0
- package/.docs/integrations/sandboxes/modal.md +40 -0
- package/.docs/models/gateways/netlify.md +3 -2
- package/.docs/models/gateways/openrouter.md +2 -5
- package/.docs/models/index.md +1 -1
- package/.docs/models/providers/above.md +2 -2
- package/.docs/models/providers/cerebras.md +1 -1
- package/.docs/models/providers/cortecs.md +3 -3
- package/.docs/models/providers/deepseek.md +9 -3
- package/.docs/models/providers/edenai.md +285 -288
- package/.docs/models/providers/fireworks-ai.md +4 -7
- package/.docs/models/providers/kilo.md +15 -18
- package/.docs/models/providers/melious.md +1 -4
- package/.docs/models/providers/nano-gpt.md +7 -1
- package/.docs/models/providers/opencode-go.md +2 -1
- package/.docs/models/providers/opencode.md +3 -2
- package/.docs/models/providers/requesty.md +2 -2
- package/.docs/models/providers/zenmux.md +8 -1
- package/.docs/reference/agents/durable-agent.md +2 -0
- package/.docs/reference/channels/channel-provider.md +20 -1
- package/.docs/reference/coding-agent/create-coding-agent.md +22 -14
- package/.docs/reference/index.md +1 -0
- package/.docs/reference/memory/observational-memory.md +8 -0
- package/.docs/reference/processors/agents-md-injector.md +2 -0
- package/.docs/reference/processors/cyber-refusal-handler.md +76 -0
- package/.docs/reference/workspace/local-sandbox.md +2 -0
- package/.docs/reference/workspace/workspace-class.md +14 -1
- package/package.json +5 -5
|
@@ -879,7 +879,7 @@ const agent = new Agent({
|
|
|
879
879
|
The retry mechanism:
|
|
880
880
|
|
|
881
881
|
- Works in `processOutputStep()` and `processInputStep()` methods
|
|
882
|
-
- Replays the step with the abort reason
|
|
882
|
+
- Replays the step with the abort reason appended verbatim to the end of the conversation as a system reminder. The retry keeps the previous request as its prefix, which preserves provider prompt caching
|
|
883
883
|
- Tracks retry count via the `retryCount` parameter
|
|
884
884
|
- Requires an explicit `maxProcessorRetries` limit on the agent or call
|
|
885
885
|
|
|
@@ -90,7 +90,9 @@ const average = scores.reduce((a, b) => a + b, 0) / scores.length
|
|
|
90
90
|
|
|
91
91
|
## Dataset experiments with an inline task
|
|
92
92
|
|
|
93
|
-
`dataset.startExperiment(
|
|
93
|
+
`dataset.startExperiment()` accepts `targetType` and `targetId` for a registered target, or an inline `task` function. For a registered memory-enabled agent, the runner creates a fresh thread per item by default. See [memory-enabled experiment targets](https://mastra.ai/docs/evals/experiments) for resource isolation and request context options.
|
|
94
|
+
|
|
95
|
+
To choose each item's thread explicitly through `agent.generate()`, use an inline `task` and store `{ threadId, resourceId }` in each item's `metadata`. The scorer pipeline still runs as normal.
|
|
94
96
|
|
|
95
97
|
```typescript
|
|
96
98
|
import { randomUUID } from 'node:crypto'
|
|
@@ -53,7 +53,9 @@ You can also delete experiments with the [Core API](https://mastra.ai/reference/
|
|
|
53
53
|
|
|
54
54
|
## Experiment targets
|
|
55
55
|
|
|
56
|
-
|
|
56
|
+
Use `targetType` and `targetId` to select an agent, workflow, or scorer registered on your Mastra instance. The target processes each dataset item. The experiment's `scorers` evaluate the target's output afterward.
|
|
57
|
+
|
|
58
|
+
For custom execution, pass an inline `task` function instead. See [dataset experiments with an inline task](https://mastra.ai/docs/evals/evals-with-memory) for an example that controls each item's memory thread.
|
|
57
59
|
|
|
58
60
|
### Registered agent
|
|
59
61
|
|
|
@@ -68,25 +70,27 @@ const summary = await dataset.startExperiment({
|
|
|
68
70
|
})
|
|
69
71
|
```
|
|
70
72
|
|
|
71
|
-
Each item's `input` is passed directly to
|
|
73
|
+
Each item's `input` is passed directly to the agent. Use a supported message format, such as a prompt string, a message object, or an array of messages. See the [`agent.generate()` reference](https://mastra.ai/reference/agents/generate) for accepted inputs.
|
|
72
74
|
|
|
73
75
|
#### Memory-enabled agents
|
|
74
76
|
|
|
75
|
-
When the target agent has its own memory
|
|
77
|
+
When the target agent has its own memory, the experiment runner injects a fresh memory thread for each item and retry. Without a resource id in the request context, each thread gets an isolated, experiment-owned resource.
|
|
78
|
+
|
|
79
|
+
To run against an existing resource, set `MASTRA_RESOURCE_ID_KEY` in the experiment's `requestContext` or in an item's `requestContext`. Item values override experiment values. The runner still creates a fresh thread for each item and retry, but all threads using the same resource id share that resource's memory.
|
|
76
80
|
|
|
77
81
|
Injected threads are tagged so you can map them back to the run: thread metadata carries the `experimentId` and the dataset item's id as `experimentItemId`. No thread title is generated for them.
|
|
78
82
|
|
|
79
|
-
|
|
83
|
+
When items share a resource, resource-scoped memory features can read and write the same state:
|
|
80
84
|
|
|
81
|
-
- Resource-scoped working memory updates persist to the resource
|
|
85
|
+
- Resource-scoped working memory updates persist to the resource and can affect other items. Items run concurrently by default, so don't rely on an execution order.
|
|
82
86
|
- Resource-scoped semantic recall can surface the resource's prior conversations to the experiment, and experiment transcripts become recallable in that resource's later conversations.
|
|
83
87
|
|
|
84
88
|
Use this approach to evaluate an agent against a real user's accumulated context. To keep experiment runs from touching real user state, use a dedicated evaluation resource id instead.
|
|
85
89
|
|
|
86
90
|
Thread injection is skipped in the following cases:
|
|
87
91
|
|
|
88
|
-
- If the request context
|
|
89
|
-
- If the agent has no memory, or the request context
|
|
92
|
+
- If the request context sets `MASTRA_THREAD_ID_KEY`, the runner uses that thread as-is. Items and retries using the same thread id share a conversation.
|
|
93
|
+
- If the agent has no memory, or the request context explicitly sets the resource id to an empty string or `null`, the runner skips thread injection.
|
|
90
94
|
|
|
91
95
|
### Registered workflow
|
|
92
96
|
|
|
@@ -101,13 +105,24 @@ const summary = await dataset.startExperiment({
|
|
|
101
105
|
})
|
|
102
106
|
```
|
|
103
107
|
|
|
104
|
-
The workflow receives each item's `input` as
|
|
108
|
+
The workflow receives each item's `input` as `inputData` in `run.start()`. Structure the input to match the workflow's input schema.
|
|
105
109
|
|
|
106
110
|
### Registered scorer
|
|
107
111
|
|
|
108
|
-
|
|
112
|
+
Use a scorer as the target to evaluate the judge itself. The runner calls `scorer.run(item.input)`, so the dataset item's `input` must contain the full payload the scorer expects.
|
|
113
|
+
|
|
114
|
+
This example assumes a registered `accuracy` scorer that accepts string `input`, `output`, and `groundTruth` fields. Adapt those fields to your scorer's expected shape:
|
|
109
115
|
|
|
110
116
|
```typescript
|
|
117
|
+
await dataset.addItem({
|
|
118
|
+
input: {
|
|
119
|
+
input: 'What is the capital of France?',
|
|
120
|
+
output: 'Paris',
|
|
121
|
+
groundTruth: 'Paris',
|
|
122
|
+
},
|
|
123
|
+
groundTruth: { score: 1 },
|
|
124
|
+
})
|
|
125
|
+
|
|
111
126
|
const summary = await dataset.startExperiment({
|
|
112
127
|
name: 'judge-accuracy-eval',
|
|
113
128
|
targetType: 'scorer',
|
|
@@ -115,7 +130,12 @@ const summary = await dataset.startExperiment({
|
|
|
115
130
|
})
|
|
116
131
|
```
|
|
117
132
|
|
|
118
|
-
The
|
|
133
|
+
The two `groundTruth` fields serve different purposes:
|
|
134
|
+
|
|
135
|
+
- `input.groundTruth` is the reference answer passed to the judge: `'Paris'`.
|
|
136
|
+
- The top-level `groundTruth` is the expected result from the judge: `{ score: 1 }`. It isn't passed to the target scorer.
|
|
137
|
+
|
|
138
|
+
The target's output contains its `score` and `reason`. To evaluate that output against the top-level ground truth, add another scorer through the experiment's `scorers` option. Without an additional scorer, the experiment records the judge's output but doesn't score its agreement with the expected result.
|
|
119
139
|
|
|
120
140
|
## Scoring results
|
|
121
141
|
|
|
@@ -278,6 +278,8 @@ Under the default persistence policy, these `running` checkpoints are only writt
|
|
|
278
278
|
|
|
279
279
|
Durable agent runs are excluded from the generic boot-time restart of active workflow runs. The only automatic recovery path for durable agent runs is `recovery.durableAgents: 'auto'`, which holds a recovery lease and registers thread runtimes before re-driving each run.
|
|
280
280
|
|
|
281
|
+
> **Warning:** Automatic and manual recovery re-run the agentic loop from the last persisted snapshot, which can reissue LLM calls (with real cost) and re-execute tool calls. A sub-agent delegation is one generated `agent-<name>` tool call, and recovery doesn't checkpoint the LLM or tool calls inside that delegation independently. If recovery replays the delegation step, the entire sub-agent run starts again and can repeat inner LLM calls and already-completed tool side effects. Make tools, including tools used by sub-agents, idempotent, or deduplicate their external side effects.
|
|
282
|
+
|
|
281
283
|
### Snapshot persistence
|
|
282
284
|
|
|
283
285
|
The `shouldPersistSnapshot` option on `createDurableAgent()` (also accepted in the agent-level `durable` config) controls which workflow snapshots a durable agent writes. The default policy:
|
|
@@ -313,8 +315,6 @@ export const mastra = new Mastra({
|
|
|
313
315
|
|
|
314
316
|
On startup, this discovers every registered durable agent with runs stuck in `running` status and re-drives them from the last persisted snapshot.
|
|
315
317
|
|
|
316
|
-
> **Warning:** Recovery re-runs the agentic loop from the last snapshot, which re-issues LLM calls (real cost) and re-executes tool calls. Make sure your tools are idempotent before enabling automatic recovery.
|
|
317
|
-
|
|
318
318
|
### Manual recovery
|
|
319
319
|
|
|
320
320
|
If you need finer control, such as gating recovery behind a leader election or running it on a schedule, call the methods directly:
|
|
@@ -191,6 +191,40 @@ const memory = new Memory({
|
|
|
191
191
|
|
|
192
192
|
`bufferOnIdle` is off by default. It's separate from `bufferTokens`: `bufferTokens` controls step-time async buffering, while `bufferOnIdle` controls end-of-turn buffering for idle turns.
|
|
193
193
|
|
|
194
|
+
### Retries and failure policy
|
|
195
|
+
|
|
196
|
+
Observer and Reflector model calls are retried on transient provider errors, and a turn aborts if they still fail. Each stage is controlled independently by:
|
|
197
|
+
|
|
198
|
+
- `maxRetries`: retries after the initial model call (default `8`).
|
|
199
|
+
- `failurePolicy`: what happens once retries are exhausted, either `'abort'` (default) or `'continue'`.
|
|
200
|
+
|
|
201
|
+
```typescript
|
|
202
|
+
const memory = new Memory({
|
|
203
|
+
options: {
|
|
204
|
+
observationalMemory: {
|
|
205
|
+
observation: {
|
|
206
|
+
maxRetries: 8,
|
|
207
|
+
failurePolicy: 'continue',
|
|
208
|
+
},
|
|
209
|
+
reflection: {
|
|
210
|
+
maxRetries: 8,
|
|
211
|
+
failurePolicy: 'continue',
|
|
212
|
+
},
|
|
213
|
+
},
|
|
214
|
+
},
|
|
215
|
+
})
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
`maxRetries` governs OM's own retry ladder; the underlying model call is configured with no provider-level retries, so the option is the single retry knob for the stage.
|
|
219
|
+
|
|
220
|
+
With `failurePolicy: 'continue'`, Mastra retries as configured, reports the failure through the existing OM diagnostics, keeps the failed input pending for a later cycle, and lets the main agent turn continue. It doesn't advance observation boundaries or discard unobserved messages. A Reflector failure under `'continue'` leaves any already-persisted observations committed and defers reflection to the next threshold crossing.
|
|
221
|
+
|
|
222
|
+
A provider outage usually takes out both stages, so set the policy on both if you want the turn to survive one.
|
|
223
|
+
|
|
224
|
+
This policy applies only to Observer and Reflector model and provider failures in synchronous, resource-scoped, and buffered observation. Persistence, indexing, transform, locking, invariant, and explicit abort failures remain fatal. It doesn't prevent the underlying provider or network error, change `blockAfter`, or change attachment handling.
|
|
225
|
+
|
|
226
|
+
> **Warning:** `'continue'` has no backstop for a sustained outage. Unobserved messages stay pending and keep accruing in the main agent's context, so a long outage moves the failure from the memory layer to the model's context limit.
|
|
227
|
+
|
|
194
228
|
See [the API reference](https://mastra.ai/reference/memory/observational-memory) for the full configuration shape.
|
|
195
229
|
|
|
196
230
|
## Benefits
|
|
@@ -100,6 +100,8 @@ Get your credentials from the [Modal dashboard](https://modal.com/settings/token
|
|
|
100
100
|
|
|
101
101
|
**onStart** (`function`): Lifecycle hook called after the sandbox reaches running status.
|
|
102
102
|
|
|
103
|
+
**volumes** (`Record<string, Volume>`): Modal Volumes to mount into the sandbox, keyed by absolute mount path. The working directory must not be at or inside a mount path.
|
|
104
|
+
|
|
103
105
|
**onStop** (`function`): Lifecycle hook called before the sandbox stops.
|
|
104
106
|
|
|
105
107
|
**onDestroy** (`function`): Lifecycle hook called before the sandbox is destroyed.
|
|
@@ -137,6 +139,44 @@ await sandbox._stop()
|
|
|
137
139
|
await sandbox._start()
|
|
138
140
|
```
|
|
139
141
|
|
|
142
|
+
## Volumes and persistent files
|
|
143
|
+
|
|
144
|
+
Mount [Modal Volumes](https://modal.com/docs/guide/volumes) with the `volumes` option, then use `ModalFilesystem` to expose a mounted path to the agent's workspace file tools. Files written under the mount persist across sandbox restarts.
|
|
145
|
+
|
|
146
|
+
```typescript
|
|
147
|
+
import { Workspace } from '@mastra/core/workspace'
|
|
148
|
+
import { ModalFilesystem, ModalSandbox } from '@mastra/modal'
|
|
149
|
+
import { ModalClient } from 'modal'
|
|
150
|
+
|
|
151
|
+
const modal = new ModalClient()
|
|
152
|
+
const privateVolume = await modal.volumes.fromName('agent-user-123', { createIfMissing: true })
|
|
153
|
+
const teamVolume = await modal.volumes.fromName('agent-team', { createIfMissing: true })
|
|
154
|
+
|
|
155
|
+
const sandbox = new ModalSandbox({
|
|
156
|
+
id: 'agent-sandbox',
|
|
157
|
+
workingDirectory: '/workspace',
|
|
158
|
+
volumes: {
|
|
159
|
+
'/mnt/agent': privateVolume,
|
|
160
|
+
'/mnt/team': teamVolume,
|
|
161
|
+
},
|
|
162
|
+
})
|
|
163
|
+
|
|
164
|
+
const workspace = new Workspace({
|
|
165
|
+
sandbox,
|
|
166
|
+
filesystem: new ModalFilesystem({ sandbox, basePath: '/mnt/agent' }),
|
|
167
|
+
})
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
`ModalFilesystem` runs file operations inside the sandbox and confines all paths to `basePath`. Set `readOnly: true` to reject writes.
|
|
171
|
+
|
|
172
|
+
Writes made by other sandboxes to the same Volume are only visible after a reload. Call `reloadVolumes()` on a long-lived sandbox to pick them up:
|
|
173
|
+
|
|
174
|
+
```typescript
|
|
175
|
+
await sandbox.reloadVolumes()
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
> **Note:** Keep the working directory outside every Volume mount (for example `/workspace`). Modal can't reload a Volume that is the current working directory, so `ModalSandbox` throws if `workingDirectory` is at or inside a mount path.
|
|
179
|
+
|
|
140
180
|
## Background processes
|
|
141
181
|
|
|
142
182
|
`ModalSandbox` includes a built-in process manager for spawning and managing background processes. Each `spawn()` call creates a new `ContainerProcess` via the Modal SDK's `Sandbox.exec()` API.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Netlify
|
|
6
6
|
|
|
7
|
-
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access
|
|
7
|
+
Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 270 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [Netlify documentation](https://docs.netlify.com/build/ai-gateway/overview/).
|
|
10
10
|
|
|
@@ -229,6 +229,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
229
229
|
| `openrouter/openrouter/free` |
|
|
230
230
|
| `openrouter/openrouter/pareto-code` |
|
|
231
231
|
| `openrouter/perceptron/perceptron-mk1` |
|
|
232
|
+
| `openrouter/perceptron/perceptron-mk1.5` |
|
|
232
233
|
| `openrouter/perplexity/sonar` |
|
|
233
234
|
| `openrouter/perplexity/sonar-deep-research` |
|
|
234
235
|
| `openrouter/perplexity/sonar-pro` |
|
|
@@ -279,6 +280,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
279
280
|
| `openrouter/thedrummer/unslopnemo-12b` |
|
|
280
281
|
| `openrouter/thinkingmachines/inkling` |
|
|
281
282
|
| `openrouter/thinkingmachines/inkling-small` |
|
|
283
|
+
| `openrouter/typesafe/jev-router` |
|
|
282
284
|
| `openrouter/undi95/remm-slerp-l2-13b` |
|
|
283
285
|
| `openrouter/upstage/solar-mini4` |
|
|
284
286
|
| `openrouter/upstage/solar-pro4` |
|
|
@@ -302,7 +304,6 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
302
304
|
| `openrouter/z-ai/glm-5` |
|
|
303
305
|
| `openrouter/z-ai/glm-5.1` |
|
|
304
306
|
| `openrouter/z-ai/glm-5.2` |
|
|
305
|
-
| `openrouter/z-ai/glm-5.2:free` |
|
|
306
307
|
| `openrouter/z-ai/glm-5.3` |
|
|
307
308
|
| `openrouter/z-ai/glm-5.3-flash` |
|
|
308
309
|
| `typesafe/jev-1.13.0` |
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# OpenRouter
|
|
6
6
|
|
|
7
|
-
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access
|
|
7
|
+
OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 384 models through Mastra's model router.
|
|
8
8
|
|
|
9
9
|
Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
|
|
10
10
|
|
|
@@ -68,7 +68,6 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
68
68
|
| `amazon/nova-premier-v1` |
|
|
69
69
|
| `amazon/nova-pro-v1` |
|
|
70
70
|
| `anthracite-org/magnum-v4-72b` |
|
|
71
|
-
| `anthropic/claude-3-haiku` |
|
|
72
71
|
| `anthropic/claude-fable-5` |
|
|
73
72
|
| `anthropic/claude-fable-5.1` |
|
|
74
73
|
| `anthropic/claude-haiku-4.5` |
|
|
@@ -214,8 +213,6 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
214
213
|
| `moonshotai/kimi-k3` |
|
|
215
214
|
| `morph/morph-v3-fast` |
|
|
216
215
|
| `morph/morph-v3-large` |
|
|
217
|
-
| `nex-agi/nex-n2.5-mini:free` |
|
|
218
|
-
| `nex-agi/nex-n2.5-pro:free` |
|
|
219
216
|
| `nousresearch/hermes-3-llama-3.1-405b` |
|
|
220
217
|
| `nousresearch/hermes-3-llama-3.1-70b` |
|
|
221
218
|
| `nousresearch/hermes-4-405b` |
|
|
@@ -298,6 +295,7 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
298
295
|
| `openrouter/fusion` |
|
|
299
296
|
| `openrouter/pareto-code` |
|
|
300
297
|
| `perceptron/perceptron-mk1` |
|
|
298
|
+
| `perceptron/perceptron-mk1.5` |
|
|
301
299
|
| `perplexity/sonar` |
|
|
302
300
|
| `perplexity/sonar-deep-research` |
|
|
303
301
|
| `perplexity/sonar-pro` |
|
|
@@ -419,7 +417,6 @@ ANTHROPIC_API_KEY=ant-...
|
|
|
419
417
|
| `z-ai/glm-5-turbo` |
|
|
420
418
|
| `z-ai/glm-5.1` |
|
|
421
419
|
| `z-ai/glm-5.2` |
|
|
422
|
-
| `z-ai/glm-5.2:free` |
|
|
423
420
|
| `z-ai/glm-5.3` |
|
|
424
421
|
| `z-ai/glm-5.3-flash` |
|
|
425
422
|
| `z-ai/glm-5.3-flashx` |
|
package/.docs/models/index.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# Model Providers
|
|
6
6
|
|
|
7
|
-
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to
|
|
7
|
+
Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7618 models from 210 providers through a single API.
|
|
8
8
|
|
|
9
9
|
## Features
|
|
10
10
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# above.dev
|
|
6
6
|
|
|
7
|
-
Access
|
|
7
|
+
Access 10 above.dev models through Mastra's model router. Authentication is handled automatically using the `ABOVE_API_KEY` environment variable.
|
|
8
8
|
|
|
9
9
|
Learn more in the [above.dev documentation](https://above.dev/docs).
|
|
10
10
|
|
|
@@ -38,8 +38,8 @@ for await (const chunk of stream) {
|
|
|
38
38
|
|
|
39
39
|
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|
|
40
40
|
| -------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
|
|
41
|
-
| `above/deepseek-v4-flash` | 1.0M | | | | | | $0.17 | $0.66 |
|
|
42
41
|
| `above/deepseek-v4-pro` | 1.0M | | | | | | $0.73 | $2 |
|
|
42
|
+
| `above/deepseek-v4.1-flash` | 1.0M | | | | | | $0.17 | $0.66 |
|
|
43
43
|
| `above/glm-5.2` | 1.0M | | | | | | $2 | $5 |
|
|
44
44
|
| `above/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
|
|
45
45
|
| `above/glm-5.3-flash` | 1.0M | | | | | | $0.17 | $0.55 |
|
|
@@ -37,7 +37,7 @@ for await (const chunk of stream) {
|
|
|
37
37
|
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|
|
38
38
|
| ----------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
|
|
39
39
|
| `cerebras/gpt-oss-120b` | 131K | | | | | | $0.35 | $0.75 |
|
|
40
|
-
| `cerebras/qwen-3.8-27b` |
|
|
40
|
+
| `cerebras/qwen-3.8-27b` | 131K | | | | | | $0.99 | $1 |
|
|
41
41
|
|
|
42
42
|
Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
|
|
43
43
|
|
|
@@ -56,7 +56,7 @@ for await (const chunk of stream) {
|
|
|
56
56
|
| `cortecs/deepseek-v4-flash-0731` | 1.0M | | | | | | $0.09 | $0.17 |
|
|
57
57
|
| `cortecs/deepseek-v4-pro` | 1.0M | | | | | | $2 | $3 |
|
|
58
58
|
| `cortecs/deepseek-v4-pro-0813` | 1.0M | | | | | | $2 | $4 |
|
|
59
|
-
| `cortecs/deepseek-v4.1-flash` | 1.0M | | | | | | $0.
|
|
59
|
+
| `cortecs/deepseek-v4.1-flash` | 1.0M | | | | | | $0.30 | $1 |
|
|
60
60
|
| `cortecs/devstral-2512` | 256K | | | | | | $0.48 | $2 |
|
|
61
61
|
| `cortecs/gemini-2.5-flash` | 1.0M | | | | | | $0.30 | $2 |
|
|
62
62
|
| `cortecs/gemini-2.5-pro` | 1.0M | | | | | | $1 | $10 |
|
|
@@ -107,7 +107,7 @@ for await (const chunk of stream) {
|
|
|
107
107
|
| `cortecs/minimax-m2.1` | 196K | | | | | | $0.36 | $1 |
|
|
108
108
|
| `cortecs/minimax-m2.5` | 196K | | | | | | $0.30 | $1 |
|
|
109
109
|
| `cortecs/minimax-m2.7` | 197K | | | | | | $0.67 | $3 |
|
|
110
|
-
| `cortecs/minimax-m3` | 1.0M | | | | | | $0.
|
|
110
|
+
| `cortecs/minimax-m3` | 1.0M | | | | | | $0.30 | $1 |
|
|
111
111
|
| `cortecs/ministral-14b-2512` | 256K | | | | | | $0.24 | $0.24 |
|
|
112
112
|
| `cortecs/ministral-3b-2512` | 256K | | | | | | $0.12 | $0.12 |
|
|
113
113
|
| `cortecs/ministral-8b-2512` | 256K | | | | | | $0.18 | $0.18 |
|
|
@@ -142,7 +142,7 @@ for await (const chunk of stream) {
|
|
|
142
142
|
| `cortecs/qwen3.6-27b` | 262K | | | | | | $0.45 | $3 |
|
|
143
143
|
| `cortecs/qwen3.6-35b-a3b` | 262K | | | | | | $0.17 | $0.56 |
|
|
144
144
|
| `cortecs/qwen3.8-2.4t-a95b` | 262K | | | | | | $3 | $6 |
|
|
145
|
-
| `cortecs/qwen3.8-27b` |
|
|
145
|
+
| `cortecs/qwen3.8-27b` | 1.0M | | | | | | $0.10 | $0.40 |
|
|
146
146
|
| `cortecs/qwen3.8-flash-next` | 262K | | | | | | $0.20 | $0.50 |
|
|
147
147
|
| `cortecs/qwen3guard-gen-0.6b` | 32K | | | | | | — | — |
|
|
148
148
|
| `cortecs/qwen3guard-gen-8b` | 32K | | | | | | — | — |
|
|
@@ -91,8 +91,14 @@ const response = await agent.generate("Hello!", {
|
|
|
91
91
|
|
|
92
92
|
### Available Options
|
|
93
93
|
|
|
94
|
-
**
|
|
94
|
+
**logprobs** (`boolean | undefined`): Whether to return log probabilities for generated tokens.
|
|
95
95
|
|
|
96
|
-
**
|
|
96
|
+
**topLogprobs** (`number | undefined`): Number of most likely tokens to return at each token position. Setting this option automatically enables logprobs.
|
|
97
97
|
|
|
98
|
-
**
|
|
98
|
+
**userId** (`string | undefined`): An opaque identifier for the end user. DeepSeek uses this identifier for content-safety tracing and request isolation. Must contain only ASCII letters, numbers, underscores, and hyphens, and must be at most 512 characters long.
|
|
99
|
+
|
|
100
|
+
**strictJsonSchema** (`boolean | undefined`): Whether to use strict JSON schema validation for structured outputs. Only applies when the serving endpoint supports JSON schema response formats (e.g. Azure). Defaults to true.
|
|
101
|
+
|
|
102
|
+
**thinking** (`{ type?: "enabled" | "disabled" | "adaptive" | undefined; } | undefined`)
|
|
103
|
+
|
|
104
|
+
**reasoningEffort** (`"low" | "high" | "max" | "medium" | "xhigh" | undefined`)
|