@mastra/mcp-docs-server 1.3.0-alpha.4 → 1.3.0-alpha.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -130,7 +130,7 @@ await observability!.deleteFeedback({
130
130
  })
131
131
  ```
132
132
 
133
- Deleted records also disappear from feedback analytics. ClickHouse uses a lightweight delete to hide rows without guaranteeing immediate physical removal, so open-source deployments must configure an [observability retention period](https://mastra.ai/reference/storage/retention) to physically purge them. Configure retention for every observability signal to expire deletion requests after the signal rows they protect. If any signal is unbounded, deletion requests also remain unbounded to prevent deleted data from being reintroduced. On ClickHouse, delete APIs leave the separate delta cursor table untouched. Its rows contain identifiers rather than feedback payloads and expire within two days.
133
+ Deleted records also disappear from feedback analytics. ClickHouse uses a lightweight delete to hide rows without guaranteeing immediate physical removal, so open-source deployments must configure an [observability retention period](https://mastra.ai/reference/storage/retention) to physically purge them. Each request is marked applied once its delete succeeds; if the delete fails, the request stays unapplied, doesn't block updates to the still-visible feedback, and you retry by calling the delete API again. Open-source deployments have no background reconciler. Configure retention for every observability signal to expire deletion requests after the signal rows they protect. If any signal is unbounded, deletion requests also remain unbounded to prevent deleted data from being reintroduced. On ClickHouse, delete APIs leave the separate delta cursor table untouched. Its rows contain identifiers rather than feedback payloads and expire within two days.
134
134
 
135
135
  ## Query feedback analytics
136
136
 
@@ -93,6 +93,10 @@ Lightweight deletion is a hide-only operation that marks rows with ClickHouse's
93
93
 
94
94
  When all five observability signals have finite retention, Mastra also applies a TTL to deletion requests so they outlive the signal rows they protect. If any signal is unbounded, deletion requests remain unbounded. See [storage retention](https://mastra.ai/reference/storage/retention) for how the deletion-request TTL is calculated.
95
95
 
96
+ Feedback review-status updates require `ALTER UPDATE` permission on `mastra_feedback_events` and, when delta polling is enabled, `INSERT` permission on `mastra_feedback_events_delta`. These updates wait for the ClickHouse mutation to finish on the server that receives the write, so latency depends on that server's mutation queue. Other replicas apply the mutation through the replication log, so a reader on a lagging replica can briefly see the previous status, and an inactive replica doesn't block the update. They modify the existing feedback row and preserve deletion masks: a concurrent review update cannot recreate deleted feedback. Successful updates remain available through delta polling. If a newer version of the same feedback event is ingested while a review update is in flight, the update is re-applied to that newer version; after repeated conflicts it fails with a conflict error (HTTP 409) and the caller retries.
97
+
98
+ The status mutation and the delta insert are separate operations. If the delta insert fails, the API returns an error even though the status may have changed, and continued delta polling doesn't recover that notification: later polls from the same cursor never return it, and starting delta mode without a cursor subscribes at the current head. To recover, retry the review update after resolving the error, or reread the feedback with a regular list query.
99
+
96
100
  ### Observability with the legacy domain
97
101
 
98
102
  `ObservabilityStorageClickhouse` is the original observability adapter and remains supported for projects that haven't migrated to the vNext schema. The configuration shape is the same as the vNext class.
@@ -368,7 +372,16 @@ const observability = new ObservabilityStorageClickhouseVNext({
368
372
  await observability.init()
369
373
  ```
370
374
 
371
- In CI/CD pipelines, set `disableInit: true` on `ClickhouseStore` and run `init()` from a deployment step that uses elevated credentials. Runtime application credentials can then be limited to read and insert.
375
+ In CI/CD pipelines, set `disableInit: true` on `ClickhouseStore` and run `init()` from a deployment step that uses elevated credentials. Runtime application credentials still need more than read and insert:
376
+
377
+ - `SELECT` and `INSERT` on the Mastra tables.
378
+ - `ALTER DELETE` on the observability tables you delete from. On ClickHouse 26.6 and earlier, lightweight deletes also require `ALTER UPDATE` on those tables; ClickHouse 26.7 removed that requirement.
379
+ - For feedback review updates, `ALTER UPDATE(reviewStatus)` on `mastra_feedback_events` and `INSERT` on `mastra_feedback_events_delta`.
380
+
381
+ ```sql
382
+ GRANT ALTER UPDATE(reviewStatus) ON <database>.mastra_feedback_events TO <runtime_user>;
383
+ GRANT INSERT ON <database>.mastra_feedback_events_delta TO <runtime_user>;
384
+ ```
372
385
 
373
386
  ## Observability
374
387
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Netlify
6
6
 
7
- Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 264 models through Mastra's model router.
7
+ Netlify AI Gateway provides unified access to multiple providers with built-in caching and observability. Access 265 models through Mastra's model router.
8
8
 
9
9
  Learn more in the [Netlify documentation](https://docs.netlify.com/build/ai-gateway/overview/).
10
10
 
@@ -153,6 +153,7 @@ ANTHROPIC_API_KEY=ant-...
153
153
  | `openrouter/deepseek/deepseek-v4-pro` |
154
154
  | `openrouter/deepseek/deepseek-v4-pro-0813` |
155
155
  | `openrouter/deepseek/deepseek-v4.1-flash` |
156
+ | `openrouter/fireworks/ember-1` |
156
157
  | `openrouter/google/gemma-2-27b-it` |
157
158
  | `openrouter/google/gemma-3-12b-it` |
158
159
  | `openrouter/google/gemma-3-27b-it` |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![OpenRouter logo](https://models.dev/logos/openrouter.svg)OpenRouter
6
6
 
7
- OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 386 models through Mastra's model router.
7
+ OpenRouter aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 384 models through Mastra's model router.
8
8
 
9
9
  Learn more in the [OpenRouter documentation](https://openrouter.ai/models).
10
10
 
@@ -211,9 +211,7 @@ ANTHROPIC_API_KEY=ant-...
211
211
  | `moonshotai/kimi-k3` |
212
212
  | `morph/morph-v3-fast` |
213
213
  | `morph/morph-v3-large` |
214
- | `nex-agi/nex-n2.5-mini` |
215
214
  | `nex-agi/nex-n2.5-mini:free` |
216
- | `nex-agi/nex-n2.5-pro` |
217
215
  | `nex-agi/nex-n2.5-pro:free` |
218
216
  | `nousresearch/hermes-3-llama-3.1-405b` |
219
217
  | `nousresearch/hermes-3-llama-3.1-70b` |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Vercel logo](https://models.dev/logos/vercel.svg)Vercel
6
6
 
7
- Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 387 models through Mastra's model router.
7
+ Vercel aggregates models from multiple providers with enhanced features like rate limiting and failover. Access 388 models through Mastra's model router.
8
8
 
9
9
  Learn more in the [Vercel documentation](https://ai-sdk.dev/providers/ai-sdk-providers).
10
10
 
@@ -71,6 +71,7 @@ ANTHROPIC_API_KEY=ant-...
71
71
  | `alibaba/qwen3.8-flash` |
72
72
  | `alibaba/qwen3.8-max` |
73
73
  | `alibaba/qwen3.8-max-0902` |
74
+ | `alibaba/qwen3.8-max-prime` |
74
75
  | `alibaba/qwen3.8-omni-flash` |
75
76
  | `alibaba/wan-v2.5-t2v-preview` |
76
77
  | `alibaba/wan-v2.6-i2v` |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # Model Providers
6
6
 
7
- Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7594 models from 210 providers through a single API.
7
+ Mastra provides a unified interface for working with LLMs across multiple providers, giving you access to 7614 models from 210 providers through a single API.
8
8
 
9
9
  ## Features
10
10
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![above.dev logo](https://models.dev/logos/above.svg)above.dev
6
6
 
7
- Access 8 above.dev models through Mastra's model router. Authentication is handled automatically using the `ABOVE_API_KEY` environment variable.
7
+ Access 9 above.dev models through Mastra's model router. Authentication is handled automatically using the `ABOVE_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [above.dev documentation](https://above.dev/docs).
10
10
 
@@ -36,16 +36,17 @@ for await (const chunk of stream) {
36
36
 
37
37
  ## Models
38
38
 
39
- | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
40
- | ------------------------------------ | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
41
- | `above/deepseek-v4-flash` | 1.0M | | | | | | $0.17 | $0.66 |
42
- | `above/deepseek-v4-flash-vision-exp` | 1.0M | | | | | | $0.24 | $0.73 |
43
- | `above/deepseek-v4-pro` | 1.0M | | | | | | $0.73 | $2 |
44
- | `above/glm-5.2` | 1.0M | | | | | | $2 | $5 |
45
- | `above/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
46
- | `above/glm-5.3-flash` | 1.0M | | | | | | $0.17 | $0.55 |
47
- | `above/mimo-v2.5-pro` | 1.0M | | | | | | $0.51 | $1 |
48
- | `above/qwen3.8-max` | 1.0M | | | | | | $2 | $7 |
39
+ | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
40
+ | -------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
41
+ | `above/deepseek-v4-flash` | 1.0M | | | | | | $0.17 | $0.66 |
42
+ | `above/deepseek-v4-pro` | 1.0M | | | | | | $0.73 | $2 |
43
+ | `above/glm-5.2` | 1.0M | | | | | | $2 | $5 |
44
+ | `above/glm-5.2-fast` | 1.0M | | | | | | $2 | $7 |
45
+ | `above/glm-5.3-flash` | 1.0M | | | | | | $0.17 | $0.55 |
46
+ | `above/mimo-v2.6-flash` | 1.0M | | | | | | $0.17 | $0.34 |
47
+ | `above/mimo-v2.6-pro` | 1.0M | | | | | | $0.51 | $1 |
48
+ | `above/mimo-v2.6-pro-ultraspeed` | 1.0M | | | | | | $5 | $10 |
49
+ | `above/qwen3.8-max` | 1.0M | | | | | | $2 | $7 |
49
50
 
50
51
  Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
51
52
 
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Cortecs logo](https://models.dev/logos/cortecs.svg)Cortecs
6
6
 
7
- Access 106 Cortecs models through Mastra's model router. Authentication is handled automatically using the `CORTECS_API_KEY` environment variable.
7
+ Access 109 Cortecs models through Mastra's model router. Authentication is handled automatically using the `CORTECS_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Cortecs documentation](https://cortecs.ai).
10
10
 
@@ -43,6 +43,7 @@ for await (const chunk of stream) {
43
43
  | `cortecs/claude-4-6-sonnet` | 1.0M | | | | | | $3 | $16 |
44
44
  | `cortecs/claude-haiku-4-5` | 200K | | | | | | $1.00 | $5 |
45
45
  | `cortecs/claude-opus-5` | 1.0M | | | | | | $6 | $27 |
46
+ | `cortecs/claude-opus-5.5` | 1.0M | | | | | | $4 | $22 |
46
47
  | `cortecs/claude-opus4-5` | 200K | | | | | | $5 | $27 |
47
48
  | `cortecs/claude-opus4-6` | 1.0M | | | | | | $5 | $27 |
48
49
  | `cortecs/claude-opus4-7` | 1.0M | | | | | | $5 | $27 |
@@ -87,8 +88,10 @@ for await (const chunk of stream) {
87
88
  | `cortecs/gpt-5.1` | 400K | | | | | | $1 | $11 |
88
89
  | `cortecs/gpt-5.4` | 1.1M | | | | | | $3 | $15 |
89
90
  | `cortecs/gpt-5.6-luna` | 1.1M | | | | | | $0.22 | $1 |
90
- | `cortecs/gpt-5.6-sol` | 1.1M | | | | | | $6 | $33 |
91
+ | `cortecs/gpt-5.6-sol` | 1.1M | | | | | | $4 | $22 |
91
92
  | `cortecs/gpt-5.6-terra` | 1.1M | | | | | | $2 | $13 |
93
+ | `cortecs/gpt-6-luna` | 1.1M | | | | | | $0.12 | $0.60 |
94
+ | `cortecs/gpt-6-sol` | 1.1M | | | | | | $2 | $12 |
92
95
  | `cortecs/gpt-oss-120b` | 131K | | | | | | $0.09 | $0.45 |
93
96
  | `cortecs/gpt-oss-20b` | 131K | | | | | | $0.04 | $0.17 |
94
97
  | `cortecs/gpt-oss-safeguard-120b` | 128K | | | | | | $0.18 | $0.70 |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Deep Infra logo](https://models.dev/logos/deepinfra.svg)Deep Infra
6
6
 
7
- Access 68 Deep Infra models through Mastra's model router. Authentication is handled automatically using the `DEEPINFRA_API_KEY` environment variable.
7
+ Access 70 Deep Infra models through Mastra's model router. Authentication is handled automatically using the `DEEPINFRA_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Deep Infra documentation](https://deepinfra.com/models).
10
10
 
@@ -91,11 +91,13 @@ for await (const chunk of stream) {
91
91
  | `deepinfra/thinkingmachines/Inkling-Small` | 524K | | | | | | $0.45 | $1 |
92
92
  | `deepinfra/XiaomiMiMo/MiMo-V2.5` | 262K | | | | | | $0.14 | $0.28 |
93
93
  | `deepinfra/XiaomiMiMo/MiMo-V2.5-Pro` | 1.0M | | | | | | $1 | $3 |
94
+ | `deepinfra/XiaomiMiMo/MiMo-V2.6-Flash` | 1.0M | | | | | | $0.14 | $0.28 |
95
+ | `deepinfra/XiaomiMiMo/MiMo-V2.6-Pro` | 1.0M | | | | | | $0.43 | $0.87 |
94
96
  | `deepinfra/zai-org/GLM-4.6` | 203K | | | | | | $0.50 | $2 |
95
97
  | `deepinfra/zai-org/GLM-4.7` | 203K | | | | | | $0.40 | $2 |
96
98
  | `deepinfra/zai-org/GLM-5.1` | 203K | | | | | | $1 | $4 |
97
99
  | `deepinfra/zai-org/GLM-5.2` | 1.0M | | | | | | $0.75 | $2 |
98
- | `deepinfra/zai-org/GLM-5.3` | 1.0M | | | | | | $1 | $4 |
100
+ | `deepinfra/zai-org/GLM-5.3` | 1.0M | | | | | | $0.90 | $4 |
99
101
  | `deepinfra/zai-org/GLM-5.3-Flash` | 1.0M | | | | | | $0.15 | $0.50 |
100
102
 
101
103
  Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Eden AI logo](https://models.dev/logos/edenai.svg)Eden AI
6
6
 
7
- Access 283 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
7
+ Access 285 Eden AI models through Mastra's model router. Authentication is handled automatically using the `EDENAI_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Eden AI documentation](https://docs.edenai.co).
10
10
 
@@ -136,11 +136,13 @@ for await (const chunk of stream) {
136
136
  | `edenai/fireworks_ai/accounts/fireworks/models/inkling` | 1.0M | | | | | | $1 | $4 |
137
137
  | `edenai/fireworks_ai/accounts/fireworks/models/muse-glimmer-30b` | 131K | | | | | | $0.35 | $2 |
138
138
  | `edenai/fireworks_ai/gpt-oss-120b` | 131K | | | | | | $0.15 | $0.60 |
139
- | `edenai/flexai/DeepSeek-V4-Flash-0731` | 1.0M | | | | | | $0.07 | $0.18 |
140
- | `edenai/flexai/gpt-oss-120b` | 131K | | | | | | $0.04 | $0.17 |
141
- | `edenai/flexai/gpt-oss-20b` | 131K | | | | | | $0.02 | $0.09 |
139
+ | `edenai/flexai/DeepSeek-V4-Flash-0731` | 1.0M | | | | | | $0.06 | $0.18 |
140
+ | `edenai/flexai/gpt-oss-120b` | 131K | | | | | | $0.03 | $0.17 |
141
+ | `edenai/flexai/gpt-oss-20b` | 131K | | | | | | $0.02 | $0.10 |
142
142
  | `edenai/flexai/Muse-Glimmer-30B` | 131K | | | | | | $0.30 | $1 |
143
143
  | `edenai/flexai/Step-3.7-Flash` | 262K | | | | | | $0.20 | $1 |
144
+ | `edenai/google/deep-research-max-preview-04-2026` | 131K | | | | | | $2 | $12 |
145
+ | `edenai/google/deep-research-preview-04-2026` | 131K | | | | | | $2 | $12 |
144
146
  | `edenai/google/gemini-2.5-flash-image` | 33K | | | | | | $0.30 | $3 |
145
147
  | `edenai/google/gemini-3-flash-preview` | 1.0M | | | | | | $0.50 | $3 |
146
148
  | `edenai/google/gemini-3-pro-image` | 66K | | | | | | $2 | $12 |
@@ -62,9 +62,9 @@ for await (const chunk of stream) {
62
62
  | `empiriolabs/kimi-k3` | 1.0M | | | | | | $3 | $15 |
63
63
  | `empiriolabs/mimo-v2-5` | 1.0M | | | | | | $0.70 | $1 |
64
64
  | `empiriolabs/mimo-v2-5-pro` | 1.0M | | | | | | $2 | $4 |
65
- | `empiriolabs/mimo-v2-6-flash` | 1.0M | | | | | | $0.70 | $1 |
66
- | `empiriolabs/mimo-v2-6-pro` | 1.0M | | | | | | $2 | $4 |
67
- | `empiriolabs/mimo-v2-6-pro-ultraspeed` | 1.0M | | | | | | $22 | $44 |
65
+ | `empiriolabs/mimo-v2-6-flash` | 1.0M | | | | | | $0.14 | $0.28 |
66
+ | `empiriolabs/mimo-v2-6-pro` | 1.0M | | | | | | $0.43 | $0.87 |
67
+ | `empiriolabs/mimo-v2-6-pro-ultraspeed` | 1.0M | | | | | | $4 | $9 |
68
68
  | `empiriolabs/minimax-m2-7` | 200K | | | | | | $0.15 | $0.60 |
69
69
  | `empiriolabs/minimax-m2-7-highspeed` | 200K | | | | | | $0.30 | $1 |
70
70
  | `empiriolabs/minimax-m3` | 1.0M | | | | | | $0.23 | $0.90 |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Fireworks AI logo](https://models.dev/logos/fireworks-ai.svg)Fireworks AI
6
6
 
7
- Access 33 Fireworks AI models through Mastra's model router. Authentication is handled automatically using the `FIREWORKS_API_KEY` environment variable.
7
+ Access 34 Fireworks AI models through Mastra's model router. Authentication is handled automatically using the `FIREWORKS_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Fireworks AI documentation](https://fireworks.ai/docs/).
10
10
 
@@ -39,6 +39,7 @@ for await (const chunk of stream) {
39
39
  | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
40
40
  | ----------------------------------------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
41
41
  | `fireworks-ai/accounts/fireworks/models/deepseek-v4p1-flash` | 1.0M | | | | | | $0.22 | $0.66 |
42
+ | `fireworks-ai/accounts/fireworks/models/ember-1` | 1.0M | | | | | | $3 | $15 |
42
43
  | `fireworks-ai/accounts/fireworks/models/glm-5p3` | 1.0M | | | | | | $1 | $4 |
43
44
  | `fireworks-ai/accounts/fireworks/models/glm-5p3-flash` | 1.0M | | | | | | $0.15 | $0.50 |
44
45
  | `fireworks-ai/accounts/fireworks/models/gpt-oss-120b` | 131K | | | | | | $0.15 | $0.60 |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Kilo Gateway logo](https://models.dev/logos/kilo.svg)Kilo Gateway
6
6
 
7
- Access 393 Kilo Gateway models through Mastra's model router. Authentication is handled automatically using the `KILO_API_KEY` environment variable.
7
+ Access 391 Kilo Gateway models through Mastra's model router. Authentication is handled automatically using the `KILO_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Kilo Gateway documentation](https://kilo.ai).
10
10
 
@@ -42,8 +42,8 @@ for await (const chunk of stream) {
42
42
  | `kilo/~anthropic/claude-haiku-latest` | 200K | | | | | | $1 | $5 |
43
43
  | `kilo/~anthropic/claude-opus-latest` | 1.0M | | | | | | $4 | $20 |
44
44
  | `kilo/~anthropic/claude-sonnet-latest` | 1.0M | | | | | | $2 | $10 |
45
- | `kilo/~deepseek/deepseek-flash-latest` | 1.0M | | | | | | $0.10 | $0.60 |
46
- | `kilo/~deepseek/deepseek-pro-latest` | 1.0M | | | | | | $0.39 | $1 |
45
+ | `kilo/~deepseek/deepseek-flash-latest` | 1.0M | | | | | | $0.05 | $0.25 |
46
+ | `kilo/~deepseek/deepseek-pro-latest` | 1.0M | | | | | | $0.39 | $3 |
47
47
  | `kilo/~deepseek/deepseek-v4-flash-latest` | 1.0M | | | | | | $0.04 | $0.55 |
48
48
  | `kilo/~google/gemini-flash-latest` | 1.0M | | | | | | $0.75 | $4 |
49
49
  | `kilo/~google/gemini-pro-latest` | 1.0M | | | | | | $2 | $12 |
@@ -54,8 +54,8 @@ for await (const chunk of stream) {
54
54
  | `kilo/~openai/gpt-sol-latest` | 1.1M | | | | | | $2 | $10 |
55
55
  | `kilo/~openai/gpt-terra-latest` | 1.1M | | | | | | $2 | $12 |
56
56
  | `kilo/~x-ai/grok-latest` | 500K | | | | | | $2 | $5 |
57
- | `kilo/~z-ai/glm-flash-latest` | 1.0M | | | | | | $0.07 | $0.25 |
58
- | `kilo/~z-ai/glm-latest` | 1.0M | | | | | | $0.56 | $2 |
57
+ | `kilo/~z-ai/glm-flash-latest` | 1.0M | | | | | | $0.04 | $0.14 |
58
+ | `kilo/~z-ai/glm-latest` | 1.0M | | | | | | $0.56 | $3 |
59
59
  | `kilo/aion-labs/aion-2.0` | 131K | | | | | | $0.80 | $2 |
60
60
  | `kilo/aion-labs/aion-3.0` | 131K | | | | | | $3 | $6 |
61
61
  | `kilo/aion-labs/aion-3.0-mini` | 131K | | | | | | $0.70 | $1 |
@@ -182,7 +182,7 @@ for await (const chunk of stream) {
182
182
  | `kilo/microsoft/wizardlm-2-8x22b` | 66K | | | | | | $0.62 | $0.62 |
183
183
  | `kilo/minimax/minimax-01` | 1.0M | | | | | | $0.20 | $1 |
184
184
  | `kilo/minimax/minimax-m1` | 1.0M | | | | | | $0.40 | $2 |
185
- | `kilo/minimax/minimax-m2` | 205K | | | | | | $0.30 | $1 |
185
+ | `kilo/minimax/minimax-m2` | 197K | | | | | | $0.30 | $1 |
186
186
  | `kilo/minimax/minimax-m2-her` | 66K | | | | | | $0.30 | $1 |
187
187
  | `kilo/minimax/minimax-m2.1` | 205K | | | | | | $0.30 | $1 |
188
188
  | `kilo/minimax/minimax-m2.5` | 200K | | | | | | $0.30 | $1 |
@@ -214,9 +214,7 @@ for await (const chunk of stream) {
214
214
  | `kilo/moonshotai/kimi-k3` | 1.0M | | | | | | $3 | $15 |
215
215
  | `kilo/morph/morph-v3-fast` | 82K | | | | | | $0.80 | $1 |
216
216
  | `kilo/morph/morph-v3-large` | 262K | | | | | | $0.90 | $2 |
217
- | `kilo/nex-agi/nex-n2.5-mini` | 262K | | | | | | $0.03 | $0.10 |
218
217
  | `kilo/nex-agi/nex-n2.5-mini:free` | 262K | | | | | | — | — |
219
- | `kilo/nex-agi/nex-n2.5-pro` | 262K | | | | | | $0.07 | $0.25 |
220
218
  | `kilo/nex-agi/nex-n2.5-pro:free` | 262K | | | | | | — | — |
221
219
  | `kilo/nousresearch/hermes-3-llama-3.1-405b` | 131K | | | | | | $1 | $1 |
222
220
  | `kilo/nousresearch/hermes-3-llama-3.1-70b` | 131K | | | | | | $0.70 | $0.70 |
@@ -385,9 +383,9 @@ for await (const chunk of stream) {
385
383
  | `kilo/stepfun/step-3.7-flash:free` | 262K | | | | | | — | — |
386
384
  | `kilo/tencent/hunyuan-a13b-instruct` | 131K | | | | | | $0.14 | $0.57 |
387
385
  | `kilo/tencent/hy-mt2-1.8b` | 8K | | | | | | $0.04 | $0.18 |
388
- | `kilo/tencent/hy-mt2-30b-a3b` | 8K | | | | | | $0.07 | $0.29 |
386
+ | `kilo/tencent/hy-mt2-30b-a3b` | 8K | | | | | | $0.07 | $0.28 |
389
387
  | `kilo/tencent/hy-mt2-7b` | 8K | | | | | | $0.07 | $0.29 |
390
- | `kilo/tencent/hy3` | 262K | | | | | | $0.08 | $0.33 |
388
+ | `kilo/tencent/hy3` | 262K | | | | | | $0.13 | $0.53 |
391
389
  | `kilo/tencent/hy3-preview` | 262K | | | | | | $0.18 | $0.60 |
392
390
  | `kilo/tencent/hy4-preview` | 1.0M | | | | | | $0.83 | $3 |
393
391
  | `kilo/thedrummer/cydonia-24b-v4.1` | 131K | | | | | | $0.30 | $0.50 |
@@ -393,8 +393,8 @@ for await (const chunk of stream) {
393
393
  | `nano-gpt/openai/gpt-5.6-terra-pro` | 1.1M | | | | | | $2 | $12 |
394
394
  | `nano-gpt/openai/gpt-6-astra` | 1.1M | | | | | | $10 | $50 |
395
395
  | `nano-gpt/openai/gpt-6-astra-pro` | 1.1M | | | | | | $10 | $50 |
396
- | `nano-gpt/openai/gpt-6-luna` | 1.1M | | | | | | $0.05 | $0.25 |
397
- | `nano-gpt/openai/gpt-6-luna-pro` | 1.1M | | | | | | $0.05 | $0.25 |
396
+ | `nano-gpt/openai/gpt-6-luna` | 1.1M | | | | | | $0.10 | $0.50 |
397
+ | `nano-gpt/openai/gpt-6-luna-pro` | 1.1M | | | | | | $0.10 | $0.50 |
398
398
  | `nano-gpt/openai/gpt-6-sol` | 1.1M | | | | | | $2 | $10 |
399
399
  | `nano-gpt/openai/gpt-6-sol-pro` | 1.1M | | | | | | $2 | $10 |
400
400
  | `nano-gpt/openai/gpt-astra-latest` | 1.1M | | | | | | $10 | $50 |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Tempr logo](https://models.dev/logos/tempr.svg)Tempr
6
6
 
7
- Access 29 Tempr models through Mastra's model router. Authentication is handled automatically using the `TEMPR_API_KEY` environment variable.
7
+ Access 39 Tempr models through Mastra's model router. Authentication is handled automatically using the `TEMPR_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Tempr documentation](https://temprhq.io/docs/gateway-reference.html).
10
10
 
@@ -67,6 +67,16 @@ for await (const chunk of stream) {
67
67
  | `tempr/google/gemini-flash-lite-latest` | 1.0M | | | | | | $0.30 | $3 |
68
68
  | `tempr/google/gemma-4-26b-a4b-it` | 262K | | | | | | — | — |
69
69
  | `tempr/google/gemma-4-31b-it` | 262K | | | | | | — | — |
70
+ | `tempr/mistral/mistral-embed` | 8K | | | | | | $0.10 | — |
71
+ | `tempr/mistral/mistral-large-2512` | 262K | | | | | | $0.50 | $2 |
72
+ | `tempr/mistral/mistral-large-latest` | 262K | | | | | | $0.50 | $2 |
73
+ | `tempr/mistral/mistral-medium-2604` | 262K | | | | | | $2 | $8 |
74
+ | `tempr/mistral/mistral-medium-latest` | 262K | | | | | | $2 | $8 |
75
+ | `tempr/mistral/mistral-small-2603` | 256K | | | | | | $0.15 | $0.60 |
76
+ | `tempr/mistral/mistral-small-latest` | 256K | | | | | | $0.15 | $0.60 |
77
+ | `tempr/mistral/voxtral-small-latest` | 32K | | | | | | $0.10 | $0.30 |
78
+ | `tempr/mistral/zai-glm-5-2` | 1.0M | | | | | | $1 | $4 |
79
+ | `tempr/mistral/zai-glm-5-3` | 1.0M | | | | | | $1 | $4 |
70
80
 
71
81
  Model availability, capabilities, context windows, and pricing are sourced from [models.dev](https://models.dev) and may change.
72
82
 
@@ -98,7 +108,7 @@ const agent = new Agent({
98
108
  model: ({ requestContext }) => {
99
109
  const useAdvanced = requestContext.task === "complex";
100
110
  return useAdvanced
101
- ? "tempr/google/gemma-4-31b-it"
111
+ ? "tempr/mistral/zai-glm-5-3"
102
112
  : "tempr/anthropic/claude-fable-5";
103
113
  }
104
114
  });
@@ -46,7 +46,7 @@ for await (const chunk of stream) {
46
46
  | `togetherai/LiquidAI/LFM2-24B-A2B` | 33K | | | | | | $0.03 | $0.12 |
47
47
  | `togetherai/meta-llama/Llama-3.3-70B-Instruct-Turbo` | 131K | | | | | | $1 | $1 |
48
48
  | `togetherai/meta-llama/Meta-Llama-3-8B-Instruct-Lite` | 8K | | | | | | $0.14 | $0.14 |
49
- | `togetherai/MiniMaxAI/MiniMax-M2.7` | 203K | | | | | | $0.30 | $1 |
49
+ | `togetherai/MiniMaxAI/MiniMax-M2.7` | 197K | | | | | | $0.30 | $1 |
50
50
  | `togetherai/MiniMaxAI/MiniMax-M3` | 524K | | | | | | $0.30 | $1 |
51
51
  | `togetherai/moonshotai/Kimi-K2.6` | 262K | | | | | | $1 | $5 |
52
52
  | `togetherai/moonshotai/Kimi-K2.7-Code` | 262K | | | | | | $0.95 | $4 |
@@ -4,7 +4,7 @@
4
4
 
5
5
  # ![Vivgrid logo](https://models.dev/logos/vivgrid.svg)Vivgrid
6
6
 
7
- Access 31 Vivgrid models through Mastra's model router. Authentication is handled automatically using the `VIVGRID_API_KEY` environment variable.
7
+ Access 34 Vivgrid models through Mastra's model router. Authentication is handled automatically using the `VIVGRID_API_KEY` environment variable.
8
8
 
9
9
  Learn more in the [Vivgrid documentation](https://docs.vivgrid.com/models).
10
10
 
@@ -41,6 +41,7 @@ for await (const chunk of stream) {
41
41
  | `vivgrid/claude-fable-5` | 1.0M | | | | | | $10 | $50 |
42
42
  | `vivgrid/claude-fable-5-1` | 1.0M | | | | | | $10 | $50 |
43
43
  | `vivgrid/claude-opus-5` | 1.0M | | | | | | $5 | $25 |
44
+ | `vivgrid/claude-opus-5-5` | 1.0M | | | | | | $4 | $20 |
44
45
  | `vivgrid/claude-sonnet-5` | 1.0M | | | | | | $2 | $10 |
45
46
  | `vivgrid/deepseek-v3.2` | 128K | | | | | | $0.28 | $0.42 |
46
47
  | `vivgrid/deepseek-v4-flash` | 1.0M | | | | | | $0.15 | $0.30 |
@@ -67,6 +68,8 @@ for await (const chunk of stream) {
67
68
  | `vivgrid/gpt-5.6-sol` | 1.1M | | | | | | $5 | $30 |
68
69
  | `vivgrid/gpt-5.6-terra` | 1.1M | | | | | | $3 | $15 |
69
70
  | `vivgrid/gpt-6-astra` | 1.1M | | | | | | $10 | $50 |
71
+ | `vivgrid/gpt-6-luna` | 1.1M | | | | | | $0.10 | $0.50 |
72
+ | `vivgrid/gpt-6-sol` | 1.1M | | | | | | $2 | $10 |
70
73
  | `vivgrid/kimi-k3` | 1.0M | | | | | | $3 | $15 |
71
74
  | `vivgrid/viv-fast` | 1.0M | | | | | | $0.13 | $0.40 |
72
75
 
@@ -138,7 +138,7 @@ await mastraClient.deleteFeedback({
138
138
 
139
139
  Returns `Promise<{ success: boolean }>` from `mastraClient.deleteFeedback()` and the HTTP route. The storage domain method returns `Promise<void>`. Requests with more than 1,000 ids return `400`, and a server running `@mastra/core` older than `1.66.0` returns `501`. ClickHouse vNext, PostgreSQL vNext, DuckDB, and the in-memory store implement deletion. Every other adapter, including LibSQL and MongoDB, throws `OBSERVABILITY_STORAGE_DELETE_FEEDBACK_NOT_IMPLEMENTED`.
140
140
 
141
- On ClickHouse, deletion uses lightweight deletes on the main feedback events table to remove rows from reads, including OLAP queries, without guaranteeing immediate physical removal, so open-source deployments must configure an [observability retention period](https://mastra.ai/reference/storage/retention) to physically purge them. ClickHouse applies a TTL to deletion requests only when all five signals have finite retention. See [ClickHouse native TTL](https://mastra.ai/reference/storage/retention). Delete APIs leave the separate delta cursor table untouched. Its rows contain identifiers rather than feedback payloads and expire within two days.
141
+ On ClickHouse, deletion uses lightweight deletes on the main feedback events table to remove rows from reads, including OLAP queries, without guaranteeing immediate physical removal, so open-source deployments must configure an [observability retention period](https://mastra.ai/reference/storage/retention) to physically purge them. Each request is marked applied once its delete succeeds; if the delete fails, the request stays unapplied, doesn't block `updateFeedbackReviewStatus()` on the still-visible feedback, and you retry by calling `deleteFeedback()` again. Open-source deployments have no background reconciler. ClickHouse applies a TTL to deletion requests only when all five signals have finite retention. See [ClickHouse native TTL](https://mastra.ai/reference/storage/retention). Delete APIs leave the separate delta cursor table untouched. Its rows contain identifiers rather than feedback payloads and expire within two days.
142
142
 
143
143
  ### `deleteScores(args)`
144
144
 
@@ -49,7 +49,7 @@ interface BatchDeleteTracesArgs {
49
49
 
50
50
  When `organizationId` or `resourceId` is provided, only records matching the scope are deleted. ClickHouse vNext, PostgreSQL vNext, DuckDB, and the in-memory store support scoped deletion. Other storage adapters throw `OBSERVABILITY_STORAGE_BATCH_DELETE_TRACES_SCOPE_NOT_SUPPORTED` rather than applying an unscoped delete.
51
51
 
52
- For ClickHouse vNext, the method records the deletion predicate and waits for the lightweight delete masks to be applied before it resolves. Lightweight deletion is hide-only through ClickHouse's `_row_exists` mask. Physical removal depends on merges and the retention TTLs you configure. See [ClickHouse trace deletion and retention](https://mastra.ai/integrations/databases/clickhouse) and [ClickHouse native TTL](https://mastra.ai/reference/storage/retention).
52
+ For ClickHouse vNext, the method records the deletion predicate and waits for the lightweight delete masks to be applied before it resolves. Lightweight deletion is hide-only through ClickHouse's `_row_exists` mask. Physical removal depends on merges and the retention TTLs you configure. Each request is marked applied once every delete succeeds; if a delete fails, the request stays unapplied and you retry by calling `batchDeleteTraces()` again. Mastra OSS has no background reconciler. See [ClickHouse trace deletion and retention](https://mastra.ai/integrations/databases/clickhouse) and [ClickHouse native TTL](https://mastra.ai/reference/storage/retention).
53
53
 
54
54
  ### `SpanTypeMap`
55
55
 
@@ -51,7 +51,7 @@ Trailing user messages are treated as new messages and kept as sent. When they c
51
51
 
52
52
  ### Stored messages as the base layer
53
53
 
54
- Loaders add stored messages with a `memory` source. When a stored message and an input message share an ID, the stored copy takes the slot with its reasoning, provider metadata, and `createdAt`, and the input's parts are layered on top. That preserves the original assistant turn while still persisting new contributions such as tool results. A tool outcome from the input only fills in a call that's still pending in the stored copy, so an echo can't overwrite a stored result.
54
+ Loaders add stored messages with a `memory` source. When a stored message and an input message share an ID, the stored copy keeps its text, reasoning, provider metadata, and `createdAt`. The only thing taken from the input is a tool outcome for a call that's still pending in the stored copy, such as a client tool result or an approval answer. Text and metadata from the input are ignored, so a client can't change a stored message by sending a different copy of it. To change a stored message, update it in storage. This also applies with `retainFullInput`.
55
55
 
56
56
  ## Related
57
57
 
@@ -4,10 +4,10 @@
4
4
 
5
5
  # TokenLimiterProcessor
6
6
 
7
- The `TokenLimiterProcessor` limits the number of tokens in messages. It can be used as an input, per-step input, and output processor:
7
+ The `TokenLimiterProcessor` limits the number of tokens in messages. Depending on `trimMode`, it acts as a prompt processor, an input processor, and an output processor:
8
8
 
9
- - **Input processor** (`processInput`): Filters historical messages to fit within the context window before the agentic loop starts, prioritizing recent messages
10
- - **Per-step input processor** (`processInputStep`): Prunes messages at each step of a multi-step agent workflow, preventing unbounded token growth when tools trigger additional LLM calls
9
+ - **Prompt processor** (`processLLMRequest`): In the default `best-fit` and `contiguous` trim modes, enforces the input budget on the provider prompt right before each model call, at every step of the agentic loop. The prompt is measured after earlier prompt processors (such as `ToolCallFilter`) have transformed it, so only tokens that actually reach the model are counted. Tool call and tool result messages are grouped so they're kept or removed together, and trimming is transient: stored messages are never modified.
10
+ - **Input processor** (`processInput`): In `memory-only` trim mode, filters historical messages to fit within the context window before the agentic loop starts, prioritizing recent messages
11
11
  - **Output processor**: Limits generated response tokens via streaming or non-streaming with configurable strategies for handling exceeded limits
12
12
 
13
13
  ## Usage example
@@ -34,7 +34,7 @@ const processor = new TokenLimiterProcessor({
34
34
 
35
35
  **options.countMode** (`'cumulative' | 'part'`): Whether to count tokens from the beginning of the stream or just the current part: 'cumulative' counts all tokens from start, 'part' only counts tokens in current part
36
36
 
37
- **options.trimMode** (`'best-fit' | 'contiguous'`): Controls how messages are trimmed when exceeding the token limit: 'best-fit' keeps as many messages as possible (may create gaps), 'contiguous' stops at the first message that does not fit, ensuring a continuous suffix of conversation history
37
+ **options.trimMode** (`'best-fit' | 'contiguous' | 'memory-only'`): Controls how the token limit is enforced: 'best-fit' trims the provider prompt while keeping as many messages as possible (may create gaps), 'contiguous' trims the provider prompt but stops at the first message that does not fit (keeping a continuous suffix of conversation history), and 'memory-only' trims stored history in processInput instead of the prompt
38
38
 
39
39
  ## Returns
40
40
 
@@ -42,9 +42,11 @@ const processor = new TokenLimiterProcessor({
42
42
 
43
43
  **name** (`string`): Optional processor display name
44
44
 
45
- **processInput** (`(args: { messages: MastraDBMessage[]; abort: (reason?: string) => never }) => Promise<MastraDBMessage[]>`): Filters input messages to fit within token limit before the agentic loop starts, prioritizing recent messages while preserving system messages
45
+ **processInput** (`(args: { messages: MastraDBMessage[]; abort: (reason?: string) => never }) => Promise<MastraDBMessage[]>`): Trims stored history to fit within the token limit before the agentic loop starts in 'memory-only' trim mode, prioritizing recent messages while preserving system messages and the current turn
46
46
 
47
- **processInputStep** (`(args: ProcessInputStepArgs) => Promise<void>`): Prunes messages at each step of the agentic loop (including tool call continuations) to keep the conversation within the token limit. Mutates the messageList directly by removing oldest messages first while preserving system messages.
47
+ **processInputStep** (`(args: ProcessInputStepArgs) => Promise<void>`): In 'memory-only' trim mode, applies stored-history trimming at each step. In the 'best-fit' and 'contiguous' trim modes, it leaves trimming to processLLMRequest when the agent runs that method for this processor. Otherwise it trims stored history, for example in generateLegacy() and streamLegacy() or when the limiter is inside a processor workflow.
48
+
49
+ **processLLMRequest** (`(args: ProcessLLMRequestArgs) => Promise<ProcessLLMRequestResult>`): Enforces the input budget on the provider prompt in the 'best-fit' and 'contiguous' trim modes. Runs after earlier prompt processors have transformed the prompt, counts that exact prompt, and returns a trimmed copy for the model call only. System messages are always preserved, and tool call and tool result messages are grouped so they are kept or removed together.
48
50
 
49
51
  **processOutputStream** (`(args: ProcessOutputStreamArgs) => Promise<ChunkType | null>`): Processes streaming output parts to limit token count during streaming. Only text and object parts count against the limit and can be withheld; lifecycle, reasoning and tool parts always pass through.
50
52
 
@@ -72,10 +74,11 @@ Images and file attachments are estimated instead of tokenized, including `file`
72
74
 
73
75
  ## Error behavior
74
76
 
75
- When used as an input processor (both `processInput` and `processInputStep`), `TokenLimiterProcessor` throws a `TripWire` error in the following cases:
77
+ When trimming input, `TokenLimiterProcessor` throws a `TripWire` error in the following cases:
76
78
 
77
- - **Empty messages**: If there are no messages to process, a TripWire is thrown because you can't send an LLM request with no messages.
79
+ - **Empty messages**: If there are no non-system messages to process, a TripWire is thrown because you can't send an LLM request with no messages.
78
80
  - **System messages exceed limit**: If system messages alone exceed the token limit, a TripWire is thrown because you can't send an LLM request with only system messages and no user/assistant messages.
81
+ - **No messages fit**: If no message fits within the remaining token budget, a TripWire is thrown because you can't send an LLM request with no messages.
79
82
 
80
83
  ```typescript
81
84
  import { TripWire } from '@mastra/core/agent'
@@ -112,9 +115,9 @@ export const agent = new Agent({
112
115
  })
113
116
  ```
114
117
 
115
- ### As a per-step input processor (limit multi-step token growth)
118
+ ### As a per-step processor (limit multi-step token growth)
116
119
 
117
- When an agent uses tools across multiple steps (e.g. `maxSteps > 1`), each step accumulates conversation history from all previous steps. Use `inputProcessors` to also limit tokens at each step of the agentic loop. The `TokenLimiterProcessor` automatically applies to both the initial input and every subsequent step:
120
+ When an agent uses tools across multiple steps (e.g. `maxSteps > 1`), each step accumulates conversation history from all previous steps. `TokenLimiterProcessor` applies its limit at every step, measuring the prompt that's about to be sent after any earlier prompt processors have run. Register prompt-shrinking processors such as `ToolCallFilter` before it, so the limiter counts the prompt the model actually receives:
118
121
 
119
122
  ```typescript
120
123
  import { Agent } from '@mastra/core/agent'
@@ -136,6 +139,8 @@ const result = await agent.generate('Research this topic using your tools', {
136
139
  })
137
140
  ```
138
141
 
142
+ Processor workflows don't run `processLLMRequest`, so a `TokenLimiterProcessor` inside a processor workflow trims stored messages in `processInputStep`, before prompt processors such as `ToolCallFilter` remove anything from the request. To count the prompt the model receives, add the limiter directly to `inputProcessors`.
143
+
139
144
  ### As an output processor (limit response length)
140
145
 
141
146
  Use `outputProcessors` to limit the length of generated responses:
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mastra/mcp-docs-server",
3
- "version": "1.3.0-alpha.4",
3
+ "version": "1.3.0-alpha.6",
4
4
  "description": "MCP server for accessing Mastra.ai documentation, changelogs, and news.",
5
5
  "type": "module",
6
6
  "main": "dist/index.js",
@@ -26,8 +26,8 @@
26
26
  "@mastra/mcp-legacy": "npm:@mastra/mcp@^1.18.0",
27
27
  "local-pkg": "^1.1.2",
28
28
  "zod": "^4.6.4",
29
- "@mastra/core": "1.70.0-alpha.2",
30
- "@mastra/mcp": "^2.1.0-alpha.0"
29
+ "@mastra/mcp": "^2.1.0-alpha.0",
30
+ "@mastra/core": "1.70.0-alpha.4"
31
31
  },
32
32
  "devDependencies": {
33
33
  "@hono/node-server": "^2.0.0",
@@ -43,9 +43,9 @@
43
43
  "tsx": "^4.23.1",
44
44
  "typescript": "^7.0.2",
45
45
  "vitest": "4.1.11",
46
- "@mastra/core": "1.70.0-alpha.2",
47
- "@internal/lint": "0.0.135",
48
- "@internal/types-builder": "0.0.110"
46
+ "@internal/types-builder": "0.0.110",
47
+ "@mastra/core": "1.70.0-alpha.4",
48
+ "@internal/lint": "0.0.135"
49
49
  },
50
50
  "homepage": "https://mastra.ai",
51
51
  "repository": {