@mastra/mcp-docs-server 1.2.18-alpha.3 → 1.2.18-alpha.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/.docs/docs/datasets/running-experiments.md +86 -1
  2. package/.docs/docs/mastra-platform/deploy.md +3 -1
  3. package/.docs/docs/mastra-platform/workspaces.md +6 -3
  4. package/.docs/docs/sandbox/filesystem.md +120 -139
  5. package/.docs/docs/sandbox/lsp.md +195 -143
  6. package/.docs/docs/sandbox/overview.md +103 -69
  7. package/.docs/docs/sandbox/search.md +172 -153
  8. package/.docs/docs/sandbox/skills.md +94 -151
  9. package/.docs/integrations/deploy/render.md +136 -89
  10. package/.docs/integrations/observability/arize.md +8 -6
  11. package/.docs/models/index.md +1 -1
  12. package/.docs/models/providers/edenai.md +2 -3
  13. package/.docs/models/providers/empiriolabs.md +1 -1
  14. package/.docs/models/providers/kilo.md +2 -2
  15. package/.docs/models/providers/llmgateway.md +2 -1
  16. package/.docs/models/providers/nano-gpt.md +3 -1
  17. package/.docs/models/providers/ofox.md +1 -1
  18. package/.docs/models/providers/opencode.md +65 -65
  19. package/.docs/reference/cli/mastra.md +2 -2
  20. package/.docs/reference/client-js/datasets.md +146 -0
  21. package/.docs/reference/configuration.md +58 -0
  22. package/.docs/reference/datasets/createExperiment.md +76 -0
  23. package/.docs/reference/datasets/finalizeExperiment.md +43 -0
  24. package/.docs/reference/datasets/runExperimentItem.md +55 -0
  25. package/.docs/reference/datasets/submitExperimentResult.md +56 -0
  26. package/.docs/reference/index.md +5 -0
  27. package/.docs/reference/observability/tracing/exporters/langfuse.md +2 -0
  28. package/.docs/reference/pubsub/redis-streams.md +11 -1
  29. package/.docs/reference/rag/metadata-filters.md +16 -8
  30. package/.docs/reference/rag/retrieval.md +113 -5
  31. package/.docs/reference/server/routes.md +111 -0
  32. package/CHANGELOG.md +14 -0
  33. package/package.json +3 -3
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ![OpenCode Zen logo](https://models.dev/logos/opencode.svg)OpenCode Zen
4
4
 
5
- Access 91 OpenCode Zen models through Mastra's model router. Authentication is handled automatically using the `OPENCODE_API_KEY` environment variable.
5
+ Access 92 OpenCode Zen models through Mastra's model router. Authentication is handled automatically using the `OPENCODE_API_KEY` environment variable.
6
6
 
7
7
  Learn more in the [OpenCode Zen documentation](https://opencode.ai/docs/zen).
8
8
 
@@ -34,70 +34,70 @@ for await (const chunk of stream) {
34
34
 
35
35
  ## Models
36
36
 
37
- | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
- | -------------------------------------- | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
- | `opencode/big-pickle` | 200K | | | | | | — | — |
40
- | `opencode/claude-fable-5` | 1.0M | | | | | | $10 | $50 |
41
- | `opencode/claude-haiku-4-5` | 200K | | | | | | $1 | $5 |
42
- | `opencode/claude-opus-4-5` | 200K | | | | | | $5 | $25 |
43
- | `opencode/claude-opus-4-6` | 1.0M | | | | | | $5 | $25 |
44
- | `opencode/claude-opus-4-7` | 1.0M | | | | | | $5 | $25 |
45
- | `opencode/claude-opus-4-8` | 1.0M | | | | | | $5 | $25 |
46
- | `opencode/claude-opus-5` | 1.0M | | | | | | $5 | $25 |
47
- | `opencode/claude-sonnet-4` | 1.0M | | | | | | $3 | $15 |
48
- | `opencode/claude-sonnet-4-5` | 1.0M | | | | | | $3 | $15 |
49
- | `opencode/claude-sonnet-4-6` | 1.0M | | | | | | $3 | $15 |
50
- | `opencode/claude-sonnet-5` | 1.0M | | | | | | $2 | $10 |
51
- | `opencode/deepseek-v4-flash` | 1.0M | | | | | | $0.14 | $0.28 |
52
- | `opencode/deepseek-v4-flash-free` | 200K | | | | | | — | — |
53
- | `opencode/deepseek-v4-pro` | 1.0M | | | | | | $2 | $4 |
54
- | `opencode/gemini-3-flash` | 1.0M | | | | | | $0.50 | $3 |
55
- | `opencode/gemini-3.1-pro` | 1.0M | | | | | | $2 | $12 |
56
- | `opencode/gemini-3.5-flash` | 1.0M | | | | | | $2 | $9 |
57
- | `opencode/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
58
- | `opencode/gemini-3.6-flash` | 1.0M | | | | | | $2 | $8 |
59
- | `opencode/gemini-3.7-flash` | 1.0M | | | | | | $2 | $8 |
60
- | `opencode/glm-5` | 205K | | | | | | $1 | $3 |
61
- | `opencode/glm-5.1` | 205K | | | | | | $1 | $4 |
62
- | `opencode/glm-5.2` | 1.0M | | | | | | $1 | $4 |
63
- | `opencode/gpt-5` | 400K | | | | | | $1 | $9 |
64
- | `opencode/gpt-5-codex` | 400K | | | | | | $1 | $9 |
65
- | `opencode/gpt-5-nano` | 400K | | | | | | $0.05 | $0.40 |
66
- | `opencode/gpt-5.1` | 400K | | | | | | $1 | $9 |
67
- | `opencode/gpt-5.1-codex` | 400K | | | | | | $1 | $9 |
68
- | `opencode/gpt-5.1-codex-max` | 400K | | | | | | $1 | $10 |
69
- | `opencode/gpt-5.1-codex-mini` | 400K | | | | | | $0.25 | $2 |
70
- | `opencode/gpt-5.2` | 400K | | | | | | $2 | $14 |
71
- | `opencode/gpt-5.2-codex` | 400K | | | | | | $2 | $14 |
72
- | `opencode/gpt-5.3-codex` | 400K | | | | | | $2 | $14 |
73
- | `opencode/gpt-5.3-codex-spark` | 128K | | | | | | $2 | $14 |
74
- | `opencode/gpt-5.4` | 1.1M | | | | | | $3 | $15 |
75
- | `opencode/gpt-5.4-mini` | 400K | | | | | | $0.75 | $5 |
76
- | `opencode/gpt-5.4-nano` | 400K | | | | | | $0.20 | $1 |
77
- | `opencode/gpt-5.4-pro` | 1.1M | | | | | | $30 | $180 |
78
- | `opencode/gpt-5.5` | 1.1M | | | | | | $5 | $30 |
79
- | `opencode/gpt-5.5-pro` | 1.1M | | | | | | $30 | $180 |
80
- | `opencode/gpt-5.6-luna` | 1.1M | | | | | | $0.20 | $1 |
81
- | `opencode/gpt-5.6-sol` | 1.1M | | | | | | $3 | $15 |
82
- | `opencode/gpt-5.6-terra` | 1.1M | | | | | | $3 | $15 |
83
- | `opencode/grok-4.5` | 500K | | | | | | $2 | $6 |
84
- | `opencode/grok-4.6` | 500K | | | | | | $2 | $6 |
85
- | `opencode/grok-build-0.1` | 256K | | | | | | $1 | $2 |
86
- | `opencode/hy3-free` | 190K | | | | | | — | — |
87
- | `opencode/kimi-k2.5` | 262K | | | | | | $0.60 | $3 |
88
- | `opencode/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
89
- | `opencode/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
90
- | `opencode/kimi-k3` | 1.0M | | | | | | $3 | $15 |
91
- | `opencode/laguna-s-2.1-free` | 256K | | | | | | — | — |
92
- | `opencode/mimo-v2.5-free` | 200K | | | | | | — | — |
93
- | `opencode/minimax-m2.5` | 205K | | | | | | $0.30 | $1 |
94
- | `opencode/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
95
- | `opencode/minimax-m3` | 512K | | | | | | $0.30 | $1 |
96
- | `opencode/muse-spark-1.2` | 1.0M | | | | | | $1 | $4 |
97
- | `opencode/nemotron-3-ultra-free` | 1.0M | | | | | | — | — |
98
- | `opencode/nemotron-3.5-lightning-free` | 262K | | | | | | — | — |
99
- | `opencode/qwen3.5-plus` | 262K | | | | | | $0.20 | $1 |
100
- | `opencode/qwen3.6-plus` | 262K | | | | | | $0.50 | $3 |
37
+ | Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
38
+ | ------------------------------------------ | ------- | ----- | --------- | ----- | ----- | ----- | ---------- | ----------- |
39
+ | `opencode/big-pickle` | 200K | | | | | | — | — |
40
+ | `opencode/claude-fable-5` | 1.0M | | | | | | $10 | $50 |
41
+ | `opencode/claude-haiku-4-5` | 200K | | | | | | $1 | $5 |
42
+ | `opencode/claude-opus-4-5` | 200K | | | | | | $5 | $25 |
43
+ | `opencode/claude-opus-4-6` | 1.0M | | | | | | $5 | $25 |
44
+ | `opencode/claude-opus-4-7` | 1.0M | | | | | | $5 | $25 |
45
+ | `opencode/claude-opus-4-8` | 1.0M | | | | | | $5 | $25 |
46
+ | `opencode/claude-opus-5` | 1.0M | | | | | | $5 | $25 |
47
+ | `opencode/claude-sonnet-4` | 1.0M | | | | | | $3 | $15 |
48
+ | `opencode/claude-sonnet-4-5` | 1.0M | | | | | | $3 | $15 |
49
+ | `opencode/claude-sonnet-4-6` | 1.0M | | | | | | $3 | $15 |
50
+ | `opencode/claude-sonnet-5` | 1.0M | | | | | | $2 | $10 |
51
+ | `opencode/deepseek-v4-flash` | 1.0M | | | | | | $0.14 | $0.28 |
52
+ | `opencode/deepseek-v4-flash-free` | 200K | | | | | | — | — |
53
+ | `opencode/deepseek-v4-pro` | 1.0M | | | | | | $2 | $4 |
54
+ | `opencode/gemini-3-flash` | 1.0M | | | | | | $0.50 | $3 |
55
+ | `opencode/gemini-3.1-pro` | 1.0M | | | | | | $2 | $12 |
56
+ | `opencode/gemini-3.5-flash` | 1.0M | | | | | | $2 | $9 |
57
+ | `opencode/gemini-3.5-flash-lite` | 1.0M | | | | | | $0.30 | $3 |
58
+ | `opencode/gemini-3.6-flash` | 1.0M | | | | | | $2 | $8 |
59
+ | `opencode/gemini-3.7-flash` | 1.0M | | | | | | $2 | $8 |
60
+ | `opencode/glm-5` | 205K | | | | | | $1 | $3 |
61
+ | `opencode/glm-5.1` | 205K | | | | | | $1 | $4 |
62
+ | `opencode/glm-5.2` | 1.0M | | | | | | $1 | $4 |
63
+ | `opencode/gpt-5` | 400K | | | | | | $1 | $9 |
64
+ | `opencode/gpt-5-codex` | 400K | | | | | | $1 | $9 |
65
+ | `opencode/gpt-5-nano` | 400K | | | | | | $0.05 | $0.40 |
66
+ | `opencode/gpt-5.1` | 400K | | | | | | $1 | $9 |
67
+ | `opencode/gpt-5.1-codex` | 400K | | | | | | $1 | $9 |
68
+ | `opencode/gpt-5.1-codex-max` | 400K | | | | | | $1 | $10 |
69
+ | `opencode/gpt-5.1-codex-mini` | 400K | | | | | | $0.25 | $2 |
70
+ | `opencode/gpt-5.2` | 400K | | | | | | $2 | $14 |
71
+ | `opencode/gpt-5.2-codex` | 400K | | | | | | $2 | $14 |
72
+ | `opencode/gpt-5.3-codex` | 400K | | | | | | $2 | $14 |
73
+ | `opencode/gpt-5.3-codex-spark` | 128K | | | | | | $2 | $14 |
74
+ | `opencode/gpt-5.4` | 1.1M | | | | | | $3 | $15 |
75
+ | `opencode/gpt-5.4-mini` | 400K | | | | | | $0.75 | $5 |
76
+ | `opencode/gpt-5.4-nano` | 400K | | | | | | $0.20 | $1 |
77
+ | `opencode/gpt-5.4-pro` | 1.1M | | | | | | $30 | $180 |
78
+ | `opencode/gpt-5.5` | 1.1M | | | | | | $5 | $30 |
79
+ | `opencode/gpt-5.5-pro` | 1.1M | | | | | | $30 | $180 |
80
+ | `opencode/gpt-5.6-luna` | 1.1M | | | | | | $0.20 | $1 |
81
+ | `opencode/gpt-5.6-sol` | 1.1M | | | | | | $3 | $15 |
82
+ | `opencode/gpt-5.6-terra` | 1.1M | | | | | | $3 | $15 |
83
+ | `opencode/grok-4.5` | 500K | | | | | | $2 | $6 |
84
+ | `opencode/grok-4.6` | 500K | | | | | | $2 | $6 |
85
+ | `opencode/grok-build-0.1` | 256K | | | | | | $1 | $2 |
86
+ | `opencode/hy3-free` | 190K | | | | | | — | — |
87
+ | `opencode/kimi-k2.5` | 262K | | | | | | $0.60 | $3 |
88
+ | `opencode/kimi-k2.6` | 262K | | | | | | $0.95 | $4 |
89
+ | `opencode/kimi-k2.7-code` | 262K | | | | | | $0.95 | $4 |
90
+ | `opencode/kimi-k3` | 1.0M | | | | | | $3 | $15 |
91
+ | `opencode/mimo-v2.5-free` | 200K | | | | | | — | — |
92
+ | `opencode/minimax-m2.5` | 205K | | | | | | $0.30 | $1 |
93
+ | `opencode/minimax-m2.7` | 205K | | | | | | $0.30 | $1 |
94
+ | `opencode/minimax-m3` | 512K | | | | | | $0.30 | $1 |
95
+ | `opencode/muse-spark-1.2` | 1.0M | | | | | | $1 | $4 |
96
+ | `opencode/muse-spark-1.2-contributor-free` | 1.0M | | | | | | — | — |
97
+ | `opencode/nemotron-3-ultra-free` | 1.0M | | | | | | — | — |
98
+ | `opencode/nemotron-3.5-lightning-free` | 262K | | | | | | — | — |
99
+ | `opencode/qwen3.5-plus` | 262K | | | | | | $0.20 | $1 |
100
+ | `opencode/qwen3.6-plus` | 262K | | | | | | $0.50 | $3 |
101
101
 
102
102
  ## Advanced configuration
103
103
 
@@ -541,11 +541,11 @@ mastra env db create --kind neon --name my-app-db --region aws-us-east-1 --share
541
541
 
542
542
  #### `--kind`
543
543
 
544
- Database provider (required). One of `turso` or `neon`.
544
+ Database provider (required). One of `turso`, `neon`, or `redis`.
545
545
 
546
546
  #### `--name`
547
547
 
548
- Database name. Defaults to a name derived from the project slug (for example `my-app-db`).
548
+ Database name. Defaults to a name derived from the project slug and provider (for example `my-app-turso` or `my-app-redis`).
549
549
 
550
550
  #### `--region`
551
551
 
@@ -0,0 +1,146 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # Datasets API
4
+
5
+ The Datasets API exposes Mastra's dataset and experiment routes from `MastraClient`. This page covers the caller-driven experiment methods, which let an orchestrator you own (for example a Temporal workflow) drive the experiment loop while Mastra acts as the system of record. Create the experiment, then either have Mastra execute each item server-side with `runExperimentItem` or ingest results you computed yourself with `submitExperimentResult`, and call finalize when the run is done.
6
+
7
+ Item runs, result submission, and finalization are safe to retry. Creation is safe to retry only when the request includes a caller-supplied `id`; without one, each retry creates a new experiment.
8
+
9
+ - `createDatasetExperiment` is idempotent on a caller-supplied `id`
10
+ - `runExperimentItem` and `submitExperimentResult` upsert on `(experimentId, itemId, attempt)`
11
+ - `finalizeExperiment` returns the stored record if the experiment is already completed
12
+
13
+ ## Usage example
14
+
15
+ ```typescript
16
+ import { MastraClient } from '@mastra/client-js'
17
+
18
+ const client = new MastraClient({
19
+ baseUrl: 'http://localhost:4111',
20
+ })
21
+
22
+ // 1. Create the experiment. Pass your orchestrator's run ID so a
23
+ // retried workflow converges on the same experiment.
24
+ const { experimentId, totalItems, datasetVersion } = await client.createDatasetExperiment({
25
+ datasetId: 'clinical-triage-evals',
26
+ id: 'temporal-wf-run-42',
27
+ targetType: 'agent',
28
+ targetId: 'triage-agent',
29
+ scorerIds: ['clinical-judge'],
30
+ })
31
+
32
+ // 2. Run one item per activity. Mastra executes the agent, runs the
33
+ // scorers, and upserts the result. Retried activities converge on
34
+ // a single row per (experimentId, itemId, attempt).
35
+ const { result, scores } = await client.runExperimentItem({
36
+ datasetId: 'clinical-triage-evals',
37
+ experimentId,
38
+ itemId: 'item-1',
39
+ })
40
+
41
+ // 3. Finalize. The server computes succeeded/failed/skipped counts
42
+ // from the persisted rows.
43
+ const experiment = await client.finalizeExperiment({
44
+ datasetId: 'clinical-triage-evals',
45
+ experimentId,
46
+ })
47
+ ```
48
+
49
+ For pure ingestion, create the experiment without `targetType`/`targetId` and replace step 2 with `submitExperimentResult`.
50
+
51
+ ## createDatasetExperiment()
52
+
53
+ Creates an experiment without starting a run. No runner is spawned, and the dataset version is pinned at creation time. Include a target when Mastra should execute items via `runExperimentItem`. Omit the target for pure ingestion via `submitExperimentResult`.
54
+
55
+ **datasetId** (`string`): ID of the dataset the experiment runs against.
56
+
57
+ **id** (`string`): Caller-supplied experiment ID (for example a workflow run ID). Makes creation idempotent on retry. Conflicting reuse of an ID returns a 409.
58
+
59
+ **targetType** (`'agent' | 'workflow' | 'scorer'`): Type of target that runExperimentItem executes. Provide together with targetId, or omit both.
60
+
61
+ **targetId** (`string`): ID of the registered target. Provide together with targetType.
62
+
63
+ **scorerIds** (`string[]`): Run-level scorer IDs resolved server-side by runExperimentItem. Requires a target.
64
+
65
+ **name** (`string`): Human-readable experiment name.
66
+
67
+ **description** (`string`): Experiment description.
68
+
69
+ **metadata** (`Record<string, unknown>`): Arbitrary metadata stored on the experiment.
70
+
71
+ **version** (`number`): Dataset version to pin. Defaults to the current dataset version.
72
+
73
+ **provenance** (`object`): Where the experiment came from (source, sourceId, sourceVersion, metadata).
74
+
75
+ **grouping** (`object`): Grouping fields (experimentSetId, comparisonId, variantId, trialIndex) for organizing related runs.
76
+
77
+ Returns `Promise<{ experimentId, status, totalItems, datasetVersion }>`.
78
+
79
+ ## runExperimentItem()
80
+
81
+ Executes one experiment item server-side: Mastra runs the experiment's target against the item, runs the resolved scorers, and upserts the result keyed by `(experimentId, itemId, attempt)`. Requires an experiment created with a target. Scorers resolve with the same precedence as native runs: experiment `scorerIds` win over item `scorerIds`, which win over dataset `scorerIds`.
82
+
83
+ **datasetId** (`string`): ID of the dataset the experiment runs against.
84
+
85
+ **experimentId** (`string`): ID of an experiment created with a target.
86
+
87
+ **itemId** (`string`): Dataset item to execute. Must exist at the pinned dataset version.
88
+
89
+ **attempt** (`number`): Zero-based repetition index for repeated trials. Defaults to 0. Retries of the same attempt converge on one row.
90
+
91
+ **requestContext** (`Record<string, unknown>`): Request context merged with the item's own request context (item values win).
92
+
93
+ Returns `Promise<{ result, scores }>` with the persisted result row and the scores produced for the item.
94
+
95
+ Calling it on a target-less experiment returns a `400`. Calling it after finalization returns a `409`.
96
+
97
+ ## submitExperimentResult()
98
+
99
+ Submits (or re-submits) one externally computed item result for a target-less experiment. Submitting the same `(experimentId, itemId, attempt)` key again updates the existing row instead of creating a duplicate. Use a different `attempt` value to record deliberate repeated trials as separate rows.
100
+
101
+ **datasetId** (`string`): ID of the dataset the experiment runs against.
102
+
103
+ **experimentId** (`string`): ID of the target-less experiment.
104
+
105
+ **itemId** (`string`): Dataset item this result belongs to. Must exist at the pinned dataset version.
106
+
107
+ **attempt** (`number`): Zero-based repetition index for repeated trials. Defaults to 0. Retries of the same attempt converge on one row.
108
+
109
+ **input** (`unknown`): Input replayed by the external runner. Defaults to the dataset item input.
110
+
111
+ **output** (`unknown`): Output produced by the external runner.
112
+
113
+ **groundTruth** (`unknown`): Ground truth. Defaults to the dataset item groundTruth.
114
+
115
+ **error** (`{ message: string; stack?: string; code?: string } | null`): Failure info when the item run failed. Counted as failed at finalization.
116
+
117
+ **startedAt** (`Date`): When the item run started.
118
+
119
+ **completedAt** (`Date`): When the item run completed.
120
+
121
+ **traceId** (`string`): Trace ID linking the result to observability data.
122
+
123
+ **scores** (`Array<{ scorerId: string; scorerName?: string; score: number; reason?: string; metadata?: Record<string, unknown> }>`): Externally computed scores. Persisted keyed by runId = experimentId, so they appear in comparisons like scorer-produced scores.
124
+
125
+ Returns `Promise<DatasetExperimentResult>`, the persisted result row.
126
+
127
+ Submitting to an experiment that has a target returns a `400`. Submitting after finalization returns a `409`.
128
+
129
+ ## finalizeExperiment()
130
+
131
+ Marks a caller-driven experiment completed. The server computes per-item counts from the persisted result rows, so the caller keeps no bookkeeping: `succeededCount` (at least one attempt without an error), `failedCount` (every attempt errored), and `skippedCount` (never submitted). Idempotent: finalizing an already-completed experiment returns the stored record.
132
+
133
+ **datasetId** (`string`): ID of the dataset the experiment runs against.
134
+
135
+ **experimentId** (`string`): ID of the experiment to finalize.
136
+
137
+ Returns `Promise<DatasetExperiment>`, the updated experiment record.
138
+
139
+ ## Related
140
+
141
+ - [Running experiments](https://mastra.ai/docs/datasets/running-experiments)
142
+ - [dataset.createExperiment()](https://mastra.ai/reference/datasets/createExperiment)
143
+ - [dataset.runExperimentItem()](https://mastra.ai/reference/datasets/runExperimentItem)
144
+ - [dataset.submitExperimentResult()](https://mastra.ai/reference/datasets/submitExperimentResult)
145
+ - [dataset.finalizeExperiment()](https://mastra.ai/reference/datasets/finalizeExperiment)
146
+ - [Server routes](https://mastra.ai/reference/server/routes)
@@ -731,6 +731,43 @@ export const mastra = new Mastra({
731
731
  })
732
732
  ```
733
733
 
734
+ ### server.handleShutdownSignals
735
+
736
+ **Type:** `boolean`\
737
+ **Default:** `true`
738
+
739
+ Whether the server generated by `mastra dev` and `mastra build` installs its own `SIGINT`/`SIGTERM` handlers. By default, receiving a signal begins graceful shutdown. The generated server stops accepting new connections and drains in-flight requests before it runs `mastra.shutdown()` and exits the process (see [`server.drainTimeout`](#serverdraintimeout)).
740
+
741
+ Set `handleShutdownSignals: false` to manage signals yourself, for example with a handler registered in your Mastra config module that calls `mastra.shutdown()`:
742
+
743
+ ```typescript
744
+ import { Mastra } from '@mastra/core'
745
+
746
+ export const mastra = new Mastra({
747
+ server: {
748
+ handleShutdownSignals: false,
749
+ },
750
+ })
751
+
752
+ let shuttingDown = false
753
+ const shutdown = async () => {
754
+ if (shuttingDown) return
755
+ shuttingDown = true
756
+ // custom shutdown coordination
757
+ await mastra.shutdown()
758
+ process.exit(0)
759
+ }
760
+
761
+ process.once('SIGINT', shutdown)
762
+ process.once('SIGTERM', shutdown)
763
+ ```
764
+
765
+ Be aware of these caveats:
766
+
767
+ - Without a handler of your own, Node.js default signal behavior applies: the process terminates immediately with no drain.
768
+ - Your handler has no access to the HTTP server handle in the generated entry file, so it can't drain in-flight HTTP requests. If you only need a longer drain window, prefer [`server.drainTimeout`](#serverdraintimeout). For full control over the server lifecycle, use a [server adapter](https://mastra.ai/docs/server/server-adapters).
769
+ - During `mastra dev`, the built-in handlers release storage resources between hot reloads (for example, DuckDB file locks). Disabling them without an equivalent handler can cause lock conflicts after a reload.
770
+
734
771
  ### server.host
735
772
 
736
773
  **Type:** `string`\
@@ -879,6 +916,27 @@ export const mastra = new Mastra({
879
916
  })
880
917
  ```
881
918
 
919
+ ### server.drainTimeout
920
+
921
+ **Type:** `number`\
922
+ **Default:** `5000` (5 seconds)
923
+
924
+ Maximum time in milliseconds to drain in-flight requests after the server receives `SIGINT` or `SIGTERM`. The value must be finite and between `0` and `2147483647`. During this window the server stops accepting new connections and waits for in-flight requests (including active streams) to finish; if the deadline passes first, remaining HTTP connections are closed. Upgraded connections such as WebSockets end when the process exits. Either way, `mastra.shutdown()` then runs (bounded by its own 5 second limit) before the process exits. Setting `0` skips the drain entirely.
925
+
926
+ Increase this when rolling deploys should let long-running agent turns finish. Configure it at least 5 seconds plus a safety margin below your platform's grace period so `mastra.shutdown()` can still complete. For example, account for `terminationGracePeriodSeconds` on Kubernetes:
927
+
928
+ ```typescript
929
+ import { Mastra } from '@mastra/core'
930
+
931
+ export const mastra = new Mastra({
932
+ server: {
933
+ drainTimeout: 600000, // 10 minutes
934
+ },
935
+ })
936
+ ```
937
+
938
+ Sending either shutdown signal again during the drain bypasses it and terminates the process immediately.
939
+
882
940
  ### server.studioBase
883
941
 
884
942
  **Type:** `string`\
@@ -0,0 +1,76 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # dataset.createExperiment()
4
+
5
+ **Added in:** `@mastra/core@1.61.0`
6
+
7
+ Creates an experiment without starting a run. Your own orchestrator (for example a Temporal workflow) drives the loop in one of two shapes:
8
+
9
+ - **With a target:** pass `targetType` and `targetId`, then call [`runExperimentItem()`](https://mastra.ai/reference/datasets/runExperimentItem) per item. Mastra executes the target and runs the scorers server-side.
10
+ - **Without a target:** omit both and run everything on your own infrastructure, ingesting each result with [`submitExperimentResult()`](https://mastra.ai/reference/datasets/submitExperimentResult).
11
+
12
+ Both shapes finish with [`finalizeExperiment()`](https://mastra.ai/reference/datasets/finalizeExperiment).
13
+
14
+ The experiment pins the dataset version at creation time, so submissions are validated against a stable set of items even if the dataset changes afterwards.
15
+
16
+ ## Usage example
17
+
18
+ ```typescript
19
+ import { Mastra } from '@mastra/core'
20
+
21
+ const mastra = new Mastra({/* storage config */})
22
+
23
+ const dataset = await mastra.datasets.get({ id: 'dataset-id' })
24
+
25
+ const { experimentId, totalItems, datasetVersion } = await dataset.createExperiment({
26
+ id: 'temporal-wf-run-42', // optional: idempotent create on retry
27
+ targetType: 'agent',
28
+ targetId: 'translation-agent',
29
+ scorers: ['accuracy'],
30
+ })
31
+ ```
32
+
33
+ Passing your own `id` makes creation idempotent: calling it again with the same `id` returns the existing experiment instead of failing, so a retried workflow activity is safe. If the `id` belongs to an experiment on another dataset or an experiment with a different target, the call throws an `EXPERIMENT_ID_CONFLICT` error.
34
+
35
+ `targetType` and `targetId` must be provided together, and the target must exist in the Mastra registry at create time. `scorers` requires a target because Mastra never scores target-less experiments; submit flat scores through `submitExperimentResult` instead.
36
+
37
+ ## Parameters
38
+
39
+ **id** (`string`): Caller-supplied experiment ID (for example a workflow run ID). Makes creation idempotent on retry.
40
+
41
+ **targetType** (`'agent' | 'workflow' | 'scorer'`): Type of target that runExperimentItem() executes. Provide together with targetId, or omit both for pure ingestion.
42
+
43
+ **targetId** (`string`): ID of the registered target. Provide together with targetType.
44
+
45
+ **scorers** (`string[]`): Run-level scorer IDs resolved server-side by runExperimentItem(). Requires a target.
46
+
47
+ **name** (`string`): Human-readable experiment name.
48
+
49
+ **description** (`string`): Experiment description.
50
+
51
+ **metadata** (`Record<string, unknown>`): Arbitrary metadata stored on the experiment.
52
+
53
+ **provenance** (`ExperimentProvenance`): Where the experiment came from (source system, ID, and version).
54
+
55
+ **grouping** (`ExperimentGrouping`): Grouping fields (experimentSetId, comparisonId, variantId, trialIndex) for organizing related runs.
56
+
57
+ **version** (`number`): Dataset version to pin. Defaults to the current dataset version.
58
+
59
+ ## Returns
60
+
61
+ **result** (`Promise<object>`): Immediate response with experiment ID.
62
+
63
+ **result.experimentId** (`string`): ID of the created (or existing) experiment.
64
+
65
+ **result.status** (`ExperimentStatus`): 'running' for a new experiment; the stored status when an existing experiment is returned.
66
+
67
+ **result.totalItems** (`number`): Number of dataset items visible at the pinned version.
68
+
69
+ **result.datasetVersion** (`number`): The pinned dataset version.
70
+
71
+ ## Related
72
+
73
+ - [dataset.runExperimentItem()](https://mastra.ai/reference/datasets/runExperimentItem)
74
+ - [dataset.submitExperimentResult()](https://mastra.ai/reference/datasets/submitExperimentResult)
75
+ - [dataset.finalizeExperiment()](https://mastra.ai/reference/datasets/finalizeExperiment)
76
+ - [Running experiments](https://mastra.ai/docs/datasets/running-experiments)
@@ -0,0 +1,43 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # dataset.finalizeExperiment()
4
+
5
+ **Added in:** `@mastra/core@1.61.0`
6
+
7
+ Finalizes a caller-driven experiment created with [`createExperiment()`](https://mastra.ai/reference/datasets/createExperiment), with or without a target. Mastra computes the final counts server-side from the persisted result rows and marks the experiment `'completed'`, so the caller never tracks completion bookkeeping.
8
+
9
+ Counts are per item, so `succeededCount + failedCount + skippedCount === totalItems` even when items have multiple attempts:
10
+
11
+ - `succeededCount`: items where at least one attempt completed without an error
12
+ - `failedCount`: items that were submitted but where every attempt errored
13
+ - `skippedCount`: dataset items that were never submitted (any attempt)
14
+
15
+ Attempt-level detail remains available through [`listExperimentResults()`](https://mastra.ai/reference/datasets/listExperimentResults).
16
+
17
+ Finalization is idempotent: calling it on an already-finalized experiment returns the stored record without recomputing.
18
+
19
+ ## Usage example
20
+
21
+ ```typescript
22
+ const experiment = await dataset.finalizeExperiment({ experimentId })
23
+
24
+ console.log(experiment.status) // 'completed'
25
+ console.log(experiment.succeededCount)
26
+ console.log(experiment.failedCount)
27
+ console.log(experiment.skippedCount)
28
+ ```
29
+
30
+ ## Parameters
31
+
32
+ **experimentId** (`string`): ID of the experiment to finalize.
33
+
34
+ ## Returns
35
+
36
+ Returns a `Promise<Experiment>`, the updated experiment record with final status, counts, and `completedAt` timestamp.
37
+
38
+ ## Related
39
+
40
+ - [dataset.createExperiment()](https://mastra.ai/reference/datasets/createExperiment)
41
+ - [dataset.runExperimentItem()](https://mastra.ai/reference/datasets/runExperimentItem)
42
+ - [dataset.submitExperimentResult()](https://mastra.ai/reference/datasets/submitExperimentResult)
43
+ - [Running experiments](https://mastra.ai/docs/datasets/running-experiments)
@@ -0,0 +1,55 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # dataset.runExperimentItem()
4
+
5
+ **Added in:** `@mastra/core@1.61.0`
6
+
7
+ Executes one experiment item server-side: runs the experiment's target against the item, runs the resolved scorers, and upserts the result row keyed by `(experimentId, itemId, attempt)`.
8
+
9
+ Built for caller-driven loops where a durable orchestrator (for example a Temporal workflow) fans out one call per item and owns retries and timeouts. A retried call converges on the same row. Requires an experiment created with a target via [`createExperiment()`](https://mastra.ai/reference/datasets/createExperiment).
10
+
11
+ Scorers resolve with the same precedence as native runs: experiment `scorers` win over item `scorerIds`, which win over dataset `scorerIds`.
12
+
13
+ ## Usage example
14
+
15
+ ```typescript
16
+ const { experimentId } = await dataset.createExperiment({
17
+ targetType: 'agent',
18
+ targetId: 'translation-agent',
19
+ scorers: ['accuracy'],
20
+ })
21
+
22
+ const { result, scores } = await dataset.runExperimentItem({
23
+ experimentId,
24
+ itemId: 'item-1',
25
+ })
26
+
27
+ console.log(result.output)
28
+ console.log(scores)
29
+ ```
30
+
31
+ Each call executes the item exactly once, with no internal retry loop. Calling it on a target-less experiment throws `EXPERIMENT_HAS_NO_TARGET`. Use [`submitExperimentResult()`](https://mastra.ai/reference/datasets/submitExperimentResult) for those. Calls after finalization throw `EXPERIMENT_ALREADY_FINALIZED`, and an `itemId` that isn't visible at the pinned dataset version throws `DATASET_ITEM_NOT_FOUND`.
32
+
33
+ ## Parameters
34
+
35
+ **experimentId** (`string`): ID of an experiment created with a target.
36
+
37
+ **itemId** (`string`): ID of the dataset item to execute. Must be visible at the pinned dataset version.
38
+
39
+ **attempt** (`number`): Attempt number for repeated trials. Defaults to 0. Same (experimentId, itemId, attempt) upserts the existing row.
40
+
41
+ **requestContext** (`Record<string, unknown>`): Request context merged with the item's own request context (item values win).
42
+
43
+ ## Returns
44
+
45
+ **result** (`Promise<object>`): The persisted result and its scores.
46
+
47
+ **result.result** (`ExperimentResult`): The persisted result row, including output, error, traceId, attempt, and timestamps.
48
+
49
+ **result.scores** (`ScorerResult[]`): Scores produced by the resolved scorers for this item.
50
+
51
+ ## Related
52
+
53
+ - [dataset.createExperiment()](https://mastra.ai/reference/datasets/createExperiment)
54
+ - [dataset.finalizeExperiment()](https://mastra.ai/reference/datasets/finalizeExperiment)
55
+ - [Running experiments](https://mastra.ai/docs/datasets/running-experiments)
@@ -0,0 +1,56 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # dataset.submitExperimentResult()
4
+
5
+ **Added in:** `@mastra/core@1.61.0`
6
+
7
+ Submits (or re-submits) one externally computed item result for a target-less experiment created with [`createExperiment()`](https://mastra.ai/reference/datasets/createExperiment).
8
+
9
+ Submissions are upserts keyed by `(experimentId, itemId, attempt)`: a retried worker that submits the same key converges on a single row instead of duplicating results. Use a different `attempt` value to record repeated trials as separate rows.
10
+
11
+ ## Usage example
12
+
13
+ ```typescript
14
+ const result = await dataset.submitExperimentResult({
15
+ experimentId,
16
+ itemId: 'item-1',
17
+ output: { translation: 'Hola' },
18
+ scores: [{ scorerId: 'accuracy', score: 0.92, reason: 'Faithful translation' }],
19
+ })
20
+ ```
21
+
22
+ The item must be visible at the experiment's pinned dataset version. Submissions to an experiment that has a target throw `EXPERIMENT_HAS_TARGET`, which prevents two writers from racing on the same rows. Submissions after finalization throw `EXPERIMENT_ALREADY_FINALIZED`.
23
+
24
+ ## Parameters
25
+
26
+ **experimentId** (`string`): ID of the target-less experiment.
27
+
28
+ **itemId** (`string`): ID of the dataset item this result belongs to.
29
+
30
+ **attempt** (`number`): Attempt number for repeated trials. Defaults to 0. Same (experimentId, itemId, attempt) upserts the existing row.
31
+
32
+ **input** (`unknown`): Input the worker ran. Defaults to the dataset item's input at the pinned version.
33
+
34
+ **output** (`unknown`): Output produced by the external worker.
35
+
36
+ **groundTruth** (`unknown`): Expected output. Defaults to the dataset item's ground truth at the pinned version.
37
+
38
+ **error** (`object`): Error details (message, optional stack and code) when the item failed.
39
+
40
+ **startedAt** (`Date`): When the worker started this item.
41
+
42
+ **completedAt** (`Date`): When the worker finished this item.
43
+
44
+ **traceId** (`string`): Trace ID linking the result to observability data.
45
+
46
+ **scores** (`array`): Externally computed scores (scorerId, score, optional scorerName, reason, metadata). Persisted to the scores store under the experiment, so they appear in comparisons alongside native scorer runs. Score persistence is best-effort and never fails the submission.
47
+
48
+ ## Returns
49
+
50
+ Returns a `Promise<ExperimentResult>`, the persisted result row, including its `id`, `attempt`, and timestamps.
51
+
52
+ ## Related
53
+
54
+ - [dataset.createExperiment()](https://mastra.ai/reference/datasets/createExperiment)
55
+ - [dataset.finalizeExperiment()](https://mastra.ai/reference/datasets/finalizeExperiment)
56
+ - [Running experiments](https://mastra.ai/docs/datasets/running-experiments)