@mastra/mcp-docs-server 1.2.18-alpha.3 → 1.2.18-alpha.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,26 +2,16 @@
2
2
 
3
3
  # Search and indexing
4
4
 
5
- Search lets agents find relevant content in indexed workspace files. When an agent needs to answer a question or find information, it can search the indexed content instead of reading every file.
5
+ Search gives agents a fast way to find relevant content without reading every file. It works with [direct filesystem access](https://mastra.ai/docs/sandbox/filesystem) and [`mounts`](https://mastra.ai/docs/sandbox/filesystem), but stores searchable content in a separate index.
6
6
 
7
- Search works with both [`mounts`](https://mastra.ai/docs/sandbox/filesystem) and a [workspace-only filesystem](https://mastra.ai/docs/sandbox/filesystem). With `mounts`, paths from every mounted filesystem are available through one composite filesystem. With `filesystem`, search uses that provider directly. In both cases, queries search the workspace index rather than reading live files on demand.
8
-
9
- ## When to use search
10
-
11
- Use workspace search when your agent needs to:
12
-
13
- - Find exact terms, filenames, or error messages with BM25 keyword search
14
- - Find conceptually related content with vector search
15
- - Combine keyword and semantic results with hybrid search
16
- - Search a large set of files without reading each file in full
17
- - Index content from files, databases, or APIs
7
+ Queries read that index, not the live filesystem. This separation also lets you index content from databases, APIs, or other application sources.
18
8
 
19
9
  ## Quickstart
20
10
 
21
- Enable BM25 search, index content, and query it through the workspace:
11
+ Enable BM25 keyword search, add a document to the index, and search it:
22
12
 
23
13
  ```typescript
24
- import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
14
+ import { LocalFilesystem, Workspace } from '@mastra/core/workspace'
25
15
 
26
16
  const workspace = new Workspace({
27
17
  filesystem: new LocalFilesystem({ basePath: './workspace' }),
@@ -31,53 +21,34 @@ const workspace = new Workspace({
31
21
  await workspace.index('/docs/guide.md', 'Reset passwords from the account settings page.')
32
22
 
33
23
  const results = await workspace.search('password reset')
34
- console.log(results)
35
- ```
36
-
37
- This configuration also gives agents tools for searching and indexing workspace content.
38
-
39
- ## How it works
40
-
41
- Workspace search has two phases: indexing and querying.
42
24
 
43
- ### Indexing
44
-
45
- Content must be indexed before it can be searched. When you index a document:
46
-
47
- - The content is tokenized (split into searchable terms)
48
- - For BM25: term frequencies and document statistics are computed
49
- - For vector: the content is embedded using your embedder function and stored in the vector store
50
-
51
- Each indexed document has:
52
-
53
- - **id** - A unique identifier (typically the file path)
54
- - **content** - The text content
55
- - **metadata** - Optional key-value data stored with the document
25
+ for (const result of results) {
26
+ console.log(`${result.id}: ${result.score}`)
27
+ console.log(result.content)
28
+ }
29
+ ```
56
30
 
57
- ### Querying
31
+ The first argument to `workspace.index()` is the document ID returned in search results. It looks like a file path in this example, but `index()` doesn't read or create that file.
58
32
 
59
- When you search:
33
+ Configuring search also gives agents search and indexing tools. See [Agent tools](#agent-tools) to control them.
60
34
 
61
- 1. The query is processed using the same tokenization/embedding as indexing
62
- 2. Documents are scored based on relevance to the query
63
- 3. Results are ranked by score and returned with the matching content
35
+ ## Choose a search mode
64
36
 
65
- Workspaces support three search modes: BM25 keyword search and vector semantic search, plus hybrid search that combines both.
37
+ Mastra supports three search modes:
66
38
 
67
- ## BM25 keyword search
39
+ | Mode | Best for | Example queries |
40
+ | -------- | ---------------------------------------- | ------------------------------------------- |
41
+ | `bm25` | Exact terms, technical queries, and code | "useState hook", "404 error", "config.yaml" |
42
+ | `vector` | Concepts and natural-language questions | "how to handle user authentication" |
43
+ | `hybrid` | A mix of exact and conceptual queries | Most agent search use cases |
68
44
 
69
- BM25 scores documents based on term frequency and document length. It works well for exact matches and specific terminology.
45
+ If you don't pass a mode to `workspace.search()`, Mastra uses hybrid search when both BM25 and vector search are configured. Otherwise, it uses the configured mode.
70
46
 
71
- ```typescript
72
- import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
47
+ ### BM25 keyword search
73
48
 
74
- const workspace = new Workspace({
75
- filesystem: new LocalFilesystem({ basePath: './workspace' }),
76
- bm25: true,
77
- })
78
- ```
49
+ BM25 scores documents by term frequency and document length. It needs no external service or embedding model. Pass `bm25: true` to use the defaults shown in the quickstart.
79
50
 
80
- For custom BM25 parameters (`k1` is term frequency saturation, `b` is document length normalization):
51
+ To tune term-frequency saturation and document-length normalization, pass `k1` and `b`:
81
52
 
82
53
  ```typescript
83
54
  const workspace = new Workspace({
@@ -89,21 +60,21 @@ const workspace = new Workspace({
89
60
  })
90
61
  ```
91
62
 
92
- ## Vector search
63
+ ### Vector search
93
64
 
94
- Vector search uses embeddings to find semantically similar content. It requires a vector store and embedder function.
65
+ Vector search uses embeddings to find semantically similar content. Configure a vector store and a function that embeds one string at a time:
95
66
 
96
67
  ```typescript
68
+ import { openai } from '@ai-sdk/openai'
97
69
  import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
98
70
  import { PineconeVector } from '@mastra/pinecone'
99
71
  import { embed } from 'ai'
100
- import { openai } from '@ai-sdk/openai'
101
72
 
102
73
  const workspace = new Workspace({
103
74
  filesystem: new LocalFilesystem({ basePath: './workspace' }),
104
75
  vectorStore: new PineconeVector({
105
- apiKey: process.env.PINECONE_API_KEY,
106
- index: 'workspace-index',
76
+ id: 'workspace-search',
77
+ apiKey: process.env.PINECONE_API_KEY!,
107
78
  }),
108
79
  embedder: async (text: string) => {
109
80
  const { embedding } = await embed({
@@ -115,30 +86,42 @@ const workspace = new Workspace({
115
86
  })
116
87
  ```
117
88
 
118
- ### Batch embedding
89
+ #### Batch embedding
90
+
91
+ A single-text embedder makes a separate provider call for each document. For large indexes, use a provider's batch API to embed several documents in one call.
92
+
93
+ A batch embedder must:
94
+
95
+ - Accept an array of strings.
96
+ - Return one embedding per string, in the same order.
97
+ - Have a `batch: true` property so Mastra can detect it at runtime.
98
+ - Optionally set `maxBatchSize` to the largest array the provider accepts.
119
99
 
120
- The embedder above takes one text at a time. Indexing a workspace with hundreds of files calls the provider hundreds of times, which is slow and expensive.
100
+ Mastra splits larger indexing sets according to `maxBatchSize` and can process the resulting groups concurrently. Set this value to the provider's documented limit:
121
101
 
122
- When the provider supports batching (for example, OpenAI's `embedMany`), pass an embedder that takes an array of texts and accepts many embeddings back in one call. To opt in, set a `batch: true` property on the function. Mastra checks for that property at runtime and switches to the batched path.
102
+ | Provider | Maximum inputs |
103
+ | -------- | -------------- |
104
+ | OpenAI | 2048 |
105
+ | Cohere | 96 |
106
+ | Voyage | 128 |
123
107
 
124
- The following example replaces the single-text embedder with a batched one. The embedder function takes an array and returns an array of embeddings in the same order, plus carries two extra properties:
108
+ If you omit `maxBatchSize`, Mastra uses its internal batch size.
125
109
 
126
- - `batch: true`: marks the function as batch-capable. Without this property, Mastra calls it one text at a time.
127
- - `maxBatchSize`: the largest array the provider accepts in one call. Mastra splits larger requests into chunks of this size and sends them in parallel. Set this to your provider's documented limit. For example, OpenAI accepts 2048 and Cohere accepts 96. Voyage accepts 128. Omit it to send every pending text in one request.
110
+ Replace the single-text embedder with a batch embedder:
128
111
 
129
112
  ```typescript
130
- import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
113
+ import { openai } from '@ai-sdk/openai'
114
+ import { LocalFilesystem, Workspace } from '@mastra/core/workspace'
131
115
  import { PineconeVector } from '@mastra/pinecone'
132
116
  import { embedMany } from 'ai'
133
- import { openai } from '@ai-sdk/openai'
134
117
 
135
118
  const model = openai.embedding('text-embedding-3-small')
136
119
 
137
120
  const workspace = new Workspace({
138
121
  filesystem: new LocalFilesystem({ basePath: './workspace' }),
139
122
  vectorStore: new PineconeVector({
140
- apiKey: process.env.PINECONE_API_KEY,
141
- index: 'workspace-index',
123
+ id: 'workspace-search',
124
+ apiKey: process.env.PINECONE_API_KEY!,
142
125
  }),
143
126
  embedder: Object.assign(
144
127
  async (texts: string[]) => {
@@ -150,13 +133,11 @@ const workspace = new Workspace({
150
133
  })
151
134
  ```
152
135
 
153
- `Object.assign` adds the `batch` and `maxBatchSize` properties to the embedder function. Mastra reads them as metadata and never passes them to the provider.
154
-
155
- Single-text embedders still work. The function signature `(text: string) => Promise<number[]>` is unchanged, so existing code keeps running without modification.
136
+ `Object.assign()` adds `batch` and `maxBatchSize` as properties on the function. Mastra reads them as metadata and doesn't pass them to the provider. Single-text embedders with the `(text: string) => Promise<number[]>` signature remain supported.
156
137
 
157
- ## Hybrid search
138
+ ### Hybrid search
158
139
 
159
- Configure both BM25 and vector search to enable hybrid mode, which combines keyword matching with semantic understanding.
140
+ Configure both BM25 and vector search to combine keyword and semantic results:
160
141
 
161
142
  ```typescript
162
143
  const workspace = new Workspace({
@@ -167,31 +148,30 @@ const workspace = new Workspace({
167
148
  })
168
149
  ```
169
150
 
170
- ## Custom index name
151
+ ### Custom index name
171
152
 
172
- By default, the search index name is derived from the workspace ID. To set a custom name, use `searchIndexName`:
153
+ The vector index name defaults to a sanitized version of the workspace ID followed by `_search`. Set `searchIndexName` when you need a stable or shared name:
173
154
 
174
155
  ```typescript
175
156
  const workspace = new Workspace({
176
157
  filesystem: new LocalFilesystem({ basePath: './workspace' }),
177
- bm25: true,
158
+ vectorStore: pineconeVector,
159
+ embedder: embedderFn,
178
160
  searchIndexName: 'my_workspace_vectors',
179
161
  })
180
162
  ```
181
163
 
182
- The index name must be a valid SQL identifier: start with a letter or underscore, contain only letters, numbers, or shows, and be at most 63 characters long.
164
+ The name must start with a letter or `_`, contain only letters, numbers, or `_`, and contain at most 63 characters.
183
165
 
184
- ## Indexing content
166
+ ## Index content
185
167
 
186
- ### Manual indexing
168
+ Use `workspace.index()` to add supplied text to the search index. Mastra tokenizes the content for BM25 search and generates an embedding when vector search is configured.
187
169
 
188
- Use `workspace.index()` to add content to the search index programmatically. The file paths become document IDs. You can also pass metadata for each document.
170
+ The document ID doesn't need to exist in the filesystem. Add metadata when results need filtering or more context:
189
171
 
190
172
  ```typescript
191
- // Basic indexing
192
173
  await workspace.index('/docs/guide.md', 'Content of the guide...')
193
174
 
194
- // Index with metadata for filtering or context
195
175
  await workspace.index('/docs/api.md', apiDocContent, {
196
176
  metadata: {
197
177
  category: 'api',
@@ -200,114 +180,81 @@ await workspace.index('/docs/api.md', apiDocContent, {
200
180
  })
201
181
  ```
202
182
 
203
- Manual indexing is useful when:
183
+ Manual indexing works well for database records and API responses, or when you need to preprocess, chunk, or annotate content yourself.
204
184
 
205
- - You're indexing content that doesn't come from files (e.g., database records, API responses)
206
- - You want to pre-process or chunk content before indexing
207
- - You need to add custom metadata to documents
185
+ ### Keep files and the index in sync
208
186
 
209
- ### Auto-indexing
187
+ Filesystem and index mutations are independent. Writing, editing, or deleting a file doesn't update its indexed content, and `workspace.index()` doesn't read or change a file.
210
188
 
211
- Configure `autoIndexPaths` to automatically index files when the workspace initializes. Each entry can be a directory path (indexed recursively) or a glob pattern for selective indexing.
212
-
213
- `autoIndexPaths` works with a static `filesystem` or with `mounts`. For mounts, include the mount prefix in each path, such as `/docs/**/*.md`. Resolver-backed filesystems aren't auto-indexed during workspace initialization because the provider is only selected for a request; index their content manually instead.
189
+ To write a file and make it searchable, perform both operations:
214
190
 
215
191
  ```typescript
216
- const workspace = new Workspace({
217
- filesystem: new LocalFilesystem({ basePath: './workspace' }),
218
- bm25: true,
219
- autoIndexPaths: ['docs', 'support/faq'],
220
- })
221
-
222
- await workspace.init()
223
- ```
192
+ const path = 'notes/launch.md'
193
+ const content = '# Launch notes\n\nShip the new dashboard on Friday.'
194
+ const filesystem = workspace.filesystem
224
195
 
225
- When `init()` is called, all matching files are read and indexed for search. The file path becomes the document ID.
226
-
227
- Glob patterns let you index specific file types:
196
+ if (!filesystem) {
197
+ throw new Error('This operation requires a static filesystem')
198
+ }
228
199
 
229
- ```typescript
230
- const workspace = new Workspace({
231
- filesystem: new LocalFilesystem({ basePath: './workspace' }),
232
- bm25: true,
233
- autoIndexPaths: ['docs/**/*.md', 'support/**/*.txt'],
234
- })
200
+ await filesystem.writeFile(path, content)
201
+ await workspace.index(path, content)
235
202
  ```
236
203
 
237
- ## Searching
238
-
239
- Use `workspace.search()` to find relevant content. Results are ranked by relevance score.
240
-
241
- ```typescript
242
- const results = await workspace.search('password reset')
243
-
244
- for (const result of results) {
245
- console.log(`${result.id}: ${result.score}`)
246
- console.log(result.content)
247
- }
248
- ```
204
+ After editing a file, index its complete updated content again. Deleting a file leaves its indexed document in place, and Mastra doesn't currently expose a public method for removing one indexed document. Use the `grep` file tool instead when an agent must search the current filesystem state without maintaining a separate index.
249
205
 
250
- ### Search options
206
+ ## Search content
251
207
 
252
- You can customize the search behavior with options:
208
+ Use `workspace.search()` to return documents ranked by relevance:
253
209
 
254
210
  ```typescript
255
211
  const results = await workspace.search('authentication flow', {
256
212
  topK: 10,
257
213
  mode: 'hybrid',
258
- minScore: 0.5,
259
214
  vectorWeight: 0.5,
260
215
  })
261
216
  ```
262
217
 
263
- | Option | Description |
264
- | -------------- | ------------------------------------------------------------------------------------------------------------- |
265
- | `topK` | Maximum number of results to return. Default: 5 |
266
- | `mode` | Search mode: `'bm25'`, `'vector'`, or `'hybrid'`. Defaults to the best available mode based on configuration. |
267
- | `minScore` | Filter out results below this score threshold (0-1). |
268
- | `vectorWeight` | In hybrid mode, how much to weight vector scores vs BM25. 0 = all BM25, 1 = all vector, 0.5 = equal. |
218
+ | Option | Description |
219
+ | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
220
+ | `topK` | Maximum number of results. Direct API calls default to 10. The agent search tool defaults to 5. |
221
+ | `mode` | Search mode: `'bm25'`, `'vector'`, or `'hybrid'`. Defaults to the best available mode for the workspace configuration. |
222
+ | `minScore` | Removes results below this score. The scale depends on the mode, vector provider, and indexed content. |
223
+ | `vectorWeight` | In hybrid mode, the weight given to vector scores. `0` gives vector scores zero weight, `1` gives BM25 scores zero weight, and `0.5` weights both equally. |
224
+ | `filter` | Vector-store metadata filter used by vector retrieval. In hybrid mode, it doesn't filter BM25-only results. The syntax depends on the configured vector store. |
269
225
 
270
- ### Search results
271
-
272
- Each result contains:
226
+ Each result contains the matching content and its score:
273
227
 
274
228
  ```typescript
275
229
  interface SearchResult {
276
- id: string // Document ID (typically file path)
277
- content: string // The matching content
278
- score: number // Relevance score (0-1)
230
+ id: string
231
+ content: string
232
+ score: number
279
233
  lineRange?: {
280
- // Lines where the match was found
281
234
  start: number
282
235
  end: number
283
236
  }
284
- metadata?: Record<string, unknown> // Metadata stored with the document
237
+ metadata?: Record<string, unknown>
285
238
  scoreDetails?: {
286
- // Score breakdown (hybrid mode only)
287
239
  vector?: number
288
240
  bm25?: number
289
241
  }
290
242
  }
291
243
  ```
292
244
 
293
- **Understanding scores:**
294
-
295
- - Scores range from 0 to 1, where 1 is a perfect match
296
- - BM25 scores are normalized based on the best match in the result set
297
- - Vector scores represent cosine similarity between query and document embeddings
298
- - In hybrid mode, scores are combined using the `vectorWeight` parameter
245
+ Score scales aren't interchangeable:
299
246
 
300
- ### When to use each mode
301
-
302
- | Mode | Best for | Example queries |
303
- | -------- | ------------------------------------ | ------------------------------------------------------------------------ |
304
- | `bm25` | Exact terms, technical queries, code | "useState hook", "404 error", "config.yaml" |
305
- | `vector` | Conceptual queries, natural language | "how to handle user authentication", "best practices for error handling" |
306
- | `hybrid` | General search, unknown query types | Most agent use cases |
247
+ - Standalone BM25 search returns raw BM25 scores, which aren't limited to 0 through 1.
248
+ - Vector score meaning depends on the configured vector store.
249
+ - Hybrid search min-max normalizes BM25 for its weighted calculation but uses the vector provider's score unchanged. The combined score isn't clamped to 0 through 1.
250
+ - `scoreDetails` keeps the raw component scores, including the unnormalized BM25 score.
251
+ - Tune `minScore` for the active mode, vector provider, and indexed content instead of reusing one threshold across indexes.
307
252
 
308
253
  ## Agent tools
309
254
 
310
- When you configure search on a workspace, agents receive `mastra_workspace_search` and `mastra_workspace_index` tools. Use `WORKSPACE_TOOLS.SEARCH` constants to configure them independently. For example, remove indexing when the agent should search existing content without changing the index:
255
+ When search is configured, agents receive `mastra_workspace_search` and `mastra_workspace_index`. Configure them independently with `WORKSPACE_TOOLS.SEARCH`.
256
+
257
+ Disable indexing when an agent should search existing content without changing the index:
311
258
 
312
259
  ```typescript
313
260
  import { LocalFilesystem, Workspace, WORKSPACE_TOOLS } from '@mastra/core/workspace'
@@ -323,10 +270,82 @@ const workspace = new Workspace({
323
270
  })
324
271
  ```
325
272
 
326
- See [workspace class reference](https://mastra.ai/reference/workspace/workspace-class) for the complete tool list and [`WorkspaceToolsConfig`](https://mastra.ai/reference/workspace/workspace-class) for shared settings.
273
+ The index tool accepts the same supplied `path`, `content`, and optional metadata as `workspace.index()`. Because it can trigger embedding calls and vector-store writes, require approval when agents shouldn't mutate the index without review:
274
+
275
+ ```typescript
276
+ const workspace = new Workspace({
277
+ filesystem: new LocalFilesystem({ basePath: './workspace' }),
278
+ bm25: true,
279
+ tools: {
280
+ [WORKSPACE_TOOLS.SEARCH.INDEX]: {
281
+ requireApproval: true,
282
+ },
283
+ },
284
+ })
285
+ ```
286
+
287
+ `requireReadBeforeWrite` applies to filesystem write tools, not index mutations. Mastra also excludes the index tool when a static filesystem is read-only, even though indexing doesn't write to that filesystem.
288
+
289
+ See the [search tool reference](https://mastra.ai/reference/workspace/workspace-class) for the complete tool list and the [tool configuration reference](https://mastra.ai/reference/workspace/workspace-class) for shared settings.
290
+
291
+ ## Auto-indexing
292
+
293
+ Use `autoIndexPaths` to build an index from static filesystem content during application startup.
294
+
295
+ > **Warning:** `autoIndexPaths` runs as part of `workspace.init()`. Passing the workspace to `new Mastra({ workspace })` registers it but doesn't initialize it or build the index. Call and await `workspace.init()` unless another runtime, such as `AgentController`, owns workspace initialization.
296
+ >
297
+ > ```typescript
298
+ > await workspace.init()
299
+ > ```
300
+ >
301
+ > When your application owns initialization, call it once during startup before serving requests.
302
+
303
+ In a normal Mastra application, construct `Mastra` first so its logger is available to the workspace. Then initialize the workspace at module startup:
304
+
305
+ ```typescript
306
+ import { Mastra } from '@mastra/core'
307
+ import { LocalFilesystem, Workspace } from '@mastra/core/workspace'
308
+
309
+ const workspace = new Workspace({
310
+ filesystem: new LocalFilesystem({ basePath: './workspace' }),
311
+ bm25: true,
312
+ autoIndexPaths: ['docs', 'support/faq'],
313
+ })
314
+
315
+ export const mastra = new Mastra({ workspace })
316
+
317
+ await workspace.init()
318
+ ```
319
+
320
+ Top-level `await` keeps module startup pending until initialization finishes. When `init()` resolves, Mastra has attempted to read and index every matching file.
321
+
322
+ Each `autoIndexPaths` entry can be a file, directory, or glob pattern. Recursive directory and glob traversal is limited to ten levels:
323
+
324
+ ```typescript
325
+ const workspace = new Workspace({
326
+ filesystem: new LocalFilesystem({ basePath: './workspace' }),
327
+ bm25: true,
328
+ autoIndexPaths: ['docs/**/*.md', 'support/**/*.txt'],
329
+ })
330
+
331
+ export const mastra = new Mastra({ workspace })
332
+
333
+ await workspace.init()
334
+ ```
335
+
336
+ Auto-indexing supports a static `filesystem` or `mounts`. Include mount prefixes in paths, such as `/docs/**/*.md`. It doesn't support a filesystem resolver because no provider has been selected when the workspace initializes. Index resolver-backed content manually instead.
337
+
338
+ Initialization performs one scan and doesn't watch for later changes. It splits large files into chunks before indexing them and generates embeddings in vector mode.
339
+
340
+ Keep the configured paths narrow and avoid dependency or output directories. Don't include binary file trees. Use a batch embedder for large vector indexes.
341
+
342
+ Calling `workspace.init()` again scans the paths again and clears the process-local BM25 index before rebuilding it. Don't call it from an agent call, tool, workflow step, or request handler.
343
+
344
+ A persistent vector index behaves differently. Auto-indexing replaces the files it reads but can retain documents for files removed before a later scan. Use a stable workspace `id` or `searchIndexName` when vectors must survive process restarts, and plan cleanup around the vector store's persistence behavior.
327
345
 
328
346
  ## Related
329
347
 
330
- - [Sandbox](https://mastra.ai/docs/sandbox/overview)
331
- - [RAG overview](https://mastra.ai/reference/rag/overview)
332
- - [Workspace class reference](https://mastra.ai/reference/workspace/workspace-class)
348
+ - [Sandboxes](https://mastra.ai/docs/sandbox/overview)
349
+ - [Filesystem](https://mastra.ai/docs/sandbox/filesystem)
350
+ - [Retrieval-Augmented Generation overview](https://mastra.ai/reference/rag/overview)
351
+ - [Workspace configuration reference](https://mastra.ai/reference/workspace/workspace-class)