@mastra/mcp-docs-server 1.2.18-alpha.3 → 1.2.18-alpha.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.docs/docs/mastra-platform/workspaces.md +6 -3
- package/.docs/docs/sandbox/filesystem.md +120 -139
- package/.docs/docs/sandbox/lsp.md +195 -143
- package/.docs/docs/sandbox/overview.md +103 -69
- package/.docs/docs/sandbox/search.md +172 -153
- package/.docs/docs/sandbox/skills.md +94 -151
- package/.docs/integrations/deploy/render.md +136 -89
- package/.docs/models/index.md +1 -1
- package/.docs/models/providers/edenai.md +2 -3
- package/.docs/models/providers/empiriolabs.md +1 -1
- package/.docs/models/providers/kilo.md +2 -2
- package/.docs/models/providers/llmgateway.md +2 -1
- package/.docs/models/providers/nano-gpt.md +3 -1
- package/.docs/models/providers/ofox.md +1 -1
- package/.docs/models/providers/opencode.md +65 -65
- package/.docs/reference/observability/tracing/exporters/langfuse.md +2 -0
- package/.docs/reference/rag/metadata-filters.md +16 -8
- package/.docs/reference/rag/retrieval.md +113 -5
- package/CHANGELOG.md +7 -0
- package/package.json +3 -3
|
@@ -2,26 +2,16 @@
|
|
|
2
2
|
|
|
3
3
|
# Search and indexing
|
|
4
4
|
|
|
5
|
-
Search
|
|
5
|
+
Search gives agents a fast way to find relevant content without reading every file. It works with [direct filesystem access](https://mastra.ai/docs/sandbox/filesystem) and [`mounts`](https://mastra.ai/docs/sandbox/filesystem), but stores searchable content in a separate index.
|
|
6
6
|
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
## When to use search
|
|
10
|
-
|
|
11
|
-
Use workspace search when your agent needs to:
|
|
12
|
-
|
|
13
|
-
- Find exact terms, filenames, or error messages with BM25 keyword search
|
|
14
|
-
- Find conceptually related content with vector search
|
|
15
|
-
- Combine keyword and semantic results with hybrid search
|
|
16
|
-
- Search a large set of files without reading each file in full
|
|
17
|
-
- Index content from files, databases, or APIs
|
|
7
|
+
Queries read that index, not the live filesystem. This separation also lets you index content from databases, APIs, or other application sources.
|
|
18
8
|
|
|
19
9
|
## Quickstart
|
|
20
10
|
|
|
21
|
-
Enable BM25 search, index
|
|
11
|
+
Enable BM25 keyword search, add a document to the index, and search it:
|
|
22
12
|
|
|
23
13
|
```typescript
|
|
24
|
-
import {
|
|
14
|
+
import { LocalFilesystem, Workspace } from '@mastra/core/workspace'
|
|
25
15
|
|
|
26
16
|
const workspace = new Workspace({
|
|
27
17
|
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
@@ -31,53 +21,34 @@ const workspace = new Workspace({
|
|
|
31
21
|
await workspace.index('/docs/guide.md', 'Reset passwords from the account settings page.')
|
|
32
22
|
|
|
33
23
|
const results = await workspace.search('password reset')
|
|
34
|
-
console.log(results)
|
|
35
|
-
```
|
|
36
|
-
|
|
37
|
-
This configuration also gives agents tools for searching and indexing workspace content.
|
|
38
|
-
|
|
39
|
-
## How it works
|
|
40
|
-
|
|
41
|
-
Workspace search has two phases: indexing and querying.
|
|
42
24
|
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
- For BM25: term frequencies and document statistics are computed
|
|
49
|
-
- For vector: the content is embedded using your embedder function and stored in the vector store
|
|
50
|
-
|
|
51
|
-
Each indexed document has:
|
|
52
|
-
|
|
53
|
-
- **id** - A unique identifier (typically the file path)
|
|
54
|
-
- **content** - The text content
|
|
55
|
-
- **metadata** - Optional key-value data stored with the document
|
|
25
|
+
for (const result of results) {
|
|
26
|
+
console.log(`${result.id}: ${result.score}`)
|
|
27
|
+
console.log(result.content)
|
|
28
|
+
}
|
|
29
|
+
```
|
|
56
30
|
|
|
57
|
-
|
|
31
|
+
The first argument to `workspace.index()` is the document ID returned in search results. It looks like a file path in this example, but `index()` doesn't read or create that file.
|
|
58
32
|
|
|
59
|
-
|
|
33
|
+
Configuring search also gives agents search and indexing tools. See [Agent tools](#agent-tools) to control them.
|
|
60
34
|
|
|
61
|
-
|
|
62
|
-
2. Documents are scored based on relevance to the query
|
|
63
|
-
3. Results are ranked by score and returned with the matching content
|
|
35
|
+
## Choose a search mode
|
|
64
36
|
|
|
65
|
-
|
|
37
|
+
Mastra supports three search modes:
|
|
66
38
|
|
|
67
|
-
|
|
39
|
+
| Mode | Best for | Example queries |
|
|
40
|
+
| -------- | ---------------------------------------- | ------------------------------------------- |
|
|
41
|
+
| `bm25` | Exact terms, technical queries, and code | "useState hook", "404 error", "config.yaml" |
|
|
42
|
+
| `vector` | Concepts and natural-language questions | "how to handle user authentication" |
|
|
43
|
+
| `hybrid` | A mix of exact and conceptual queries | Most agent search use cases |
|
|
68
44
|
|
|
69
|
-
|
|
45
|
+
If you don't pass a mode to `workspace.search()`, Mastra uses hybrid search when both BM25 and vector search are configured. Otherwise, it uses the configured mode.
|
|
70
46
|
|
|
71
|
-
|
|
72
|
-
import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
|
|
47
|
+
### BM25 keyword search
|
|
73
48
|
|
|
74
|
-
|
|
75
|
-
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
76
|
-
bm25: true,
|
|
77
|
-
})
|
|
78
|
-
```
|
|
49
|
+
BM25 scores documents by term frequency and document length. It needs no external service or embedding model. Pass `bm25: true` to use the defaults shown in the quickstart.
|
|
79
50
|
|
|
80
|
-
|
|
51
|
+
To tune term-frequency saturation and document-length normalization, pass `k1` and `b`:
|
|
81
52
|
|
|
82
53
|
```typescript
|
|
83
54
|
const workspace = new Workspace({
|
|
@@ -89,21 +60,21 @@ const workspace = new Workspace({
|
|
|
89
60
|
})
|
|
90
61
|
```
|
|
91
62
|
|
|
92
|
-
|
|
63
|
+
### Vector search
|
|
93
64
|
|
|
94
|
-
Vector search uses embeddings to find semantically similar content.
|
|
65
|
+
Vector search uses embeddings to find semantically similar content. Configure a vector store and a function that embeds one string at a time:
|
|
95
66
|
|
|
96
67
|
```typescript
|
|
68
|
+
import { openai } from '@ai-sdk/openai'
|
|
97
69
|
import { Workspace, LocalFilesystem } from '@mastra/core/workspace'
|
|
98
70
|
import { PineconeVector } from '@mastra/pinecone'
|
|
99
71
|
import { embed } from 'ai'
|
|
100
|
-
import { openai } from '@ai-sdk/openai'
|
|
101
72
|
|
|
102
73
|
const workspace = new Workspace({
|
|
103
74
|
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
104
75
|
vectorStore: new PineconeVector({
|
|
105
|
-
|
|
106
|
-
|
|
76
|
+
id: 'workspace-search',
|
|
77
|
+
apiKey: process.env.PINECONE_API_KEY!,
|
|
107
78
|
}),
|
|
108
79
|
embedder: async (text: string) => {
|
|
109
80
|
const { embedding } = await embed({
|
|
@@ -115,30 +86,42 @@ const workspace = new Workspace({
|
|
|
115
86
|
})
|
|
116
87
|
```
|
|
117
88
|
|
|
118
|
-
|
|
89
|
+
#### Batch embedding
|
|
90
|
+
|
|
91
|
+
A single-text embedder makes a separate provider call for each document. For large indexes, use a provider's batch API to embed several documents in one call.
|
|
92
|
+
|
|
93
|
+
A batch embedder must:
|
|
94
|
+
|
|
95
|
+
- Accept an array of strings.
|
|
96
|
+
- Return one embedding per string, in the same order.
|
|
97
|
+
- Have a `batch: true` property so Mastra can detect it at runtime.
|
|
98
|
+
- Optionally set `maxBatchSize` to the largest array the provider accepts.
|
|
119
99
|
|
|
120
|
-
|
|
100
|
+
Mastra splits larger indexing sets according to `maxBatchSize` and can process the resulting groups concurrently. Set this value to the provider's documented limit:
|
|
121
101
|
|
|
122
|
-
|
|
102
|
+
| Provider | Maximum inputs |
|
|
103
|
+
| -------- | -------------- |
|
|
104
|
+
| OpenAI | 2048 |
|
|
105
|
+
| Cohere | 96 |
|
|
106
|
+
| Voyage | 128 |
|
|
123
107
|
|
|
124
|
-
|
|
108
|
+
If you omit `maxBatchSize`, Mastra uses its internal batch size.
|
|
125
109
|
|
|
126
|
-
|
|
127
|
-
- `maxBatchSize`: the largest array the provider accepts in one call. Mastra splits larger requests into chunks of this size and sends them in parallel. Set this to your provider's documented limit. For example, OpenAI accepts 2048 and Cohere accepts 96. Voyage accepts 128. Omit it to send every pending text in one request.
|
|
110
|
+
Replace the single-text embedder with a batch embedder:
|
|
128
111
|
|
|
129
112
|
```typescript
|
|
130
|
-
import {
|
|
113
|
+
import { openai } from '@ai-sdk/openai'
|
|
114
|
+
import { LocalFilesystem, Workspace } from '@mastra/core/workspace'
|
|
131
115
|
import { PineconeVector } from '@mastra/pinecone'
|
|
132
116
|
import { embedMany } from 'ai'
|
|
133
|
-
import { openai } from '@ai-sdk/openai'
|
|
134
117
|
|
|
135
118
|
const model = openai.embedding('text-embedding-3-small')
|
|
136
119
|
|
|
137
120
|
const workspace = new Workspace({
|
|
138
121
|
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
139
122
|
vectorStore: new PineconeVector({
|
|
140
|
-
|
|
141
|
-
|
|
123
|
+
id: 'workspace-search',
|
|
124
|
+
apiKey: process.env.PINECONE_API_KEY!,
|
|
142
125
|
}),
|
|
143
126
|
embedder: Object.assign(
|
|
144
127
|
async (texts: string[]) => {
|
|
@@ -150,13 +133,11 @@ const workspace = new Workspace({
|
|
|
150
133
|
})
|
|
151
134
|
```
|
|
152
135
|
|
|
153
|
-
`Object.assign` adds
|
|
154
|
-
|
|
155
|
-
Single-text embedders still work. The function signature `(text: string) => Promise<number[]>` is unchanged, so existing code keeps running without modification.
|
|
136
|
+
`Object.assign()` adds `batch` and `maxBatchSize` as properties on the function. Mastra reads them as metadata and doesn't pass them to the provider. Single-text embedders with the `(text: string) => Promise<number[]>` signature remain supported.
|
|
156
137
|
|
|
157
|
-
|
|
138
|
+
### Hybrid search
|
|
158
139
|
|
|
159
|
-
Configure both BM25 and vector search to
|
|
140
|
+
Configure both BM25 and vector search to combine keyword and semantic results:
|
|
160
141
|
|
|
161
142
|
```typescript
|
|
162
143
|
const workspace = new Workspace({
|
|
@@ -167,31 +148,30 @@ const workspace = new Workspace({
|
|
|
167
148
|
})
|
|
168
149
|
```
|
|
169
150
|
|
|
170
|
-
|
|
151
|
+
### Custom index name
|
|
171
152
|
|
|
172
|
-
|
|
153
|
+
The vector index name defaults to a sanitized version of the workspace ID followed by `_search`. Set `searchIndexName` when you need a stable or shared name:
|
|
173
154
|
|
|
174
155
|
```typescript
|
|
175
156
|
const workspace = new Workspace({
|
|
176
157
|
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
177
|
-
|
|
158
|
+
vectorStore: pineconeVector,
|
|
159
|
+
embedder: embedderFn,
|
|
178
160
|
searchIndexName: 'my_workspace_vectors',
|
|
179
161
|
})
|
|
180
162
|
```
|
|
181
163
|
|
|
182
|
-
The
|
|
164
|
+
The name must start with a letter or `_`, contain only letters, numbers, or `_`, and contain at most 63 characters.
|
|
183
165
|
|
|
184
|
-
##
|
|
166
|
+
## Index content
|
|
185
167
|
|
|
186
|
-
|
|
168
|
+
Use `workspace.index()` to add supplied text to the search index. Mastra tokenizes the content for BM25 search and generates an embedding when vector search is configured.
|
|
187
169
|
|
|
188
|
-
|
|
170
|
+
The document ID doesn't need to exist in the filesystem. Add metadata when results need filtering or more context:
|
|
189
171
|
|
|
190
172
|
```typescript
|
|
191
|
-
// Basic indexing
|
|
192
173
|
await workspace.index('/docs/guide.md', 'Content of the guide...')
|
|
193
174
|
|
|
194
|
-
// Index with metadata for filtering or context
|
|
195
175
|
await workspace.index('/docs/api.md', apiDocContent, {
|
|
196
176
|
metadata: {
|
|
197
177
|
category: 'api',
|
|
@@ -200,114 +180,81 @@ await workspace.index('/docs/api.md', apiDocContent, {
|
|
|
200
180
|
})
|
|
201
181
|
```
|
|
202
182
|
|
|
203
|
-
Manual indexing
|
|
183
|
+
Manual indexing works well for database records and API responses, or when you need to preprocess, chunk, or annotate content yourself.
|
|
204
184
|
|
|
205
|
-
|
|
206
|
-
- You want to pre-process or chunk content before indexing
|
|
207
|
-
- You need to add custom metadata to documents
|
|
185
|
+
### Keep files and the index in sync
|
|
208
186
|
|
|
209
|
-
|
|
187
|
+
Filesystem and index mutations are independent. Writing, editing, or deleting a file doesn't update its indexed content, and `workspace.index()` doesn't read or change a file.
|
|
210
188
|
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
`autoIndexPaths` works with a static `filesystem` or with `mounts`. For mounts, include the mount prefix in each path, such as `/docs/**/*.md`. Resolver-backed filesystems aren't auto-indexed during workspace initialization because the provider is only selected for a request; index their content manually instead.
|
|
189
|
+
To write a file and make it searchable, perform both operations:
|
|
214
190
|
|
|
215
191
|
```typescript
|
|
216
|
-
const
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
autoIndexPaths: ['docs', 'support/faq'],
|
|
220
|
-
})
|
|
221
|
-
|
|
222
|
-
await workspace.init()
|
|
223
|
-
```
|
|
192
|
+
const path = 'notes/launch.md'
|
|
193
|
+
const content = '# Launch notes\n\nShip the new dashboard on Friday.'
|
|
194
|
+
const filesystem = workspace.filesystem
|
|
224
195
|
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
|
|
196
|
+
if (!filesystem) {
|
|
197
|
+
throw new Error('This operation requires a static filesystem')
|
|
198
|
+
}
|
|
228
199
|
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
232
|
-
bm25: true,
|
|
233
|
-
autoIndexPaths: ['docs/**/*.md', 'support/**/*.txt'],
|
|
234
|
-
})
|
|
200
|
+
await filesystem.writeFile(path, content)
|
|
201
|
+
await workspace.index(path, content)
|
|
235
202
|
```
|
|
236
203
|
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
Use `workspace.search()` to find relevant content. Results are ranked by relevance score.
|
|
240
|
-
|
|
241
|
-
```typescript
|
|
242
|
-
const results = await workspace.search('password reset')
|
|
243
|
-
|
|
244
|
-
for (const result of results) {
|
|
245
|
-
console.log(`${result.id}: ${result.score}`)
|
|
246
|
-
console.log(result.content)
|
|
247
|
-
}
|
|
248
|
-
```
|
|
204
|
+
After editing a file, index its complete updated content again. Deleting a file leaves its indexed document in place, and Mastra doesn't currently expose a public method for removing one indexed document. Use the `grep` file tool instead when an agent must search the current filesystem state without maintaining a separate index.
|
|
249
205
|
|
|
250
|
-
|
|
206
|
+
## Search content
|
|
251
207
|
|
|
252
|
-
|
|
208
|
+
Use `workspace.search()` to return documents ranked by relevance:
|
|
253
209
|
|
|
254
210
|
```typescript
|
|
255
211
|
const results = await workspace.search('authentication flow', {
|
|
256
212
|
topK: 10,
|
|
257
213
|
mode: 'hybrid',
|
|
258
|
-
minScore: 0.5,
|
|
259
214
|
vectorWeight: 0.5,
|
|
260
215
|
})
|
|
261
216
|
```
|
|
262
217
|
|
|
263
|
-
| Option | Description
|
|
264
|
-
| -------------- |
|
|
265
|
-
| `topK` | Maximum number of results to
|
|
266
|
-
| `mode` | Search mode: `'bm25'`, `'vector'`, or `'hybrid'`. Defaults to the best available mode
|
|
267
|
-
| `minScore` |
|
|
268
|
-
| `vectorWeight` | In hybrid mode,
|
|
218
|
+
| Option | Description |
|
|
219
|
+
| -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
220
|
+
| `topK` | Maximum number of results. Direct API calls default to 10. The agent search tool defaults to 5. |
|
|
221
|
+
| `mode` | Search mode: `'bm25'`, `'vector'`, or `'hybrid'`. Defaults to the best available mode for the workspace configuration. |
|
|
222
|
+
| `minScore` | Removes results below this score. The scale depends on the mode, vector provider, and indexed content. |
|
|
223
|
+
| `vectorWeight` | In hybrid mode, the weight given to vector scores. `0` gives vector scores zero weight, `1` gives BM25 scores zero weight, and `0.5` weights both equally. |
|
|
224
|
+
| `filter` | Vector-store metadata filter used by vector retrieval. In hybrid mode, it doesn't filter BM25-only results. The syntax depends on the configured vector store. |
|
|
269
225
|
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
Each result contains:
|
|
226
|
+
Each result contains the matching content and its score:
|
|
273
227
|
|
|
274
228
|
```typescript
|
|
275
229
|
interface SearchResult {
|
|
276
|
-
id: string
|
|
277
|
-
content: string
|
|
278
|
-
score: number
|
|
230
|
+
id: string
|
|
231
|
+
content: string
|
|
232
|
+
score: number
|
|
279
233
|
lineRange?: {
|
|
280
|
-
// Lines where the match was found
|
|
281
234
|
start: number
|
|
282
235
|
end: number
|
|
283
236
|
}
|
|
284
|
-
metadata?: Record<string, unknown>
|
|
237
|
+
metadata?: Record<string, unknown>
|
|
285
238
|
scoreDetails?: {
|
|
286
|
-
// Score breakdown (hybrid mode only)
|
|
287
239
|
vector?: number
|
|
288
240
|
bm25?: number
|
|
289
241
|
}
|
|
290
242
|
}
|
|
291
243
|
```
|
|
292
244
|
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
- Scores range from 0 to 1, where 1 is a perfect match
|
|
296
|
-
- BM25 scores are normalized based on the best match in the result set
|
|
297
|
-
- Vector scores represent cosine similarity between query and document embeddings
|
|
298
|
-
- In hybrid mode, scores are combined using the `vectorWeight` parameter
|
|
245
|
+
Score scales aren't interchangeable:
|
|
299
246
|
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
| `vector` | Conceptual queries, natural language | "how to handle user authentication", "best practices for error handling" |
|
|
306
|
-
| `hybrid` | General search, unknown query types | Most agent use cases |
|
|
247
|
+
- Standalone BM25 search returns raw BM25 scores, which aren't limited to 0 through 1.
|
|
248
|
+
- Vector score meaning depends on the configured vector store.
|
|
249
|
+
- Hybrid search min-max normalizes BM25 for its weighted calculation but uses the vector provider's score unchanged. The combined score isn't clamped to 0 through 1.
|
|
250
|
+
- `scoreDetails` keeps the raw component scores, including the unnormalized BM25 score.
|
|
251
|
+
- Tune `minScore` for the active mode, vector provider, and indexed content instead of reusing one threshold across indexes.
|
|
307
252
|
|
|
308
253
|
## Agent tools
|
|
309
254
|
|
|
310
|
-
When
|
|
255
|
+
When search is configured, agents receive `mastra_workspace_search` and `mastra_workspace_index`. Configure them independently with `WORKSPACE_TOOLS.SEARCH`.
|
|
256
|
+
|
|
257
|
+
Disable indexing when an agent should search existing content without changing the index:
|
|
311
258
|
|
|
312
259
|
```typescript
|
|
313
260
|
import { LocalFilesystem, Workspace, WORKSPACE_TOOLS } from '@mastra/core/workspace'
|
|
@@ -323,10 +270,82 @@ const workspace = new Workspace({
|
|
|
323
270
|
})
|
|
324
271
|
```
|
|
325
272
|
|
|
326
|
-
|
|
273
|
+
The index tool accepts the same supplied `path`, `content`, and optional metadata as `workspace.index()`. Because it can trigger embedding calls and vector-store writes, require approval when agents shouldn't mutate the index without review:
|
|
274
|
+
|
|
275
|
+
```typescript
|
|
276
|
+
const workspace = new Workspace({
|
|
277
|
+
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
278
|
+
bm25: true,
|
|
279
|
+
tools: {
|
|
280
|
+
[WORKSPACE_TOOLS.SEARCH.INDEX]: {
|
|
281
|
+
requireApproval: true,
|
|
282
|
+
},
|
|
283
|
+
},
|
|
284
|
+
})
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
`requireReadBeforeWrite` applies to filesystem write tools, not index mutations. Mastra also excludes the index tool when a static filesystem is read-only, even though indexing doesn't write to that filesystem.
|
|
288
|
+
|
|
289
|
+
See the [search tool reference](https://mastra.ai/reference/workspace/workspace-class) for the complete tool list and the [tool configuration reference](https://mastra.ai/reference/workspace/workspace-class) for shared settings.
|
|
290
|
+
|
|
291
|
+
## Auto-indexing
|
|
292
|
+
|
|
293
|
+
Use `autoIndexPaths` to build an index from static filesystem content during application startup.
|
|
294
|
+
|
|
295
|
+
> **Warning:** `autoIndexPaths` runs as part of `workspace.init()`. Passing the workspace to `new Mastra({ workspace })` registers it but doesn't initialize it or build the index. Call and await `workspace.init()` unless another runtime, such as `AgentController`, owns workspace initialization.
|
|
296
|
+
>
|
|
297
|
+
> ```typescript
|
|
298
|
+
> await workspace.init()
|
|
299
|
+
> ```
|
|
300
|
+
>
|
|
301
|
+
> When your application owns initialization, call it once during startup before serving requests.
|
|
302
|
+
|
|
303
|
+
In a normal Mastra application, construct `Mastra` first so its logger is available to the workspace. Then initialize the workspace at module startup:
|
|
304
|
+
|
|
305
|
+
```typescript
|
|
306
|
+
import { Mastra } from '@mastra/core'
|
|
307
|
+
import { LocalFilesystem, Workspace } from '@mastra/core/workspace'
|
|
308
|
+
|
|
309
|
+
const workspace = new Workspace({
|
|
310
|
+
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
311
|
+
bm25: true,
|
|
312
|
+
autoIndexPaths: ['docs', 'support/faq'],
|
|
313
|
+
})
|
|
314
|
+
|
|
315
|
+
export const mastra = new Mastra({ workspace })
|
|
316
|
+
|
|
317
|
+
await workspace.init()
|
|
318
|
+
```
|
|
319
|
+
|
|
320
|
+
Top-level `await` keeps module startup pending until initialization finishes. When `init()` resolves, Mastra has attempted to read and index every matching file.
|
|
321
|
+
|
|
322
|
+
Each `autoIndexPaths` entry can be a file, directory, or glob pattern. Recursive directory and glob traversal is limited to ten levels:
|
|
323
|
+
|
|
324
|
+
```typescript
|
|
325
|
+
const workspace = new Workspace({
|
|
326
|
+
filesystem: new LocalFilesystem({ basePath: './workspace' }),
|
|
327
|
+
bm25: true,
|
|
328
|
+
autoIndexPaths: ['docs/**/*.md', 'support/**/*.txt'],
|
|
329
|
+
})
|
|
330
|
+
|
|
331
|
+
export const mastra = new Mastra({ workspace })
|
|
332
|
+
|
|
333
|
+
await workspace.init()
|
|
334
|
+
```
|
|
335
|
+
|
|
336
|
+
Auto-indexing supports a static `filesystem` or `mounts`. Include mount prefixes in paths, such as `/docs/**/*.md`. It doesn't support a filesystem resolver because no provider has been selected when the workspace initializes. Index resolver-backed content manually instead.
|
|
337
|
+
|
|
338
|
+
Initialization performs one scan and doesn't watch for later changes. It splits large files into chunks before indexing them and generates embeddings in vector mode.
|
|
339
|
+
|
|
340
|
+
Keep the configured paths narrow and avoid dependency or output directories. Don't include binary file trees. Use a batch embedder for large vector indexes.
|
|
341
|
+
|
|
342
|
+
Calling `workspace.init()` again scans the paths again and clears the process-local BM25 index before rebuilding it. Don't call it from an agent call, tool, workflow step, or request handler.
|
|
343
|
+
|
|
344
|
+
A persistent vector index behaves differently. Auto-indexing replaces the files it reads but can retain documents for files removed before a later scan. Use a stable workspace `id` or `searchIndexName` when vectors must survive process restarts, and plan cleanup around the vector store's persistence behavior.
|
|
327
345
|
|
|
328
346
|
## Related
|
|
329
347
|
|
|
330
|
-
- [
|
|
331
|
-
- [
|
|
332
|
-
- [
|
|
348
|
+
- [Sandboxes](https://mastra.ai/docs/sandbox/overview)
|
|
349
|
+
- [Filesystem](https://mastra.ai/docs/sandbox/filesystem)
|
|
350
|
+
- [Retrieval-Augmented Generation overview](https://mastra.ai/reference/rag/overview)
|
|
351
|
+
- [Workspace configuration reference](https://mastra.ai/reference/workspace/workspace-class)
|