@mastra/mongodb 1.18.10-alpha.0 → 1.19.0-alpha.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/docs/SKILL.md +1 -1
- package/dist/docs/assets/SOURCE_MAP.json +1 -1
- package/dist/docs/references/reference-rag-vector-databases.md +23 -0
- package/dist/docs/references/reference-vectors-mongodb.md +90 -5
- package/dist/index.cjs +233 -64
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +233 -64
- package/dist/index.js.map +1 -1
- package/dist/vector/index.d.ts +89 -5
- package/dist/vector/index.d.ts.map +1 -1
- package/package.json +6 -5
package/dist/docs/SKILL.md
CHANGED
|
@@ -3,7 +3,7 @@ name: mastra-mongodb
|
|
|
3
3
|
description: Documentation for @mastra/mongodb. Use when working with @mastra/mongodb APIs, configuration, or implementation.
|
|
4
4
|
metadata:
|
|
5
5
|
package: "@mastra/mongodb"
|
|
6
|
-
version: "1.
|
|
6
|
+
version: "1.19.0-alpha.0"
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
## When to use
|
|
@@ -54,6 +54,29 @@ const results = await store.hybridQuery({
|
|
|
54
54
|
|
|
55
55
|
See the [MongoDB vector reference](https://mastra.ai/reference/vectors/mongodb) for details on `createSearchIndex()`, `textQuery()`, and `hybridQuery()`.
|
|
56
56
|
|
|
57
|
+
### Automated Embedding
|
|
58
|
+
|
|
59
|
+
MongoDB can generate the embeddings itself, so you don't need an embedding provider in your application. Create the index with `autoEmbed` and a Voyage AI model, write plain text, and search with a query string:
|
|
60
|
+
|
|
61
|
+
```ts
|
|
62
|
+
await store.createIndex({
|
|
63
|
+
indexName: 'myCollection',
|
|
64
|
+
autoEmbed: { model: 'voyage-4' },
|
|
65
|
+
})
|
|
66
|
+
await store.upsert({
|
|
67
|
+
indexName: 'myCollection',
|
|
68
|
+
documents: chunks.map(chunk => chunk.text),
|
|
69
|
+
metadata: chunks.map(chunk => ({ text: chunk.text })),
|
|
70
|
+
})
|
|
71
|
+
const results = await store.query({
|
|
72
|
+
indexName: 'myCollection',
|
|
73
|
+
queryText: 'search terms',
|
|
74
|
+
topK: 10,
|
|
75
|
+
})
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
Automated Embedding is a MongoDB Preview feature and needs a deployment where it's available. See the [MongoDB vector reference](https://mastra.ai/reference/vectors/mongodb) for the supported options and requirements.
|
|
79
|
+
|
|
57
80
|
**PgVector**:
|
|
58
81
|
|
|
59
82
|
```ts
|
|
@@ -87,7 +87,9 @@ Creates a new vector index (collection) in MongoDB.
|
|
|
87
87
|
|
|
88
88
|
**indexName** (`string`): Name of the collection to create
|
|
89
89
|
|
|
90
|
-
**dimension** (`number`): Vector dimension (must match your embedding model)
|
|
90
|
+
**dimension** (`number`): Vector dimension (must match your embedding model). Required unless autoEmbed is set, where the embedding model determines the dimension. Passing both is an error.
|
|
91
|
+
|
|
92
|
+
**autoEmbed** (`MongoDBAutoEmbedConfig`): Generate the embeddings in MongoDB instead of supplying vectors. See Automated Embedding for the supported fields.
|
|
91
93
|
|
|
92
94
|
**metric** (`'cosine' | 'euclidean' | 'dotproduct'`): Distance metric for similarity search (Default: `cosine`)
|
|
93
95
|
|
|
@@ -115,13 +117,13 @@ Adds or updates vectors and their metadata in the collection. On a bring-your-ow
|
|
|
115
117
|
|
|
116
118
|
**indexName** (`string`): Name of the collection to insert into
|
|
117
119
|
|
|
118
|
-
**vectors** (`number[][]`): Array of embedding vectors
|
|
120
|
+
**vectors** (`number[][]`): Array of embedding vectors. Required unless the index was created with autoEmbed, where MongoDB generates them from documents and supplying vectors is an error.
|
|
119
121
|
|
|
120
122
|
**metadata** (`Record<string, any>[]`): Metadata for each vector
|
|
121
123
|
|
|
122
124
|
**ids** (`string[]`): Optional vector IDs (auto-generated if not provided)
|
|
123
125
|
|
|
124
|
-
**documents** (`string[]`):
|
|
126
|
+
**documents** (`string[]`): Document text content to store alongside vectors. On an autoEmbed index this is the text MongoDB embeds, and it is required.
|
|
125
127
|
|
|
126
128
|
### `query()`
|
|
127
129
|
|
|
@@ -129,7 +131,11 @@ Searches for similar vectors with optional metadata filtering.
|
|
|
129
131
|
|
|
130
132
|
**indexName** (`string`): Name of the collection to search in
|
|
131
133
|
|
|
132
|
-
**queryVector** (`number[]`): Query vector to find similar vectors for
|
|
134
|
+
**queryVector** (`number[]`): Query vector to find similar vectors for. Supply this or queryText, not both.
|
|
135
|
+
|
|
136
|
+
**queryText** (`string`): Text for MongoDB to embed at query time. autoEmbed indexes only, and mutually exclusive with queryVector. For full-text matching use textQuery() instead.
|
|
137
|
+
|
|
138
|
+
**model** (`string`): Embedding model for this query, overriding the index's. Requires queryText, and must be compatible with the index's model.
|
|
133
139
|
|
|
134
140
|
**topK** (`number`): Number of results to return (Default: `10`)
|
|
135
141
|
|
|
@@ -226,7 +232,11 @@ Runs a hybrid search that fuses vector similarity with full-text results through
|
|
|
226
232
|
|
|
227
233
|
**indexName** (`string`): Name of the Mastra index to search
|
|
228
234
|
|
|
229
|
-
**queryVector** (`number[]`): Query vector for
|
|
235
|
+
**queryVector** (`number[]`): Query vector for the vector branch. Supply this or queryText, not both.
|
|
236
|
+
|
|
237
|
+
**queryText** (`string`): Text for MongoDB to embed for the vector branch. autoEmbed indexes only. Independent of query, so each branch can search for something different.
|
|
238
|
+
|
|
239
|
+
**model** (`string`): Embedding model for the vector branch, overriding the index's. Requires queryText.
|
|
230
240
|
|
|
231
241
|
**query** (`string`): Full-text search query string
|
|
232
242
|
|
|
@@ -273,6 +283,8 @@ interface IndexStats {
|
|
|
273
283
|
}
|
|
274
284
|
```
|
|
275
285
|
|
|
286
|
+
On an `autoEmbed` index, `dimension` is read from the index definition and reports MongoDB's default of `1024` when the index doesn't pin one. `count` counts documents that carry the embedded text field, since the generated vectors aren't stored on your documents.
|
|
287
|
+
|
|
276
288
|
### `deleteIndex()`
|
|
277
289
|
|
|
278
290
|
Deletes a vector index. Behavior depends on how the index was created:
|
|
@@ -367,6 +379,79 @@ try {
|
|
|
367
379
|
}
|
|
368
380
|
```
|
|
369
381
|
|
|
382
|
+
## Automated Embedding
|
|
383
|
+
|
|
384
|
+
MongoDB can generate the embeddings for you. Create the index with `autoEmbed` and a Voyage AI model, write plain text through `documents`, and search with `queryText`. No embedding provider runs in your application, and no vectors travel through it.
|
|
385
|
+
|
|
386
|
+
```typescript
|
|
387
|
+
import { MongoDBVector } from '@mastra/mongodb'
|
|
388
|
+
|
|
389
|
+
const store = new MongoDBVector({
|
|
390
|
+
id: 'mongodb-vector',
|
|
391
|
+
uri: process.env.MONGODB_URI,
|
|
392
|
+
dbName: process.env.MONGODB_DB_NAME,
|
|
393
|
+
})
|
|
394
|
+
|
|
395
|
+
// No `dimension`: the embedding model determines it.
|
|
396
|
+
await store.createIndex({
|
|
397
|
+
indexName: 'movies',
|
|
398
|
+
autoEmbed: { model: 'voyage-4' },
|
|
399
|
+
filterFields: ['year'],
|
|
400
|
+
})
|
|
401
|
+
|
|
402
|
+
// Automated Embedding indexes build slower than client-embedded ones.
|
|
403
|
+
await store.waitForIndexReady({ indexName: 'movies', timeoutMs: 300000 })
|
|
404
|
+
|
|
405
|
+
// No `vectors`: MongoDB embeds the text as documents are written.
|
|
406
|
+
await store.upsert({
|
|
407
|
+
indexName: 'movies',
|
|
408
|
+
documents: [
|
|
409
|
+
'A lonely astronaut adrift near a strange ocean planet.',
|
|
410
|
+
'A heist crew robs a bank vault in Paris.',
|
|
411
|
+
],
|
|
412
|
+
metadata: [{ year: 1972 }, { year: 2001 }],
|
|
413
|
+
})
|
|
414
|
+
|
|
415
|
+
// No `queryVector`: MongoDB embeds the query string with the same model.
|
|
416
|
+
const results = await store.query({
|
|
417
|
+
indexName: 'movies',
|
|
418
|
+
queryText: 'space opera about isolation',
|
|
419
|
+
topK: 5,
|
|
420
|
+
filter: { year: { $gt: 1970 } },
|
|
421
|
+
})
|
|
422
|
+
```
|
|
423
|
+
|
|
424
|
+
### `autoEmbed` options
|
|
425
|
+
|
|
426
|
+
**model** (`string`): Voyage AI model to embed with, for example voyage-4. MongoDB rejects a name it does not support and lists the ones it does.
|
|
427
|
+
|
|
428
|
+
**path** (`string`): Text field to embed. Defaults to the managed document field that upsert({ documents }) writes. Point it at a field of your own to index an existing collection in place. (Default: `document`)
|
|
429
|
+
|
|
430
|
+
**similarity** (`'cosine' | 'dotProduct' | 'euclidean'`): Vector similarity function. Defaults to MongoDB's own default when omitted.
|
|
431
|
+
|
|
432
|
+
**numDimensions** (`256 | 512 | 1024 | 2048`): Length of the generated embeddings. (Default: `1024`)
|
|
433
|
+
|
|
434
|
+
**quantization** (`'float' | 'scalar' | 'binary' | 'binaryNoRescore'`): Storage format for the generated vectors. (Default: `scalar`)
|
|
435
|
+
|
|
436
|
+
**indexingMethod** (`'hnsw' | 'flat'`): Index structure for the vector field. (Default: `hnsw`)
|
|
437
|
+
|
|
438
|
+
**hnswOptions** (`{ maxEdges?: number; numEdgeCandidates?: number }`): Tuning for the HNSW graph. MongoDB's defaults suit most workloads.
|
|
439
|
+
|
|
440
|
+
Any other field the `autoEmbed` index definition accepts is forwarded to MongoDB as given, so options added after this release work without a package update. Anything omitted keeps MongoDB's default.
|
|
441
|
+
|
|
442
|
+
**Important notes:**
|
|
443
|
+
|
|
444
|
+
- Automated Embedding is a **MongoDB Preview feature**. It requires an Atlas cluster with Automated Embedding available, or the `mongodb/mongodb-atlas-local:preview` image for local development. Self-managed deployments need `mongot` configured with a Voyage AI API key.
|
|
445
|
+
- The Voyage AI API key must be provisioned through the Atlas UI. A key issued directly by Voyage AI is rejected by the default embedding endpoint.
|
|
446
|
+
- **MongoDB stores the generated vectors outside your collection**, so documents carry text only. `includeVector: true` isn't supported on an `autoEmbed` index and throws.
|
|
447
|
+
- Embeddings are generated asynchronously. A query can legitimately return no results for a short time after a write, even once the index reports ready.
|
|
448
|
+
- An index declares either a vector field or an `autoEmbed` field, never both. Omit `autoEmbed` to keep supplying your own vectors. The client-side path is unchanged.
|
|
449
|
+
- Passing `vectors` to an `autoEmbed` index is an error, since they would be written to a field no index reads.
|
|
450
|
+
- When `document` is the embedded field it's no longer a declared filter field, so `documentFilter` uses the `$match` pre-filter automatically. Metadata filters declared through `filterFields` still push into `$vectorSearch`.
|
|
451
|
+
- `updateVector()` rejects a vector update on an `autoEmbed` index. Upsert the document with new text instead, and MongoDB re-embeds it.
|
|
452
|
+
- `queryVector` still works against an `autoEmbed` index, but the vector must come from a compatible model and match the index `quantization`. MongoDB rejects a mismatch.
|
|
453
|
+
- Embedding generation is billed per token, for both documents and queries.
|
|
454
|
+
|
|
370
455
|
## Indexing an existing collection
|
|
371
456
|
|
|
372
457
|
You can create a vector index on an existing operational collection instead of using a managed collection. This is useful when you want to add vector search capabilities to documents that already exist in your MongoDB database.
|