@mastra/pg 1.22.0-alpha.0 → 1.22.0-alpha.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +110 -0
- package/dist/docs/SKILL.md +1 -1
- package/dist/docs/assets/SOURCE_MAP.json +1 -1
- package/dist/docs/references/docs-deployment-workers.md +2 -2
- package/dist/docs/references/docs-storage.md +2 -0
- package/dist/docs/references/integrations-databases-postgresql.md +26 -0
- package/dist/docs/references/reference-rag-vector-databases.md +4 -4
- package/dist/docs/references/reference-vectors-pg.md +2 -0
- package/dist/index.cjs +515 -133
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +515 -133
- package/dist/index.js.map +1 -1
- package/dist/storage/db/sanitize-json.d.ts +11 -0
- package/dist/storage/db/sanitize-json.d.ts.map +1 -0
- package/dist/storage/domains/background-tasks/index.d.ts +3 -1
- package/dist/storage/domains/background-tasks/index.d.ts.map +1 -1
- package/dist/storage/domains/experiments/index.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/ddl.d.ts +17 -0
- package/dist/storage/domains/observability/v-next/ddl.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/discovery.d.ts +14 -0
- package/dist/storage/domains/observability/v-next/discovery.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/helpers.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/index.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/signal-schema.d.ts +5 -1
- package/dist/storage/domains/observability/v-next/signal-schema.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/sql.d.ts +29 -4
- package/dist/storage/domains/observability/v-next/sql.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/traces.d.ts.map +1 -1
- package/dist/storage/domains/observability/v-next/tracing.d.ts +4 -3
- package/dist/storage/domains/observability/v-next/tracing.d.ts.map +1 -1
- package/dist/storage/domains/workflows/index.d.ts +2 -10
- package/dist/storage/domains/workflows/index.d.ts.map +1 -1
- package/dist/storage/factory-storage.d.ts.map +1 -1
- package/dist/vector/index.d.ts +60 -5
- package/dist/vector/index.d.ts.map +1 -1
- package/package.json +3 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,115 @@
|
|
|
1
1
|
# @mastra/pg
|
|
2
2
|
|
|
3
|
+
## 1.22.0-alpha.4
|
|
4
|
+
|
|
5
|
+
### Patch Changes
|
|
6
|
+
|
|
7
|
+
- Stop observability filter discovery from re-scanning all history on every refresh ([#22136](https://github.com/mastra-ai/mastra/pull/22136))
|
|
8
|
+
|
|
9
|
+
Discovery queries that build Studio's Traces, Logs, and Metrics filter suggestions scanned every span, metric, and log event each time the cache went stale, which grew unbounded with retained data. Refreshes are now bounded to the last 30 days by default (configurable via `observability.discovery.lookbackSeconds`, `0` restores the previous unbounded behaviour), so the planner can prune partitions instead of reading all of them. Only one process refreshes a given cache entry at a time, so running several server instances no longer multiplies the work, and Studio holds discovery results for five minutes instead of refetching on every page mount.
|
|
10
|
+
|
|
11
|
+
- Experiment results now include isolated metadata snapshots from the dataset items that ran. ([#22005](https://github.com/mastra-ai/mastra/pull/22005))
|
|
12
|
+
|
|
13
|
+
- Fix `PgFactoryStorage` reads of `json` columns holding a JSON value that isn't an object. ([#22258](https://github.com/mastra-ai/mastra/pull/22258))
|
|
14
|
+
|
|
15
|
+
node-pg parses JSONB through its own type parsers, so the value reaching row deserialization is already a JS value. Deserialization parsed it a second time when it was a string, which threw and failed the whole read. Objects and arrays came back as JS objects and never hit that branch, so the defect stayed hidden until a caller stored a string.
|
|
16
|
+
|
|
17
|
+
This surfaced through Factory credential encryption, which stores secrets as an opaque envelope string: saving or reading any credential on Postgres threw `Unexpected token ... is not valid JSON`, affecting model provider credentials, integrations, and custom providers. libsql was unaffected, since it stores `json` columns as text where parsing on read is correct.
|
|
18
|
+
|
|
19
|
+
- Enforce atomic conditional background task state updates so cancellation cannot be overwritten during dispatch. Background task storage is no longer exposed by Cloudflare KV or ClickHouse, which cannot provide the required compare-and-set semantics. ([#22228](https://github.com/mastra-ai/mastra/pull/22228))
|
|
20
|
+
|
|
21
|
+
- Updated dependencies [[`aa3a85d`](https://github.com/mastra-ai/mastra/commit/aa3a85daf094c683bb97efdf4b6a696d2e474af5), [`d29d06f`](https://github.com/mastra-ai/mastra/commit/d29d06fe00bbd35b4571150ea04c59d2ed783c71), [`e6516df`](https://github.com/mastra-ai/mastra/commit/e6516dfcdae4f4ac0e7971d84359a81385ee602f), [`0b2a3d1`](https://github.com/mastra-ai/mastra/commit/0b2a3d1783875c5b97b7b36ab3d03d7360e0dde7), [`6bb5d71`](https://github.com/mastra-ai/mastra/commit/6bb5d7193fe9166b219f0fccae17db7a5ae86e65), [`57de7d6`](https://github.com/mastra-ai/mastra/commit/57de7d644ba7146edb4e9e6111ec4fa98c3a59e9), [`e8e299c`](https://github.com/mastra-ai/mastra/commit/e8e299cc6abdfc39947e2fec25803493015d3882), [`edfc548`](https://github.com/mastra-ai/mastra/commit/edfc548886bc7bae17b681f8b6b41a47eb32bcd2), [`a8a4871`](https://github.com/mastra-ai/mastra/commit/a8a4871215f51da95c47129602157ce5372f634a), [`5165cdc`](https://github.com/mastra-ai/mastra/commit/5165cdcdcf50e144bb8113278535196cc9b07065), [`6bb5d71`](https://github.com/mastra-ai/mastra/commit/6bb5d7193fe9166b219f0fccae17db7a5ae86e65), [`9ee8120`](https://github.com/mastra-ai/mastra/commit/9ee8120ce17f76b9f617489e05a283353742690a), [`d975e92`](https://github.com/mastra-ai/mastra/commit/d975e924d4936f46c386bd3dee39c671720289f6), [`1cfa878`](https://github.com/mastra-ai/mastra/commit/1cfa8784d8da0dfaa0317e5048bc48b6084a5ea5), [`c118318`](https://github.com/mastra-ai/mastra/commit/c1183181c9804303db4b511c2e2648f8b714712b), [`fc07c64`](https://github.com/mastra-ai/mastra/commit/fc07c6465043e08e99193a6751a01c56ffc2e7a1), [`542dee2`](https://github.com/mastra-ai/mastra/commit/542dee254167f974ff8cbbbfc0ce10f9a2616a7b), [`a58483c`](https://github.com/mastra-ai/mastra/commit/a58483cff1a9d41fce7c931843f48cb0ac450f64), [`a58483c`](https://github.com/mastra-ai/mastra/commit/a58483cff1a9d41fce7c931843f48cb0ac450f64), [`895e9df`](https://github.com/mastra-ai/mastra/commit/895e9dfc17d6f34299eca64e317ded9e5f5e5ef8)]:
|
|
22
|
+
- @mastra/core@1.62.0-alpha.8
|
|
23
|
+
|
|
24
|
+
## 1.22.0-alpha.3
|
|
25
|
+
|
|
26
|
+
### Minor Changes
|
|
27
|
+
|
|
28
|
+
- Added namespace isolation to PgVector operations so applications can safely reuse vector indexes across tenants. Existing vectors remain available in the default namespace. ([#22149](https://github.com/mastra-ai/mastra/pull/22149))
|
|
29
|
+
|
|
30
|
+
```ts
|
|
31
|
+
await pgVector.upsert({
|
|
32
|
+
indexName: 'documents',
|
|
33
|
+
vectors,
|
|
34
|
+
ids,
|
|
35
|
+
namespace: 'tenant-123',
|
|
36
|
+
});
|
|
37
|
+
|
|
38
|
+
const results = await pgVector.query({
|
|
39
|
+
indexName: 'documents',
|
|
40
|
+
queryVector,
|
|
41
|
+
namespace: 'tenant-123',
|
|
42
|
+
});
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
### Patch Changes
|
|
46
|
+
|
|
47
|
+
- Fixed `PgVector` scanning every vector table on startup. Constructing a `PgVector` warms an index cache in the background, and that warmup asked for full index statistics, which include `SELECT COUNT(*)` per table. On a large index that is a full table scan per index, per process start, and the warmup never used the count it paid for. `query()`, `upsert()`, `updateVector()` and the "has this index changed?" check in `createIndex()` paid for the same count. ([#22180](https://github.com/mastra-ai/mastra/pull/22180))
|
|
48
|
+
|
|
49
|
+
These paths now read only the index metadata they use (dimension, metric, index type, vector type, index configuration), all of which comes from the Postgres catalog at a cost that does not grow with the size of the table.
|
|
50
|
+
|
|
51
|
+
`describeIndex()` is unchanged and still returns an exact `count`:
|
|
52
|
+
|
|
53
|
+
```ts
|
|
54
|
+
const stats = await pgVector.describeIndex({ indexName: 'embeddings' });
|
|
55
|
+
console.log(stats.count); // exact row count, as before
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
Concurrent callers on a cold cache also no longer duplicate the lookup: the first call is shared with everyone waiting on it, and a failed lookup is not cached.
|
|
59
|
+
|
|
60
|
+
Fixes [#21952](https://github.com/mastra-ai/mastra/issues/21952).
|
|
61
|
+
|
|
62
|
+
- Updated dependencies [[`c8e4cea`](https://github.com/mastra-ai/mastra/commit/c8e4ceac9a390d78c8327dff3cdb2861dd71957f), [`ed01e9a`](https://github.com/mastra-ai/mastra/commit/ed01e9a807514a904374bf687a7b8f18750f6f78), [`4e9a228`](https://github.com/mastra-ai/mastra/commit/4e9a2283d5fd6ed1b70a2751eb3dc2cbf82ada20), [`63041eb`](https://github.com/mastra-ai/mastra/commit/63041eb4c50b520a0a80e03d4cd6ea99f67715a0)]:
|
|
63
|
+
- @mastra/core@1.62.0-alpha.6
|
|
64
|
+
|
|
65
|
+
## 1.22.0-alpha.2
|
|
66
|
+
|
|
67
|
+
### Minor Changes
|
|
68
|
+
|
|
69
|
+
- **In-progress traces now appear in Studio with `PostgresStoreVNext`** ([#22137](https://github.com/mastra-ai/mastra/pull/22137))
|
|
70
|
+
|
|
71
|
+
`PostgresStoreVNext` previously persisted a span only after it finished, so a long agent run stayed invisible until it completed. It now uses the `event-sourced` tracing strategy: one row is written when a span starts and another when it ends, and reads collapse those rows back into a single span. Traces show up in Studio while the run is executing, and filtering by `running` status works.
|
|
72
|
+
|
|
73
|
+
This also fixes duplicate traces from durable runs. A run that suspends and resumes opens a second root span on the same trace, which used to render as two separate entries; the trace list now shows the current root only.
|
|
74
|
+
|
|
75
|
+
Writes stay append-only, so throughput is unchanged. The span table gains an `isPending` column, added automatically on `init()` — no manual migration needed. Closes #22054.
|
|
76
|
+
|
|
77
|
+
### Patch Changes
|
|
78
|
+
|
|
79
|
+
- Updated dependencies [[`79f04a7`](https://github.com/mastra-ai/mastra/commit/79f04a7f6c6829da541139f638f2f1d267916e08), [`fd4d5fe`](https://github.com/mastra-ai/mastra/commit/fd4d5fe4f943699b85db5e74404f190d5a6b8c2a), [`f591643`](https://github.com/mastra-ai/mastra/commit/f591643becdf0be9bddce6ba1748e64bc30d77f1), [`b1ad324`](https://github.com/mastra-ai/mastra/commit/b1ad324d657f3544b0701332aef7eb10e9a36258), [`61c566d`](https://github.com/mastra-ai/mastra/commit/61c566dd2f2cde2b23ed8f139924e530d4202214)]:
|
|
80
|
+
- @mastra/core@1.62.0-alpha.4
|
|
81
|
+
|
|
82
|
+
## 1.22.0-alpha.1
|
|
83
|
+
|
|
84
|
+
### Minor Changes
|
|
85
|
+
|
|
86
|
+
- Added native application collection counts so totals no longer load matching rows. ([#22021](https://github.com/mastra-ai/mastra/pull/22021))
|
|
87
|
+
|
|
88
|
+
**Before**
|
|
89
|
+
|
|
90
|
+
```ts
|
|
91
|
+
const total = (await storage.ops.findMany('jobs', { status: 'failed' })).length;
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
**After**
|
|
95
|
+
|
|
96
|
+
```ts
|
|
97
|
+
const total = await storage.ops.count?.('jobs', { status: 'failed' });
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
### Patch Changes
|
|
101
|
+
|
|
102
|
+
- Workflow snapshot upserts no longer overwrite a previously stored `resourceId` with NULL when a run is re-persisted without one (for example during resume). ([#22105](https://github.com/mastra-ai/mastra/pull/22105))
|
|
103
|
+
|
|
104
|
+
- Fix PostgreSQL observability writes failing on NUL characters and unpaired Unicode surrogates ([#21728](https://github.com/mastra-ai/mastra/pull/21728))
|
|
105
|
+
|
|
106
|
+
Span serialization truncated strings by UTF-16 code unit, so a cut inside an emoji left a lone surrogate that PostgreSQL rejected on the jsonb cast (`22P02`). NUL characters were rejected as well (`22P05`). Because observability events are inserted as a single multi-row statement, one malformed field discarded the entire batch.
|
|
107
|
+
|
|
108
|
+
Truncation now preserves complete surrogate pairs, and the v-next PostgreSQL observability encoder sanitizes NUL and unpaired surrogates before the jsonb cast, using the same sanitizer that workflow snapshots already rely on. Valid Unicode, including complete emoji, is preserved.
|
|
109
|
+
|
|
110
|
+
- Updated dependencies [[`2c85f42`](https://github.com/mastra-ai/mastra/commit/2c85f428e04ccd63ea31a7ec80b5b327afdad555), [`11bbeb9`](https://github.com/mastra-ai/mastra/commit/11bbeb9b108ef2264e05acefc6dafb9cbb342921), [`1a485f3`](https://github.com/mastra-ai/mastra/commit/1a485f3538f5ec64d58bd8b5e1e99de0c695c87b), [`0d37487`](https://github.com/mastra-ai/mastra/commit/0d37487d9f349388a3f1cef6a536cf9dcc4b6273), [`8661d7d`](https://github.com/mastra-ai/mastra/commit/8661d7d7179f0a024456aabdd8679bcecd09ac28), [`575e343`](https://github.com/mastra-ai/mastra/commit/575e343900451021d96110916497d334af7bc252), [`cacb839`](https://github.com/mastra-ai/mastra/commit/cacb8392d9e74189b56d857290b0615f98a2683d), [`b47b26e`](https://github.com/mastra-ai/mastra/commit/b47b26e6fe95cb8a3482be2c5e52de157fe59d0b), [`0d37487`](https://github.com/mastra-ai/mastra/commit/0d37487d9f349388a3f1cef6a536cf9dcc4b6273), [`c46eb09`](https://github.com/mastra-ai/mastra/commit/c46eb09ce4987509af57a0ac582c61241a6dd2f1), [`30ed33e`](https://github.com/mastra-ai/mastra/commit/30ed33ee14084a26019aba15fceadda6d6ddefaf), [`91ad69d`](https://github.com/mastra-ai/mastra/commit/91ad69d64994c89199b0c55399e64ed91c61df2f), [`8dc408d`](https://github.com/mastra-ai/mastra/commit/8dc408d34438f9e13297f792c11a5cfd6cf952e1), [`c92def1`](https://github.com/mastra-ai/mastra/commit/c92def10a13c822972c96f0a4ca6ffc1f4258aed), [`c5eaec5`](https://github.com/mastra-ai/mastra/commit/c5eaec5a860d80d0e3805e67db0414b87ac8cbed), [`e66b2ba`](https://github.com/mastra-ai/mastra/commit/e66b2ba100db63eaeab6e21e1ea34b113f2ec781)]:
|
|
111
|
+
- @mastra/core@1.62.0-alpha.3
|
|
112
|
+
|
|
3
113
|
## 1.22.0-alpha.0
|
|
4
114
|
|
|
5
115
|
### Minor Changes
|
package/dist/docs/SKILL.md
CHANGED
|
@@ -29,7 +29,7 @@ Subscribes to workflow events on the [PubSub](https://mastra.ai/docs/server/pubs
|
|
|
29
29
|
|
|
30
30
|
In a split deployment, the orchestration worker pulls events from a distributed PubSub backend and delegates step execution back to the API over HTTP. In-process, it runs steps directly.
|
|
31
31
|
|
|
32
|
-
The orchestration worker requires a PubSub backend that supports pull mode (e.g., [`RedisStreamsPubSub`](https://mastra.ai/reference/pubsub/redis-streams) or [`GoogleCloudPubSub`](https://mastra.ai/reference/pubsub/google-cloud-pubsub)).
|
|
32
|
+
The orchestration worker requires a PubSub backend that supports pull mode (e.g., [`RedisStreamsPubSub`](https://mastra.ai/reference/pubsub/redis-streams), [`ValkeyStreamsPubSub`](https://mastra.ai/reference/pubsub/valkey-streams), or [`GoogleCloudPubSub`](https://mastra.ai/reference/pubsub/google-cloud-pubsub)).
|
|
33
33
|
|
|
34
34
|
### Scheduler worker
|
|
35
35
|
|
|
@@ -104,7 +104,7 @@ Any [supported storage backend](https://mastra.ai/reference/workers/overview) wo
|
|
|
104
104
|
|
|
105
105
|
Run the same build artifact in multiple containers, each with a different [`MASTRA_WORKERS`](https://mastra.ai/reference/workers/overview) value to control which worker starts in each process.
|
|
106
106
|
|
|
107
|
-
Split deployments require a distributed PubSub backend ([`RedisStreamsPubSub`](https://mastra.ai/reference/pubsub/redis-streams) or [`GoogleCloudPubSub`](https://mastra.ai/reference/pubsub/google-cloud-pubsub)), a shared [storage backend](https://mastra.ai/reference/workers/overview), and network connectivity between the orchestration worker and the API.
|
|
107
|
+
Split deployments require a distributed PubSub backend ([`RedisStreamsPubSub`](https://mastra.ai/reference/pubsub/redis-streams), [`ValkeyStreamsPubSub`](https://mastra.ai/reference/pubsub/valkey-streams), or [`GoogleCloudPubSub`](https://mastra.ai/reference/pubsub/google-cloud-pubsub)), a shared [storage backend](https://mastra.ai/reference/workers/overview), and network connectivity between the orchestration worker and the API.
|
|
108
108
|
|
|
109
109
|
### Select workers
|
|
110
110
|
|
|
@@ -197,6 +197,7 @@ Each provider page includes installation instructions, configuration parameters,
|
|
|
197
197
|
- [Convex](https://mastra.ai/integrations/databases/convex)
|
|
198
198
|
- [DuckDB](https://mastra.ai/integrations/databases/duckdb)
|
|
199
199
|
- [DynamoDB](https://mastra.ai/integrations/databases/dynamodb)
|
|
200
|
+
- [Elasticsearch](https://mastra.ai/integrations/databases/elasticsearch)
|
|
200
201
|
- [Google Cloud Spanner](https://mastra.ai/integrations/databases/spanner)
|
|
201
202
|
- [LanceDB](https://mastra.ai/integrations/databases/lancedb)
|
|
202
203
|
- [libSQL](https://mastra.ai/integrations/databases/libsql)
|
|
@@ -207,6 +208,7 @@ Each provider page includes installation instructions, configuration parameters,
|
|
|
207
208
|
- [OracleDB](https://mastra.ai/integrations/databases/oracledb)
|
|
208
209
|
- [PostgreSQL](https://mastra.ai/integrations/databases/postgresql)
|
|
209
210
|
- [Redis](https://mastra.ai/integrations/databases/redis)
|
|
211
|
+
- [Valkey](https://mastra.ai/integrations/databases/valkey)
|
|
210
212
|
- [Upstash](https://mastra.ai/integrations/databases/upstash)
|
|
211
213
|
|
|
212
214
|
## Next steps
|
|
@@ -146,6 +146,32 @@ PostgreSQL supports observability and can handle low trace volumes. Throughput c
|
|
|
146
146
|
- Setting up table partitioning for efficient data retention
|
|
147
147
|
- Migrating observability to [ClickHouse via composite storage](https://mastra.ai/reference/storage/composite) if you need to scale further
|
|
148
148
|
|
|
149
|
+
`PostgresStoreVNext` uses the `event-sourced` [tracing strategy](https://mastra.ai/docs/observability/integrations/exporters/mastra-storage) instead. It writes one row when a span starts and another when it ends, never updating a row in place, and collapses those rows when a trace is read. Writes stay append-only, and traces appear in Studio while the run is still executing.
|
|
150
|
+
|
|
151
|
+
#### Filter suggestions
|
|
152
|
+
|
|
153
|
+
Studio's Traces, Logs, and Metrics pages offer filter values (tags, service names, environments, entity names, metric names and labels) discovered from your observability data. Those values are cached in the database and recomputed in the background when the cache goes stale.
|
|
154
|
+
|
|
155
|
+
To keep the refresh cheap on large datasets, it only scans events from the last 30 days. Values that appear exclusively in older events won't be suggested, though filtering by them still works. Use `discovery` to change the window or how often it refreshes:
|
|
156
|
+
|
|
157
|
+
```typescript
|
|
158
|
+
import { PostgresStoreVNext } from '@mastra/pg'
|
|
159
|
+
|
|
160
|
+
const storage = new PostgresStoreVNext({
|
|
161
|
+
id: 'pg-storage',
|
|
162
|
+
connectionString: process.env.DATABASE_URL,
|
|
163
|
+
observability: {
|
|
164
|
+
connectionString: process.env.OBSERVABILITY_DATABASE_URL,
|
|
165
|
+
discovery: {
|
|
166
|
+
lookbackSeconds: 7 * 24 * 60 * 60, // scan the last 7 days; 0 scans all history
|
|
167
|
+
ttlSeconds: 15 * 60, // refresh at most every 15 minutes
|
|
168
|
+
},
|
|
169
|
+
},
|
|
170
|
+
})
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
Lower `lookbackSeconds` if refreshes are slow, and raise `ttlSeconds` if they run more often than your filter values change. Only one process refreshes a given cache entry at a time, so adding server instances doesn't multiply the work.
|
|
174
|
+
|
|
149
175
|
### Initialization
|
|
150
176
|
|
|
151
177
|
When you pass storage to the Mastra class, `init()` is called automatically before any storage operation:
|
|
@@ -27,9 +27,9 @@ await store.upsert({
|
|
|
27
27
|
})
|
|
28
28
|
```
|
|
29
29
|
|
|
30
|
-
### Using MongoDB
|
|
30
|
+
### Using MongoDB Vector Search
|
|
31
31
|
|
|
32
|
-
For detailed setup instructions and best practices, see the [official MongoDB
|
|
32
|
+
MongoDB Vector Search is a good solution for teams who want to consolidate vector search, full-text search, and operational data in a single database to minimize infrastructure complexity and maintain production-grade performance. For detailed setup instructions and best practices, see the [official MongoDB Vector Search documentation](https://www.mongodb.com/docs/atlas/atlas-vector-search/vector-search-overview/?utm_campaign=devrel\&utm_source=third-party-content\&utm_medium=cta\&utm_content=mastra-docs).
|
|
33
33
|
|
|
34
34
|
### Using VoyageAI with MongoDB
|
|
35
35
|
|
|
@@ -37,7 +37,7 @@ MongoDB works seamlessly with VoyageAI's embedding models, which are optimized f
|
|
|
37
37
|
|
|
38
38
|
### Hybrid Search (Vector + Full-Text)
|
|
39
39
|
|
|
40
|
-
MongoDB supports hybrid search that fuses vector similarity with BM25 full-text search using server-side `$rankFusion` (requires MongoDB >= 8.0; generally available from 8.1, and enabled on Atlas 8.0.x). This is useful when you want to combine semantic and keyword-based retrieval:
|
|
40
|
+
MongoDB supports hybrid search that fuses vector similarity with BM25 full-text search using server-side `$rankFusion` (requires MongoDB >= 8.0; generally available from 8.1, and enabled on MongoDB Atlas 8.0.x). This is useful when you want to combine semantic and keyword-based retrieval:
|
|
41
41
|
|
|
42
42
|
```ts
|
|
43
43
|
await store.createSearchIndex({ indexName: 'myCollection', fields: ['text'] })
|
|
@@ -420,7 +420,7 @@ Each vector database enforces specific naming conventions for indexes and collec
|
|
|
420
420
|
|
|
421
421
|
**MongoDB**:
|
|
422
422
|
|
|
423
|
-
Collection
|
|
423
|
+
Collection and index names must:
|
|
424
424
|
|
|
425
425
|
- Start with a letter or underscore
|
|
426
426
|
- Be up to 120 bytes long
|
|
@@ -173,6 +173,8 @@ interface PGIndexStats {
|
|
|
173
173
|
}
|
|
174
174
|
```
|
|
175
175
|
|
|
176
|
+
`count` is an exact `SELECT COUNT(*)`, which scans the whole table, so avoid calling `describeIndex()` on a hot path for a large index. Reads and writes never pay for it: they only use the index metadata, which comes from the Postgres catalog.
|
|
177
|
+
|
|
176
178
|
### `deleteIndex()`
|
|
177
179
|
|
|
178
180
|
**indexName** (`string`): Name of the index to delete
|