@nxgt/mongo-meilisearch 0.1.6 → 0.1.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +25 -9
- package/docs/README.md +13 -0
- package/docs/guide/boundaries.md +131 -0
- package/docs/guide/following-changes.md +195 -0
- package/docs/guide/reindex.md +132 -0
- package/docs/guide/sync-lifecycle.md +234 -0
- package/docs/roadmap.md +62 -0
- package/docs/troubleshooting.md +380 -0
- package/package.json +6 -5
package/README.md
CHANGED
|
@@ -225,10 +225,11 @@ out of range, an empty `name`, or a `transform` that is not a function.
|
|
|
225
225
|
|
|
226
226
|
## Not included
|
|
227
227
|
|
|
228
|
-
- **Several processes sharing one sync.** There is no lock
|
|
229
|
-
follower per sync name. Two do the same writes twice, and a
|
|
230
|
-
one while the other follows removes documents the follower has
|
|
231
|
-
indexed and will not send again.
|
|
228
|
+
- **Several processes sharing one sync — not yet.** There is no lock today,
|
|
229
|
+
so run **one** follower per sync name. Two do the same writes twice, and a
|
|
230
|
+
`reindex` in one while the other follows removes documents the follower has
|
|
231
|
+
already indexed and will not send again. A lease on a sync name is being
|
|
232
|
+
worked on: [the roadmap](docs/roadmap.md) says where it stands.
|
|
232
233
|
- **Partial updates.** A change sends the whole document the transform
|
|
233
234
|
gives, never a patch.
|
|
234
235
|
- **Keeping the index's settings.** That is `@nxgt/meilisearch`'s `sync`.
|
|
@@ -289,8 +290,9 @@ that is `{ code, sync, cause? }`; both types are exported.
|
|
|
289
290
|
Each is a `@ts-expect-error` case in this package's type tests.
|
|
290
291
|
|
|
291
292
|
- A transform that reads a field the collection's schema does not have.
|
|
292
|
-
- A transform that
|
|
293
|
-
id of the wrong type (an `ObjectId` for a string id).
|
|
293
|
+
- A transform that leaves out a field the index's document has, or gives an
|
|
294
|
+
id of the wrong type (an `ObjectId` for a string id). A field the document
|
|
295
|
+
does not have, beside all the ones it does, is passed on to Meilisearch.
|
|
294
296
|
- A transform that gives something other than a document or `null`.
|
|
295
297
|
- No `toIndexId` when the index's ids are not strings; one that gives
|
|
296
298
|
another type than the index's ids, or takes another than the collection's.
|
|
@@ -318,9 +320,10 @@ Each is a `@ts-expect-error` case in this package's type tests.
|
|
|
318
320
|
- **Without post-images, a change carries the document as it is now**, not
|
|
319
321
|
as the change left it (`@nxgt/mongo`'s change streams). For an index, where
|
|
320
322
|
only the latest state counts, that is what you want.
|
|
321
|
-
- **Only one process per sync name
|
|
322
|
-
Inside one process this package refuses it: `reindex()` and a
|
|
323
|
-
`start()` throw `RUNNING` while a sync of the same object is
|
|
323
|
+
- **Only one process per sync name**, while there is no lock; see *Not
|
|
324
|
+
included*. Inside one process this package refuses it: `reindex()` and a
|
|
325
|
+
second `start()` throw `RUNNING` while a sync of the same object is
|
|
326
|
+
following.
|
|
324
327
|
- **A dropped collection stops the sync** (`'invalidated'`) and leaves the
|
|
325
328
|
index as it was. What the collection had is removed by the reindex the
|
|
326
329
|
next `start` runs.
|
|
@@ -335,6 +338,19 @@ Each is a `@ts-expect-error` case in this package's type tests.
|
|
|
335
338
|
`documents.delete` and `tasks.get`, plus `indexes.create` unless the index
|
|
336
339
|
already exists — the first write to an index that does not creates it.
|
|
337
340
|
|
|
341
|
+
## Documentation
|
|
342
|
+
|
|
343
|
+
- [Guide index](docs/README.md) — every page, and when to read it.
|
|
344
|
+
- [The sync's lifecycle](docs/guide/sync-lifecycle.md) — the two definitions,
|
|
345
|
+
the transform, and every option with its default.
|
|
346
|
+
- [Reindexing](docs/guide/reindex.md) — the full fill, and what it removes.
|
|
347
|
+
- [Following changes](docs/guide/following-changes.md) — batches, the resume
|
|
348
|
+
point, and how a sync stops.
|
|
349
|
+
- [What it leaves out](docs/guide/boundaries.md) — the settings, the joins,
|
|
350
|
+
and the one follower per sync name.
|
|
351
|
+
- [Troubleshooting](docs/troubleshooting.md) — the errors, by their message.
|
|
352
|
+
- [Roadmap](docs/roadmap.md) — what is next, and what is not planned.
|
|
353
|
+
|
|
338
354
|
## License
|
|
339
355
|
|
|
340
356
|
MIT
|
package/docs/README.md
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# `@nxgt/mongo-meilisearch` documentation
|
|
2
|
+
|
|
3
|
+
The [README](../README.md) is the short version: what the package is, and one
|
|
4
|
+
example per area. These pages are the long one.
|
|
5
|
+
|
|
6
|
+
| Page | Read it when |
|
|
7
|
+
| --- | --- |
|
|
8
|
+
| [The sync's lifecycle](guide/sync-lifecycle.md) | you are wiring a collection to an index for the first time, or looking for an option and its default |
|
|
9
|
+
| [Reindexing](guide/reindex.md) | the index has to be filled, or refilled, from the collection |
|
|
10
|
+
| [Following changes](guide/following-changes.md) | a process has to keep the index in step, and stop cleanly — or it stopped and you need to know why |
|
|
11
|
+
| [What it leaves out](guide/boundaries.md) | you are about to run two followers, join another collection in, or wonder who applies the index settings |
|
|
12
|
+
| [Troubleshooting](troubleshooting.md) | something threw, and you have the message |
|
|
13
|
+
| [Roadmap](roadmap.md) | you want to know what is coming, and what will not |
|
|
@@ -0,0 +1,131 @@
|
|
|
1
|
+
# What it leaves out
|
|
2
|
+
|
|
3
|
+
The sync writes documents into one index from one collection. Everything
|
|
4
|
+
below is deliberately somebody else's job — either another package's, or the
|
|
5
|
+
caller's — so that the sync stays a function of the collection it was given.
|
|
6
|
+
|
|
7
|
+
```ts
|
|
8
|
+
import { bindIndex } from '@nxgt/meilisearch';
|
|
9
|
+
import { getCollection } from '@nxgt/mongo';
|
|
10
|
+
import { createSearchSync } from '@nxgt/mongo-meilisearch';
|
|
11
|
+
import { db, meili } from './clients'; // a driver Db, an SDK client
|
|
12
|
+
import { articles, articleIndex } from './search';
|
|
13
|
+
|
|
14
|
+
const articleSearch = createSearchSync({
|
|
15
|
+
collection: getCollection(db, articles),
|
|
16
|
+
index: bindIndex(meili, articleIndex),
|
|
17
|
+
transform: (article) => ({ id: String(article._id), title: article.title }),
|
|
18
|
+
});
|
|
19
|
+
|
|
20
|
+
// One follower, in one process: everything below is what that costs.
|
|
21
|
+
const running = await articleSearch.start();
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
## Several processes sharing one sync — not yet
|
|
25
|
+
|
|
26
|
+
There is **no lock today**: run one follower per sync name. Two processes
|
|
27
|
+
under one name do the same writes twice, and both are wrong the moment one of
|
|
28
|
+
them reindexes: a `reindex()` in one while the other follows removes
|
|
29
|
+
documents the follower has already indexed and will never send again.
|
|
30
|
+
|
|
31
|
+
Inside **one** process the package does refuse it: a second `start()`, or a
|
|
32
|
+
`reindex()`, throws `SearchSyncError` with the code `RUNNING`.
|
|
33
|
+
|
|
34
|
+
Give two syncs different `name`s when they are meant to be independent — the
|
|
35
|
+
default, `'<collection>:<index uid>'`, is shared by any two syncs over the
|
|
36
|
+
same pair.
|
|
37
|
+
|
|
38
|
+
A lease on a sync name, renewed while a follower runs, is being worked on;
|
|
39
|
+
[the roadmap](../roadmap.md) says where it stands. Until it ships, one writer
|
|
40
|
+
per name is the rule.
|
|
41
|
+
|
|
42
|
+
## Partial updates
|
|
43
|
+
|
|
44
|
+
A change sends the **whole** document the transform gives, never a patch:
|
|
45
|
+
|
|
46
|
+
```ts
|
|
47
|
+
transform: (article) => ({
|
|
48
|
+
id: String(article._id),
|
|
49
|
+
title: article.title,
|
|
50
|
+
body: article.body,
|
|
51
|
+
});
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
That is what makes a change safe to apply twice, which it may be — the resume
|
|
55
|
+
point moves after a batch is applied, so a restart re-sends what was in
|
|
56
|
+
flight.
|
|
57
|
+
|
|
58
|
+
## The index's settings
|
|
59
|
+
|
|
60
|
+
`searchableAttributes`, `filterableAttributes`, the ranking rules: they are
|
|
61
|
+
`@nxgt/meilisearch`'s, applied as a deployment step.
|
|
62
|
+
|
|
63
|
+
```ts
|
|
64
|
+
import { bindIndex, syncIndexes } from '@nxgt/meilisearch';
|
|
65
|
+
|
|
66
|
+
const index = bindIndex(meili, articleIndex);
|
|
67
|
+
await index.sync(); // this one index
|
|
68
|
+
await syncIndexes(meili, [articleIndex]); // or every index of the app at once
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
This package writes documents, and creates nothing but them — the first write
|
|
72
|
+
to an index that does not exist is what creates it, with no settings of its
|
|
73
|
+
own.
|
|
74
|
+
|
|
75
|
+
## Joining other collections in
|
|
76
|
+
|
|
77
|
+
The transform may read another collection, but a change to **that**
|
|
78
|
+
collection does not reach the index: the sync follows the one it was given.
|
|
79
|
+
|
|
80
|
+
```ts
|
|
81
|
+
export const articleIndex = defineIndex<{
|
|
82
|
+
id: string;
|
|
83
|
+
title: string;
|
|
84
|
+
author: string;
|
|
85
|
+
}>()({ uid: 'articles', primaryKey: 'id' });
|
|
86
|
+
|
|
87
|
+
import { authors } from './search'; // the authors collection's definition
|
|
88
|
+
|
|
89
|
+
createSearchSync({
|
|
90
|
+
collection: getCollection(db, articles),
|
|
91
|
+
index: bindIndex(meili, articleIndex),
|
|
92
|
+
// The author's name is read here, so it is right when the article is
|
|
93
|
+
// written — and stale from the moment the author is renamed: nothing is
|
|
94
|
+
// following the authors.
|
|
95
|
+
transform: async (article) => ({
|
|
96
|
+
id: String(article._id),
|
|
97
|
+
title: article.title,
|
|
98
|
+
author:
|
|
99
|
+
(await getCollection(db, authors).findById(article.authorId))?.name ?? '',
|
|
100
|
+
}),
|
|
101
|
+
});
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
Run a sync per collection that has to move the index, or denormalise the
|
|
105
|
+
field into the collection the sync follows and let its own writes carry it.
|
|
106
|
+
|
|
107
|
+
## One index, one collection
|
|
108
|
+
|
|
109
|
+
A reindex removes every document the collection does not give it, whoever
|
|
110
|
+
wrote it. Two collections pointed at one index, or an index another process
|
|
111
|
+
also writes to, lose documents at the first reindex. The sync owns its index.
|
|
112
|
+
|
|
113
|
+
## What the caller still owns
|
|
114
|
+
|
|
115
|
+
- **Restarting a sync that stopped.** `closed` rejects with a
|
|
116
|
+
`SearchSyncError`; the supervisor is what brings the process back. See
|
|
117
|
+
[Following changes](following-changes.md).
|
|
118
|
+
- **Choosing when a full reindex runs**, with `onHistoryLost: 'fail'`.
|
|
119
|
+
- **The two clients.** This package creates neither the `MongoClient` nor the
|
|
120
|
+
Meilisearch client, and closes neither.
|
|
121
|
+
- **The permissions.** The MongoDB user needs `find` and `changeStream` on
|
|
122
|
+
the collection and `find`, `insert`, `update` and `delete` on
|
|
123
|
+
`stateCollection`; the Meilisearch key needs `documents.add`,
|
|
124
|
+
`documents.get`, `documents.delete` and `tasks.get`, plus `indexes.create`
|
|
125
|
+
unless the index already exists.
|
|
126
|
+
|
|
127
|
+
## Next
|
|
128
|
+
|
|
129
|
+
- [The sync's lifecycle](sync-lifecycle.md) — the options behind each of
|
|
130
|
+
these.
|
|
131
|
+
- [Roadmap](../roadmap.md) — what is being worked on, and what is not planned.
|
|
@@ -0,0 +1,195 @@
|
|
|
1
|
+
# Following changes
|
|
2
|
+
|
|
3
|
+
`start()` follows the collection's change stream into the index, from where
|
|
4
|
+
the last run stopped, and gives back the running sync.
|
|
5
|
+
|
|
6
|
+
```ts
|
|
7
|
+
import { bindIndex } from '@nxgt/meilisearch';
|
|
8
|
+
import { getCollection, type ReadDocumentOf } from '@nxgt/mongo';
|
|
9
|
+
import { createSearchSync } from '@nxgt/mongo-meilisearch';
|
|
10
|
+
import { db, meili } from './clients'; // a driver Db, an SDK client
|
|
11
|
+
import { articles, articleIndex } from './search';
|
|
12
|
+
|
|
13
|
+
const collection = getCollection(db, articles);
|
|
14
|
+
const index = bindIndex(meili, articleIndex);
|
|
15
|
+
const transform = (article: ReadDocumentOf<typeof articles>) =>
|
|
16
|
+
article.draft ? null : { id: String(article._id), title: article.title };
|
|
17
|
+
|
|
18
|
+
const articleSearch = createSearchSync({ collection, index, transform });
|
|
19
|
+
|
|
20
|
+
const running = await articleSearch.start();
|
|
21
|
+
await running.ready; // already resolved: start waits for the stream to open
|
|
22
|
+
await running.flush(); // sends what is waiting now, and records how far it got
|
|
23
|
+
await running.close(); // flushes, then stops
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
`RunningSearchSync` is `AsyncDisposable`, so
|
|
27
|
+
`await using running = await articleSearch.start()` closes it at the end of
|
|
28
|
+
the block.
|
|
29
|
+
|
|
30
|
+
| Member | Type | Effect |
|
|
31
|
+
| --- | --- | --- |
|
|
32
|
+
| `ready` | `Promise<void>` | Resolves once changes are being heard. `start` has already awaited it |
|
|
33
|
+
| `closed` | `Promise<CloseReason>` | `'closed'` after `close()`, `'invalidated'` when the collection was dropped or renamed. **Rejects** with a `SearchSyncError` when an error stopped it |
|
|
34
|
+
| `flush()` | `Promise<void>` | Sends what is waiting, and records the point it covers |
|
|
35
|
+
| `close()` | `Promise<void>` | Flushes, then stops |
|
|
36
|
+
|
|
37
|
+
## Where it starts from
|
|
38
|
+
|
|
39
|
+
- **Nothing recorded** — the first run under this `name` — it
|
|
40
|
+
[reindexes](reindex.md), then follows from the position the reindex saved.
|
|
41
|
+
- **Something recorded**: it resumes there, so a change made while no process
|
|
42
|
+
was following is applied now.
|
|
43
|
+
- **A position the server's history no longer reaches**: see
|
|
44
|
+
[When the history is gone](#when-the-history-is-gone).
|
|
45
|
+
|
|
46
|
+
## What reaches the index
|
|
47
|
+
|
|
48
|
+
Every create, update, soft delete, restore and hard delete of the collection
|
|
49
|
+
runs through the transform:
|
|
50
|
+
|
|
51
|
+
```ts
|
|
52
|
+
await collection.create({ title: 'written' }); // added to the index
|
|
53
|
+
await collection.update(id, { draft: true }); // transform returns null → removed
|
|
54
|
+
await collection.update(id, { draft: false }); // back in
|
|
55
|
+
await collection.delete(id); // soft delete → removed
|
|
56
|
+
await collection.hardDelete(id); // removed
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Changes are sent in **batches**: one waits `flushIntervalMs` (1000 ms) for
|
|
60
|
+
others, or goes at once as soon as `batchSize` (500) are waiting. Several
|
|
61
|
+
changes to one document within a batch send only the last — the batch holds
|
|
62
|
+
one entry per id, so adds and deletes never race.
|
|
63
|
+
|
|
64
|
+
The resume point is recorded **after** the batch is applied. What was not
|
|
65
|
+
sent when a process stops is sent again by the next `start`, so a change may
|
|
66
|
+
reach the index twice; the transform is a function of the document, so
|
|
67
|
+
applying it twice is applying it once.
|
|
68
|
+
|
|
69
|
+
## A quiet collection keeps its position fresh
|
|
70
|
+
|
|
71
|
+
Every `positionIntervalMs` (60 s), a sync with nothing to send records where
|
|
72
|
+
the stream is anyway — one small write per interval, and none at all while
|
|
73
|
+
the stream itself does not move.
|
|
74
|
+
|
|
75
|
+
Without it, a collection nothing writes to for longer than the server's
|
|
76
|
+
history covers would need a full reindex at the next start.
|
|
77
|
+
|
|
78
|
+
## When the history is gone
|
|
79
|
+
|
|
80
|
+
MongoDB keeps a bounded history of changes, the oplog. A sync stopped for
|
|
81
|
+
longer than it covers cannot resume from its recorded point. By default
|
|
82
|
+
`start` reindexes and follows from there. `onHistoryLost: 'fail'` hands the
|
|
83
|
+
decision back:
|
|
84
|
+
|
|
85
|
+
```ts
|
|
86
|
+
import { SearchSyncError } from '@nxgt/mongo-meilisearch';
|
|
87
|
+
|
|
88
|
+
const articleSearch = createSearchSync({
|
|
89
|
+
collection,
|
|
90
|
+
index,
|
|
91
|
+
transform,
|
|
92
|
+
onHistoryLost: 'fail',
|
|
93
|
+
});
|
|
94
|
+
|
|
95
|
+
let running: RunningSearchSync;
|
|
96
|
+
try {
|
|
97
|
+
running = await articleSearch.start();
|
|
98
|
+
} catch (error) {
|
|
99
|
+
if (error instanceof SearchSyncError && error.code === 'HISTORY_LOST') {
|
|
100
|
+
await articleSearch.reindex(); // when it suits you
|
|
101
|
+
running = await articleSearch.start();
|
|
102
|
+
} else throw error;
|
|
103
|
+
}
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
`onHistoryLost` applies to `start` alone. A history that runs out **under** a
|
|
107
|
+
running sync stops it with `FAILED`; the next `start` is where the choice is
|
|
108
|
+
made again.
|
|
109
|
+
|
|
110
|
+
## When it stops
|
|
111
|
+
|
|
112
|
+
```ts
|
|
113
|
+
running.closed.then(
|
|
114
|
+
(reason) => console.info(`sync stopped: ${reason}`), // 'closed' | 'invalidated'
|
|
115
|
+
(error: SearchSyncError) => {
|
|
116
|
+
console.error({ sync: error.sync, code: error.code, cause: error.cause });
|
|
117
|
+
process.exit(1); // let the supervisor restart it
|
|
118
|
+
},
|
|
119
|
+
);
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
- **`close()`** — `closed` resolves `'closed'`. What was waiting is flushed
|
|
123
|
+
and recorded first.
|
|
124
|
+
- **The collection is dropped or renamed** — `closed` resolves
|
|
125
|
+
`'invalidated'`, and the recorded point is forgotten, since nothing could
|
|
126
|
+
resume from inside a collection that no longer exists. The next `start`
|
|
127
|
+
reindexes what the recreated collection holds.
|
|
128
|
+
- **An error** — `closed` rejects with a `SearchSyncError`, and so do `flush`
|
|
129
|
+
and `close` from then on. Nothing past the last applied batch was recorded,
|
|
130
|
+
so the next `start` sends it again.
|
|
131
|
+
|
|
132
|
+
**Await `closed`, or catch it.** A rejection nobody handles ends the process.
|
|
133
|
+
|
|
134
|
+
| `code` | When |
|
|
135
|
+
| --- | --- |
|
|
136
|
+
| `HISTORY_LOST` | `start` with `onHistoryLost: 'fail'`, and the recorded point is older than the server's history. `cause` is `@nxgt/mongo`'s `DataError`, `serverCode` 286 or 280 |
|
|
137
|
+
| `ID_MISMATCH` | the transform gave a document whose primary key is not its index id |
|
|
138
|
+
| `RUNNING` | a second `start()`, or a `reindex()`, while this sync is already following in this process |
|
|
139
|
+
| `FAILED` | anything else: the transform threw, MongoDB or Meilisearch refused. The message says what the sync was doing, and `cause` carries the original error |
|
|
140
|
+
|
|
141
|
+
[Troubleshooting](../troubleshooting.md) has each message with its fix.
|
|
142
|
+
|
|
143
|
+
## In a worker
|
|
144
|
+
|
|
145
|
+
```ts
|
|
146
|
+
const running = await articleSearch.start();
|
|
147
|
+
|
|
148
|
+
const stop = async () => {
|
|
149
|
+
await running.close(); // flushes and records before exiting
|
|
150
|
+
await mongo.close();
|
|
151
|
+
process.exit(0);
|
|
152
|
+
};
|
|
153
|
+
process.on('SIGTERM', stop);
|
|
154
|
+
process.on('SIGINT', stop);
|
|
155
|
+
|
|
156
|
+
await running.closed; // rejects if the sync falls over; the supervisor restarts it
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
One process per sync name: there is no lock, and two followers do the same
|
|
160
|
+
writes twice. See [What it leaves out](boundaries.md).
|
|
161
|
+
|
|
162
|
+
## Signatures
|
|
163
|
+
|
|
164
|
+
```ts
|
|
165
|
+
interface SearchSync {
|
|
166
|
+
start(): Promise<RunningSearchSync>;
|
|
167
|
+
}
|
|
168
|
+
|
|
169
|
+
interface RunningSearchSync extends AsyncDisposable {
|
|
170
|
+
readonly ready: Promise<void>;
|
|
171
|
+
readonly closed: Promise<CloseReason>; // 'closed' | 'invalidated'
|
|
172
|
+
flush(): Promise<void>;
|
|
173
|
+
close(): Promise<void>;
|
|
174
|
+
}
|
|
175
|
+
|
|
176
|
+
class SearchSyncError extends Error {
|
|
177
|
+
readonly code: SearchSyncErrorCode; // 'HISTORY_LOST' | 'ID_MISMATCH' | 'RUNNING' | 'FAILED'
|
|
178
|
+
readonly sync: string;
|
|
179
|
+
constructor(message: string, options: SearchSyncErrorOptions);
|
|
180
|
+
}
|
|
181
|
+
|
|
182
|
+
interface SearchSyncErrorOptions {
|
|
183
|
+
code: SearchSyncErrorCode;
|
|
184
|
+
sync: string;
|
|
185
|
+
cause?: unknown;
|
|
186
|
+
}
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
`CloseReason` is `@nxgt/mongo`'s; its third value, `'failed'`, never occurs
|
|
190
|
+
here — a failure rejects instead.
|
|
191
|
+
|
|
192
|
+
## Next
|
|
193
|
+
|
|
194
|
+
- [Reindexing](reindex.md) — the fill `start` runs the first time.
|
|
195
|
+
- [What it leaves out](boundaries.md).
|
|
@@ -0,0 +1,132 @@
|
|
|
1
|
+
# Reindexing
|
|
2
|
+
|
|
3
|
+
`reindex()` fills the index from the collection: every live document through
|
|
4
|
+
the transform, and then out of the index everything the collection no longer
|
|
5
|
+
gives it.
|
|
6
|
+
|
|
7
|
+
```ts
|
|
8
|
+
import { bindIndex } from '@nxgt/meilisearch';
|
|
9
|
+
import { getCollection } from '@nxgt/mongo';
|
|
10
|
+
import { createSearchSync } from '@nxgt/mongo-meilisearch';
|
|
11
|
+
import { db, meili } from './clients'; // a driver Db, an SDK client
|
|
12
|
+
import { articles, articleIndex } from './search';
|
|
13
|
+
|
|
14
|
+
const articleSearch = createSearchSync({
|
|
15
|
+
collection: getCollection(db, articles),
|
|
16
|
+
index: bindIndex(meili, articleIndex),
|
|
17
|
+
transform: (article) =>
|
|
18
|
+
article.draft ? null : { id: String(article._id), title: article.title },
|
|
19
|
+
});
|
|
20
|
+
|
|
21
|
+
const report = await articleSearch.reindex();
|
|
22
|
+
// { indexed: 1204, skipped: 17, removed: 3 }
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
| Field | Type | Effect |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| `indexed` | `number` | Documents sent to the index |
|
|
28
|
+
| `skipped` | `number` | Live documents the transform returned `null` for |
|
|
29
|
+
| `removed` | `number` | Documents the index held and the collection no longer gives it |
|
|
30
|
+
|
|
31
|
+
## What it does, in order
|
|
32
|
+
|
|
33
|
+
1. **Takes the collection's current position** — a change stream's first read
|
|
34
|
+
answers with it, without waiting for a change.
|
|
35
|
+
2. **Reads every live document**, `pageSize` at a time (100 by default,
|
|
36
|
+
lowered to the collection's `maxPageSize` when it asks for more), runs the
|
|
37
|
+
transform, and sends what it gives in batches of `batchSize` (500).
|
|
38
|
+
Soft-deleted documents are not live, so they are never sent.
|
|
39
|
+
3. **Pages the whole index** and deletes every document whose id the
|
|
40
|
+
collection did not just give it: deleted, turned away by the transform, or
|
|
41
|
+
never from this collection at all.
|
|
42
|
+
4. **Records the position from step 1** as the resume point, and stamps
|
|
43
|
+
`reindexedAt`.
|
|
44
|
+
|
|
45
|
+
Step 1 before step 2 is what makes it safe to reindex a live collection: a
|
|
46
|
+
change made while the documents are read is followed again from that
|
|
47
|
+
position, so nothing falls between the reindex and the stream. A reindex that
|
|
48
|
+
throws records nothing, and the next one starts over.
|
|
49
|
+
|
|
50
|
+
```ts
|
|
51
|
+
const before = await articleSearch.state(); // undefined, the first time
|
|
52
|
+
await articleSearch.reindex();
|
|
53
|
+
const after = await articleSearch.state();
|
|
54
|
+
after?.reindexedAt; // a Date
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
## When to run it
|
|
58
|
+
|
|
59
|
+
- **The first `start`** runs it for you, when nothing is recorded under the
|
|
60
|
+
sync's name.
|
|
61
|
+
- **After the transform changes** — a new field, a different rule for `null`
|
|
62
|
+
— since the index still holds what the old one produced.
|
|
63
|
+
- **After a stop longer than the server's change history**, which `start`
|
|
64
|
+
does itself unless `onHistoryLost: 'fail'` asks to decide.
|
|
65
|
+
See [Following changes](following-changes.md).
|
|
66
|
+
|
|
67
|
+
It is a **deployment step, not a request-time one**: it holds one id per live
|
|
68
|
+
document in memory and scans the whole index, which on a collection of
|
|
69
|
+
millions is hundreds of megabytes and a full pass over the index.
|
|
70
|
+
|
|
71
|
+
## Not beside its own follower
|
|
72
|
+
|
|
73
|
+
```ts
|
|
74
|
+
const running = await articleSearch.start();
|
|
75
|
+
await articleSearch.reindex();
|
|
76
|
+
// SearchSyncError: Search sync "articles:articles" is already following
|
|
77
|
+
// changes in this process: close it before you reindex.
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
The code is `RUNNING`. A reindex removes what the follower has just indexed,
|
|
81
|
+
and the follower, having already handled that change, would never send it
|
|
82
|
+
again. Close the follower, reindex, start again:
|
|
83
|
+
|
|
84
|
+
```ts
|
|
85
|
+
await running.close();
|
|
86
|
+
await articleSearch.reindex();
|
|
87
|
+
const fresh = await articleSearch.start();
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
The check is per object and per process: the package has no lock across
|
|
91
|
+
processes, and [does not pretend to](boundaries.md).
|
|
92
|
+
|
|
93
|
+
## Errors
|
|
94
|
+
|
|
95
|
+
Anything that goes wrong while reindexing comes back as a `SearchSyncError`
|
|
96
|
+
with the code `FAILED` — MongoDB refused a read, Meilisearch refused a batch,
|
|
97
|
+
the transform threw — and the original error as its `cause`:
|
|
98
|
+
|
|
99
|
+
```ts
|
|
100
|
+
try {
|
|
101
|
+
await articleSearch.reindex();
|
|
102
|
+
} catch (error) {
|
|
103
|
+
if (error instanceof SearchSyncError) {
|
|
104
|
+
console.error(error.code, error.sync, error.cause);
|
|
105
|
+
}
|
|
106
|
+
throw error;
|
|
107
|
+
}
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
`ID_MISMATCH` is the one exception that means something more precise: the
|
|
111
|
+
transform gave a document whose primary key is not that document's index id.
|
|
112
|
+
[Troubleshooting](../troubleshooting.md) has both messages.
|
|
113
|
+
|
|
114
|
+
## Signatures
|
|
115
|
+
|
|
116
|
+
```ts
|
|
117
|
+
interface SearchSync {
|
|
118
|
+
reindex(): Promise<ReindexReport>;
|
|
119
|
+
}
|
|
120
|
+
|
|
121
|
+
interface ReindexReport {
|
|
122
|
+
indexed: number;
|
|
123
|
+
skipped: number;
|
|
124
|
+
removed: number;
|
|
125
|
+
}
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
## Next
|
|
129
|
+
|
|
130
|
+
- [Following changes](following-changes.md) — what runs after the fill.
|
|
131
|
+
- [The sync's lifecycle](sync-lifecycle.md) — `pageSize`, `batchSize` and the
|
|
132
|
+
rest.
|
|
@@ -0,0 +1,234 @@
|
|
|
1
|
+
# The sync's lifecycle
|
|
2
|
+
|
|
3
|
+
A sync is the pairing of one MongoDB collection with one Meilisearch index:
|
|
4
|
+
`createSearchSync` describes it, and then `reindex` fills the index and
|
|
5
|
+
`start` keeps it in step.
|
|
6
|
+
|
|
7
|
+
```ts
|
|
8
|
+
import { bindIndex } from '@nxgt/meilisearch';
|
|
9
|
+
import { getCollection } from '@nxgt/mongo';
|
|
10
|
+
import { createSearchSync } from '@nxgt/mongo-meilisearch';
|
|
11
|
+
import { db, meili } from './clients'; // a driver Db, an SDK client
|
|
12
|
+
import { articles, articleIndex } from './search';
|
|
13
|
+
|
|
14
|
+
const articleSearch = createSearchSync({
|
|
15
|
+
collection: getCollection(db, articles),
|
|
16
|
+
index: bindIndex(meili, articleIndex),
|
|
17
|
+
transform: (article) =>
|
|
18
|
+
article.draft
|
|
19
|
+
? null
|
|
20
|
+
: { id: String(article._id), title: article.title, body: article.body },
|
|
21
|
+
});
|
|
22
|
+
|
|
23
|
+
const running = await articleSearch.start(); // reindexes the first time
|
|
24
|
+
// …
|
|
25
|
+
await running.close();
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
`createSearchSync` **sends nothing**: it validates its options, resolves the
|
|
29
|
+
defaults, and gives back an object. Every write comes from
|
|
30
|
+
[`reindex`](reindex.md) or [`start`](following-changes.md).
|
|
31
|
+
|
|
32
|
+
## The two definitions
|
|
33
|
+
|
|
34
|
+
The collection is `@nxgt/mongo`'s and the index is `@nxgt/meilisearch`'s;
|
|
35
|
+
this package creates neither, and neither knows about it.
|
|
36
|
+
|
|
37
|
+
```ts
|
|
38
|
+
import { defineIndex } from '@nxgt/meilisearch';
|
|
39
|
+
import { defineCollection, id } from '@nxgt/mongo';
|
|
40
|
+
import { z } from 'zod';
|
|
41
|
+
|
|
42
|
+
export const articles = defineCollection({
|
|
43
|
+
name: 'articles',
|
|
44
|
+
schema: z.object({
|
|
45
|
+
_id: id(),
|
|
46
|
+
title: z.string(),
|
|
47
|
+
body: z.string(),
|
|
48
|
+
draft: z.boolean().default(false),
|
|
49
|
+
}),
|
|
50
|
+
timestamps: true,
|
|
51
|
+
softDelete: true,
|
|
52
|
+
});
|
|
53
|
+
|
|
54
|
+
export interface ArticleHit {
|
|
55
|
+
id: string;
|
|
56
|
+
title: string;
|
|
57
|
+
body: string;
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
export const articleIndex = defineIndex<ArticleHit>()({
|
|
61
|
+
uid: 'articles',
|
|
62
|
+
primaryKey: 'id',
|
|
63
|
+
settings: { searchableAttributes: ['title', 'body'] },
|
|
64
|
+
});
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
The index's **settings** are `@nxgt/meilisearch`'s to apply, as a deployment
|
|
68
|
+
step beside the collection's own sync:
|
|
69
|
+
|
|
70
|
+
```ts
|
|
71
|
+
const index = bindIndex(meili, articleIndex);
|
|
72
|
+
await index.sync(); // creates the index and applies its settings
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
## The transform
|
|
76
|
+
|
|
77
|
+
It takes the collection's document, typed by its schema, and gives the
|
|
78
|
+
index's document, typed by its definition. `null` keeps a document **out** of
|
|
79
|
+
the index — and takes it out if it was in.
|
|
80
|
+
|
|
81
|
+
```ts
|
|
82
|
+
const transform = (article: ReadDocumentOf<typeof articles>) =>
|
|
83
|
+
article.draft
|
|
84
|
+
? null
|
|
85
|
+
: { id: String(article._id), title: article.title, body: article.body };
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
It runs for every document of a reindex and for every change that is
|
|
89
|
+
followed, so a document that stops qualifying is removed the moment it does.
|
|
90
|
+
It may be `async` — the type allows a promise — and it should be a function
|
|
91
|
+
of its document alone: a change may be applied twice, and a transform that
|
|
92
|
+
throws stops the sync.
|
|
93
|
+
|
|
94
|
+
The document it returns must carry its **index id** under the index's primary
|
|
95
|
+
key. One under another id could never be taken out again, so it is refused
|
|
96
|
+
with `ID_MISMATCH`.
|
|
97
|
+
|
|
98
|
+
## Options
|
|
99
|
+
|
|
100
|
+
| Option | Type | Default | Effect |
|
|
101
|
+
| --- | --- | --- | --- |
|
|
102
|
+
| `collection` | `TypedCollection<C>` | — | Where the documents are: from `getCollection` |
|
|
103
|
+
| `index` | `TypedIndex<I>` | — | Where they go: from `bindIndex` |
|
|
104
|
+
| `transform` | `Transform<C, I>` | — | The document as the index holds it, or `null` to keep it out |
|
|
105
|
+
| `toIndexId` | `ToIndexId<C, I>` | `String` | The index id of a `_id`. Optional while the index's ids are strings, required otherwise |
|
|
106
|
+
| `name` | `string` | `'<collection>:<index uid>'` | What the resume point is recorded under. Two syncs under one name share it |
|
|
107
|
+
| `stateCollection` | `string` | `'nxgt_search_sync'` | Where the resume point is kept, in the collection's database |
|
|
108
|
+
| `batchSize` | `number` | `500` | Changes, or documents, sent to Meilisearch at once |
|
|
109
|
+
| `flushIntervalMs` | `number` | `1000` | How long a change waits for others; `0` sends at the next tick |
|
|
110
|
+
| `positionIntervalMs` | `number` | `60000` | How often a sync with nothing to send records where the stream is |
|
|
111
|
+
| `pageSize` | `number` | `100` | Documents a reindex reads per page; above the collection's `maxPageSize`, lowered to it |
|
|
112
|
+
| `onHistoryLost` | `'reindex' \| 'fail'` | `'reindex'` | What `start` does when the resume point is older than the server's change history |
|
|
113
|
+
|
|
114
|
+
### `toIndexId`
|
|
115
|
+
|
|
116
|
+
`String(_id)` by default, which suits an `ObjectId`. When the index's primary
|
|
117
|
+
key is not a string, the option stops being optional:
|
|
118
|
+
|
|
119
|
+
```ts
|
|
120
|
+
createSearchSync({
|
|
121
|
+
collection: getCollection(db, counters), // _id: z.int()
|
|
122
|
+
index: bindIndex(meili, counterIndex), // primaryKey: 'n', a number
|
|
123
|
+
toIndexId: (id) => id,
|
|
124
|
+
transform: (counter) => ({ n: counter._id, label: counter.label }),
|
|
125
|
+
});
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
### `name` and `stateCollection`
|
|
129
|
+
|
|
130
|
+
The resume point is a document in the collection's own database, so a restart
|
|
131
|
+
picks up where the last run stopped:
|
|
132
|
+
|
|
133
|
+
```ts
|
|
134
|
+
const sync = createSearchSync({
|
|
135
|
+
collection,
|
|
136
|
+
index,
|
|
137
|
+
transform,
|
|
138
|
+
name: 'articles:public',
|
|
139
|
+
});
|
|
140
|
+
|
|
141
|
+
await sync.state();
|
|
142
|
+
// { _id: 'articles:public', resumeToken: …, updatedAt: …, reindexedAt: … }
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
`state()` is `undefined` until the first reindex records something.
|
|
146
|
+
|
|
147
|
+
### `batchSize`, `flushIntervalMs`, `pageSize`
|
|
148
|
+
|
|
149
|
+
`batchSize`, `positionIntervalMs` and `pageSize` must be whole numbers above
|
|
150
|
+
zero and `flushIntervalMs` a whole number of milliseconds, or
|
|
151
|
+
`createSearchSync` throws a `TypeError` before anything runs — as it does for
|
|
152
|
+
an empty `name` or a `transform` that is not a function.
|
|
153
|
+
|
|
154
|
+
```ts
|
|
155
|
+
createSearchSync({ collection, index, transform, batchSize: 0 });
|
|
156
|
+
// TypeError: createSearchSync: batchSize must be a whole number above 0, not 0
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
## A worker process
|
|
160
|
+
|
|
161
|
+
Wiring, settings, first fill and shutdown, in the order they happen:
|
|
162
|
+
|
|
163
|
+
```ts
|
|
164
|
+
import { bindIndex } from '@nxgt/meilisearch';
|
|
165
|
+
import { connectMongo, getCollection } from '@nxgt/mongo';
|
|
166
|
+
import { createSearchSync, SearchSyncError } from '@nxgt/mongo-meilisearch';
|
|
167
|
+
import { Meilisearch } from 'meilisearch';
|
|
168
|
+
import { articles, articleIndex } from './search';
|
|
169
|
+
|
|
170
|
+
const mongo = await connectMongo(process.env.MONGO_URI!);
|
|
171
|
+
const meili = new Meilisearch({
|
|
172
|
+
host: process.env.MEILI_HOST!,
|
|
173
|
+
apiKey: process.env.MEILI_KEY!,
|
|
174
|
+
});
|
|
175
|
+
|
|
176
|
+
const index = bindIndex(meili, articleIndex);
|
|
177
|
+
await index.sync();
|
|
178
|
+
|
|
179
|
+
const articleSearch = createSearchSync({
|
|
180
|
+
collection: getCollection(mongo.db, articles),
|
|
181
|
+
index,
|
|
182
|
+
transform: (article) =>
|
|
183
|
+
article.draft
|
|
184
|
+
? null
|
|
185
|
+
: { id: String(article._id), title: article.title, body: article.body },
|
|
186
|
+
});
|
|
187
|
+
|
|
188
|
+
const running = await articleSearch.start();
|
|
189
|
+
process.on('SIGTERM', () => {
|
|
190
|
+
void running.close().then(() => mongo.close());
|
|
191
|
+
});
|
|
192
|
+
|
|
193
|
+
// Rejects with a SearchSyncError if the sync stops on an error.
|
|
194
|
+
await running.closed.catch((error: SearchSyncError) => {
|
|
195
|
+
console.error({ sync: error.sync, code: error.code, cause: error.cause });
|
|
196
|
+
process.exit(1); // let the supervisor restart it
|
|
197
|
+
});
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
## Signatures
|
|
201
|
+
|
|
202
|
+
```ts
|
|
203
|
+
function createSearchSync<
|
|
204
|
+
C extends AnyCollectionDefinition,
|
|
205
|
+
I extends AnyIndexDefinition,
|
|
206
|
+
>(options: SearchSyncOptions<C, I>): SearchSync;
|
|
207
|
+
|
|
208
|
+
type Transform<C, I> = (
|
|
209
|
+
document: ReadDocumentOf<C>,
|
|
210
|
+
) => DocumentOf<I> | null | Promise<DocumentOf<I> | null>;
|
|
211
|
+
|
|
212
|
+
type ToIndexId<C, I> = (id: IdOf<C>) => IdOf<I>;
|
|
213
|
+
|
|
214
|
+
interface SearchSync {
|
|
215
|
+
readonly name: string;
|
|
216
|
+
reindex(): Promise<ReindexReport>;
|
|
217
|
+
start(): Promise<RunningSearchSync>;
|
|
218
|
+
state(): Promise<SearchSyncState | undefined>;
|
|
219
|
+
}
|
|
220
|
+
|
|
221
|
+
interface SearchSyncState {
|
|
222
|
+
_id: string;
|
|
223
|
+
resumeToken: ResumeToken;
|
|
224
|
+
updatedAt: Date;
|
|
225
|
+
reindexedAt: Date | undefined;
|
|
226
|
+
}
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
## Next
|
|
230
|
+
|
|
231
|
+
- [Reindexing](reindex.md) — the full fill, and what it removes.
|
|
232
|
+
- [Following changes](following-changes.md) — the change stream, its batches
|
|
233
|
+
and its errors.
|
|
234
|
+
- [What it leaves out](boundaries.md).
|
package/docs/roadmap.md
ADDED
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# Roadmap
|
|
2
|
+
|
|
3
|
+
Where `@nxgt/mongo-meilisearch` is going. A direction, not a commitment: the
|
|
4
|
+
version an item shipped in is the only number on this page.
|
|
5
|
+
|
|
6
|
+
## Now
|
|
7
|
+
|
|
8
|
+
- **A lease on a sync name, renewed while it runs** — a follower holds the
|
|
9
|
+
sync's name for a bounded time and renews it as it works, so a process that
|
|
10
|
+
dies is taken over once its lease expires rather than leaving the index
|
|
11
|
+
behind, and a name nothing is following no longer blocks the next start.
|
|
12
|
+
|
|
13
|
+
## Next
|
|
14
|
+
|
|
15
|
+
_Nothing queued._
|
|
16
|
+
|
|
17
|
+
## Later
|
|
18
|
+
|
|
19
|
+
_Nothing queued._
|
|
20
|
+
|
|
21
|
+
## Not planned
|
|
22
|
+
|
|
23
|
+
- **Partial updates** — a change sends the whole document the transform gives,
|
|
24
|
+
never a patch. The transform is a function of the document alone, which is
|
|
25
|
+
also what makes a change safe to apply twice.
|
|
26
|
+
- **Keeping the index's settings** — that is `@nxgt/meilisearch`'s `syncIndex`
|
|
27
|
+
/ `syncIndexes`, run as a deployment step. This package writes documents.
|
|
28
|
+
- **Joining other collections into a document** — the transform may read them,
|
|
29
|
+
but a change to one of them does not reach the index: the sync follows the
|
|
30
|
+
collection it was given. Run a sync per collection that has to move the
|
|
31
|
+
index.
|
|
32
|
+
- **One index fed by two collections, or by another writer** — a reindex
|
|
33
|
+
removes every document the collection does not give it, whoever wrote it.
|
|
34
|
+
The sync owns its index.
|
|
35
|
+
|
|
36
|
+
## Shipped
|
|
37
|
+
|
|
38
|
+
- **Documentation that travels with the package** — a guide page for the
|
|
39
|
+
sync's lifecycle, following a collection's changes, `reindex` and where the
|
|
40
|
+
package's boundaries are, a troubleshooting page whose headings are the
|
|
41
|
+
exact error text, and this roadmap, installed in `docs/` rather than left on
|
|
42
|
+
GitHub — 0.1.7.
|
|
43
|
+
- **`@nxgt/mongo` 0.15.0** — a deduplicated file write two callers cannot
|
|
44
|
+
both win — 0.1.6.
|
|
45
|
+
- **`@nxgt/mongo` 0.14.0** — files under its `./gridfs` subpath — 0.1.5.
|
|
46
|
+
- **`@nxgt/mongo` 0.13.0** — `upsert` in one round trip — 0.1.4.
|
|
47
|
+
- **A reindex's cost written down** — it holds one id per live document in
|
|
48
|
+
memory and pages the whole index, so it is a deployment step — 0.1.3.
|
|
49
|
+
- **`@nxgt/mongo` 0.12.0** — strings from outside read from the schema —
|
|
50
|
+
0.1.2.
|
|
51
|
+
- **A peer range that matches what it is built against** — `@nxgt/mongo`
|
|
52
|
+
`^0.11.0`, which has the `position` this package records while a collection
|
|
53
|
+
is quiet — 0.1.1.
|
|
54
|
+
- **First release** — `createSearchSync` with a transform typed by both
|
|
55
|
+
definitions (`null` keeps a document out), `reindex`, and `start`, which
|
|
56
|
+
follows the collection's changes in batches and resumes from a point
|
|
57
|
+
recorded in MongoDB, reindexing when the server's history no longer reaches
|
|
58
|
+
it; `SearchSyncError` with `HISTORY_LOST`, `ID_MISMATCH`, `RUNNING` and
|
|
59
|
+
`FAILED` — 0.1.0.
|
|
60
|
+
|
|
61
|
+
Everything released is in [`CHANGELOG.md`](https://github.com/softistx/nxgt-data/blob/develop/packages/mongo-meilisearch/CHANGELOG.md) — it is not in
|
|
62
|
+
the published package, only in the repository.
|
|
@@ -0,0 +1,380 @@
|
|
|
1
|
+
# Troubleshooting
|
|
2
|
+
|
|
3
|
+
Every heading is the text the error prints, so the page can be searched with
|
|
4
|
+
what you have in front of you. Stacks, ids and paths are cut, and a sync's
|
|
5
|
+
name is written as it comes out by default — `<collection>:<index uid>`, here
|
|
6
|
+
`articles:articles`.
|
|
7
|
+
|
|
8
|
+
Everything this package throws is a `SearchSyncError` carrying a `code`
|
|
9
|
+
(`HISTORY_LOST`, `ID_MISMATCH`, `RUNNING`, `FAILED`), the sync's `name`, and
|
|
10
|
+
the original error as `cause` — except the options, which are refused with a
|
|
11
|
+
`TypeError` before anything is opened.
|
|
12
|
+
|
|
13
|
+
| Area | Entries |
|
|
14
|
+
| --- | --- |
|
|
15
|
+
| [Install](#install) | [ERESOLVE](#npm-error-eresolve-unable-to-resolve-dependency-tree) · [incorrect peer dependency](#warn-incorrect-peer-dependency-nxgtmongo0140) · [TS2307](#error-ts2307-cannot-find-module-nxgtmongo-or-its-corresponding-type-declarations) |
|
|
16
|
+
| [Options](#options) | [transform](#createsearchsync-transform-must-be-a-function) · [batchSize](#createsearchsync-batchsize-must-be-a-whole-number-above-0-not-0) · [flushIntervalMs](#createsearchsync-flushintervalms-must-be-a-whole-number-of-milliseconds-not--1) · [name](#createsearchsync-name-must-not-be-empty) |
|
|
17
|
+
| [Starting](#starting) | [no replica set](#search-sync-articlesarticles-failed-starting-the-changestream-stage-is-only-supported-on-replica-sets) · [privileges](#search-sync-articlesarticles-failed-reindexing-not-authorized-on-app-to-execute-command--aggregate-articles-pipeline---changestream-----) · [history lost](#search-sync-articlesarticles-was-last-at-a-point-the-servers-change-history-no-longer-reaches-reindex-it-or-start-it-with-onhistorylost-reindex) · [already running](#search-sync-articlesarticles-is-already-following-changes-in-this-process-close-it-before-you-start-it-twice) |
|
|
18
|
+
| [The transform](#the-transform) | [an id that is not the index's](#search-sync-articlesarticles-transform-gave-id-other-for-the-document--whose-index-id-is-) · [not a document](#transform-must-return-a-document-or-null-not-nope) · [it threw](#search-sync-articlesarticles-failed-following-changes-boom) |
|
|
19
|
+
| [Sending](#sending) | [an id Meilisearch refuses](#search-sync-articlesarticles-failed-sending-changes-task-3-documentadditionorupdate-on-index-articles-failed-document-identifier--is-invalid) |
|
|
20
|
+
| [Stopping](#stopping) | [a dropped collection](#the-sync-stops-and-closed-resolves-with-invalidated) |
|
|
21
|
+
|
|
22
|
+
## Install
|
|
23
|
+
|
|
24
|
+
### `npm error ERESOLVE unable to resolve dependency tree`
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
npm error Found: @nxgt/mongo@0.14.0
|
|
28
|
+
npm error Could not resolve dependency:
|
|
29
|
+
npm error peer @nxgt/mongo@"^0.15.0" from @nxgt/mongo-meilisearch@0.1.6
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
**When:** `npm install`, before anything is downloaded.
|
|
33
|
+
|
|
34
|
+
**Why:** this package is a bridge: `@nxgt/mongo` and `@nxgt/meilisearch` are
|
|
35
|
+
**required peers**, and the ranges are carets on 0.x versions, so each accepts
|
|
36
|
+
one minor only. The collection and the index are built by those two packages
|
|
37
|
+
and handed here, so they have to be the copies this version was built against.
|
|
38
|
+
`mongodb`, `meilisearch` and `typescript` are peers on the same terms.
|
|
39
|
+
|
|
40
|
+
**Fix:**
|
|
41
|
+
|
|
42
|
+
```sh
|
|
43
|
+
npm install @nxgt/mongo@^0.15.0 @nxgt/meilisearch@^0.1.0 @nxgt/mongo-meilisearch mongodb meilisearch zod
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
Raise the siblings rather than install past the conflict: `--force` and
|
|
47
|
+
`--legacy-peer-deps` leave two copies of a sibling in the tree, and the
|
|
48
|
+
definitions one of them builds are not the ones this package reads.
|
|
49
|
+
|
|
50
|
+
### `warn: incorrect peer dependency "@nxgt/mongo@0.14.0"`
|
|
51
|
+
|
|
52
|
+
**When:** `bun install`, which prints it and carries on.
|
|
53
|
+
|
|
54
|
+
**Why:** bun does not fail on an unmet peer — it warns and installs what the
|
|
55
|
+
manifest asked for. The bridge then runs against a sibling it was not built
|
|
56
|
+
against, and what is missing shows up at the first call that needs it rather
|
|
57
|
+
than at install.
|
|
58
|
+
|
|
59
|
+
**Fix:**
|
|
60
|
+
|
|
61
|
+
```sh
|
|
62
|
+
bun add @nxgt/mongo@^0.15.0 @nxgt/meilisearch@^0.1.0
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
Treat that warning as an error. `bun pm ls` shows which versions were resolved.
|
|
66
|
+
|
|
67
|
+
### `error TS2307: Cannot find module '@nxgt/mongo' or its corresponding type declarations.`
|
|
68
|
+
|
|
69
|
+
At run time the same tree gives
|
|
70
|
+
`Error [ERR_MODULE_NOT_FOUND]: Cannot find package '@nxgt/mongo'` under Node,
|
|
71
|
+
and `error: Cannot find module '@nxgt/mongo'` under Bun.
|
|
72
|
+
|
|
73
|
+
**When:** the first build or the first import of your own code, typically
|
|
74
|
+
under pnpm or npm with a strict `node_modules` layout.
|
|
75
|
+
|
|
76
|
+
**Why:** a peer installed *for* this package is not a dependency of **yours**.
|
|
77
|
+
It is resolved under the bridge alone, so `@nxgt/mongo-meilisearch` finds it
|
|
78
|
+
and your own `import { getCollection } from '@nxgt/mongo'` does not.
|
|
79
|
+
|
|
80
|
+
**Fix:**
|
|
81
|
+
|
|
82
|
+
```sh
|
|
83
|
+
bun add @nxgt/mongo @nxgt/meilisearch @nxgt/mongo-meilisearch mongodb meilisearch zod
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
You import both siblings yourself — `getCollection` and `bindIndex` are what
|
|
87
|
+
build the two arguments this package takes.
|
|
88
|
+
|
|
89
|
+
## Options
|
|
90
|
+
|
|
91
|
+
`createSearchSync` sends nothing and opens nothing: everything below throws
|
|
92
|
+
where the sync is described.
|
|
93
|
+
|
|
94
|
+
### `createSearchSync: transform must be a function`
|
|
95
|
+
|
|
96
|
+
**When:** calling `createSearchSync`.
|
|
97
|
+
|
|
98
|
+
**Why:** `transform` is the one required function of the options, and it is
|
|
99
|
+
missing or is not callable — usually an object destructured from a config, or
|
|
100
|
+
an `await import` whose default was not unwrapped.
|
|
101
|
+
|
|
102
|
+
**Fix:**
|
|
103
|
+
|
|
104
|
+
```ts
|
|
105
|
+
createSearchSync({
|
|
106
|
+
collection,
|
|
107
|
+
index,
|
|
108
|
+
transform: (article) => ({ id: String(article._id), title: article.title }),
|
|
109
|
+
});
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
### `createSearchSync: batchSize must be a whole number above 0, not 0`
|
|
113
|
+
|
|
114
|
+
The same refusal covers `positionIntervalMs` and `pageSize`.
|
|
115
|
+
|
|
116
|
+
**When:** calling `createSearchSync`.
|
|
117
|
+
|
|
118
|
+
**Why:** those three count documents or milliseconds, and `0`, a fraction,
|
|
119
|
+
`NaN` and a negative number have no meaning for any of them. A value read from
|
|
120
|
+
the environment is a string until it is parsed, and `Number('')` is `0`.
|
|
121
|
+
|
|
122
|
+
**Fix:**
|
|
123
|
+
|
|
124
|
+
```ts
|
|
125
|
+
createSearchSync({ collection, index, transform, batchSize: 500 }); // the default
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
### `createSearchSync: flushIntervalMs must be a whole number of milliseconds, not -1`
|
|
129
|
+
|
|
130
|
+
**When:** calling `createSearchSync`.
|
|
131
|
+
|
|
132
|
+
**Why:** `flushIntervalMs` is the one option that accepts `0` — send every
|
|
133
|
+
change as it comes — so it is checked apart from the three above. A negative
|
|
134
|
+
number or a fraction is still refused.
|
|
135
|
+
|
|
136
|
+
**Fix:**
|
|
137
|
+
|
|
138
|
+
```ts
|
|
139
|
+
createSearchSync({ collection, index, transform, flushIntervalMs: 1000 }); // the default
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
### `createSearchSync: name must not be empty`
|
|
143
|
+
|
|
144
|
+
**When:** calling `createSearchSync` with `name: ''`.
|
|
145
|
+
|
|
146
|
+
**Why:** the name is the `_id` of the document the resume point is recorded
|
|
147
|
+
under, so it cannot be empty. Left out, it is `<collection>:<index uid>`.
|
|
148
|
+
|
|
149
|
+
**Fix:**
|
|
150
|
+
|
|
151
|
+
```ts
|
|
152
|
+
createSearchSync({ collection, index, transform, name: 'articles-search' });
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
Give a name when two syncs would otherwise share the default one — and keep it
|
|
156
|
+
stable, because changing it loses the recorded point and the next `start`
|
|
157
|
+
reindexes.
|
|
158
|
+
|
|
159
|
+
## Starting
|
|
160
|
+
|
|
161
|
+
### `Search sync "articles:articles" failed starting: The $changeStream stage is only supported on replica sets`
|
|
162
|
+
|
|
163
|
+
Code `FAILED`; the `cause` is the driver's `MongoServerError`, code 40573. The
|
|
164
|
+
same server message comes out of a first `reindex()` as
|
|
165
|
+
`… failed reindexing: …`, since a reindex takes the stream's position first.
|
|
166
|
+
|
|
167
|
+
**When:** `start()` or `reindex()`, against a standalone `mongod`.
|
|
168
|
+
|
|
169
|
+
**Why:** the sync follows the collection with a change stream, and MongoDB
|
|
170
|
+
serves change streams on a replica set or a sharded cluster only.
|
|
171
|
+
|
|
172
|
+
**Fix:** run a replica set — a single node is enough, and is what this
|
|
173
|
+
package's own specs use:
|
|
174
|
+
|
|
175
|
+
```sh
|
|
176
|
+
mongod --replSet rs0 --dbpath ./data # then, once: rs.initiate()
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
### `Search sync "articles:articles" failed reindexing: not authorized on app to execute command { aggregate: "articles", pipeline: [ { $changeStream: {} } ], … }`
|
|
180
|
+
|
|
181
|
+
The same shape appears for the state collection:
|
|
182
|
+
`… not authorized on app to execute command { update: "nxgt_search_sync", … }`.
|
|
183
|
+
|
|
184
|
+
**When:** `reindex()` or `start()`, against a server with authentication.
|
|
185
|
+
|
|
186
|
+
**Why:** a sync needs more than `find`. The MongoDB user needs `find` **and**
|
|
187
|
+
`changeStream` on the collection, and `find`, `insert`, `update` and `delete`
|
|
188
|
+
on the state collection (`nxgt_search_sync` by default).
|
|
189
|
+
|
|
190
|
+
**Fix:** grant them, and name the state collection if it lives elsewhere:
|
|
191
|
+
|
|
192
|
+
```ts
|
|
193
|
+
createSearchSync({ collection, index, transform, stateCollection: 'search_state' });
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
The Meilisearch key needs `documents.add`, `documents.get`, `documents.delete`
|
|
197
|
+
and `tasks.get`, plus `indexes.create` unless the index already exists.
|
|
198
|
+
|
|
199
|
+
### `Search sync "articles:articles" was last at a point the server's change history no longer reaches. Reindex it, or start it with onHistoryLost: 'reindex'.`
|
|
200
|
+
|
|
201
|
+
Code `HISTORY_LOST`; the `cause` carries the server's code, 286 or 280.
|
|
202
|
+
|
|
203
|
+
**When:** `start()`, on a sync that was stopped for longer than the server's
|
|
204
|
+
oplog covers.
|
|
205
|
+
|
|
206
|
+
**Why:** the resume point is a token into the oplog. Once the oplog has rolled
|
|
207
|
+
past it, MongoDB cannot say what happened in between, so following from there
|
|
208
|
+
would silently miss changes.
|
|
209
|
+
|
|
210
|
+
**Fix:** let it reindex, which is the default:
|
|
211
|
+
|
|
212
|
+
```ts
|
|
213
|
+
createSearchSync({ collection, index, transform, onHistoryLost: 'reindex' });
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
Keep `'fail'` where a reindex is too expensive to run unattended — then
|
|
217
|
+
`reindex()` it yourself when the alert comes in.
|
|
218
|
+
|
|
219
|
+
### `Search sync "articles:articles" is already following changes in this process: close it before you start it twice.`
|
|
220
|
+
|
|
221
|
+
Code `RUNNING`. `reindex()` on a sync that is following ends
|
|
222
|
+
`… before you reindex.`
|
|
223
|
+
|
|
224
|
+
**When:** a second `start()`, or a `reindex()`, on a sync object that is
|
|
225
|
+
already following.
|
|
226
|
+
|
|
227
|
+
**Why:** a reindex removes what the index holds and the collection no longer
|
|
228
|
+
gives it — including the documents the running follower has just indexed, which
|
|
229
|
+
it will never send again. The refusal covers one process only: there is **no
|
|
230
|
+
lock**, so two processes following one name is yours to prevent.
|
|
231
|
+
|
|
232
|
+
**Fix:**
|
|
233
|
+
|
|
234
|
+
```ts
|
|
235
|
+
const running = await articleSearch.start();
|
|
236
|
+
// …
|
|
237
|
+
await running.close(); // then reindex, or start again
|
|
238
|
+
await articleSearch.reindex();
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
## The transform
|
|
242
|
+
|
|
243
|
+
### `Search sync "articles:articles": transform gave "id" "other" for the document …, whose index id is …`
|
|
244
|
+
|
|
245
|
+
The message ends: *A document under another id could never be taken out of the
|
|
246
|
+
index again.* Code `ID_MISMATCH`.
|
|
247
|
+
|
|
248
|
+
**When:** `reindex()`, or a change the follower handles.
|
|
249
|
+
|
|
250
|
+
**Why:** the index document's primary key has to be the id this sync derives
|
|
251
|
+
from the Mongo `_id` — `String(_id)` by default, or whatever `toIndexId`
|
|
252
|
+
returns. Under any other id, a later delete would look for a document that is
|
|
253
|
+
not there and leave the wrong one in the index for good.
|
|
254
|
+
|
|
255
|
+
**Fix:** build the key from the document's own `_id`:
|
|
256
|
+
|
|
257
|
+
```ts
|
|
258
|
+
transform: (article) => ({ id: String(article._id), title: article.title });
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
With a non-string primary key, give `toIndexId` and use the same value:
|
|
262
|
+
|
|
263
|
+
```ts
|
|
264
|
+
createSearchSync({
|
|
265
|
+
collection,
|
|
266
|
+
index, // primaryKey: 'n', a number
|
|
267
|
+
toIndexId: (id) => Number(String(id).slice(0, 8)),
|
|
268
|
+
transform: (article) => ({ n: Number(String(article._id).slice(0, 8)), title: article.title }),
|
|
269
|
+
});
|
|
270
|
+
```
|
|
271
|
+
|
|
272
|
+
### `transform must return a document or null, not nope`
|
|
273
|
+
|
|
274
|
+
Reaches the caller wrapped: `Search sync "articles:articles" failed
|
|
275
|
+
reindexing: transform must return a document or null, not nope`.
|
|
276
|
+
|
|
277
|
+
**When:** `reindex()`, or a change the follower handles.
|
|
278
|
+
|
|
279
|
+
**Why:** the transform gave something that is not a plain object and is not
|
|
280
|
+
`null` — a string, a number, an array, or an implicit `undefined` from a branch
|
|
281
|
+
that returns nothing.
|
|
282
|
+
|
|
283
|
+
**Fix:** return `null` for a document that should stay out of the index, and
|
|
284
|
+
make every branch return:
|
|
285
|
+
|
|
286
|
+
```ts
|
|
287
|
+
transform: (article) => (article.draft ? null : { id: String(article._id), title: article.title });
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
`null` also **removes** a document that was in the index, which is how a soft
|
|
291
|
+
delete or an unpublish reaches search.
|
|
292
|
+
|
|
293
|
+
### `Search sync "articles:articles" failed following changes: boom`
|
|
294
|
+
|
|
295
|
+
Code `FAILED`; the `cause` is the error your transform threw, and `boom` is its
|
|
296
|
+
message.
|
|
297
|
+
|
|
298
|
+
**When:** while following, as soon as a change reaches a transform that throws.
|
|
299
|
+
`closed` rejects with it, `flush()` and `close()` reject with the same error,
|
|
300
|
+
and nothing past that change is recorded.
|
|
301
|
+
|
|
302
|
+
**Why:** the sync cannot skip a document it could not build: the next `start`
|
|
303
|
+
resumes from the last recorded point, so the same change arrives again and
|
|
304
|
+
stops it again.
|
|
305
|
+
|
|
306
|
+
**Fix:** keep the transform total — catch what you can, and return `null` for a
|
|
307
|
+
document you cannot index:
|
|
308
|
+
|
|
309
|
+
```ts
|
|
310
|
+
transform: (article) => {
|
|
311
|
+
try {
|
|
312
|
+
return { id: String(article._id), title: render(article.body) };
|
|
313
|
+
} catch {
|
|
314
|
+
return null; // out of the index rather than stopping the sync
|
|
315
|
+
}
|
|
316
|
+
};
|
|
317
|
+
```
|
|
318
|
+
|
|
319
|
+
And take the rejection, whatever you do with it: a rejected `closed` that
|
|
320
|
+
nobody handles ends the process.
|
|
321
|
+
|
|
322
|
+
```ts
|
|
323
|
+
running.closed.catch((error) => log.error(error));
|
|
324
|
+
```
|
|
325
|
+
|
|
326
|
+
## Sending
|
|
327
|
+
|
|
328
|
+
### `Search sync "articles:articles" failed sending changes: Task 3 (documentAdditionOrUpdate) on index "articles" failed: Document identifier … is invalid`
|
|
329
|
+
|
|
330
|
+
Meilisearch's own text follows: *A document identifier can be of type integer
|
|
331
|
+
or string, only composed of alphanumeric characters (a-z A-Z 0-9), hyphens (-)
|
|
332
|
+
and underscores (\_), and can not be more than 511 bytes.* Code `FAILED`; the
|
|
333
|
+
`cause` is `@nxgt/meilisearch`'s `SearchIndexError`, carrying the failed task.
|
|
334
|
+
|
|
335
|
+
**When:** a batch is sent — on `reindex()`, on a `flush()`, or from the
|
|
336
|
+
follower's own timer, in which case it surfaces on `closed`.
|
|
337
|
+
|
|
338
|
+
**Why:** Meilisearch has its own rules for ids, and the whole batch fails if
|
|
339
|
+
one document breaks them. An `ObjectId`'s 24-character hex string is fine; a
|
|
340
|
+
slug, an email or anything built with a space is not.
|
|
341
|
+
|
|
342
|
+
**Fix:** derive the index id from the `_id` and nothing else:
|
|
343
|
+
|
|
344
|
+
```ts
|
|
345
|
+
createSearchSync({
|
|
346
|
+
collection,
|
|
347
|
+
index,
|
|
348
|
+
toIndexId: (id) => String(id), // the default
|
|
349
|
+
transform: (article) => ({ id: String(article._id), title: article.title }),
|
|
350
|
+
});
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
A failed batch is not skipped: fix the transform and reindex, or the same
|
|
354
|
+
change stops the sync again.
|
|
355
|
+
|
|
356
|
+
## Stopping
|
|
357
|
+
|
|
358
|
+
### The sync stops and `closed` resolves with `'invalidated'`
|
|
359
|
+
|
|
360
|
+
Not an error: `closed` **resolves**, and the index is left as it was.
|
|
361
|
+
|
|
362
|
+
**When:** the collection is dropped or renamed under a running sync.
|
|
363
|
+
|
|
364
|
+
**Why:** MongoDB invalidates the change stream, and the point it stopped at
|
|
365
|
+
belongs to a collection that no longer exists. The recorded point is forgotten,
|
|
366
|
+
so the next `start()` reindexes from scratch rather than resuming into nothing.
|
|
367
|
+
|
|
368
|
+
**Fix:** watch `closed` for the reason, not only for a rejection — a dropped
|
|
369
|
+
collection is silent if you only handle failures:
|
|
370
|
+
|
|
371
|
+
```ts
|
|
372
|
+
const running = await articleSearch.start();
|
|
373
|
+
running.closed.then(
|
|
374
|
+
(reason) => { if (reason === 'invalidated') log.warn('collection dropped'); },
|
|
375
|
+
(error) => log.error(error),
|
|
376
|
+
);
|
|
377
|
+
```
|
|
378
|
+
|
|
379
|
+
Whatever the collection holds after it is recreated is indexed by the reindex
|
|
380
|
+
the next `start()` runs.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@nxgt/mongo-meilisearch",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.7",
|
|
4
4
|
"description": "Keeps a Meilisearch index in step with a MongoDB collection: a typed transform, a full reindex, and a change stream that resumes where it stopped",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -8,6 +8,7 @@
|
|
|
8
8
|
"types": "./dist/index.d.ts",
|
|
9
9
|
"files": [
|
|
10
10
|
"dist",
|
|
11
|
+
"docs",
|
|
11
12
|
"README.md",
|
|
12
13
|
"package.json",
|
|
13
14
|
"LICENSE"
|
|
@@ -48,8 +49,8 @@
|
|
|
48
49
|
]
|
|
49
50
|
},
|
|
50
51
|
"devDependencies": {
|
|
51
|
-
"@nxgt/meilisearch": "^0.1.
|
|
52
|
-
"@nxgt/mongo": "^0.15.
|
|
52
|
+
"@nxgt/meilisearch": "^0.1.1",
|
|
53
|
+
"@nxgt/mongo": "^0.15.1",
|
|
53
54
|
"@types/bun": "^1.4.0",
|
|
54
55
|
"meilisearch": "0.62.0",
|
|
55
56
|
"mongodb": "7.6.0",
|
|
@@ -57,8 +58,8 @@
|
|
|
57
58
|
"zod": "4.6.5"
|
|
58
59
|
},
|
|
59
60
|
"peerDependencies": {
|
|
60
|
-
"@nxgt/meilisearch": "^0.1.
|
|
61
|
-
"@nxgt/mongo": "^0.15.
|
|
61
|
+
"@nxgt/meilisearch": "^0.1.1",
|
|
62
|
+
"@nxgt/mongo": "^0.15.1",
|
|
62
63
|
"meilisearch": ">=0.62.0 <1",
|
|
63
64
|
"mongodb": ">=7.0.0 <8",
|
|
64
65
|
"typescript": "^6.0.3"
|