@alexify/migronaut 2.0.0 → 2.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +436 -0
- package/README.md +235 -6
- package/bullmq.d.ts +860 -0
- package/bullmq.js +1 -0
- package/index.d.ts +888 -19
- package/migronaut.schema.json +238 -1
- package/package.json +21 -5
- package/src/bullmq/index.js +55 -0
- package/src/bullmq/jobs.js +454 -0
- package/src/bullmq/processor.js +632 -0
- package/src/bullmq/producer.js +427 -0
- package/src/bullmq/service.js +653 -0
- package/src/bullmq/wait.js +124 -0
- package/src/cli/args.js +12 -2
- package/src/cli/commands/converge.js +188 -0
- package/src/cli/commands/down.js +2 -0
- package/src/cli/commands/lock.js +2 -1
- package/src/cli/commands/redo.js +8 -1
- package/src/cli/commands/up.js +14 -1
- package/src/cli/exit-codes.js +9 -2
- package/src/cli/index.js +2 -0
- package/src/cli/shared.js +14 -4
- package/src/cli/table.js +164 -0
- package/src/core/audit.js +88 -3
- package/src/core/changelog.js +71 -6
- package/src/core/collections.js +396 -0
- package/src/core/config.js +130 -25
- package/src/core/converge-log.js +47 -0
- package/src/core/converge-plan.js +686 -0
- package/src/core/converge-search-run.js +440 -0
- package/src/core/converge-search.js +404 -0
- package/src/core/converge.js +1024 -0
- package/src/core/index-spec.js +507 -0
- package/src/core/lock-wait.js +260 -0
- package/src/core/lock.js +95 -28
- package/src/core/migrator.js +600 -287
- package/src/core/options.js +266 -0
- package/src/core/run-recorder.js +157 -0
- package/src/core/run.js +58 -90
- package/src/core/search-index-spec.js +758 -0
- package/src/core/sequence.js +134 -0
- package/src/core/server-info.js +63 -0
- package/src/errors/index.js +60 -0
- package/src/index.js +8 -0
- package/src/utils/actor.js +48 -0
- package/src/utils/canonical.js +212 -0
- package/src/utils/collection-name.js +21 -0
- package/src/utils/error.js +18 -1
- package/src/utils/id.js +77 -0
- package/src/utils/loader.js +39 -21
- package/src/utils/migration-name.js +32 -0
- package/src/utils/redact.js +21 -1
- package/src/utils/telemetry.js +410 -0
- package/src/utils/template.js +43 -2
package/CHANGELOG.md
CHANGED
|
@@ -3,6 +3,442 @@
|
|
|
3
3
|
All notable changes to this project will be documented in this file.
|
|
4
4
|
Release headings carry the publish date (`## vX.Y.Z — YYYY-MM-DD`).
|
|
5
5
|
|
|
6
|
+
## v2.2.0 — 2026-10-05
|
|
7
|
+
|
|
8
|
+
Atlas Search and Vector Search indexes in declared collections. Additive: a definition without
|
|
9
|
+
`searchIndexes` behaves exactly as in 2.1 — converge never even asks the server about Search for
|
|
10
|
+
it.
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- **Declared search indexes** — a collection definition takes `searchIndexes: [{ name?, type?,
|
|
15
|
+
definition }]`: Atlas Search (`type: 'search'`, the default) and Atlas Vector Search
|
|
16
|
+
(`'vectorSearch'`) indexes, with the definition exactly as Atlas documents it — automated
|
|
17
|
+
embedding (`autoEmbed`) fields included. `converge` keeps them in step with the same lock,
|
|
18
|
+
history, events, re-plan and fixed-point check as regular indexes. Experimental.
|
|
19
|
+
- **Rows**: a new target, `searchIndex`. A missing index is created (one `createSearchIndexes`
|
|
20
|
+
per collection), a changed definition is updated **in place** (`modify` — the old definition
|
|
21
|
+
serves until the new one is built), an undeclared one is kept, or dropped last under
|
|
22
|
+
`prune`. A definition that declares only `searchIndexes` is valid, and its collection is
|
|
23
|
+
created when missing.
|
|
24
|
+
- **Never a rebuild**: a `$search` against a missing index returns nothing rather than fail, so
|
|
25
|
+
converge never drops a search index to build it again. A change no update can make — the
|
|
26
|
+
type, or an `autoEmbed` field's path, model, `numDimensions`, quantization or modality — is a
|
|
27
|
+
`conflict` naming the new-name recipe, and so is a declared index the server is still
|
|
28
|
+
deleting.
|
|
29
|
+
- **Comparison**: definitions compare whole, key order ignored, with the defaults the server
|
|
30
|
+
writes into what it reports filled in on both sides — top-level (`analyzer`,
|
|
31
|
+
`searchAnalyzer`, `dynamic`, `storedSource`, `numPartitions`), per field mapping (`string`,
|
|
32
|
+
`number`, `autocomplete`, `token`, `geo`, `document`, nested fields and `multi` included) and
|
|
33
|
+
per vector field (`quantization`, `indexingMethod`, `hnswOptions`; `autoEmbed`
|
|
34
|
+
`numDimensions` and `quantization`). Vector `fields`, and a field indexed as several types,
|
|
35
|
+
compare as sets. A server that reports no `type` (a self-managed `mongot`) has it inferred
|
|
36
|
+
from the definition, and its `latestVersion` is read as the definition version.
|
|
37
|
+
- **Server-only options**: an option the server reports that the declaration does not set,
|
|
38
|
+
and whose default migronaut does not know (a newer `mongot`'s), is left out of the
|
|
39
|
+
comparison — named on the row (`ignored`) and in one warning — instead of making every
|
|
40
|
+
converge update, and the server rebuild, the index. Only option objects are trimmed: a
|
|
41
|
+
field, a mapping type or a vector field only the server has is still a difference. The
|
|
42
|
+
cost: removing such an option from a declaration goes unnoticed; declare the value wanted.
|
|
43
|
+
- **What differs**: a `modify` row names it down to the option
|
|
44
|
+
(`mappings.fields.title.norms`) — the first five paths and how many more. A search index
|
|
45
|
+
that still differs after an update is `unstable`, with a warning that every update builds it
|
|
46
|
+
again.
|
|
47
|
+
- **Raw commands** (`createSearchIndexes`, `updateSearchIndex`, `dropSearchIndex`,
|
|
48
|
+
`$listSearchIndexes`), so every driver in the peer range works — the driver's helpers start
|
|
49
|
+
at 5.6. A vector index update is retried once with its type when a self-managed `mongot`
|
|
50
|
+
asks for it; where the server refuses that too (an Atlas CLI local deployment on 8.0 and
|
|
51
|
+
8.3), the step fails with the new-name recipe as its hint.
|
|
52
|
+
- **Build state**: rows carry `build` (`{ status, queryable, message?, updating? }`), and the
|
|
53
|
+
result `search` (`{ available, evidence, notReady, wait? }`) — how Search availability was
|
|
54
|
+
told, the declared indexes still building, updating, stale or failed, and how a wait ended.
|
|
55
|
+
A build under way does not count against `inSync`; a FAILED one fails `converge --check`
|
|
56
|
+
(exit 28). A STALE index (queryable, no longer replicating) is reported as such — table,
|
|
57
|
+
closing warning, `audit` — not as still building. A build message is kept to 500 characters.
|
|
58
|
+
- **`waitForSearchIndexes`** (config, `MIGRONAUT_WAIT_FOR_SEARCH_INDEXES`, `converge({
|
|
59
|
+
waitForSearchIndexes })`, `converge --wait-search` / `--no-wait-search`) — hold the converge
|
|
60
|
+
until every declared search index serves its declaration; a FAILED build of an index the run
|
|
61
|
+
created or changed, or `searchIndexWaitTimeoutMs` (`MIGRONAUT_SEARCH_INDEX_WAIT_TIMEOUT_MS`,
|
|
62
|
+
default 10 minutes), fails it with `ConvergeFailedError` `phase: 'wait'`. Off by default.
|
|
63
|
+
- **Without the lock**: the wait only reads, so the migration lock is released when it starts
|
|
64
|
+
(`lock:released` with `early: true`) — the next deploy or a queue's next job need not wait
|
|
65
|
+
out a build. The run itself (its id, span, history entry, `converge:end`) ends with the
|
|
66
|
+
wait.
|
|
67
|
+
- **Only what the run started holds it**: an index that failed or went STALE before the run,
|
|
68
|
+
its definition unchanged, is named and warned about instead — converge cannot fix it, and a
|
|
69
|
+
deploy is not held up by it. `--check` still fails on a FAILED build.
|
|
70
|
+
- **Resilient**: a list that fails with a network or failover blip is read again at the next
|
|
71
|
+
poll, up to three in a row (then `reason: 'unreadable'`); a stop cuts the pause between polls
|
|
72
|
+
short.
|
|
73
|
+
- **Observable**: `converge:wait` events (`started`, `progress` every 30 s, then `ready`,
|
|
74
|
+
`failed`, `timeout`, `unreadable` or `aborted`), `result.search.wait`, the history entry, a
|
|
75
|
+
queue job's log lines and `search-wait` progress phase, and — with `telemetry` — the
|
|
76
|
+
histogram `migronaut.converge.search.wait.duration` by `migronaut.converge.search.wait.outcome`.
|
|
77
|
+
- **`onSearchUnavailable`** (`MIGRONAUT_ON_SEARCH_UNAVAILABLE`) — declared search indexes on a
|
|
78
|
+
server without Atlas Search: `'fail'` (default) refuses the run before any write, with a hint;
|
|
79
|
+
`'skip'` converges everything else and reports `skip` rows. Detected once per run from the
|
|
80
|
+
server's own answer (checked against MongoDB 5.0, 6.0, 7.0, 8.0 and 8.2).
|
|
81
|
+
- **`migronaut audit`** — a `search` check, when the definitions declare search indexes: whether
|
|
82
|
+
the server has Atlas Search, and whether a declared index failed to build.
|
|
83
|
+
- **Types** — `SearchIndexDefinition`, `SearchDefinition`, `VectorSearchDefinition`,
|
|
84
|
+
`VectorSearchField`, `SearchIndexType`, `SearchIndexStatus`, `SearchIndexBuild`,
|
|
85
|
+
`SearchIndexNotReady`, `ConvergeSearchSummary`, `ConvergeWaitOutcome`, `ConvergeWaitEvent`;
|
|
86
|
+
`ConvergeTarget` gains `'searchIndex'`, `ConvergeActionKind` `'skip'`, `ConvergeAction`
|
|
87
|
+
`ignored`, `ConvergeHistoryEntry` `search`, `LockEvent` `early`, the event map
|
|
88
|
+
`'converge:wait'`; in `bullmq.d.ts` `ConvergeJobResult.search`, and `MigrationJobProgress`
|
|
89
|
+
gains the `'search-wait'` phase and `searchIndexes`. An exhaustive `switch` over one of these
|
|
90
|
+
unions needs a branch for the new member.
|
|
91
|
+
- **Queue** — a converge job's result carries `search`, and its log lines name search indexes.
|
|
92
|
+
Whether a job waits for builds is the worker kit's `waitForSearchIndexes`; the job payload is
|
|
93
|
+
unchanged.
|
|
94
|
+
|
|
95
|
+
### Changed
|
|
96
|
+
|
|
97
|
+
- A definition with none of `indexes`, `searchIndexes` and `validator` is refused with "declares
|
|
98
|
+
no indexes, searchIndexes or validator — nothing to manage" (was "declares neither indexes nor
|
|
99
|
+
a validator").
|
|
100
|
+
- The converge table's drop/rebuild count includes dropped search indexes, and `converge --yes`
|
|
101
|
+
is required for them in `--json` mode, like an index drop.
|
|
102
|
+
- `ConvergeFailedError`'s documentation names every phase: `plan`, `replan` (already thrown by
|
|
103
|
+
2.1, undocumented), `apply` and the new `wait`. A search index list that cannot be read is
|
|
104
|
+
reported in the phase that read it — `plan` only before the first write — and a run that stops
|
|
105
|
+
for any reason settles every row and carries the result so far in `context.converge`.
|
|
106
|
+
- `runWithLock` (internal) hands the work a `control` whose `release()` gives the lock up early;
|
|
107
|
+
`lock:released` may now come before `run:end`, with `early: true`.
|
|
108
|
+
|
|
109
|
+
### Tooling
|
|
110
|
+
|
|
111
|
+
- `tests/integration/search-atlas.test.js` — an opt-in, manual suite against
|
|
112
|
+
`mongodb/mongodb-atlas-local` (`MIGRONAUT_TEST_ATLAS_URI`): every scenario ends at a fixed
|
|
113
|
+
point. Passes against `mongodb/mongodb-atlas-local` 8.0 (8.0.32) and `latest` (8.3.11). CI does
|
|
114
|
+
not run it; the coverage gate comes from the unit tier's fake, which answers the search
|
|
115
|
+
commands with lag, normalization and build progress on demand. The suite stays strict about
|
|
116
|
+
server-only options (production tolerates them): one there is a default the tables in
|
|
117
|
+
`search-index-spec.js` lack. 9/9 on 8.0.32 and 8.3.11, the lock released during the wait
|
|
118
|
+
included.
|
|
119
|
+
- `src/core/converge.js` is split: `converge-search-run.js` (the search half of a run) and
|
|
120
|
+
`server-info.js` (read options, read pace, server version, not-found codes).
|
|
121
|
+
|
|
122
|
+
## v2.1.0 — 2026-10-04
|
|
123
|
+
|
|
124
|
+
Migrations as a queue, ids in your own format, OpenTelemetry, and declared collections. Additive:
|
|
125
|
+
nothing changes for anyone who uses none of them, with the narrow exceptions listed under
|
|
126
|
+
**Changed**.
|
|
127
|
+
|
|
128
|
+
### Added
|
|
129
|
+
|
|
130
|
+
- **`@alexify/migronaut/bullmq`** — a new entry point that runs migrations as
|
|
131
|
+
[BullMQ](https://docs.bullmq.io/) jobs, **one migration per job**, so migronaut can be a
|
|
132
|
+
migration service: enqueue from an HTTP handler, a schedule or a deploy hook, and let a worker
|
|
133
|
+
apply them in order.
|
|
134
|
+
- `createMigrationQueue(options)` — the facade: `enqueueUp` / `enqueueDown` (returning a group
|
|
135
|
+
handle with `wait()`), `startWorker`, `status` / `pending` / `audit` / `lockInfo`, `getJob`,
|
|
136
|
+
`pause` / `resume`, `schedule` / `unschedule`, `close`.
|
|
137
|
+
- `createMigrationProcessor(options)` — the processor on its own, for a Worker you construct
|
|
138
|
+
(NestJS, BullMQ Pro), plus `enqueueUp` / `enqueueDown` / `planUpJobs` / `planDownJobs` /
|
|
139
|
+
`waitForGroup` for a Queue you own.
|
|
140
|
+
- **BullMQ is injected, never depended on** — `bullmq: { Queue, Worker, QueueEvents }` from
|
|
141
|
+
your own install. The package still has no `dependencies` and gains no peer; `src/` never
|
|
142
|
+
imports `bullmq` (a test enforces it).
|
|
143
|
+
- **Order comes from MongoDB, not from Redis.** Every job is a single-file run under the usual
|
|
144
|
+
lock; it refuses while an earlier migration is still pending, so a failed migration stops the
|
|
145
|
+
line (`MIGRATION_BLOCKED` for the jobs behind it). Jobs get one attempt on purpose — a BullMQ
|
|
146
|
+
retry re-queues behind the waiting jobs — and a held lock is waited out inside the job.
|
|
147
|
+
- **One batch per enqueue**, so `down` still rolls back a whole deploy; duplicate enqueues are
|
|
148
|
+
deduplicated, and a job whose migration is already applied completes as `skipped`.
|
|
149
|
+
- Job payloads are validated as untrusted input; messages, stacks and job logs are redacted. A
|
|
150
|
+
queue job's target must be a file of the migration sequence — a payload can never make the
|
|
151
|
+
worker import a dotfile, a declaration file or a helper module next to the migrations.
|
|
152
|
+
- **Shutdown puts unstarted work back.** A job that a closing worker stops before its migration
|
|
153
|
+
starts (waiting for the lock, or fetched during shutdown) is moved back to the head of the
|
|
154
|
+
queue instead of failing — so a rolling deploy no longer fails the rest of the enqueue it
|
|
155
|
+
interrupts as `MIGRATION_BLOCKED`. A migration already running always finishes.
|
|
156
|
+
- **A versioned, strict job contract.** A worker accepts every job data version from
|
|
157
|
+
`MIN_JOB_DATA_VERSION` (exported) up to its own and refuses newer ones and unknown fields
|
|
158
|
+
rather than ignoring what they mean — roll workers out before producers. Jobs always state
|
|
159
|
+
`ordered`, and may carry the plan-time `checksum` of their file.
|
|
160
|
+
- **Scheduler ticks are bounded**: they carry the queue's `jobOptions`, and keep the last 100
|
|
161
|
+
completed / 500 failed jobs when those set no retention.
|
|
162
|
+
- **Correlation**: a job's `runId` is on its progress (`completed` and `failed`), its failure log
|
|
163
|
+
row and its error's `context` (with `jobId` and `groupId`) — a failed job has no return value;
|
|
164
|
+
the worker's failure log line names the group, migration and run.
|
|
165
|
+
- An injected `Queue` (or `QueueEvents`) on another name or prefix than the facade's is rejected,
|
|
166
|
+
and the facade takes an injected queue's prefix by default; `startWorker()` can be retried
|
|
167
|
+
after a failed start; a closed queue refuses every further call.
|
|
168
|
+
- **A worker decides what a job may ask for** — the `allow` option (`{ down: true, force:
|
|
169
|
+
false, unordered: false }` by default) on `createMigrationProcessor` and
|
|
170
|
+
`createMigrationQueue`: a payload asking to re-run an applied migration or to skip the order
|
|
171
|
+
guard is refused (`QUEUE_JOB_INVALID`, `context.permission`) unless allowed. The facade's
|
|
172
|
+
`enqueue*` calls follow the same policy.
|
|
173
|
+
- **A job runs only the file it was planned with**: an `up` job carries the plan-time checksum,
|
|
174
|
+
and a worker with another version of the file fails it (`CHECKSUM_MISMATCH`,
|
|
175
|
+
`context.planned`) instead of applying it.
|
|
176
|
+
- **Several workers without global concurrency**: a job blocked only by earlier migrations that
|
|
177
|
+
have not failed (one may be in flight on another worker) waits for them within its lock-wait
|
|
178
|
+
budget instead of failing its group; `MigrationBlockedError` carries `context.failed`.
|
|
179
|
+
- `wait()` holds every await to its one `timeoutMs` budget, decides a timeout by the clock, and
|
|
180
|
+
its `QueueJobFailedError` carries the job's own typed `code`; an empty group can be waited for
|
|
181
|
+
without QueueEvents. A long lock wait reports progress every few seconds, not every poll.
|
|
182
|
+
- What the queue stores about a failure — `failedReason`, stack, job logs — masks the values a
|
|
183
|
+
duplicate-key error quotes, on top of credentials. `schedule({ every })` needs at least 1000 ms.
|
|
184
|
+
- Dedup ids encode file names reversibly (two names can no longer share one and absorb each
|
|
185
|
+
other's job), and a forced re-run has a dedup id of its own.
|
|
186
|
+
- **A schedule holds a failed migration**: a `sync` tick whose next migration failed, with its
|
|
187
|
+
file unchanged since, enqueues nothing (`returnvalue.held`, and a warning) instead of re-running
|
|
188
|
+
it every tick; a changed file or an explicit `enqueueUp(name)` resumes it.
|
|
189
|
+
- Jobs carry `requestedBy` / `reason` (`enqueueUp` / `enqueueDown` / `enqueueConverge` options).
|
|
190
|
+
- **`bullmq.d.ts`** — hand-written types for the entry point, with structural `BullMQ*Like`
|
|
191
|
+
interfaces instead of an import of `bullmq`, generic over the classes you inject.
|
|
192
|
+
- **`up(file, { batch })`** — stamp an explicit batch number instead of the next free one, and
|
|
193
|
+
**`MigratorKit.nextBatch()`** to peek at it: together they let several single-file runs form one
|
|
194
|
+
batch.
|
|
195
|
+
- **`up(file, { ordered: true })` / `down(file, { ordered: true })`** — refuse a single-file run
|
|
196
|
+
that would go out of sequence (`MigrationBlockedError`); an ordered `up` also applies the
|
|
197
|
+
`strict` drift check and the `onOutOfOrder` policy a bulk run would, and only targets a file of
|
|
198
|
+
the migration sequence.
|
|
199
|
+
- **`up(file, { checksum })`** — refuse (`CHECKSUM_MISMATCH`, `context.planned`) to apply any other
|
|
200
|
+
version of the file than the one with this SHA-256; `dryRun('up')` rows carry each file's
|
|
201
|
+
`checksum` for it.
|
|
202
|
+
- **`list(filter, { checksums: false })`** / **`status({ checksums: false })`** — skip hashing the
|
|
203
|
+
applied files, for a caller that needs names and dates only.
|
|
204
|
+
- **Who asked, and why** — `up`, `down`, `redo` and `converge` take `requestedBy` (≤ 128
|
|
205
|
+
characters) and `reason` (≤ 512), and `migronaut up` / `down` / `redo` / `converge` take
|
|
206
|
+
`--reason`. They are stamped on the changelog (`requestedBy` / `reason` on an apply,
|
|
207
|
+
`revertRequestedBy` / `revertReason` on a revert; a later apply that says nothing clears the old
|
|
208
|
+
ones) and on the converge history; `status()` rows show them. `executedBy` stays the OS user —
|
|
209
|
+
on a queue worker, the container's — which is why the requester has fields of its own. A failed
|
|
210
|
+
attempt's trace now also records the checksum of the file version that failed
|
|
211
|
+
(`StatusRow.failedChecksum`).
|
|
212
|
+
- **`runMigrations(config, { signal })`** — an `AbortSignal` (wired to SIGTERM) stops a wait for
|
|
213
|
+
the lock between polls, and a run that holds it between migrations.
|
|
214
|
+
- **`LockInfo.runId` and `LockInfo.ttlMs`** — which run holds the lock (also on `migronaut lock`),
|
|
215
|
+
and the holder's TTL, which the lock document now records.
|
|
216
|
+
- **Three error codes**: `MIGRATION_BLOCKED` (exit 24), `QUEUE_JOB_INVALID` (25),
|
|
217
|
+
`QUEUE_JOB_FAILED` (26), with `MigrationBlockedError`, `QueueJobInvalidError` and
|
|
218
|
+
`QueueJobFailedError` exported from the package root.
|
|
219
|
+
- **Runnable example** — `examples/migration-service`: a queue, a worker and a plain `node:http`
|
|
220
|
+
API (not published to npm).
|
|
221
|
+
- **`generateId` config option** — your own identifier format (ULID, CUID, nanoid, UUIDv7, …)
|
|
222
|
+
everywhere migronaut used to call `crypto.randomUUID()`: the `runId` of every run — on changelog
|
|
223
|
+
records, events, log lines and the lock's owner token — and the `groupId` of every queue
|
|
224
|
+
enqueue. One option covers the kit, the CLI (through `migronaut.config.js`/`.ts`),
|
|
225
|
+
`runMigrations` and the queue adapter.
|
|
226
|
+
- **Injected, like the logger** — migronaut ships no generator but the default. It is called
|
|
227
|
+
with no arguments, so third-party functions pass straight through: `generateId: ulid`.
|
|
228
|
+
- **Checked on every call** — it must synchronously return a non-empty string of at most 128
|
|
229
|
+
characters. A throw, a promise or anything else fails the run with `CONFIG_INVALID` before a
|
|
230
|
+
migration starts (and an enqueue before a job is added).
|
|
231
|
+
- Code-only: no environment variable and no place in a JSON config, like `logger` and `hooks`.
|
|
232
|
+
- **`MigratorKit.generateId()`** — a new id in the kit's configured format, for code that wants its
|
|
233
|
+
own ids to match (it is how the queue adapter mints group ids). Resolves the config; does not
|
|
234
|
+
connect.
|
|
235
|
+
- **`IdGenerator` type** — `() => string`, exported from the package root.
|
|
236
|
+
- **`telemetry` config option** — OpenTelemetry traces and metrics through a tracer and/or a meter
|
|
237
|
+
from your own `@opentelemetry/api`: `telemetry: { tracer, meter }`. One option covers the kit,
|
|
238
|
+
the CLI (through `migronaut.config.js`/`.ts`), `runMigrations` and the queue adapter.
|
|
239
|
+
- **Spans**: `migronaut.run` for every run that acquired the lock, and a child
|
|
240
|
+
`migronaut.migration` for every migration executed. The migration's span is the *active* one
|
|
241
|
+
while its hooks, its `up`/`down` and its changelog write run — so an instrumented MongoDB
|
|
242
|
+
driver nests its command spans under the migration that issued them. That is what lifecycle
|
|
243
|
+
events cannot do, and why this lives in the kit: at application startup there is no ambient
|
|
244
|
+
span, and the driver instrumentation records nothing without a parent.
|
|
245
|
+
- **Metrics**: `migronaut.run.duration`, `migronaut.migration.duration`,
|
|
246
|
+
`migronaut.lock.acquire.duration` and `migronaut.lock.wait.duration` (one point per wait for a
|
|
247
|
+
held lock, by `migronaut.lock.wait.outcome`: `acquired`, `timeout`, `aborted`) — histograms,
|
|
248
|
+
in seconds, with boundaries from 10ms to an hour — and the counters `migronaut.lock.refused`
|
|
249
|
+
and `migronaut.lock.lost`.
|
|
250
|
+
- **Dimensions**: every span and metric point carries `db.namespace` (the database name), plus
|
|
251
|
+
the caller's own static attributes from `telemetry.attributes` (at most 20 scalars).
|
|
252
|
+
- **Failures** set the span's status to `ERROR` with a redacted message (credentials and the
|
|
253
|
+
values a duplicate-key error quotes masked, at most 1 KB), and `error.type` — on
|
|
254
|
+
the span and on the metric point — to the typed error code. No exception event is recorded: it
|
|
255
|
+
would carry the unredacted message and stack.
|
|
256
|
+
- **A run that never got the lock emits no span.** A caller polling for a busy lock retries the
|
|
257
|
+
whole run every few hundred milliseconds; the refusals are counted
|
|
258
|
+
(`migronaut.lock.refused`) instead.
|
|
259
|
+
- **Injected, like the logger** — `@opentelemetry/api` is neither a dependency nor a peer, and
|
|
260
|
+
`src/` never imports it (a test enforces it). The types are structural: `MigronautTracer`,
|
|
261
|
+
`MigronautSpan`, `MigronautMeter`, `MigronautHistogram`, `MigronautCounter`,
|
|
262
|
+
`MigronautMetricOptions`, `MigronautAttributes` and `MigronautTelemetry`, exported from the
|
|
263
|
+
package root.
|
|
264
|
+
- **It can never fail a run.** Every call into the tracer, a span, the meter and an instrument
|
|
265
|
+
is guarded — a promise one of them returns included, so a rejecting SDK cannot surface as an
|
|
266
|
+
unhandled rejection — and a tracer that throws before or after running the work, or runs it
|
|
267
|
+
twice, still gets each migration executed exactly once.
|
|
268
|
+
- Code-only, like `logger` and `generateId`. The span, attribute and metric names are new and
|
|
269
|
+
should be treated as experimental.
|
|
270
|
+
- **`bullmq.telemetry`** — `createMigrationQueue({ bullmq: { Queue, Worker, telemetry } })` hands
|
|
271
|
+
BullMQ's own telemetry object (`new BullMQOtel({ tracerName })` from `bullmq-otel`) to the Queue
|
|
272
|
+
and the Worker it constructs. Together with the kit's `telemetry` it gives one trace from the
|
|
273
|
+
request that enqueued, across Redis, to the MongoDB commands in the worker. Until now the facade
|
|
274
|
+
built its Queue without it, so the enqueuing side of that trace could not be joined.
|
|
275
|
+
`workerOptions.telemetry` and `startWorker({ telemetry })` override it for the worker alone.
|
|
276
|
+
- **OpenTelemetry in the example** — `examples/migration-service` gains `tracing.js` and a Jaeger
|
|
277
|
+
in its `docker-compose.yml`: set `OTEL_EXPORTER_OTLP_ENDPOINT` and every enqueue is one trace.
|
|
278
|
+
- **Declared collections** — indexes and validators declared as an end state, and applied by
|
|
279
|
+
`migronaut converge`, with no migration file per change. For what only ever has a current value
|
|
280
|
+
(which indexes a collection has, which validator guards it); migrations stay the tool for
|
|
281
|
+
changes with an order and a history.
|
|
282
|
+
- **`collections` config option** — an array of definitions, `{ name, indexes?, validator?,
|
|
283
|
+
validationLevel?, validationAction?, prune? }`, each index in the driver's own flat
|
|
284
|
+
`createIndexes` shape. Works in `migronaut.config.{ts,js,json}` (and the JSON Schema) and in
|
|
285
|
+
`new MigratorKit({ collections })`.
|
|
286
|
+
- **`collectionsDir` config option** (`MIGRONAUT_COLLECTIONS_DIR`) — one definition file per
|
|
287
|
+
collection (`.ts`/`.js` default export, or `.json`; the name defaults to the file name).
|
|
288
|
+
Opt-in, combined with `collections`, and loaded only when a converge runs — a broken file never
|
|
289
|
+
blocks `status` or an emergency `down`. A collection declared twice is `CONFIG_INVALID`.
|
|
290
|
+
- **Validated strictly**: an unknown definition key or index option is an error with its path
|
|
291
|
+
(`collections[2].indexes[0].uniqe`) — the driver silently drops an option it does not know, so
|
|
292
|
+
a typo would otherwise build the wrong index and then look in sync forever.
|
|
293
|
+
- **`MigratorKit.converge({ dryRun?, prune?, noLock?, ordered? })`** and **`migronaut converge`**
|
|
294
|
+
(`--dry-run`, `--check`, `--prune`, `--yes`). Stateless: every run reads `listCollections` and
|
|
295
|
+
`listIndexes`, plans, and carries the plan out one step at a time under the migration lock —
|
|
296
|
+
create the collection, set the validator, create indexes, `collMod` a TTL or `hidden` in place,
|
|
297
|
+
rebuild what changed otherwise, drop last. Nothing is recorded.
|
|
298
|
+
- **Safe by default.** An index you did not declare is kept and reported (`keep`), and dropped
|
|
299
|
+
only with `prune` (per definition, or for the definitions that do not decide). An identical
|
|
300
|
+
index under another name is accepted as is rather than rebuilt; a different one covering the
|
|
301
|
+
same key is a conflict that refuses the run before the first write. A rebuild whose new index
|
|
302
|
+
fails to build puts the old one back — and says so, with the reason, when it cannot.
|
|
303
|
+
- **A unique index is never rebuilt unasked.** Dropping it opens a window with no constraint,
|
|
304
|
+
and a duplicate written in that window leaves neither index buildable. Such a rebuild is a
|
|
305
|
+
`conflict` unless `converge({ rebuildUnique: true })` / `--rebuild-unique`; a converge after
|
|
306
|
+
`up` and a queue job never pass it.
|
|
307
|
+
- **Comparisons follow what the server stores** — text indexes in their `_fts` form, collations
|
|
308
|
+
field by field against the expanded spec (`strength`, `caseLevel` and `numericOrdering` at
|
|
309
|
+
their universal defaults when left out), the collection's default collation,
|
|
310
|
+
`{ locale: 'simple' }`, a flag stored as `1`. A compound `Map` key may hold an integer-like
|
|
311
|
+
field only first — the driver reads keys back as plain objects.
|
|
312
|
+
Anything applied that still compares as changed is reported under `unstable` instead of being
|
|
313
|
+
rebuilt on every run.
|
|
314
|
+
- **The CLI plans first** and asks before any drop or rebuild, and before changing the validator
|
|
315
|
+
of a collection that holds data; `--json` refuses such a plan without `--yes` (and applies a
|
|
316
|
+
purely additive one); a plan with a conflict is refused without asking. `--check` exits `28` on
|
|
317
|
+
drift — a CI gate; `--ordered` refuses while a migration is pending.
|
|
318
|
+
- **Results say what changed**: every row that changes, drops or keeps something carries
|
|
319
|
+
`from` and `to` — the live and the declared index or validator, as plain JSON.
|
|
320
|
+
- **History** — every converge that changes something or fails appends an entry to
|
|
321
|
+
`_migronaut_converge` (`convergeLogCollection`, `MIGRONAUT_CONVERGE_LOG_COLLECTION`): when,
|
|
322
|
+
the trigger, the run id, who ran it and where, who asked and why, and every row it touched
|
|
323
|
+
with its `from` / `to`. Read it with `MigratorKit.convergeHistory({ limit })` or
|
|
324
|
+
`migronaut converge --history [--limit n] [--json]`. Best-effort, and a database that never
|
|
325
|
+
converges never gets the collection.
|
|
326
|
+
- **In place where the server can**: making an index unique (MongoDB 7.0+, `collMod`
|
|
327
|
+
`prepareUnique` then `unique` — duplicates leave the index as it was) and adding a TTL to a
|
|
328
|
+
single-field index (5.1+) are `modify`, not a rebuild. (6.0 accepts the unique conversion but
|
|
329
|
+
did not enforce it in our tests, so it rebuilds.)
|
|
330
|
+
- **Sharded clusters**: behind a `mongos`, prune never drops the index backing a shard key (read
|
|
331
|
+
from `config.collections`, or kept with a warning when the server refuses the drop).
|
|
332
|
+
- **Re-planned before each collection**: every collection after the first is read and planned
|
|
333
|
+
again right before its turn; one that changed meanwhile into a conflict or a new drop/rebuild
|
|
334
|
+
stops the run (`ConvergeFailedError`, `phase: 'replan'`) before it is touched.
|
|
335
|
+
- **Regular expressions** in a validator or partial filter must use flags the driver stores as
|
|
336
|
+
written (`i`, `m`; a `BSONRegExp` for server options) — `g` would become dotAll and `s`, `u`,
|
|
337
|
+
`y` vanish — and compare in their stored form. A `__proto__` key in a JSON definition stays a
|
|
338
|
+
key.
|
|
339
|
+
- **At scale**: all declared collections are read with one `listCollections` and a bounded
|
|
340
|
+
fan-out of `listIndexes`; the new indexes of a collection are built by one `createIndexes`
|
|
341
|
+
(one pass over the data); a connection that fails mid-rebuild is reported, never "repaired"
|
|
342
|
+
by restoring the old index next to a build the server may still be running.
|
|
343
|
+
- **`convergeAfterUp` config option** (`MIGRONAUT_CONVERGE_AFTER_UP`) — a bulk `up` (no file, no
|
|
344
|
+
`to`; also `runMigrations`) ends by converging under the same lock, even when nothing was
|
|
345
|
+
pending, so a failed converge is retried by the next deploy. `up(undefined, { converge })` and
|
|
346
|
+
`migronaut up --converge` / `--no-converge` decide per run. `up` still returns its migration
|
|
347
|
+
rows; `runMigrations` adds `summary.converge`.
|
|
348
|
+
- **`MigratorKit.convergesAfterUp()`** — whether a bulk `up` on the kit ends by converging.
|
|
349
|
+
- **Events**: `converge:start`, `converge:action` (per step: `'started'` before it runs — an index
|
|
350
|
+
build can take hours — then `'applied'` or `'failed'`) and `converge:end` (with the full
|
|
351
|
+
result), for real runs. An index build also logs which index it is starting on. The run itself is an ordinary `run:start`/`run:end` with
|
|
352
|
+
`command: 'converge'`, and an ordinary `migronaut.run` span.
|
|
353
|
+
- **Types**: `CollectionDefinition`, `CollectionDefinitionFile`, `IndexDefinition`,
|
|
354
|
+
`IndexKeyDirection`, `IndexCollation`, `ValidationLevel`, `ValidationAction`, `ConvergeOptions`,
|
|
355
|
+
`ConvergeResult`, `CollectionConvergeResult`, `ConvergeAction`, `ConvergeActionKind`,
|
|
356
|
+
`ConvergeActionStatus`, `ConvergeTarget`, `ConvergeUnstable`, `ConvergeTrigger` and the three
|
|
357
|
+
event payloads, exported from the package root.
|
|
358
|
+
- The definition shape, the result shape and the queue contract below are new and should be
|
|
359
|
+
treated as experimental.
|
|
360
|
+
- **Converge jobs in the queue adapter** — `JOB_NAMES.CONVERGE` (`'converge'`), `enqueueConverge()`
|
|
361
|
+
and `MigrationQueue.enqueueConverge()`, and `schedule({ job: 'converge' })` with its own default
|
|
362
|
+
id, `DEFAULT_CONVERGE_SCHEDULER_ID` (`'migronaut-converge'`). With `convergeAfterUp`,
|
|
363
|
+
`enqueueUp` ends a group that reaches the newest migration with a converge job (or adds a
|
|
364
|
+
converge-only job when nothing is pending but a dry run finds drift), and an idle `sync` tick does
|
|
365
|
+
the same. A converge job refuses as `MIGRATION_BLOCKED` while a migration is pending, is keyed for
|
|
366
|
+
deduplication on the migration it follows, and carries no `prune` — what may be dropped comes
|
|
367
|
+
from the worker's own definitions. New types: `ConvergeJobData`, `ConvergeJobResult`,
|
|
368
|
+
`ConvergeJobSpec`, `ConvergeHandle`, `EnqueueConvergeOptions`.
|
|
369
|
+
- **Two exit codes**: `CONVERGE_FAILED` (27, with `ConvergeFailedError` exported from the package
|
|
370
|
+
root) and the CLI-only `COLLECTIONS_DRIFT` (28, from `converge --check`).
|
|
371
|
+
|
|
372
|
+
### Changed
|
|
373
|
+
|
|
374
|
+
- **`MigronautErrorCode` gained four members** (above). TypeScript consumers with an exhaustive
|
|
375
|
+
`switch` over the code union need a `default` branch or the new cases.
|
|
376
|
+
- **`collections`, `collectionsDir` and `convergeAfterUp` config keys are now validated**
|
|
377
|
+
(`CONFIG_INVALID`). They were previously unknown and ignored, like any stray key.
|
|
378
|
+
- **The CLI arg parser makes a `--x` / `--no-x` pair tri-state**, as commander does: when a command
|
|
379
|
+
declares both, neither given leaves the option unset instead of defaulting to `true`. Only
|
|
380
|
+
`up --converge` / `--no-converge` uses this.
|
|
381
|
+
- **`dryRun('up')` now applies the out-of-order policy** of the run it previews: under
|
|
382
|
+
`onOutOfOrder: 'error'` a bulk preview refuses with `MIGRATION_OUT_OF_ORDER` instead of listing
|
|
383
|
+
rows the run would reject; under `'warn'` it logs the warning. A single-file preview is exempt,
|
|
384
|
+
as the single-file run is.
|
|
385
|
+
- **`connect()` is safe to call concurrently** — overlapping first calls on one `MigratorKit`
|
|
386
|
+
share a single connection instead of each opening (and all but one leaking) a client. Matters
|
|
387
|
+
for a long-lived kit serving several callers.
|
|
388
|
+
- The lock-wait loop of `runMigrations` moved to `src/core/lock-wait.js`, shared with the queue
|
|
389
|
+
processor. Two changes to how it waits: **polls back off**, doubling from `lockPollIntervalMs`
|
|
390
|
+
up to 5 s (and at most a quarter of the budget), so a fleet waiting out a long deploy no longer
|
|
391
|
+
hammers the lock document; and **the default `lockWaitTimeoutMs` follows the holder's TTL** —
|
|
392
|
+
`max(90 s, 1.5 × lockTTLSeconds)` — so a holder with a long TTL, whose heartbeat moves the lock
|
|
393
|
+
only every TTL/2, is no longer mistaken for a stalled one. An explicit `lockWaitTimeoutMs` is
|
|
394
|
+
used as given (with a warning when it is shorter than the holder's heartbeat). `waitedMs` is
|
|
395
|
+
now measured by the clock from the first refusal, and a wait that times out rethrows the
|
|
396
|
+
refusal with `context.timedOut`, `attempts` and `waitedMs`.
|
|
397
|
+
- **`runMigrations` validates `onLockHeld`** (`CONFIG_INVALID`): a value other than `'throw'` or
|
|
398
|
+
`'wait'` — `'Wait'`, say — used to behave as `'throw'` without a word. A poll interval above the
|
|
399
|
+
largest timer (2³¹−1 ms, which Node fires after 1 ms) is refused too.
|
|
400
|
+
- **The lock document gained a `nonce` field**, minted by migronaut on every acquire and matched
|
|
401
|
+
alongside `owner` when the lock is confirmed, renewed and released. The owner token is the run
|
|
402
|
+
id, whose format `generateId` now decides; the nonce keeps mutual exclusion independent of it,
|
|
403
|
+
so a generator that repeats an id can blur correlation but never let two runs hold the lock.
|
|
404
|
+
Older releases ignore the field and can share a database with this one. The nonce is never
|
|
405
|
+
exposed; the owner — the holder's run id — now is, as `LockInfo.runId` (`lock`, `lockInfo()`,
|
|
406
|
+
`LockAlreadyHeldError` context), next to the holder's `ttlMs`, also written since this release:
|
|
407
|
+
"which run holds the lock?" is the first question about a stuck one, and the nonce, not the
|
|
408
|
+
owner, is what proves ownership.
|
|
409
|
+
- **A `generateId` config key that is not a function is now rejected** (`CONFIG_INVALID`). The key
|
|
410
|
+
was previously unknown and ignored, like any stray key.
|
|
411
|
+
- **A `telemetry` config key that is not usable is now rejected** (`CONFIG_INVALID`): it must be an
|
|
412
|
+
object whose `tracer` has `startActiveSpan` and whose `meter` has `createHistogram` and
|
|
413
|
+
`createCounter`. Absent, `null` and `{}` all mean "off". The key was previously unknown and
|
|
414
|
+
ignored.
|
|
415
|
+
- **A scheduled `sync` job no longer joins the trace that registered its schedule.** Its template
|
|
416
|
+
now carries `telemetry: { omitContext: true }`: BullMQ builds each scheduler iteration from the
|
|
417
|
+
previous job's options, so with telemetry on, every tick would otherwise have been appended to
|
|
418
|
+
one ever-growing trace. Each tick is a trace of its own; no effect on a queue without telemetry.
|
|
419
|
+
|
|
420
|
+
### Fixed
|
|
421
|
+
|
|
422
|
+
- `docs/reference/cli.md` lists exit code `23` (`MIGRATION_OUT_OF_ORDER`), missing since v2.0.0.
|
|
423
|
+
- The events table in `docs/guide/hooks.md` had fallen behind the payloads: it now lists
|
|
424
|
+
`migration:skipped`, the `command` / `durationMs` / count fields of `run:start` and `run:end`,
|
|
425
|
+
and `ttlMs` / `acquireMs` on `lock:acquired`.
|
|
426
|
+
|
|
427
|
+
### Tooling
|
|
428
|
+
|
|
429
|
+
- A unit test pins `src/utils/id.js` as the only module that mints an identifier, so no id can
|
|
430
|
+
bypass `generateId`.
|
|
431
|
+
- A unit test pins that nothing under `src/` or `bin/`, and neither declaration file, imports an
|
|
432
|
+
`@opentelemetry/*` package or `bullmq-otel`; they are devDependencies only. The structural types
|
|
433
|
+
are checked part by part against the real `@opentelemetry/api`, and one integration test runs the
|
|
434
|
+
real `@opentelemetry/instrumentation-mongodb` to prove the driver's spans nest under a migration.
|
|
435
|
+
- Declared collections are pinned by a table-driven planner test, a fixed-point integration matrix
|
|
436
|
+
(every kind of index converges, then plans as unchanged on a real server), and the shared queue
|
|
437
|
+
scenarios on both the fake and the real BullMQ.
|
|
438
|
+
- An in-tree fake BullMQ carries the adapter's unit and integration tests; the same scenarios run
|
|
439
|
+
against the real `bullmq` package when `MIGRONAUT_TEST_REDIS_URL` is set, which CI now does
|
|
440
|
+
(a Redis service on the `test` job). `bullmq` and `ioredis` are devDependencies for that only.
|
|
441
|
+
|
|
6
442
|
## v2.0.0 — 2026-08-30
|
|
7
443
|
|
|
8
444
|
A major bump for three narrow contract changes (below); everything else is
|