@alexify/migronaut 2.0.0 → 2.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/CHANGELOG.md +436 -0
  2. package/README.md +235 -6
  3. package/bullmq.d.ts +860 -0
  4. package/bullmq.js +1 -0
  5. package/index.d.ts +888 -19
  6. package/migronaut.schema.json +238 -1
  7. package/package.json +21 -5
  8. package/src/bullmq/index.js +55 -0
  9. package/src/bullmq/jobs.js +454 -0
  10. package/src/bullmq/processor.js +632 -0
  11. package/src/bullmq/producer.js +427 -0
  12. package/src/bullmq/service.js +653 -0
  13. package/src/bullmq/wait.js +124 -0
  14. package/src/cli/args.js +12 -2
  15. package/src/cli/commands/converge.js +188 -0
  16. package/src/cli/commands/down.js +2 -0
  17. package/src/cli/commands/lock.js +2 -1
  18. package/src/cli/commands/redo.js +8 -1
  19. package/src/cli/commands/up.js +14 -1
  20. package/src/cli/exit-codes.js +9 -2
  21. package/src/cli/index.js +2 -0
  22. package/src/cli/shared.js +14 -4
  23. package/src/cli/table.js +164 -0
  24. package/src/core/audit.js +88 -3
  25. package/src/core/changelog.js +71 -6
  26. package/src/core/collections.js +396 -0
  27. package/src/core/config.js +130 -25
  28. package/src/core/converge-log.js +47 -0
  29. package/src/core/converge-plan.js +686 -0
  30. package/src/core/converge-search-run.js +440 -0
  31. package/src/core/converge-search.js +404 -0
  32. package/src/core/converge.js +1024 -0
  33. package/src/core/index-spec.js +507 -0
  34. package/src/core/lock-wait.js +260 -0
  35. package/src/core/lock.js +95 -28
  36. package/src/core/migrator.js +600 -287
  37. package/src/core/options.js +266 -0
  38. package/src/core/run-recorder.js +157 -0
  39. package/src/core/run.js +58 -90
  40. package/src/core/search-index-spec.js +758 -0
  41. package/src/core/sequence.js +134 -0
  42. package/src/core/server-info.js +63 -0
  43. package/src/errors/index.js +60 -0
  44. package/src/index.js +8 -0
  45. package/src/utils/actor.js +48 -0
  46. package/src/utils/canonical.js +212 -0
  47. package/src/utils/collection-name.js +21 -0
  48. package/src/utils/error.js +18 -1
  49. package/src/utils/id.js +77 -0
  50. package/src/utils/loader.js +39 -21
  51. package/src/utils/migration-name.js +32 -0
  52. package/src/utils/redact.js +21 -1
  53. package/src/utils/telemetry.js +410 -0
  54. package/src/utils/template.js +43 -2
package/CHANGELOG.md CHANGED
@@ -3,6 +3,442 @@
3
3
  All notable changes to this project will be documented in this file.
4
4
  Release headings carry the publish date (`## vX.Y.Z — YYYY-MM-DD`).
5
5
 
6
+ ## v2.2.0 — 2026-10-05
7
+
8
+ Atlas Search and Vector Search indexes in declared collections. Additive: a definition without
9
+ `searchIndexes` behaves exactly as in 2.1 — converge never even asks the server about Search for
10
+ it.
11
+
12
+ ### Added
13
+
14
+ - **Declared search indexes** — a collection definition takes `searchIndexes: [{ name?, type?,
15
+ definition }]`: Atlas Search (`type: 'search'`, the default) and Atlas Vector Search
16
+ (`'vectorSearch'`) indexes, with the definition exactly as Atlas documents it — automated
17
+ embedding (`autoEmbed`) fields included. `converge` keeps them in step with the same lock,
18
+ history, events, re-plan and fixed-point check as regular indexes. Experimental.
19
+ - **Rows**: a new target, `searchIndex`. A missing index is created (one `createSearchIndexes`
20
+ per collection), a changed definition is updated **in place** (`modify` — the old definition
21
+ serves until the new one is built), an undeclared one is kept, or dropped last under
22
+ `prune`. A definition that declares only `searchIndexes` is valid, and its collection is
23
+ created when missing.
24
+ - **Never a rebuild**: a `$search` against a missing index returns nothing rather than fail, so
25
+ converge never drops a search index to build it again. A change no update can make — the
26
+ type, or an `autoEmbed` field's path, model, `numDimensions`, quantization or modality — is a
27
+ `conflict` naming the new-name recipe, and so is a declared index the server is still
28
+ deleting.
29
+ - **Comparison**: definitions compare whole, key order ignored, with the defaults the server
30
+ writes into what it reports filled in on both sides — top-level (`analyzer`,
31
+ `searchAnalyzer`, `dynamic`, `storedSource`, `numPartitions`), per field mapping (`string`,
32
+ `number`, `autocomplete`, `token`, `geo`, `document`, nested fields and `multi` included) and
33
+ per vector field (`quantization`, `indexingMethod`, `hnswOptions`; `autoEmbed`
34
+ `numDimensions` and `quantization`). Vector `fields`, and a field indexed as several types,
35
+ compare as sets. A server that reports no `type` (a self-managed `mongot`) has it inferred
36
+ from the definition, and its `latestVersion` is read as the definition version.
37
+ - **Server-only options**: an option the server reports that the declaration does not set,
38
+ and whose default migronaut does not know (a newer `mongot`'s), is left out of the
39
+ comparison — named on the row (`ignored`) and in one warning — instead of making every
40
+ converge update, and the server rebuild, the index. Only option objects are trimmed: a
41
+ field, a mapping type or a vector field only the server has is still a difference. The
42
+ cost: removing such an option from a declaration goes unnoticed; declare the value wanted.
43
+ - **What differs**: a `modify` row names it down to the option
44
+ (`mappings.fields.title.norms`) — the first five paths and how many more. A search index
45
+ that still differs after an update is `unstable`, with a warning that every update builds it
46
+ again.
47
+ - **Raw commands** (`createSearchIndexes`, `updateSearchIndex`, `dropSearchIndex`,
48
+ `$listSearchIndexes`), so every driver in the peer range works — the driver's helpers start
49
+ at 5.6. A vector index update is retried once with its type when a self-managed `mongot`
50
+ asks for it; where the server refuses that too (an Atlas CLI local deployment on 8.0 and
51
+ 8.3), the step fails with the new-name recipe as its hint.
52
+ - **Build state**: rows carry `build` (`{ status, queryable, message?, updating? }`), and the
53
+ result `search` (`{ available, evidence, notReady, wait? }`) — how Search availability was
54
+ told, the declared indexes still building, updating, stale or failed, and how a wait ended.
55
+ A build under way does not count against `inSync`; a FAILED one fails `converge --check`
56
+ (exit 28). A STALE index (queryable, no longer replicating) is reported as such — table,
57
+ closing warning, `audit` — not as still building. A build message is kept to 500 characters.
58
+ - **`waitForSearchIndexes`** (config, `MIGRONAUT_WAIT_FOR_SEARCH_INDEXES`, `converge({
59
+ waitForSearchIndexes })`, `converge --wait-search` / `--no-wait-search`) — hold the converge
60
+ until every declared search index serves its declaration; a FAILED build of an index the run
61
+ created or changed, or `searchIndexWaitTimeoutMs` (`MIGRONAUT_SEARCH_INDEX_WAIT_TIMEOUT_MS`,
62
+ default 10 minutes), fails it with `ConvergeFailedError` `phase: 'wait'`. Off by default.
63
+ - **Without the lock**: the wait only reads, so the migration lock is released when it starts
64
+ (`lock:released` with `early: true`) — the next deploy or a queue's next job need not wait
65
+ out a build. The run itself (its id, span, history entry, `converge:end`) ends with the
66
+ wait.
67
+ - **Only what the run started holds it**: an index that failed or went STALE before the run,
68
+ its definition unchanged, is named and warned about instead — converge cannot fix it, and a
69
+ deploy is not held up by it. `--check` still fails on a FAILED build.
70
+ - **Resilient**: a list that fails with a network or failover blip is read again at the next
71
+ poll, up to three in a row (then `reason: 'unreadable'`); a stop cuts the pause between polls
72
+ short.
73
+ - **Observable**: `converge:wait` events (`started`, `progress` every 30 s, then `ready`,
74
+ `failed`, `timeout`, `unreadable` or `aborted`), `result.search.wait`, the history entry, a
75
+ queue job's log lines and `search-wait` progress phase, and — with `telemetry` — the
76
+ histogram `migronaut.converge.search.wait.duration` by `migronaut.converge.search.wait.outcome`.
77
+ - **`onSearchUnavailable`** (`MIGRONAUT_ON_SEARCH_UNAVAILABLE`) — declared search indexes on a
78
+ server without Atlas Search: `'fail'` (default) refuses the run before any write, with a hint;
79
+ `'skip'` converges everything else and reports `skip` rows. Detected once per run from the
80
+ server's own answer (checked against MongoDB 5.0, 6.0, 7.0, 8.0 and 8.2).
81
+ - **`migronaut audit`** — a `search` check, when the definitions declare search indexes: whether
82
+ the server has Atlas Search, and whether a declared index failed to build.
83
+ - **Types** — `SearchIndexDefinition`, `SearchDefinition`, `VectorSearchDefinition`,
84
+ `VectorSearchField`, `SearchIndexType`, `SearchIndexStatus`, `SearchIndexBuild`,
85
+ `SearchIndexNotReady`, `ConvergeSearchSummary`, `ConvergeWaitOutcome`, `ConvergeWaitEvent`;
86
+ `ConvergeTarget` gains `'searchIndex'`, `ConvergeActionKind` `'skip'`, `ConvergeAction`
87
+ `ignored`, `ConvergeHistoryEntry` `search`, `LockEvent` `early`, the event map
88
+ `'converge:wait'`; in `bullmq.d.ts` `ConvergeJobResult.search`, and `MigrationJobProgress`
89
+ gains the `'search-wait'` phase and `searchIndexes`. An exhaustive `switch` over one of these
90
+ unions needs a branch for the new member.
91
+ - **Queue** — a converge job's result carries `search`, and its log lines name search indexes.
92
+ Whether a job waits for builds is the worker kit's `waitForSearchIndexes`; the job payload is
93
+ unchanged.
94
+
95
+ ### Changed
96
+
97
+ - A definition with none of `indexes`, `searchIndexes` and `validator` is refused with "declares
98
+ no indexes, searchIndexes or validator — nothing to manage" (was "declares neither indexes nor
99
+ a validator").
100
+ - The converge table's drop/rebuild count includes dropped search indexes, and `converge --yes`
101
+ is required for them in `--json` mode, like an index drop.
102
+ - `ConvergeFailedError`'s documentation names every phase: `plan`, `replan` (already thrown by
103
+ 2.1, undocumented), `apply` and the new `wait`. A search index list that cannot be read is
104
+ reported in the phase that read it — `plan` only before the first write — and a run that stops
105
+ for any reason settles every row and carries the result so far in `context.converge`.
106
+ - `runWithLock` (internal) hands the work a `control` whose `release()` gives the lock up early;
107
+ `lock:released` may now come before `run:end`, with `early: true`.
108
+
109
+ ### Tooling
110
+
111
+ - `tests/integration/search-atlas.test.js` — an opt-in, manual suite against
112
+ `mongodb/mongodb-atlas-local` (`MIGRONAUT_TEST_ATLAS_URI`): every scenario ends at a fixed
113
+ point. Passes against `mongodb/mongodb-atlas-local` 8.0 (8.0.32) and `latest` (8.3.11). CI does
114
+ not run it; the coverage gate comes from the unit tier's fake, which answers the search
115
+ commands with lag, normalization and build progress on demand. The suite stays strict about
116
+ server-only options (production tolerates them): one there is a default the tables in
117
+ `search-index-spec.js` lack. 9/9 on 8.0.32 and 8.3.11, the lock released during the wait
118
+ included.
119
+ - `src/core/converge.js` is split: `converge-search-run.js` (the search half of a run) and
120
+ `server-info.js` (read options, read pace, server version, not-found codes).
121
+
122
+ ## v2.1.0 — 2026-10-04
123
+
124
+ Migrations as a queue, ids in your own format, OpenTelemetry, and declared collections. Additive:
125
+ nothing changes for anyone who uses none of them, with the narrow exceptions listed under
126
+ **Changed**.
127
+
128
+ ### Added
129
+
130
+ - **`@alexify/migronaut/bullmq`** — a new entry point that runs migrations as
131
+ [BullMQ](https://docs.bullmq.io/) jobs, **one migration per job**, so migronaut can be a
132
+ migration service: enqueue from an HTTP handler, a schedule or a deploy hook, and let a worker
133
+ apply them in order.
134
+ - `createMigrationQueue(options)` — the facade: `enqueueUp` / `enqueueDown` (returning a group
135
+ handle with `wait()`), `startWorker`, `status` / `pending` / `audit` / `lockInfo`, `getJob`,
136
+ `pause` / `resume`, `schedule` / `unschedule`, `close`.
137
+ - `createMigrationProcessor(options)` — the processor on its own, for a Worker you construct
138
+ (NestJS, BullMQ Pro), plus `enqueueUp` / `enqueueDown` / `planUpJobs` / `planDownJobs` /
139
+ `waitForGroup` for a Queue you own.
140
+ - **BullMQ is injected, never depended on** — `bullmq: { Queue, Worker, QueueEvents }` from
141
+ your own install. The package still has no `dependencies` and gains no peer; `src/` never
142
+ imports `bullmq` (a test enforces it).
143
+ - **Order comes from MongoDB, not from Redis.** Every job is a single-file run under the usual
144
+ lock; it refuses while an earlier migration is still pending, so a failed migration stops the
145
+ line (`MIGRATION_BLOCKED` for the jobs behind it). Jobs get one attempt on purpose — a BullMQ
146
+ retry re-queues behind the waiting jobs — and a held lock is waited out inside the job.
147
+ - **One batch per enqueue**, so `down` still rolls back a whole deploy; duplicate enqueues are
148
+ deduplicated, and a job whose migration is already applied completes as `skipped`.
149
+ - Job payloads are validated as untrusted input; messages, stacks and job logs are redacted. A
150
+ queue job's target must be a file of the migration sequence — a payload can never make the
151
+ worker import a dotfile, a declaration file or a helper module next to the migrations.
152
+ - **Shutdown puts unstarted work back.** A job that a closing worker stops before its migration
153
+ starts (waiting for the lock, or fetched during shutdown) is moved back to the head of the
154
+ queue instead of failing — so a rolling deploy no longer fails the rest of the enqueue it
155
+ interrupts as `MIGRATION_BLOCKED`. A migration already running always finishes.
156
+ - **A versioned, strict job contract.** A worker accepts every job data version from
157
+ `MIN_JOB_DATA_VERSION` (exported) up to its own and refuses newer ones and unknown fields
158
+ rather than ignoring what they mean — roll workers out before producers. Jobs always state
159
+ `ordered`, and may carry the plan-time `checksum` of their file.
160
+ - **Scheduler ticks are bounded**: they carry the queue's `jobOptions`, and keep the last 100
161
+ completed / 500 failed jobs when those set no retention.
162
+ - **Correlation**: a job's `runId` is on its progress (`completed` and `failed`), its failure log
163
+ row and its error's `context` (with `jobId` and `groupId`) — a failed job has no return value;
164
+ the worker's failure log line names the group, migration and run.
165
+ - An injected `Queue` (or `QueueEvents`) on another name or prefix than the facade's is rejected,
166
+ and the facade takes an injected queue's prefix by default; `startWorker()` can be retried
167
+ after a failed start; a closed queue refuses every further call.
168
+ - **A worker decides what a job may ask for** — the `allow` option (`{ down: true, force:
169
+ false, unordered: false }` by default) on `createMigrationProcessor` and
170
+ `createMigrationQueue`: a payload asking to re-run an applied migration or to skip the order
171
+ guard is refused (`QUEUE_JOB_INVALID`, `context.permission`) unless allowed. The facade's
172
+ `enqueue*` calls follow the same policy.
173
+ - **A job runs only the file it was planned with**: an `up` job carries the plan-time checksum,
174
+ and a worker with another version of the file fails it (`CHECKSUM_MISMATCH`,
175
+ `context.planned`) instead of applying it.
176
+ - **Several workers without global concurrency**: a job blocked only by earlier migrations that
177
+ have not failed (one may be in flight on another worker) waits for them within its lock-wait
178
+ budget instead of failing its group; `MigrationBlockedError` carries `context.failed`.
179
+ - `wait()` holds every await to its one `timeoutMs` budget, decides a timeout by the clock, and
180
+ its `QueueJobFailedError` carries the job's own typed `code`; an empty group can be waited for
181
+ without QueueEvents. A long lock wait reports progress every few seconds, not every poll.
182
+ - What the queue stores about a failure — `failedReason`, stack, job logs — masks the values a
183
+ duplicate-key error quotes, on top of credentials. `schedule({ every })` needs at least 1000 ms.
184
+ - Dedup ids encode file names reversibly (two names can no longer share one and absorb each
185
+ other's job), and a forced re-run has a dedup id of its own.
186
+ - **A schedule holds a failed migration**: a `sync` tick whose next migration failed, with its
187
+ file unchanged since, enqueues nothing (`returnvalue.held`, and a warning) instead of re-running
188
+ it every tick; a changed file or an explicit `enqueueUp(name)` resumes it.
189
+ - Jobs carry `requestedBy` / `reason` (`enqueueUp` / `enqueueDown` / `enqueueConverge` options).
190
+ - **`bullmq.d.ts`** — hand-written types for the entry point, with structural `BullMQ*Like`
191
+ interfaces instead of an import of `bullmq`, generic over the classes you inject.
192
+ - **`up(file, { batch })`** — stamp an explicit batch number instead of the next free one, and
193
+ **`MigratorKit.nextBatch()`** to peek at it: together they let several single-file runs form one
194
+ batch.
195
+ - **`up(file, { ordered: true })` / `down(file, { ordered: true })`** — refuse a single-file run
196
+ that would go out of sequence (`MigrationBlockedError`); an ordered `up` also applies the
197
+ `strict` drift check and the `onOutOfOrder` policy a bulk run would, and only targets a file of
198
+ the migration sequence.
199
+ - **`up(file, { checksum })`** — refuse (`CHECKSUM_MISMATCH`, `context.planned`) to apply any other
200
+ version of the file than the one with this SHA-256; `dryRun('up')` rows carry each file's
201
+ `checksum` for it.
202
+ - **`list(filter, { checksums: false })`** / **`status({ checksums: false })`** — skip hashing the
203
+ applied files, for a caller that needs names and dates only.
204
+ - **Who asked, and why** — `up`, `down`, `redo` and `converge` take `requestedBy` (≤ 128
205
+ characters) and `reason` (≤ 512), and `migronaut up` / `down` / `redo` / `converge` take
206
+ `--reason`. They are stamped on the changelog (`requestedBy` / `reason` on an apply,
207
+ `revertRequestedBy` / `revertReason` on a revert; a later apply that says nothing clears the old
208
+ ones) and on the converge history; `status()` rows show them. `executedBy` stays the OS user —
209
+ on a queue worker, the container's — which is why the requester has fields of its own. A failed
210
+ attempt's trace now also records the checksum of the file version that failed
211
+ (`StatusRow.failedChecksum`).
212
+ - **`runMigrations(config, { signal })`** — an `AbortSignal` (wired to SIGTERM) stops a wait for
213
+ the lock between polls, and a run that holds it between migrations.
214
+ - **`LockInfo.runId` and `LockInfo.ttlMs`** — which run holds the lock (also on `migronaut lock`),
215
+ and the holder's TTL, which the lock document now records.
216
+ - **Three error codes**: `MIGRATION_BLOCKED` (exit 24), `QUEUE_JOB_INVALID` (25),
217
+ `QUEUE_JOB_FAILED` (26), with `MigrationBlockedError`, `QueueJobInvalidError` and
218
+ `QueueJobFailedError` exported from the package root.
219
+ - **Runnable example** — `examples/migration-service`: a queue, a worker and a plain `node:http`
220
+ API (not published to npm).
221
+ - **`generateId` config option** — your own identifier format (ULID, CUID, nanoid, UUIDv7, …)
222
+ everywhere migronaut used to call `crypto.randomUUID()`: the `runId` of every run — on changelog
223
+ records, events, log lines and the lock's owner token — and the `groupId` of every queue
224
+ enqueue. One option covers the kit, the CLI (through `migronaut.config.js`/`.ts`),
225
+ `runMigrations` and the queue adapter.
226
+ - **Injected, like the logger** — migronaut ships no generator but the default. It is called
227
+ with no arguments, so third-party functions pass straight through: `generateId: ulid`.
228
+ - **Checked on every call** — it must synchronously return a non-empty string of at most 128
229
+ characters. A throw, a promise or anything else fails the run with `CONFIG_INVALID` before a
230
+ migration starts (and an enqueue before a job is added).
231
+ - Code-only: no environment variable and no place in a JSON config, like `logger` and `hooks`.
232
+ - **`MigratorKit.generateId()`** — a new id in the kit's configured format, for code that wants its
233
+ own ids to match (it is how the queue adapter mints group ids). Resolves the config; does not
234
+ connect.
235
+ - **`IdGenerator` type** — `() => string`, exported from the package root.
236
+ - **`telemetry` config option** — OpenTelemetry traces and metrics through a tracer and/or a meter
237
+ from your own `@opentelemetry/api`: `telemetry: { tracer, meter }`. One option covers the kit,
238
+ the CLI (through `migronaut.config.js`/`.ts`), `runMigrations` and the queue adapter.
239
+ - **Spans**: `migronaut.run` for every run that acquired the lock, and a child
240
+ `migronaut.migration` for every migration executed. The migration's span is the *active* one
241
+ while its hooks, its `up`/`down` and its changelog write run — so an instrumented MongoDB
242
+ driver nests its command spans under the migration that issued them. That is what lifecycle
243
+ events cannot do, and why this lives in the kit: at application startup there is no ambient
244
+ span, and the driver instrumentation records nothing without a parent.
245
+ - **Metrics**: `migronaut.run.duration`, `migronaut.migration.duration`,
246
+ `migronaut.lock.acquire.duration` and `migronaut.lock.wait.duration` (one point per wait for a
247
+ held lock, by `migronaut.lock.wait.outcome`: `acquired`, `timeout`, `aborted`) — histograms,
248
+ in seconds, with boundaries from 10ms to an hour — and the counters `migronaut.lock.refused`
249
+ and `migronaut.lock.lost`.
250
+ - **Dimensions**: every span and metric point carries `db.namespace` (the database name), plus
251
+ the caller's own static attributes from `telemetry.attributes` (at most 20 scalars).
252
+ - **Failures** set the span's status to `ERROR` with a redacted message (credentials and the
253
+ values a duplicate-key error quotes masked, at most 1 KB), and `error.type` — on
254
+ the span and on the metric point — to the typed error code. No exception event is recorded: it
255
+ would carry the unredacted message and stack.
256
+ - **A run that never got the lock emits no span.** A caller polling for a busy lock retries the
257
+ whole run every few hundred milliseconds; the refusals are counted
258
+ (`migronaut.lock.refused`) instead.
259
+ - **Injected, like the logger** — `@opentelemetry/api` is neither a dependency nor a peer, and
260
+ `src/` never imports it (a test enforces it). The types are structural: `MigronautTracer`,
261
+ `MigronautSpan`, `MigronautMeter`, `MigronautHistogram`, `MigronautCounter`,
262
+ `MigronautMetricOptions`, `MigronautAttributes` and `MigronautTelemetry`, exported from the
263
+ package root.
264
+ - **It can never fail a run.** Every call into the tracer, a span, the meter and an instrument
265
+ is guarded — a promise one of them returns included, so a rejecting SDK cannot surface as an
266
+ unhandled rejection — and a tracer that throws before or after running the work, or runs it
267
+ twice, still gets each migration executed exactly once.
268
+ - Code-only, like `logger` and `generateId`. The span, attribute and metric names are new and
269
+ should be treated as experimental.
270
+ - **`bullmq.telemetry`** — `createMigrationQueue({ bullmq: { Queue, Worker, telemetry } })` hands
271
+ BullMQ's own telemetry object (`new BullMQOtel({ tracerName })` from `bullmq-otel`) to the Queue
272
+ and the Worker it constructs. Together with the kit's `telemetry` it gives one trace from the
273
+ request that enqueued, across Redis, to the MongoDB commands in the worker. Until now the facade
274
+ built its Queue without it, so the enqueuing side of that trace could not be joined.
275
+ `workerOptions.telemetry` and `startWorker({ telemetry })` override it for the worker alone.
276
+ - **OpenTelemetry in the example** — `examples/migration-service` gains `tracing.js` and a Jaeger
277
+ in its `docker-compose.yml`: set `OTEL_EXPORTER_OTLP_ENDPOINT` and every enqueue is one trace.
278
+ - **Declared collections** — indexes and validators declared as an end state, and applied by
279
+ `migronaut converge`, with no migration file per change. For what only ever has a current value
280
+ (which indexes a collection has, which validator guards it); migrations stay the tool for
281
+ changes with an order and a history.
282
+ - **`collections` config option** — an array of definitions, `{ name, indexes?, validator?,
283
+ validationLevel?, validationAction?, prune? }`, each index in the driver's own flat
284
+ `createIndexes` shape. Works in `migronaut.config.{ts,js,json}` (and the JSON Schema) and in
285
+ `new MigratorKit({ collections })`.
286
+ - **`collectionsDir` config option** (`MIGRONAUT_COLLECTIONS_DIR`) — one definition file per
287
+ collection (`.ts`/`.js` default export, or `.json`; the name defaults to the file name).
288
+ Opt-in, combined with `collections`, and loaded only when a converge runs — a broken file never
289
+ blocks `status` or an emergency `down`. A collection declared twice is `CONFIG_INVALID`.
290
+ - **Validated strictly**: an unknown definition key or index option is an error with its path
291
+ (`collections[2].indexes[0].uniqe`) — the driver silently drops an option it does not know, so
292
+ a typo would otherwise build the wrong index and then look in sync forever.
293
+ - **`MigratorKit.converge({ dryRun?, prune?, noLock?, ordered? })`** and **`migronaut converge`**
294
+ (`--dry-run`, `--check`, `--prune`, `--yes`). Stateless: every run reads `listCollections` and
295
+ `listIndexes`, plans, and carries the plan out one step at a time under the migration lock —
296
+ create the collection, set the validator, create indexes, `collMod` a TTL or `hidden` in place,
297
+ rebuild what changed otherwise, drop last. Nothing is recorded.
298
+ - **Safe by default.** An index you did not declare is kept and reported (`keep`), and dropped
299
+ only with `prune` (per definition, or for the definitions that do not decide). An identical
300
+ index under another name is accepted as is rather than rebuilt; a different one covering the
301
+ same key is a conflict that refuses the run before the first write. A rebuild whose new index
302
+ fails to build puts the old one back — and says so, with the reason, when it cannot.
303
+ - **A unique index is never rebuilt unasked.** Dropping it opens a window with no constraint,
304
+ and a duplicate written in that window leaves neither index buildable. Such a rebuild is a
305
+ `conflict` unless `converge({ rebuildUnique: true })` / `--rebuild-unique`; a converge after
306
+ `up` and a queue job never pass it.
307
+ - **Comparisons follow what the server stores** — text indexes in their `_fts` form, collations
308
+ field by field against the expanded spec (`strength`, `caseLevel` and `numericOrdering` at
309
+ their universal defaults when left out), the collection's default collation,
310
+ `{ locale: 'simple' }`, a flag stored as `1`. A compound `Map` key may hold an integer-like
311
+ field only first — the driver reads keys back as plain objects.
312
+ Anything applied that still compares as changed is reported under `unstable` instead of being
313
+ rebuilt on every run.
314
+ - **The CLI plans first** and asks before any drop or rebuild, and before changing the validator
315
+ of a collection that holds data; `--json` refuses such a plan without `--yes` (and applies a
316
+ purely additive one); a plan with a conflict is refused without asking. `--check` exits `28` on
317
+ drift — a CI gate; `--ordered` refuses while a migration is pending.
318
+ - **Results say what changed**: every row that changes, drops or keeps something carries
319
+ `from` and `to` — the live and the declared index or validator, as plain JSON.
320
+ - **History** — every converge that changes something or fails appends an entry to
321
+ `_migronaut_converge` (`convergeLogCollection`, `MIGRONAUT_CONVERGE_LOG_COLLECTION`): when,
322
+ the trigger, the run id, who ran it and where, who asked and why, and every row it touched
323
+ with its `from` / `to`. Read it with `MigratorKit.convergeHistory({ limit })` or
324
+ `migronaut converge --history [--limit n] [--json]`. Best-effort, and a database that never
325
+ converges never gets the collection.
326
+ - **In place where the server can**: making an index unique (MongoDB 7.0+, `collMod`
327
+ `prepareUnique` then `unique` — duplicates leave the index as it was) and adding a TTL to a
328
+ single-field index (5.1+) are `modify`, not a rebuild. (6.0 accepts the unique conversion but
329
+ did not enforce it in our tests, so it rebuilds.)
330
+ - **Sharded clusters**: behind a `mongos`, prune never drops the index backing a shard key (read
331
+ from `config.collections`, or kept with a warning when the server refuses the drop).
332
+ - **Re-planned before each collection**: every collection after the first is read and planned
333
+ again right before its turn; one that changed meanwhile into a conflict or a new drop/rebuild
334
+ stops the run (`ConvergeFailedError`, `phase: 'replan'`) before it is touched.
335
+ - **Regular expressions** in a validator or partial filter must use flags the driver stores as
336
+ written (`i`, `m`; a `BSONRegExp` for server options) — `g` would become dotAll and `s`, `u`,
337
+ `y` vanish — and compare in their stored form. A `__proto__` key in a JSON definition stays a
338
+ key.
339
+ - **At scale**: all declared collections are read with one `listCollections` and a bounded
340
+ fan-out of `listIndexes`; the new indexes of a collection are built by one `createIndexes`
341
+ (one pass over the data); a connection that fails mid-rebuild is reported, never "repaired"
342
+ by restoring the old index next to a build the server may still be running.
343
+ - **`convergeAfterUp` config option** (`MIGRONAUT_CONVERGE_AFTER_UP`) — a bulk `up` (no file, no
344
+ `to`; also `runMigrations`) ends by converging under the same lock, even when nothing was
345
+ pending, so a failed converge is retried by the next deploy. `up(undefined, { converge })` and
346
+ `migronaut up --converge` / `--no-converge` decide per run. `up` still returns its migration
347
+ rows; `runMigrations` adds `summary.converge`.
348
+ - **`MigratorKit.convergesAfterUp()`** — whether a bulk `up` on the kit ends by converging.
349
+ - **Events**: `converge:start`, `converge:action` (per step: `'started'` before it runs — an index
350
+ build can take hours — then `'applied'` or `'failed'`) and `converge:end` (with the full
351
+ result), for real runs. An index build also logs which index it is starting on. The run itself is an ordinary `run:start`/`run:end` with
352
+ `command: 'converge'`, and an ordinary `migronaut.run` span.
353
+ - **Types**: `CollectionDefinition`, `CollectionDefinitionFile`, `IndexDefinition`,
354
+ `IndexKeyDirection`, `IndexCollation`, `ValidationLevel`, `ValidationAction`, `ConvergeOptions`,
355
+ `ConvergeResult`, `CollectionConvergeResult`, `ConvergeAction`, `ConvergeActionKind`,
356
+ `ConvergeActionStatus`, `ConvergeTarget`, `ConvergeUnstable`, `ConvergeTrigger` and the three
357
+ event payloads, exported from the package root.
358
+ - The definition shape, the result shape and the queue contract below are new and should be
359
+ treated as experimental.
360
+ - **Converge jobs in the queue adapter** — `JOB_NAMES.CONVERGE` (`'converge'`), `enqueueConverge()`
361
+ and `MigrationQueue.enqueueConverge()`, and `schedule({ job: 'converge' })` with its own default
362
+ id, `DEFAULT_CONVERGE_SCHEDULER_ID` (`'migronaut-converge'`). With `convergeAfterUp`,
363
+ `enqueueUp` ends a group that reaches the newest migration with a converge job (or adds a
364
+ converge-only job when nothing is pending but a dry run finds drift), and an idle `sync` tick does
365
+ the same. A converge job refuses as `MIGRATION_BLOCKED` while a migration is pending, is keyed for
366
+ deduplication on the migration it follows, and carries no `prune` — what may be dropped comes
367
+ from the worker's own definitions. New types: `ConvergeJobData`, `ConvergeJobResult`,
368
+ `ConvergeJobSpec`, `ConvergeHandle`, `EnqueueConvergeOptions`.
369
+ - **Two exit codes**: `CONVERGE_FAILED` (27, with `ConvergeFailedError` exported from the package
370
+ root) and the CLI-only `COLLECTIONS_DRIFT` (28, from `converge --check`).
371
+
372
+ ### Changed
373
+
374
+ - **`MigronautErrorCode` gained four members** (above). TypeScript consumers with an exhaustive
375
+ `switch` over the code union need a `default` branch or the new cases.
376
+ - **`collections`, `collectionsDir` and `convergeAfterUp` config keys are now validated**
377
+ (`CONFIG_INVALID`). They were previously unknown and ignored, like any stray key.
378
+ - **The CLI arg parser makes a `--x` / `--no-x` pair tri-state**, as commander does: when a command
379
+ declares both, neither given leaves the option unset instead of defaulting to `true`. Only
380
+ `up --converge` / `--no-converge` uses this.
381
+ - **`dryRun('up')` now applies the out-of-order policy** of the run it previews: under
382
+ `onOutOfOrder: 'error'` a bulk preview refuses with `MIGRATION_OUT_OF_ORDER` instead of listing
383
+ rows the run would reject; under `'warn'` it logs the warning. A single-file preview is exempt,
384
+ as the single-file run is.
385
+ - **`connect()` is safe to call concurrently** — overlapping first calls on one `MigratorKit`
386
+ share a single connection instead of each opening (and all but one leaking) a client. Matters
387
+ for a long-lived kit serving several callers.
388
+ - The lock-wait loop of `runMigrations` moved to `src/core/lock-wait.js`, shared with the queue
389
+ processor. Two changes to how it waits: **polls back off**, doubling from `lockPollIntervalMs`
390
+ up to 5 s (and at most a quarter of the budget), so a fleet waiting out a long deploy no longer
391
+ hammers the lock document; and **the default `lockWaitTimeoutMs` follows the holder's TTL** —
392
+ `max(90 s, 1.5 × lockTTLSeconds)` — so a holder with a long TTL, whose heartbeat moves the lock
393
+ only every TTL/2, is no longer mistaken for a stalled one. An explicit `lockWaitTimeoutMs` is
394
+ used as given (with a warning when it is shorter than the holder's heartbeat). `waitedMs` is
395
+ now measured by the clock from the first refusal, and a wait that times out rethrows the
396
+ refusal with `context.timedOut`, `attempts` and `waitedMs`.
397
+ - **`runMigrations` validates `onLockHeld`** (`CONFIG_INVALID`): a value other than `'throw'` or
398
+ `'wait'` — `'Wait'`, say — used to behave as `'throw'` without a word. A poll interval above the
399
+ largest timer (2³¹−1 ms, which Node fires after 1 ms) is refused too.
400
+ - **The lock document gained a `nonce` field**, minted by migronaut on every acquire and matched
401
+ alongside `owner` when the lock is confirmed, renewed and released. The owner token is the run
402
+ id, whose format `generateId` now decides; the nonce keeps mutual exclusion independent of it,
403
+ so a generator that repeats an id can blur correlation but never let two runs hold the lock.
404
+ Older releases ignore the field and can share a database with this one. The nonce is never
405
+ exposed; the owner — the holder's run id — now is, as `LockInfo.runId` (`lock`, `lockInfo()`,
406
+ `LockAlreadyHeldError` context), next to the holder's `ttlMs`, also written since this release:
407
+ "which run holds the lock?" is the first question about a stuck one, and the nonce, not the
408
+ owner, is what proves ownership.
409
+ - **A `generateId` config key that is not a function is now rejected** (`CONFIG_INVALID`). The key
410
+ was previously unknown and ignored, like any stray key.
411
+ - **A `telemetry` config key that is not usable is now rejected** (`CONFIG_INVALID`): it must be an
412
+ object whose `tracer` has `startActiveSpan` and whose `meter` has `createHistogram` and
413
+ `createCounter`. Absent, `null` and `{}` all mean "off". The key was previously unknown and
414
+ ignored.
415
+ - **A scheduled `sync` job no longer joins the trace that registered its schedule.** Its template
416
+ now carries `telemetry: { omitContext: true }`: BullMQ builds each scheduler iteration from the
417
+ previous job's options, so with telemetry on, every tick would otherwise have been appended to
418
+ one ever-growing trace. Each tick is a trace of its own; no effect on a queue without telemetry.
419
+
420
+ ### Fixed
421
+
422
+ - `docs/reference/cli.md` lists exit code `23` (`MIGRATION_OUT_OF_ORDER`), missing since v2.0.0.
423
+ - The events table in `docs/guide/hooks.md` had fallen behind the payloads: it now lists
424
+ `migration:skipped`, the `command` / `durationMs` / count fields of `run:start` and `run:end`,
425
+ and `ttlMs` / `acquireMs` on `lock:acquired`.
426
+
427
+ ### Tooling
428
+
429
+ - A unit test pins `src/utils/id.js` as the only module that mints an identifier, so no id can
430
+ bypass `generateId`.
431
+ - A unit test pins that nothing under `src/` or `bin/`, and neither declaration file, imports an
432
+ `@opentelemetry/*` package or `bullmq-otel`; they are devDependencies only. The structural types
433
+ are checked part by part against the real `@opentelemetry/api`, and one integration test runs the
434
+ real `@opentelemetry/instrumentation-mongodb` to prove the driver's spans nest under a migration.
435
+ - Declared collections are pinned by a table-driven planner test, a fixed-point integration matrix
436
+ (every kind of index converges, then plans as unchanged on a real server), and the shared queue
437
+ scenarios on both the fake and the real BullMQ.
438
+ - An in-tree fake BullMQ carries the adapter's unit and integration tests; the same scenarios run
439
+ against the real `bullmq` package when `MIGRONAUT_TEST_REDIS_URL` is set, which CI now does
440
+ (a Redis service on the `test` job). `bullmq` and `ioredis` are devDependencies for that only.
441
+
6
442
  ## v2.0.0 — 2026-08-30
7
443
 
8
444
  A major bump for three narrow contract changes (below); everything else is