@alexify/migronaut 1.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/CHANGELOG.md +409 -1
  2. package/README.md +248 -24
  3. package/bin/migronaut.js +11 -3
  4. package/bullmq.d.ts +845 -0
  5. package/bullmq.js +1 -0
  6. package/index.d.ts +757 -29
  7. package/migronaut.schema.json +191 -1
  8. package/package.json +27 -6
  9. package/src/bullmq/index.js +55 -0
  10. package/src/bullmq/jobs.js +454 -0
  11. package/src/bullmq/processor.js +608 -0
  12. package/src/bullmq/producer.js +424 -0
  13. package/src/bullmq/service.js +653 -0
  14. package/src/bullmq/wait.js +124 -0
  15. package/src/cli/args.js +12 -2
  16. package/src/cli/commands/baseline.js +45 -0
  17. package/src/cli/commands/converge.js +160 -0
  18. package/src/cli/commands/down.js +2 -0
  19. package/src/cli/commands/lock.js +2 -1
  20. package/src/cli/commands/redo.js +8 -1
  21. package/src/cli/commands/unlock.js +12 -2
  22. package/src/cli/commands/up.js +14 -1
  23. package/src/cli/exit-codes.js +10 -2
  24. package/src/cli/index.js +4 -0
  25. package/src/cli/shared.js +29 -7
  26. package/src/cli/table.js +105 -0
  27. package/src/core/audit.js +17 -3
  28. package/src/core/baseline.js +80 -0
  29. package/src/core/changelog.js +140 -24
  30. package/src/core/collections.js +372 -0
  31. package/src/core/config.js +125 -27
  32. package/src/core/converge-log.js +47 -0
  33. package/src/core/converge-plan.js +483 -0
  34. package/src/core/converge.js +867 -0
  35. package/src/core/import-runner.js +34 -6
  36. package/src/core/import.js +14 -7
  37. package/src/core/index-spec.js +496 -0
  38. package/src/core/lock-wait.js +260 -0
  39. package/src/core/lock.js +71 -20
  40. package/src/core/migrator.js +805 -304
  41. package/src/core/options.js +251 -0
  42. package/src/core/run-recorder.js +157 -0
  43. package/src/core/run.js +71 -71
  44. package/src/core/runner.js +70 -20
  45. package/src/core/sequence.js +134 -0
  46. package/src/errors/index.js +71 -1
  47. package/src/index.js +16 -0
  48. package/src/utils/actor.js +48 -0
  49. package/src/utils/canonical.js +179 -0
  50. package/src/utils/collection-name.js +21 -0
  51. package/src/utils/error.js +18 -1
  52. package/src/utils/id.js +77 -0
  53. package/src/utils/loader.js +39 -21
  54. package/src/utils/logger.js +30 -12
  55. package/src/utils/migration-name.js +32 -0
  56. package/src/utils/redact.js +57 -4
  57. package/src/utils/sanitize.js +8 -3
  58. package/src/utils/telemetry.js +393 -0
  59. package/src/utils/template.js +60 -12
package/CHANGELOG.md CHANGED
@@ -1,8 +1,416 @@
1
1
  # Changelog
2
2
 
3
3
  All notable changes to this project will be documented in this file.
4
+ Release headings carry the publish date (`## vX.Y.Z — YYYY-MM-DD`).
4
5
 
5
- ## v1.0.0
6
+ ## v2.1.0 — 2026-10-04
7
+
8
+ Migrations as a queue, ids in your own format, OpenTelemetry, and declared collections. Additive:
9
+ nothing changes for anyone who uses none of them, with the narrow exceptions listed under
10
+ **Changed**.
11
+
12
+ ### Added
13
+
14
+ - **`@alexify/migronaut/bullmq`** — a new entry point that runs migrations as
15
+ [BullMQ](https://docs.bullmq.io/) jobs, **one migration per job**, so migronaut can be a
16
+ migration service: enqueue from an HTTP handler, a schedule or a deploy hook, and let a worker
17
+ apply them in order.
18
+ - `createMigrationQueue(options)` — the facade: `enqueueUp` / `enqueueDown` (returning a group
19
+ handle with `wait()`), `startWorker`, `status` / `pending` / `audit` / `lockInfo`, `getJob`,
20
+ `pause` / `resume`, `schedule` / `unschedule`, `close`.
21
+ - `createMigrationProcessor(options)` — the processor on its own, for a Worker you construct
22
+ (NestJS, BullMQ Pro), plus `enqueueUp` / `enqueueDown` / `planUpJobs` / `planDownJobs` /
23
+ `waitForGroup` for a Queue you own.
24
+ - **BullMQ is injected, never depended on** — `bullmq: { Queue, Worker, QueueEvents }` from
25
+ your own install. The package still has no `dependencies` and gains no peer; `src/` never
26
+ imports `bullmq` (a test enforces it).
27
+ - **Order comes from MongoDB, not from Redis.** Every job is a single-file run under the usual
28
+ lock; it refuses while an earlier migration is still pending, so a failed migration stops the
29
+ line (`MIGRATION_BLOCKED` for the jobs behind it). Jobs get one attempt on purpose — a BullMQ
30
+ retry re-queues behind the waiting jobs — and a held lock is waited out inside the job.
31
+ - **One batch per enqueue**, so `down` still rolls back a whole deploy; duplicate enqueues are
32
+ deduplicated, and a job whose migration is already applied completes as `skipped`.
33
+ - Job payloads are validated as untrusted input; messages, stacks and job logs are redacted. A
34
+ queue job's target must be a file of the migration sequence — a payload can never make the
35
+ worker import a dotfile, a declaration file or a helper module next to the migrations.
36
+ - **Shutdown puts unstarted work back.** A job that a closing worker stops before its migration
37
+ starts (waiting for the lock, or fetched during shutdown) is moved back to the head of the
38
+ queue instead of failing — so a rolling deploy no longer fails the rest of the enqueue it
39
+ interrupts as `MIGRATION_BLOCKED`. A migration already running always finishes.
40
+ - **A versioned, strict job contract.** A worker accepts every job data version from
41
+ `MIN_JOB_DATA_VERSION` (exported) up to its own and refuses newer ones and unknown fields
42
+ rather than ignoring what they mean — roll workers out before producers. Jobs always state
43
+ `ordered`, and may carry the plan-time `checksum` of their file.
44
+ - **Scheduler ticks are bounded**: they carry the queue's `jobOptions`, and keep the last 100
45
+ completed / 500 failed jobs when those set no retention.
46
+ - **Correlation**: a job's `runId` is on its progress (`completed` and `failed`), its failure log
47
+ row and its error's `context` (with `jobId` and `groupId`) — a failed job has no return value;
48
+ the worker's failure log line names the group, migration and run.
49
+ - An injected `Queue` (or `QueueEvents`) on another name or prefix than the facade's is rejected,
50
+ and the facade takes an injected queue's prefix by default; `startWorker()` can be retried
51
+ after a failed start; a closed queue refuses every further call.
52
+ - **A worker decides what a job may ask for** — the `allow` option (`{ down: true, force:
53
+ false, unordered: false }` by default) on `createMigrationProcessor` and
54
+ `createMigrationQueue`: a payload asking to re-run an applied migration or to skip the order
55
+ guard is refused (`QUEUE_JOB_INVALID`, `context.permission`) unless allowed. The facade's
56
+ `enqueue*` calls follow the same policy.
57
+ - **A job runs only the file it was planned with**: an `up` job carries the plan-time checksum,
58
+ and a worker with another version of the file fails it (`CHECKSUM_MISMATCH`,
59
+ `context.planned`) instead of applying it.
60
+ - **Several workers without global concurrency**: a job blocked only by earlier migrations that
61
+ have not failed (one may be in flight on another worker) waits for them within its lock-wait
62
+ budget instead of failing its group; `MigrationBlockedError` carries `context.failed`.
63
+ - `wait()` holds every await to its one `timeoutMs` budget, decides a timeout by the clock, and
64
+ its `QueueJobFailedError` carries the job's own typed `code`; an empty group can be waited for
65
+ without QueueEvents. A long lock wait reports progress every few seconds, not every poll.
66
+ - What the queue stores about a failure — `failedReason`, stack, job logs — masks the values a
67
+ duplicate-key error quotes, on top of credentials. `schedule({ every })` needs at least 1000 ms.
68
+ - Dedup ids encode file names reversibly (two names can no longer share one and absorb each
69
+ other's job), and a forced re-run has a dedup id of its own.
70
+ - **A schedule holds a failed migration**: a `sync` tick whose next migration failed, with its
71
+ file unchanged since, enqueues nothing (`returnvalue.held`, and a warning) instead of re-running
72
+ it every tick; a changed file or an explicit `enqueueUp(name)` resumes it.
73
+ - Jobs carry `requestedBy` / `reason` (`enqueueUp` / `enqueueDown` / `enqueueConverge` options).
74
+ - **`bullmq.d.ts`** — hand-written types for the entry point, with structural `BullMQ*Like`
75
+ interfaces instead of an import of `bullmq`, generic over the classes you inject.
76
+ - **`up(file, { batch })`** — stamp an explicit batch number instead of the next free one, and
77
+ **`MigratorKit.nextBatch()`** to peek at it: together they let several single-file runs form one
78
+ batch.
79
+ - **`up(file, { ordered: true })` / `down(file, { ordered: true })`** — refuse a single-file run
80
+ that would go out of sequence (`MigrationBlockedError`); an ordered `up` also applies the
81
+ `strict` drift check and the `onOutOfOrder` policy a bulk run would, and only targets a file of
82
+ the migration sequence.
83
+ - **`up(file, { checksum })`** — refuse (`CHECKSUM_MISMATCH`, `context.planned`) to apply any other
84
+ version of the file than the one with this SHA-256; `dryRun('up')` rows carry each file's
85
+ `checksum` for it.
86
+ - **`list(filter, { checksums: false })`** / **`status({ checksums: false })`** — skip hashing the
87
+ applied files, for a caller that needs names and dates only.
88
+ - **Who asked, and why** — `up`, `down`, `redo` and `converge` take `requestedBy` (≤ 128
89
+ characters) and `reason` (≤ 512), and `migronaut up` / `down` / `redo` / `converge` take
90
+ `--reason`. They are stamped on the changelog (`requestedBy` / `reason` on an apply,
91
+ `revertRequestedBy` / `revertReason` on a revert; a later apply that says nothing clears the old
92
+ ones) and on the converge history; `status()` rows show them. `executedBy` stays the OS user —
93
+ on a queue worker, the container's — which is why the requester has fields of its own. A failed
94
+ attempt's trace now also records the checksum of the file version that failed
95
+ (`StatusRow.failedChecksum`).
96
+ - **`runMigrations(config, { signal })`** — an `AbortSignal` (wired to SIGTERM) stops a wait for
97
+ the lock between polls, and a run that holds it between migrations.
98
+ - **`LockInfo.runId` and `LockInfo.ttlMs`** — which run holds the lock (also on `migronaut lock`),
99
+ and the holder's TTL, which the lock document now records.
100
+ - **Three error codes**: `MIGRATION_BLOCKED` (exit 24), `QUEUE_JOB_INVALID` (25),
101
+ `QUEUE_JOB_FAILED` (26), with `MigrationBlockedError`, `QueueJobInvalidError` and
102
+ `QueueJobFailedError` exported from the package root.
103
+ - **Runnable example** — `examples/migration-service`: a queue, a worker and a plain `node:http`
104
+ API (not published to npm).
105
+ - **`generateId` config option** — your own identifier format (ULID, CUID, nanoid, UUIDv7, …)
106
+ everywhere migronaut used to call `crypto.randomUUID()`: the `runId` of every run — on changelog
107
+ records, events, log lines and the lock's owner token — and the `groupId` of every queue
108
+ enqueue. One option covers the kit, the CLI (through `migronaut.config.js`/`.ts`),
109
+ `runMigrations` and the queue adapter.
110
+ - **Injected, like the logger** — migronaut ships no generator but the default. It is called
111
+ with no arguments, so third-party functions pass straight through: `generateId: ulid`.
112
+ - **Checked on every call** — it must synchronously return a non-empty string of at most 128
113
+ characters. A throw, a promise or anything else fails the run with `CONFIG_INVALID` before a
114
+ migration starts (and an enqueue before a job is added).
115
+ - Code-only: no environment variable and no place in a JSON config, like `logger` and `hooks`.
116
+ - **`MigratorKit.generateId()`** — a new id in the kit's configured format, for code that wants its
117
+ own ids to match (it is how the queue adapter mints group ids). Resolves the config; does not
118
+ connect.
119
+ - **`IdGenerator` type** — `() => string`, exported from the package root.
120
+ - **`telemetry` config option** — OpenTelemetry traces and metrics through a tracer and/or a meter
121
+ from your own `@opentelemetry/api`: `telemetry: { tracer, meter }`. One option covers the kit,
122
+ the CLI (through `migronaut.config.js`/`.ts`), `runMigrations` and the queue adapter.
123
+ - **Spans**: `migronaut.run` for every run that acquired the lock, and a child
124
+ `migronaut.migration` for every migration executed. The migration's span is the *active* one
125
+ while its hooks, its `up`/`down` and its changelog write run — so an instrumented MongoDB
126
+ driver nests its command spans under the migration that issued them. That is what lifecycle
127
+ events cannot do, and why this lives in the kit: at application startup there is no ambient
128
+ span, and the driver instrumentation records nothing without a parent.
129
+ - **Metrics**: `migronaut.run.duration`, `migronaut.migration.duration`,
130
+ `migronaut.lock.acquire.duration` and `migronaut.lock.wait.duration` (one point per wait for a
131
+ held lock, by `migronaut.lock.wait.outcome`: `acquired`, `timeout`, `aborted`) — histograms,
132
+ in seconds, with boundaries from 10ms to an hour — and the counters `migronaut.lock.refused`
133
+ and `migronaut.lock.lost`.
134
+ - **Dimensions**: every span and metric point carries `db.namespace` (the database name), plus
135
+ the caller's own static attributes from `telemetry.attributes` (at most 20 scalars).
136
+ - **Failures** set the span's status to `ERROR` with a redacted message (credentials and the
137
+ values a duplicate-key error quotes masked, at most 1 KB), and `error.type` — on
138
+ the span and on the metric point — to the typed error code. No exception event is recorded: it
139
+ would carry the unredacted message and stack.
140
+ - **A run that never got the lock emits no span.** A caller polling for a busy lock retries the
141
+ whole run every few hundred milliseconds; the refusals are counted
142
+ (`migronaut.lock.refused`) instead.
143
+ - **Injected, like the logger** — `@opentelemetry/api` is neither a dependency nor a peer, and
144
+ `src/` never imports it (a test enforces it). The types are structural: `MigronautTracer`,
145
+ `MigronautSpan`, `MigronautMeter`, `MigronautHistogram`, `MigronautCounter`,
146
+ `MigronautMetricOptions`, `MigronautAttributes` and `MigronautTelemetry`, exported from the
147
+ package root.
148
+ - **It can never fail a run.** Every call into the tracer, a span, the meter and an instrument
149
+ is guarded — a promise one of them returns included, so a rejecting SDK cannot surface as an
150
+ unhandled rejection — and a tracer that throws before or after running the work, or runs it
151
+ twice, still gets each migration executed exactly once.
152
+ - Code-only, like `logger` and `generateId`. The span, attribute and metric names are new and
153
+ should be treated as experimental.
154
+ - **`bullmq.telemetry`** — `createMigrationQueue({ bullmq: { Queue, Worker, telemetry } })` hands
155
+ BullMQ's own telemetry object (`new BullMQOtel({ tracerName })` from `bullmq-otel`) to the Queue
156
+ and the Worker it constructs. Together with the kit's `telemetry` it gives one trace from the
157
+ request that enqueued, across Redis, to the MongoDB commands in the worker. Until now the facade
158
+ built its Queue without it, so the enqueuing side of that trace could not be joined.
159
+ `workerOptions.telemetry` and `startWorker({ telemetry })` override it for the worker alone.
160
+ - **OpenTelemetry in the example** — `examples/migration-service` gains `tracing.js` and a Jaeger
161
+ in its `docker-compose.yml`: set `OTEL_EXPORTER_OTLP_ENDPOINT` and every enqueue is one trace.
162
+ - **Declared collections** — indexes and validators declared as an end state, and applied by
163
+ `migronaut converge`, with no migration file per change. For what only ever has a current value
164
+ (which indexes a collection has, which validator guards it); migrations stay the tool for
165
+ changes with an order and a history.
166
+ - **`collections` config option** — an array of definitions, `{ name, indexes?, validator?,
167
+ validationLevel?, validationAction?, prune? }`, each index in the driver's own flat
168
+ `createIndexes` shape. Works in `migronaut.config.{ts,js,json}` (and the JSON Schema) and in
169
+ `new MigratorKit({ collections })`.
170
+ - **`collectionsDir` config option** (`MIGRONAUT_COLLECTIONS_DIR`) — one definition file per
171
+ collection (`.ts`/`.js` default export, or `.json`; the name defaults to the file name).
172
+ Opt-in, combined with `collections`, and loaded only when a converge runs — a broken file never
173
+ blocks `status` or an emergency `down`. A collection declared twice is `CONFIG_INVALID`.
174
+ - **Validated strictly**: an unknown definition key or index option is an error with its path
175
+ (`collections[2].indexes[0].uniqe`) — the driver silently drops an option it does not know, so
176
+ a typo would otherwise build the wrong index and then look in sync forever.
177
+ - **`MigratorKit.converge({ dryRun?, prune?, noLock?, ordered? })`** and **`migronaut converge`**
178
+ (`--dry-run`, `--check`, `--prune`, `--yes`). Stateless: every run reads `listCollections` and
179
+ `listIndexes`, plans, and carries the plan out one step at a time under the migration lock —
180
+ create the collection, set the validator, create indexes, `collMod` a TTL or `hidden` in place,
181
+ rebuild what changed otherwise, drop last. Nothing is recorded.
182
+ - **Safe by default.** An index you did not declare is kept and reported (`keep`), and dropped
183
+ only with `prune` (per definition, or for the definitions that do not decide). An identical
184
+ index under another name is accepted as is rather than rebuilt; a different one covering the
185
+ same key is a conflict that refuses the run before the first write. A rebuild whose new index
186
+ fails to build puts the old one back — and says so, with the reason, when it cannot.
187
+ - **A unique index is never rebuilt unasked.** Dropping it opens a window with no constraint,
188
+ and a duplicate written in that window leaves neither index buildable. Such a rebuild is a
189
+ `conflict` unless `converge({ rebuildUnique: true })` / `--rebuild-unique`; a converge after
190
+ `up` and a queue job never pass it.
191
+ - **Comparisons follow what the server stores** — text indexes in their `_fts` form, collations
192
+ field by field against the expanded spec (`strength`, `caseLevel` and `numericOrdering` at
193
+ their universal defaults when left out), the collection's default collation,
194
+ `{ locale: 'simple' }`, a flag stored as `1`. A compound `Map` key may hold an integer-like
195
+ field only first — the driver reads keys back as plain objects.
196
+ Anything applied that still compares as changed is reported under `unstable` instead of being
197
+ rebuilt on every run.
198
+ - **The CLI plans first** and asks before any drop or rebuild, and before changing the validator
199
+ of a collection that holds data; `--json` refuses such a plan without `--yes` (and applies a
200
+ purely additive one); a plan with a conflict is refused without asking. `--check` exits `28` on
201
+ drift — a CI gate; `--ordered` refuses while a migration is pending.
202
+ - **Results say what changed**: every row that changes, drops or keeps something carries
203
+ `from` and `to` — the live and the declared index or validator, as plain JSON.
204
+ - **History** — every converge that changes something or fails appends an entry to
205
+ `_migronaut_converge` (`convergeLogCollection`, `MIGRONAUT_CONVERGE_LOG_COLLECTION`): when,
206
+ the trigger, the run id, who ran it and where, who asked and why, and every row it touched
207
+ with its `from` / `to`. Read it with `MigratorKit.convergeHistory({ limit })` or
208
+ `migronaut converge --history [--limit n] [--json]`. Best-effort, and a database that never
209
+ converges never gets the collection.
210
+ - **In place where the server can**: making an index unique (MongoDB 7.0+, `collMod`
211
+ `prepareUnique` then `unique` — duplicates leave the index as it was) and adding a TTL to a
212
+ single-field index (5.1+) are `modify`, not a rebuild. (6.0 accepts the unique conversion but
213
+ did not enforce it in our tests, so it rebuilds.)
214
+ - **Sharded clusters**: behind a `mongos`, prune never drops the index backing a shard key (read
215
+ from `config.collections`, or kept with a warning when the server refuses the drop).
216
+ - **Re-planned before each collection**: every collection after the first is read and planned
217
+ again right before its turn; one that changed meanwhile into a conflict or a new drop/rebuild
218
+ stops the run (`ConvergeFailedError`, `phase: 'replan'`) before it is touched.
219
+ - **Regular expressions** in a validator or partial filter must use flags the driver stores as
220
+ written (`i`, `m`; a `BSONRegExp` for server options) — `g` would become dotAll and `s`, `u`,
221
+ `y` vanish — and compare in their stored form. A `__proto__` key in a JSON definition stays a
222
+ key.
223
+ - **At scale**: all declared collections are read with one `listCollections` and a bounded
224
+ fan-out of `listIndexes`; the new indexes of a collection are built by one `createIndexes`
225
+ (one pass over the data); a connection that fails mid-rebuild is reported, never "repaired"
226
+ by restoring the old index next to a build the server may still be running.
227
+ - **`convergeAfterUp` config option** (`MIGRONAUT_CONVERGE_AFTER_UP`) — a bulk `up` (no file, no
228
+ `to`; also `runMigrations`) ends by converging under the same lock, even when nothing was
229
+ pending, so a failed converge is retried by the next deploy. `up(undefined, { converge })` and
230
+ `migronaut up --converge` / `--no-converge` decide per run. `up` still returns its migration
231
+ rows; `runMigrations` adds `summary.converge`.
232
+ - **`MigratorKit.convergesAfterUp()`** — whether a bulk `up` on the kit ends by converging.
233
+ - **Events**: `converge:start`, `converge:action` (per step: `'started'` before it runs — an index
234
+ build can take hours — then `'applied'` or `'failed'`) and `converge:end` (with the full
235
+ result), for real runs. An index build also logs which index it is starting on. The run itself is an ordinary `run:start`/`run:end` with
236
+ `command: 'converge'`, and an ordinary `migronaut.run` span.
237
+ - **Types**: `CollectionDefinition`, `CollectionDefinitionFile`, `IndexDefinition`,
238
+ `IndexKeyDirection`, `IndexCollation`, `ValidationLevel`, `ValidationAction`, `ConvergeOptions`,
239
+ `ConvergeResult`, `CollectionConvergeResult`, `ConvergeAction`, `ConvergeActionKind`,
240
+ `ConvergeActionStatus`, `ConvergeTarget`, `ConvergeUnstable`, `ConvergeTrigger` and the three
241
+ event payloads, exported from the package root.
242
+ - The definition shape, the result shape and the queue contract below are new and should be
243
+ treated as experimental.
244
+ - **Converge jobs in the queue adapter** — `JOB_NAMES.CONVERGE` (`'converge'`), `enqueueConverge()`
245
+ and `MigrationQueue.enqueueConverge()`, and `schedule({ job: 'converge' })` with its own default
246
+ id, `DEFAULT_CONVERGE_SCHEDULER_ID` (`'migronaut-converge'`). With `convergeAfterUp`,
247
+ `enqueueUp` ends a group that reaches the newest migration with a converge job (or adds a
248
+ converge-only job when nothing is pending but a dry run finds drift), and an idle `sync` tick does
249
+ the same. A converge job refuses as `MIGRATION_BLOCKED` while a migration is pending, is keyed for
250
+ deduplication on the migration it follows, and carries no `prune` — what may be dropped comes
251
+ from the worker's own definitions. New types: `ConvergeJobData`, `ConvergeJobResult`,
252
+ `ConvergeJobSpec`, `ConvergeHandle`, `EnqueueConvergeOptions`.
253
+ - **Two exit codes**: `CONVERGE_FAILED` (27, with `ConvergeFailedError` exported from the package
254
+ root) and the CLI-only `COLLECTIONS_DRIFT` (28, from `converge --check`).
255
+
256
+ ### Changed
257
+
258
+ - **`MigronautErrorCode` gained four members** (above). TypeScript consumers with an exhaustive
259
+ `switch` over the code union need a `default` branch or the new cases.
260
+ - **`collections`, `collectionsDir` and `convergeAfterUp` config keys are now validated**
261
+ (`CONFIG_INVALID`). They were previously unknown and ignored, like any stray key.
262
+ - **The CLI arg parser makes a `--x` / `--no-x` pair tri-state**, as commander does: when a command
263
+ declares both, neither given leaves the option unset instead of defaulting to `true`. Only
264
+ `up --converge` / `--no-converge` uses this.
265
+ - **`dryRun('up')` now applies the out-of-order policy** of the run it previews: under
266
+ `onOutOfOrder: 'error'` a bulk preview refuses with `MIGRATION_OUT_OF_ORDER` instead of listing
267
+ rows the run would reject; under `'warn'` it logs the warning. A single-file preview is exempt,
268
+ as the single-file run is.
269
+ - **`connect()` is safe to call concurrently** — overlapping first calls on one `MigratorKit`
270
+ share a single connection instead of each opening (and all but one leaking) a client. Matters
271
+ for a long-lived kit serving several callers.
272
+ - The lock-wait loop of `runMigrations` moved to `src/core/lock-wait.js`, shared with the queue
273
+ processor. Two changes to how it waits: **polls back off**, doubling from `lockPollIntervalMs`
274
+ up to 5 s (and at most a quarter of the budget), so a fleet waiting out a long deploy no longer
275
+ hammers the lock document; and **the default `lockWaitTimeoutMs` follows the holder's TTL** —
276
+ `max(90 s, 1.5 × lockTTLSeconds)` — so a holder with a long TTL, whose heartbeat moves the lock
277
+ only every TTL/2, is no longer mistaken for a stalled one. An explicit `lockWaitTimeoutMs` is
278
+ used as given (with a warning when it is shorter than the holder's heartbeat). `waitedMs` is
279
+ now measured by the clock from the first refusal, and a wait that times out rethrows the
280
+ refusal with `context.timedOut`, `attempts` and `waitedMs`.
281
+ - **`runMigrations` validates `onLockHeld`** (`CONFIG_INVALID`): a value other than `'throw'` or
282
+ `'wait'` — `'Wait'`, say — used to behave as `'throw'` without a word. A poll interval above the
283
+ largest timer (2³¹−1 ms, which Node fires after 1 ms) is refused too.
284
+ - **The lock document gained a `nonce` field**, minted by migronaut on every acquire and matched
285
+ alongside `owner` when the lock is confirmed, renewed and released. The owner token is the run
286
+ id, whose format `generateId` now decides; the nonce keeps mutual exclusion independent of it,
287
+ so a generator that repeats an id can blur correlation but never let two runs hold the lock.
288
+ Older releases ignore the field and can share a database with this one. The nonce is never
289
+ exposed; the owner — the holder's run id — now is, as `LockInfo.runId` (`lock`, `lockInfo()`,
290
+ `LockAlreadyHeldError` context), next to the holder's `ttlMs`, also written since this release:
291
+ "which run holds the lock?" is the first question about a stuck one, and the nonce, not the
292
+ owner, is what proves ownership.
293
+ - **A `generateId` config key that is not a function is now rejected** (`CONFIG_INVALID`). The key
294
+ was previously unknown and ignored, like any stray key.
295
+ - **A `telemetry` config key that is not usable is now rejected** (`CONFIG_INVALID`): it must be an
296
+ object whose `tracer` has `startActiveSpan` and whose `meter` has `createHistogram` and
297
+ `createCounter`. Absent, `null` and `{}` all mean "off". The key was previously unknown and
298
+ ignored.
299
+ - **A scheduled `sync` job no longer joins the trace that registered its schedule.** Its template
300
+ now carries `telemetry: { omitContext: true }`: BullMQ builds each scheduler iteration from the
301
+ previous job's options, so with telemetry on, every tick would otherwise have been appended to
302
+ one ever-growing trace. Each tick is a trace of its own; no effect on a queue without telemetry.
303
+
304
+ ### Fixed
305
+
306
+ - `docs/reference/cli.md` lists exit code `23` (`MIGRATION_OUT_OF_ORDER`), missing since v2.0.0.
307
+ - The events table in `docs/guide/hooks.md` had fallen behind the payloads: it now lists
308
+ `migration:skipped`, the `command` / `durationMs` / count fields of `run:start` and `run:end`,
309
+ and `ttlMs` / `acquireMs` on `lock:acquired`.
310
+
311
+ ### Tooling
312
+
313
+ - A unit test pins `src/utils/id.js` as the only module that mints an identifier, so no id can
314
+ bypass `generateId`.
315
+ - A unit test pins that nothing under `src/` or `bin/`, and neither declaration file, imports an
316
+ `@opentelemetry/*` package or `bullmq-otel`; they are devDependencies only. The structural types
317
+ are checked part by part against the real `@opentelemetry/api`, and one integration test runs the
318
+ real `@opentelemetry/instrumentation-mongodb` to prove the driver's spans nest under a migration.
319
+ - Declared collections are pinned by a table-driven planner test, a fixed-point integration matrix
320
+ (every kind of index converges, then plans as unchanged on a real server), and the shared queue
321
+ scenarios on both the fake and the real BullMQ.
322
+ - An in-tree fake BullMQ carries the adapter's unit and integration tests; the same scenarios run
323
+ against the real `bullmq` package when `MIGRONAUT_TEST_REDIS_URL` is set, which CI now does
324
+ (a Redis service on the `test` job). `bullmq` and `ioredis` are devDependencies for that only.
325
+
326
+ ## v2.0.0 — 2026-08-30
327
+
328
+ A major bump for three narrow contract changes (below); everything else is
329
+ additive. Upgrading is a no-op for the most commonly scripted surface:
330
+ `pendingMigrations()` / `list('pending')` still report a file as `pending`
331
+ even when a failed attempt was recorded for it.
332
+
333
+ ### Breaking changes
334
+
335
+ - **`unlock --json` now requires `--yes`.** Previously `--json` alone
336
+ force-released the lock; it now refuses with `CONFIG_INVALID` (exit 6), because
337
+ a non-interactive mode must never assume consent to a destructive action —
338
+ force-releasing a live run's lock enables exactly the concurrent migration the
339
+ lock exists to prevent. **Migration:** add `--yes` to any scripted
340
+ `migronaut unlock --json`.
341
+ - **`StatusRow.status` and `MigrationStatus` gained `'failed'`.** A recorded
342
+ failed attempt now renders as `failed` in `status()` / `list('all')` instead of
343
+ `pending`. **Migration:** TypeScript consumers with an exhaustive `switch` over
344
+ the status must handle the new member; treat `'failed'` as outstanding work.
345
+ - **`MigronautErrorCode` gained `'MIGRATION_OUT_OF_ORDER'`** (and `EXIT_CODES`
346
+ the matching key, exit 23). **Migration:** exhaustive `switch` statements over
347
+ the code union must handle it.
348
+
349
+ ### Added
350
+
351
+ - **`migronaut baseline`** — adopt an existing database with no prior migration tool: mark
352
+ migration files on disk as applied (checksum from disk, one shared batch, `origin: 'baseline'`)
353
+ without executing them. Forward-only, idempotent, confirmation-gated.
354
+ - **Out-of-order detection** — a bulk `up` flags pending migrations that sort before the newest
355
+ applied one (a file merged late from a parallel branch). New scalar config `onOutOfOrder:
356
+ 'warn' | 'error' | 'allow'` (default `'warn'`, env `MIGRONAUT_ON_OUT_OF_ORDER`), a new
357
+ `MIGRATION_OUT_OF_ORDER` error code (exit 23), `outOfOrder` on `StatusRow`, and an `ordering`
358
+ check in `audit`.
359
+ - **Failed-attempt traces** — a failing `up` now leaves a best-effort `status: 'failed'` record
360
+ (error text, `failedAt`, `runId`) in the changelog; `status` renders it as `failed` while every
361
+ run path still retries the file; a successful apply clears the trace. A forced re-run's failure
362
+ never demotes an `applied` record.
363
+ - **Audit-trail read surface** — `StatusRow` now carries `executedBy`, `environment`, `runId`,
364
+ `revertedAt` and `origin` from the changelog record, so `status --json` can answer "who ran
365
+ migration X, from which run, and was it ever reverted?".
366
+ - **`onKit` option on `runMigrations`** — receive the internally-constructed `MigratorKit` before
367
+ connect and subscribe to its lifecycle events (metrics without log parsing).
368
+ - **`createLogger` export** — the default console logger (with level control) for programmatic
369
+ callers.
370
+ - **`cwd` option on `MigratorKit`** — scope config discovery, `.env` loading and a relative
371
+ `migrationsDir` to a project root, for processes hosting kits for several projects.
372
+ - **ESM-interop integration test** pinning the "ESM consumers still work" promise.
373
+
374
+ ### Changed
375
+
376
+ - **Server-time changelog stamps** — `appliedAt`/`revertedAt`/`failedAt` are stamped with the
377
+ server clock (`$currentDate`, matching the lock's `$$NOW` discipline) unless the record carries
378
+ an explicit `appliedAt` (import), so `redo`/`down --steps` ordering is immune to host clock skew.
379
+ - **Non-transactional changelog-write failures are reported distinctly** — when a migration's own
380
+ writes committed but recording them failed, the error now says exactly that (context
381
+ `phase: 'changelog-write'`, `bodySucceeded: true`) instead of the generic "migration failed"
382
+ that invited re-running committed writes.
383
+ - **A timed-out migration is told to stop** — `timeoutMs` now aborts `ctx.signal` (with the
384
+ `MigrationTimeoutError` as reason), so cooperative bodies stop writing instead of racing the
385
+ next lock holder.
386
+ - **`lockWaitTimeoutMs` bounds stall, not total wait** — while a waiting `runMigrations` observes
387
+ the holder's heartbeat advancing, the deadline re-arms; only a stalled holder times peers out.
388
+ - **`baseline` requires `--yes` in `--json` mode** — the same non-interactive
389
+ confirmation policy `up --force --json` and `unlock --json` follow.
390
+ - **Failure telemetry carries timing** — `migration:error` events, error result rows and error
391
+ contexts now include `durationMs`/`batch`; the run ends with a `✔ Done N applied in Xms`
392
+ summary line.
393
+ - Signals during the CLI's pre-connect now abort the run (exit 11) instead of being silently
394
+ dropped; partial-results lists survive hook/load failures and lock-release failures; lock
395
+ warn lines carry `runId`; import interruptions report progress and that a `--force` re-run
396
+ resumes idempotently.
397
+
398
+ ### Fixed / hardened
399
+
400
+ - URI redaction now masks query-string secrets (`tlsCertificateKeyFilePassword`, `proxyPassword`,
401
+ `sslKeyPassword`, secret `authMechanismProperties` values) and empty-username passwords —
402
+ in logs, errors, `--json`, events, and `init`-generated config files (which now warn about
403
+ query-string secrets too).
404
+ - Terminal sanitization strips the whole C1 control block (DCS/OSC/PM/APC, not just CSI), and
405
+ `bin/migronaut.js`'s last-resort handlers sanitize their output.
406
+ - `MigrationLock.release()` without a held owner token is a no-op instead of an unscoped delete;
407
+ an uncontended `acquire()` takes one round trip instead of two.
408
+ - Import checksum resolution is concurrency-bounded (no EMFILE on thousands-record changelogs);
409
+ the strict drift check reuses the instance checksum cache; `dry-run up` fetches applied names
410
+ instead of full records; a warn/error-only injected logger keeps its output instead of being
411
+ silenced entirely.
412
+
413
+ ## v1.0.0 — 2026-07-28
6
414
 
7
415
  Initial release. Requires **Node.js ≥ 22.18**.
8
416