mongo-data-anonymizer 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/README.md +268 -18
  2. package/dist/anonymization/anonymize.d.ts +50 -8
  3. package/dist/anonymization/anonymize.js +747 -96
  4. package/dist/anonymization/anonymize.js.map +1 -1
  5. package/dist/anonymization/collections.d.ts +12 -0
  6. package/dist/anonymization/collections.js +14 -0
  7. package/dist/anonymization/collections.js.map +1 -0
  8. package/dist/anonymization/database.d.ts +69 -9
  9. package/dist/anonymization/database.js +233 -61
  10. package/dist/anonymization/database.js.map +1 -1
  11. package/dist/anonymization/randomizer.d.ts +8 -0
  12. package/dist/anonymization/randomizer.js +32 -0
  13. package/dist/anonymization/randomizer.js.map +1 -0
  14. package/dist/anonymization/rules.d.ts +60 -0
  15. package/dist/anonymization/rules.js +193 -0
  16. package/dist/anonymization/rules.js.map +1 -0
  17. package/dist/cli.js +11 -4
  18. package/dist/cli.js.map +1 -1
  19. package/dist/config.d.ts +7 -0
  20. package/dist/config.js +182 -0
  21. package/dist/config.js.map +1 -0
  22. package/dist/index.d.ts +5 -0
  23. package/dist/index.js +6 -0
  24. package/dist/index.js.map +1 -0
  25. package/dist/logger.d.ts +7 -0
  26. package/dist/logger.js +14 -0
  27. package/dist/logger.js.map +1 -0
  28. package/dist/progress.d.ts +25 -0
  29. package/dist/progress.js +74 -0
  30. package/dist/progress.js.map +1 -0
  31. package/dist/run.d.ts +62 -0
  32. package/dist/run.js +365 -0
  33. package/dist/run.js.map +1 -0
  34. package/package.json +48 -30
  35. package/.eslintignore +0 -1
  36. package/.eslintrc.js +0 -25
  37. package/.github/workflows/release.yml +0 -37
  38. package/.github/workflows/test.yml +0 -22
  39. package/.nvmrc +0 -1
  40. package/.prettierrc +0 -7
  41. package/dist/main.d.ts +0 -1
  42. package/dist/main.js +0 -95
  43. package/dist/main.js.map +0 -1
  44. package/dist/utils/field-utils.d.ts +0 -1
  45. package/dist/utils/field-utils.js +0 -18
  46. package/dist/utils/field-utils.js.map +0 -1
  47. package/jest.config.js +0 -8
  48. package/src/anonymization/anonymize.ts +0 -116
  49. package/src/anonymization/database.ts +0 -50
  50. package/src/cli.ts +0 -4
  51. package/src/main.ts +0 -98
  52. package/src/utils/field-utils.ts +0 -16
  53. package/test/anonymizer/anonymizer.test.ts +0 -164
  54. package/test/utils/field-utils.test.ts +0 -31
  55. package/tsconfig.json +0 -14
package/README.md CHANGED
@@ -1,37 +1,287 @@
1
- # MongoDB Anonymizer
1
+ # MongoDB Data Anonymizer
2
2
 
3
- This package allows you to anonymize specified fields in MongoDB collections and copy them to a target database. It is based on the [mongodb-anonymizer](https://github.com/rap2hpoutre/mongodb-anonymizer) package by [rap2hpoutre](https://github.com/rap2hpoutre).
3
+ Copies a MongoDB database to another database, replacing personal data with realistic fake values. Fake values are **deterministic**: the same original value always becomes the same fake value, in every collection and on every run with the same secret. So references between collections still line up, and the output can be reproduced.
4
+
5
+ Based on [mongodb-anonymizer](https://github.com/rap2hpoutre/mongodb-anonymizer) by [rap2hpoutre](https://github.com/rap2hpoutre).
4
6
 
5
7
  ## Installation
6
8
 
7
- You can install the package globally using npm:
9
+ Requires Node.js 22.13 or newer.
8
10
 
9
11
  ```bash
10
- npm install -g mongodb-anonymizer
12
+ npm install -g mongo-data-anonymizer
11
13
  ```
12
14
 
13
- ## Usage
15
+ Or run it without installing:
14
16
 
15
- To use the package, you need to provide the source and target MongoDB URIs, the fields to anonymize, and optionally, the collections to anonymize or ignore, and the batch size. Here is an example:
17
+ ```bash
18
+ npx mongo-data-anonymizer --help
19
+ ```
20
+
21
+ ## Usage
16
22
 
17
23
  ```bash
18
- mongodb-anonymizer --sourceUri "mongodb://localhost:27017/source" --targetUri "mongodb://localhost:27017/target" --fieldList "email,password" --collectionList "users,admins" --ignoreCollections "logs" --batchSize 500
24
+ export ANONYMIZER_SECRET="$(openssl rand -hex 32)" # keep it to reproduce the same output later
25
+
26
+ mongo-data-anonymizer \
27
+ --sourceUri "mongodb://localhost:27017/production" \
28
+ --targetUri "mongodb://localhost:27017/staging" \
29
+ --fieldList "+users.password:null,+*guestEmail,-description" \
30
+ --ignoreCollections "logs" \
31
+ --copyNonAnonymized \
32
+ --dropTarget
19
33
  ```
20
34
 
21
- In this example, the `email` and `password` fields in the `users` and `admins` collections will be anonymized, the `logs` collection will be ignored, and the batch size for processing documents is set to `500`.
35
+ With these options:
36
+
37
+ - Every collection is anonymized except `logs`, which is copied as-is (`--copyNonAnonymized`).
38
+ - The [default fields](#default-fields) are anonymized, plus `password` in `users` (set to `null`) and every key ending in `guestEmail`, minus `description`. Email addresses in any other string are replaced too (`--scrubEmails`, on by default).
39
+ - Collections that already exist in the target are dropped and rewritten (`--dropTarget`).
40
+
41
+ To see what would happen without writing anything, run it first with `--dryRun`. For every collection, it shows:
42
+
43
+ - whether the collection would be anonymized, copied or skipped;
44
+ - roughly how many documents it has;
45
+ - which keys match no rule but look like personal data by name, such as `SSNNumber` or `guestPhoneList`. It checks the first 1000 documents of each collection (`--sampleSize`, 0 for all), so rare keys can be missed, and it ignores boolean and numeric values. Those keys would be written unchanged, so add them to `--fieldList` if needed. With `--no-scrubEmails`, keys whose text contains an email address are listed too.
46
+
47
+ A dry run also runs the [safety checks](#safety-checks), except the probe collection, which would be a write. So it fails too if, for example, the target already has the collections and `--dropTarget` isn't set.
22
48
 
23
49
  ## Options
24
50
 
25
- Here are the available options:
51
+ | Option | Default | Description |
52
+ | ------------------------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
53
+ | `--sourceUri` | _required_ | URI of the source database, including the database name. It is only read from; see [Safety checks](#safety-checks). |
54
+ | `--targetUri` | _required_ | URI of the target database, including the database name. The run refuses to start if it is the same database as the source. |
55
+ | `--fieldList` | see below | Comma-separated [field rules](#field-rules). |
56
+ | `--collectionList` | all | Comma-separated collections to anonymize. |
57
+ | `--ignoreCollections` | none | Comma-separated collections never to anonymize. |
58
+ | `--copyNonAnonymized` | `false` | Copy the collections that are not anonymized (not in `--collectionList`, or in `--ignoreCollections`) as-is. Without it they are skipped. |
59
+ | `--dropTarget` | `false` | Drop target collections that already exist. Without it the run stops before writing anything if any of them exists. |
60
+ | `--dryRun` | `false` | Print what would be done for every collection, without writing anything. |
61
+ | `--strict` | `false` | Fail instead of warning when a matched value can't be anonymized (ObjectId, Binary, Decimal128, invalid dates, ...), so no personal data is left behind silently. |
62
+ | `--scrubEmails` | `true` | Replace email addresses found in any string, including fields no rule matches and free text, with the same fake email they get everywhere else. Turn off with `--no-scrubEmails`. |
63
+ | `--sampleSize` | `1000` | Documents per collection a dry run checks for unmatched personal-looking keys. `0` checks every document. |
64
+ | `--assumeDifferentTarget` | `false` | Allow a target that shares collection UUIDs or servers with the source, such as a copy restored with `mongorestore --preserveUUID`. The probe collection still checks that writes to the target don't reach the source. |
65
+ | `--indexCommitQuorum` | `1` | Commit quorum for index builds when the target is a replica set: a number of members, `majority` or `votingMembers`. See [Replica sets and large databases](#replica-sets-and-large-databases). |
66
+ | `--indexTimeout` | `900` | Seconds to wait for a collection's indexes to build. When it runs out, those indexes are skipped with a warning. `0` means no limit. |
67
+ | `--batchSize` | `1000` | Documents read and inserted per batch. A batch is also cut at 16 MB, so large documents don't use much memory. |
68
+ | `--secret` | random | Secret for [deterministic anonymization](#deterministic-anonymization). Prefer the `ANONYMIZER_SECRET` environment variable, which stays out of your shell history. |
69
+
70
+ Every option can also be set with an environment variable: `ANONYMIZER_` followed by the option name in upper snake case, e.g. `ANONYMIZER_SOURCE_URI`, `ANONYMIZER_TARGET_URI`, `ANONYMIZER_SECRET`. Boolean variables accept `true`/`false`, `1`/`0`, `yes`/`no` and `on`/`off`; any other value is an error. Command-line options take precedence over environment variables, and a repeated option isn't merged: the last value wins. Other `ANONYMIZER_*` variables are ignored.
71
+
72
+ The process exits with a non-zero code when the run fails.
73
+
74
+ ## Safety checks
75
+
76
+ Before writing anything, the run stops if:
77
+
78
+ - **The source and target are the same database.** Two checks cover this:
79
+ - **Similarity check.** The tool compares collection UUIDs, and the database name on the same servers. A match refuses the run.
80
+ - This check can also match two different databases: a target restored from the source with `mongorestore --preserveUUID` or from a disk snapshot, or two identical docker stacks. If you are sure the target is a different database, pass `--assumeDifferentTarget` to skip it.
81
+ - **Probe collection.** The tool creates an empty collection named `anonymizer_probe_<random>` in the target, checks whether it appears in the source, then drops it right away.
82
+ - This catches the same server reached through another address, such as an SSH tunnel or a different port mapping.
83
+ - It always runs, even with `--assumeDifferentTarget`, but only in a real run, not in a dry run.
84
+ - The target user needs permission to create and drop collections, which the copy needs anyway.
85
+ - If a run is killed at exactly that moment, the probe collection can stay behind. The next run warns about it, and you can drop it.
86
+ - **A URI doesn't name a database, or names an invalid one.** Without a database name, the driver would silently use `test`.
87
+ - **A collection named in the options doesn't exist in the source.** This covers `--collectionList`, `--ignoreCollections` and `collection.field` rules. It prevents, for example, a typo in `--collectionList` combined with `--copyNonAnonymized` from copying data unanonymized.
88
+ - **No collection would be anonymized** and everything selected would only be copied as-is. Views don't count as anonymized collections. If everything is skipped, the run writes nothing and warns.
89
+ - **The target already has one of the collections** and `--dropTarget` isn't set.
90
+ - **A replacement can't be applied.** Examples: invalid JSON, an unknown Faker method, or a Faker method that needs arguments.
91
+
92
+ ## Field rules
93
+
94
+ Each item of `--fieldList` has the form `[collection.]field[:replacement]`.
95
+
96
+ - **`field`** is matched against keys **at any depth**, including inside subdocuments and arrays, ignoring case, `_` and `-`. `email` matches `email`, `Email`, `e_mail`, `profile.email` and `contacts[].email`; `firstname` matches `firstName` and `first_name`.
97
+ - **`*`** in `field` matches any characters: `*email` matches `email`, `orderEmail` and `guestEmail`, but not `emailTemplate`.
98
+ - **`collection.`** limits the rule to one collection (`users.email`). The collection must exist in the source.
99
+ - **Precedence:** a collection's own rules, exact names or patterns, come before global rules; within each, an exact field name comes before a pattern, and a rule you add wins over a default for the same field (`+first_name:REDACTED`). So `users.*name:REDACTED` also overrides the global `name` in `users`.
100
+ - **`:replacement`** sets what the value becomes:
101
+ - `faker.<category>.<method>`, e.g. `faker.person.jobTitle`: any [Faker](https://fakerjs.dev/api/) method that can be called without arguments. It is still deterministic, and its result is used whatever the original type, so `phone:faker.string.uuid` turns numbers and dates into UUID strings too. As with any rule, `null`, booleans, empty strings, `NaN` and `Infinity` are still kept.
102
+ - `null`, `[]`, `{}`, or JSON such as `{"a":1}`. Encode commas as `%2C` inside JSON, since `,` separates rules.
103
+ - `keep` leaves the field as it is. Use it with a collection to override a broader rule: `images.name:keep` keeps file names in `images`, while `name` is still anonymized everywhere else. Everything inside a kept field is kept too, including subdocuments such as `address.city`; only email addresses are still replaced (`--scrubEmails`). The dry run still reports personal-looking keys inside a kept field.
104
+ - `keep` and `null` are reserved words and must be written in lower case; `KEEP` or `Null` is rejected as a likely typo.
105
+ - Any other text is used literally, e.g. `REDACTED`.
106
+
107
+ Without a replacement, the fake value is chosen from the field name: emails, first, last and full names, usernames, addresses, streets, cities, countries, zip codes, phone numbers, dates, birthdates, company names, IP addresses, IBANs, passwords and tokens (random strings), and free text (`description`, `comment`, `note`, `message`, `body`, ...). A value that is an email address always gets a fake email, whatever the field is called, so `recipient` and `users.email` stay consistent.
108
+
109
+ ### Default fields
110
+
111
+ The defaults are chosen so that a first run on a typical application database is already safe. The patterns end with the personal word, so `orderEmail` and `mainGuest.phoneNo` are covered, while `emailTemplate`, `emailVerified` and `productName` are not.
112
+
113
+ | Kind | Rules |
114
+ | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
115
+ | Emails | `*email`, `*emails`, `recipient`, `recipients`, `sender` |
116
+ | Names | `name`, `firstname`, `lastname`, `middlename`, `fullname`, `surname`, `displayname`, `nickname`, `username` |
117
+ | Phones | `*phone`, `*phones`, `*phoneno`, `*phonenumber`, `*mobile`, `fax` |
118
+ | Addresses | `*address`, `street`, `city`, `country`, `*zip`, `*zipcode`, `*postcode`, `*postalcode` |
119
+ | Dates of birth | `birthdate`, `birthday`, `dateofbirth`, `dob` |
120
+ | Free text | `description`, `comment`, `comments`, `note`, `notes` |
121
+ | Identifiers and credentials | `ip`, `ssn`, `iban`, `passport`, `*password`, `password*`, `*passwordhash`, `*token`, `*tokenhash`, `*secret`, `*secretkey`, `*apikey`, `*privatekey`, `salt` |
122
+
123
+ Items can adjust the defaults instead of replacing them:
124
+
125
+ | `--fieldList` | Result |
126
+ | --------------------------- | -------------------------------------------- |
127
+ | _(not set)_ | the default fields |
128
+ | `email,ssn` | only `email` and `ssn` |
129
+ | `+taxId,+users.apiKey:null` | the defaults, plus `taxId` and `apiKey` |
130
+ | `-description,-comment` | the defaults, without those two |
131
+ | `-*token` | the defaults, without the `*token` rule |
132
+ | `+images.name:keep` | the defaults, but `name` is kept in `images` |
133
+
134
+ A removal must name an existing rule exactly, including its `*`: `-*email` removes the default `*email` rule, while `-email` is an error.
135
+
136
+ ### How values are replaced
137
+
138
+ - **Types are kept.** Strings stay strings, numbers stay numbers with the same number of digits, and a `Date` stays a `Date`. `null`, booleans, empty strings, `NaN` and `Infinity` are kept as they are.
139
+ - **Arrays**, including nested arrays, are anonymized element by element.
140
+ - **A matched subdocument** has every value inside it anonymized. For example, with `address` matched, `address.street`, `address.city` and `address.zip` are each replaced with a fitting fake value. A nested key with a rule of its own follows that rule: with `address` and `city:REDACTED`, `address.city` becomes `REDACTED`.
141
+ - **`_id`** is never matched as a whole. Personal data inside a compound `_id`, such as `{ _id: { email } }`, is anonymized like any other subdocument. With `--scrubEmails`, an `_id` that is an email address is replaced as well. Because values are deterministic, the same `_id` still becomes the same new `_id` everywhere, including in the fields that reference it.
142
+ - **Email addresses in any string value** are replaced by default, even in fields no rule matches and inside free text such as `"Write to jane@x.com"`.
143
+ - The address becomes the same fake email it gets everywhere else, and the text around it is kept.
144
+ - A value such as `Jane Doe <jane@x.com>` gets a fake name and that same fake email. This holds in any field, such as `to`, `cc` or `headers.From`. A list such as `Jane <jane@x.com>, bob@y.com` is handled item by item, and `mailto:` is kept.
145
+ - Addresses that are part of a URL or a connection string are left alone: `git@github.com:org/repo.git`, `https://user@host/...`, `mongodb://user:pass@host`. An address followed by a colon and a space, a tag or punctuation, as in `jane@x.com: urgent` or `jane@x.com:<br>`, is still replaced; one followed by a colon and text, as in `jane@x.com:8080` or `jane@x.com:notes`, is taken for a host and left alone.
146
+ - Domains are recognised in ASCII only, so an address such as `jane@münchen.de` is not replaced.
147
+ - Email addresses used as object keys are not replaced.
148
+ - Turn this off with `--no-scrubEmails`.
149
+ - **Emails** always get the reserved `example.com` domain, so they never reach a real inbox. They also get a suffix derived from the original value, which makes collisions on unique indexes practically impossible.
150
+ - **Unsupported values** are left unchanged, with a warning. This covers BSON types that have no sensible fake (ObjectId, Binary, Decimal128, ...) and invalid dates. Use `--strict` to fail instead, or a replacement such as `field:null` to overwrite them.
151
+
152
+ ### Known over-matches
153
+
154
+ The default patterns can also match keys that aren't personal data:
155
+
156
+ - `*token` matches pagination tokens such as `nextToken`;
157
+ - `*address` matches `macAddress` or `serverAddress`;
158
+ - `*phone` matches `microphone`;
159
+ - `name` matches the names of roles, categories or products, which can break lookups by name.
160
+
161
+ - `password*` matches `passwordPolicy` (whose settings would be replaced) and `passwordUpdatedAt` (a date that becomes a random date);
162
+ - `salt` matches recipe or nutrition data.
163
+
164
+ Use `:keep` for a collection (`+roles.name:keep`), or remove a pattern (`-*token`).
26
165
 
27
- - `--sourceUri`: The MongoDB URI of the source database.
28
- - `--targetUri`: The MongoDB URI of the target database.
29
- - `--fieldList`: A comma-separated list of fields to anonymize. You can use the `+` or `-` modifiers to add or remove fields from the default list respectively. For example, `+age` will add `age` to the default fields, and `-email` will remove `email` from the default fields.
30
- - `--collectionList`: (Optional) A comma-separated list of collections to anonymize. If not provided, all collections will be anonymized.
31
- - `--ignoreCollections`: (Optional) A comma-separated list of collections to ignore during the anonymization process.
32
- - `--batchSize`: (Optional) The number of documents to process at a time. Defaults to `1000`.
33
- - `--copyNonAnonymized`: (Optional) If set, non-anonymized collections will be copied as-is to the target database. By default, non-anonymized collections are not copied.
166
+ Inside a matched subdocument, some structure is preserved:
34
167
 
35
- ## Note
168
+ - `type`, `kind` and `__typename` values are kept;
169
+ - GeoJSON objects keep their type. A `Point` gets a valid fake position; other geometries (a GPS track, a home area...) become a small valid shape of the same type around a fake position. The object's other keys, such as `formattedAddress`, are anonymized as usual. So `2dsphere` indexes still build;
170
+ - a legacy `[longitude, latitude]` pair under a geo key (`coordinates`, `loc`, `location`, `geo...`, `position`, `point`) gets a valid fake position, so `2d` and `2dsphere` indexes still build. Pairs of numbers under other keys are anonymized number by number;
171
+ - `lat`/`lng` values stay within valid ranges. A legacy `{lat, lng}` object stores latitude first, which a `2dsphere` index reads the wrong way round; a fake longitude there can make the index fail to build.
36
172
 
37
- This package does not delete any data from the source database. It creates a copy of the data in the target database with the specified fields anonymized. Please ensure that you have the necessary permissions and storage capacity before using this package.
173
+ ## Deterministic anonymization
174
+
175
+ Every fake value is derived from `HMAC-SHA256(secret, original value)`. As a result:
176
+
177
+ - **Consistent references:** the same email in `users.email` and in `orders.customer.email` becomes the same fake email, so joins and lookups still work.
178
+ - **Reproducible output:** running again with the same secret and the same version of this package produces identical data. Generated dates are relative to a fixed reference date, not to today. The Faker version is pinned for this reason; a new major version of this package may change the fake values.
179
+ - **Not reversible:** without the secret, an original value can't be recovered by hashing guesses. Keep the secret private, like a password.
180
+
181
+ Matching is exact. `John@x.com` and `john@x.com` are different values and get different fake values.
182
+
183
+ If no secret is given, a random one is generated for the run. The output is then consistent within that run, but differs from run to run.
184
+
185
+ ## Before using on production data
186
+
187
+ The defaults and email scrubbing cover the usual cases, but every database has its own field names. Before handing an anonymized copy to anyone:
188
+
189
+ 1. **Set a secret** and keep it private: `export ANONYMIZER_SECRET="$(openssl rand -hex 32)"`.
190
+ 2. **Run `--dryRun --sampleSize 0`** and read every warning. Add the keys it lists to `--fieldList`, with `+` or with a pattern such as `+*guestPhone`.
191
+ - The defaults also replace business content that happens to use the same field names, such as a product's `name` or `description` or a venue's `address`. Keep what your developers need with collection rules such as `+products.name:keep,+venues.address:keep`.
192
+ 3. **Think about data the dry run can't recognise:**
193
+ - free text that mentions names or phone numbers: add its field, for example `+contact.message`;
194
+ - fields named in another language;
195
+ - personal data stored in numbers or IDs.
196
+ 4. **Run with `--strict`**, so a matched value that can't be anonymized stops the run instead of producing a warning.
197
+ - When a rule matches a field that holds references, such as `sender: ObjectId(...)` in a messages collection, remove the rule (`-sender`) or keep the field for that collection (`+messages.sender:keep`). Replacing it with `null` would break the reference.
198
+ 5. **Spot-check the result.** Open a few documents of the collections that hold customer data, and search the target for a real email address or phone number you know is in the source.
199
+
200
+ ## Replica sets and large databases
201
+
202
+ **Progress.**
203
+
204
+ - Every 5 seconds, the run logs how far each collection has got, for example `users: 120,000/2,300,000 (5%), 9,800 docs/s, ETA 3m42s`.
205
+ - If nothing completes for 30 seconds, it warns and names the operation it is waiting for, for example `No progress for 30s: createIndexes on bookings (...)`.
206
+
207
+ **Index builds on a replica set.**
208
+
209
+ - On a replica set, MongoDB normally waits for every voting member to finish an index build.
210
+ - If a secondary is down, lagging or stuck, the build waits forever. With an arbiter, even `majority` can't be reached without the secondary.
211
+ - So the tool builds indexes with `--indexCommitQuorum 1` by default: the primary's vote is enough, and the secondaries still build the index as they catch up.
212
+ - If the build still doesn't finish within `--indexTimeout` seconds, the server aborts it. The collection's indexes are then skipped with a warning; build them by hand once the replica set is healthy.
213
+
214
+ **Write concern** comes from the target URI. For a throwaway target on a replica set whose default is `majority`, adding `w=1` to the target URI makes inserts faster.
215
+
216
+ **Open files.**
217
+
218
+ - MongoDB keeps one file per collection and per index open. Dropped collections keep theirs until the next checkpoint.
219
+ - So several `--dropTarget` reruns in a row can briefly need a few hundred extra file handles.
220
+ - The target's `mongod` should run with the open-files limit MongoDB recommends (`ulimit -n` 64000 or more). The default of 1024 in many containers can make WiredTiger fail with `Too many open files`.
221
+ - In Docker Compose:
222
+
223
+ ```yaml
224
+ services:
225
+ mongodb:
226
+ ulimits:
227
+ nofile:
228
+ soft: 64000
229
+ hard: 64000
230
+ ```
231
+
232
+ ## What else is copied
233
+
234
+ For every collection it writes, the tool also copies:
235
+
236
+ - the collection options: validators, collation, capped and time series settings;
237
+ - all indexes, including unique ones. Indexes are built after the data is inserted; an index that can't be built produces a warning instead of failing the run;
238
+ - views, recreated after the collections.
239
+
240
+ System collections (`system.*`) are skipped.
241
+
242
+ Validators are copied as they are, so they may reject fake values. For example, a `$jsonSchema` pattern that requires your company's email domain rejects the `example.com` addresses. The run then stops with `Document failed validation`. To avoid this, give the field a replacement the validator accepts, or relax the validator in the target.
243
+
244
+ ## Upgrading from 0.2
245
+
246
+ - **Node.js 22.13 or newer** is required, and the package is ESM-only.
247
+ - **Reruns need `--dropTarget`.** Rerunning into a target that already has the collections used to fail halfway through with duplicate key errors. Now the run refuses to start, unless `--dropTarget` is set.
248
+ - **`--copyNonAnonymized` no longer copies everything raw.** In 0.2, using it without `--collectionList` copied every collection without anonymizing anything. Now it only affects collections that are not selected for anonymization.
249
+ - **Fields are matched at any depth**, including inside subdocuments and arrays. Dates, numbers, `null` and arrays keep their types; in 0.2 they were turned into `{}`.
250
+ - **Fake values are deterministic.** Set `ANONYMIZER_SECRET` to get the same output on every run.
251
+ - **`+` and `-` apply to each item** of `--fieldList`, so `+age,-*email` works as expected. Removing a rule that doesn't exist is an error.
252
+ - **Wider defaults.** The default fields now use patterns (`*email`, `*phone`, `*address`, ...) and cover more kinds of data. Email addresses in any string are replaced too (`--scrubEmails`). To remove a default, use its exact rule, for example `-*email` instead of `-email`.
253
+ - **Stricter checks:** a URI must name a database, and collection names used in options must exist in the source. See [Safety checks](#safety-checks).
254
+ - **Logs are plain text on stderr**, no longer bunyan JSON. Failures now exit with a non-zero code.
255
+
256
+ ## Programmatic use
257
+
258
+ ```ts
259
+ import { run } from 'mongo-data-anonymizer';
260
+
261
+ const report = await run({
262
+ sourceUri: 'mongodb://localhost:27017/production',
263
+ targetUri: 'mongodb://localhost:27017/staging',
264
+ fieldList: ['email', 'name', 'users.password:null'],
265
+ collectionList: [],
266
+ ignoreCollections: [],
267
+ batchSize: 1000,
268
+ copyNonAnonymized: false,
269
+ dropTarget: true,
270
+ dryRun: false,
271
+ strict: true,
272
+ secret: process.env.ANONYMIZER_SECRET,
273
+ });
274
+
275
+ for (const warning of report.warnings) console.warn(warning);
276
+ ```
277
+
278
+ ## Development
279
+
280
+ ```bash
281
+ npm ci
282
+ npm test # unit + integration tests; integration tests start an in-memory MongoDB
283
+ npm run lint
284
+ npm run typecheck
285
+ npm run build
286
+ npm run dev -- --help # runs src/cli.ts directly, without building
287
+ ```
@@ -1,9 +1,51 @@
1
- export declare class Anonymize {
2
- anonymizeBatch(batch: any[], collectionName: string, list: string[]): any[];
3
- private getKeysToAnonymize;
4
- private anonymizeDocument;
5
- private anonymizeValue;
6
- private applyReplacement;
7
- private getFakerValue;
8
- private getFakerValueForField;
1
+ import { CollectionRules, type FieldRule } from './rules.ts';
2
+ /** Replacement that keeps a field unchanged, e.g. `images.name:keep` to override a global `name` rule. */
3
+ export declare const KEEP = "keep";
4
+ export interface AnonymizerOptions {
5
+ /**
6
+ * Secret used to derive fake values. The same secret always maps the same
7
+ * original value to the same fake value (across collections and runs), and
8
+ * without it the mapping cannot be reversed by hashing guesses.
9
+ * A random secret is used when none is given.
10
+ */
11
+ secret?: string;
12
+ /**
13
+ * Throw instead of warning when a matched value can't be anonymized
14
+ * (ObjectId, Binary, Decimal128, invalid dates, ...), so no personal data
15
+ * is left behind silently.
16
+ */
17
+ strict?: boolean;
18
+ /**
19
+ * Replace email addresses found in any string, including fields no rule
20
+ * matches and free text, with the same fake email they get everywhere else.
21
+ * On by default.
22
+ */
23
+ scrubEmails?: boolean;
24
+ }
25
+ export declare class Anonymizer {
26
+ #private;
27
+ constructor(options?: AnonymizerOptions);
28
+ onWarning(handler: (message: string) => void): this;
29
+ /**
30
+ * Throws for rules whose replacement can never be applied, before any data
31
+ * is written. Faker replacements are called once, since some methods only
32
+ * fail when called without arguments.
33
+ */
34
+ validateRules(rules: Iterable<FieldRule>): void;
35
+ /**
36
+ * Anonymizes every document in the batch. `rules` are the rules that apply
37
+ * to the batch's collection, keyed by lower-cased field name.
38
+ */
39
+ anonymizeBatch<T>(batch: T[], rules: CollectionRules): T[];
40
+ /**
41
+ * Anonymizes one document. `_id` is never matched as a whole, but values
42
+ * inside a compound `_id` are anonymized like any other subdocument.
43
+ */
44
+ anonymizeDocument<T>(document: T, rules: CollectionRules): T;
45
+ /**
46
+ * Dotted paths of keys that no rule covers but that look like personal
47
+ * data, either by name (`profile.phoneNumber`) or because their text
48
+ * contains an email address (`notes`, `message`). Array indexes are left out.
49
+ */
50
+ findUnmatchedPersonalKeys(document: unknown, rules: CollectionRules): Set<string>;
9
51
  }