exwiw 1.1.6 → 1.1.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/README.md CHANGED
@@ -1,25 +1,8 @@
1
1
  # Exwiw
2
2
 
3
- Export What I Want (Exwiw) is a Ruby gem that allows you to export records from a database to a dump file(to specifically, the full list of INSERT sql) on the specified conditions.
3
+ Export What I Want (exwiw) exports part of a database as SQL `INSERT` files: the rows related to the records you name, with sensitive columns masked. It is meant for building a development database that looks like production, without copying all of production or maintaining hand-made seed data.
4
4
 
5
- ## When to use
6
-
7
- Most of case in developing a software, There is no better choice than the same data in production.
8
- You might make well-crafted data, but it's very very hard to maintain.
9
-
10
- If you find the way to maintain the data for develoment env, then exwiw might be a solution for that.
11
-
12
- - Export the full database and mask data and import to another database.
13
- - Setup some system to replicate and mask data in real-time to another database.
14
-
15
-
16
- You want to export only the data you want to export.
17
-
18
- ## Features
19
-
20
- - Export the full list of INSERT sql for the specified conditions.
21
- - Provide serveral masking options for sensitive columns.
22
- - Provide config generator for ActiveRecord, for Mongoid, and from a live database connection (any application, any language).
5
+ Each table is described in a JSON schema config: its columns, how to mask them, and its `belongs_to` relations. Given a target table and ids, exwiw follows those relations to decide which rows of every other table to export. The schema config can be generated from ActiveRecord or Mongoid models, or from a live database.
23
6
 
24
7
  ## Installation
25
8
 
@@ -27,46 +10,29 @@ You want to export only the data you want to export.
27
10
  bundle add exwiw
28
11
  ```
29
12
 
30
- Most of cases, you want to add 'require: false' to the Gemfile.
31
-
32
- If bundler is not being used to manage dependencies, install the gem by executing:
33
-
34
- ```bash
35
- gem install exwiw
36
- ```
37
-
38
- ## Supported Databases
13
+ You usually want `require: false` on the Gemfile entry. Without bundler, run `gem install exwiw`.
39
14
 
40
- - mysql
41
- - postgresql
42
- - sqlite
43
- - mongodb (see [MongoDB support](docs/mongodb.md))
15
+ ## Supported databases
44
16
 
45
- For MySQL, exwiw connects through whichever of the `mysql2` or `trilogy` gem is
46
- available (preferring `mysql2`), so an app on either driver works without any
47
- extra setup. There is no separate `trilogy` adapter name — pass `--adapter=mysql`
48
- either way.
17
+ - MySQL
18
+ - PostgreSQL
19
+ - SQLite
20
+ - MongoDB (see [MongoDB support](docs/mongodb.md))
49
21
 
50
- Set `EXWIW_MYSQL_DRIVER=trilogy` (or `mysql2`) to force a specific driver. This
51
- is useful when the `mysql2` gem is linked against a `libmysqlclient` that can no
52
- longer load the server's auth plugin — e.g. a MySQL 9.x client drops the
53
- `mysql_native_password` plugin and raises `Authentication plugin
54
- 'mysql_native_password' cannot be loaded` on connect. The pure-Ruby `trilogy`
55
- driver implements that auth handshake itself and sidesteps the issue.
22
+ For MySQL, exwiw uses the `mysql2` gem if it is available and `trilogy` otherwise; pass `--adapter=mysql` either way. Set `EXWIW_MYSQL_DRIVER=trilogy` (or `mysql2`) to choose one. `trilogy` helps when `mysql2` is linked against a client library that cannot load the server's auth plugin, such as a MySQL 9.x client connecting to a server that uses `mysql_native_password`.
56
23
 
57
24
  ## Usage
58
25
 
59
26
  exwiw has three subcommands:
60
27
 
61
- - `export` (default) — generate INSERT/COPY SQL files. If the subcommand is omitted, `export` is assumed.
62
- - `explain` — print each query `export` would run together with its `EXPLAIN` output. SQL adapters compile the SELECT without executing it; mongodb runs the server's explain (defaulting to the execution-free `queryPlanner`).
63
- - `schema generate|check|tidy --from-db` — maintain the schema config by reading a live database, for applications that cannot be loaded to generate it from their models. See [Non-Rails applications](#non-rails-applications-exwiw-schema----from-db).
28
+ - `export` (the default): write the dump files.
29
+ - `explain`: print the queries `export` would run, with their `EXPLAIN` output.
30
+ - `schema generate|check|tidy --from-db`: maintain the schema config from a live database. See [Non-Rails applications](#non-rails-applications-exwiw-schema----from-db).
64
31
 
65
32
  ### `exwiw export`
66
33
 
67
34
  ```bash
68
- # dump & masking all records from database to dump.sql based on schema.json
69
- # pass database password as an environment variable 'DATABASE_PASSWORD'
35
+ # The database password is read from DATABASE_PASSWORD.
70
36
  exwiw \
71
37
  --adapter=mysql \
72
38
  --host=localhost \
@@ -75,62 +41,42 @@ exwiw \
75
41
  --database=app_production \
76
42
  --schema-dir=exwiw/schema \
77
43
  --target-table=shops \
78
- --ids=1 \ # comma separated ids
79
- --output-dir=dump \
80
- --log-level=info
44
+ --ids=1,2 \
45
+ --output-dir=dump
81
46
  ```
82
47
 
83
- By default `--ids` are matched against the target table's primary key. If the target table declares a per-table `scope_column`, exwiw runs in [scope-column mode](#scope-column-mode) instead — `--ids` are then values of that shared column, and the table is scoped like any other rather than anchored by primary key. In scope-column mode `--target-table` can be omitted: passing only `--ids` runs in scope-column mode as long as some table in the schema declares a `scope_column`.
48
+ This exports the `shops` rows with id 1 and 2 and the rows of other tables related to them. `--schema-dir` reads every JSON file in the directory.
84
49
 
85
- | `--target-table` | `--ids` | Mode |
50
+ | `--target-table` | `--ids` | What is exported |
86
51
  |---|---|---|
87
- | given | given | Scope-column mode if the table declares a `scope_column`, single-target mode otherwise |
88
- | omitted | given | Scope-column mode if any table declares a `scope_column`, an error otherwise (SQL adapters only) |
89
- | omitted | omitted | Every table is dumped in full |
52
+ | given | given | The target rows by primary key, and their related rows. If the table declares a `scope_column`, [scope-column mode](#scope-column-mode) is used instead |
53
+ | omitted | given | [Scope-column mode](#scope-column-mode). An error if no table declares a `scope_column` (SQL adapters only) |
54
+ | omitted | omitted | Every table in full |
90
55
 
91
- When `--target-table` and `--ids` are omitted, exwiw dumps all tables defined in `--schema-dir`:
56
+ The output directory (`dump/` by default) is emptied before each run. When it already has files and stdin is a terminal, exwiw asks before removing them.
92
57
 
93
- ```bash
94
- # dump all tables
95
- exwiw \
96
- --adapter=postgresql \
97
- --host=localhost \
98
- --port=5432 \
99
- --user=reader \
100
- --database=app_production \
101
- --schema-dir=exwiw/schema \
102
- --output-dir=dump
103
- ```
58
+ The output files are:
104
59
 
105
- This command will generate sql files in the `dump` directory.
60
+ - `insert-000-schema.sql`: `CREATE TABLE IF NOT EXISTS ...` for every table. Run it first to create an empty database.
61
+ - `insert-{idx}-{table}.sql`: one per table. A file may depend on files with a smaller `idx`, so import them in order.
106
62
 
107
- The output dir is emptied before each export so it never mixes files from a previous run (defaulting to `dump/` when `--output-dir` is omitted). When run interactively (stdin is a tty) and the dir already contains files, exwiw asks for confirmation before removing them; in non-interactive contexts (CI, pipes) it proceeds without prompting.
63
+ exwiw writes no `DELETE` statements. Import into an empty database, or clear the target's rows yourself.
108
64
 
109
- - `dump/insert-000-schema.sql` — idempotent `CREATE TABLE IF NOT EXISTS ...` for every table in scope. Apply this first to provision an empty database.
110
- - `dump/insert-{idx}-{table_name}.sql`
65
+ ### Restoring the dump
111
66
 
112
- idx means the order of the dump. bigger idx might depend on smaller idx,
113
- so you should import the dump in order.
67
+ `insert-000-schema.sql` is created with the database's own tools (`mysqldump`, `pg_dump`, or the sqlite3 driver), so `mysqldump` or `pg_dump` must be on `PATH`. Set `EXWIW_MYSQLDUMP` to use a specific `mysqldump`, for example an 8.0 one when a 9.x `mysqldump` cannot authenticate against the server.
114
68
 
115
- exwiw generates INSERT statements only — it does not generate DELETE statements. Import into an empty database (`insert-000-schema.sql` provisions one), or clear the target's rows yourself before importing.
69
+ The schema file is rewritten so that running it again is harmless (`IF NOT EXISTS`, and PostgreSQL constraints and triggers that are skipped if they already exist). For MySQL, `DEFINER` clauses are removed so that a managed MySQL instance accepts the views and triggers.
116
70
 
117
- `insert-000-schema.sql` is generated by shelling out to the database client tools (`mysqldump` for `mysql`, `pg_dump` for `postgresql`, and the sqlite3 driver for `sqlite`), so the corresponding client must be available on PATH when running exwiw. For `mysql`, set `EXWIW_MYSQLDUMP` to point at a specific `mysqldump` binary when the one on PATH is incompatible with the server (e.g. a MySQL 9.x `mysqldump` cannot load `mysql_native_password` against a server still using that auth plugin — `EXWIW_MYSQLDUMP=/path/to/mysql@8.0/bin/mysqldump`). The output is post-processed to make it idempotent: `CREATE TABLE IF NOT EXISTS`, `CREATE INDEX IF NOT EXISTS` (where the engine supports it), and PostgreSQL's `ALTER TABLE ... ADD CONSTRAINT` and `CREATE TRIGGER` statements are wrapped in `DO $$ ... EXCEPTION WHEN duplicate_object`. For `mysql`, the source server's `DEFINER=user@host` stamp on views and triggers is stripped too, so restoring into a managed MySQL instance (which usually can't grant the privilege to recreate someone else's `DEFINER`) does not fail.
71
+ MySQL data files turn off foreign key checks (`FOREIGN_KEY_CHECKS=0`). PostgreSQL data files set `session_replication_role = 'replica'`, which turns off both foreign key checks and triggers. This setting needs superuser (`rds_superuser` on RDS); without it a `WARNING` is printed and triggers fire. The setting stays on for the rest of the connection, so run `SET session_replication_role = 'origin'` if you keep using that connection. SQLite loads with its triggers active.
118
72
 
119
- The schema file carries the source's triggers. They are suppressed while the target is loaded: each `insert-NNN-<table>.sql` opens with a block that sets `session_replication_role = 'replica'` for the connection, which turns off both user triggers and foreign-key enforcement for the statements that follow (the PostgreSQL counterpart of the `FOREIGN_KEY_CHECKS=0` `mysql` dumps already carry). Setting it requires superuser (`rds_superuser` on RDS); if the restoring role lacks the privilege the block reports a `WARNING` and the load proceeds with triggers firing. The setting applies to the connection and is not reset at the end of each file — every file re-arms it itself, so `psql -f` per file and `cat insert-*.sql | psql` both work, but a session that sources these files and then goes on to do other work stays in replica mode; reset it yourself (`SET session_replication_role = 'origin'`) in that case. `sqlite` has no equivalent and loads with its triggers active.
120
-
121
- For `postgresql`, the extensions a managed platform installs to run the source instance itself are treated as out of target and left out of the dump entirely — currently `google_vacuum_mgmt` (Cloud SQL / AlloyDB adaptive autovacuum), `google_columnar_engine` and `google_db_advisor` (AlloyDB). They serve the source instance's operation (vacuum tuning, the in-memory columnar cache, index advice), hold no application data, are referenced by nothing in the application's own schema, and ship only with the managed platform, so a restore target outside it can never create them. Their schemas are dropped via `pg_dump --exclude-schema` and their `CREATE EXTENSION` / `COMMENT ON EXTENSION` statements — which are not schema-qualified, so no `pg_dump` filter reaches them — are removed from the output; whatever was excluded is named in the run's log.
122
-
123
- The list is exact names, not a `google_*` prefix match: those prefixes are not reserved, so a prefix rule would also drop a schema an application legitimately owns (`google_calendar` for a Google Calendar integration) together with its tables. Every other extension is kept and wrapped in the usual warn-and-skip `DO` block, including two kinds that are also managed-platform-only:
124
-
125
- - a third-party extension pulled in as a dependency of an excluded one (`google_db_advisor` requires `hypopg`), since that one *is* installable on a plain PostgreSQL, and
126
- - an application-facing platform extension (`google_ml_integration`, `alloydb_scann`, `alloydb_ai_nl`), which the application's own SQL and DDL can name (a ScaNN index is `USING scann`) — removing its `CREATE` would strand whatever refers to it, so it warns and skips instead.
73
+ On PostgreSQL, extensions that only exist to run a managed instance (`google_vacuum_mgmt`, `google_columnar_engine`, `google_db_advisor`) are left out of the dump, because a database outside that platform cannot create them. Other extensions are kept, and are skipped with a warning when the target cannot create them.
127
74
 
128
75
  ### `exwiw explain`
129
76
 
130
- Print the query each `export` would run together with its `EXPLAIN` output, to stdout. For the SQL adapters (`mysql`, `postgresql`, `sqlite`) this is the compiled SELECT plus its `EXPLAIN` (estimate-only; `EXPLAIN QUERY PLAN` on SQLite) — no SELECT is executed. For `mongodb` it is the `find` description plus the server's explain document as JSON.
77
+ Prints the query `export` would run for each table, with its `EXPLAIN` output. For the SQL adapters the SELECT is not executed. For MongoDB, see [`exwiw explain` verbosity](docs/mongodb.md#exwiw-explain-verbosity).
131
78
 
132
79
  ```bash
133
- # preview the queries exwiw would run, without executing the SELECTs
134
80
  exwiw explain \
135
81
  --adapter=postgresql \
136
82
  --host=localhost --port=5432 --user=reader \
@@ -139,515 +85,45 @@ exwiw explain \
139
85
  --target-table=shops --ids=1
140
86
  ```
141
87
 
142
- The `--output-dir`, `--output-format`, and `--after-insert-hook` options are dump-specific and rejected when used with `explain`.
143
-
144
- MongoDB-specific explain behavior — the configurable verbosity (`queryPlanner` / `executionStats` / `allPlansExecution`) and how scoped collections are shown — is described in [MongoDB support](docs/mongodb.md#exwiw-explain-verbosity).
145
-
146
- ### How each table is narrowed — the six scoping paths
147
-
148
- Only the dump target itself is filtered by `--ids` directly. Every *other* table must be **scoped** — narrowed to just the rows related to the target — some other way, and a table that cannot be scoped at all is dumped in full (or, in scope-column mode, aborts the run). exwiw resolves each table through the **first** of these six paths that applies:
149
-
150
- | # | Path | When it applies | Resulting query shape |
151
- |---|------|-----------------|-----------------------|
152
- | 1 | **Direct filter** | The table is the `--target-table` itself; or, in [scope-column mode](#scope-column-mode), it declares a `scope_column` | `WHERE pk IN (ids)` / `WHERE scope_column IN (ids)` |
153
- | 2 | **`belongs_to` join walk** | The table reaches the target (or a scope-column table) by following its `belongs_to` edges | `WHERE fk IN (ids)` for a single hop; a chain of `JOIN`s for longer paths |
154
- | 3 | **Referenced-by (automatic reverse)** | No `belongs_to` path of its own, but **exactly one** already-constrained table points at it by foreign key | Constrained to the ids that referencer's own query selects |
155
- | 4 | **`reverse_scope` (declared reverse)** | Referenced by **many** scoped tables — typically a global-identity table like `users` — and the referencers are enumerated in its config | Constrained to the `UNION` of the enumerated referencers' ids |
156
- | 5 | **Scoped-parent cascade** | No path or referencer, but a `belongs_to` parent is itself scoped (by any path above) | Constrained to the parent's in-scope primary keys; cascades over multiple hops |
157
- | 6 | **Full dump** | Nothing relates the table to the target | All rows. In scope-column mode this **aborts** unless the table opts in with `scope_exempt: true` |
158
-
159
- How the paths behave and interact:
160
-
161
- 1. **Direct filter.** In the default single-target mode the target is anchored on its primary key (or a custom field via the mongodb-only `--ids-field`). In [scope-column mode](#scope-column-mode) there is no single anchor: every table that declares a `scope_column` is filtered on that column directly.
162
- 2. **`belongs_to` join walk** — the "normal join" path. exwiw BFS-walks `belongs_to` edges to the nearest terminus (the target table, or a directly scoped table in scope-column mode) and compiles the shortest path into `INNER JOIN`s. A [polymorphic `belongs_to`](#polymorphic-belongs_to) hop additionally pins the type column, and is resolved for **every** concrete arm with the arms `UNION`ed (see [Every arm is extracted](#every-arm-is-extracted)).
163
- 3. **Referenced-by** handles a table with no outgoing path that is pointed *at* by a constrained child — `active_storage_blobs`, referenced by `active_storage_attachments.blob_id`, is the canonical case (see [ActiveStorage](#activestorage-has_one_attached--has_many_attached)). It is automatic but deliberately narrow: it requires a single, non-polymorphic referencer. With two or more referencers it steps aside to path 5, then 6, unless you declare `reverse_scope`.
164
- 4. **[`reverse_scope`](#reverse-scope-for-multi-referencer-tables-reverse_scope)** is the declared, multi-referencer form of path 3: the config enumerates which referencers' (already scoped) queries feed the id set. Unscoped arms are skipped with a warning rather than widening the dump.
165
- 5. **Scoped-parent cascade** rescues satellites: a table whose only link is a `belongs_to` toward a hub that is itself scoped (e.g. via referenced-by or `reverse_scope`) is constrained to that parent's in-scope ids. The cascade recurses hop by hop (each level requires a single unambiguous scopable parent) and stops on `belongs_to` cycles.
166
- 6. **Full dump** is the fallback for a genuinely unrelated table — intended for reference/master data. Single-target mode dumps it in full (with a warning when an ambiguous cascade was the reason); scope-column mode refuses to run instead, unless the table is explicitly marked [`scope_exempt: true`](#scope_exempt-intentional-full-dump) (Rails-managed tables are exempt automatically).
167
-
168
- Paths 3–5 all materialize their id set once and probe it via a `JOIN` on a `SELECT DISTINCT` derived table rather than `IN (subquery)` — see [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery). Scope-column mode classifies every table up front with these same paths (`:direct` / `:via_path` / `:referenced_by` / `:via_scoped_parent` / `:exempt` / `:unscopable` in `QueryAstBuilder#scope_category`) and aborts before extracting anything if any table lands on `:unscopable`. The MongoDB adapter follows the same model, except id sets are captured at runtime while parent collections stream instead of being expressed as SQL subqueries — see [MongoDB support](docs/mongodb.md).
169
-
170
- ### Scope-column mode
171
-
172
- The default `--target-table` extraction assumes the schema converges on a single
173
- root: every table is reached by walking `belongs_to` toward that one table. Some
174
- schemas are not shaped that way — many independent top-level tables each carry the
175
- *same* scope/tenant column (e.g. `tenant_id`, `business_entity_id`), and a foreign
176
- key that **cannot be joined** (most importantly a cross-database `belongs_to`,
177
- whose join is impossible but whose FK column is still filterable) is not reached at
178
- all. Choosing one table as `--target-table` would leave the others unrelated to it,
179
- and an unrelated table is dumped in full — a problem if it holds personal data.
180
-
181
- Scope-column mode handles this shape: instead of anchoring on one table's primary
182
- key, **every table is filtered by a shared column** whose values are `--ids` (tables
183
- keyed by an unrelated kind of id can use their own [ID space](#per-table-scope_column-and-id-spaces)).
184
- Declare that column per table in the schema config with `scope_column:`:
185
-
186
- ```json
187
- {
188
- "name": "shops",
189
- "primary_key": "id",
190
- "scope_column": "business_entity_id",
191
- "columns": [{ "name": "id" }, { "name": "name" }, { "name": "business_entity_id" }]
192
- }
193
- ```
194
-
195
- Then pass the scope values as `--ids`, without `--target-table`:
196
-
197
- ```bash
198
- exwiw \
199
- --adapter=postgresql \
200
- --host=localhost --port=5432 --user=reader \
201
- --database=app_production \
202
- --schema-dir=exwiw/schema \
203
- --ids=42,43 \
204
- --output-dir=dump
205
- ```
206
-
207
- Because the schema declares a `scope_column`, exwiw runs in scope-column mode: the
208
- `--ids` (`42,43`) are **`business_entity_id` values, not shop primary keys**, and
209
- `shops` is scoped by `business_entity_id IN (42,43)` like every other scoped table.
210
- Where extraction starts is decided by the schema, so `--target-table` is not
211
- needed. If no table declares a `scope_column`, `--ids` alone is an error: name the
212
- table the ids belong to with `--target-table` (single-target mode), or declare a
213
- `scope_column`. Tables marked `ignore: true` do not count.
214
-
215
- Naming a scoped table as `--target-table` (`--target-table=shops --ids=42,43`)
216
- selects the same mode and extracts the same rows; the target is *not* used as a
217
- primary-key anchor. (A table that declares a `scope_column` therefore can no longer
218
- be single-extracted by primary key.)
219
-
220
- Each table is resolved as follows:
221
-
222
- - **Declares the scope column** (`scope_column:`, or carries the global column of
223
- the deprecated `--scope-column` flag) → `WHERE scope_column IN (ids)`.
224
- - **Does not, but `belongs_to` reaches a table that does** → exwiw joins up to the
225
- nearest such table and applies the scope filter there (the same join machinery
226
- the single-target mode uses).
227
- - **`belongs_to` a parent that is itself scoped but carries no scope column of its
228
- own** → exwiw constrains this table to the parent's in-scope ids by joining it to
229
- the parent's scoped query, materialized as a derived table
230
- (`JOIN (SELECT DISTINCT parent.pk … FROM <parent's scoped query>) … ON fk = …`).
231
- This covers a *hub* table that has no scope column and is scoped only because an
232
- extractable child references it (see referenced-by below): the hub's other
233
- `belongs_to` children ride along to just the in-scope rows instead of being dumped
234
- in full. The parent itself may be scoped the same way, so this **cascades across
235
- multiple hops** (each a single unambiguous scopable parent) and the derived-table
236
- JOINs nest correspondingly; the recursion terminates on a genuine `belongs_to`
237
- cycle (a table already on the path is left `:unscopable` rather than looped on).
238
- (See [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery) for the
239
- materialization rationale.)
240
- - **Cannot be scoped at all** (no scope column and no path to one) → exwiw
241
- **aborts** and lists the offending tables, so an unscoped table is never silently
242
- dumped in full. For each, either declare a `scope_column`, add a `belongs_to`
243
- path, set `ignore: true` to skip it, or mark it `scope_exempt: true` (below) to
244
- export it in full.
245
-
246
- > **Note — referenced-by is preferred over the hub cascade.** A table that is
247
- > *both* `belongs_to` a scoped hub *and* referenced-by a constrained child is
248
- > scoped to the (narrower) referenced-by id-set, not the hub cascade, so the hub's
249
- > other children the child does not reference are dropped (under-scoping). To force
250
- > the broader hub cascade, set `ignore: true` on the child's `belongs_to` edge that
251
- > points at this table.
252
-
253
- Scope-column mode is SQL-only (mysql / postgresql / sqlite); with the mongodb
254
- adapter, `--ids` still requires `--target-collection`. It works with `exwiw
255
- explain` too, which is the recommended way to preview the queries before exporting.
256
-
257
- #### Cross-database foreign keys
258
-
259
- The motivating case for declaring a `scope_column` is a foreign key that cannot be
260
- joined: when a `belongs_to` target lives in a different database (see the
261
- cross-database `belongs_to` note under the generator), that join is impossible, but
262
- the foreign-key *column* is still present and can be filtered directly. Declaring
263
- `scope_column: "<that foreign key>"` on the owning table scopes it by the column
264
- value, with no join — `schema:generate` points this out in the ignored relation's
265
- `comment`.
266
-
267
- #### `scope_exempt` (intentional full dump)
268
-
269
- A genuine reference/master table (no personal data) that has no scope linkage can
270
- opt out of the strict check and be exported in full:
271
-
272
- ```json
273
- {
274
- "name": "countries",
275
- "primary_key": "id",
276
- "scope_exempt": true,
277
- "columns": [{ "name": "id" }, { "name": "code" }]
278
- }
279
- ```
280
-
281
- Rails-managed tables (`schema_migrations`, `ar_internal_metadata`) are treated as
282
- exempt automatically.
283
-
284
- #### Per-table `scope_column` and ID spaces
285
-
286
- The values of every `scope_column` belong to an **ID space**, and `--ids` gives the
287
- values of each space. A table that does not declare `id_space` is in the `default`
288
- space, which a plain `--ids=1,2` fills, so without any `id_space` every scoped table
289
- is filtered by the same `--ids`. Each table names its own column, so a table that
290
- stores that same value under a differently named column simply declares that name:
291
-
292
- ```json
293
- {
294
- "name": "legacy_orders",
295
- "primary_key": "id",
296
- "scope_column": "legacy_tenant_id",
297
- "columns": [{ "name": "id" }, { "name": "legacy_tenant_id" }]
298
- }
299
- ```
300
-
301
- When one database holds two groups of tables that no foreign key connects, each
302
- keyed by a different kind of id, put one group in a named ID space. Here `tenants`
303
- are identified by integer ids and `organizations` by UUIDs (each object is its own
304
- schema file):
305
-
306
- ```json
307
- { "name": "tenants", "primary_key": "id", "scope_column": "id", "columns": [{ "name": "id" }] }
308
- { "name": "organizations", "primary_key": "id", "scope_column": "id", "id_space": "org", "columns": [{ "name": "id" }] }
309
- ```
310
-
311
- ```bash
312
- exwiw ... --target-table=tenants \
313
- --ids=1,2 \
314
- --ids=org=0b6f4c1e-0000-4000-8000-000000000001
315
- ```
316
-
317
- `--ids=1,2` is the same as `--ids=default=1,2`. The part before the first `=` is
318
- read as an ID space name only when it is shaped like one (a lowercase letter
319
- followed by lowercase letters, digits or underscores), so an id that contains `=`
320
- itself is passed with an explicit `default=`. Each ID space may be given once. A
321
- table without a scope column is filtered by the ID space of the table it is scoped
322
- through (a `belongs_to` join, `reverse_scope`, referenced-by or the parent cascade).
323
-
324
- The run aborts before extracting anything when a table is filtered by an ID space
325
- that was given no values, when values are given for an ID space no table uses, or
326
- when a table reaches scoped tables of more than one ID space (otherwise the
327
- `belongs_to` walk would silently settle on the nearest one).
328
-
329
- Single `--target-table` mode and the MongoDB adapter use only the `default` space
330
- and abort when a named one is given. In the config file, `ids:` takes either a list
331
- (the `default` space) or a mapping from ID space name to values, such as
332
- `ids: { default: [1, 2], org: [0b6f4c1e-0000-4000-8000-000000000001] }`; as with
333
- the other keys, it is used only when `--ids` is not passed at all.
334
-
335
- `scope_exempt`, `scope_column` and `id_space` are user-maintained and preserved
336
- across `schema:generate` regeneration (the generators never emit them).
337
-
338
- #### Deprecated: the `--scope-column` flag
339
-
340
- Before per-table declarations, scope-column mode was selected with a global
341
- `--scope-column=COLUMN` flag (every table filtered by that one column, `--ids` its
342
- values, no `--target-table`). The flag still works — SQL-only and mutually
343
- exclusive with `--target-table` — but is **deprecated** and emits a warning; prefer
344
- declaring a per-table `scope_column` and dropping the flag. A per-table
345
- `scope_column` takes precedence over the flag for any table that sets both.
88
+ `--output-dir`, `--output-format` and `--after-insert-hook` cannot be used with `explain`.
346
89
 
347
90
  ### Config file (`exwiw.yml`)
348
91
 
349
- Options you would otherwise repeat on every run can be kept in a YAML config file. Pass it with `--config=PATH`; when `--config` is omitted, exwiw automatically loads `exwiw.yml` (or `exwiw.yaml`) from the current directory if present.
350
-
351
- **Options passed on the CLI always take precedence over the config file** — the config only fills in options you did not pass. This lets you commit the stable settings (which schema to read, output format, ...) while still varying the environment-specific connection details per invocation.
92
+ Options can be kept in a YAML file passed with `--config=PATH`. Without `--config`, `exwiw.yml` (or `exwiw.yaml`) in the current directory is used if it exists. Options passed on the command line take precedence.
352
93
 
353
94
  ```yaml
354
- # exwiw.yml — keep at the project root, alongside exwiw/schema/
355
95
  adapter: postgresql
356
96
  schema_dir: exwiw/schema
357
97
  output_dir: dump
358
98
  output_format: insert # insert | copy
359
99
  after_insert_hook: hooks/seed.rb
360
100
  log_level: info # debug | info
361
- # target_table / ids / ids_field / scope_column may also be set here
362
- # (ids may map ID spaces to values; see "Per-table scope_column and ID spaces")
363
- # mongodb_query_timeout_ms: 30000 # global query timeout (mongodb only)
101
+ # target_table, ids, ids_field and scope_column can also be set here.
364
102
  ```
365
103
 
366
- With the file above, only the connection details need to be supplied on the CLI:
367
-
368
104
  ```bash
369
105
  DATABASE_PASSWORD=... exwiw \
370
106
  --host=localhost --port=5432 --user=reader --database=app_production \
371
107
  --target-table=shops --ids=1
372
108
  ```
373
109
 
374
- Notes:
110
+ - Connection settings (`host`, `port`, `user`, `database`, `uri`, `password`) are rejected, so they stay out of a committed file. `adapter` is allowed.
111
+ - Relative paths are resolved from the config file's directory, not the current directory.
112
+ - Unknown keys are rejected. Keys that only apply to `export` are ignored by `explain` and `schema`, so all subcommands can share one file.
113
+ - MongoDB-only keys (`explain_verbosity`, `mongodb_query_timeout_ms`, `parallel_workers`) are described in [MongoDB support](docs/mongodb.md).
375
114
 
376
- - **Database connection settings stay on the CLI/environment.** `host`, `port`, `user`, `database`, `uri`, and `password` are **rejected** in the config file (exwiw exits with an error). `adapter` is the one connection-related key that *is* allowed in the file.
377
- - **Relative paths in the config (`schema_dir`, `output_dir`, `after_insert_hook`) are resolved relative to the config file's own directory**, not the current working directory. So with the config at the project root, `schema_dir: exwiw/schema` reads naturally, and an absolute `--config=/path/to/exwiw.yml` works no matter where you run from. (CLI path flags remain relative to the current directory — each source resolves relative to where it is written.) Absolute paths are used as-is.
378
- - Unknown keys are rejected so a typo surfaces immediately. (`insert_only`, whose behavior was removed, is the one grandfathered key: accepted and ignored with a warning.)
379
- - Export-only keys (`output_dir`, `output_format`, `after_insert_hook`) are ignored when running `explain` or `schema`, so a single config file can be shared by every subcommand.
380
- - `explain_verbosity` sets the mongodb `explain` verbosity (`queryPlanner` | `executionStats` | `allPlansExecution`, default `queryPlanner`); the `EXWIW_MONGODB_EXPLAIN_VERBOSITY` env var overrides it. Ignored by the SQL adapters and by `export`. See [MongoDB support](docs/mongodb.md#exwiw-explain-verbosity).
381
- - `mongodb_query_timeout_ms` sets the global, server-enforced query timeout (mongodb only); the `--mongodb-query-timeout-ms` CLI flag overrides it. Ignored by the SQL adapters. See [MongoDB support](docs/mongodb.md).
382
-
383
- ### Generator
384
-
385
- The config generator is provided as a Rake task.
386
-
387
- ```bash
388
- # generate table schema under exwiw/schema/
389
- bundle exec rake exwiw:schema:generate
390
- ```
391
-
392
- The output directory is resolved in this order:
393
-
394
- 1. the `EXWIW_SCHEMA_DIR_PATH` environment variable, if set;
395
- 2. otherwise `schema_dir` from the config file (`exwiw.yml` / `exwiw.yaml` in the current directory), so the generator and the `exwiw` CLI share one location without repeating the path;
396
- 3. otherwise the `exwiw/schema` default.
397
-
398
- ```sh
399
- EXWIW_SCHEMA_DIR_PATH=custom_directory bundle exec rake exwiw:schema:generate
400
- ```
401
-
402
- As with the CLI, a relative `schema_dir` in the config file is resolved relative to the config file's own directory.
403
-
404
- An application with more than one schema source — ActiveRecord models and Mongoid documents, say —
405
- should give each source its own directory (`EXWIW_SCHEMA_DIR_PATH`), because every task judges the
406
- whole directory against the one source it reads. Left sharing a directory, `tidy` / `check` see the
407
- other source's configs as belonging to tables and collections that no longer exist, and report or
408
- remove them.
409
-
410
- #### Safe mode (masking new columns by default)
411
-
412
- A migration that adds a column would otherwise leave `schema:generate` emitting it unmasked, so
413
- it starts being exported the moment the config is regenerated — before anyone has judged whether
414
- it holds personal data. So `schema:generate` runs in **safe mode by default**: every column the
415
- config does not have yet is emitted **masked** and flagged
416
- [`needs_mask_decision: true`](#needs_mask_decision).
417
-
418
- Columns already in the config keep whatever they say — the merge that preserves `replace_with` /
419
- `comment` / `ignore` preserves a resolved decision too — so in practice this marks exactly the
420
- columns a migration just added.
421
-
422
- ```bash
423
- bundle exec rake exwiw:schema:generate # safe mode
424
- EXWIW_NEW_COLUMNS=plain bundle exec rake exwiw:schema:generate # opt out
425
- ```
426
-
427
- Opting out is for the **first-time bootstrap** of a config, where every column of every table is
428
- new and safe mode would flag the whole thing at once. Use it nowhere else: a column committed
429
- under `plain` carries no flag, so nothing afterwards can tell it apart from one whose masking was
430
- decided.
431
-
432
- A column that has a **default of its own** is masked with that default: it is a value the column
433
- provably holds, and it is what the application treats as neutral, so masking a `default: true`
434
- flag does not quietly turn the feature off for every row in the dump. A default the database
435
- computes (`now()`) is not a constant and does not count, and neither does a JSON object — `{...}`
436
- in a mask is a column placeholder, so those fall back to `{}`. Otherwise the mask depends on the column
437
- type: `masked-{primary key}` for text (with `@example.com` appended when the column name mentions
438
- mail, so it stays a valid address), `0` for numbers, `false` for booleans, a fixed date/timestamp,
439
- and `{}` for JSON. Text always takes the template rather than its default, since the mask has to
440
- vary per row. Three kinds of
441
- column are flagged but deliberately **not** masked:
442
-
443
- - **The primary key, and the foreign keys/types the `belongs_tos` join on.** Masking them
444
- would break the joins and leave the dump referencing rows that were never exported.
445
- - **Types no constant safely fits** — `uuid`, `binary`, enums, array columns (which report their
446
- member type, so a scalar default would not fit), and text columns too short to hold the masked
447
- value. An invalid default would fail the restore the dump feeds, which is worse than exporting
448
- the column while the flag keeps the change from being merged.
449
- - **Columns covered by a unique index**, unless the mask varies per row (the text masks do, via
450
- the primary key). A constant would collapse every row onto one value and break the restore with
451
- a duplicate key.
452
-
453
- `schema:generate_mongoid` runs in safe mode too, on the same `EXWIW_NEW_COLUMNS=plain` opt-out. The
454
- masks come from the Mongoid field type (`String`, `Integer`, `Float`, `BigDecimal`,
455
- `Mongoid::Boolean`, `Date`, `Time` / `DateTime` / `ActiveSupport::TimeWithZone`); a field of any
456
- other type — `Hash`, `Array`, a typeless field, a BSON type — is flagged but not masked, as is a
457
- field covered by a unique index unless its mask varies per document. The structural fields are the
458
- `_id` primary key, the STI discriminator (`_type`) and every `belongs_to` foreign key of the
459
- collection: flagged, never masked. The foreign keys are read from the models, so the ones the
460
- config itself drops (a polymorphic `belongs_to`, a `belongs_to` on an embedded document) are
461
- covered too.
462
-
463
- #### Tidying stale config (`schema:tidy`)
464
-
465
- `schema:generate` adds and updates config files for the tables it finds, but it never deletes the config file of a table that has been dropped from the application. To reconcile the existing config against the current schema, run:
466
-
467
- ```bash
468
- bundle exec rake exwiw:schema:tidy
469
- ```
470
-
471
- `schema:tidy` compares the config files already on disk with the **live database** (read through the database connection, not the models) and removes only what no longer exists there:
472
-
473
- - a config file whose table has been dropped from the database is **deleted**, and
474
- - columns recorded in a surviving table's config that the table no longer has are **dropped** from that file.
475
-
476
- Because it reads the database directly, a table that still exists in the database but has lost (or never had) an ActiveRecord model is **kept** — only a table that is genuinely gone is removed. (This is the deliberate counterpart to `generate`, which is model-driven and only ever adds what the models know about.)
477
-
478
- It respects `EXWIW_SCHEMA_DIR_PATH` and the per-database subdirectory layout in the same way as `schema:generate`. Unlike `generate`, `tidy` never adds or regenerates entries — every surviving table/column (including hand-edited `comment` / `ignore` / `replace_with`) is left untouched, so it is safe to run on a customized config. The task prints which tables and columns it removed (or that the config was already tidy). Stale `belongs_tos` are not pruned by `tidy`; rerun `schema:generate` to refresh those.
479
-
480
- #### Checking the config against the schema
481
-
482
- `schema:check` reports how the committed config differs from what the application would
483
- generate now — without writing anything, so it can run on a working tree it must not modify:
484
-
485
- ```bash
486
- bundle exec rake exwiw:schema:check
487
- ```
488
-
489
- It regenerates into a throwaway copy of the config directory (safe mode + `tidy`) and prints
490
- the comparison as JSON, then exits non-zero when anything needs attention:
491
-
492
- ```json
493
- {
494
- "added_tables": [],
495
- "added_columns": ["users.contact_email"],
496
- "removed_tables": [],
497
- "removed_columns": ["orders.legacy_flag"],
498
- "changed_tables": ["orders", "users"],
499
- "needs_mask_decision": ["orders.memo"],
500
- "stale_tables": [],
501
- "stale_columns": ["orders.legacy_flag"]
502
- }
503
- ```
504
-
505
- `added_*` / `removed_*` / `changed_tables` mean the config no longer matches the schema — run
506
- `schema:generate` and `schema:tidy` to reconcile it. `needs_mask_decision` lists the columns
507
- whose masking nobody has decided on yet (see [the flag](#needs_mask_decision)). `stale_tables` /
508
- `stale_columns` are the subset of the removals an extraction would actually trip over — a
509
- non-ignored config still naming a table or column the schema no longer has, so the export's
510
- SELECT would fail; removals of `ignore: true` entries (and of a rails-managed table's columns,
511
- which are dumped as `SELECT *`) stay out of them. They drive the exit code only under
512
- [`--fail-on=stale`](#non-rails-applications-exwiw-schema----from-db). The exit code
513
- makes it usable as a CI check that keeps a schema change from being merged until both are
514
- resolved; the JSON is stable and sorted, so it can be posted as-is. In a multi-database app each
515
- entry is prefixed with its database (`primary/users.email`), so the same table name in two
516
- databases stays distinct.
517
-
518
- Set `EXWIW_SCHEMA_CHECK_OUTPUT=<path>` to have the same JSON written to a file, which spares a
519
- caller from assuming stdout carries nothing else (application boot is free to print).
520
-
521
- A Mongoid config directory has its own task, `schema:check_mongoid`, with the same output,
522
- the same `EXWIW_SCHEMA_CHECK_OUTPUT` file and the same exit code — it just regenerates through
523
- `MongoidSchemaGenerator` (safe mode + `tidy_mongoid`) instead. Collections and fields are
524
- reported under the same keys as tables and columns. An application that cannot be loaded to
525
- generate from its models at all can run the same check against its database instead: see
526
- [Non-Rails applications](#non-rails-applications-exwiw-schema----from-db).
527
-
528
- #### Multiple databases
529
-
530
- If the application uses Rails' multiple-database support (`connects_to`), `schema:generate` buckets models by the database they connect to and writes each database's config files into its own subdirectory of the output directory, named after the database config name (`primary`, `analytics`, ...):
531
-
532
- ```
533
- exwiw/schema/
534
- primary/
535
- shops.json
536
- users.json
537
- schema_migrations.json
538
- analytics/
539
- analytics_events.json
540
- ```
541
-
542
- Each database keeps its own Rails migration history, so a `schema_migrations` (and `ar_internal_metadata`) entry is emitted under every database that contains one — the example above shows `primary/schema_migrations.json` and would also produce `analytics/schema_migrations.json` when the analytics database has its own migration table. Single-database applications are unaffected and continue to write files flat into the output directory.
543
-
544
- A `belongs_to` whose target model lives in a *different* database (e.g. a `primary` model referencing an `analytics` one) cannot be joined: each database is exported on its own connection and into its own subdirectory, so the target table is absent from the directory this config is loaded with. `schema:generate` detects such a relation (by comparing the owning and target models' database config names) and emits it with `ignore: true` and `ignore_type: "cross_database"`, recording why in the `comment`; the relation is then dropped from extraction at load time, while the foreign-key column itself is still exported as a plain column. Polymorphic associations are handled per target, so only the targets that cross a database boundary are ignored. The task also prints a summary of every cross-database `belongs_to` it ignored. **To extract across such a boundary, declare `scope_column: "<foreign_key>"` on the owning table (see [scope-column mode](#scope-column-mode)) so its rows are filtered by the foreign-key value directly** — there is no join, so the cross-database boundary is not a problem there.
545
-
546
- **Limitations**
547
-
548
- - The rails-managed table *names* are resolved from the global `ActiveRecord::Base.schema_migrations_table_name` / `internal_metadata_table_name` accessors, which are shared across all connections. A per-database override of these names is not detected, so such a table will be missing from that database's generated configs.
549
-
550
- #### Mongoid applications
551
-
552
- For MongoDB applications backed by [Mongoid](https://www.mongodb.com/docs/mongoid/), a separate rake task introspects Mongoid document models and emits `MongodbCollectionConfig` files:
553
-
554
- ```bash
555
- bundle exec rake exwiw:schema:generate_mongoid
556
- bundle exec rake exwiw:schema:tidy_mongoid # delete the config of a collection no model stores into
557
- bundle exec rake exwiw:schema:check_mongoid # report the difference without changing anything
558
- ```
559
-
560
- What it derives from each model (fields, `belongs_tos`, `embedded_in`, STI handling), how to annotate constructs exwiw cannot represent with `ignore` / `ignore_type`, and the `EXWIW_SKIP_UNSUPPORTED=1` bootstrap flag are all documented in [MongoDB support](docs/mongodb.md#generating-config-from-mongoid-models).
561
-
562
- `tidy_mongoid` reconciles against the *models* rather than a live connection (MongoDB has no schema to read, and a collection exists only once something is written to it): a config file whose collection no model stores into any more is deleted, and nothing else is touched — fields already track the models through `generate_mongoid`. `check_mongoid` runs both into a throwaway copy, so it never writes to the working tree.
563
-
564
- #### Non-Rails applications (`exwiw schema ... --from-db`)
565
-
566
- The rake tasks above read the application's models, which requires loading the application — so
567
- they are only available where that is possible. For an application written in any other language
568
- (or a Ruby one exwiw cannot boot), the same three operations are available on the CLI, reading
569
- the **database** instead of the models:
570
-
571
- ```bash
572
- # generate / refresh the config from the live schema
573
- exwiw schema generate --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
574
-
575
- # report how the committed config differs from the database (exits 1 when it needs work)
576
- exwiw schema check --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
577
-
578
- # remove tables/columns/relations the database no longer has
579
- exwiw schema tidy --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
580
- ```
581
-
582
- - `--from-db` is **required**: the schema source is always stated explicitly rather than inferred.
583
- - mysql and postgresql only. sqlite is not supported, and a MongoDB schema lives in the
584
- application rather than the database (use `schema:generate_mongoid`).
585
- - The usual connection flags and `DATABASE_PASSWORD` apply, and `adapter` / `schema_dir` may come
586
- from [the config file](#config-file-exwiwyml) instead (`--schema-dir` wins). Unlike `export`,
587
- an empty or absent `DATABASE_PASSWORD` is accepted, since these commands are commonly pointed
588
- at a CI database that runs with trust authentication.
589
- - `generate` creates the schema directory if it does not exist; `check` and `tidy` require it.
590
- - Everything else behaves as the rake tasks do: [safe mode](#safe-mode-masking-new-columns-by-default)
591
- is on unless `EXWIW_NEW_COLUMNS=plain`, `check` prints the same JSON report, honours
592
- `EXWIW_SCHEMA_CHECK_OUTPUT`, and exits 1 when the config needs attention. A check that could
593
- not *run* (an unreachable database, a malformed config) exits with a different status, so CI
594
- can tell the two apart.
595
- - `check` also accepts `--fail-on=stale` for use as a pre-extraction gate: the exit code then
596
- tracks only the report's `stale_tables` / `stale_columns` — a non-ignored config still naming
597
- a table or column the schema no longer has, which is exactly the drift that would fail the
598
- export's SELECT. Additions and unresolved `needs_mask_decision` flags stay visible in the
599
- report but do not stop the run, so a schema migration that merely *adds* a column does not
600
- block extraction. The default (`--fail-on=any`) is the CI behavior above, unchanged.
601
- `--fail-on` is command-line only (not a config-file key — a gate flag belongs to the
602
- invocation, not the committed config) and is meaningful only on `schema check`: the other
603
- schema verbs reject it, and `export` ignores it like the other schema-only flags.
604
- - One run covers one database — the connection addresses one — so the files are written flat into
605
- the schema directory. There is no per-database subdirectory layout here; a second database is a
606
- second run against a second connection.
607
-
608
- The database is read through its catalog only: tables, columns and their types/defaults, primary
609
- keys, unique indexes and foreign keys. Views are skipped, since they hold no rows of their own,
610
- and a table with no primary key is emitted with `ignore: true` and a comment saying what to add
611
- to export it (exwiw identifies and joins rows by primary key). Following that advice sticks: a
612
- `primary_key` written by hand — on a table the database reports none for, or one column of a
613
- composite key that identifies a row on its own — is kept by later runs, which then treat the table
614
- as an ordinary one rather than re-imposing the signpost `type` / `comment`.
615
-
616
- **`belongs_tos` are only ever added, never rewritten.** A foreign-key constraint is weaker
617
- evidence than an application model: plenty of schemas express a relation only in application code,
618
- and a `belongs_to` is the path extraction follows to reach a table, so silently dropping one
619
- narrows the dump. Regeneration therefore keeps every relation the config already declares —
620
- in its existing order, with its `comment` / `ignore` / `ignore_type` — and appends only the
621
- foreign-key-backed relations that are not there yet. A hand-written relation's foreign-key column
622
- is treated as structural too, so safe mode never masks it. Removing a relation is `tidy`'s job:
623
- it drops a `belongs_to` whose target table no longer exists in the database (an `ignore: true`
624
- entry is kept, since it records a decision), alongside the tables and columns that are gone.
625
-
626
- So a relation the database does not know about is declared once, by hand, and survives from then
627
- on:
115
+ ### Output format
628
116
 
629
- ```json
630
- {
631
- "name": "orders",
632
- "primary_key": "id",
633
- "belongs_tos": [{
634
- "table_name": "buyers",
635
- "foreign_key": "buyer_id",
636
- "comment": "enforced in the application; no foreign key in the database"
637
- }]
638
- }
639
- ```
117
+ For PostgreSQL, `--output-format=copy` writes `COPY ... FROM stdin` instead of `INSERT`, which loads much faster. Import it with `psql -d app_dev -f dump/insert-001-shops.sql`.
640
118
 
641
- ### Configuration
119
+ ## Schema config
642
120
 
643
- This is an example of the one table schema:
121
+ Each table has one JSON file:
644
122
 
645
123
  ```json
646
124
  {
647
125
  "name": "users",
648
126
  "primary_key": "id",
649
- "filter": "users.id > 0",
650
- "bulk_insert_chunk_size": 1000,
651
127
  "belongs_tos": [{
652
128
  "table_name": "companies",
653
129
  "foreign_key": "company_id"
@@ -663,97 +139,26 @@ This is an example of the one table schema:
663
139
  }
664
140
  ```
665
141
 
666
- `--schema-dir` will use all json files in the specified directory.
667
-
668
- #### Unknown keys are rejected
142
+ `belongs_tos` decides which rows are exported (see [How each table is narrowed](#how-each-table-is-narrowed)), and each column can be masked (see [Masking](#masking)). The other table-level keys are:
669
143
 
670
- Loading a table/collection config with a key that no declared attribute accepts is an **error** (`Exwiw::UnknownConfigKeyError`, an `ArgumentError` subclass) naming the key, the table/collection, the offending file, and the allowed keys. This also applies to the nested `belongs_tos` / `columns` / `fields` / `reverse_scope` / `embedded_in` / `replace_with_fake_data` entries. Previously such keys were silently dropped, which turned a typo (`reverse_scop`) — or a key another adapter supports but this one does not (e.g. `raw_sql` on a MongoDB field) — into a silent no-op: the config loaded, the dump ran, and the requested masking/scoping simply never happened.
144
+ - `filter`: an SQL condition added to the table's query, such as `"access_logs.created_at > '2025-01-01'"`. It is added to every query that joins this table, so it also narrows the tables that depend on it, which can leave their foreign keys pointing at rows that were not exported. Qualify column names with the table name. A filter reduces the rows returned, not necessarily the rows read; see [Batched extraction](#batched-extraction-batch_scope) for that.
145
+ - `bulk_insert_chunk_size`: the maximum rows per `INSERT` statement (10,000 by default), to stay under limits such as MySQL's `max_allowed_packet`.
146
+ - `ignore`, `comment`: see below.
671
147
 
672
- For free-form annotations, use the `comment` key — it is a declared, documentation-only attribute on table/collection configs and on their `belongs_tos` / `columns` / `fields` entries, so it always passes (see [Ignore / annotate a column or `belongs_to`](#ignore--annotate-a-column-or-belongs_to)).
673
-
674
- ### Output format
148
+ ### Unknown keys are rejected
675
149
 
676
- By default, exwiw generates `INSERT` statements. For PostgreSQL, you can pass `--output-format=copy` to generate `COPY FROM stdin` format instead, which is significantly faster for bulk loading.
677
-
678
- The generated file uses tab-separated values with PostgreSQL's text-format escaping (`\N` for NULL, `\\` for backslash, etc.). Import with `psql`:
679
-
680
- ```bash
681
- psql -d app_dev -f dump/insert-001-shops.sql
682
- ```
683
-
684
- `--output-format=copy` is only supported with the `postgresql` adapter.
685
-
686
- ### After-insert hook
687
-
688
- `--after-insert-hook=PATH` runs a post-processing hook **after** all per-table insert files have been written. The hook can be either a Ruby file (`.rb`) or any executable script (e.g. `.sh`).
689
-
690
- **Ruby hook (`.rb`)**: provides a tiny DSL with these builtins:
691
-
692
- - `cli_options` — Hash of all parsed CLI options (e.g. `cli_options.fetch(:ids)` returns the `--ids` array of the `default` ID space)
693
- - `ids_for(id_space = "default")` — the values the run was scoped by for that ID space (e.g. `ids_for("org")` for `--ids=org=...`)
694
- - `insert_sql(template)` — appends an ERB-rendered string to a buffer. After the hook finishes, the buffer is concatenated and written to `insert-{N+1}-after_insert.{ext}` where `{N+1}` is one past the last per-table insert file. For the MongoDB adapter the equivalent alias `insert_jsonl(template)` is available; output goes to `insert-{N+1}-after_insert.jsonl`. Multiple `insert_sql` calls in a single hook are joined with `"\n"` into the same file. If no `insert_sql` call is made, no file is created.
695
- - `insert_jsonl(collection, template)` — **MongoDB adapter only**. SQL statements name their table in-band, but JSONL documents do not — the import convention derives the target collection from the filename — so the two-argument form writes the ERB-rendered extended-JSON lines to the named collection's own `insert-NNN-<collection>.jsonl` file, importable with the same `mongoimport --collection <collection>` convention as the per-collection dump files. Multiple calls targeting the same collection are appended (joined with `"\n"`) into that collection's file; distinct collections get one file each, numbered sequentially after the last per-collection dump file (the collection-less `after_insert` buffer, when also used, keeps `{N+1}` and the collection files follow it). Calling this form with a SQL adapter raises an error.
696
-
697
- Example `hooks/seed_default_users.rb`:
698
-
699
- ```ruby
700
- insert_sql <<~SQL
701
- -- seed default users for tenants <%= cli_options.fetch(:ids).join(',') %>
702
- <%- cli_options.fetch(:ids).each do |tenant_id| -%>
703
- INSERT INTO users (tenant_id, email) VALUES (<%= tenant_id %>, 'default@example.com');
704
- <%- end -%>
705
- SQL
706
- ```
707
-
708
- MongoDB example seeding two collections (`insert-{N+1}-users.jsonl` and `insert-{N+2}-posts.jsonl`):
709
-
710
- ```ruby
711
- insert_jsonl 'users', <<~JSONL
712
- <%- cli_options.fetch(:ids).each do |shop_id| -%>
713
- {"shop_id":{"$oid":"<%= shop_id %>"},"email":"default@example.com"}
714
- <%- end -%>
715
- JSONL
716
- insert_jsonl 'posts', '{"title":"welcome"}'
717
- ```
718
-
719
- **Shell hook**: anything other than `.rb` is exec'd as a child process. It is a pure side-effect hook — exwiw does not capture its stdout. The hook receives these env vars and inherits `DATABASE_PASSWORD` from the parent:
720
-
721
- - `EXWIW_OUTPUT_DIR`, `EXWIW_SCHEMA_DIR`
722
- - `EXWIW_DATABASE_ADAPTER`, `EXWIW_DATABASE_HOST`, `EXWIW_DATABASE_PORT`, `EXWIW_DATABASE_USER`, `EXWIW_DATABASE_NAME`
723
- - `EXWIW_TARGET_TABLE`, `EXWIW_IDS` (comma-separated, the `default` ID space), `EXWIW_OUTPUT_FORMAT`
724
- - `EXWIW_IDS_<ID_SPACE>` for each ID space given values, with the name uppercased (`--ids=org=...` becomes `EXWIW_IDS_ORG`). `EXWIW_IDS_DEFAULT` is set only when the `default` space is given, while `EXWIW_IDS` is always set
725
-
726
- A non-zero exit code from the shell hook aborts exwiw.
727
-
728
- Note: Ruby hooks are evaluated via `instance_eval` inside the exwiw process — only pass paths you trust.
150
+ A config with a key exwiw does not know fails to load, with an error naming the key and the file. A typo such as `reverse_scop` would otherwise silently disable what it was meant to do. Use `comment` for notes.
729
151
 
730
152
  ### Ignore a table
731
153
 
732
- Set `"ignore": true` on a table's config JSON to exclude it from data extraction. The table's DDL is still emitted into `insert-000-schema.{sql,js}` so the schema stays consistent, but no `insert-*` files are generated for it and the table is never queried.
733
-
734
- ```json
735
- {
736
- "name": "audit_logs",
737
- "primary_key": "id",
738
- "ignore": true,
739
- "belongs_tos": [],
740
- "columns": [{ "name": "id" }]
741
- }
742
- ```
743
-
744
- Constraints:
154
+ `"ignore": true` on a table stops its data from being exported. Its `CREATE TABLE` is still written to `insert-000-schema.sql`.
745
155
 
746
- - If another non-ignored table has a `belongs_to` entry pointing at an ignored table, exwiw raises `ArgumentError` on load. Remove the `belongs_to` entry on the referencing table, or unset `ignore` on the referenced table.
747
- - Specifying an ignored table as `--target-table` raises `ArgumentError`.
748
- - `ignore: true` is preserved by `exwiw:schema:generate` regenerations (the receiver value wins over the auto-generated config).
749
- - On a MongoDB [embedded config](docs/mongodb.md#excluding-an-embedded-path-ignore-true), which is masked through its parent rather than dumped on its own, `ignore: true` excludes the embedded path from the parent's projection — the subdocument is left out of the dump entirely.
156
+ - A table that has a `belongs_to` to an ignored table fails to load. Remove that `belongs_to`, or ignore it too.
157
+ - An ignored table cannot be the `--target-table`.
750
158
 
751
159
  ### Ignore / annotate a column or `belongs_to`
752
160
 
753
- Individual `columns` (SQL) / `fields` (MongoDB) and `belongs_tos` entries accept two optional, **user-owned** keys:
754
-
755
- - `comment` — a free-form note. Purely informational; exwiw never reads it.
756
- - `ignore: true` — drops that entry from extraction. An ignored column/field is excluded from the `SELECT` and the generated `INSERT` (the column still exists in the target schema, since the DDL comes from the source database — exwiw just does not copy its data). An ignored `belongs_to` is removed from dependency ordering and query building, so the relation is not traversed.
161
+ Entries in `columns` and `belongs_tos` accept `comment` (a note exwiw never reads) and `ignore: true`. An ignored column is left out of the `SELECT` and the `INSERT`, so the restored rows get the column's database default, or `NULL` if it has none. On a `NOT NULL` column without a default, the `INSERT` can fail. An ignored `belongs_to` is not followed.
757
162
 
758
163
  ```json
759
164
  {
@@ -761,7 +166,7 @@ Individual `columns` (SQL) / `fields` (MongoDB) and `belongs_tos` entries accept
761
166
  "primary_key": "id",
762
167
  "belongs_tos": [
763
168
  { "table_name": "companies", "foreign_key": "company_id" },
764
- { "table_name": "audit_logs", "foreign_key": "log_id", "ignore": true, "comment": "huge table, not needed for this export" }
169
+ { "table_name": "audit_logs", "foreign_key": "log_id", "ignore": true, "comment": "not needed in development" }
765
170
  ],
766
171
  "columns": [
767
172
  { "name": "id" },
@@ -770,153 +175,70 @@ Individual `columns` (SQL) / `fields` (MongoDB) and `belongs_tos` entries accept
770
175
  }
771
176
  ```
772
177
 
773
- On a MongoDB [embedded config](docs/mongodb.md#embedded-documents) an ignored field cannot be left out of the query — the subdocuments arrive as part of whatever the parent collection fetched — so it is removed from each subdocument during masking instead. The result is the same: the field is absent from the dump. A MongoDB config marking its primary key `ignore: true` is rejected on load — the dump has to keep the identifier, or a restore would assign fresh ids and break every reference pointing at the documents.
178
+ ### Hand-edited keys survive regeneration
774
179
 
775
- The ignored entries are removed only at runtime, right after the config is loaded from file; the JSON on disk keeps them. Both `comment` and `ignore` are **preserved across `exwiw:schema:generate` / `exwiw:mongoid:schema:generate` regenerations** (the hand-edited value wins over the auto-generated config), just like `replace_with`. This applies to the MongoDB `MongodbCollectionConfig` (`fields` / `belongs_tos`) as well.
180
+ Regenerating the config (see [Generating the schema config](#generating-the-schema-config)) keeps what you wrote by hand. A column already in the config is kept as it is, `comment` and `ignore` on a table or `belongs_to` are kept, and so are the keys the generators never write (`filter`, `bulk_insert_chunk_size`, `scope_column`, `id_space`, `scope_exempt`, `reverse_scope`, `batch_scope`).
776
181
 
777
- ### `needs_mask_decision`
182
+ ## How each table is narrowed
778
183
 
779
- A column/field may also carry `needs_mask_decision: true`, marking a column whose masking
780
- nobody has decided on yet:
184
+ Only the target table is filtered by `--ids` directly. exwiw narrows every other table to the rows related to the target, using the first of these rules that applies:
781
185
 
782
- ```json
783
- { "name": "contact_email", "replace_with": "masked-{id}@example.com", "needs_mask_decision": true }
784
- ```
186
+ | # | Rule | When it applies | Result |
187
+ |---|------|-----------------|--------|
188
+ | 1 | Direct filter | The table is the `--target-table`, or declares a `scope_column` in [scope-column mode](#scope-column-mode) | `WHERE pk IN (ids)` or `WHERE scope_column IN (ids)` |
189
+ | 2 | `belongs_to` path | Following `belongs_to` reaches a table of rule 1 | Joined along the shortest path |
190
+ | 3 | Referenced by one table | No path, but exactly one narrowed table has a foreign key to it | Only the rows that table points at |
191
+ | 4 | [`reverse_scope`](#reverse-scope-for-multi-referencer-tables-reverse_scope) | Several narrowed tables point at it, and the config lists them | Only the rows the listed tables point at |
192
+ | 5 | Narrowed parent | No path, but a `belongs_to` parent is narrowed by a rule above | Only the rows belonging to the parent's exported rows, over any number of hops |
193
+ | 6 | Full dump | Nothing relates the table to the target | All rows, meant for master data. Scope-column mode stops with an error instead, unless the table has [`scope_exempt: true`](#scope_exempt-intentional-full-dump) |
785
194
 
786
- Extraction ignores the key entirely — what the column exports is whatever `replace_with` /
787
- `ignore` say. It exists so the decision can be tracked and required:
788
- [safe mode](#safe-mode-masking-new-columns-by-default) attaches it to every newly discovered
789
- column/field together with a default mask, and
790
- [`schema:check`](#checking-the-config-against-the-schema) — `schema:check_mongoid` for a Mongoid
791
- config — reports the ones that still carry it, so CI can keep a pull request red until each is
792
- resolved. Resolving it means removing the key — after keeping the mask (ideally recording why
793
- in `comment`), replacing it with a real masking rule, dropping `replace_with` to export the
794
- raw value, or setting `ignore: true`.
195
+ - Rule 3 covers tables like `active_storage_blobs`, which nothing in the table itself links to the target. It applies only when the referencing `belongs_to` is not polymorphic.
196
+ - Rule 5 applies only when exactly one parent is narrowed, and stops at a `belongs_to` cycle.
197
+ - If a table matches both rule 3 and rule 5, rule 3 wins and the result can miss rows that rule 5 would have kept; set `ignore: true` on the referencing `belongs_to` to use rule 5.
198
+ - Whichever rule narrows a table with a `belongs_to` to itself, the ancestors of its rows are added (see [Self-referencing `belongs_to`](#self-referencing-belongs_to-tree-tables)).
199
+ - MongoDB follows the same rules, collecting the ids while it reads the parent collections instead of using SQL subqueries. It does not add ancestors.
795
200
 
796
- Like `comment` / `ignore`, the on-disk state wins over regeneration: once removed,
797
- `schema:generate` does not bring it back.
201
+ When a rule picks the full dump because the relation was ambiguous, exwiw logs a warning. `exwiw explain` is the easiest way to see what each table resolved to. Rules 3 to 5 appear in it as a `JOIN` on a `SELECT DISTINCT` subquery; [`docs/scope-id-set-join-notes.md`](docs/scope-id-set-join-notes.md) explains why.
798
202
 
799
203
  ### Polymorphic `belongs_to`
800
204
 
801
- A Rails polymorphic association (`belongs_to :reviewable, polymorphic: true`) does not point at a single table — the target row is selected at runtime by a type column. exwiw models this as **one `belongs_to` entry per concrete target table**, each carrying two extra fields:
802
-
803
- - `foreign_type` — the type column on *this* table (e.g. `reviewable_type`).
804
- - `type_value` — the value stored in that column for this target (e.g. `"Product"`), i.e. the target model's `polymorphic_name`.
205
+ A polymorphic association (`belongs_to :reviewable, polymorphic: true`) is written as one `belongs_to` per target table, each with the type column (`foreign_type`) and the value it holds for that target (`type_value`):
805
206
 
806
207
  ```json
807
208
  {
808
209
  "name": "reviews",
809
210
  "primary_key": "id",
810
211
  "belongs_tos": [
811
- {
812
- "table_name": "products",
813
- "foreign_key": "reviewable_id",
814
- "foreign_type": "reviewable_type",
815
- "type_value": "Product"
816
- },
817
- {
818
- "table_name": "shops",
819
- "foreign_key": "reviewable_id",
820
- "foreign_type": "reviewable_type",
821
- "type_value": "Shop"
822
- }
212
+ { "table_name": "products", "foreign_key": "reviewable_id", "foreign_type": "reviewable_type", "type_value": "Product" },
213
+ { "table_name": "shops", "foreign_key": "reviewable_id", "foreign_type": "reviewable_type", "type_value": "Shop" }
823
214
  ],
824
215
  "columns": [{ "name": "id" }, { "name": "reviewable_type" }, { "name": "reviewable_id" }]
825
216
  }
826
217
  ```
827
218
 
828
- `exwiw:schema:generate` expands a polymorphic `belongs_to` automatically: it finds every model that registers the association as a target via `has_many` / `has_one ..., as: :reviewable` and emits one entry per target table (ordered by table name so the output is stable across Ruby versions). A plain (non-polymorphic) `belongs_to` simply omits `foreign_type` / `type_value`.
829
-
830
- At dump time, when a polymorphic `belongs_to` lies on the path to the dump target, exwiw constrains **both** the foreign key and the type column, so only rows of the matching type are extracted. For example, dumping `products` pulls only reviews whose `reviewable_type = 'Product'`:
219
+ `exwiw:schema:generate` writes these entries from the models' `has_many ..., as:` declarations. A non-polymorphic `belongs_to` leaves out `foreign_type` and `type_value`.
831
220
 
832
- ```sql
833
- SELECT reviews.* FROM reviews
834
- WHERE reviews.reviewable_id IN (/* products subquery */)
835
- AND reviews.reviewable_type = 'Product'
836
- ```
837
-
838
- The same type filter is applied on the join path when the polymorphic table is an intermediate hop rather than the directly-dumped table.
221
+ Following such a `belongs_to` also checks the type column, so dumping `products` exports only the reviews with `reviewable_type = 'Product'`.
839
222
 
840
223
  #### Every arm is extracted
841
224
 
842
- A polymorphic `belongs_to` is several `belongs_to` entries — one per concrete target — that a row selects between via its type column. A single JOIN can only follow **one** of them, so a join table reached through such a hop would come out holding only the rows of that one `type_value`. exwiw therefore resolves **every** arm of the group and constrains the table to the union of the ids the arms keep. In [scope-column mode](#scope-column-mode):
843
-
844
- ```sql
845
- SELECT comments.* FROM comments
846
- JOIN (
847
- SELECT DISTINCT exwiw_scope_src_0.id AS exwiw_scope_id
848
- FROM (
849
- SELECT comments.id FROM comments
850
- JOIN posts ON comments.commentable_id = posts.id AND comments.commentable_type = 'Post'
851
- JOIN shops ON posts.shop_id = shops.id AND shops.tenant_id = 't1'
852
- UNION
853
- SELECT comments.id FROM comments
854
- JOIN pages ON comments.commentable_id = pages.id AND comments.commentable_type = 'Page'
855
- JOIN shops ON pages.shop_id = shops.id AND shops.tenant_id = 't1'
856
- ) AS exwiw_scope_src_0
857
- ) AS exwiw_scope_ids_0
858
- ON comments.id = exwiw_scope_ids_0.exwiw_scope_id
859
- ```
860
-
861
- `UNION`, not `OR`, because each arm joins a *different* table: OR-ing them in one `WHERE` would need outer joins, whereas each arm is a self-contained query of exactly the shape a single-arm table already produces. It rides on the existing scope id-set machinery, so the id set is materialized once (see [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery)) and, on mysql, into a session `TEMPORARY TABLE`.
862
-
863
- Notes:
864
-
865
- - **Only arms that reach the scope are included.** An arm whose target has no scope of its own is dropped, never widened — an unscoped arm would pull in every tenant's rows. An arm marked `"ignore": true` is dropped as usual, before any of this.
866
- - An arm's target does not need a `belongs_to` path to a scoped table: if it is scoped by other means (referenced-by, [`reverse_scope`](#reverse-scope-for-multi-referencer-tables-reverse_scope), or the parent cascade) the arm probes that query's ids instead, still pinned by the type column.
867
- - An arm whose target is scoped **through this same table** (e.g. `active_storage_blobs`, narrowed by referenced-by from `active_storage_attachments`, appearing as an `ActiveStorage::Blob` arm of those same attachments) is dropped: adopting it would make the two tables scope each other and leave the referenced table short of rows the join table kept — a dangling foreign key on import.
868
- - **Arms are grouped by the type column (`foreign_type`), not by the foreign key.** Rails keeps every arm's id in one `<name>_id` column, which is what the generators emit, but a hand-written config may give each arm its own foreign key while one type column still selects between them — e.g. `owner_type` choosing between `user_id` and `team_id`. Those are arms of the same discriminator, so each is resolved and joined on its own key. Two *independent* polymorphic associations on one table have distinct type columns and therefore stay in distinct groups.
869
- - **The route the walk picks decides whether the arms are unioned.** The route is the shortest one; in single-target mode a `belongs_to` pointing at the dump target itself is taken first. When the route leaves through a polymorphic arm, every arm of that type column is unioned. When it leaves through a non-polymorphic `belongs_to`, the table is joined along that route alone and the polymorphic arms are not unioned in, so a row reachable only through an arm is not extracted. A plain `belongs_to` is not preferred over a shorter polymorphic route. A single arm is emitted as a plain JOIN.
870
- - A table with no route of its own falls back to the parent cascade, which does prefer plain parents. A scopable plain parent is used when there is exactly one. Only when there is none are the polymorphic `belongs_to`s consulted: the arms of one type column that point at tables scoped by other means (referenced-by, `reverse_scope`, or the cascade) are unioned, each pinned by the type column. Multiple independent polymorphic associations that are all scopable are as ambiguous as multiple plain parents, so the table is left unscopable (dumped in full with a warning in single-target mode). A table scoped this way also counts as a constrained child in the [referenced-by](#how-each-table-is-narrowed--the-six-scoping-paths) detection of the tables it points at, so a parent that had a single constrained referencer may now have two and be narrowed by the next path instead.
871
-
872
- In single `--target-table` mode the same union is built. An arm that points at the dump target itself compares its foreign key with `--ids`, and the other arms join or probe their way to the target as above. One difference: an arm target that is constrained **only** by the automatic referenced-by detection does not count. Such a parent is extracted just to keep a child's foreign key valid; it does not own the rows that point at it. So dumping `products` still pulls only `reviewable_type = 'Product'` reviews, even though the product's shop is extracted too, while dumping `shops` pulls the shop's own reviews **and** the reviews of its products:
873
-
874
- ```sql
875
- SELECT reviews.* FROM reviews
876
- JOIN (
877
- SELECT DISTINCT exwiw_scope_src_0.id AS exwiw_scope_id
878
- FROM (
879
- SELECT reviews.id FROM reviews
880
- JOIN products ON reviews.reviewable_id = products.id
881
- AND products.shop_id = 1 AND reviews.reviewable_type = 'Product'
882
- UNION
883
- SELECT reviews.id FROM reviews
884
- WHERE reviews.reviewable_id = 1 AND reviews.reviewable_type = 'Shop'
885
- ) AS exwiw_scope_src_0
886
- ) AS exwiw_scope_ids_0
887
- ON reviews.id = exwiw_scope_ids_0.exwiw_scope_id
888
- ```
225
+ When a table's path to the target goes through a polymorphic `belongs_to`, exwiw follows every entry with the same `foreign_type` and exports the union of their rows. Dumping `shops` exports the shop's own reviews and the reviews of its products.
889
226
 
890
- A declared [`reverse_scope`](#reverse-scope-for-multi-referencer-tables-reverse_scope) does count, since declaring it states that the referencers own the rows.
227
+ - An entry whose target table is not narrowed is skipped, so a polymorphic `belongs_to` never widens the dump.
228
+ - If the shortest path leaves through a non-polymorphic `belongs_to`, only that path is used.
229
+ - In single-target mode, an entry is skipped when its table is narrowed only by rule 3. Those rows are exported to keep foreign keys valid, not because they own the rows that point at them. So dumping `products` still exports only the `Product` reviews, even though the product's shop is exported too. An entry whose table has a `reverse_scope` is followed.
891
230
 
892
231
  ### ActiveStorage (`has_one_attached` / `has_many_attached`)
893
232
 
894
- ActiveStorage is handled automatically — no ActiveStorage-specific configuration is required. The `has_one_attached` / `has_many_attached` macros don't add a column to the owning model; they generate ordinary associations that exwiw already understands:
233
+ ActiveStorage needs no configuration:
895
234
 
896
- - **`active_storage_attachments`** is the polymorphic join row (`belongs_to :record, polymorphic: true` + `belongs_to :blob`). `exwiw:schema:generate` expands the polymorphic `record` into one `belongs_to` per model that declared `has_*_attached` (found via the generated `has_* ..., as: :record` reflections), exactly like any other [polymorphic `belongs_to`](#polymorphic-belongs_to). So only the attachments whose owner is among the dumped rows are extracted. Every owner type that reaches the scope is extracted (see [Every arm is extracted](#every-arm-is-extracted)); before that, only the single owner type the walk happened to settle on came out.
897
- - **`active_storage_blobs`** has no `belongs_to` of its own (attachments point *at* it), so it has no path to the dump target. exwiw narrows it via **reverse / "referenced_by" extraction**: a parent table referenced by exactly one constrained, non-polymorphic child is constrained to just the referenced ids instead of dumping every row. The id set is materialized once and joined back (see [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery)):
898
-
899
- ```sql
900
- SELECT active_storage_blobs.* FROM active_storage_blobs
901
- JOIN (
902
- SELECT DISTINCT exwiw_scope_src_0.blob_id AS exwiw_scope_id
903
- FROM (
904
- SELECT active_storage_attachments.blob_id FROM active_storage_attachments
905
- WHERE active_storage_attachments.record_id IN (/* owner subquery */)
906
- AND active_storage_attachments.record_type = '...'
907
- ) AS exwiw_scope_src_0
908
- ) AS exwiw_scope_ids_0
909
- ON active_storage_blobs.id = exwiw_scope_ids_0.exwiw_scope_id
910
- ```
911
-
912
- `active_storage_variant_records` also references blobs, but since it has no path of its own to the dump target it doesn't constrain anything and is ignored as a referencer — blobs stays narrowed to the attachment-referenced ids. (A parent referenced by *multiple* constrained children currently falls back to dumping all of its rows.)
913
- - **`active_storage_variant_records`** holds derivative variant-tracking rows that ActiveStorage regenerates lazily, and it too has no path to the dump target — left alone it would land in the "no relation → dump all" branch and, worse, its `blob_id` could point at blobs outside the narrowed set above (a foreign-key violation on import). `exwiw:schema:generate` therefore emits it with **`ignore: true`** (and drops it from the attachments `record` polymorphic expansion so nothing carries a dangling reference to it), so its data is skipped while the DDL is still written. Remove `ignore` from the generated config if you really need to export it.
235
+ - `active_storage_attachments` is a polymorphic `belongs_to :record`, so only the attachments of exported records are exported.
236
+ - `active_storage_blobs` is narrowed by rule 3 to the blobs those attachments point at.
237
+ - `active_storage_variant_records` is generated with `ignore: true`, because ActiveStorage recreates its rows when needed and its `blob_id` could point at blobs that were not exported. Remove `ignore` if you need it.
914
238
 
915
239
  ### Reverse scope for multi-referencer tables (`reverse_scope`)
916
240
 
917
- The automatic reverse extraction above narrows a table referenced by **exactly one** constrained child. A table referenced by **two or more** constrained children falls back to dumping every row — fine for `active_storage_blobs`, but a problem for a **global-identity table** such as `users`: it carries no scope/tenant column and has no `belongs_to` of its own, yet dozens of scoped tables point *at* it. Dumping it (and everything that hangs off it) in full pulls in every tenant's identities.
918
-
919
- `reverse_scope` opts such a table into **multi-referencer** reverse scoping: you enumerate the referencers whose own (already scoped) extraction queries should be `UNION`'d into the id set the table is constrained to. It is a user-owned key (never emitted by `schema:generate`, preserved across regeneration like `scope_exempt`/`scope_column`):
241
+ A table such as `users` often has no `belongs_to` toward the target, while many narrowed tables point at it. Rule 3 does not apply when there is more than one, so it would be dumped in full, with every tenant's users. `reverse_scope` lists the tables and columns that point at it, and the table is narrowed to the values those tables export:
920
242
 
921
243
  ```json
922
244
  {
@@ -926,121 +248,100 @@ The automatic reverse extraction above narrows a table referenced by **exactly o
926
248
  "via": [
927
249
  { "table": "customers", "column": "user_id" },
928
250
  { "table": "staff", "column": "user_id" },
929
- { "table": "business_entity_customers", "column": "kantan_yoyaku_user_id" }
251
+ { "table": "members", "column": "legacy_user_id" }
930
252
  ]
931
253
  },
932
254
  "columns": [{ "name": "id" }, { "name": "name" }]
933
255
  }
934
256
  ```
935
257
 
936
- produces (each arm reuses that referencer's own scope, so a per-tenant run keeps only that tenant's ids; the `UNION` id set is materialized once and joined back — see [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery)):
258
+ - `column` is given explicitly, so a column with a non-standard name, or one without a `belongs_to`, works.
259
+ - List only narrowed tables. A table in `via` that is not narrowed would add every value, so it is skipped with a warning.
260
+ - A table in `via` may itself be narrowed by its own `reverse_scope`.
261
+ - By default the values are matched against the primary key. Set `reverse_scope.column` to match another column, such as `{ "column": "code", "via": [{ "table": "contracts", "column": "rate_code" }] }`.
262
+ - Tables that `belongs_to` the reverse-scoped table are narrowed by rule 5 and need no config.
937
263
 
938
- ```sql
939
- SELECT users.* FROM users
940
- JOIN (
941
- SELECT DISTINCT exwiw_scope_src_0.user_id AS exwiw_scope_id
942
- FROM (
943
- SELECT customers.user_id FROM customers WHERE <customers' scope> AND customers.user_id IS NOT NULL
944
- UNION
945
- SELECT staff.user_id FROM staff WHERE <staff' scope> AND staff.user_id IS NOT NULL
946
- UNION
947
- SELECT business_entity_customers.kantan_yoyaku_user_id FROM business_entity_customers
948
- WHERE <…' scope> AND business_entity_customers.kantan_yoyaku_user_id IS NOT NULL
949
- ) AS exwiw_scope_src_0
950
- ) AS exwiw_scope_ids_0
951
- ON users.id = exwiw_scope_ids_0.exwiw_scope_id
952
- ```
264
+ ### Self-referencing `belongs_to` (tree tables)
953
265
 
954
- Notes:
266
+ When a table has a `belongs_to` to itself, such as `categories.parent_id`, narrowing it would drop the parents of the rows it keeps. exwiw also keeps every ancestor of those rows, up to the root. Declaring the `belongs_to` is enough.
955
267
 
956
- - **`column` is explicit**, so a *non-default* foreign key (e.g. `kantan_yoyaku_user_id`, or `organization_admins.id` which itself references `users.id`) is honored, and even a column with no declared `belongs_to` edge can be enumerated.
957
- - **Only scoped referencers belong in `via`.** Each arm's query must come out constrained; an unconstrained referencer (e.g. a `scope_exempt` table, or one with no path to a scope) would project *every* id and union the whole table back — so such an arm is **skipped with a warning** rather than silently widening the dump. An unknown table is likewise skipped with a warning. If no arm survives, the table stays unscopable and (in [scope-column mode](#scope-column-mode)) the run aborts via `validate_scope!`.
958
- - **Declarations chain.** A referencer that is itself scoped only by its *own* `reverse_scope` counts as scoped: the arm nests that declaration's `UNION` inside its query, so a normalized side table referenced by a document-style hub that is in turn reverse-scoped through a join table resolves end to end (`attachments <- documents <- join rows <- the target`). A cycle of declarations is cut — the repeated table's arm comes out unconstrained and is dropped with the warning above. The chaining is deliberately limited to `reverse_scope` arms: the *automatic* single-referencer detection still treats a child scoped solely by its own declaration as unconstrained, so such a child never rescues (or, by widening the candidate set past one, costs) a parent that relies on the automatic detection — declare `reverse_scope` on the parent if you need that shape. Chains deeper than three declarations are flagged with a warning, since each level re-embeds its referencers' subqueries and the generated SQL grows exponentially with depth.
959
- - **NULLs are excluded** per arm (`IS NOT NULL`).
960
- - **`column` picks the key the union is matched against.** By default the arms' projected values are compared with this table's `primary_key`. When the referencers point at another key of the table — say `rate_cards` are looked up by a `code` that `contracts.rate_code` carries, not by `rate_cards.id` — set `"reverse_scope": { "column": "code", "via": [{ "table": "contracts", "column": "rate_code" }] }` and the clause becomes `ON rate_cards.code = …`. The column must be declared in `columns`.
961
- - **Satellites need no config.** A table that `belongs_to` the reverse-scoped table (e.g. `end_users.id → users.id`, or `identities.user_id → users.id`) tightens to the kept ids automatically through the normal cascade — only the reverse-scoped table itself declares `reverse_scope`. The cascade is **multi-hop**, so a table several `belongs_to` hops below the reverse-scoped table (e.g. `end_user_profiles → end_users → users`) also tightens automatically, with no config of its own.
962
- - Works in both single-target and scope-column mode. In single-target mode there is no scope-column pre-flight (`validate_scope!`), so a satellite the cascade cannot resolve to a single scopable parent (e.g. it `belongs_to` two scopable hubs) is dumped in full with a warning rather than aborting. Polymorphic foreign keys are not eligible as anchors (the named `column` is always a concrete column).
963
- - **The MongoDB adapter supports `reverse_scope` too** — same config shape and semantics, but the id set is captured at runtime instead of being emitted as a `UNION` subquery. See [`reverse_scope` on collections](docs/mongodb.md#reverse_scope-on-collections) under MongoDB support.
268
+ - Ancestors are kept regardless of the table's `filter` and scope, so foreign keys stay valid. If a tree spans tenants, the dump can include another tenant's ancestors.
269
+ - Tables below the tree (`category_notes`) also keep the rows of the ancestors, and tables the tree points at keep what the ancestors point at. An ancestor with many rows hanging off it brings them all; set `ignore: true` on that `belongs_to` to leave them out.
270
+ - The walk stops at a `NULL` or missing parent and at cycles in the data. Polymorphic self-references are not supported.
271
+ - On MySQL this needs 8.0 or later, and a tree deeper than `cte_max_recursion_depth` (1000 by default) fails.
272
+ - [Batched](#batched-extraction-batch_scope) tables do not get ancestors added; exwiw warns about this. MongoDB does not add ancestors either.
964
273
 
965
- ### Why a JOIN, not `IN (subquery)`
274
+ ### Scope-column mode
966
275
 
967
- Every scope id-set above — the multi-referencer `reverse_scope` `UNION`, the single-referencer reverse extraction, and the multi-hop forward cascade — is emitted as a `JOIN` to a `SELECT DISTINCT` derived table rather than `<col> IN (<subquery>)`:
276
+ Single-target mode assumes every table reaches one target table through `belongs_to`. In many multi-tenant schemas, tables instead each carry a tenant column (`tenant_id`) and are not all connected to one root. Picking one table as the target would dump the unrelated tables in full.
968
277
 
969
- ```sql
970
- … JOIN (SELECT DISTINCT src.<id> AS exwiw_scope_id FROM (<id-set subquery>) AS src) AS ids
971
- ON <table>.<col> = ids.exwiw_scope_id
278
+ In scope-column mode, each table names the column that holds the tenant id, and `--ids` are values of that column:
279
+
280
+ ```json
281
+ {
282
+ "name": "shops",
283
+ "primary_key": "id",
284
+ "scope_column": "tenant_id",
285
+ "columns": [{ "name": "id" }, { "name": "name" }, { "name": "tenant_id" }]
286
+ }
287
+ ```
288
+
289
+ ```bash
290
+ exwiw \
291
+ --adapter=postgresql \
292
+ --host=localhost --port=5432 --user=reader \
293
+ --database=app_production \
294
+ --schema-dir=exwiw/schema \
295
+ --ids=42,43 \
296
+ --output-dir=dump
972
297
  ```
973
298
 
974
- Both forms select the **same rows** — the `DISTINCT` dedups, so the join never fans out — but the query plans differ sharply on a large table. As `<col> IN (… UNION …)`, MySQL cannot turn a `UNION` subquery into a materialized semi-join and falls back to its IN-to-`EXISTS` rewrite: a **correlated `DEPENDENT SUBQUERY`** re-evaluated for every outer row, i.e. a full scan of the (potentially huge) outer table multiplied by the cost of the union. The derived-table form forces the engine to evaluate the id set **once** (the `DISTINCT` makes the derived table non-mergeable, hence materialized) and then probe the outer table by its primary key. On a global-identity table such as `users` this is the difference between a full table scan and an index lookup; the cascade nests the same way, so each level is materialized once instead of being re-evaluated by the level above.
299
+ Here `42,43` are `tenant_id` values, not shop ids. Passing `--target-table=shops` gives the same result.
975
300
 
976
- All three SQL adapters (mysql / postgresql / sqlite) emit this shape. PostgreSQL additionally reconciles a `uuid`/`varchar` type mismatch by casting the join key and the projected id to `text`, exactly as the old `IN` form did.
301
+ Tables without a `scope_column` are narrowed by the rules above. A table that none of them can narrow stops the run before anything is exported, and the error lists those tables. For each, declare a `scope_column`, add a `belongs_to`, set `ignore: true`, or set `scope_exempt: true`.
977
302
 
978
- ### Rails-managed tables (special `type` values)
303
+ Scope-column mode is for the SQL adapters only. Use `exwiw explain` to check the queries first.
979
304
 
980
- Some tables are owned by Rails itself rather than the application — they have no ActiveRecord model and Rails reserves the right to evolve their column shape between versions (e.g. `schema_migrations`, `ar_internal_metadata`). exwiw treats them as a distinct category via the `type` field on a table config:
305
+ #### Cross-database foreign keys
981
306
 
982
- - `type: "rails_managed_schema_migrations"` — Rails' migration history table (`ActiveRecord::Base.schema_migrations_table_name`).
983
- - `type: "rails_managed_internal_metadata"` — Rails' internal metadata table (`ActiveRecord::Base.internal_metadata_table_name`).
307
+ A `belongs_to` to a table in another database cannot be joined, so `schema:generate` writes it with `ignore: true`. The foreign key column is still there, so declaring `scope_column: "<that foreign key>"` on the table narrows it without a join.
984
308
 
985
- `exwiw:schema:generate` emits these entries automatically when the corresponding tables exist on the connection — they are NOT pulled from `ActiveRecord::Base.descendants` because they have no model class.
309
+ #### `scope_exempt` (intentional full dump)
986
310
 
987
- A rails-managed entry has a minimal shape (no `primary_key`, no `belongs_tos`, no `columns`):
311
+ A master table with no personal data and no relation to the tenant can be exported in full:
988
312
 
989
313
  ```json
990
- {
991
- "name": "schema_migrations",
992
- "type": "rails_managed_schema_migrations",
993
- "comment": "Managed internally by Rails. Tracks applied schema migrations."
994
- }
314
+ { "name": "countries", "primary_key": "id", "scope_exempt": true, "columns": [{ "name": "id" }, { "name": "code" }] }
995
315
  ```
996
316
 
997
- Behavior at dump time:
998
-
999
- - Extraction uses `SELECT *` so the dump is robust against Rails-side column additions.
1000
- - `INSERT` statements omit the column list (`INSERT INTO schema_migrations VALUES (...)`). For PostgreSQL `--output-format=copy`, the `COPY` header similarly omits the column list (`COPY schema_migrations FROM stdin;`).
317
+ `schema_migrations` and `ar_internal_metadata` are exempt automatically.
1001
318
 
1002
- Constraints:
1003
-
1004
- - Defining `primary_key`, `columns`, or `belongs_tos` on a rails-managed entry is rejected with `ArgumentError` on load.
1005
- - A rails-managed table cannot be used as `--target-table`.
1006
- - In multi-database setups, the rails-managed entry is emitted under whichever database's connection actually contains the table (see [Multiple databases](#multiple-databases)). The table name itself is still derived from the global `ActiveRecord::Base.schema_migrations_table_name` / `internal_metadata_table_name` (prefix/suffix) accessors.
1007
-
1008
- ### Composite primary keys (unsupported)
319
+ #### Per-table `scope_column` and ID spaces
1009
320
 
1010
- exwiw does not yet support tables with a composite primary key. When `exwiw:schema:generate` encounters a model whose `primary_key` is an array, it still emits a config entry so the table is not silently dropped, but marks it `ignore: true`, tags it `type: "unsupported_composite_primary_key"`, and records the key columns in a `comment`:
321
+ Each table names its own column, so tables that store the tenant id under different names work together. When one database has two groups of tables keyed by different kinds of id, give one group an `id_space` and pass its values separately:
1011
322
 
1012
323
  ```json
1013
- {
1014
- "name": "composite_pk_records",
1015
- "type": "unsupported_composite_primary_key",
1016
- "ignore": true,
1017
- "comment": "exwiw does not support composite primary keys (organization_id, location_id); data extraction is skipped.",
1018
- "belongs_tos": [],
1019
- "columns": [{ "name": "organization_id" }, { "name": "location_id" }, { "name": "name" }]
1020
- }
324
+ { "name": "tenants", "primary_key": "id", "scope_column": "id", "columns": [{ "name": "id" }] }
325
+ { "name": "organizations", "primary_key": "id", "scope_column": "id", "id_space": "org", "columns": [{ "name": "id" }] }
1021
326
  ```
1022
327
 
1023
- Unlike rails-managed entries, `columns` and `belongs_tos` are retained so the entry is ready to wire up once composite-key support lands. The `type` is purely a marker — `ignore: true` is what actually excludes the table from extraction, so removing `ignore` (and supplying a workable `primary_key`) lets you opt the table back in manually.
1024
-
1025
- ### Bulk insert chunk size
328
+ ```bash
329
+ exwiw ... --ids=1,2 --ids=org=0b6f4c1e-0000-4000-8000-000000000001
330
+ ```
1026
331
 
1027
- `bulk_insert_chunk_size` splits the generated `INSERT` statement into multiple statements, each containing at most the specified number of rows. This is useful when the number of records per table is large enough to hit limits like MySQL's `max_allowed_packet`.
332
+ - Tables without `id_space`, and `--ids` without a prefix, use the `default` ID space. Write `--ids=default=...` when an id itself contains `=`.
333
+ - A table without a scope column uses the ID space of the table it is narrowed through.
334
+ - The run stops before exporting when an ID space a table needs has no values, when values are given for an unused ID space, or when a table reaches tables of more than one ID space.
335
+ - Single-target mode and MongoDB support only the `default` ID space.
336
+ - In the config file, `ids:` takes a list, or a mapping such as `ids: { default: [1, 2], org: [...] }`.
1028
337
 
1029
- If omitted, the adapter default applies: 10,000 rows per statement for the SQL adapters (1,000 documents per chunk for MongoDB). Tables at or below the chunk size still produce a single `INSERT` statement. To force a single statement regardless of table size, set a value larger than the table's row count.
338
+ The older global `--scope-column=COLUMN` flag still works but is deprecated; declare `scope_column` per table instead.
1030
339
 
1031
340
  ### Batched extraction (`batch_scope`)
1032
341
 
1033
- A scoped table is normally extracted with one query, whose scope filter sits on the table it joins up to:
1034
-
1035
- ```sql
1036
- SELECT activities.* FROM activities
1037
- JOIN customers ON activities.customer_id = customers.id
1038
- AND customers.tenant_id IN ('t1')
1039
- ```
1040
-
1041
- That is index-driven while the scope keeps few `customers`. Past some number of them the planner's estimate of "probe the foreign-key index once per customer" exceeds its estimate of "scan the table once", and it switches to a **sequential scan of the whole table** — for a result set that is a small fraction of it. On a table of hundreds of millions of rows the scan then exceeds the server's `statement_timeout`, or simply runs for hours. Note that no `filter` on the extracted table fixes this: a predicate that reduces the *output* does not reduce the *work* once the plan is a scan (it may not even change the plan).
342
+ A table with hundreds of millions of rows can be slow to export even when few of its rows are kept. Once the scope covers many parent rows, the database may decide that scanning the whole table is cheaper than using the foreign key index, and the query runs for hours or hits `statement_timeout`.
1042
343
 
1043
- `batch_scope` removes the choice instead of arguing with the estimate. It names the scoped table this one reaches — the **batch table** — and exwiw resolves that table's in-scope primary keys once, then extracts one `size`-sized slice of those ids at a time:
344
+ `batch_scope` avoids this by exporting the table in batches. exwiw first fetches the ids of a narrowed table on the path (the batch table), then runs one query per `size` ids, with the ids written into the query:
1044
345
 
1045
346
  ```json
1046
347
  {
@@ -1052,224 +353,222 @@ That is index-driven while the scope keeps few `customers`. Past some number of
1052
353
  }
1053
354
  ```
1054
355
 
1055
- Each batch runs with that slice's ids in place of the scope filter:
1056
-
1057
356
  ```sql
1058
357
  SELECT activities.* FROM activities
1059
358
  JOIN customers ON activities.customer_id = customers.id
1060
359
  AND customers.id IN (/* 1000 ids */)
1061
360
  ```
1062
361
 
1063
- An explicit id list of that size is exactly estimated and selective, so the foreign-key index is unambiguously the cheapest plan for every batch, and total work is proportional to the rows the table actually keeps rather than to the table's size.
362
+ - The exported rows are the same as without batching. The batch table's ids are sorted, so the output is the same on every run.
363
+ - `size` defaults to 1000.
364
+ - The batch table can be several hops up the path.
365
+ - A table with its own `scope_column` can name itself. This only helps when the scope column is indexed.
366
+ - `batch_scope` requires scope-column mode, and a table narrowed by rule 1 or rule 2. Other cases are rejected before anything is written.
367
+ - With `--output-format=copy`, all batches are held in memory at once. Use the default `INSERT` format for very large results.
368
+ - `exwiw explain` also shows the query that fetches the batch table's ids.
1064
369
 
1065
- - **The dumped rows are the same as the unbatched query's** (in batch-by-batch order). The slices partition the id set — every id is in exactly one batch — so no row is dropped or emitted twice. The ids are sorted (in exwiw, not with `ORDER BY` — the id-set query stays cheap on the source DB) before slicing, so batch composition, and the dump, is reproducible run to run.
1066
- - **`size` defaults to 1000** ids per batch.
1067
- - The batch table's ids come from **its own extraction query**, so it is narrowed by exactly the filter it would carry in the unbatched query. They are held in memory for the extraction: one scope's worth of primary keys, orders of magnitude smaller than the table being batched.
1068
- - The batch table may be **any number of hops up** the path — a table two hops below it (`activity_orders → activities → customers`) names `customers` too, and the batch ids are applied where the path meets the scope, bounding the whole join chain.
1069
- - A table that **carries the scope column itself** batches by naming itself; each batch then filters `WHERE <pk> IN (<ids>)` directly. Note that the id-set query is then the same scope predicate over the same table, so this shape only avoids the scan when the scope column is indexed (ideally index-only) — the join shape above is the one that genuinely removes the planner's choice.
1070
- - `bulk_insert_chunk_size` is independent: batches are query boundaries, chunks are `INSERT` statement boundaries.
1071
- - With `--output-format=copy`, batching bounds each query's cost but not memory: COPY builds the whole table's body in memory, so all batches' rows are resident at once. Use the default INSERT format (which streams) when the kept rows themselves are huge.
370
+ ## Masking
1072
371
 
1073
- **Supported shapes.** A batch key only splits an extraction correctly when *every* row the table keeps is selected through the batch table's scope filter — otherwise a route the batch key does not constrain would keep the same rows in every batch, and the dump would repeat them (a primary-key conflict on import). So `batch_scope` requires [scope-column mode](#scope-column-mode) and one of:
372
+ Each column can be masked with one of the following keys.
1074
373
 
1075
- - the table is **directly scoped** (`scope_column`) and names itself, or
1076
- - the table reaches the scope through a **single `belongs_to` join path** (path 2 in [the six scoping paths](#how-each-table-is-narrowed--the-six-scoping-paths)) whose scoped terminus is the named table.
374
+ ### `replace_with`
1077
375
 
1078
- Every other shape — polymorphic arm `UNION`s, `reverse_scope`, referenced-by, the parent cascade, `scope_exempt` (on the batched table *or* the batch table, whose id set would then not be scoped), and single `--target-table` mode — is **rejected with an explanation** rather than silently mis-sliced, before any output is written. (In single-target mode the extraction is already anchored on a caller-supplied id list, so batching it means running exwiw once per slice of `--ids`.)
376
+ Replaces the value with a string. `{column}` is replaced with that column's value, so for a row with `id` 1, `"user{id}@example.com"` becomes `user1@example.com`. `{}` is kept as is, so `"replace_with": "{}"` gives an empty JSON object.
1079
377
 
1080
- `exwiw explain` prints the id-set query and its `EXPLAIN` after a batched table's own query, since that query is the part of a batched export the table's query does not show. It cannot show a batch's literal id list — `explain` resolves no ids, because it executes no extraction SELECT.
378
+ A number or boolean is used as it is, so non-text columns keep their type:
1081
379
 
1082
- Like `scope_column` / `scope_exempt` / `reverse_scope`, `batch_scope` is user-maintained: never emitted by `schema:generate`, and preserved across regeneration.
380
+ ```jsonc
381
+ { "name": "score", "replace_with": 0 }
382
+ { "name": "active", "replace_with": false }
383
+ ```
1083
384
 
1084
- ### Filter
385
+ `NULL` stays `NULL` (an empty string is still replaced).
1085
386
 
1086
- Some case, you don't need full records related to target. e.g. dump user access logs only for the last year.
1087
- `filter` is here for that. Be careful to use this option, as it will be:
387
+ ### `raw_sql`
1088
388
 
1089
- - injected as it is in table condition(e.g. WHERE on mysql), so you are recommended to clearify table name of column to avoid ambiguity.
1090
- - injected to every where / join clause, so it affects to all tables depends on filterted target-table. it results to data inconsistency.
1091
- - a way to reduce the rows returned, which is **not** necessarily a way to reduce the work: on a large table the engine may keep (or switch to) a full scan and evaluate the filter per row. See [batched extraction](#batched-extraction-batch_scope) when the goal is to bound how much of the table is read.
389
+ An SQL expression used in place of the column, such as `"CONCAT('user', shops.id, '@example.com')"`. Use it when a database function is needed. Qualify column names with the table name. `replace_with` is ignored when both are set. SQL adapters only.
1092
390
 
1093
- ### Masking
391
+ ### `map`
1094
392
 
1095
- `exwiw` provides several options for masking value.
393
+ Ruby code that returns a `Proc`. The proc is called with each row, and its return value (a `String`, a number or `nil`) replaces the column:
1096
394
 
1097
- #### `replace_with`
395
+ ```jsonc
396
+ { "name": "email", "map": "proc { |r| 'user' + r['id'].to_s + '@example.com' }" }
397
+ ```
1098
398
 
1099
- It will replace the value with the specified string,
1100
- and you can use the column name with `{}` to replace the value with the column value.
399
+ - `r['column']` reads any column of the row, after the database-side masking of other columns.
400
+ - `NULL` is not kept automatically; the proc receives `nil` and decides.
401
+ - It runs in the exwiw process, so `explain` does not show it. SQL adapters only.
402
+ - Because the config runs arbitrary Ruby, only load configs you trust.
1101
403
 
1102
- For example, Let assume we have the record which id is 1,
1103
- then "user{id}@example.com" will be replaced with "user1@example.com".
404
+ Prefer `replace_with` or `raw_sql` when they can do the job.
1104
405
 
1105
- `replace_with` **preserves NULL**: a source value that is `NULL` (or, for MongoDB, an
1106
- absent field) is left as-is instead of being replaced by the masked literal, so the
1107
- "not set" signal survives into the dump. Only true `NULL`/absent is preserved — an empty
1108
- string is a real value and is still masked. Because of this you do not need to hand-write a
1109
- `raw_sql` `CASE WHEN ... IS NOT NULL ...` to keep NULLs.
406
+ ### `replace_with_fake_data`
1110
407
 
1111
- A **non-String** value (number or boolean) is used verbatim instead of being rendered as a
1112
- template, so a column that is not text keeps its type:
408
+ Replaces the value with realistic fake data from the [faker](https://github.com/faker-ruby/faker) gem. The value is chosen from a seed column, so the same seed always gives the same value, across tables, runs and adapters:
1113
409
 
1114
410
  ```jsonc
1115
- { "name": "score", "replace_with": 0 } // integer column -> SELECT emits the literal 0
1116
- { "name": "active", "replace_with": false } // boolean column
1117
- { "name": "email", "replace_with": "masked-{id}@example.com" } // template, as above
411
+ { "name": "name", "replace_with_fake_data": { "seed": "users.id", "type": "human_name", "locale": "ja" } }
1118
412
  ```
1119
413
 
1120
- The SQL adapters emit it as a typed literal (not concatenated into text) and the MongoDB
1121
- adapter assigns it as-is, so the field keeps its BSON type. NULL preservation applies to both
1122
- forms.
414
+ | type | example (en) | example (`locale: ja`) |
415
+ |------|--------------|------------------------|
416
+ | `human_name` | `Adrianna Kilback` | `山田 太郎` |
417
+ | `first_name` | `Adrianna` | `太郎` |
418
+ | `last_name` | `Kilback` | `山田` |
419
+ | `human_name_kana` | (ja only) | `ヤマダ タロウ` |
420
+ | `first_name_kana` | (ja only) | `タロウ` |
421
+ | `last_name_kana` | (ja only) | `ヤマダ` |
422
+ | `phone_number` | `(555) 123-4567` | |
423
+ | `address` | `282 Kevin Brook, Imogeneborough, CA 58517` | |
424
+ | `company_name` | `Hirthe-Ritchie` | |
425
+ | `email` | `cliff.fay.9d6b804eff5a3f57@example.com` | |
426
+ | `username` | `cliff.fay_9d6b804eff5a3f57` | |
1123
427
 
1124
- In the String form, a `{...}` placeholder must name a column: an empty brace pair (`{}`) names
1125
- nothing, so it is emitted literally — which is what makes `"replace_with": "{}"` a usable
1126
- empty-JSON mask, on every adapter.
428
+ - `seed` is a column of the same table, with or without the table name. Use a stable id such as the primary key.
429
+ - For one seed, the name types describe the same person: `human_name` is `last_name` + `first_name`, and the kana matches the kanji.
430
+ - Different seeds can get the same name. `email` and `username` include a token from the seed, so they stay unique.
431
+ - Values change when `locale`, the faker version, or exwiw's bundled Japanese name list changes.
432
+ - `NULL` stays `NULL`.
433
+ - Add `gem "faker"` to your Gemfile. A config that uses only `ja` name types does not need it.
434
+ - It runs in the exwiw process, so `explain` does not show it. The cost is small; see [`docs/row-transform-masking-notes.md`](docs/row-transform-masking-notes.md).
1127
435
 
1128
- #### `raw_sql`
436
+ Only one masking key can be set on a column.
1129
437
 
1130
- It will used instead of the original value.
438
+ ## Generating the schema config
1131
439
 
1132
- For example, `"raw_sql": "CONCAT('user', shops.id, '@example.com')"` is equivalent to
1133
- `"replace_with": "user{id}@example.com"`.
1134
- This is useful when you want to transform with functions provided by the database.
440
+ In a Rails application, a rake task writes the schema config from the models:
1135
441
 
1136
- Notice that you are recommended to clearify table name of column to avoid ambiguity.
442
+ ```bash
443
+ bundle exec rake exwiw:schema:generate
444
+ ```
1137
445
 
1138
- If it used with `replace_with`, `replace_with` will be ignored.
446
+ The files go to `EXWIW_SCHEMA_DIR_PATH` if set, otherwise `schema_dir` from `exwiw.yml`, otherwise `exwiw/schema`. If the application has more than one source of models, such as ActiveRecord and Mongoid, give each its own directory; `check` and `tidy` treat files they do not recognize as stale.
1139
447
 
1140
- #### `map`
448
+ ### Safe mode (masking new columns by default)
1141
449
 
1142
- The value is evaluated as Ruby code once (per table, at dump time), must yield a
1143
- `Proc`, and the proc is called for every fetched row. Its return value replaces
1144
- the column value in the dump:
450
+ A new column could hold personal data, so `schema:generate` writes every column that is not yet in the config as masked, and marks it with [`needs_mask_decision: true`](#needs_mask_decision). Columns already in the config are left as they are.
1145
451
 
1146
- ```jsonc
1147
- { "name": "email", "map": "proc { |r| 'user' + r['id'].to_s + '@example.com' }" }
452
+ The mask is the column's default value if it has a constant one, otherwise a value by type: `masked-{primary key}` for text (with `@example.com` when the name mentions mail), `0`, `false`, a fixed date, or `{}` for JSON. Some columns are marked but left unmasked, because masking them would break the dump or the restore:
453
+
454
+ - the primary key and the columns `belongs_to` joins on
455
+ - types no constant fits, such as `uuid`, binary, enums, arrays, and text too short for the mask
456
+ - columns under a unique index, unless the mask differs per row
457
+
458
+ For the first config of an application, where every column is new, run with `EXWIW_NEW_COLUMNS=plain` to turn safe mode off. Do not use it afterwards: those columns get no mark, so nothing tells them apart from reviewed ones.
459
+
460
+ ### `needs_mask_decision`
461
+
462
+ `needs_mask_decision: true` marks a column whose masking nobody has decided yet. Export ignores it. [`schema:check`](#checking-the-config-against-the-schema) reports these columns, so CI can block a pull request until each is decided. To decide, keep the mask (ideally with a `comment` saying why), change it, remove `replace_with` to export the real value, or set `ignore: true`, and then remove the key.
463
+
464
+ ### Tidying stale config (`schema:tidy`)
465
+
466
+ `schema:generate` never deletes anything. `schema:tidy` removes the config files of tables that no longer exist in the database, and the columns those tables no longer have. It reads the database, not the models, so a table without a model is kept. It does not touch anything else, and does not remove stale `belongs_tos`; run `schema:generate` for those.
467
+
468
+ ```bash
469
+ bundle exec rake exwiw:schema:tidy
1148
470
  ```
1149
471
 
1150
- which is equivalent to `"replace_with": "user{id}@example.com"`.
1151
-
1152
- - `r['column_name']` reads any column of the current row — the value as fetched
1153
- from the database (i.e. after SQL-side masking such as another column's
1154
- `replace_with`, before Ruby-side transforms). `r` is only valid inside the
1155
- call; do not retain it.
1156
- - Return a `String`, `Numeric`, or `nil`. Unlike `replace_with` there is **no
1157
- automatic NULL preservation** — the proc receives `nil` and decides.
1158
- - `map` is exclusive with the other masking keys on the same column
1159
- (`raw_sql` / `replace_with` / `replace_with_fake_data`).
1160
- - SQL adapters only. On the MongoDB adapter the key is rejected on load, like
1161
- `raw_sql` (see [Unknown keys are rejected](#unknown-keys-are-rejected)).
1162
- Because the transform runs in the exwiw process, it is invisible to
1163
- `explain`.
1164
-
1165
- **Security note**: `map` executes arbitrary Ruby from the schema config. Treat
1166
- config files with the same trust as your Gemfile — only load trusted configs.
1167
-
1168
- This is the most powerful option, but it runs per row in the exwiw process
1169
- rather than in the database. The measured dispatch cost is small, though
1170
- (~0.6–0.8µs/row plus whatever the proc body does — see
1171
- [`docs/row-transform-masking-notes.md`](docs/row-transform-masking-notes.md)).
1172
- Prefer `replace_with`/`raw_sql` when they can express the transform; reach for
1173
- `map` when they cannot.
1174
-
1175
- #### `replace_with_fake_data`
1176
-
1177
- Replaces the value with realistic-looking fake data generated by the
1178
- [faker](https://github.com/faker-ruby/faker) gem, picked **deterministically**
1179
- from the value of a seed column — the same seed value always maps to the same
1180
- fake value, across tables, runs, and adapters:
472
+ ### Checking the config against the schema
1181
473
 
1182
- ```jsonc
474
+ `schema:check` reports how the config differs from what `generate` and `tidy` would produce, without writing anything. It exits non-zero when something needs attention, so it can run in CI:
475
+
476
+ ```bash
477
+ bundle exec rake exwiw:schema:check
478
+ ```
479
+
480
+ ```json
1183
481
  {
1184
- "name": "name",
1185
- "replace_with_fake_data": { "seed": "users.id", "type": "human_name" }
482
+ "added_tables": [],
483
+ "added_columns": ["users.contact_email"],
484
+ "removed_tables": [],
485
+ "removed_columns": ["orders.legacy_flag"],
486
+ "changed_tables": ["orders", "users"],
487
+ "needs_mask_decision": ["orders.memo"],
488
+ "stale_tables": [],
489
+ "stale_columns": ["orders.legacy_flag"]
1186
490
  }
1187
491
  ```
1188
492
 
1189
- - `seed` names a column of the same table, bare (`"id"`) or table-qualified
1190
- (`"users.id"`). The seed value is hashed (SHA-256, after `to_s`
1191
- normalization, so sqlite's integer `123` and postgres/mysql's string `"123"`
1192
- agree) and the hash picks the fake value. Use a stable identifier (integer or
1193
- string primary key) as the seed; float/decimal/binary columns are discouraged
1194
- because their text forms differ per adapter. A `NULL` seed value hashes `""`
1195
- (still deterministic).
1196
- - Like `replace_with`, it **preserves NULL** in the target column.
1197
- - `locale` (optional) sets the locale used to build the candidate values, e.g.
1198
- `{ "seed": "id", "type": "human_name", "locale": "ja" }` produces Japanese
1199
- names.
1200
- - Supported `type`s:
1201
-
1202
- | type | example output (en) | example output (`locale: ja`) |
1203
- |------|----------------|----------------|
1204
- | `human_name` | `Adrianna Kilback` | `山田 太郎` |
1205
- | `first_name` | `Adrianna` | `太郎` |
1206
- | `last_name` | `Kilback` | `山田` |
1207
- | `human_name_kana` | — (ja only) | `ヤマダ タロウ` |
1208
- | `first_name_kana` | — (ja only) | `タロウ` |
1209
- | `last_name_kana` | — (ja only) | `ヤマダ` |
1210
- | `phone_number` | `(555) 123-4567` | |
1211
- | `address` | `282 Kevin Brook, Imogeneborough, CA 58517` | |
1212
- | `company_name` | `Hirthe-Ritchie` | |
1213
- | `email` | `cliff.fay.9d6b804eff5a3f57@example.com` | |
1214
- | `username` | `cliff.fay_9d6b804eff5a3f57` | |
1215
-
1216
- - **Coherent identity across the name family.** The person-family types
1217
- (`human_name`, `first_name`, `last_name` and their `*_kana` counterparts) all
1218
- draw from a single shared pool of people per locale, keyed by the same seed —
1219
- so for one seed value the last name, first name, full name, and every kana
1220
- reading belong to the **same person**: `human_name` always equals
1221
- `last_name` + `first_name`, and `human_name_kana` matches `human_name`. Full
1222
- names are ordered per locale (`姓 名` for `ja`, `First Last` otherwise).
1223
- - **Kana (`*_kana`) types require `locale: ja`.** faker's `ja` locale ships
1224
- kanji names with no reading, so exwiw bundles its own paired (kanji, katakana)
1225
- dataset for `ja`; this is what lets a fake person carry a kana reading that
1226
- actually matches its kanji. Requesting a `*_kana` type with any other locale
1227
- raises a clear error at build time.
1228
- - Values are drawn from a pre-generated pool per (type, locale), so distinct
1229
- seeds can share a fake value. The name family uses one shared **person** pool
1230
- of 20,000 identities — for `ja` these are 20,000 *distinct* people enumerated
1231
- from the bundled (kanji, kana) dataset (142 surnames × 142 given names); the
1232
- other types use an independent 10,000-candidate pool. The
1233
- uniqueness-sensitive types (`email`, `username`) additionally embed a 64-bit
1234
- hex token derived from the seed hash, so they stay collision-free under a
1235
- unique index even at millions of rows (collision probability at 5M distinct
1236
- seeds ≈ 7e-7) and always use the `example.com` domain.
1237
- - **Determinism caveat**: values are stable for a given locale plus the version
1238
- of the value source — the faker gem for the non-`ja` name family and the
1239
- independent types, and exwiw's bundled dataset for the `ja` name family.
1240
- Upgrading that source (or changing `locale`) regenerates the pool and maps
1241
- seeds to different values. The seed→value mapping itself never changes within
1242
- one version.
1243
- - The faker gem is **not** a runtime dependency of exwiw — add `gem "faker"` to
1244
- your Gemfile to use this mode (exwiw raises a clear error otherwise). A config
1245
- that uses **only** `ja` person types needs no faker (that pool is built
1246
- entirely from the bundled dataset); faker is required for every other type
1247
- and locale.
1248
- - Exclusive with the other masking keys on the same column, and invisible to
1249
- `explain`. Also supported by the MongoDB adapter on a `MongodbField` (seed
1250
- names a field of the collection, or `_id`), where it is applied document-side
1251
- after `replace_with` — see [MongoDB support](docs/mongodb.md#masking).
1252
-
1253
- **Performance**: this is a per-row Ruby transform, measured at ~1.5–1.6µs/row
1254
- per fake column (so ≈ +8s per 5M rows per column; ~+40% against a local sqlite
1255
- fetch — the worst case — and proportionally less against a network database,
1256
- where the fetch dominates). Values are drawn from a pool pre-generated once,
1257
- not by calling faker per row (which would be ~20× slower). Memory is
1258
- unaffected: the transform streams with the dump. See
1259
- [`docs/row-transform-masking-notes.md`](docs/row-transform-masking-notes.md)
1260
- for the benchmark, and `script/bench_row_transform.rb` to measure on your data.
1261
-
1262
- ### MongoDB
1263
-
1264
- exwiw can export MongoDB databases too (`--adapter=mongodb`): JSONL output importable with `mongoimport`, schema/index DDL for `mongosh`, masking inside embedded documents, `reverse_scope` on collections, Mongoid-based config generation, a server-enforced query timeout, and parallel dump workers. Everything MongoDB-specific is documented in [docs/mongodb.md](docs/mongodb.md).
1265
-
1266
- ## How it works
1267
-
1268
- - Load the table information from the specified config file.
1269
- - Calculate the dependency between tables.
1270
- - Generate the full list of INSERT sql based on the specified conditions.
1271
- - If the processing table has no relation with target tables, then dump all records.
1272
- - If the processing table has relation with target tables, then dump the records which are related to the target tables.
493
+ - `added_*`, `removed_*` and `changed_tables`: run `schema:generate` and `schema:tidy`.
494
+ - `needs_mask_decision`: columns still waiting for a decision.
495
+ - `stale_*`: removed tables and columns that the config still exports, which would make the export fail. Used by `--fail-on=stale` (see [below](#non-rails-applications-exwiw-schema----from-db)).
496
+
497
+ With multiple databases each entry starts with the database name (`primary/users.email`). Set `EXWIW_SCHEMA_CHECK_OUTPUT=<path>` to also write the JSON to a file.
498
+
499
+ ### Multiple databases
500
+
501
+ With Rails' multiple databases, `schema:generate` writes each database's files into a subdirectory named after it (`exwiw/schema/primary/`, `exwiw/schema/analytics/`). Each database is exported by a separate run.
502
+
503
+ A `belongs_to` to a model in another database is written with `ignore: true` and `ignore_type: "cross_database"`; see [Cross-database foreign keys](#cross-database-foreign-keys).
504
+
505
+ ### Rails-managed tables and composite primary keys
506
+
507
+ `schema_migrations` and `ar_internal_metadata` get a config with a `type` such as `rails_managed_schema_migrations` and no columns. They are exported with `SELECT *` and `INSERT` without a column list, so new Rails versions do not break them. They cannot be the `--target-table`.
508
+
509
+ Composite primary keys are not supported. Such a table is generated with `ignore: true` and `type: "unsupported_composite_primary_key"`.
510
+
511
+ ### Mongoid applications
512
+
513
+ ```bash
514
+ bundle exec rake exwiw:schema:generate_mongoid
515
+ bundle exec rake exwiw:schema:tidy_mongoid
516
+ bundle exec rake exwiw:schema:check_mongoid
517
+ ```
518
+
519
+ These work like the ActiveRecord tasks. See [Generating config from Mongoid models](docs/mongodb.md#generating-config-from-mongoid-models).
520
+
521
+ ### Non-Rails applications (`exwiw schema ... --from-db`)
522
+
523
+ For applications exwiw cannot load, the same three operations read the database instead of the models:
524
+
525
+ ```bash
526
+ exwiw schema generate --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
527
+ exwiw schema check --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
528
+ exwiw schema tidy --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
529
+ ```
530
+
531
+ - MySQL and PostgreSQL only.
532
+ - `DATABASE_PASSWORD` may be empty, for CI databases without a password.
533
+ - Safe mode and the `check` report work as above. `check` exits 1 when the config needs attention, and with another status when it could not run.
534
+ - `check --fail-on=stale` fails only on `stale_*`, so it can run before an export without blocking on newly added columns.
535
+ - One run covers one database, and the files are written directly into the schema directory.
536
+
537
+ Foreign keys in the database become `belongs_tos`. A table without a primary key is generated with `ignore: true` and a comment; once you add a `primary_key` by hand, it is kept.
538
+
539
+ Regeneration only adds `belongs_tos`, because many relations exist only in application code. Add those by hand; they are kept from then on, and their foreign key columns are never masked. `tidy` removes a `belongs_to` whose table no longer exists.
540
+
541
+ ## After-insert hook
542
+
543
+ `--after-insert-hook=PATH` runs a script after all data files are written, to add rows of your own.
544
+
545
+ A Ruby hook (`.rb`) can use:
546
+
547
+ - `cli_options`: the parsed options, such as `cli_options.fetch(:ids)`.
548
+ - `ids_for(id_space = "default")`: the ids of an [ID space](#per-table-scope_column-and-id-spaces).
549
+ - `insert_sql(template)`: renders an ERB template and writes it to `insert-{N+1}-after_insert.sql`, after the last data file. Multiple calls go into the same file.
550
+ - `insert_jsonl(collection, template)`: MongoDB only; see [MongoDB support](docs/mongodb.md).
551
+
552
+ ```ruby
553
+ insert_sql <<~SQL
554
+ <%- cli_options.fetch(:ids).each do |tenant_id| -%>
555
+ INSERT INTO users (tenant_id, email) VALUES (<%= tenant_id %>, 'default@example.com');
556
+ <%- end -%>
557
+ SQL
558
+ ```
559
+
560
+ Ruby hooks run inside the exwiw process, so only use hooks you trust.
561
+
562
+ Any other file is run as a command. Its output is not captured, and a non-zero exit stops exwiw. It gets `DATABASE_PASSWORD` and these environment variables:
563
+
564
+ - `EXWIW_OUTPUT_DIR`, `EXWIW_SCHEMA_DIR`
565
+ - `EXWIW_DATABASE_ADAPTER`, `EXWIW_DATABASE_HOST`, `EXWIW_DATABASE_PORT`, `EXWIW_DATABASE_USER`, `EXWIW_DATABASE_NAME`
566
+ - `EXWIW_TARGET_TABLE`, `EXWIW_IDS` (comma-separated, the `default` ID space), `EXWIW_OUTPUT_FORMAT`
567
+ - `EXWIW_IDS_<ID_SPACE>` for each ID space given values (`--ids=org=...` becomes `EXWIW_IDS_ORG`)
568
+
569
+ ## MongoDB
570
+
571
+ `--adapter=mongodb` exports JSON Lines for `mongoimport`. Setup, options and the differences from the SQL adapters are in [docs/mongodb.md](docs/mongodb.md).
1273
572
 
1274
573
  ## Development
1275
574