exwiw 1.1.6 → 1.1.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +6 -0
- data/README.md +317 -1018
- data/docs/mongodb.md +4 -0
- data/docs/scope-id-set-join-notes.md +12 -0
- data/lib/exwiw/adapter/mongodb_adapter.rb +12 -0
- data/lib/exwiw/adapter/mysql_adapter.rb +18 -15
- data/lib/exwiw/adapter/postgresql_adapter.rb +26 -16
- data/lib/exwiw/adapter/sqlite_adapter.rb +16 -13
- data/lib/exwiw/adapter.rb +30 -1
- data/lib/exwiw/query_ast.rb +34 -1
- data/lib/exwiw/query_ast_builder.rb +177 -44
- data/lib/exwiw/version.rb +1 -1
- metadata +2 -1
data/README.md
CHANGED
|
@@ -1,25 +1,8 @@
|
|
|
1
1
|
# Exwiw
|
|
2
2
|
|
|
3
|
-
Export What I Want (
|
|
3
|
+
Export What I Want (exwiw) exports part of a database as SQL `INSERT` files: the rows related to the records you name, with sensitive columns masked. It is meant for building a development database that looks like production, without copying all of production or maintaining hand-made seed data.
|
|
4
4
|
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
Most of case in developing a software, There is no better choice than the same data in production.
|
|
8
|
-
You might make well-crafted data, but it's very very hard to maintain.
|
|
9
|
-
|
|
10
|
-
If you find the way to maintain the data for develoment env, then exwiw might be a solution for that.
|
|
11
|
-
|
|
12
|
-
- Export the full database and mask data and import to another database.
|
|
13
|
-
- Setup some system to replicate and mask data in real-time to another database.
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
You want to export only the data you want to export.
|
|
17
|
-
|
|
18
|
-
## Features
|
|
19
|
-
|
|
20
|
-
- Export the full list of INSERT sql for the specified conditions.
|
|
21
|
-
- Provide serveral masking options for sensitive columns.
|
|
22
|
-
- Provide config generator for ActiveRecord, for Mongoid, and from a live database connection (any application, any language).
|
|
5
|
+
Each table is described in a JSON schema config: its columns, how to mask them, and its `belongs_to` relations. Given a target table and ids, exwiw follows those relations to decide which rows of every other table to export. The schema config can be generated from ActiveRecord or Mongoid models, or from a live database.
|
|
23
6
|
|
|
24
7
|
## Installation
|
|
25
8
|
|
|
@@ -27,46 +10,29 @@ You want to export only the data you want to export.
|
|
|
27
10
|
bundle add exwiw
|
|
28
11
|
```
|
|
29
12
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
If bundler is not being used to manage dependencies, install the gem by executing:
|
|
33
|
-
|
|
34
|
-
```bash
|
|
35
|
-
gem install exwiw
|
|
36
|
-
```
|
|
37
|
-
|
|
38
|
-
## Supported Databases
|
|
13
|
+
You usually want `require: false` on the Gemfile entry. Without bundler, run `gem install exwiw`.
|
|
39
14
|
|
|
40
|
-
|
|
41
|
-
- postgresql
|
|
42
|
-
- sqlite
|
|
43
|
-
- mongodb (see [MongoDB support](docs/mongodb.md))
|
|
15
|
+
## Supported databases
|
|
44
16
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
17
|
+
- MySQL
|
|
18
|
+
- PostgreSQL
|
|
19
|
+
- SQLite
|
|
20
|
+
- MongoDB (see [MongoDB support](docs/mongodb.md))
|
|
49
21
|
|
|
50
|
-
Set `EXWIW_MYSQL_DRIVER=trilogy` (or `mysql2`) to
|
|
51
|
-
is useful when the `mysql2` gem is linked against a `libmysqlclient` that can no
|
|
52
|
-
longer load the server's auth plugin — e.g. a MySQL 9.x client drops the
|
|
53
|
-
`mysql_native_password` plugin and raises `Authentication plugin
|
|
54
|
-
'mysql_native_password' cannot be loaded` on connect. The pure-Ruby `trilogy`
|
|
55
|
-
driver implements that auth handshake itself and sidesteps the issue.
|
|
22
|
+
For MySQL, exwiw uses the `mysql2` gem if it is available and `trilogy` otherwise; pass `--adapter=mysql` either way. Set `EXWIW_MYSQL_DRIVER=trilogy` (or `mysql2`) to choose one. `trilogy` helps when `mysql2` is linked against a client library that cannot load the server's auth plugin, such as a MySQL 9.x client connecting to a server that uses `mysql_native_password`.
|
|
56
23
|
|
|
57
24
|
## Usage
|
|
58
25
|
|
|
59
26
|
exwiw has three subcommands:
|
|
60
27
|
|
|
61
|
-
- `export` (default)
|
|
62
|
-
- `explain
|
|
63
|
-
- `schema generate|check|tidy --from-db
|
|
28
|
+
- `export` (the default): write the dump files.
|
|
29
|
+
- `explain`: print the queries `export` would run, with their `EXPLAIN` output.
|
|
30
|
+
- `schema generate|check|tidy --from-db`: maintain the schema config from a live database. See [Non-Rails applications](#non-rails-applications-exwiw-schema----from-db).
|
|
64
31
|
|
|
65
32
|
### `exwiw export`
|
|
66
33
|
|
|
67
34
|
```bash
|
|
68
|
-
#
|
|
69
|
-
# pass database password as an environment variable 'DATABASE_PASSWORD'
|
|
35
|
+
# The database password is read from DATABASE_PASSWORD.
|
|
70
36
|
exwiw \
|
|
71
37
|
--adapter=mysql \
|
|
72
38
|
--host=localhost \
|
|
@@ -75,62 +41,42 @@ exwiw \
|
|
|
75
41
|
--database=app_production \
|
|
76
42
|
--schema-dir=exwiw/schema \
|
|
77
43
|
--target-table=shops \
|
|
78
|
-
--ids=1 \
|
|
79
|
-
--output-dir=dump
|
|
80
|
-
--log-level=info
|
|
44
|
+
--ids=1,2 \
|
|
45
|
+
--output-dir=dump
|
|
81
46
|
```
|
|
82
47
|
|
|
83
|
-
|
|
48
|
+
This exports the `shops` rows with id 1 and 2 and the rows of other tables related to them. `--schema-dir` reads every JSON file in the directory.
|
|
84
49
|
|
|
85
|
-
| `--target-table` | `--ids` |
|
|
50
|
+
| `--target-table` | `--ids` | What is exported |
|
|
86
51
|
|---|---|---|
|
|
87
|
-
| given | given |
|
|
88
|
-
| omitted | given | Scope-column mode if
|
|
89
|
-
| omitted | omitted | Every table
|
|
52
|
+
| given | given | The target rows by primary key, and their related rows. If the table declares a `scope_column`, [scope-column mode](#scope-column-mode) is used instead |
|
|
53
|
+
| omitted | given | [Scope-column mode](#scope-column-mode). An error if no table declares a `scope_column` (SQL adapters only) |
|
|
54
|
+
| omitted | omitted | Every table in full |
|
|
90
55
|
|
|
91
|
-
|
|
56
|
+
The output directory (`dump/` by default) is emptied before each run. When it already has files and stdin is a terminal, exwiw asks before removing them.
|
|
92
57
|
|
|
93
|
-
|
|
94
|
-
# dump all tables
|
|
95
|
-
exwiw \
|
|
96
|
-
--adapter=postgresql \
|
|
97
|
-
--host=localhost \
|
|
98
|
-
--port=5432 \
|
|
99
|
-
--user=reader \
|
|
100
|
-
--database=app_production \
|
|
101
|
-
--schema-dir=exwiw/schema \
|
|
102
|
-
--output-dir=dump
|
|
103
|
-
```
|
|
58
|
+
The output files are:
|
|
104
59
|
|
|
105
|
-
|
|
60
|
+
- `insert-000-schema.sql`: `CREATE TABLE IF NOT EXISTS ...` for every table. Run it first to create an empty database.
|
|
61
|
+
- `insert-{idx}-{table}.sql`: one per table. A file may depend on files with a smaller `idx`, so import them in order.
|
|
106
62
|
|
|
107
|
-
|
|
63
|
+
exwiw writes no `DELETE` statements. Import into an empty database, or clear the target's rows yourself.
|
|
108
64
|
|
|
109
|
-
|
|
110
|
-
- `dump/insert-{idx}-{table_name}.sql`
|
|
65
|
+
### Restoring the dump
|
|
111
66
|
|
|
112
|
-
|
|
113
|
-
so you should import the dump in order.
|
|
67
|
+
`insert-000-schema.sql` is created with the database's own tools (`mysqldump`, `pg_dump`, or the sqlite3 driver), so `mysqldump` or `pg_dump` must be on `PATH`. Set `EXWIW_MYSQLDUMP` to use a specific `mysqldump`, for example an 8.0 one when a 9.x `mysqldump` cannot authenticate against the server.
|
|
114
68
|
|
|
115
|
-
|
|
69
|
+
The schema file is rewritten so that running it again is harmless (`IF NOT EXISTS`, and PostgreSQL constraints and triggers that are skipped if they already exist). For MySQL, `DEFINER` clauses are removed so that a managed MySQL instance accepts the views and triggers.
|
|
116
70
|
|
|
117
|
-
|
|
71
|
+
MySQL data files turn off foreign key checks (`FOREIGN_KEY_CHECKS=0`). PostgreSQL data files set `session_replication_role = 'replica'`, which turns off both foreign key checks and triggers. This setting needs superuser (`rds_superuser` on RDS); without it a `WARNING` is printed and triggers fire. The setting stays on for the rest of the connection, so run `SET session_replication_role = 'origin'` if you keep using that connection. SQLite loads with its triggers active.
|
|
118
72
|
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
For `postgresql`, the extensions a managed platform installs to run the source instance itself are treated as out of target and left out of the dump entirely — currently `google_vacuum_mgmt` (Cloud SQL / AlloyDB adaptive autovacuum), `google_columnar_engine` and `google_db_advisor` (AlloyDB). They serve the source instance's operation (vacuum tuning, the in-memory columnar cache, index advice), hold no application data, are referenced by nothing in the application's own schema, and ship only with the managed platform, so a restore target outside it can never create them. Their schemas are dropped via `pg_dump --exclude-schema` and their `CREATE EXTENSION` / `COMMENT ON EXTENSION` statements — which are not schema-qualified, so no `pg_dump` filter reaches them — are removed from the output; whatever was excluded is named in the run's log.
|
|
122
|
-
|
|
123
|
-
The list is exact names, not a `google_*` prefix match: those prefixes are not reserved, so a prefix rule would also drop a schema an application legitimately owns (`google_calendar` for a Google Calendar integration) together with its tables. Every other extension is kept and wrapped in the usual warn-and-skip `DO` block, including two kinds that are also managed-platform-only:
|
|
124
|
-
|
|
125
|
-
- a third-party extension pulled in as a dependency of an excluded one (`google_db_advisor` requires `hypopg`), since that one *is* installable on a plain PostgreSQL, and
|
|
126
|
-
- an application-facing platform extension (`google_ml_integration`, `alloydb_scann`, `alloydb_ai_nl`), which the application's own SQL and DDL can name (a ScaNN index is `USING scann`) — removing its `CREATE` would strand whatever refers to it, so it warns and skips instead.
|
|
73
|
+
On PostgreSQL, extensions that only exist to run a managed instance (`google_vacuum_mgmt`, `google_columnar_engine`, `google_db_advisor`) are left out of the dump, because a database outside that platform cannot create them. Other extensions are kept, and are skipped with a warning when the target cannot create them.
|
|
127
74
|
|
|
128
75
|
### `exwiw explain`
|
|
129
76
|
|
|
130
|
-
|
|
77
|
+
Prints the query `export` would run for each table, with its `EXPLAIN` output. For the SQL adapters the SELECT is not executed. For MongoDB, see [`exwiw explain` verbosity](docs/mongodb.md#exwiw-explain-verbosity).
|
|
131
78
|
|
|
132
79
|
```bash
|
|
133
|
-
# preview the queries exwiw would run, without executing the SELECTs
|
|
134
80
|
exwiw explain \
|
|
135
81
|
--adapter=postgresql \
|
|
136
82
|
--host=localhost --port=5432 --user=reader \
|
|
@@ -139,515 +85,45 @@ exwiw explain \
|
|
|
139
85
|
--target-table=shops --ids=1
|
|
140
86
|
```
|
|
141
87
|
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
MongoDB-specific explain behavior — the configurable verbosity (`queryPlanner` / `executionStats` / `allPlansExecution`) and how scoped collections are shown — is described in [MongoDB support](docs/mongodb.md#exwiw-explain-verbosity).
|
|
145
|
-
|
|
146
|
-
### How each table is narrowed — the six scoping paths
|
|
147
|
-
|
|
148
|
-
Only the dump target itself is filtered by `--ids` directly. Every *other* table must be **scoped** — narrowed to just the rows related to the target — some other way, and a table that cannot be scoped at all is dumped in full (or, in scope-column mode, aborts the run). exwiw resolves each table through the **first** of these six paths that applies:
|
|
149
|
-
|
|
150
|
-
| # | Path | When it applies | Resulting query shape |
|
|
151
|
-
|---|------|-----------------|-----------------------|
|
|
152
|
-
| 1 | **Direct filter** | The table is the `--target-table` itself; or, in [scope-column mode](#scope-column-mode), it declares a `scope_column` | `WHERE pk IN (ids)` / `WHERE scope_column IN (ids)` |
|
|
153
|
-
| 2 | **`belongs_to` join walk** | The table reaches the target (or a scope-column table) by following its `belongs_to` edges | `WHERE fk IN (ids)` for a single hop; a chain of `JOIN`s for longer paths |
|
|
154
|
-
| 3 | **Referenced-by (automatic reverse)** | No `belongs_to` path of its own, but **exactly one** already-constrained table points at it by foreign key | Constrained to the ids that referencer's own query selects |
|
|
155
|
-
| 4 | **`reverse_scope` (declared reverse)** | Referenced by **many** scoped tables — typically a global-identity table like `users` — and the referencers are enumerated in its config | Constrained to the `UNION` of the enumerated referencers' ids |
|
|
156
|
-
| 5 | **Scoped-parent cascade** | No path or referencer, but a `belongs_to` parent is itself scoped (by any path above) | Constrained to the parent's in-scope primary keys; cascades over multiple hops |
|
|
157
|
-
| 6 | **Full dump** | Nothing relates the table to the target | All rows. In scope-column mode this **aborts** unless the table opts in with `scope_exempt: true` |
|
|
158
|
-
|
|
159
|
-
How the paths behave and interact:
|
|
160
|
-
|
|
161
|
-
1. **Direct filter.** In the default single-target mode the target is anchored on its primary key (or a custom field via the mongodb-only `--ids-field`). In [scope-column mode](#scope-column-mode) there is no single anchor: every table that declares a `scope_column` is filtered on that column directly.
|
|
162
|
-
2. **`belongs_to` join walk** — the "normal join" path. exwiw BFS-walks `belongs_to` edges to the nearest terminus (the target table, or a directly scoped table in scope-column mode) and compiles the shortest path into `INNER JOIN`s. A [polymorphic `belongs_to`](#polymorphic-belongs_to) hop additionally pins the type column, and is resolved for **every** concrete arm with the arms `UNION`ed (see [Every arm is extracted](#every-arm-is-extracted)).
|
|
163
|
-
3. **Referenced-by** handles a table with no outgoing path that is pointed *at* by a constrained child — `active_storage_blobs`, referenced by `active_storage_attachments.blob_id`, is the canonical case (see [ActiveStorage](#activestorage-has_one_attached--has_many_attached)). It is automatic but deliberately narrow: it requires a single, non-polymorphic referencer. With two or more referencers it steps aside to path 5, then 6, unless you declare `reverse_scope`.
|
|
164
|
-
4. **[`reverse_scope`](#reverse-scope-for-multi-referencer-tables-reverse_scope)** is the declared, multi-referencer form of path 3: the config enumerates which referencers' (already scoped) queries feed the id set. Unscoped arms are skipped with a warning rather than widening the dump.
|
|
165
|
-
5. **Scoped-parent cascade** rescues satellites: a table whose only link is a `belongs_to` toward a hub that is itself scoped (e.g. via referenced-by or `reverse_scope`) is constrained to that parent's in-scope ids. The cascade recurses hop by hop (each level requires a single unambiguous scopable parent) and stops on `belongs_to` cycles.
|
|
166
|
-
6. **Full dump** is the fallback for a genuinely unrelated table — intended for reference/master data. Single-target mode dumps it in full (with a warning when an ambiguous cascade was the reason); scope-column mode refuses to run instead, unless the table is explicitly marked [`scope_exempt: true`](#scope_exempt-intentional-full-dump) (Rails-managed tables are exempt automatically).
|
|
167
|
-
|
|
168
|
-
Paths 3–5 all materialize their id set once and probe it via a `JOIN` on a `SELECT DISTINCT` derived table rather than `IN (subquery)` — see [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery). Scope-column mode classifies every table up front with these same paths (`:direct` / `:via_path` / `:referenced_by` / `:via_scoped_parent` / `:exempt` / `:unscopable` in `QueryAstBuilder#scope_category`) and aborts before extracting anything if any table lands on `:unscopable`. The MongoDB adapter follows the same model, except id sets are captured at runtime while parent collections stream instead of being expressed as SQL subqueries — see [MongoDB support](docs/mongodb.md).
|
|
169
|
-
|
|
170
|
-
### Scope-column mode
|
|
171
|
-
|
|
172
|
-
The default `--target-table` extraction assumes the schema converges on a single
|
|
173
|
-
root: every table is reached by walking `belongs_to` toward that one table. Some
|
|
174
|
-
schemas are not shaped that way — many independent top-level tables each carry the
|
|
175
|
-
*same* scope/tenant column (e.g. `tenant_id`, `business_entity_id`), and a foreign
|
|
176
|
-
key that **cannot be joined** (most importantly a cross-database `belongs_to`,
|
|
177
|
-
whose join is impossible but whose FK column is still filterable) is not reached at
|
|
178
|
-
all. Choosing one table as `--target-table` would leave the others unrelated to it,
|
|
179
|
-
and an unrelated table is dumped in full — a problem if it holds personal data.
|
|
180
|
-
|
|
181
|
-
Scope-column mode handles this shape: instead of anchoring on one table's primary
|
|
182
|
-
key, **every table is filtered by a shared column** whose values are `--ids` (tables
|
|
183
|
-
keyed by an unrelated kind of id can use their own [ID space](#per-table-scope_column-and-id-spaces)).
|
|
184
|
-
Declare that column per table in the schema config with `scope_column:`:
|
|
185
|
-
|
|
186
|
-
```json
|
|
187
|
-
{
|
|
188
|
-
"name": "shops",
|
|
189
|
-
"primary_key": "id",
|
|
190
|
-
"scope_column": "business_entity_id",
|
|
191
|
-
"columns": [{ "name": "id" }, { "name": "name" }, { "name": "business_entity_id" }]
|
|
192
|
-
}
|
|
193
|
-
```
|
|
194
|
-
|
|
195
|
-
Then pass the scope values as `--ids`, without `--target-table`:
|
|
196
|
-
|
|
197
|
-
```bash
|
|
198
|
-
exwiw \
|
|
199
|
-
--adapter=postgresql \
|
|
200
|
-
--host=localhost --port=5432 --user=reader \
|
|
201
|
-
--database=app_production \
|
|
202
|
-
--schema-dir=exwiw/schema \
|
|
203
|
-
--ids=42,43 \
|
|
204
|
-
--output-dir=dump
|
|
205
|
-
```
|
|
206
|
-
|
|
207
|
-
Because the schema declares a `scope_column`, exwiw runs in scope-column mode: the
|
|
208
|
-
`--ids` (`42,43`) are **`business_entity_id` values, not shop primary keys**, and
|
|
209
|
-
`shops` is scoped by `business_entity_id IN (42,43)` like every other scoped table.
|
|
210
|
-
Where extraction starts is decided by the schema, so `--target-table` is not
|
|
211
|
-
needed. If no table declares a `scope_column`, `--ids` alone is an error: name the
|
|
212
|
-
table the ids belong to with `--target-table` (single-target mode), or declare a
|
|
213
|
-
`scope_column`. Tables marked `ignore: true` do not count.
|
|
214
|
-
|
|
215
|
-
Naming a scoped table as `--target-table` (`--target-table=shops --ids=42,43`)
|
|
216
|
-
selects the same mode and extracts the same rows; the target is *not* used as a
|
|
217
|
-
primary-key anchor. (A table that declares a `scope_column` therefore can no longer
|
|
218
|
-
be single-extracted by primary key.)
|
|
219
|
-
|
|
220
|
-
Each table is resolved as follows:
|
|
221
|
-
|
|
222
|
-
- **Declares the scope column** (`scope_column:`, or carries the global column of
|
|
223
|
-
the deprecated `--scope-column` flag) → `WHERE scope_column IN (ids)`.
|
|
224
|
-
- **Does not, but `belongs_to` reaches a table that does** → exwiw joins up to the
|
|
225
|
-
nearest such table and applies the scope filter there (the same join machinery
|
|
226
|
-
the single-target mode uses).
|
|
227
|
-
- **`belongs_to` a parent that is itself scoped but carries no scope column of its
|
|
228
|
-
own** → exwiw constrains this table to the parent's in-scope ids by joining it to
|
|
229
|
-
the parent's scoped query, materialized as a derived table
|
|
230
|
-
(`JOIN (SELECT DISTINCT parent.pk … FROM <parent's scoped query>) … ON fk = …`).
|
|
231
|
-
This covers a *hub* table that has no scope column and is scoped only because an
|
|
232
|
-
extractable child references it (see referenced-by below): the hub's other
|
|
233
|
-
`belongs_to` children ride along to just the in-scope rows instead of being dumped
|
|
234
|
-
in full. The parent itself may be scoped the same way, so this **cascades across
|
|
235
|
-
multiple hops** (each a single unambiguous scopable parent) and the derived-table
|
|
236
|
-
JOINs nest correspondingly; the recursion terminates on a genuine `belongs_to`
|
|
237
|
-
cycle (a table already on the path is left `:unscopable` rather than looped on).
|
|
238
|
-
(See [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery) for the
|
|
239
|
-
materialization rationale.)
|
|
240
|
-
- **Cannot be scoped at all** (no scope column and no path to one) → exwiw
|
|
241
|
-
**aborts** and lists the offending tables, so an unscoped table is never silently
|
|
242
|
-
dumped in full. For each, either declare a `scope_column`, add a `belongs_to`
|
|
243
|
-
path, set `ignore: true` to skip it, or mark it `scope_exempt: true` (below) to
|
|
244
|
-
export it in full.
|
|
245
|
-
|
|
246
|
-
> **Note — referenced-by is preferred over the hub cascade.** A table that is
|
|
247
|
-
> *both* `belongs_to` a scoped hub *and* referenced-by a constrained child is
|
|
248
|
-
> scoped to the (narrower) referenced-by id-set, not the hub cascade, so the hub's
|
|
249
|
-
> other children the child does not reference are dropped (under-scoping). To force
|
|
250
|
-
> the broader hub cascade, set `ignore: true` on the child's `belongs_to` edge that
|
|
251
|
-
> points at this table.
|
|
252
|
-
|
|
253
|
-
Scope-column mode is SQL-only (mysql / postgresql / sqlite); with the mongodb
|
|
254
|
-
adapter, `--ids` still requires `--target-collection`. It works with `exwiw
|
|
255
|
-
explain` too, which is the recommended way to preview the queries before exporting.
|
|
256
|
-
|
|
257
|
-
#### Cross-database foreign keys
|
|
258
|
-
|
|
259
|
-
The motivating case for declaring a `scope_column` is a foreign key that cannot be
|
|
260
|
-
joined: when a `belongs_to` target lives in a different database (see the
|
|
261
|
-
cross-database `belongs_to` note under the generator), that join is impossible, but
|
|
262
|
-
the foreign-key *column* is still present and can be filtered directly. Declaring
|
|
263
|
-
`scope_column: "<that foreign key>"` on the owning table scopes it by the column
|
|
264
|
-
value, with no join — `schema:generate` points this out in the ignored relation's
|
|
265
|
-
`comment`.
|
|
266
|
-
|
|
267
|
-
#### `scope_exempt` (intentional full dump)
|
|
268
|
-
|
|
269
|
-
A genuine reference/master table (no personal data) that has no scope linkage can
|
|
270
|
-
opt out of the strict check and be exported in full:
|
|
271
|
-
|
|
272
|
-
```json
|
|
273
|
-
{
|
|
274
|
-
"name": "countries",
|
|
275
|
-
"primary_key": "id",
|
|
276
|
-
"scope_exempt": true,
|
|
277
|
-
"columns": [{ "name": "id" }, { "name": "code" }]
|
|
278
|
-
}
|
|
279
|
-
```
|
|
280
|
-
|
|
281
|
-
Rails-managed tables (`schema_migrations`, `ar_internal_metadata`) are treated as
|
|
282
|
-
exempt automatically.
|
|
283
|
-
|
|
284
|
-
#### Per-table `scope_column` and ID spaces
|
|
285
|
-
|
|
286
|
-
The values of every `scope_column` belong to an **ID space**, and `--ids` gives the
|
|
287
|
-
values of each space. A table that does not declare `id_space` is in the `default`
|
|
288
|
-
space, which a plain `--ids=1,2` fills, so without any `id_space` every scoped table
|
|
289
|
-
is filtered by the same `--ids`. Each table names its own column, so a table that
|
|
290
|
-
stores that same value under a differently named column simply declares that name:
|
|
291
|
-
|
|
292
|
-
```json
|
|
293
|
-
{
|
|
294
|
-
"name": "legacy_orders",
|
|
295
|
-
"primary_key": "id",
|
|
296
|
-
"scope_column": "legacy_tenant_id",
|
|
297
|
-
"columns": [{ "name": "id" }, { "name": "legacy_tenant_id" }]
|
|
298
|
-
}
|
|
299
|
-
```
|
|
300
|
-
|
|
301
|
-
When one database holds two groups of tables that no foreign key connects, each
|
|
302
|
-
keyed by a different kind of id, put one group in a named ID space. Here `tenants`
|
|
303
|
-
are identified by integer ids and `organizations` by UUIDs (each object is its own
|
|
304
|
-
schema file):
|
|
305
|
-
|
|
306
|
-
```json
|
|
307
|
-
{ "name": "tenants", "primary_key": "id", "scope_column": "id", "columns": [{ "name": "id" }] }
|
|
308
|
-
{ "name": "organizations", "primary_key": "id", "scope_column": "id", "id_space": "org", "columns": [{ "name": "id" }] }
|
|
309
|
-
```
|
|
310
|
-
|
|
311
|
-
```bash
|
|
312
|
-
exwiw ... --target-table=tenants \
|
|
313
|
-
--ids=1,2 \
|
|
314
|
-
--ids=org=0b6f4c1e-0000-4000-8000-000000000001
|
|
315
|
-
```
|
|
316
|
-
|
|
317
|
-
`--ids=1,2` is the same as `--ids=default=1,2`. The part before the first `=` is
|
|
318
|
-
read as an ID space name only when it is shaped like one (a lowercase letter
|
|
319
|
-
followed by lowercase letters, digits or underscores), so an id that contains `=`
|
|
320
|
-
itself is passed with an explicit `default=`. Each ID space may be given once. A
|
|
321
|
-
table without a scope column is filtered by the ID space of the table it is scoped
|
|
322
|
-
through (a `belongs_to` join, `reverse_scope`, referenced-by or the parent cascade).
|
|
323
|
-
|
|
324
|
-
The run aborts before extracting anything when a table is filtered by an ID space
|
|
325
|
-
that was given no values, when values are given for an ID space no table uses, or
|
|
326
|
-
when a table reaches scoped tables of more than one ID space (otherwise the
|
|
327
|
-
`belongs_to` walk would silently settle on the nearest one).
|
|
328
|
-
|
|
329
|
-
Single `--target-table` mode and the MongoDB adapter use only the `default` space
|
|
330
|
-
and abort when a named one is given. In the config file, `ids:` takes either a list
|
|
331
|
-
(the `default` space) or a mapping from ID space name to values, such as
|
|
332
|
-
`ids: { default: [1, 2], org: [0b6f4c1e-0000-4000-8000-000000000001] }`; as with
|
|
333
|
-
the other keys, it is used only when `--ids` is not passed at all.
|
|
334
|
-
|
|
335
|
-
`scope_exempt`, `scope_column` and `id_space` are user-maintained and preserved
|
|
336
|
-
across `schema:generate` regeneration (the generators never emit them).
|
|
337
|
-
|
|
338
|
-
#### Deprecated: the `--scope-column` flag
|
|
339
|
-
|
|
340
|
-
Before per-table declarations, scope-column mode was selected with a global
|
|
341
|
-
`--scope-column=COLUMN` flag (every table filtered by that one column, `--ids` its
|
|
342
|
-
values, no `--target-table`). The flag still works — SQL-only and mutually
|
|
343
|
-
exclusive with `--target-table` — but is **deprecated** and emits a warning; prefer
|
|
344
|
-
declaring a per-table `scope_column` and dropping the flag. A per-table
|
|
345
|
-
`scope_column` takes precedence over the flag for any table that sets both.
|
|
88
|
+
`--output-dir`, `--output-format` and `--after-insert-hook` cannot be used with `explain`.
|
|
346
89
|
|
|
347
90
|
### Config file (`exwiw.yml`)
|
|
348
91
|
|
|
349
|
-
Options
|
|
350
|
-
|
|
351
|
-
**Options passed on the CLI always take precedence over the config file** — the config only fills in options you did not pass. This lets you commit the stable settings (which schema to read, output format, ...) while still varying the environment-specific connection details per invocation.
|
|
92
|
+
Options can be kept in a YAML file passed with `--config=PATH`. Without `--config`, `exwiw.yml` (or `exwiw.yaml`) in the current directory is used if it exists. Options passed on the command line take precedence.
|
|
352
93
|
|
|
353
94
|
```yaml
|
|
354
|
-
# exwiw.yml — keep at the project root, alongside exwiw/schema/
|
|
355
95
|
adapter: postgresql
|
|
356
96
|
schema_dir: exwiw/schema
|
|
357
97
|
output_dir: dump
|
|
358
98
|
output_format: insert # insert | copy
|
|
359
99
|
after_insert_hook: hooks/seed.rb
|
|
360
100
|
log_level: info # debug | info
|
|
361
|
-
# target_table
|
|
362
|
-
# (ids may map ID spaces to values; see "Per-table scope_column and ID spaces")
|
|
363
|
-
# mongodb_query_timeout_ms: 30000 # global query timeout (mongodb only)
|
|
101
|
+
# target_table, ids, ids_field and scope_column can also be set here.
|
|
364
102
|
```
|
|
365
103
|
|
|
366
|
-
With the file above, only the connection details need to be supplied on the CLI:
|
|
367
|
-
|
|
368
104
|
```bash
|
|
369
105
|
DATABASE_PASSWORD=... exwiw \
|
|
370
106
|
--host=localhost --port=5432 --user=reader --database=app_production \
|
|
371
107
|
--target-table=shops --ids=1
|
|
372
108
|
```
|
|
373
109
|
|
|
374
|
-
|
|
110
|
+
- Connection settings (`host`, `port`, `user`, `database`, `uri`, `password`) are rejected, so they stay out of a committed file. `adapter` is allowed.
|
|
111
|
+
- Relative paths are resolved from the config file's directory, not the current directory.
|
|
112
|
+
- Unknown keys are rejected. Keys that only apply to `export` are ignored by `explain` and `schema`, so all subcommands can share one file.
|
|
113
|
+
- MongoDB-only keys (`explain_verbosity`, `mongodb_query_timeout_ms`, `parallel_workers`) are described in [MongoDB support](docs/mongodb.md).
|
|
375
114
|
|
|
376
|
-
|
|
377
|
-
- **Relative paths in the config (`schema_dir`, `output_dir`, `after_insert_hook`) are resolved relative to the config file's own directory**, not the current working directory. So with the config at the project root, `schema_dir: exwiw/schema` reads naturally, and an absolute `--config=/path/to/exwiw.yml` works no matter where you run from. (CLI path flags remain relative to the current directory — each source resolves relative to where it is written.) Absolute paths are used as-is.
|
|
378
|
-
- Unknown keys are rejected so a typo surfaces immediately. (`insert_only`, whose behavior was removed, is the one grandfathered key: accepted and ignored with a warning.)
|
|
379
|
-
- Export-only keys (`output_dir`, `output_format`, `after_insert_hook`) are ignored when running `explain` or `schema`, so a single config file can be shared by every subcommand.
|
|
380
|
-
- `explain_verbosity` sets the mongodb `explain` verbosity (`queryPlanner` | `executionStats` | `allPlansExecution`, default `queryPlanner`); the `EXWIW_MONGODB_EXPLAIN_VERBOSITY` env var overrides it. Ignored by the SQL adapters and by `export`. See [MongoDB support](docs/mongodb.md#exwiw-explain-verbosity).
|
|
381
|
-
- `mongodb_query_timeout_ms` sets the global, server-enforced query timeout (mongodb only); the `--mongodb-query-timeout-ms` CLI flag overrides it. Ignored by the SQL adapters. See [MongoDB support](docs/mongodb.md).
|
|
382
|
-
|
|
383
|
-
### Generator
|
|
384
|
-
|
|
385
|
-
The config generator is provided as a Rake task.
|
|
386
|
-
|
|
387
|
-
```bash
|
|
388
|
-
# generate table schema under exwiw/schema/
|
|
389
|
-
bundle exec rake exwiw:schema:generate
|
|
390
|
-
```
|
|
391
|
-
|
|
392
|
-
The output directory is resolved in this order:
|
|
393
|
-
|
|
394
|
-
1. the `EXWIW_SCHEMA_DIR_PATH` environment variable, if set;
|
|
395
|
-
2. otherwise `schema_dir` from the config file (`exwiw.yml` / `exwiw.yaml` in the current directory), so the generator and the `exwiw` CLI share one location without repeating the path;
|
|
396
|
-
3. otherwise the `exwiw/schema` default.
|
|
397
|
-
|
|
398
|
-
```sh
|
|
399
|
-
EXWIW_SCHEMA_DIR_PATH=custom_directory bundle exec rake exwiw:schema:generate
|
|
400
|
-
```
|
|
401
|
-
|
|
402
|
-
As with the CLI, a relative `schema_dir` in the config file is resolved relative to the config file's own directory.
|
|
403
|
-
|
|
404
|
-
An application with more than one schema source — ActiveRecord models and Mongoid documents, say —
|
|
405
|
-
should give each source its own directory (`EXWIW_SCHEMA_DIR_PATH`), because every task judges the
|
|
406
|
-
whole directory against the one source it reads. Left sharing a directory, `tidy` / `check` see the
|
|
407
|
-
other source's configs as belonging to tables and collections that no longer exist, and report or
|
|
408
|
-
remove them.
|
|
409
|
-
|
|
410
|
-
#### Safe mode (masking new columns by default)
|
|
411
|
-
|
|
412
|
-
A migration that adds a column would otherwise leave `schema:generate` emitting it unmasked, so
|
|
413
|
-
it starts being exported the moment the config is regenerated — before anyone has judged whether
|
|
414
|
-
it holds personal data. So `schema:generate` runs in **safe mode by default**: every column the
|
|
415
|
-
config does not have yet is emitted **masked** and flagged
|
|
416
|
-
[`needs_mask_decision: true`](#needs_mask_decision).
|
|
417
|
-
|
|
418
|
-
Columns already in the config keep whatever they say — the merge that preserves `replace_with` /
|
|
419
|
-
`comment` / `ignore` preserves a resolved decision too — so in practice this marks exactly the
|
|
420
|
-
columns a migration just added.
|
|
421
|
-
|
|
422
|
-
```bash
|
|
423
|
-
bundle exec rake exwiw:schema:generate # safe mode
|
|
424
|
-
EXWIW_NEW_COLUMNS=plain bundle exec rake exwiw:schema:generate # opt out
|
|
425
|
-
```
|
|
426
|
-
|
|
427
|
-
Opting out is for the **first-time bootstrap** of a config, where every column of every table is
|
|
428
|
-
new and safe mode would flag the whole thing at once. Use it nowhere else: a column committed
|
|
429
|
-
under `plain` carries no flag, so nothing afterwards can tell it apart from one whose masking was
|
|
430
|
-
decided.
|
|
431
|
-
|
|
432
|
-
A column that has a **default of its own** is masked with that default: it is a value the column
|
|
433
|
-
provably holds, and it is what the application treats as neutral, so masking a `default: true`
|
|
434
|
-
flag does not quietly turn the feature off for every row in the dump. A default the database
|
|
435
|
-
computes (`now()`) is not a constant and does not count, and neither does a JSON object — `{...}`
|
|
436
|
-
in a mask is a column placeholder, so those fall back to `{}`. Otherwise the mask depends on the column
|
|
437
|
-
type: `masked-{primary key}` for text (with `@example.com` appended when the column name mentions
|
|
438
|
-
mail, so it stays a valid address), `0` for numbers, `false` for booleans, a fixed date/timestamp,
|
|
439
|
-
and `{}` for JSON. Text always takes the template rather than its default, since the mask has to
|
|
440
|
-
vary per row. Three kinds of
|
|
441
|
-
column are flagged but deliberately **not** masked:
|
|
442
|
-
|
|
443
|
-
- **The primary key, and the foreign keys/types the `belongs_tos` join on.** Masking them
|
|
444
|
-
would break the joins and leave the dump referencing rows that were never exported.
|
|
445
|
-
- **Types no constant safely fits** — `uuid`, `binary`, enums, array columns (which report their
|
|
446
|
-
member type, so a scalar default would not fit), and text columns too short to hold the masked
|
|
447
|
-
value. An invalid default would fail the restore the dump feeds, which is worse than exporting
|
|
448
|
-
the column while the flag keeps the change from being merged.
|
|
449
|
-
- **Columns covered by a unique index**, unless the mask varies per row (the text masks do, via
|
|
450
|
-
the primary key). A constant would collapse every row onto one value and break the restore with
|
|
451
|
-
a duplicate key.
|
|
452
|
-
|
|
453
|
-
`schema:generate_mongoid` runs in safe mode too, on the same `EXWIW_NEW_COLUMNS=plain` opt-out. The
|
|
454
|
-
masks come from the Mongoid field type (`String`, `Integer`, `Float`, `BigDecimal`,
|
|
455
|
-
`Mongoid::Boolean`, `Date`, `Time` / `DateTime` / `ActiveSupport::TimeWithZone`); a field of any
|
|
456
|
-
other type — `Hash`, `Array`, a typeless field, a BSON type — is flagged but not masked, as is a
|
|
457
|
-
field covered by a unique index unless its mask varies per document. The structural fields are the
|
|
458
|
-
`_id` primary key, the STI discriminator (`_type`) and every `belongs_to` foreign key of the
|
|
459
|
-
collection: flagged, never masked. The foreign keys are read from the models, so the ones the
|
|
460
|
-
config itself drops (a polymorphic `belongs_to`, a `belongs_to` on an embedded document) are
|
|
461
|
-
covered too.
|
|
462
|
-
|
|
463
|
-
#### Tidying stale config (`schema:tidy`)
|
|
464
|
-
|
|
465
|
-
`schema:generate` adds and updates config files for the tables it finds, but it never deletes the config file of a table that has been dropped from the application. To reconcile the existing config against the current schema, run:
|
|
466
|
-
|
|
467
|
-
```bash
|
|
468
|
-
bundle exec rake exwiw:schema:tidy
|
|
469
|
-
```
|
|
470
|
-
|
|
471
|
-
`schema:tidy` compares the config files already on disk with the **live database** (read through the database connection, not the models) and removes only what no longer exists there:
|
|
472
|
-
|
|
473
|
-
- a config file whose table has been dropped from the database is **deleted**, and
|
|
474
|
-
- columns recorded in a surviving table's config that the table no longer has are **dropped** from that file.
|
|
475
|
-
|
|
476
|
-
Because it reads the database directly, a table that still exists in the database but has lost (or never had) an ActiveRecord model is **kept** — only a table that is genuinely gone is removed. (This is the deliberate counterpart to `generate`, which is model-driven and only ever adds what the models know about.)
|
|
477
|
-
|
|
478
|
-
It respects `EXWIW_SCHEMA_DIR_PATH` and the per-database subdirectory layout in the same way as `schema:generate`. Unlike `generate`, `tidy` never adds or regenerates entries — every surviving table/column (including hand-edited `comment` / `ignore` / `replace_with`) is left untouched, so it is safe to run on a customized config. The task prints which tables and columns it removed (or that the config was already tidy). Stale `belongs_tos` are not pruned by `tidy`; rerun `schema:generate` to refresh those.
|
|
479
|
-
|
|
480
|
-
#### Checking the config against the schema
|
|
481
|
-
|
|
482
|
-
`schema:check` reports how the committed config differs from what the application would
|
|
483
|
-
generate now — without writing anything, so it can run on a working tree it must not modify:
|
|
484
|
-
|
|
485
|
-
```bash
|
|
486
|
-
bundle exec rake exwiw:schema:check
|
|
487
|
-
```
|
|
488
|
-
|
|
489
|
-
It regenerates into a throwaway copy of the config directory (safe mode + `tidy`) and prints
|
|
490
|
-
the comparison as JSON, then exits non-zero when anything needs attention:
|
|
491
|
-
|
|
492
|
-
```json
|
|
493
|
-
{
|
|
494
|
-
"added_tables": [],
|
|
495
|
-
"added_columns": ["users.contact_email"],
|
|
496
|
-
"removed_tables": [],
|
|
497
|
-
"removed_columns": ["orders.legacy_flag"],
|
|
498
|
-
"changed_tables": ["orders", "users"],
|
|
499
|
-
"needs_mask_decision": ["orders.memo"],
|
|
500
|
-
"stale_tables": [],
|
|
501
|
-
"stale_columns": ["orders.legacy_flag"]
|
|
502
|
-
}
|
|
503
|
-
```
|
|
504
|
-
|
|
505
|
-
`added_*` / `removed_*` / `changed_tables` mean the config no longer matches the schema — run
|
|
506
|
-
`schema:generate` and `schema:tidy` to reconcile it. `needs_mask_decision` lists the columns
|
|
507
|
-
whose masking nobody has decided on yet (see [the flag](#needs_mask_decision)). `stale_tables` /
|
|
508
|
-
`stale_columns` are the subset of the removals an extraction would actually trip over — a
|
|
509
|
-
non-ignored config still naming a table or column the schema no longer has, so the export's
|
|
510
|
-
SELECT would fail; removals of `ignore: true` entries (and of a rails-managed table's columns,
|
|
511
|
-
which are dumped as `SELECT *`) stay out of them. They drive the exit code only under
|
|
512
|
-
[`--fail-on=stale`](#non-rails-applications-exwiw-schema----from-db). The exit code
|
|
513
|
-
makes it usable as a CI check that keeps a schema change from being merged until both are
|
|
514
|
-
resolved; the JSON is stable and sorted, so it can be posted as-is. In a multi-database app each
|
|
515
|
-
entry is prefixed with its database (`primary/users.email`), so the same table name in two
|
|
516
|
-
databases stays distinct.
|
|
517
|
-
|
|
518
|
-
Set `EXWIW_SCHEMA_CHECK_OUTPUT=<path>` to have the same JSON written to a file, which spares a
|
|
519
|
-
caller from assuming stdout carries nothing else (application boot is free to print).
|
|
520
|
-
|
|
521
|
-
A Mongoid config directory has its own task, `schema:check_mongoid`, with the same output,
|
|
522
|
-
the same `EXWIW_SCHEMA_CHECK_OUTPUT` file and the same exit code — it just regenerates through
|
|
523
|
-
`MongoidSchemaGenerator` (safe mode + `tidy_mongoid`) instead. Collections and fields are
|
|
524
|
-
reported under the same keys as tables and columns. An application that cannot be loaded to
|
|
525
|
-
generate from its models at all can run the same check against its database instead: see
|
|
526
|
-
[Non-Rails applications](#non-rails-applications-exwiw-schema----from-db).
|
|
527
|
-
|
|
528
|
-
#### Multiple databases
|
|
529
|
-
|
|
530
|
-
If the application uses Rails' multiple-database support (`connects_to`), `schema:generate` buckets models by the database they connect to and writes each database's config files into its own subdirectory of the output directory, named after the database config name (`primary`, `analytics`, ...):
|
|
531
|
-
|
|
532
|
-
```
|
|
533
|
-
exwiw/schema/
|
|
534
|
-
primary/
|
|
535
|
-
shops.json
|
|
536
|
-
users.json
|
|
537
|
-
schema_migrations.json
|
|
538
|
-
analytics/
|
|
539
|
-
analytics_events.json
|
|
540
|
-
```
|
|
541
|
-
|
|
542
|
-
Each database keeps its own Rails migration history, so a `schema_migrations` (and `ar_internal_metadata`) entry is emitted under every database that contains one — the example above shows `primary/schema_migrations.json` and would also produce `analytics/schema_migrations.json` when the analytics database has its own migration table. Single-database applications are unaffected and continue to write files flat into the output directory.
|
|
543
|
-
|
|
544
|
-
A `belongs_to` whose target model lives in a *different* database (e.g. a `primary` model referencing an `analytics` one) cannot be joined: each database is exported on its own connection and into its own subdirectory, so the target table is absent from the directory this config is loaded with. `schema:generate` detects such a relation (by comparing the owning and target models' database config names) and emits it with `ignore: true` and `ignore_type: "cross_database"`, recording why in the `comment`; the relation is then dropped from extraction at load time, while the foreign-key column itself is still exported as a plain column. Polymorphic associations are handled per target, so only the targets that cross a database boundary are ignored. The task also prints a summary of every cross-database `belongs_to` it ignored. **To extract across such a boundary, declare `scope_column: "<foreign_key>"` on the owning table (see [scope-column mode](#scope-column-mode)) so its rows are filtered by the foreign-key value directly** — there is no join, so the cross-database boundary is not a problem there.
|
|
545
|
-
|
|
546
|
-
**Limitations**
|
|
547
|
-
|
|
548
|
-
- The rails-managed table *names* are resolved from the global `ActiveRecord::Base.schema_migrations_table_name` / `internal_metadata_table_name` accessors, which are shared across all connections. A per-database override of these names is not detected, so such a table will be missing from that database's generated configs.
|
|
549
|
-
|
|
550
|
-
#### Mongoid applications
|
|
551
|
-
|
|
552
|
-
For MongoDB applications backed by [Mongoid](https://www.mongodb.com/docs/mongoid/), a separate rake task introspects Mongoid document models and emits `MongodbCollectionConfig` files:
|
|
553
|
-
|
|
554
|
-
```bash
|
|
555
|
-
bundle exec rake exwiw:schema:generate_mongoid
|
|
556
|
-
bundle exec rake exwiw:schema:tidy_mongoid # delete the config of a collection no model stores into
|
|
557
|
-
bundle exec rake exwiw:schema:check_mongoid # report the difference without changing anything
|
|
558
|
-
```
|
|
559
|
-
|
|
560
|
-
What it derives from each model (fields, `belongs_tos`, `embedded_in`, STI handling), how to annotate constructs exwiw cannot represent with `ignore` / `ignore_type`, and the `EXWIW_SKIP_UNSUPPORTED=1` bootstrap flag are all documented in [MongoDB support](docs/mongodb.md#generating-config-from-mongoid-models).
|
|
561
|
-
|
|
562
|
-
`tidy_mongoid` reconciles against the *models* rather than a live connection (MongoDB has no schema to read, and a collection exists only once something is written to it): a config file whose collection no model stores into any more is deleted, and nothing else is touched — fields already track the models through `generate_mongoid`. `check_mongoid` runs both into a throwaway copy, so it never writes to the working tree.
|
|
563
|
-
|
|
564
|
-
#### Non-Rails applications (`exwiw schema ... --from-db`)
|
|
565
|
-
|
|
566
|
-
The rake tasks above read the application's models, which requires loading the application — so
|
|
567
|
-
they are only available where that is possible. For an application written in any other language
|
|
568
|
-
(or a Ruby one exwiw cannot boot), the same three operations are available on the CLI, reading
|
|
569
|
-
the **database** instead of the models:
|
|
570
|
-
|
|
571
|
-
```bash
|
|
572
|
-
# generate / refresh the config from the live schema
|
|
573
|
-
exwiw schema generate --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
|
|
574
|
-
|
|
575
|
-
# report how the committed config differs from the database (exits 1 when it needs work)
|
|
576
|
-
exwiw schema check --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
|
|
577
|
-
|
|
578
|
-
# remove tables/columns/relations the database no longer has
|
|
579
|
-
exwiw schema tidy --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
|
|
580
|
-
```
|
|
581
|
-
|
|
582
|
-
- `--from-db` is **required**: the schema source is always stated explicitly rather than inferred.
|
|
583
|
-
- mysql and postgresql only. sqlite is not supported, and a MongoDB schema lives in the
|
|
584
|
-
application rather than the database (use `schema:generate_mongoid`).
|
|
585
|
-
- The usual connection flags and `DATABASE_PASSWORD` apply, and `adapter` / `schema_dir` may come
|
|
586
|
-
from [the config file](#config-file-exwiwyml) instead (`--schema-dir` wins). Unlike `export`,
|
|
587
|
-
an empty or absent `DATABASE_PASSWORD` is accepted, since these commands are commonly pointed
|
|
588
|
-
at a CI database that runs with trust authentication.
|
|
589
|
-
- `generate` creates the schema directory if it does not exist; `check` and `tidy` require it.
|
|
590
|
-
- Everything else behaves as the rake tasks do: [safe mode](#safe-mode-masking-new-columns-by-default)
|
|
591
|
-
is on unless `EXWIW_NEW_COLUMNS=plain`, `check` prints the same JSON report, honours
|
|
592
|
-
`EXWIW_SCHEMA_CHECK_OUTPUT`, and exits 1 when the config needs attention. A check that could
|
|
593
|
-
not *run* (an unreachable database, a malformed config) exits with a different status, so CI
|
|
594
|
-
can tell the two apart.
|
|
595
|
-
- `check` also accepts `--fail-on=stale` for use as a pre-extraction gate: the exit code then
|
|
596
|
-
tracks only the report's `stale_tables` / `stale_columns` — a non-ignored config still naming
|
|
597
|
-
a table or column the schema no longer has, which is exactly the drift that would fail the
|
|
598
|
-
export's SELECT. Additions and unresolved `needs_mask_decision` flags stay visible in the
|
|
599
|
-
report but do not stop the run, so a schema migration that merely *adds* a column does not
|
|
600
|
-
block extraction. The default (`--fail-on=any`) is the CI behavior above, unchanged.
|
|
601
|
-
`--fail-on` is command-line only (not a config-file key — a gate flag belongs to the
|
|
602
|
-
invocation, not the committed config) and is meaningful only on `schema check`: the other
|
|
603
|
-
schema verbs reject it, and `export` ignores it like the other schema-only flags.
|
|
604
|
-
- One run covers one database — the connection addresses one — so the files are written flat into
|
|
605
|
-
the schema directory. There is no per-database subdirectory layout here; a second database is a
|
|
606
|
-
second run against a second connection.
|
|
607
|
-
|
|
608
|
-
The database is read through its catalog only: tables, columns and their types/defaults, primary
|
|
609
|
-
keys, unique indexes and foreign keys. Views are skipped, since they hold no rows of their own,
|
|
610
|
-
and a table with no primary key is emitted with `ignore: true` and a comment saying what to add
|
|
611
|
-
to export it (exwiw identifies and joins rows by primary key). Following that advice sticks: a
|
|
612
|
-
`primary_key` written by hand — on a table the database reports none for, or one column of a
|
|
613
|
-
composite key that identifies a row on its own — is kept by later runs, which then treat the table
|
|
614
|
-
as an ordinary one rather than re-imposing the signpost `type` / `comment`.
|
|
615
|
-
|
|
616
|
-
**`belongs_tos` are only ever added, never rewritten.** A foreign-key constraint is weaker
|
|
617
|
-
evidence than an application model: plenty of schemas express a relation only in application code,
|
|
618
|
-
and a `belongs_to` is the path extraction follows to reach a table, so silently dropping one
|
|
619
|
-
narrows the dump. Regeneration therefore keeps every relation the config already declares —
|
|
620
|
-
in its existing order, with its `comment` / `ignore` / `ignore_type` — and appends only the
|
|
621
|
-
foreign-key-backed relations that are not there yet. A hand-written relation's foreign-key column
|
|
622
|
-
is treated as structural too, so safe mode never masks it. Removing a relation is `tidy`'s job:
|
|
623
|
-
it drops a `belongs_to` whose target table no longer exists in the database (an `ignore: true`
|
|
624
|
-
entry is kept, since it records a decision), alongside the tables and columns that are gone.
|
|
625
|
-
|
|
626
|
-
So a relation the database does not know about is declared once, by hand, and survives from then
|
|
627
|
-
on:
|
|
115
|
+
### Output format
|
|
628
116
|
|
|
629
|
-
|
|
630
|
-
{
|
|
631
|
-
"name": "orders",
|
|
632
|
-
"primary_key": "id",
|
|
633
|
-
"belongs_tos": [{
|
|
634
|
-
"table_name": "buyers",
|
|
635
|
-
"foreign_key": "buyer_id",
|
|
636
|
-
"comment": "enforced in the application; no foreign key in the database"
|
|
637
|
-
}]
|
|
638
|
-
}
|
|
639
|
-
```
|
|
117
|
+
For PostgreSQL, `--output-format=copy` writes `COPY ... FROM stdin` instead of `INSERT`, which loads much faster. Import it with `psql -d app_dev -f dump/insert-001-shops.sql`.
|
|
640
118
|
|
|
641
|
-
|
|
119
|
+
## Schema config
|
|
642
120
|
|
|
643
|
-
|
|
121
|
+
Each table has one JSON file:
|
|
644
122
|
|
|
645
123
|
```json
|
|
646
124
|
{
|
|
647
125
|
"name": "users",
|
|
648
126
|
"primary_key": "id",
|
|
649
|
-
"filter": "users.id > 0",
|
|
650
|
-
"bulk_insert_chunk_size": 1000,
|
|
651
127
|
"belongs_tos": [{
|
|
652
128
|
"table_name": "companies",
|
|
653
129
|
"foreign_key": "company_id"
|
|
@@ -663,97 +139,26 @@ This is an example of the one table schema:
|
|
|
663
139
|
}
|
|
664
140
|
```
|
|
665
141
|
|
|
666
|
-
|
|
667
|
-
|
|
668
|
-
#### Unknown keys are rejected
|
|
142
|
+
`belongs_tos` decides which rows are exported (see [How each table is narrowed](#how-each-table-is-narrowed)), and each column can be masked (see [Masking](#masking)). The other table-level keys are:
|
|
669
143
|
|
|
670
|
-
|
|
144
|
+
- `filter`: an SQL condition added to the table's query, such as `"access_logs.created_at > '2025-01-01'"`. It is added to every query that joins this table, so it also narrows the tables that depend on it, which can leave their foreign keys pointing at rows that were not exported. Qualify column names with the table name. A filter reduces the rows returned, not necessarily the rows read; see [Batched extraction](#batched-extraction-batch_scope) for that.
|
|
145
|
+
- `bulk_insert_chunk_size`: the maximum rows per `INSERT` statement (10,000 by default), to stay under limits such as MySQL's `max_allowed_packet`.
|
|
146
|
+
- `ignore`, `comment`: see below.
|
|
671
147
|
|
|
672
|
-
|
|
673
|
-
|
|
674
|
-
### Output format
|
|
148
|
+
### Unknown keys are rejected
|
|
675
149
|
|
|
676
|
-
|
|
677
|
-
|
|
678
|
-
The generated file uses tab-separated values with PostgreSQL's text-format escaping (`\N` for NULL, `\\` for backslash, etc.). Import with `psql`:
|
|
679
|
-
|
|
680
|
-
```bash
|
|
681
|
-
psql -d app_dev -f dump/insert-001-shops.sql
|
|
682
|
-
```
|
|
683
|
-
|
|
684
|
-
`--output-format=copy` is only supported with the `postgresql` adapter.
|
|
685
|
-
|
|
686
|
-
### After-insert hook
|
|
687
|
-
|
|
688
|
-
`--after-insert-hook=PATH` runs a post-processing hook **after** all per-table insert files have been written. The hook can be either a Ruby file (`.rb`) or any executable script (e.g. `.sh`).
|
|
689
|
-
|
|
690
|
-
**Ruby hook (`.rb`)**: provides a tiny DSL with these builtins:
|
|
691
|
-
|
|
692
|
-
- `cli_options` — Hash of all parsed CLI options (e.g. `cli_options.fetch(:ids)` returns the `--ids` array of the `default` ID space)
|
|
693
|
-
- `ids_for(id_space = "default")` — the values the run was scoped by for that ID space (e.g. `ids_for("org")` for `--ids=org=...`)
|
|
694
|
-
- `insert_sql(template)` — appends an ERB-rendered string to a buffer. After the hook finishes, the buffer is concatenated and written to `insert-{N+1}-after_insert.{ext}` where `{N+1}` is one past the last per-table insert file. For the MongoDB adapter the equivalent alias `insert_jsonl(template)` is available; output goes to `insert-{N+1}-after_insert.jsonl`. Multiple `insert_sql` calls in a single hook are joined with `"\n"` into the same file. If no `insert_sql` call is made, no file is created.
|
|
695
|
-
- `insert_jsonl(collection, template)` — **MongoDB adapter only**. SQL statements name their table in-band, but JSONL documents do not — the import convention derives the target collection from the filename — so the two-argument form writes the ERB-rendered extended-JSON lines to the named collection's own `insert-NNN-<collection>.jsonl` file, importable with the same `mongoimport --collection <collection>` convention as the per-collection dump files. Multiple calls targeting the same collection are appended (joined with `"\n"`) into that collection's file; distinct collections get one file each, numbered sequentially after the last per-collection dump file (the collection-less `after_insert` buffer, when also used, keeps `{N+1}` and the collection files follow it). Calling this form with a SQL adapter raises an error.
|
|
696
|
-
|
|
697
|
-
Example `hooks/seed_default_users.rb`:
|
|
698
|
-
|
|
699
|
-
```ruby
|
|
700
|
-
insert_sql <<~SQL
|
|
701
|
-
-- seed default users for tenants <%= cli_options.fetch(:ids).join(',') %>
|
|
702
|
-
<%- cli_options.fetch(:ids).each do |tenant_id| -%>
|
|
703
|
-
INSERT INTO users (tenant_id, email) VALUES (<%= tenant_id %>, 'default@example.com');
|
|
704
|
-
<%- end -%>
|
|
705
|
-
SQL
|
|
706
|
-
```
|
|
707
|
-
|
|
708
|
-
MongoDB example seeding two collections (`insert-{N+1}-users.jsonl` and `insert-{N+2}-posts.jsonl`):
|
|
709
|
-
|
|
710
|
-
```ruby
|
|
711
|
-
insert_jsonl 'users', <<~JSONL
|
|
712
|
-
<%- cli_options.fetch(:ids).each do |shop_id| -%>
|
|
713
|
-
{"shop_id":{"$oid":"<%= shop_id %>"},"email":"default@example.com"}
|
|
714
|
-
<%- end -%>
|
|
715
|
-
JSONL
|
|
716
|
-
insert_jsonl 'posts', '{"title":"welcome"}'
|
|
717
|
-
```
|
|
718
|
-
|
|
719
|
-
**Shell hook**: anything other than `.rb` is exec'd as a child process. It is a pure side-effect hook — exwiw does not capture its stdout. The hook receives these env vars and inherits `DATABASE_PASSWORD` from the parent:
|
|
720
|
-
|
|
721
|
-
- `EXWIW_OUTPUT_DIR`, `EXWIW_SCHEMA_DIR`
|
|
722
|
-
- `EXWIW_DATABASE_ADAPTER`, `EXWIW_DATABASE_HOST`, `EXWIW_DATABASE_PORT`, `EXWIW_DATABASE_USER`, `EXWIW_DATABASE_NAME`
|
|
723
|
-
- `EXWIW_TARGET_TABLE`, `EXWIW_IDS` (comma-separated, the `default` ID space), `EXWIW_OUTPUT_FORMAT`
|
|
724
|
-
- `EXWIW_IDS_<ID_SPACE>` for each ID space given values, with the name uppercased (`--ids=org=...` becomes `EXWIW_IDS_ORG`). `EXWIW_IDS_DEFAULT` is set only when the `default` space is given, while `EXWIW_IDS` is always set
|
|
725
|
-
|
|
726
|
-
A non-zero exit code from the shell hook aborts exwiw.
|
|
727
|
-
|
|
728
|
-
Note: Ruby hooks are evaluated via `instance_eval` inside the exwiw process — only pass paths you trust.
|
|
150
|
+
A config with a key exwiw does not know fails to load, with an error naming the key and the file. A typo such as `reverse_scop` would otherwise silently disable what it was meant to do. Use `comment` for notes.
|
|
729
151
|
|
|
730
152
|
### Ignore a table
|
|
731
153
|
|
|
732
|
-
|
|
733
|
-
|
|
734
|
-
```json
|
|
735
|
-
{
|
|
736
|
-
"name": "audit_logs",
|
|
737
|
-
"primary_key": "id",
|
|
738
|
-
"ignore": true,
|
|
739
|
-
"belongs_tos": [],
|
|
740
|
-
"columns": [{ "name": "id" }]
|
|
741
|
-
}
|
|
742
|
-
```
|
|
743
|
-
|
|
744
|
-
Constraints:
|
|
154
|
+
`"ignore": true` on a table stops its data from being exported. Its `CREATE TABLE` is still written to `insert-000-schema.sql`.
|
|
745
155
|
|
|
746
|
-
-
|
|
747
|
-
-
|
|
748
|
-
- `ignore: true` is preserved by `exwiw:schema:generate` regenerations (the receiver value wins over the auto-generated config).
|
|
749
|
-
- On a MongoDB [embedded config](docs/mongodb.md#excluding-an-embedded-path-ignore-true), which is masked through its parent rather than dumped on its own, `ignore: true` excludes the embedded path from the parent's projection — the subdocument is left out of the dump entirely.
|
|
156
|
+
- A table that has a `belongs_to` to an ignored table fails to load. Remove that `belongs_to`, or ignore it too.
|
|
157
|
+
- An ignored table cannot be the `--target-table`.
|
|
750
158
|
|
|
751
159
|
### Ignore / annotate a column or `belongs_to`
|
|
752
160
|
|
|
753
|
-
|
|
754
|
-
|
|
755
|
-
- `comment` — a free-form note. Purely informational; exwiw never reads it.
|
|
756
|
-
- `ignore: true` — drops that entry from extraction. An ignored column/field is excluded from the `SELECT` and the generated `INSERT` (the column still exists in the target schema, since the DDL comes from the source database — exwiw just does not copy its data). An ignored `belongs_to` is removed from dependency ordering and query building, so the relation is not traversed.
|
|
161
|
+
Entries in `columns` and `belongs_tos` accept `comment` (a note exwiw never reads) and `ignore: true`. An ignored column is left out of the `SELECT` and the `INSERT`, so the restored rows get the column's database default, or `NULL` if it has none. On a `NOT NULL` column without a default, the `INSERT` can fail. An ignored `belongs_to` is not followed.
|
|
757
162
|
|
|
758
163
|
```json
|
|
759
164
|
{
|
|
@@ -761,7 +166,7 @@ Individual `columns` (SQL) / `fields` (MongoDB) and `belongs_tos` entries accept
|
|
|
761
166
|
"primary_key": "id",
|
|
762
167
|
"belongs_tos": [
|
|
763
168
|
{ "table_name": "companies", "foreign_key": "company_id" },
|
|
764
|
-
{ "table_name": "audit_logs", "foreign_key": "log_id", "ignore": true, "comment": "
|
|
169
|
+
{ "table_name": "audit_logs", "foreign_key": "log_id", "ignore": true, "comment": "not needed in development" }
|
|
765
170
|
],
|
|
766
171
|
"columns": [
|
|
767
172
|
{ "name": "id" },
|
|
@@ -770,153 +175,70 @@ Individual `columns` (SQL) / `fields` (MongoDB) and `belongs_tos` entries accept
|
|
|
770
175
|
}
|
|
771
176
|
```
|
|
772
177
|
|
|
773
|
-
|
|
178
|
+
### Hand-edited keys survive regeneration
|
|
774
179
|
|
|
775
|
-
|
|
180
|
+
Regenerating the config (see [Generating the schema config](#generating-the-schema-config)) keeps what you wrote by hand. A column already in the config is kept as it is, `comment` and `ignore` on a table or `belongs_to` are kept, and so are the keys the generators never write (`filter`, `bulk_insert_chunk_size`, `scope_column`, `id_space`, `scope_exempt`, `reverse_scope`, `batch_scope`).
|
|
776
181
|
|
|
777
|
-
|
|
182
|
+
## How each table is narrowed
|
|
778
183
|
|
|
779
|
-
|
|
780
|
-
nobody has decided on yet:
|
|
184
|
+
Only the target table is filtered by `--ids` directly. exwiw narrows every other table to the rows related to the target, using the first of these rules that applies:
|
|
781
185
|
|
|
782
|
-
|
|
783
|
-
|
|
784
|
-
|
|
186
|
+
| # | Rule | When it applies | Result |
|
|
187
|
+
|---|------|-----------------|--------|
|
|
188
|
+
| 1 | Direct filter | The table is the `--target-table`, or declares a `scope_column` in [scope-column mode](#scope-column-mode) | `WHERE pk IN (ids)` or `WHERE scope_column IN (ids)` |
|
|
189
|
+
| 2 | `belongs_to` path | Following `belongs_to` reaches a table of rule 1 | Joined along the shortest path |
|
|
190
|
+
| 3 | Referenced by one table | No path, but exactly one narrowed table has a foreign key to it | Only the rows that table points at |
|
|
191
|
+
| 4 | [`reverse_scope`](#reverse-scope-for-multi-referencer-tables-reverse_scope) | Several narrowed tables point at it, and the config lists them | Only the rows the listed tables point at |
|
|
192
|
+
| 5 | Narrowed parent | No path, but a `belongs_to` parent is narrowed by a rule above | Only the rows belonging to the parent's exported rows, over any number of hops |
|
|
193
|
+
| 6 | Full dump | Nothing relates the table to the target | All rows, meant for master data. Scope-column mode stops with an error instead, unless the table has [`scope_exempt: true`](#scope_exempt-intentional-full-dump) |
|
|
785
194
|
|
|
786
|
-
|
|
787
|
-
|
|
788
|
-
|
|
789
|
-
|
|
790
|
-
|
|
791
|
-
config — reports the ones that still carry it, so CI can keep a pull request red until each is
|
|
792
|
-
resolved. Resolving it means removing the key — after keeping the mask (ideally recording why
|
|
793
|
-
in `comment`), replacing it with a real masking rule, dropping `replace_with` to export the
|
|
794
|
-
raw value, or setting `ignore: true`.
|
|
195
|
+
- Rule 3 covers tables like `active_storage_blobs`, which nothing in the table itself links to the target. It applies only when the referencing `belongs_to` is not polymorphic.
|
|
196
|
+
- Rule 5 applies only when exactly one parent is narrowed, and stops at a `belongs_to` cycle.
|
|
197
|
+
- If a table matches both rule 3 and rule 5, rule 3 wins and the result can miss rows that rule 5 would have kept; set `ignore: true` on the referencing `belongs_to` to use rule 5.
|
|
198
|
+
- Whichever rule narrows a table with a `belongs_to` to itself, the ancestors of its rows are added (see [Self-referencing `belongs_to`](#self-referencing-belongs_to-tree-tables)).
|
|
199
|
+
- MongoDB follows the same rules, collecting the ids while it reads the parent collections instead of using SQL subqueries. It does not add ancestors.
|
|
795
200
|
|
|
796
|
-
|
|
797
|
-
`schema:generate` does not bring it back.
|
|
201
|
+
When a rule picks the full dump because the relation was ambiguous, exwiw logs a warning. `exwiw explain` is the easiest way to see what each table resolved to. Rules 3 to 5 appear in it as a `JOIN` on a `SELECT DISTINCT` subquery; [`docs/scope-id-set-join-notes.md`](docs/scope-id-set-join-notes.md) explains why.
|
|
798
202
|
|
|
799
203
|
### Polymorphic `belongs_to`
|
|
800
204
|
|
|
801
|
-
A
|
|
802
|
-
|
|
803
|
-
- `foreign_type` — the type column on *this* table (e.g. `reviewable_type`).
|
|
804
|
-
- `type_value` — the value stored in that column for this target (e.g. `"Product"`), i.e. the target model's `polymorphic_name`.
|
|
205
|
+
A polymorphic association (`belongs_to :reviewable, polymorphic: true`) is written as one `belongs_to` per target table, each with the type column (`foreign_type`) and the value it holds for that target (`type_value`):
|
|
805
206
|
|
|
806
207
|
```json
|
|
807
208
|
{
|
|
808
209
|
"name": "reviews",
|
|
809
210
|
"primary_key": "id",
|
|
810
211
|
"belongs_tos": [
|
|
811
|
-
{
|
|
812
|
-
|
|
813
|
-
"foreign_key": "reviewable_id",
|
|
814
|
-
"foreign_type": "reviewable_type",
|
|
815
|
-
"type_value": "Product"
|
|
816
|
-
},
|
|
817
|
-
{
|
|
818
|
-
"table_name": "shops",
|
|
819
|
-
"foreign_key": "reviewable_id",
|
|
820
|
-
"foreign_type": "reviewable_type",
|
|
821
|
-
"type_value": "Shop"
|
|
822
|
-
}
|
|
212
|
+
{ "table_name": "products", "foreign_key": "reviewable_id", "foreign_type": "reviewable_type", "type_value": "Product" },
|
|
213
|
+
{ "table_name": "shops", "foreign_key": "reviewable_id", "foreign_type": "reviewable_type", "type_value": "Shop" }
|
|
823
214
|
],
|
|
824
215
|
"columns": [{ "name": "id" }, { "name": "reviewable_type" }, { "name": "reviewable_id" }]
|
|
825
216
|
}
|
|
826
217
|
```
|
|
827
218
|
|
|
828
|
-
`exwiw:schema:generate`
|
|
829
|
-
|
|
830
|
-
At dump time, when a polymorphic `belongs_to` lies on the path to the dump target, exwiw constrains **both** the foreign key and the type column, so only rows of the matching type are extracted. For example, dumping `products` pulls only reviews whose `reviewable_type = 'Product'`:
|
|
219
|
+
`exwiw:schema:generate` writes these entries from the models' `has_many ..., as:` declarations. A non-polymorphic `belongs_to` leaves out `foreign_type` and `type_value`.
|
|
831
220
|
|
|
832
|
-
|
|
833
|
-
SELECT reviews.* FROM reviews
|
|
834
|
-
WHERE reviews.reviewable_id IN (/* products subquery */)
|
|
835
|
-
AND reviews.reviewable_type = 'Product'
|
|
836
|
-
```
|
|
837
|
-
|
|
838
|
-
The same type filter is applied on the join path when the polymorphic table is an intermediate hop rather than the directly-dumped table.
|
|
221
|
+
Following such a `belongs_to` also checks the type column, so dumping `products` exports only the reviews with `reviewable_type = 'Product'`.
|
|
839
222
|
|
|
840
223
|
#### Every arm is extracted
|
|
841
224
|
|
|
842
|
-
|
|
843
|
-
|
|
844
|
-
```sql
|
|
845
|
-
SELECT comments.* FROM comments
|
|
846
|
-
JOIN (
|
|
847
|
-
SELECT DISTINCT exwiw_scope_src_0.id AS exwiw_scope_id
|
|
848
|
-
FROM (
|
|
849
|
-
SELECT comments.id FROM comments
|
|
850
|
-
JOIN posts ON comments.commentable_id = posts.id AND comments.commentable_type = 'Post'
|
|
851
|
-
JOIN shops ON posts.shop_id = shops.id AND shops.tenant_id = 't1'
|
|
852
|
-
UNION
|
|
853
|
-
SELECT comments.id FROM comments
|
|
854
|
-
JOIN pages ON comments.commentable_id = pages.id AND comments.commentable_type = 'Page'
|
|
855
|
-
JOIN shops ON pages.shop_id = shops.id AND shops.tenant_id = 't1'
|
|
856
|
-
) AS exwiw_scope_src_0
|
|
857
|
-
) AS exwiw_scope_ids_0
|
|
858
|
-
ON comments.id = exwiw_scope_ids_0.exwiw_scope_id
|
|
859
|
-
```
|
|
860
|
-
|
|
861
|
-
`UNION`, not `OR`, because each arm joins a *different* table: OR-ing them in one `WHERE` would need outer joins, whereas each arm is a self-contained query of exactly the shape a single-arm table already produces. It rides on the existing scope id-set machinery, so the id set is materialized once (see [Why a JOIN, not `IN (subquery)`](#why-a-join-not-in-subquery)) and, on mysql, into a session `TEMPORARY TABLE`.
|
|
862
|
-
|
|
863
|
-
Notes:
|
|
864
|
-
|
|
865
|
-
- **Only arms that reach the scope are included.** An arm whose target has no scope of its own is dropped, never widened — an unscoped arm would pull in every tenant's rows. An arm marked `"ignore": true` is dropped as usual, before any of this.
|
|
866
|
-
- An arm's target does not need a `belongs_to` path to a scoped table: if it is scoped by other means (referenced-by, [`reverse_scope`](#reverse-scope-for-multi-referencer-tables-reverse_scope), or the parent cascade) the arm probes that query's ids instead, still pinned by the type column.
|
|
867
|
-
- An arm whose target is scoped **through this same table** (e.g. `active_storage_blobs`, narrowed by referenced-by from `active_storage_attachments`, appearing as an `ActiveStorage::Blob` arm of those same attachments) is dropped: adopting it would make the two tables scope each other and leave the referenced table short of rows the join table kept — a dangling foreign key on import.
|
|
868
|
-
- **Arms are grouped by the type column (`foreign_type`), not by the foreign key.** Rails keeps every arm's id in one `<name>_id` column, which is what the generators emit, but a hand-written config may give each arm its own foreign key while one type column still selects between them — e.g. `owner_type` choosing between `user_id` and `team_id`. Those are arms of the same discriminator, so each is resolved and joined on its own key. Two *independent* polymorphic associations on one table have distinct type columns and therefore stay in distinct groups.
|
|
869
|
-
- **The route the walk picks decides whether the arms are unioned.** The route is the shortest one; in single-target mode a `belongs_to` pointing at the dump target itself is taken first. When the route leaves through a polymorphic arm, every arm of that type column is unioned. When it leaves through a non-polymorphic `belongs_to`, the table is joined along that route alone and the polymorphic arms are not unioned in, so a row reachable only through an arm is not extracted. A plain `belongs_to` is not preferred over a shorter polymorphic route. A single arm is emitted as a plain JOIN.
|
|
870
|
-
- A table with no route of its own falls back to the parent cascade, which does prefer plain parents. A scopable plain parent is used when there is exactly one. Only when there is none are the polymorphic `belongs_to`s consulted: the arms of one type column that point at tables scoped by other means (referenced-by, `reverse_scope`, or the cascade) are unioned, each pinned by the type column. Multiple independent polymorphic associations that are all scopable are as ambiguous as multiple plain parents, so the table is left unscopable (dumped in full with a warning in single-target mode). A table scoped this way also counts as a constrained child in the [referenced-by](#how-each-table-is-narrowed--the-six-scoping-paths) detection of the tables it points at, so a parent that had a single constrained referencer may now have two and be narrowed by the next path instead.
|
|
871
|
-
|
|
872
|
-
In single `--target-table` mode the same union is built. An arm that points at the dump target itself compares its foreign key with `--ids`, and the other arms join or probe their way to the target as above. One difference: an arm target that is constrained **only** by the automatic referenced-by detection does not count. Such a parent is extracted just to keep a child's foreign key valid; it does not own the rows that point at it. So dumping `products` still pulls only `reviewable_type = 'Product'` reviews, even though the product's shop is extracted too, while dumping `shops` pulls the shop's own reviews **and** the reviews of its products:
|
|
873
|
-
|
|
874
|
-
```sql
|
|
875
|
-
SELECT reviews.* FROM reviews
|
|
876
|
-
JOIN (
|
|
877
|
-
SELECT DISTINCT exwiw_scope_src_0.id AS exwiw_scope_id
|
|
878
|
-
FROM (
|
|
879
|
-
SELECT reviews.id FROM reviews
|
|
880
|
-
JOIN products ON reviews.reviewable_id = products.id
|
|
881
|
-
AND products.shop_id = 1 AND reviews.reviewable_type = 'Product'
|
|
882
|
-
UNION
|
|
883
|
-
SELECT reviews.id FROM reviews
|
|
884
|
-
WHERE reviews.reviewable_id = 1 AND reviews.reviewable_type = 'Shop'
|
|
885
|
-
) AS exwiw_scope_src_0
|
|
886
|
-
) AS exwiw_scope_ids_0
|
|
887
|
-
ON reviews.id = exwiw_scope_ids_0.exwiw_scope_id
|
|
888
|
-
```
|
|
225
|
+
When a table's path to the target goes through a polymorphic `belongs_to`, exwiw follows every entry with the same `foreign_type` and exports the union of their rows. Dumping `shops` exports the shop's own reviews and the reviews of its products.
|
|
889
226
|
|
|
890
|
-
|
|
227
|
+
- An entry whose target table is not narrowed is skipped, so a polymorphic `belongs_to` never widens the dump.
|
|
228
|
+
- If the shortest path leaves through a non-polymorphic `belongs_to`, only that path is used.
|
|
229
|
+
- In single-target mode, an entry is skipped when its table is narrowed only by rule 3. Those rows are exported to keep foreign keys valid, not because they own the rows that point at them. So dumping `products` still exports only the `Product` reviews, even though the product's shop is exported too. An entry whose table has a `reverse_scope` is followed.
|
|
891
230
|
|
|
892
231
|
### ActiveStorage (`has_one_attached` / `has_many_attached`)
|
|
893
232
|
|
|
894
|
-
ActiveStorage
|
|
233
|
+
ActiveStorage needs no configuration:
|
|
895
234
|
|
|
896
|
-
-
|
|
897
|
-
-
|
|
898
|
-
|
|
899
|
-
```sql
|
|
900
|
-
SELECT active_storage_blobs.* FROM active_storage_blobs
|
|
901
|
-
JOIN (
|
|
902
|
-
SELECT DISTINCT exwiw_scope_src_0.blob_id AS exwiw_scope_id
|
|
903
|
-
FROM (
|
|
904
|
-
SELECT active_storage_attachments.blob_id FROM active_storage_attachments
|
|
905
|
-
WHERE active_storage_attachments.record_id IN (/* owner subquery */)
|
|
906
|
-
AND active_storage_attachments.record_type = '...'
|
|
907
|
-
) AS exwiw_scope_src_0
|
|
908
|
-
) AS exwiw_scope_ids_0
|
|
909
|
-
ON active_storage_blobs.id = exwiw_scope_ids_0.exwiw_scope_id
|
|
910
|
-
```
|
|
911
|
-
|
|
912
|
-
`active_storage_variant_records` also references blobs, but since it has no path of its own to the dump target it doesn't constrain anything and is ignored as a referencer — blobs stays narrowed to the attachment-referenced ids. (A parent referenced by *multiple* constrained children currently falls back to dumping all of its rows.)
|
|
913
|
-
- **`active_storage_variant_records`** holds derivative variant-tracking rows that ActiveStorage regenerates lazily, and it too has no path to the dump target — left alone it would land in the "no relation → dump all" branch and, worse, its `blob_id` could point at blobs outside the narrowed set above (a foreign-key violation on import). `exwiw:schema:generate` therefore emits it with **`ignore: true`** (and drops it from the attachments `record` polymorphic expansion so nothing carries a dangling reference to it), so its data is skipped while the DDL is still written. Remove `ignore` from the generated config if you really need to export it.
|
|
235
|
+
- `active_storage_attachments` is a polymorphic `belongs_to :record`, so only the attachments of exported records are exported.
|
|
236
|
+
- `active_storage_blobs` is narrowed by rule 3 to the blobs those attachments point at.
|
|
237
|
+
- `active_storage_variant_records` is generated with `ignore: true`, because ActiveStorage recreates its rows when needed and its `blob_id` could point at blobs that were not exported. Remove `ignore` if you need it.
|
|
914
238
|
|
|
915
239
|
### Reverse scope for multi-referencer tables (`reverse_scope`)
|
|
916
240
|
|
|
917
|
-
|
|
918
|
-
|
|
919
|
-
`reverse_scope` opts such a table into **multi-referencer** reverse scoping: you enumerate the referencers whose own (already scoped) extraction queries should be `UNION`'d into the id set the table is constrained to. It is a user-owned key (never emitted by `schema:generate`, preserved across regeneration like `scope_exempt`/`scope_column`):
|
|
241
|
+
A table such as `users` often has no `belongs_to` toward the target, while many narrowed tables point at it. Rule 3 does not apply when there is more than one, so it would be dumped in full, with every tenant's users. `reverse_scope` lists the tables and columns that point at it, and the table is narrowed to the values those tables export:
|
|
920
242
|
|
|
921
243
|
```json
|
|
922
244
|
{
|
|
@@ -926,121 +248,100 @@ The automatic reverse extraction above narrows a table referenced by **exactly o
|
|
|
926
248
|
"via": [
|
|
927
249
|
{ "table": "customers", "column": "user_id" },
|
|
928
250
|
{ "table": "staff", "column": "user_id" },
|
|
929
|
-
{ "table": "
|
|
251
|
+
{ "table": "members", "column": "legacy_user_id" }
|
|
930
252
|
]
|
|
931
253
|
},
|
|
932
254
|
"columns": [{ "name": "id" }, { "name": "name" }]
|
|
933
255
|
}
|
|
934
256
|
```
|
|
935
257
|
|
|
936
|
-
|
|
258
|
+
- `column` is given explicitly, so a column with a non-standard name, or one without a `belongs_to`, works.
|
|
259
|
+
- List only narrowed tables. A table in `via` that is not narrowed would add every value, so it is skipped with a warning.
|
|
260
|
+
- A table in `via` may itself be narrowed by its own `reverse_scope`.
|
|
261
|
+
- By default the values are matched against the primary key. Set `reverse_scope.column` to match another column, such as `{ "column": "code", "via": [{ "table": "contracts", "column": "rate_code" }] }`.
|
|
262
|
+
- Tables that `belongs_to` the reverse-scoped table are narrowed by rule 5 and need no config.
|
|
937
263
|
|
|
938
|
-
|
|
939
|
-
SELECT users.* FROM users
|
|
940
|
-
JOIN (
|
|
941
|
-
SELECT DISTINCT exwiw_scope_src_0.user_id AS exwiw_scope_id
|
|
942
|
-
FROM (
|
|
943
|
-
SELECT customers.user_id FROM customers WHERE <customers' scope> AND customers.user_id IS NOT NULL
|
|
944
|
-
UNION
|
|
945
|
-
SELECT staff.user_id FROM staff WHERE <staff' scope> AND staff.user_id IS NOT NULL
|
|
946
|
-
UNION
|
|
947
|
-
SELECT business_entity_customers.kantan_yoyaku_user_id FROM business_entity_customers
|
|
948
|
-
WHERE <…' scope> AND business_entity_customers.kantan_yoyaku_user_id IS NOT NULL
|
|
949
|
-
) AS exwiw_scope_src_0
|
|
950
|
-
) AS exwiw_scope_ids_0
|
|
951
|
-
ON users.id = exwiw_scope_ids_0.exwiw_scope_id
|
|
952
|
-
```
|
|
264
|
+
### Self-referencing `belongs_to` (tree tables)
|
|
953
265
|
|
|
954
|
-
|
|
266
|
+
When a table has a `belongs_to` to itself, such as `categories.parent_id`, narrowing it would drop the parents of the rows it keeps. exwiw also keeps every ancestor of those rows, up to the root. Declaring the `belongs_to` is enough.
|
|
955
267
|
|
|
956
|
-
-
|
|
957
|
-
-
|
|
958
|
-
-
|
|
959
|
-
-
|
|
960
|
-
-
|
|
961
|
-
- **Satellites need no config.** A table that `belongs_to` the reverse-scoped table (e.g. `end_users.id → users.id`, or `identities.user_id → users.id`) tightens to the kept ids automatically through the normal cascade — only the reverse-scoped table itself declares `reverse_scope`. The cascade is **multi-hop**, so a table several `belongs_to` hops below the reverse-scoped table (e.g. `end_user_profiles → end_users → users`) also tightens automatically, with no config of its own.
|
|
962
|
-
- Works in both single-target and scope-column mode. In single-target mode there is no scope-column pre-flight (`validate_scope!`), so a satellite the cascade cannot resolve to a single scopable parent (e.g. it `belongs_to` two scopable hubs) is dumped in full with a warning rather than aborting. Polymorphic foreign keys are not eligible as anchors (the named `column` is always a concrete column).
|
|
963
|
-
- **The MongoDB adapter supports `reverse_scope` too** — same config shape and semantics, but the id set is captured at runtime instead of being emitted as a `UNION` subquery. See [`reverse_scope` on collections](docs/mongodb.md#reverse_scope-on-collections) under MongoDB support.
|
|
268
|
+
- Ancestors are kept regardless of the table's `filter` and scope, so foreign keys stay valid. If a tree spans tenants, the dump can include another tenant's ancestors.
|
|
269
|
+
- Tables below the tree (`category_notes`) also keep the rows of the ancestors, and tables the tree points at keep what the ancestors point at. An ancestor with many rows hanging off it brings them all; set `ignore: true` on that `belongs_to` to leave them out.
|
|
270
|
+
- The walk stops at a `NULL` or missing parent and at cycles in the data. Polymorphic self-references are not supported.
|
|
271
|
+
- On MySQL this needs 8.0 or later, and a tree deeper than `cte_max_recursion_depth` (1000 by default) fails.
|
|
272
|
+
- [Batched](#batched-extraction-batch_scope) tables do not get ancestors added; exwiw warns about this. MongoDB does not add ancestors either.
|
|
964
273
|
|
|
965
|
-
###
|
|
274
|
+
### Scope-column mode
|
|
966
275
|
|
|
967
|
-
|
|
276
|
+
Single-target mode assumes every table reaches one target table through `belongs_to`. In many multi-tenant schemas, tables instead each carry a tenant column (`tenant_id`) and are not all connected to one root. Picking one table as the target would dump the unrelated tables in full.
|
|
968
277
|
|
|
969
|
-
|
|
970
|
-
|
|
971
|
-
|
|
278
|
+
In scope-column mode, each table names the column that holds the tenant id, and `--ids` are values of that column:
|
|
279
|
+
|
|
280
|
+
```json
|
|
281
|
+
{
|
|
282
|
+
"name": "shops",
|
|
283
|
+
"primary_key": "id",
|
|
284
|
+
"scope_column": "tenant_id",
|
|
285
|
+
"columns": [{ "name": "id" }, { "name": "name" }, { "name": "tenant_id" }]
|
|
286
|
+
}
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
```bash
|
|
290
|
+
exwiw \
|
|
291
|
+
--adapter=postgresql \
|
|
292
|
+
--host=localhost --port=5432 --user=reader \
|
|
293
|
+
--database=app_production \
|
|
294
|
+
--schema-dir=exwiw/schema \
|
|
295
|
+
--ids=42,43 \
|
|
296
|
+
--output-dir=dump
|
|
972
297
|
```
|
|
973
298
|
|
|
974
|
-
|
|
299
|
+
Here `42,43` are `tenant_id` values, not shop ids. Passing `--target-table=shops` gives the same result.
|
|
975
300
|
|
|
976
|
-
|
|
301
|
+
Tables without a `scope_column` are narrowed by the rules above. A table that none of them can narrow stops the run before anything is exported, and the error lists those tables. For each, declare a `scope_column`, add a `belongs_to`, set `ignore: true`, or set `scope_exempt: true`.
|
|
977
302
|
|
|
978
|
-
|
|
303
|
+
Scope-column mode is for the SQL adapters only. Use `exwiw explain` to check the queries first.
|
|
979
304
|
|
|
980
|
-
|
|
305
|
+
#### Cross-database foreign keys
|
|
981
306
|
|
|
982
|
-
|
|
983
|
-
- `type: "rails_managed_internal_metadata"` — Rails' internal metadata table (`ActiveRecord::Base.internal_metadata_table_name`).
|
|
307
|
+
A `belongs_to` to a table in another database cannot be joined, so `schema:generate` writes it with `ignore: true`. The foreign key column is still there, so declaring `scope_column: "<that foreign key>"` on the table narrows it without a join.
|
|
984
308
|
|
|
985
|
-
|
|
309
|
+
#### `scope_exempt` (intentional full dump)
|
|
986
310
|
|
|
987
|
-
A
|
|
311
|
+
A master table with no personal data and no relation to the tenant can be exported in full:
|
|
988
312
|
|
|
989
313
|
```json
|
|
990
|
-
{
|
|
991
|
-
"name": "schema_migrations",
|
|
992
|
-
"type": "rails_managed_schema_migrations",
|
|
993
|
-
"comment": "Managed internally by Rails. Tracks applied schema migrations."
|
|
994
|
-
}
|
|
314
|
+
{ "name": "countries", "primary_key": "id", "scope_exempt": true, "columns": [{ "name": "id" }, { "name": "code" }] }
|
|
995
315
|
```
|
|
996
316
|
|
|
997
|
-
|
|
998
|
-
|
|
999
|
-
- Extraction uses `SELECT *` so the dump is robust against Rails-side column additions.
|
|
1000
|
-
- `INSERT` statements omit the column list (`INSERT INTO schema_migrations VALUES (...)`). For PostgreSQL `--output-format=copy`, the `COPY` header similarly omits the column list (`COPY schema_migrations FROM stdin;`).
|
|
317
|
+
`schema_migrations` and `ar_internal_metadata` are exempt automatically.
|
|
1001
318
|
|
|
1002
|
-
|
|
1003
|
-
|
|
1004
|
-
- Defining `primary_key`, `columns`, or `belongs_tos` on a rails-managed entry is rejected with `ArgumentError` on load.
|
|
1005
|
-
- A rails-managed table cannot be used as `--target-table`.
|
|
1006
|
-
- In multi-database setups, the rails-managed entry is emitted under whichever database's connection actually contains the table (see [Multiple databases](#multiple-databases)). The table name itself is still derived from the global `ActiveRecord::Base.schema_migrations_table_name` / `internal_metadata_table_name` (prefix/suffix) accessors.
|
|
1007
|
-
|
|
1008
|
-
### Composite primary keys (unsupported)
|
|
319
|
+
#### Per-table `scope_column` and ID spaces
|
|
1009
320
|
|
|
1010
|
-
|
|
321
|
+
Each table names its own column, so tables that store the tenant id under different names work together. When one database has two groups of tables keyed by different kinds of id, give one group an `id_space` and pass its values separately:
|
|
1011
322
|
|
|
1012
323
|
```json
|
|
1013
|
-
{
|
|
1014
|
-
|
|
1015
|
-
"type": "unsupported_composite_primary_key",
|
|
1016
|
-
"ignore": true,
|
|
1017
|
-
"comment": "exwiw does not support composite primary keys (organization_id, location_id); data extraction is skipped.",
|
|
1018
|
-
"belongs_tos": [],
|
|
1019
|
-
"columns": [{ "name": "organization_id" }, { "name": "location_id" }, { "name": "name" }]
|
|
1020
|
-
}
|
|
324
|
+
{ "name": "tenants", "primary_key": "id", "scope_column": "id", "columns": [{ "name": "id" }] }
|
|
325
|
+
{ "name": "organizations", "primary_key": "id", "scope_column": "id", "id_space": "org", "columns": [{ "name": "id" }] }
|
|
1021
326
|
```
|
|
1022
327
|
|
|
1023
|
-
|
|
1024
|
-
|
|
1025
|
-
|
|
328
|
+
```bash
|
|
329
|
+
exwiw ... --ids=1,2 --ids=org=0b6f4c1e-0000-4000-8000-000000000001
|
|
330
|
+
```
|
|
1026
331
|
|
|
1027
|
-
|
|
332
|
+
- Tables without `id_space`, and `--ids` without a prefix, use the `default` ID space. Write `--ids=default=...` when an id itself contains `=`.
|
|
333
|
+
- A table without a scope column uses the ID space of the table it is narrowed through.
|
|
334
|
+
- The run stops before exporting when an ID space a table needs has no values, when values are given for an unused ID space, or when a table reaches tables of more than one ID space.
|
|
335
|
+
- Single-target mode and MongoDB support only the `default` ID space.
|
|
336
|
+
- In the config file, `ids:` takes a list, or a mapping such as `ids: { default: [1, 2], org: [...] }`.
|
|
1028
337
|
|
|
1029
|
-
|
|
338
|
+
The older global `--scope-column=COLUMN` flag still works but is deprecated; declare `scope_column` per table instead.
|
|
1030
339
|
|
|
1031
340
|
### Batched extraction (`batch_scope`)
|
|
1032
341
|
|
|
1033
|
-
A
|
|
1034
|
-
|
|
1035
|
-
```sql
|
|
1036
|
-
SELECT activities.* FROM activities
|
|
1037
|
-
JOIN customers ON activities.customer_id = customers.id
|
|
1038
|
-
AND customers.tenant_id IN ('t1')
|
|
1039
|
-
```
|
|
1040
|
-
|
|
1041
|
-
That is index-driven while the scope keeps few `customers`. Past some number of them the planner's estimate of "probe the foreign-key index once per customer" exceeds its estimate of "scan the table once", and it switches to a **sequential scan of the whole table** — for a result set that is a small fraction of it. On a table of hundreds of millions of rows the scan then exceeds the server's `statement_timeout`, or simply runs for hours. Note that no `filter` on the extracted table fixes this: a predicate that reduces the *output* does not reduce the *work* once the plan is a scan (it may not even change the plan).
|
|
342
|
+
A table with hundreds of millions of rows can be slow to export even when few of its rows are kept. Once the scope covers many parent rows, the database may decide that scanning the whole table is cheaper than using the foreign key index, and the query runs for hours or hits `statement_timeout`.
|
|
1042
343
|
|
|
1043
|
-
`batch_scope`
|
|
344
|
+
`batch_scope` avoids this by exporting the table in batches. exwiw first fetches the ids of a narrowed table on the path (the batch table), then runs one query per `size` ids, with the ids written into the query:
|
|
1044
345
|
|
|
1045
346
|
```json
|
|
1046
347
|
{
|
|
@@ -1052,224 +353,222 @@ That is index-driven while the scope keeps few `customers`. Past some number of
|
|
|
1052
353
|
}
|
|
1053
354
|
```
|
|
1054
355
|
|
|
1055
|
-
Each batch runs with that slice's ids in place of the scope filter:
|
|
1056
|
-
|
|
1057
356
|
```sql
|
|
1058
357
|
SELECT activities.* FROM activities
|
|
1059
358
|
JOIN customers ON activities.customer_id = customers.id
|
|
1060
359
|
AND customers.id IN (/* 1000 ids */)
|
|
1061
360
|
```
|
|
1062
361
|
|
|
1063
|
-
|
|
362
|
+
- The exported rows are the same as without batching. The batch table's ids are sorted, so the output is the same on every run.
|
|
363
|
+
- `size` defaults to 1000.
|
|
364
|
+
- The batch table can be several hops up the path.
|
|
365
|
+
- A table with its own `scope_column` can name itself. This only helps when the scope column is indexed.
|
|
366
|
+
- `batch_scope` requires scope-column mode, and a table narrowed by rule 1 or rule 2. Other cases are rejected before anything is written.
|
|
367
|
+
- With `--output-format=copy`, all batches are held in memory at once. Use the default `INSERT` format for very large results.
|
|
368
|
+
- `exwiw explain` also shows the query that fetches the batch table's ids.
|
|
1064
369
|
|
|
1065
|
-
|
|
1066
|
-
- **`size` defaults to 1000** ids per batch.
|
|
1067
|
-
- The batch table's ids come from **its own extraction query**, so it is narrowed by exactly the filter it would carry in the unbatched query. They are held in memory for the extraction: one scope's worth of primary keys, orders of magnitude smaller than the table being batched.
|
|
1068
|
-
- The batch table may be **any number of hops up** the path — a table two hops below it (`activity_orders → activities → customers`) names `customers` too, and the batch ids are applied where the path meets the scope, bounding the whole join chain.
|
|
1069
|
-
- A table that **carries the scope column itself** batches by naming itself; each batch then filters `WHERE <pk> IN (<ids>)` directly. Note that the id-set query is then the same scope predicate over the same table, so this shape only avoids the scan when the scope column is indexed (ideally index-only) — the join shape above is the one that genuinely removes the planner's choice.
|
|
1070
|
-
- `bulk_insert_chunk_size` is independent: batches are query boundaries, chunks are `INSERT` statement boundaries.
|
|
1071
|
-
- With `--output-format=copy`, batching bounds each query's cost but not memory: COPY builds the whole table's body in memory, so all batches' rows are resident at once. Use the default INSERT format (which streams) when the kept rows themselves are huge.
|
|
370
|
+
## Masking
|
|
1072
371
|
|
|
1073
|
-
|
|
372
|
+
Each column can be masked with one of the following keys.
|
|
1074
373
|
|
|
1075
|
-
|
|
1076
|
-
- the table reaches the scope through a **single `belongs_to` join path** (path 2 in [the six scoping paths](#how-each-table-is-narrowed--the-six-scoping-paths)) whose scoped terminus is the named table.
|
|
374
|
+
### `replace_with`
|
|
1077
375
|
|
|
1078
|
-
|
|
376
|
+
Replaces the value with a string. `{column}` is replaced with that column's value, so for a row with `id` 1, `"user{id}@example.com"` becomes `user1@example.com`. `{}` is kept as is, so `"replace_with": "{}"` gives an empty JSON object.
|
|
1079
377
|
|
|
1080
|
-
|
|
378
|
+
A number or boolean is used as it is, so non-text columns keep their type:
|
|
1081
379
|
|
|
1082
|
-
|
|
380
|
+
```jsonc
|
|
381
|
+
{ "name": "score", "replace_with": 0 }
|
|
382
|
+
{ "name": "active", "replace_with": false }
|
|
383
|
+
```
|
|
1083
384
|
|
|
1084
|
-
|
|
385
|
+
`NULL` stays `NULL` (an empty string is still replaced).
|
|
1085
386
|
|
|
1086
|
-
|
|
1087
|
-
`filter` is here for that. Be careful to use this option, as it will be:
|
|
387
|
+
### `raw_sql`
|
|
1088
388
|
|
|
1089
|
-
|
|
1090
|
-
- injected to every where / join clause, so it affects to all tables depends on filterted target-table. it results to data inconsistency.
|
|
1091
|
-
- a way to reduce the rows returned, which is **not** necessarily a way to reduce the work: on a large table the engine may keep (or switch to) a full scan and evaluate the filter per row. See [batched extraction](#batched-extraction-batch_scope) when the goal is to bound how much of the table is read.
|
|
389
|
+
An SQL expression used in place of the column, such as `"CONCAT('user', shops.id, '@example.com')"`. Use it when a database function is needed. Qualify column names with the table name. `replace_with` is ignored when both are set. SQL adapters only.
|
|
1092
390
|
|
|
1093
|
-
###
|
|
391
|
+
### `map`
|
|
1094
392
|
|
|
1095
|
-
`
|
|
393
|
+
Ruby code that returns a `Proc`. The proc is called with each row, and its return value (a `String`, a number or `nil`) replaces the column:
|
|
1096
394
|
|
|
1097
|
-
|
|
395
|
+
```jsonc
|
|
396
|
+
{ "name": "email", "map": "proc { |r| 'user' + r['id'].to_s + '@example.com' }" }
|
|
397
|
+
```
|
|
1098
398
|
|
|
1099
|
-
|
|
1100
|
-
|
|
399
|
+
- `r['column']` reads any column of the row, after the database-side masking of other columns.
|
|
400
|
+
- `NULL` is not kept automatically; the proc receives `nil` and decides.
|
|
401
|
+
- It runs in the exwiw process, so `explain` does not show it. SQL adapters only.
|
|
402
|
+
- Because the config runs arbitrary Ruby, only load configs you trust.
|
|
1101
403
|
|
|
1102
|
-
|
|
1103
|
-
then "user{id}@example.com" will be replaced with "user1@example.com".
|
|
404
|
+
Prefer `replace_with` or `raw_sql` when they can do the job.
|
|
1104
405
|
|
|
1105
|
-
|
|
1106
|
-
absent field) is left as-is instead of being replaced by the masked literal, so the
|
|
1107
|
-
"not set" signal survives into the dump. Only true `NULL`/absent is preserved — an empty
|
|
1108
|
-
string is a real value and is still masked. Because of this you do not need to hand-write a
|
|
1109
|
-
`raw_sql` `CASE WHEN ... IS NOT NULL ...` to keep NULLs.
|
|
406
|
+
### `replace_with_fake_data`
|
|
1110
407
|
|
|
1111
|
-
|
|
1112
|
-
template, so a column that is not text keeps its type:
|
|
408
|
+
Replaces the value with realistic fake data from the [faker](https://github.com/faker-ruby/faker) gem. The value is chosen from a seed column, so the same seed always gives the same value, across tables, runs and adapters:
|
|
1113
409
|
|
|
1114
410
|
```jsonc
|
|
1115
|
-
{ "name": "
|
|
1116
|
-
{ "name": "active", "replace_with": false } // boolean column
|
|
1117
|
-
{ "name": "email", "replace_with": "masked-{id}@example.com" } // template, as above
|
|
411
|
+
{ "name": "name", "replace_with_fake_data": { "seed": "users.id", "type": "human_name", "locale": "ja" } }
|
|
1118
412
|
```
|
|
1119
413
|
|
|
1120
|
-
|
|
1121
|
-
|
|
1122
|
-
|
|
414
|
+
| type | example (en) | example (`locale: ja`) |
|
|
415
|
+
|------|--------------|------------------------|
|
|
416
|
+
| `human_name` | `Adrianna Kilback` | `山田 太郎` |
|
|
417
|
+
| `first_name` | `Adrianna` | `太郎` |
|
|
418
|
+
| `last_name` | `Kilback` | `山田` |
|
|
419
|
+
| `human_name_kana` | (ja only) | `ヤマダ タロウ` |
|
|
420
|
+
| `first_name_kana` | (ja only) | `タロウ` |
|
|
421
|
+
| `last_name_kana` | (ja only) | `ヤマダ` |
|
|
422
|
+
| `phone_number` | `(555) 123-4567` | |
|
|
423
|
+
| `address` | `282 Kevin Brook, Imogeneborough, CA 58517` | |
|
|
424
|
+
| `company_name` | `Hirthe-Ritchie` | |
|
|
425
|
+
| `email` | `cliff.fay.9d6b804eff5a3f57@example.com` | |
|
|
426
|
+
| `username` | `cliff.fay_9d6b804eff5a3f57` | |
|
|
1123
427
|
|
|
1124
|
-
|
|
1125
|
-
|
|
1126
|
-
|
|
428
|
+
- `seed` is a column of the same table, with or without the table name. Use a stable id such as the primary key.
|
|
429
|
+
- For one seed, the name types describe the same person: `human_name` is `last_name` + `first_name`, and the kana matches the kanji.
|
|
430
|
+
- Different seeds can get the same name. `email` and `username` include a token from the seed, so they stay unique.
|
|
431
|
+
- Values change when `locale`, the faker version, or exwiw's bundled Japanese name list changes.
|
|
432
|
+
- `NULL` stays `NULL`.
|
|
433
|
+
- Add `gem "faker"` to your Gemfile. A config that uses only `ja` name types does not need it.
|
|
434
|
+
- It runs in the exwiw process, so `explain` does not show it. The cost is small; see [`docs/row-transform-masking-notes.md`](docs/row-transform-masking-notes.md).
|
|
1127
435
|
|
|
1128
|
-
|
|
436
|
+
Only one masking key can be set on a column.
|
|
1129
437
|
|
|
1130
|
-
|
|
438
|
+
## Generating the schema config
|
|
1131
439
|
|
|
1132
|
-
|
|
1133
|
-
`"replace_with": "user{id}@example.com"`.
|
|
1134
|
-
This is useful when you want to transform with functions provided by the database.
|
|
440
|
+
In a Rails application, a rake task writes the schema config from the models:
|
|
1135
441
|
|
|
1136
|
-
|
|
442
|
+
```bash
|
|
443
|
+
bundle exec rake exwiw:schema:generate
|
|
444
|
+
```
|
|
1137
445
|
|
|
1138
|
-
|
|
446
|
+
The files go to `EXWIW_SCHEMA_DIR_PATH` if set, otherwise `schema_dir` from `exwiw.yml`, otherwise `exwiw/schema`. If the application has more than one source of models, such as ActiveRecord and Mongoid, give each its own directory; `check` and `tidy` treat files they do not recognize as stale.
|
|
1139
447
|
|
|
1140
|
-
|
|
448
|
+
### Safe mode (masking new columns by default)
|
|
1141
449
|
|
|
1142
|
-
|
|
1143
|
-
`Proc`, and the proc is called for every fetched row. Its return value replaces
|
|
1144
|
-
the column value in the dump:
|
|
450
|
+
A new column could hold personal data, so `schema:generate` writes every column that is not yet in the config as masked, and marks it with [`needs_mask_decision: true`](#needs_mask_decision). Columns already in the config are left as they are.
|
|
1145
451
|
|
|
1146
|
-
|
|
1147
|
-
|
|
452
|
+
The mask is the column's default value if it has a constant one, otherwise a value by type: `masked-{primary key}` for text (with `@example.com` when the name mentions mail), `0`, `false`, a fixed date, or `{}` for JSON. Some columns are marked but left unmasked, because masking them would break the dump or the restore:
|
|
453
|
+
|
|
454
|
+
- the primary key and the columns `belongs_to` joins on
|
|
455
|
+
- types no constant fits, such as `uuid`, binary, enums, arrays, and text too short for the mask
|
|
456
|
+
- columns under a unique index, unless the mask differs per row
|
|
457
|
+
|
|
458
|
+
For the first config of an application, where every column is new, run with `EXWIW_NEW_COLUMNS=plain` to turn safe mode off. Do not use it afterwards: those columns get no mark, so nothing tells them apart from reviewed ones.
|
|
459
|
+
|
|
460
|
+
### `needs_mask_decision`
|
|
461
|
+
|
|
462
|
+
`needs_mask_decision: true` marks a column whose masking nobody has decided yet. Export ignores it. [`schema:check`](#checking-the-config-against-the-schema) reports these columns, so CI can block a pull request until each is decided. To decide, keep the mask (ideally with a `comment` saying why), change it, remove `replace_with` to export the real value, or set `ignore: true`, and then remove the key.
|
|
463
|
+
|
|
464
|
+
### Tidying stale config (`schema:tidy`)
|
|
465
|
+
|
|
466
|
+
`schema:generate` never deletes anything. `schema:tidy` removes the config files of tables that no longer exist in the database, and the columns those tables no longer have. It reads the database, not the models, so a table without a model is kept. It does not touch anything else, and does not remove stale `belongs_tos`; run `schema:generate` for those.
|
|
467
|
+
|
|
468
|
+
```bash
|
|
469
|
+
bundle exec rake exwiw:schema:tidy
|
|
1148
470
|
```
|
|
1149
471
|
|
|
1150
|
-
|
|
1151
|
-
|
|
1152
|
-
- `r['column_name']` reads any column of the current row — the value as fetched
|
|
1153
|
-
from the database (i.e. after SQL-side masking such as another column's
|
|
1154
|
-
`replace_with`, before Ruby-side transforms). `r` is only valid inside the
|
|
1155
|
-
call; do not retain it.
|
|
1156
|
-
- Return a `String`, `Numeric`, or `nil`. Unlike `replace_with` there is **no
|
|
1157
|
-
automatic NULL preservation** — the proc receives `nil` and decides.
|
|
1158
|
-
- `map` is exclusive with the other masking keys on the same column
|
|
1159
|
-
(`raw_sql` / `replace_with` / `replace_with_fake_data`).
|
|
1160
|
-
- SQL adapters only. On the MongoDB adapter the key is rejected on load, like
|
|
1161
|
-
`raw_sql` (see [Unknown keys are rejected](#unknown-keys-are-rejected)).
|
|
1162
|
-
Because the transform runs in the exwiw process, it is invisible to
|
|
1163
|
-
`explain`.
|
|
1164
|
-
|
|
1165
|
-
**Security note**: `map` executes arbitrary Ruby from the schema config. Treat
|
|
1166
|
-
config files with the same trust as your Gemfile — only load trusted configs.
|
|
1167
|
-
|
|
1168
|
-
This is the most powerful option, but it runs per row in the exwiw process
|
|
1169
|
-
rather than in the database. The measured dispatch cost is small, though
|
|
1170
|
-
(~0.6–0.8µs/row plus whatever the proc body does — see
|
|
1171
|
-
[`docs/row-transform-masking-notes.md`](docs/row-transform-masking-notes.md)).
|
|
1172
|
-
Prefer `replace_with`/`raw_sql` when they can express the transform; reach for
|
|
1173
|
-
`map` when they cannot.
|
|
1174
|
-
|
|
1175
|
-
#### `replace_with_fake_data`
|
|
1176
|
-
|
|
1177
|
-
Replaces the value with realistic-looking fake data generated by the
|
|
1178
|
-
[faker](https://github.com/faker-ruby/faker) gem, picked **deterministically**
|
|
1179
|
-
from the value of a seed column — the same seed value always maps to the same
|
|
1180
|
-
fake value, across tables, runs, and adapters:
|
|
472
|
+
### Checking the config against the schema
|
|
1181
473
|
|
|
1182
|
-
|
|
474
|
+
`schema:check` reports how the config differs from what `generate` and `tidy` would produce, without writing anything. It exits non-zero when something needs attention, so it can run in CI:
|
|
475
|
+
|
|
476
|
+
```bash
|
|
477
|
+
bundle exec rake exwiw:schema:check
|
|
478
|
+
```
|
|
479
|
+
|
|
480
|
+
```json
|
|
1183
481
|
{
|
|
1184
|
-
"
|
|
1185
|
-
"
|
|
482
|
+
"added_tables": [],
|
|
483
|
+
"added_columns": ["users.contact_email"],
|
|
484
|
+
"removed_tables": [],
|
|
485
|
+
"removed_columns": ["orders.legacy_flag"],
|
|
486
|
+
"changed_tables": ["orders", "users"],
|
|
487
|
+
"needs_mask_decision": ["orders.memo"],
|
|
488
|
+
"stale_tables": [],
|
|
489
|
+
"stale_columns": ["orders.legacy_flag"]
|
|
1186
490
|
}
|
|
1187
491
|
```
|
|
1188
492
|
|
|
1189
|
-
- `
|
|
1190
|
-
|
|
1191
|
-
|
|
1192
|
-
|
|
1193
|
-
|
|
1194
|
-
|
|
1195
|
-
|
|
1196
|
-
|
|
1197
|
-
|
|
1198
|
-
|
|
1199
|
-
|
|
1200
|
-
|
|
1201
|
-
|
|
1202
|
-
|
|
1203
|
-
|
|
1204
|
-
|
|
1205
|
-
|
|
1206
|
-
|
|
1207
|
-
|
|
1208
|
-
|
|
1209
|
-
|
|
1210
|
-
|
|
1211
|
-
|
|
1212
|
-
|
|
1213
|
-
|
|
1214
|
-
|
|
1215
|
-
|
|
1216
|
-
|
|
1217
|
-
|
|
1218
|
-
|
|
1219
|
-
|
|
1220
|
-
|
|
1221
|
-
|
|
1222
|
-
|
|
1223
|
-
-
|
|
1224
|
-
|
|
1225
|
-
|
|
1226
|
-
|
|
1227
|
-
|
|
1228
|
-
-
|
|
1229
|
-
|
|
1230
|
-
|
|
1231
|
-
|
|
1232
|
-
|
|
1233
|
-
|
|
1234
|
-
|
|
1235
|
-
|
|
1236
|
-
|
|
1237
|
-
-
|
|
1238
|
-
|
|
1239
|
-
|
|
1240
|
-
|
|
1241
|
-
|
|
1242
|
-
|
|
1243
|
-
-
|
|
1244
|
-
|
|
1245
|
-
|
|
1246
|
-
|
|
1247
|
-
|
|
1248
|
-
|
|
1249
|
-
|
|
1250
|
-
|
|
1251
|
-
|
|
1252
|
-
|
|
1253
|
-
|
|
1254
|
-
|
|
1255
|
-
|
|
1256
|
-
|
|
1257
|
-
|
|
1258
|
-
|
|
1259
|
-
|
|
1260
|
-
|
|
1261
|
-
|
|
1262
|
-
|
|
1263
|
-
|
|
1264
|
-
|
|
1265
|
-
|
|
1266
|
-
|
|
1267
|
-
|
|
1268
|
-
- Load the table information from the specified config file.
|
|
1269
|
-
- Calculate the dependency between tables.
|
|
1270
|
-
- Generate the full list of INSERT sql based on the specified conditions.
|
|
1271
|
-
- If the processing table has no relation with target tables, then dump all records.
|
|
1272
|
-
- If the processing table has relation with target tables, then dump the records which are related to the target tables.
|
|
493
|
+
- `added_*`, `removed_*` and `changed_tables`: run `schema:generate` and `schema:tidy`.
|
|
494
|
+
- `needs_mask_decision`: columns still waiting for a decision.
|
|
495
|
+
- `stale_*`: removed tables and columns that the config still exports, which would make the export fail. Used by `--fail-on=stale` (see [below](#non-rails-applications-exwiw-schema----from-db)).
|
|
496
|
+
|
|
497
|
+
With multiple databases each entry starts with the database name (`primary/users.email`). Set `EXWIW_SCHEMA_CHECK_OUTPUT=<path>` to also write the JSON to a file.
|
|
498
|
+
|
|
499
|
+
### Multiple databases
|
|
500
|
+
|
|
501
|
+
With Rails' multiple databases, `schema:generate` writes each database's files into a subdirectory named after it (`exwiw/schema/primary/`, `exwiw/schema/analytics/`). Each database is exported by a separate run.
|
|
502
|
+
|
|
503
|
+
A `belongs_to` to a model in another database is written with `ignore: true` and `ignore_type: "cross_database"`; see [Cross-database foreign keys](#cross-database-foreign-keys).
|
|
504
|
+
|
|
505
|
+
### Rails-managed tables and composite primary keys
|
|
506
|
+
|
|
507
|
+
`schema_migrations` and `ar_internal_metadata` get a config with a `type` such as `rails_managed_schema_migrations` and no columns. They are exported with `SELECT *` and `INSERT` without a column list, so new Rails versions do not break them. They cannot be the `--target-table`.
|
|
508
|
+
|
|
509
|
+
Composite primary keys are not supported. Such a table is generated with `ignore: true` and `type: "unsupported_composite_primary_key"`.
|
|
510
|
+
|
|
511
|
+
### Mongoid applications
|
|
512
|
+
|
|
513
|
+
```bash
|
|
514
|
+
bundle exec rake exwiw:schema:generate_mongoid
|
|
515
|
+
bundle exec rake exwiw:schema:tidy_mongoid
|
|
516
|
+
bundle exec rake exwiw:schema:check_mongoid
|
|
517
|
+
```
|
|
518
|
+
|
|
519
|
+
These work like the ActiveRecord tasks. See [Generating config from Mongoid models](docs/mongodb.md#generating-config-from-mongoid-models).
|
|
520
|
+
|
|
521
|
+
### Non-Rails applications (`exwiw schema ... --from-db`)
|
|
522
|
+
|
|
523
|
+
For applications exwiw cannot load, the same three operations read the database instead of the models:
|
|
524
|
+
|
|
525
|
+
```bash
|
|
526
|
+
exwiw schema generate --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
|
|
527
|
+
exwiw schema check --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
|
|
528
|
+
exwiw schema tidy --from-db -a postgresql -h db.example.com -p 5432 -u app --database=app --schema-dir=exwiw/schema
|
|
529
|
+
```
|
|
530
|
+
|
|
531
|
+
- MySQL and PostgreSQL only.
|
|
532
|
+
- `DATABASE_PASSWORD` may be empty, for CI databases without a password.
|
|
533
|
+
- Safe mode and the `check` report work as above. `check` exits 1 when the config needs attention, and with another status when it could not run.
|
|
534
|
+
- `check --fail-on=stale` fails only on `stale_*`, so it can run before an export without blocking on newly added columns.
|
|
535
|
+
- One run covers one database, and the files are written directly into the schema directory.
|
|
536
|
+
|
|
537
|
+
Foreign keys in the database become `belongs_tos`. A table without a primary key is generated with `ignore: true` and a comment; once you add a `primary_key` by hand, it is kept.
|
|
538
|
+
|
|
539
|
+
Regeneration only adds `belongs_tos`, because many relations exist only in application code. Add those by hand; they are kept from then on, and their foreign key columns are never masked. `tidy` removes a `belongs_to` whose table no longer exists.
|
|
540
|
+
|
|
541
|
+
## After-insert hook
|
|
542
|
+
|
|
543
|
+
`--after-insert-hook=PATH` runs a script after all data files are written, to add rows of your own.
|
|
544
|
+
|
|
545
|
+
A Ruby hook (`.rb`) can use:
|
|
546
|
+
|
|
547
|
+
- `cli_options`: the parsed options, such as `cli_options.fetch(:ids)`.
|
|
548
|
+
- `ids_for(id_space = "default")`: the ids of an [ID space](#per-table-scope_column-and-id-spaces).
|
|
549
|
+
- `insert_sql(template)`: renders an ERB template and writes it to `insert-{N+1}-after_insert.sql`, after the last data file. Multiple calls go into the same file.
|
|
550
|
+
- `insert_jsonl(collection, template)`: MongoDB only; see [MongoDB support](docs/mongodb.md).
|
|
551
|
+
|
|
552
|
+
```ruby
|
|
553
|
+
insert_sql <<~SQL
|
|
554
|
+
<%- cli_options.fetch(:ids).each do |tenant_id| -%>
|
|
555
|
+
INSERT INTO users (tenant_id, email) VALUES (<%= tenant_id %>, 'default@example.com');
|
|
556
|
+
<%- end -%>
|
|
557
|
+
SQL
|
|
558
|
+
```
|
|
559
|
+
|
|
560
|
+
Ruby hooks run inside the exwiw process, so only use hooks you trust.
|
|
561
|
+
|
|
562
|
+
Any other file is run as a command. Its output is not captured, and a non-zero exit stops exwiw. It gets `DATABASE_PASSWORD` and these environment variables:
|
|
563
|
+
|
|
564
|
+
- `EXWIW_OUTPUT_DIR`, `EXWIW_SCHEMA_DIR`
|
|
565
|
+
- `EXWIW_DATABASE_ADAPTER`, `EXWIW_DATABASE_HOST`, `EXWIW_DATABASE_PORT`, `EXWIW_DATABASE_USER`, `EXWIW_DATABASE_NAME`
|
|
566
|
+
- `EXWIW_TARGET_TABLE`, `EXWIW_IDS` (comma-separated, the `default` ID space), `EXWIW_OUTPUT_FORMAT`
|
|
567
|
+
- `EXWIW_IDS_<ID_SPACE>` for each ID space given values (`--ids=org=...` becomes `EXWIW_IDS_ORG`)
|
|
568
|
+
|
|
569
|
+
## MongoDB
|
|
570
|
+
|
|
571
|
+
`--adapter=mongodb` exports JSON Lines for `mongoimport`. Setup, options and the differences from the SQL adapters are in [docs/mongodb.md](docs/mongodb.md).
|
|
1273
572
|
|
|
1274
573
|
## Development
|
|
1275
574
|
|