db-purger 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 4025ea48c6e9d97505720a8450017e70f2fd527408b518cc8837ddf512ce043d
4
- data.tar.gz: 3cb576faf87f8a5d2735fd8edc973e04063469a7edae92148dac7d7a80688be1
3
+ metadata.gz: 863444350ef64942e0669feacdf4a119514c16369a31eeb99e31a5f5435028ef
4
+ data.tar.gz: 66b063187fdaa33232a6fffe76f9496264e8bd74e2a835013139b4e19f3e045e
5
5
  SHA512:
6
- metadata.gz: 3685ddf0b11aff29c97d618026abe46a59c1d21a8310f6d8c038f913b2892e30b77468c1d5d8946978eaf14e4003278815967417e7ec5ec98cc0257b00dbce84
7
- data.tar.gz: b6dacd558e0b1a7b4b22c9ed332df7071da6300e43e9eb0476cbcf32899c998c39197aad7005e78e38f86ed821d728bf61e8f9d62fe945695987abe614443d0e
6
+ metadata.gz: dfd50964ce367548606fe758df1d1cfb48d6891d9ac09993079b851a78658807fa4a5ff8633709a49f42b9d67c9367260721cd2d0f1454dbd7da197f3363f54b
7
+ data.tar.gz: ccc2bee89229fad63fd2dc7d20682b07b008b29040cb66a92aa1cd4c48bd8f09c5f6b91834ead524be8f3394d54c8af08ce8df5d61bc5067a6229b14187c7ede
data/ARCHITECTURE.md ADDED
@@ -0,0 +1,160 @@
1
+ # Architecture
2
+
3
+ db-purger is small (~900 lines) and splits cleanly into three layers: **describe** a purge (plan + DSL),
4
+ **check** it (validator), and **run** it (purgers + instrumentation).
5
+
6
+ ```
7
+ plan file / PlanBuilder.build { ... }
8
+ │
9
+ ▼
10
+ ┌──────────────┐ ┌─────────┐ ┌─────────┐
11
+ │ PlanBuilder │──▶│ Plan │──▶│ Table │──┐ each Table owns a nested Plan
12
+ │ (DSL) │ │ │ │ │◀─┘ (recursive tree)
13
+ └──────────────┘ └─────────┘ └─────────┘
14
+ │
15
+ ┌─────────────────┼──────────────────┐
16
+ ▼ ▼ ▼
17
+ PlanValidator Executor ─────▶ PurgeTable / PurgeTableScanner
18
+ (schema check) (entry point) │ (PurgeTableHelper)
19
+ ▼
20
+ ActiveSupport::Notifications
21
+ (*.db_purger events)
22
+ │
23
+ ▼
24
+ MetricSubscriber ─▶ Metrics
25
+ ```
26
+
27
+ ## Components
28
+
29
+ | File | Responsibility |
30
+ |---|---|
31
+ | `lib/db-purger.rb` | Autoloads everything; holds the global `DBPurger.config`. |
32
+ | `config.rb` | Global options: `explain?`, `explain_file`, `datetime_format`. |
33
+ | `table.rb` | Value object for one table in the plan: name, match field, options, and a lazily created nested `Plan`. |
34
+ | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, search and ignored tables. `#purge!` is the run entry point. |
35
+ | `plan_builder.rb` | The DSL. `instance_eval`s a plan file or block against a `Plan`; nested blocks get a new builder bound to that table's nested plan. |
36
+ | `plan_validator.rb` | `ActiveModel::Validations` over plan vs. schema: missing tables, unknown tables, unknown columns. |
37
+ | `executor.rb` | Convenience façade: loads a plan file, applies config options, `verify!`, `purge!`. |
38
+ | `purge_table.rb` | Purges one table by `field = value(s)` in primary-key batches. Recurses into nested tables. |
39
+ | `purge_table_scanner.rb` | Purges a `purge_table_search` table: full `find_in_batches` scan filtered through the user's `search_proc`. |
40
+ | `purge_table_helper.rb` | Shared behaviour for both purgers: nested-table recursion, delete vs. soft delete vs. explain, transactions. |
41
+ | `metrics.rb` / `metric_subscriber.rb` | Aggregate timing and row counts per table from the notification events. |
42
+ | `dynamic_plan_builder.rb` | Generates plan-file source from `has_many` associations; a bootstrap tool, not used at purge time. |
43
+
44
+ ## The plan tree
45
+
46
+ `PlanBuilder` always attaches tables to the *current* plan's `base_table.nested_plan` once a base table
47
+ exists. So a plan file is really:
48
+
49
+ ```
50
+ Plan (root)
51
+ └── base_table: companies(:id)
52
+ └── nested Plan
53
+ ├── parent_tables: [company_tags(:company_id)]
54
+ ├── child_tables: [employments(:company_id) ─▶ nested Plan ..., websites(:id, fk: website_id) ...]
55
+ ├── search_tables: [users(:id)]
56
+ └── ignore_tables
57
+ ```
58
+
59
+ `Plan#tables` flattens this tree for validation; `Table#foreign_keys` collects the `foreign_key:` columns of
60
+ a table's direct children so the purger can `SELECT` them alongside the primary key.
61
+
62
+ ## Purge algorithm
63
+
64
+ `Plan#purge!` resets metrics and starts a `PurgeTable` on the base table with the purge value. Each
65
+ `PurgeTable#purge!` does:
66
+
67
+ ```
68
+ each_batch = loop:
69
+ batch = SELECT pk, <child foreign_keys> FROM t
70
+ WHERE field IN (values) [AND conditions] [AND pk > last_pk]
71
+ ORDER BY pk LIMIT batch_size
72
+ break if batch empty
73
+
74
+ purge_children(batch) =
75
+ for each child_table without foreign_key: # rows pointing at us
76
+ PurgeTable(child, child.field, batch.pks).purge! (recursive)
77
+
78
+ delete_rows(batch) =
79
+ if any child has foreign_key: # rows we point at
80
+ TRANSACTION
81
+ delete batch by pk
82
+ for each fk child: PurgeTable(child, child.field, batch.<fk values>).purge!
83
+ else
84
+ delete batch by pk
85
+
86
+ if table has a primary key:
87
+ if no parent_tables:
88
+ each_batch: purge_children(batch); delete_rows(batch)
89
+ else: # two passes
90
+ each_batch: purge_children(batch)
91
+ for each parent_table: # siblings sharing the key
92
+ PurgeTable(parent, parent.field, original purge value).purge!
93
+ each_batch: delete_rows(batch)
94
+ else:
95
+ raise if nested child/parent tables # no ids to propagate
96
+ single DELETE WHERE field = value [AND conditions]
97
+
98
+ for each search_table:
99
+ PurgeTableScanner(search_table).purge!
100
+ ```
101
+
102
+ Key properties:
103
+
104
+ - **Depth-first, children first.** A row is only deleted after everything referencing it, so FK
105
+ constraints hold without `ON DELETE CASCADE`.
106
+ - **Two passes when there are parent tables.** A parent table can reference this table (e.g.
107
+ `company_tags.company_id → companies.id`) *and* be referenced by one of its children, so it is purged
108
+ between the child pass and the delete pass. Tables without parent tables keep the single pass.
109
+ - **Keyset pagination** (`pk > last_pk`) rather than `OFFSET`, so batches stay cheap on large tables and still
110
+ advance in explain mode where nothing is actually deleted.
111
+ - **Bounded memory.** Only one batch of ids per level of the tree is held at a time.
112
+ - **Idempotent re-runs.** Every step re-derives its rows from the database; there is no saved cursor, so a
113
+ crashed purge is resumed by running it again.
114
+ - **Parent vs. child** is about *which value* is propagated: children get the enclosing batch's primary keys,
115
+ parents get the original purge value.
116
+
117
+ `PurgeTableScanner` follows the same shape but sources batches from `find_in_batches` over the entire table
118
+ (plus `conditions`) and narrows each batch with `search_proc` before recursing and deleting.
119
+
120
+ ## Delete strategies
121
+
122
+ `PurgeTableHelper#delete_records_with_instrumentation` picks one of three actions for every scope:
123
+
124
+ 1. **explain** (`DBPurger.config.explain?`) — rewrite `scope.to_sql` into `DELETE`/`UPDATE ... SET`, write it
125
+ to `explain_file`, return `scope.count`.
126
+ 2. **soft delete** (`mark_deleted_field`) — `scope.update_all(field => value)`.
127
+ 3. **hard delete** — `scope.delete_all`.
128
+
129
+ All three go through ActiveRecord's `*_all` methods: no model callbacks or validations run.
130
+
131
+ ## Instrumentation
132
+
133
+ Every unit of work is wrapped in `ActiveSupport::Notifications.instrument` under the `db_purger` namespace
134
+ (`purge`, `next_batch`, `delete_records`, `search_filter`). The purgers never talk to `Metrics` directly;
135
+ `MetricSubscriber` (an `ActiveSupport::Subscriber`) translates events into `Metrics` counters. This keeps the
136
+ purge code free of reporting concerns and lets callers attach their own subscribers (StatsD, logs, progress
137
+ bars) without changes to the library.
138
+
139
+ `PurgeTableScanner` uses the lower-level `instrumenter.start/finish` pair for `next_batch` because the batch
140
+ fetch happens inside `find_in_batches` rather than in a block the scanner controls.
141
+
142
+ ## Dependencies and boundaries
143
+
144
+ - **ActiveRecord** is the only database interface. The library relies on `where`, `select`, `order`,
145
+ `limit`, `find_in_batches`, `delete_all`, `update_all`, `transaction`, `to_sql` and
146
+ `connection.quote*`, so it is adapter-agnostic (specs use SQLite).
147
+ - **dynamic-active-model** supplies the `database` object. The purge code only calls `database.models` and
148
+ matches on `model.table_name`, so any object with that shape works. `DynamicPlanBuilder` additionally
149
+ relies on `reflect_on_all_associations`.
150
+ - **Config scoping.** `DBPurger.config` returns the config set by `DBPurger.with_config` for the current
151
+ thread, falling back to a process-wide default. `Executor#purge!` wraps the run in `with_config`, so executors
152
+ with different explain settings can't interfere. `MetricSubscriber.metrics` is still process-wide.
153
+
154
+ ## Testing
155
+
156
+ - `spec/db-purger/*` — unit specs for the builder, validator, executor and purger.
157
+ - `spec/integrations/*` — end-to-end purges over the schema in `spec/support/db/schema.rb`, asserting row-count
158
+ deltas per table and, for explain mode, the exact SQL in `spec/fixtures/delete_plan.sql`.
159
+ - `spec/support/test_db.rb` builds a fresh SQLite database and dynamic models (`TestDB::*`) for each run.
160
+ - `spec/fixtures/*.plan.rb` — plan files used for loading and validation cases.
data/LICENSE.txt ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2019 Douglas Youch
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,271 @@
1
+ # db-purger
2
+
3
+ [![CI](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml/badge.svg?branch=master)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
4
+ [![Coverage](https://raw.githubusercontent.com/dougyouch/db-purger/badges/coverage.svg)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
5
+ [![Gem Version](https://img.shields.io/gem/v/db-purger)](https://rubygems.org/gems/db-purger)
6
+
7
+ Purge every row tied to a single top-level record — a company, an account, a tenant — across all of the
8
+ tables that reference it, in batches, from a declarative Ruby plan.
9
+
10
+ ```ruby
11
+ executor = DBPurger::Executor.new(database, 'config/company.plan.rb')
12
+ executor.verify! # fail fast if the plan doesn't cover the schema
13
+ executor.purge!(42) # delete company 42 and everything that hangs off it
14
+ ```
15
+
16
+ ## Why
17
+
18
+ Deleting a tenant from a relational database is rarely one `DELETE`. Rows are spread across dozens of
19
+ tables, some keyed directly on the tenant id, some several joins away, some polymorphic, some that must be
20
+ soft-deleted rather than removed. `ON DELETE CASCADE` is often absent, and a single giant delete will lock
21
+ tables and blow out replication.
22
+
23
+ db-purger lets you describe those relationships once, in a plan file, and then:
24
+
25
+ - deletes in **primary-key batches** (default 10,000) so no single statement gets too large
26
+ - deletes rows **only after the rows that reference them**, so foreign-key constraints hold without `ON DELETE CASCADE`
27
+ - **validates** the plan against the live schema, so a newly added table can't be silently forgotten
28
+ - supports **soft deletes** (`UPDATE ... SET deleted_at = ...`) per table
29
+ - has an **explain mode** that prints the SQL it would run instead of running it
30
+ - emits **ActiveSupport::Notifications** events, with a built-in subscriber that collects timing and row counts
31
+
32
+ ## Installation
33
+
34
+ ```ruby
35
+ # Gemfile
36
+ gem 'db-purger'
37
+ ```
38
+
39
+ Requires Ruby >= 3.2 and ActiveRecord >= 7.0. Models are supplied by
40
+ [dynamic-active-model](https://github.com/dougyouch/dynamic-active-model), which builds ActiveRecord classes
41
+ directly from the database schema.
42
+
43
+ ## Quick start
44
+
45
+ ### 1. Load the database
46
+
47
+ db-purger works against a `DynamicActiveModel::Database`. Anything that responds to `#models` (returning
48
+ ActiveRecord classes) will do.
49
+
50
+ ```ruby
51
+ require 'active_record'
52
+ require 'dynamic-active-model'
53
+ require 'db-purger'
54
+
55
+ module PurgeDB; end
56
+
57
+ database = DynamicActiveModel::Explorer.explore(
58
+ PurgeDB,
59
+ { adapter: 'mysql2', host: 'localhost', database: 'app', username: 'app' },
60
+ %w[schema_migrations ar_internal_metadata] # tables to skip entirely
61
+ )
62
+ ```
63
+
64
+ ### 2. Write a plan
65
+
66
+ A plan is Ruby, evaluated with the plan DSL. Given this schema:
67
+
68
+ ```
69
+ companies (id, website_id, ...)
70
+ employments (id, company_id, user_id, ...)
71
+ employment_notes (id, employment_id, ...)
72
+ company_tags (company_id, tag_id) -- no primary key
73
+ events (id, model_type, model_id, ...) -- polymorphic
74
+ websites (id, content_id, ...)
75
+ contents (id, ...)
76
+ tags, jobs -- shared lookup tables, never purged
77
+ ```
78
+
79
+ a plan to purge one company looks like:
80
+
81
+ ```ruby
82
+ # config/company.plan.rb
83
+ base_table(:companies, :id)
84
+
85
+ # Tables keyed directly on the purge value (company_id = 42)
86
+ parent_table(:company_tags, :company_id)
87
+
88
+ # Tables keyed on the base table's primary key, purged batch-by-batch
89
+ child_table(:employments, :company_id) do
90
+ child_table(:employment_notes, :employment_id)
91
+ child_table(:events, :model_id, conditions: { model_type: 'PurgeDB::Employment' })
92
+ end
93
+
94
+ child_table(:events, :model_id, conditions: { model_type: 'PurgeDB::Company' })
95
+
96
+ # The company row points at the website (companies.website_id -> websites.id)
97
+ child_table(:websites, :id, foreign_key: :website_id) do
98
+ child_table(:contents, :id, foreign_key: :content_id)
99
+ end
100
+
101
+ # Shared tables that are intentionally left alone
102
+ ignore_table :tags
103
+ ignore_table :jobs
104
+ ignore_table(/\Atmp_/) # regexps are allowed
105
+ ```
106
+
107
+ ### 3. Verify and purge
108
+
109
+ ```ruby
110
+ executor = DBPurger::Executor.new(database, 'config/company.plan.rb')
111
+ executor.verify! # raises 'purge plan failed verification', errors printed to $stderr
112
+ deleted = executor.purge!(42)
113
+ ```
114
+
115
+ `purge!` returns the number of base-table rows deleted.
116
+
117
+ ## The plan DSL
118
+
119
+ | Method | Meaning |
120
+ |---|---|
121
+ | `base_table(table, field, opts = {}, &block)` | The root of the purge. Rows where `field = purge_value` are purged. Declare it **first** — every subsequent top-level call nests under it. |
122
+ | `child_table(table, field, opts = {}, &block)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch. Purged before that batch is deleted. |
123
+ | `child_table(table, :id, foreign_key: :col, &block)` | Inverted relationship: the *enclosing* table holds `col` pointing at this table's `id`. Deleted in the same transaction, right after the enclosing batch. |
124
+ | `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
125
+ | `purge_table_search(table, field, opts = {}) { \|batch\| ... }` | Scans the whole table in batches; the block receives each batch and returns the records to purge. For orphans that can't be reached by a key. |
126
+ | `ignore_table(name_or_regexp)` | Exclude a table from validation. |
127
+
128
+ Blocks nest arbitrarily deep. Because `purge_table_search` uses its block as the filter, nest tables under it
129
+ with `.nested_plan`:
130
+
131
+ ```ruby
132
+ purge_table_search(:users, :id) do |users|
133
+ users = users.index_by(&:id)
134
+ PurgeDB::Employment.where(user_id: users.keys).pluck(:user_id).each { |id| users.delete(id) }
135
+ users.values # users with no remaining employments
136
+ end.nested_plan do
137
+ child_table(:events, :model_id, conditions: { model_type: 'PurgeDB::User' })
138
+ end
139
+ ```
140
+
141
+ ### Table options
142
+
143
+ | Option | Default | Description |
144
+ |---|---|---|
145
+ | `batch_size:` | `10_000` | Rows fetched and deleted per batch. |
146
+ | `conditions:` | none | Extra `where` applied to every query for this table (hash or SQL string). |
147
+ | `foreign_key:` | none | See `child_table` above. |
148
+ | `mark_deleted_field:` | none | Soft delete: `UPDATE table SET field = value` instead of `DELETE`. |
149
+ | `mark_deleted_value:` | `1` | Value written to `mark_deleted_field`. `Time` values are formatted with `datetime_format` in explain output. |
150
+
151
+ Tables **without a primary key** (e.g. join tables) are purged with a single unbatched `DELETE ... WHERE
152
+ field = value`; nested tables are not supported under them.
153
+
154
+ ### Building a plan in code
155
+
156
+ ```ruby
157
+ plan = DBPurger::PlanBuilder.build do
158
+ base_table(:companies, :id)
159
+ child_table(:employments, :company_id)
160
+ end
161
+
162
+ DBPurger::Executor.new(database, plan).purge!(42)
163
+ # or, without the executor:
164
+ plan.purge!(database, 42)
165
+ ```
166
+
167
+ ### Generating a starting plan
168
+
169
+ `DynamicPlanBuilder` walks the `has_many` associations dynamic-active-model discovered and emits a plan file,
170
+ listing every unreachable table as `ignore_table`. Treat the output as a first draft: it only knows about
171
+ conventional `<singular_table>_id` foreign keys and cannot infer polymorphic, soft-delete, or search rules.
172
+
173
+ ```ruby
174
+ puts DBPurger::DynamicPlanBuilder.new(database).build(:companies, :id)
175
+ ```
176
+
177
+ ## Validation
178
+
179
+ `Executor#verify!` (or `DBPurger::PlanValidator.new(database, plan).valid?`) checks that:
180
+
181
+ - every table in the database is either in the plan or ignored (`missing_tables`)
182
+ - every table in the plan exists in the database (`unknown_tables`)
183
+ - every field, `foreign_key` and `mark_deleted_field` named in the plan is a real column
184
+ - the plan has a `base_table`, and every `batch_size` is positive
185
+ - tables without a primary key have no nested child or parent tables (there would be no ids to propagate)
186
+
187
+ Run it in CI against your schema so a new table can't ship without a purge decision.
188
+
189
+ ## Explain mode (dry run)
190
+
191
+ ```ruby
192
+ File.open('purge.sql', 'w') do |io|
193
+ executor = DBPurger::Executor.new(database, 'config/company.plan.rb', explain: true, explain_file: io)
194
+ executor.purge!(42)
195
+ end
196
+ ```
197
+
198
+ Nothing is deleted; each `DELETE`/`UPDATE` is written to `explain_file` (default `$stdout`). Lookups still run
199
+ against the database, so the output reflects real row ids.
200
+
201
+ | Executor option | Default |
202
+ |---|---|
203
+ | `explain:` | `false` |
204
+ | `explain_file:` | `$stdout` |
205
+ | `datetime_format:` | `'%Y-%m-%d %H:%M:%S'` |
206
+
207
+ Each executor keeps its own settings and applies them only for the duration of its `purge!` (per thread), so
208
+ creating another executor can't turn a dry run into a live one. `explain:` must be `true`, `false` or `nil`;
209
+ anything else (such as the string `'true'`) raises `ArgumentError` rather than running for real.
210
+
211
+ ## Metrics and instrumentation
212
+
213
+ Attach the built-in subscriber once at boot:
214
+
215
+ ```ruby
216
+ DBPurger::MetricSubscriber.auto_attach
217
+
218
+ executor.purge!(42)
219
+ DBPurger::MetricSubscriber.metrics.as_json
220
+ # => { took: 12.4, started_at: ..., finished_at: ...,
221
+ # purge_stats: { employments: { duration:, num_purges:, num_records: } },
222
+ # delete_stats: { employments: { duration:, num_delete_queries:, num_deleted:, num_expected_to_delete: } },
223
+ # lookup_stats: { ... }, filter_stats: { ... } }
224
+ ```
225
+
226
+ Metrics are reset at the start of each `Plan#purge!`. To feed your own telemetry, subscribe to the raw events
227
+ (all in the `db_purger` namespace):
228
+
229
+ | Event | Payload |
230
+ |---|---|
231
+ | `purge.db_purger` | `table_name`, `purge_field`, `deleted` |
232
+ | `next_batch.db_purger` | `table_name`, `start_id`, `num_records` |
233
+ | `delete_records.db_purger` | `table_name`, `num_records`, `records_deleted`, `deleted` |
234
+ | `search_filter.db_purger` | `table_name`, `num_records`, `num_records_selected` |
235
+
236
+ ## Caveats
237
+
238
+ - **Only the base table is the entry point.** Top-level `parent_table`/`child_table` calls made before
239
+ `base_table` are ignored by `Plan#purge!`.
240
+ - **Not one big transaction.** Each batch is its own set of statements (foreign-key children share a
241
+ transaction with their parent batch). An interrupted purge is safe to re-run with the same value.
242
+ - **Soft-deleted rows still match.** A `mark_deleted_field` table is not filtered on that field; add
243
+ `conditions:` if re-runs should skip already-marked rows.
244
+ - Always run explain mode against a copy of production before the first real purge with a new plan.
245
+
246
+ ## Development
247
+
248
+ ```sh
249
+ bundle install
250
+ bundle exec rspec # specs run against a throwaway SQLite database
251
+ bundle exec rubocop
252
+ script/console
253
+ ```
254
+
255
+ CI (`.github/workflows/ci.yml`) runs RuboCop and the specs on Ruby 4.0 for every push and pull request.
256
+ The HTML coverage report is attached to each run as the `coverage` artifact, and pushes to `master` refresh
257
+ the coverage badge on the `badges` branch.
258
+
259
+ See [ARCHITECTURE.md](ARCHITECTURE.md) for how the pieces fit together.
260
+
261
+ ## Releasing
262
+
263
+ 1. Bump `s.version` in `db-purger.gemspec` and merge to `master`.
264
+ 2. Tag and push: `git tag v0.5.0 && git push origin v0.5.0`
265
+
266
+ `.github/workflows/release.yml` re-runs CI, checks the tag matches the gemspec version, publishes to RubyGems
267
+ via trusted publishing (no API key), and creates a GitHub release with the `.gem` attached.
268
+
269
+ ## License
270
+
271
+ MIT — see [LICENSE.txt](LICENSE.txt).
@@ -5,16 +5,30 @@ module DBPurger
5
5
  class Config
6
6
  DEFAULT_DATETIME_FORMAT = '%Y-%m-%d %H:%M:%S'
7
7
 
8
- attr_writer :explain,
9
- :explain_file,
8
+ attr_writer :explain_file,
10
9
  :datetime_format
11
10
 
11
+ def initialize(options = {})
12
+ self.explain = options[:explain]
13
+ @explain_file = options[:explain_file]
14
+ @datetime_format = options[:datetime_format]
15
+ end
16
+
17
+ # Fail closed: a value like 'true' must not silently mean "run for real"
18
+ def explain=(value)
19
+ unless [true, false, nil].include?(value)
20
+ raise(ArgumentError, "explain must be true, false or nil, got #{value.inspect}")
21
+ end
22
+
23
+ @explain = value
24
+ end
25
+
12
26
  def explain?
13
27
  @explain == true
14
28
  end
15
29
 
16
30
  def explain_file
17
- (@explain_file || $stdout)
31
+ @explain_file || $stdout
18
32
  end
19
33
 
20
34
  def datetime_format
@@ -14,7 +14,6 @@ module DBPurger
14
14
  @tables = []
15
15
  end
16
16
 
17
- # rubocop:disable Metrics/AbcSize
18
17
  def build(base_table_name, field)
19
18
  write_table('base', base_table_name.to_s, field, [], nil)
20
19
  line_break
@@ -32,7 +31,6 @@ module DBPurger
32
31
  ignore_missing_tables
33
32
  @output
34
33
  end
35
- # rubocop:enable Metrics/AbcSize
36
34
 
37
35
  private
38
36
 
@@ -41,7 +39,7 @@ module DBPurger
41
39
  end
42
40
 
43
41
  def write(str)
44
- @output << (INDENT * @indent_depth) + str + "\n"
42
+ @output << "#{INDENT * @indent_depth}#{str}\n"
45
43
  end
46
44
 
47
45
  def line_break
@@ -49,7 +47,7 @@ module DBPurger
49
47
  end
50
48
 
51
49
  def add_parent_tables(base_table_name, field)
52
- @database.models.each do |model|
50
+ sorted_models.each do |model|
53
51
  next if model.table_name == base_table_name.to_s
54
52
  next unless column?(model, field)
55
53
 
@@ -72,7 +70,13 @@ module DBPurger
72
70
  end
73
71
 
74
72
  def find_child_models(model, field)
75
- model_has_many_associations(model).map(&:klass).select { |m| column?(m, field) }
73
+ model_has_many_associations(model).map(&:klass).select { |m| column?(m, field) }.sort_by(&:table_name)
74
+ end
75
+
76
+ # database.models order depends on how the adapter lists tables, which varies by platform;
77
+ # sort so the generated plan is deterministic
78
+ def sorted_models
79
+ @sorted_models ||= @database.models.sort_by(&:table_name)
76
80
  end
77
81
 
78
82
  def model_has_many_associations(model)
@@ -82,7 +86,7 @@ module DBPurger
82
86
  end
83
87
 
84
88
  def foreign_key_name(model)
85
- model.table_name.singularize + '_id'
89
+ "#{model.table_name.singularize}_id"
86
90
  end
87
91
 
88
92
  def column?(model, field)
@@ -101,7 +105,7 @@ module DBPurger
101
105
  end
102
106
 
103
107
  def ignore_missing_tables
104
- missing_tables = @database.models.map(&:table_name) - @tables
108
+ missing_tables = sorted_models.map(&:table_name) - @tables
105
109
  return if missing_tables.empty?
106
110
 
107
111
  line_break
@@ -8,14 +8,14 @@ module DBPurger
8
8
  def initialize(database, plan, options = {})
9
9
  @database = database
10
10
  @plan = plan.is_a?(Plan) ? plan : load_plan(plan)
11
- setup_config(options)
11
+ @config = Config.new(options)
12
12
  @error_io = $stderr
13
13
  end
14
14
 
15
15
  def purge!(purge_value)
16
16
  raise('purge_value is nil') if purge_value.nil?
17
17
 
18
- @plan.purge!(@database, purge_value)
18
+ ::DBPurger.with_config(@config) { @plan.purge!(@database, purge_value) }
19
19
  end
20
20
 
21
21
  def verify!
@@ -31,12 +31,6 @@ module DBPurger
31
31
  @plan_validator ||= PlanValidator.new(@database, @plan)
32
32
  end
33
33
 
34
- def setup_config(options)
35
- ::DBPurger.config.explain = options[:explain]
36
- ::DBPurger.config.explain_file = options[:explain_file]
37
- ::DBPurger.config.datetime_format = options[:datetime_format]
38
- end
39
-
40
34
  def load_plan(file)
41
35
  PlanBuilder
42
36
  .new(Plan.new)
@@ -23,7 +23,7 @@ module DBPurger
23
23
  self.class.metrics.update_purge_stats(
24
24
  event.payload[:table_name],
25
25
  event.duration,
26
- event.payload[:deleted]
26
+ event.payload[:deleted] || 0
27
27
  )
28
28
  end
29
29
 
@@ -31,7 +31,7 @@ module DBPurger
31
31
  self.class.metrics.update_delete_records_stats(
32
32
  event.payload[:table_name],
33
33
  event.duration,
34
- event.payload[:records_deleted],
34
+ event.payload[:records_deleted] || 0,
35
35
  event.payload[:num_records]
36
36
  )
37
37
  end
@@ -40,7 +40,7 @@ module DBPurger
40
40
  self.class.metrics.update_lookup_stats(
41
41
  event.payload[:table_name],
42
42
  event.duration,
43
- event.payload[:num_records]
43
+ event.payload[:num_records] || 0
44
44
  )
45
45
  end
46
46
 
@@ -48,8 +48,8 @@ module DBPurger
48
48
  self.class.metrics.update_search_filter_stats(
49
49
  event.payload[:table_name],
50
50
  event.duration,
51
- event.payload[:num_records],
52
- event.payload[:num_records_selected]
51
+ event.payload[:num_records] || 0,
52
+ event.payload[:num_records_selected] || 0
53
53
  )
54
54
  end
55
55
  end
@@ -18,6 +18,8 @@ module DBPurger
18
18
  end
19
19
 
20
20
  def purge!(database, purge_value)
21
+ raise('plan has no base_table') unless @base_table
22
+
21
23
  MetricSubscriber.reset!
22
24
  num_deleted = PurgeTable.new(database, @base_table, @base_table.field, purge_value).purge!
23
25
  MetricSubscriber.finished!
@@ -7,6 +7,7 @@ module DBPurger
7
7
  class PlanValidator
8
8
  include ActiveModel::Validations
9
9
 
10
+ validate :validate_base_table
10
11
  validate :validate_no_missing_tables
11
12
  validate :validate_no_unknown_tables
12
13
  validate :validate_tables
@@ -28,6 +29,10 @@ module DBPurger
28
29
 
29
30
  private
30
31
 
32
+ def validate_base_table
33
+ errors.add(:base_table, 'is required') unless @plan.base_table
34
+ end
35
+
31
36
  def validate_no_missing_tables
32
37
  errors.add(:missing_tables, missing_tables.sort.join(',')) unless missing_tables.empty?
33
38
  end
@@ -51,6 +56,26 @@ module DBPurger
51
56
  errors.add(:table, "#{table.name}.#{field} is missing in the database")
52
57
  end
53
58
  end
59
+
60
+ validate_mark_deleted_field(table, model)
61
+ validate_batch_size(table)
62
+ validate_nested_tables_have_primary_key(table, model)
63
+ end
64
+
65
+ def validate_mark_deleted_field(table, model)
66
+ return if table.mark_deleted_field.nil? || model.column_names.include?(table.mark_deleted_field.to_s)
67
+
68
+ errors.add(:table, "#{table.name}.#{table.mark_deleted_field} (mark_deleted_field) is missing in the database")
69
+ end
70
+
71
+ def validate_batch_size(table)
72
+ errors.add(:table, "#{table.name} batch_size must be positive") unless table.batch_size.to_i.positive?
73
+ end
74
+
75
+ def validate_nested_tables_have_primary_key(table, model)
76
+ return if model.primary_key || !table.nested_key_tables?
77
+
78
+ errors.add(:table, "#{table.name} has no primary key and cannot have nested child or parent tables")
54
79
  end
55
80
 
56
81
  def find_model_for_table(table)
@@ -21,9 +21,11 @@ module DBPurger
21
21
  ActiveSupport::Notifications.instrument('purge.db_purger',
22
22
  table_name: @table.name,
23
23
  purge_field: @purge_field) do |payload|
24
+ payload[:deleted] = @num_deleted
24
25
  if model.primary_key
25
26
  purge_in_batches!
26
27
  else
28
+ ensure_no_nested_key_tables!
27
29
  purge_all!
28
30
  end
29
31
  purge_search_tables
@@ -34,6 +36,13 @@ module DBPurger
34
36
 
35
37
  private
36
38
 
39
+ # without a primary key there are no batch ids to propagate, so nested tables would be silently skipped
40
+ def ensure_no_nested_key_tables!
41
+ return unless @table.nested_key_tables?
42
+
43
+ raise("#{@table.name} has no primary key and cannot have nested child or parent tables")
44
+ end
45
+
37
46
  def purge_all!
38
47
  scope = model.where(@purge_field => @purge_value)
39
48
  scope = scope.where(@table.conditions) if @table.conditions
@@ -41,13 +50,27 @@ module DBPurger
41
50
  end
42
51
 
43
52
  def purge_in_batches!
53
+ unless @table.parent_tables?
54
+ each_batch do |batch|
55
+ purge_nested_tables(batch) if @table.nested_tables?
56
+ delete_records(batch)
57
+ end
58
+ return
59
+ end
60
+
61
+ # Parent tables may reference this table's rows and be referenced by its child tables,
62
+ # so purge them after the children but before this table's rows.
63
+ each_batch { |batch| purge_nested_tables(batch) }
64
+ purge_parent_tables
65
+ each_batch { |batch| delete_records(batch) }
66
+ end
67
+
68
+ def each_batch
44
69
  start_id = nil
45
70
  until (batch = next_batch(start_id)).empty?
46
- start_id = batch.last.send(model.primary_key)
47
- purge_nested_tables(batch) if @table.nested_tables?
48
- delete_records(batch)
71
+ start_id = batch.last[model.primary_key]
72
+ yield batch
49
73
  end
50
- purge_parent_tables
51
74
  end
52
75
 
53
76
  def next_batch(start_id)
@@ -38,8 +38,9 @@ module DBPurger
38
38
  end
39
39
  end
40
40
 
41
+ # record[] reads the column; send would call a same-named method instead (e.g. a column called "reload")
41
42
  def batch_values(batch, field)
42
- batch.map { |record| record.send(field) }.compact
43
+ batch.map { |record| record[field] }.compact
43
44
  end
44
45
 
45
46
  def foreign_tables?
@@ -72,7 +73,7 @@ module DBPurger
72
73
  else
73
74
  scope.to_sql.sub(/SELECT .*?FROM/, 'DELETE FROM')
74
75
  end
75
- ::DBPurger.config.explain_file.puts(sql + ';')
76
+ ::DBPurger.config.explain_file.puts("#{sql};")
76
77
  scope.count
77
78
  end
78
79
 
@@ -32,6 +32,15 @@ module DBPurger
32
32
  @nested_plan != nil
33
33
  end
34
34
 
35
+ def parent_tables?
36
+ nested_tables? && !@nested_plan.parent_tables.empty?
37
+ end
38
+
39
+ # nested tables that depend on this table's rows (search tables scan independently)
40
+ def nested_key_tables?
41
+ nested_tables? && !(@nested_plan.child_tables.empty? && @nested_plan.parent_tables.empty?)
42
+ end
43
+
35
44
  def tables
36
45
  @nested_plan ? @nested_plan.tables : []
37
46
  end
@@ -45,7 +54,7 @@ module DBPurger
45
54
  end
46
55
 
47
56
  def mark_deleted_value
48
- @mark_deleted_value || 1
57
+ @mark_deleted_value.nil? ? 1 : @mark_deleted_value
49
58
  end
50
59
  end
51
60
  end
data/lib/db-purger.rb CHANGED
@@ -15,7 +15,22 @@ module DBPurger
15
15
  autoload :PlanValidator, 'db-purger/plan_validator'
16
16
  autoload :Table, 'db-purger/table'
17
17
 
18
+ # The config in effect for the current thread: the one set by with_config, else the global default
18
19
  def self.config
19
- @config ||= Config.new
20
+ Thread.current[:db_purger_config] || default_config
21
+ end
22
+
23
+ def self.default_config
24
+ @default_config ||= Config.new
25
+ end
26
+
27
+ # Runs the block with config in effect for this thread only, so concurrent or
28
+ # later executors can't change explain mode out from under a running purge
29
+ def self.with_config(config)
30
+ previous = Thread.current[:db_purger_config]
31
+ Thread.current[:db_purger_config] = config
32
+ yield
33
+ ensure
34
+ Thread.current[:db_purger_config] = previous
20
35
  end
21
36
  end
metadata CHANGED
@@ -1,49 +1,55 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: db-purger
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.4.0
4
+ version: 0.5.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Doug Youch
8
- autorequire:
9
8
  bindir: bin
10
9
  cert_chain: []
11
- date: 2022-03-16 00:00:00.000000000 Z
10
+ date: 1980-01-02 00:00:00.000000000 Z
12
11
  dependencies:
13
12
  - !ruby/object:Gem::Dependency
14
- name: dynamic-active-model
13
+ name: activerecord
15
14
  requirement: !ruby/object:Gem::Requirement
16
15
  requirements:
17
16
  - - ">="
18
17
  - !ruby/object:Gem::Version
19
- version: '0'
18
+ version: '7.0'
20
19
  type: :runtime
21
20
  prerelease: false
22
21
  version_requirements: !ruby/object:Gem::Requirement
23
22
  requirements:
24
23
  - - ">="
25
24
  - !ruby/object:Gem::Version
26
- version: '0'
25
+ version: '7.0'
27
26
  - !ruby/object:Gem::Dependency
28
- name: inheritance-helper
27
+ name: dynamic-active-model
29
28
  requirement: !ruby/object:Gem::Requirement
30
29
  requirements:
31
- - - ">="
30
+ - - "~>"
32
31
  - !ruby/object:Gem::Version
33
- version: '0'
32
+ version: '0.7'
34
33
  type: :runtime
35
34
  prerelease: false
36
35
  version_requirements: !ruby/object:Gem::Requirement
37
36
  requirements:
38
- - - ">="
37
+ - - "~>"
39
38
  - !ruby/object:Gem::Version
40
- version: '0'
41
- description: Purge database tables by top level id in batches
39
+ version: '0.7'
40
+ description: DB Purger deletes (or soft-deletes) every row related to a single top-level
41
+ record (e.g. a company or account) using a declarative Ruby purge plan. Tables are
42
+ purged in primary-key batches, child tables before their parents, with plan validation
43
+ against the live schema, an explain (dry-run) mode that emits SQL, and ActiveSupport::Notifications
44
+ instrumentation for metrics.
42
45
  email: dougyouch@gmail.com
43
46
  executables: []
44
47
  extensions: []
45
48
  extra_rdoc_files: []
46
49
  files:
50
+ - ARCHITECTURE.md
51
+ - LICENSE.txt
52
+ - README.md
47
53
  - lib/db-purger.rb
48
54
  - lib/db-purger/config.rb
49
55
  - lib/db-purger/dynamic_plan_builder.rb
@@ -60,8 +66,10 @@ files:
60
66
  homepage: https://github.com/dougyouch/db-purger
61
67
  licenses:
62
68
  - MIT
63
- metadata: {}
64
- post_install_message:
69
+ metadata:
70
+ source_code_uri: https://github.com/dougyouch/db-purger
71
+ bug_tracker_uri: https://github.com/dougyouch/db-purger/issues
72
+ rubygems_mfa_required: 'true'
65
73
  rdoc_options: []
66
74
  require_paths:
67
75
  - lib
@@ -69,15 +77,14 @@ required_ruby_version: !ruby/object:Gem::Requirement
69
77
  requirements:
70
78
  - - ">="
71
79
  - !ruby/object:Gem::Version
72
- version: '0'
80
+ version: '3.2'
73
81
  required_rubygems_version: !ruby/object:Gem::Requirement
74
82
  requirements:
75
83
  - - ">="
76
84
  - !ruby/object:Gem::Version
77
85
  version: '0'
78
86
  requirements: []
79
- rubygems_version: 3.2.3
80
- signing_key:
87
+ rubygems_version: 4.0.20
81
88
  specification_version: 4
82
- summary: DB Purger by top level id
89
+ summary: Purge all data tied to a top-level id across related tables, in batches
83
90
  test_files: []