db-purger 0.4.1 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 4673722a3af60f2772b9276dfe38e7595946c97080a4db8c7428a35982d11b22
4
- data.tar.gz: 2cdb3d92f389b887994f842cc7a0271f72ec3fab79fa858de11d55c5568638b5
3
+ metadata.gz: 7b756bb20c0bcea96d050a3772cb2606676891d485de6db0bf9058253f0092c3
4
+ data.tar.gz: 9dfe62f982ecd99ef96dadd535a7496d1802d0392083b8f63f9166f354bb9368
5
5
  SHA512:
6
- metadata.gz: 692efaffd2f2eb20fe4627717aafd59175e22e1f391f2d827398f69a86eb851b3ee8d40e556b0647f5b0fc2ab59bee757a5bc1dd448d7383cb4dde1d80d7bb0a
7
- data.tar.gz: 460ad2ea2ce280890bf7de9bf945b148a65953e8c930cbae75b8bdda6ce861e1bfcbc89ac42e6dee566a8242140ee018fc505057f74633bf497d43a2429dae20
6
+ metadata.gz: 2a365e0a0c25b33eb79060930289c2820823e9e25fefbf4edcdf13ccfad6e55595e3d380d5da97a76476262051b9ee36095cdf51e30ebf3248bcf43dfd8f0fa2
7
+ data.tar.gz: 9171105e396d405cacde9debad9b6923d1d318c24ebbf269a9f872acbf6e087a0ea90506e9d24c0ab99407e1874fcd9d46e74afb73a736203b1f1da178791970
data/ARCHITECTURE.md ADDED
@@ -0,0 +1,179 @@
1
+ # Architecture
2
+
3
+ db-purger is small (~900 lines) and splits cleanly into three layers: **describe** a purge (plan + DSL),
4
+ **check** it (validator), and **run** it (purgers + instrumentation).
5
+
6
+ ```
7
+ plan file / PlanBuilder.build { ... }
8
+ │
9
+ ▼
10
+ ┌──────────────┐ ┌─────────┐ ┌─────────┐
11
+ │ PlanBuilder │──▶│ Plan │──▶│ Table │──┐ each Table owns a nested Plan
12
+ │ (DSL) │ │ │ │ │◀─┘ (recursive tree)
13
+ └──────────────┘ └─────────┘ └─────────┘
14
+ │
15
+ ┌─────────────────┼──────────────────┐
16
+ ▼ ▼ ▼
17
+ PlanValidator Executor ─────▶ PurgeTable / PurgeTableScanner
18
+ (schema check) (entry point) │ (PurgeTableHelper)
19
+ ▼
20
+ ActiveSupport::Notifications
21
+ (*.db_purger events)
22
+ │
23
+ ▼
24
+ MetricSubscriber ─▶ Metrics
25
+ ```
26
+
27
+ ## Components
28
+
29
+ | File | Responsibility |
30
+ |---|---|
31
+ | `lib/db-purger.rb` | Autoloads everything; holds the global `DBPurger.config`. |
32
+ | `config.rb` | Global options: `explain?`, `explain_file`, `datetime_format`. |
33
+ | `table.rb` | Value object for one table in the plan: name, match field, options, and a lazily created nested `Plan`. |
34
+ | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, search and ignored tables. `#purge!` is the run entry point; `#root_tables` are the tables it starts from. |
35
+ | `plan_builder.rb` | The DSL. `instance_eval`s a plan file or block against a `Plan`; nested blocks get a new builder bound to that table's nested plan. |
36
+ | `plan_validator.rb` | `ActiveModel::Validations` over plan vs. schema: missing tables, unknown tables, unknown columns. |
37
+ | `executor.rb` | Convenience façade: loads a plan file, applies config options, `verify!`, `purge!`. |
38
+ | `purge_table.rb` | Purges one table by `field = value(s)` in primary-key batches. Recurses into nested tables. |
39
+ | `purge_table_scanner.rb` | Purges a `purge_table_search` table: full `find_in_batches` scan filtered through the user's `search_proc`. |
40
+ | `purge_table_helper.rb` | Shared behaviour for both purgers: nested-table recursion, delete vs. soft delete vs. explain, transactions. |
41
+ | `metrics.rb` / `metric_subscriber.rb` | Aggregate timing and row counts per table from the notification events. |
42
+ | `dynamic_plan_builder.rb` | Generates plan-file source (`build` for a base table, `build_for` for several roots); a bootstrap tool, not used at purge time. |
43
+ | `association_graph.rb` | For the generator: which tables reference a model, and by which column, from its `has_many`/`has_one`/HABTM reflections. |
44
+ | `plan_writer.rb` | For the generator: renders plan DSL text and records which tables it wrote. |
45
+
46
+ ## The plan tree
47
+
48
+ `PlanBuilder` always attaches tables to the *current* plan's `base_table.nested_plan` once a base table
49
+ exists. So a plan file is really:
50
+
51
+ ```
52
+ Plan (root)
53
+ └── base_table: companies(:id)
54
+ └── nested Plan
55
+ ├── parent_tables: [company_tags(:company_id)]
56
+ ├── child_tables: [employments(:company_id) ─▶ nested Plan ..., websites(:id, fk: website_id) ...]
57
+ ├── search_tables: [users(:id)]
58
+ └── ignore_tables
59
+ ```
60
+
61
+ Without a `base_table`, top-level `parent_table`s stay in the root plan's `parent_tables` and each one is a
62
+ root:
63
+
64
+ ```
65
+ Plan (root)
66
+ ├── parent_tables: [calls(:oid) ─▶ nested Plan ..., emails(:oid) ─▶ nested Plan ..., sms_messages(:oid) ...]
67
+ ├── search_tables
68
+ └── ignore_tables
69
+ ```
70
+
71
+ `Plan#tables` flattens this tree for validation; `Table#foreign_keys` collects the `foreign_key:` columns of
72
+ a table's direct children so the purger can `SELECT` them alongside the primary key.
73
+
74
+ ## Purge algorithm
75
+
76
+ `Plan#purge!` resets metrics and starts a `PurgeTable` with the purge value on each root table
77
+ (`root_tables` = the `base_table`, if any, followed by top-level `parent_tables`, in declaration order), then runs
78
+ any top-level search tables. With a `base_table`, `root_tables` is just `[base_table]`, which is the original
79
+ single-root algorithm. Each
80
+ `PurgeTable#purge!` does:
81
+
82
+ ```
83
+ each_batch = loop:
84
+ batch = SELECT pk, <child foreign_keys> FROM t
85
+ WHERE field IN (values) [AND conditions] [AND pk > last_pk]
86
+ ORDER BY pk LIMIT batch_size
87
+ break if batch empty
88
+
89
+ purge_children(batch) =
90
+ for each child_table without foreign_key: # rows pointing at us
91
+ PurgeTable(child, child.field, batch.pks).purge! (recursive)
92
+
93
+ delete_rows(batch) =
94
+ if any child has foreign_key: # rows we point at
95
+ TRANSACTION
96
+ delete batch by pk
97
+ for each fk child: PurgeTable(child, child.field, batch.<fk values>).purge!
98
+ else
99
+ delete batch by pk
100
+
101
+ if table has a primary key:
102
+ if no parent_tables:
103
+ each_batch: purge_children(batch); delete_rows(batch)
104
+ else: # two passes
105
+ each_batch: purge_children(batch)
106
+ for each parent_table: # siblings sharing the key
107
+ PurgeTable(parent, parent.field, original purge value).purge!
108
+ each_batch: delete_rows(batch)
109
+ else:
110
+ raise if nested child/parent tables # no ids to propagate
111
+ single DELETE WHERE field = value [AND conditions]
112
+
113
+ for each search_table:
114
+ PurgeTableScanner(search_table).purge!
115
+ ```
116
+
117
+ Key properties:
118
+
119
+ - **Depth-first, children first.** A row is only deleted after everything referencing it, so FK
120
+ constraints hold without `ON DELETE CASCADE`.
121
+ - **Two passes when there are parent tables.** A parent table can reference this table (e.g.
122
+ `company_tags.company_id → companies.id`) *and* be referenced by one of its children, so it is purged
123
+ between the child pass and the delete pass. Tables without parent tables keep the single pass.
124
+ - **Keyset pagination** (`pk > last_pk`) rather than `OFFSET`, so batches stay cheap on large tables and still
125
+ advance in explain mode where nothing is actually deleted.
126
+ - **Bounded memory.** Only one batch of ids per level of the tree is held at a time.
127
+ - **Idempotent re-runs.** Every step re-derives its rows from the database; there is no saved cursor, so a
128
+ crashed purge is resumed by running it again.
129
+ - **Parent vs. child** is about *which value* is propagated: children get the enclosing batch's primary keys,
130
+ parents get the original purge value.
131
+
132
+ `PurgeTableScanner` follows the same shape but sources batches from `find_in_batches` over the entire table
133
+ (plus `conditions`) and narrows each batch with `search_proc` before recursing and deleting.
134
+
135
+ ## Delete strategies
136
+
137
+ `PurgeTableHelper#delete_records_with_instrumentation` picks one of three actions for every scope:
138
+
139
+ 1. **explain** (`DBPurger.config.explain?`) — rewrite `scope.to_sql` into `DELETE`/`UPDATE ... SET`, write it
140
+ to `explain_file`, return `scope.count`.
141
+ 2. **soft delete** (`mark_deleted_field`) — `scope.update_all(field => value)`.
142
+ 3. **hard delete** — `scope.delete_all`.
143
+
144
+ All three go through ActiveRecord's `*_all` methods: no model callbacks or validations run.
145
+
146
+ ## Instrumentation
147
+
148
+ Every unit of work is wrapped in `ActiveSupport::Notifications.instrument` under the `db_purger` namespace
149
+ (`purge`, `next_batch`, `delete_records`, `search_filter`). The purgers never talk to `Metrics` directly;
150
+ `MetricSubscriber` (an `ActiveSupport::Subscriber`) translates events into `Metrics` counters. This keeps the
151
+ purge code free of reporting concerns and lets callers attach their own subscribers (StatsD, logs, progress
152
+ bars) without changes to the library.
153
+
154
+ `PurgeTableScanner` uses the lower-level `instrumenter.start/finish` pair for `next_batch` because the batch
155
+ fetch happens inside `find_in_batches` rather than in a block the scanner controls.
156
+
157
+ ## Dependencies and boundaries
158
+
159
+ - **ActiveRecord** is the only database interface. The library relies on `where`, `select`, `order`,
160
+ `limit`, `find_in_batches`, `delete_all`, `update_all`, `transaction`, `to_sql` and
161
+ `connection.quote*`, so it is adapter-agnostic (specs use SQLite).
162
+ - **dynamic-active-model** supplies the `database` object. The purge code only calls `database.models` and
163
+ matches on `model.table_name`, so any object with that shape works. `DynamicPlanBuilder` additionally
164
+ relies on `reflect_on_all_associations`.
165
+ - **Config scoping.** `DBPurger.config` returns the config set by `DBPurger.with_config` for the current
166
+ thread, falling back to a process-wide default. `Executor#purge!` wraps the run in `with_config`, so executors
167
+ with different explain settings can't interfere. `MetricSubscriber.metrics` is still process-wide.
168
+
169
+ ## Testing
170
+
171
+ - `spec/db-purger/*` — unit specs for the builder, validator, executor and purger.
172
+ - `spec/integrations/*` — end-to-end purges over the schema in `spec/support/db/schema.rb`, asserting row-count
173
+ deltas per table and, for explain mode, the exact SQL in `spec/fixtures/delete_plan.sql`.
174
+ - `spec/support/test_db.rb` builds a fresh SQLite database and dynamic models (`TestDB::*`) for each run.
175
+ - `spec/integrations/multi_root_plan_spec.rb` purges one org from a separate outreach schema (`spec/support/outreach_db.rb`,
176
+ real FOREIGN KEY constraints) seeded with two orgs by `OutreachSeeder`, and asserts the exact surviving rows of
177
+ every table for the generated plan, a hand-written small-batch plan and the equivalent `base_table` plan.
178
+ - `spec/support/throwaway_db.rb` builds standalone SQLite databases from raw SQL for schema-specific specs.
179
+ - `spec/fixtures/*.plan.rb` — plan files used for loading and validation cases.
data/LICENSE.txt ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2019 Douglas Youch
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,313 @@
1
+ # db-purger
2
+
3
+ [![CI](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml/badge.svg?branch=master)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
4
+ [![Coverage](https://raw.githubusercontent.com/dougyouch/db-purger/badges/coverage.svg)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
5
+ [![Branch coverage](https://raw.githubusercontent.com/dougyouch/db-purger/badges/branch-coverage.svg)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
6
+ [![Gem Version](https://img.shields.io/gem/v/db-purger)](https://rubygems.org/gems/db-purger)
7
+
8
+ Purge every row tied to a single top-level record — a company, an account, a tenant — across all of the
9
+ tables that reference it, in batches, from a declarative Ruby plan.
10
+
11
+ ```ruby
12
+ executor = DBPurger::Executor.new(database, 'config/company.plan.rb')
13
+ executor.verify! # fail fast if the plan doesn't cover the schema
14
+ executor.purge!(42) # delete company 42 and everything that hangs off it
15
+ ```
16
+
17
+ ## Why
18
+
19
+ Deleting a tenant from a relational database is rarely one `DELETE`. Rows are spread across dozens of
20
+ tables, some keyed directly on the tenant id, some several joins away, some polymorphic, some that must be
21
+ soft-deleted rather than removed. `ON DELETE CASCADE` is often absent, and a single giant delete will lock
22
+ tables and blow out replication.
23
+
24
+ db-purger lets you describe those relationships once, in a plan file, and then:
25
+
26
+ - deletes in **primary-key batches** (default 10,000) so no single statement gets too large
27
+ - deletes rows **only after the rows that reference them**, so foreign-key constraints hold without `ON DELETE CASCADE`
28
+ - **validates** the plan against the live schema, so a newly added table can't be silently forgotten
29
+ - supports **soft deletes** (`UPDATE ... SET deleted_at = ...`) per table
30
+ - has an **explain mode** that prints the SQL it would run instead of running it
31
+ - emits **ActiveSupport::Notifications** events, with a built-in subscriber that collects timing and row counts
32
+
33
+ ## Installation
34
+
35
+ ```ruby
36
+ # Gemfile
37
+ gem 'db-purger'
38
+ ```
39
+
40
+ Requires Ruby >= 3.2 and ActiveRecord >= 7.0. Models are supplied by
41
+ [dynamic-active-model](https://github.com/dougyouch/dynamic-active-model), which builds ActiveRecord classes
42
+ directly from the database schema.
43
+
44
+ ## Quick start
45
+
46
+ ### 1. Load the database
47
+
48
+ db-purger works against a `DynamicActiveModel::Database`. Anything that responds to `#models` (returning
49
+ ActiveRecord classes) will do.
50
+
51
+ ```ruby
52
+ require 'active_record'
53
+ require 'dynamic-active-model'
54
+ require 'db-purger'
55
+
56
+ module PurgeDB; end
57
+
58
+ database = DynamicActiveModel::Explorer.explore(
59
+ PurgeDB,
60
+ { adapter: 'mysql2', host: 'localhost', database: 'app', username: 'app' },
61
+ %w[schema_migrations ar_internal_metadata] # tables to skip entirely
62
+ )
63
+ ```
64
+
65
+ ### 2. Write a plan
66
+
67
+ A plan is Ruby, evaluated with the plan DSL. Given this schema:
68
+
69
+ ```
70
+ companies (id, website_id, ...)
71
+ employments (id, company_id, user_id, ...)
72
+ employment_notes (id, employment_id, ...)
73
+ company_tags (company_id, tag_id) -- no primary key
74
+ events (id, model_type, model_id, ...) -- polymorphic
75
+ websites (id, content_id, ...)
76
+ contents (id, ...)
77
+ tags, jobs -- shared lookup tables, never purged
78
+ ```
79
+
80
+ a plan to purge one company looks like:
81
+
82
+ ```ruby
83
+ # config/company.plan.rb
84
+ base_table(:companies, :id)
85
+
86
+ # Tables keyed directly on the purge value (company_id = 42)
87
+ parent_table(:company_tags, :company_id)
88
+
89
+ # Tables keyed on the base table's primary key, purged batch-by-batch
90
+ child_table(:employments, :company_id) do
91
+ child_table(:employment_notes, :employment_id)
92
+ child_table(:events, :model_id, conditions: { model_type: 'PurgeDB::Employment' })
93
+ end
94
+
95
+ child_table(:events, :model_id, conditions: { model_type: 'PurgeDB::Company' })
96
+
97
+ # The company row points at the website (companies.website_id -> websites.id)
98
+ child_table(:websites, :id, foreign_key: :website_id) do
99
+ child_table(:contents, :id, foreign_key: :content_id)
100
+ end
101
+
102
+ # Shared tables that are intentionally left alone
103
+ ignore_table :tags
104
+ ignore_table :jobs
105
+ ignore_table(/\Atmp_/) # regexps are allowed
106
+ ```
107
+
108
+ ### 3. Verify and purge
109
+
110
+ ```ruby
111
+ executor = DBPurger::Executor.new(database, 'config/company.plan.rb')
112
+ executor.verify! # raises 'purge plan failed verification', errors printed to $stderr
113
+ deleted = executor.purge!(42)
114
+ ```
115
+
116
+ `purge!` returns the number of root-table rows deleted (the base table, plus any top-level `parent_table`s).
117
+
118
+ ### Plans with several top-level tables
119
+
120
+ `base_table` is shorthand for "one root table, with everything after it nested underneath". When several tables
121
+ are equally top-level (an outreach product's `emails`, `sms_messages` and `calls`, all keyed by `oid`), leave
122
+ `base_table` out and declare each root as a top-level `parent_table`:
123
+
124
+ ```ruby
125
+ # config/outreach.plan.rb
126
+ parent_table(:calls, :oid) do # calls.email_id -> emails.id, so calls go first
127
+ child_table(:call_notes, :call_id)
128
+ child_table(:call_recordings, :call_id)
129
+ child_table(:call_tags, :call_id)
130
+ end
131
+
132
+ parent_table(:emails, :oid) do
133
+ child_table(:email_attachments, :email_id)
134
+ end
135
+
136
+ parent_table(:sms_messages, :oid) do
137
+ child_table(:sms_deliveries, :sms_message_id)
138
+ end
139
+
140
+ ignore_table :users
141
+ ```
142
+
143
+ Each root is purged by `oid = purge_value`, children first, **in declaration order**: when one root's rows
144
+ reference another's, declare the referencing root first. Without a `base_table`, top-level `child_table`s are
145
+ an error (there is no enclosing batch to take ids from). Existing `base_table` plans run exactly as before.
146
+
147
+ ## The plan DSL
148
+
149
+ | Method | Meaning |
150
+ |---|---|
151
+ | `base_table(table, field, opts = {}, &block)` | Optional single root. Rows where `field = purge_value` are purged. Declare it **first** — every subsequent top-level call nests under it. |
152
+ | `child_table(table, field, opts = {}, &block)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch. Purged before that batch is deleted. |
153
+ | `child_table(table, :id, foreign_key: :col, &block)` | Inverted relationship: the *enclosing* table holds `col` pointing at this table's `id`. Deleted in the same transaction, right after the enclosing batch. |
154
+ | `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. At the top level of a plan without a `base_table`, each one is a root, purged in declaration order. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
155
+ | `purge_table_search(table, field, opts = {}) { \|batch\| ... }` | Scans the whole table in batches; the block receives each batch and returns the records to purge. For orphans that can't be reached by a key. |
156
+ | `ignore_table(name_or_regexp)` | Exclude a table from validation. |
157
+
158
+ Blocks nest arbitrarily deep. Because `purge_table_search` uses its block as the filter, nest tables under it
159
+ with `.nested_plan`:
160
+
161
+ ```ruby
162
+ purge_table_search(:users, :id) do |users|
163
+ users = users.index_by(&:id)
164
+ PurgeDB::Employment.where(user_id: users.keys).pluck(:user_id).each { |id| users.delete(id) }
165
+ users.values # users with no remaining employments
166
+ end.nested_plan do
167
+ child_table(:events, :model_id, conditions: { model_type: 'PurgeDB::User' })
168
+ end
169
+ ```
170
+
171
+ ### Table options
172
+
173
+ | Option | Default | Description |
174
+ |---|---|---|
175
+ | `batch_size:` | `10_000` | Rows fetched and deleted per batch. |
176
+ | `conditions:` | none | Extra `where` applied to every query for this table (hash or SQL string). |
177
+ | `foreign_key:` | none | See `child_table` above. |
178
+ | `mark_deleted_field:` | none | Soft delete: `UPDATE table SET field = value` instead of `DELETE`. |
179
+ | `mark_deleted_value:` | `1` | Value written to `mark_deleted_field`. `Time` values are formatted with `datetime_format` in explain output. |
180
+
181
+ Tables **without a primary key** (e.g. join tables) are purged with a single unbatched `DELETE ... WHERE
182
+ field = value`; nested tables are not supported under them.
183
+
184
+ ### Building a plan in code
185
+
186
+ ```ruby
187
+ plan = DBPurger::PlanBuilder.build do
188
+ base_table(:companies, :id)
189
+ child_table(:employments, :company_id)
190
+ end
191
+
192
+ DBPurger::Executor.new(database, plan).purge!(42)
193
+ # or, without the executor:
194
+ plan.purge!(database, 42)
195
+ ```
196
+
197
+ ### Generating a starting plan
198
+
199
+ `DynamicPlanBuilder` walks the `has_many`, `has_one` and `has_and_belongs_to_many` associations
200
+ dynamic-active-model discovered, using each association's real foreign key, and emits a plan file listing every
201
+ unreachable table as `ignore_table`.
202
+
203
+ ```ruby
204
+ builder = DBPurger::DynamicPlanBuilder.new(database)
205
+ puts builder.build(:companies, :id) # single base_table plan
206
+ puts builder.build_for(:oid) # one top-level parent_table per table holding oid
207
+ ```
208
+
209
+ `build_for` makes **every** table holding the field a root, so rows with a null foreign key to another root
210
+ (an `email_recipients` row without an email) are still purged, and orders the roots so a root referencing
211
+ another root's rows comes first. HABTM join tables are emitted as leaves, never walking into the shared table on
212
+ the other side, and nothing is nested under a table without a primary key. A foreign-key cycle is written as a
213
+ comment instead of recursing.
214
+
215
+ Treat the output as a first draft: it cannot infer polymorphic (`as:`), soft-delete, `belongs_to`-owned
216
+ (`foreign_key:`) or search rules.
217
+
218
+ ## Validation
219
+
220
+ `Executor#verify!` (or `DBPurger::PlanValidator.new(database, plan).valid?`) checks that:
221
+
222
+ - every table in the database is either in the plan or ignored (`missing_tables`)
223
+ - every table in the plan exists in the database (`unknown_tables`)
224
+ - every field, `foreign_key` and `mark_deleted_field` named in the plan is a real column
225
+ - the plan has a `base_table` or at least one top-level `parent_table`, no top-level `child_table` is left
226
+ unreachable, and every `batch_size` is positive
227
+ - tables without a primary key have no nested child or parent tables (there would be no ids to propagate)
228
+
229
+ Run it in CI against your schema so a new table can't ship without a purge decision.
230
+
231
+ ## Explain mode (dry run)
232
+
233
+ ```ruby
234
+ File.open('purge.sql', 'w') do |io|
235
+ executor = DBPurger::Executor.new(database, 'config/company.plan.rb', explain: true, explain_file: io)
236
+ executor.purge!(42)
237
+ end
238
+ ```
239
+
240
+ Nothing is deleted; each `DELETE`/`UPDATE` is written to `explain_file` (default `$stdout`). Lookups still run
241
+ against the database, so the output reflects real row ids.
242
+
243
+ | Executor option | Default |
244
+ |---|---|
245
+ | `explain:` | `false` |
246
+ | `explain_file:` | `$stdout` |
247
+ | `datetime_format:` | `'%Y-%m-%d %H:%M:%S'` |
248
+
249
+ Each executor keeps its own settings and applies them only for the duration of its `purge!` (per thread), so
250
+ creating another executor can't turn a dry run into a live one. `explain:` must be `true`, `false` or `nil`;
251
+ anything else (such as the string `'true'`) raises `ArgumentError` rather than running for real.
252
+
253
+ ## Metrics and instrumentation
254
+
255
+ Attach the built-in subscriber once at boot:
256
+
257
+ ```ruby
258
+ DBPurger::MetricSubscriber.auto_attach
259
+
260
+ executor.purge!(42)
261
+ DBPurger::MetricSubscriber.metrics.as_json
262
+ # => { took: 12.4, started_at: ..., finished_at: ...,
263
+ # purge_stats: { employments: { duration:, num_purges:, num_records: } },
264
+ # delete_stats: { employments: { duration:, num_delete_queries:, num_deleted:, num_expected_to_delete: } },
265
+ # lookup_stats: { ... }, filter_stats: { ... } }
266
+ ```
267
+
268
+ Metrics are reset at the start of each `Plan#purge!`. To feed your own telemetry, subscribe to the raw events
269
+ (all in the `db_purger` namespace):
270
+
271
+ | Event | Payload |
272
+ |---|---|
273
+ | `purge.db_purger` | `table_name`, `purge_field`, `deleted` |
274
+ | `next_batch.db_purger` | `table_name`, `start_id`, `num_records` |
275
+ | `delete_records.db_purger` | `table_name`, `num_records`, `records_deleted`, `deleted` |
276
+ | `search_filter.db_purger` | `table_name`, `num_records`, `num_records_selected` |
277
+
278
+ ## Caveats
279
+
280
+ - **Declare `base_table` first.** A `child_table` declared before it is never reached (the validator reports
281
+ it); a `parent_table` declared before it becomes a separate root, purged after the base table.
282
+ - **Not one big transaction.** Each batch is its own set of statements (foreign-key children share a
283
+ transaction with their parent batch). An interrupted purge is safe to re-run with the same value.
284
+ - **Soft-deleted rows still match.** A `mark_deleted_field` table is not filtered on that field; add
285
+ `conditions:` if re-runs should skip already-marked rows.
286
+ - Always run explain mode against a copy of production before the first real purge with a new plan.
287
+
288
+ ## Development
289
+
290
+ ```sh
291
+ bundle install
292
+ bundle exec rspec # specs run against a throwaway SQLite database
293
+ bundle exec rubocop
294
+ script/console
295
+ ```
296
+
297
+ CI (`.github/workflows/ci.yml`) runs RuboCop and the specs on Ruby 4.0 for every push and pull request.
298
+ The HTML coverage report is attached to each run as the `coverage` artifact, and pushes to `master` refresh
299
+ the line and branch coverage badges on the `badges` branch.
300
+
301
+ See [ARCHITECTURE.md](ARCHITECTURE.md) for how the pieces fit together.
302
+
303
+ ## Releasing
304
+
305
+ 1. Bump `s.version` in `db-purger.gemspec` and merge to `master`.
306
+ 2. Tag and push: `git tag v0.6.0 && git push origin v0.6.0`
307
+
308
+ `.github/workflows/release.yml` re-runs CI, checks the tag matches the gemspec version, publishes to RubyGems
309
+ via trusted publishing (no API key), and creates a GitHub release with the `.gem` attached.
310
+
311
+ ## License
312
+
313
+ MIT — see [LICENSE.txt](LICENSE.txt).
@@ -0,0 +1,76 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::AssociationGraph answers "which tables reference this model, and by which column" from the
5
+ # has_many, has_one and has_and_belongs_to_many associations dynamic-active-model discovered
6
+ class AssociationGraph
7
+ # a table holding foreign_key that points at the parent model's primary key
8
+ Edge = Struct.new(:model, :foreign_key)
9
+
10
+ def initialize(database)
11
+ @database = database
12
+ @edges = {}
13
+ end
14
+
15
+ # database.models order depends on how the adapter lists tables, which varies by platform;
16
+ # sort so generated plans are deterministic
17
+ def models
18
+ @models ||= @database.models.sort_by(&:table_name)
19
+ end
20
+
21
+ def model_for(table_name)
22
+ models.detect { |model| model.table_name == table_name.to_s }
23
+ end
24
+
25
+ # one edge per (table, foreign key); a habtm join table also reached by a has_many appears once
26
+ def edges(model)
27
+ @edges[model] ||= model.reflect_on_all_associations
28
+ .filter_map { |reflection| edge_for(reflection) }
29
+ .uniq { |edge| edge_key(edge) }
30
+ .sort_by { |edge| edge_key(edge) }
31
+ end
32
+
33
+ # every model reachable from model through edges, excluding model itself unless there is a cycle
34
+ def reachable(model, seen = Set.new)
35
+ edges(model).each do |edge|
36
+ next if seen.include?(edge.model)
37
+
38
+ seen << edge.model
39
+ reachable(edge.model, seen)
40
+ end
41
+ seen
42
+ end
43
+
44
+ def column?(model, field)
45
+ model.column_names.include?(field.to_s)
46
+ end
47
+
48
+ private
49
+
50
+ def edge_for(reflection)
51
+ return if skip_reflection?(reflection)
52
+
53
+ case reflection
54
+ when ActiveRecord::Reflection::HasAndBelongsToManyReflection
55
+ join_table_edge(reflection)
56
+ when ActiveRecord::Reflection::HasManyReflection, ActiveRecord::Reflection::HasOneReflection
57
+ Edge.new(reflection.klass, reflection.foreign_key.to_s)
58
+ end
59
+ end
60
+
61
+ # through associations are reached via their own direct associations; polymorphic (as:) ones need a
62
+ # type condition the generator cannot infer
63
+ def skip_reflection?(reflection)
64
+ reflection.options[:through] || reflection.options[:as]
65
+ end
66
+
67
+ def join_table_edge(reflection)
68
+ join_model = model_for(reflection.join_table)
69
+ Edge.new(join_model, reflection.foreign_key.to_s) if join_model
70
+ end
71
+
72
+ def edge_key(edge)
73
+ [edge.model.table_name, edge.foreign_key]
74
+ end
75
+ end
76
+ end
@@ -5,16 +5,30 @@ module DBPurger
5
5
  class Config
6
6
  DEFAULT_DATETIME_FORMAT = '%Y-%m-%d %H:%M:%S'
7
7
 
8
- attr_writer :explain,
9
- :explain_file,
8
+ attr_writer :explain_file,
10
9
  :datetime_format
11
10
 
11
+ def initialize(options = {})
12
+ self.explain = options[:explain]
13
+ @explain_file = options[:explain_file]
14
+ @datetime_format = options[:datetime_format]
15
+ end
16
+
17
+ # Fail closed: a value like 'true' must not silently mean "run for real"
18
+ def explain=(value)
19
+ unless [true, false, nil].include?(value)
20
+ raise(ArgumentError, "explain must be true, false or nil, got #{value.inspect}")
21
+ end
22
+
23
+ @explain = value
24
+ end
25
+
12
26
  def explain?
13
27
  @explain == true
14
28
  end
15
29
 
16
30
  def explain_file
17
- (@explain_file || $stdout)
31
+ @explain_file || $stdout
18
32
  end
19
33
 
20
34
  def datetime_format
@@ -3,111 +3,124 @@
3
3
  module DBPurger
4
4
  # DBPurger::DynamicPlanBuilder generates a purge plan based on the database relations
5
5
  class DynamicPlanBuilder
6
- INDENT = ' '
7
-
8
- attr_reader :output
9
-
10
6
  def initialize(database)
11
- @database = database
12
- @output = ''.dup
13
- @indent_depth = 0
14
- @tables = []
7
+ @graph = AssociationGraph.new(database)
8
+ @writer = PlanWriter.new
9
+ end
10
+
11
+ def output
12
+ @writer.output
15
13
  end
16
14
 
17
- # rubocop:disable Metrics/AbcSize
15
+ # plan rooted at a single base table
18
16
  def build(base_table_name, field)
19
- write_table('base', base_table_name.to_s, field, [], nil)
20
- line_break
21
- model = find_model_for_table(base_table_name)
22
- foreign_key = foreign_key_name(model)
17
+ model = @graph.model_for(base_table_name)
18
+ @writer.table('base', model.table_name, field)
19
+ @writer.line_break
23
20
  if model.primary_key == field.to_s
24
- add_parent_tables(base_table_name, foreign_key)
21
+ add_referencing_parent_tables(model)
25
22
  else
26
- add_parent_tables(base_table_name, field)
27
- unless (child_models = find_child_models(model, foreign_key)).empty?
28
- line_break unless field == :id
29
- add_child_tables(child_models, foreign_key, 0)
30
- end
23
+ add_sibling_parent_tables(model, field)
24
+ add_base_child_tables(model)
31
25
  end
32
- ignore_missing_tables
33
- @output
26
+ finish
27
+ end
28
+
29
+ # plan with one top-level parent_table per table holding field (e.g. :oid), each purged by field
30
+ # directly so rows with a null foreign key are not missed; the tables referencing each root are nested
31
+ # under it, and roots are ordered so a root referencing another root's rows is purged first
32
+ def build_for(field)
33
+ @root_field = field.to_s
34
+ ordered_root_models.each_with_index do |model, idx|
35
+ @writer.line_break if idx.positive?
36
+ write_table('parent', model, field, [])
37
+ end
38
+ finish
34
39
  end
35
- # rubocop:enable Metrics/AbcSize
36
40
 
37
41
  private
38
42
 
39
- def find_model_for_table(base_table_name)
40
- @database.models.detect { |m| m.table_name == base_table_name.to_s }
43
+ def finish
44
+ @writer.ignore_tables(@graph.models.map(&:table_name) - @writer.table_names)
45
+ output
41
46
  end
42
47
 
43
- def write(str)
44
- @output << (INDENT * @indent_depth) + str + "\n"
48
+ # base_table(:companies, :id): tables holding companies.id are keyed directly on the purge value
49
+ def add_referencing_parent_tables(model)
50
+ @graph.edges(model).each { |edge| write_edge('parent', edge, [model]) }
45
51
  end
46
52
 
47
- def line_break
48
- @output << "\n"
53
+ # base_table(:employments, :company_id): other tables holding company_id share the purge value
54
+ def add_sibling_parent_tables(model, field)
55
+ @graph.models.each do |sibling|
56
+ next if sibling == model || !@graph.column?(sibling, field)
57
+
58
+ write_table('parent', sibling, field, [model])
59
+ end
49
60
  end
50
61
 
51
- def add_parent_tables(base_table_name, field)
52
- @database.models.each do |model|
53
- next if model.table_name == base_table_name.to_s
54
- next unless column?(model, field)
62
+ def add_base_child_tables(model)
63
+ edges = nestable_edges(model)
64
+ return if edges.empty?
55
65
 
56
- foreign_key = foreign_key_name(model)
57
- write_table('parent', model.table_name, field, find_child_models(model, foreign_key), foreign_key)
58
- end
66
+ @writer.line_break
67
+ edges.each { |edge| write_edge('child', edge, [model]) }
59
68
  end
60
69
 
61
- def add_child_tables(child_models, field, change_indent_by = 1)
62
- @indent_depth += change_indent_by
63
- child_models.each do |model|
64
- add_child_table(model, field)
70
+ def write_edge(table_type, edge, ancestors)
71
+ if ancestors.include?(edge.model)
72
+ @writer.comment("#{table_type}_table(#{edge.model.table_name.to_sym.inspect}, " \
73
+ "#{edge.foreign_key.to_sym.inspect}) skipped: cycle back to #{edge.model.table_name}")
74
+ else
75
+ write_table(table_type, edge.model, edge.foreign_key, ancestors)
65
76
  end
66
- @indent_depth -= change_indent_by
67
77
  end
68
78
 
69
- def add_child_table(model, field)
70
- foreign_key = foreign_key_name(model)
71
- write_table('child', model.table_name, field, find_child_models(model, foreign_key), foreign_key)
79
+ def write_table(table_type, model, field, ancestors)
80
+ edges = nestable_edges(model)
81
+ if edges.empty?
82
+ @writer.table(table_type, model.table_name, field)
83
+ warn_unnestable(model)
84
+ else
85
+ @writer.table_block(table_type, model.table_name, field) do
86
+ edges.each { |edge| write_edge('child', edge, ancestors + [model]) }
87
+ end
88
+ end
72
89
  end
73
90
 
74
- def find_child_models(model, field)
75
- model_has_many_associations(model).map(&:klass).select { |m| column?(m, field) }
76
- end
91
+ # purging nested tables needs this table's primary keys to propagate
92
+ def nestable_edges(model)
93
+ return [] unless model.primary_key
77
94
 
78
- def model_has_many_associations(model)
79
- model.reflect_on_all_associations.select do |assoc|
80
- assoc.is_a?(ActiveRecord::Reflection::HasManyReflection)
81
- end
95
+ @graph.edges(model).reject { |edge| root_model?(edge.model) }
82
96
  end
83
97
 
84
- def foreign_key_name(model)
85
- model.table_name.singularize + '_id'
98
+ def root_model?(model)
99
+ @root_field && @graph.column?(model, @root_field)
86
100
  end
87
101
 
88
- def column?(model, field)
89
- model.columns.detect { |c| c.name == field.to_s } != nil
102
+ def warn_unnestable(model)
103
+ return if model.primary_key || (edges = @graph.edges(model)).empty?
104
+
105
+ @writer.comment("#{model.table_name} has no primary key; cannot nest " \
106
+ "#{edges.map { |edge| edge.model.table_name }.join(', ')}")
90
107
  end
91
108
 
92
- def write_table(table_type, table_name, field, child_models, foreign_key)
93
- @tables << table_name
94
- if child_models.empty?
95
- write("#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect})")
96
- else
97
- write("#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect}) do")
98
- add_child_tables(child_models, foreign_key)
99
- write('end')
109
+ # repeatedly take the first root (by name) whose referencing roots are already written; on a cycle,
110
+ # fall back to name order so no root is dropped
111
+ def ordered_root_models
112
+ remaining = @graph.models.select { |model| root_model?(model) }
113
+ ordered = []
114
+ until remaining.empty?
115
+ model = remaining.detect { |root| (purged_first(root) & remaining).empty? } || remaining.first
116
+ ordered << remaining.delete(model)
100
117
  end
118
+ ordered
101
119
  end
102
120
 
103
- def ignore_missing_tables
104
- missing_tables = @database.models.map(&:table_name) - @tables
105
- return if missing_tables.empty?
106
-
107
- line_break
108
- missing_tables.each do |table_name|
109
- write("ignore_table #{table_name.to_sym.inspect}")
110
- end
121
+ # roots holding rows that reference root's rows (directly or through nested tables)
122
+ def purged_first(root)
123
+ @graph.reachable(root).select { |model| model != root && root_model?(model) }
111
124
  end
112
125
  end
113
126
  end
@@ -8,14 +8,14 @@ module DBPurger
8
8
  def initialize(database, plan, options = {})
9
9
  @database = database
10
10
  @plan = plan.is_a?(Plan) ? plan : load_plan(plan)
11
- setup_config(options)
11
+ @config = Config.new(options)
12
12
  @error_io = $stderr
13
13
  end
14
14
 
15
15
  def purge!(purge_value)
16
16
  raise('purge_value is nil') if purge_value.nil?
17
17
 
18
- @plan.purge!(@database, purge_value)
18
+ ::DBPurger.with_config(@config) { @plan.purge!(@database, purge_value) }
19
19
  end
20
20
 
21
21
  def verify!
@@ -31,12 +31,6 @@ module DBPurger
31
31
  @plan_validator ||= PlanValidator.new(@database, @plan)
32
32
  end
33
33
 
34
- def setup_config(options)
35
- ::DBPurger.config.explain = options[:explain]
36
- ::DBPurger.config.explain_file = options[:explain_file]
37
- ::DBPurger.config.datetime_format = options[:datetime_format]
38
- end
39
-
40
34
  def load_plan(file)
41
35
  PlanBuilder
42
36
  .new(Plan.new)
@@ -23,7 +23,7 @@ module DBPurger
23
23
  self.class.metrics.update_purge_stats(
24
24
  event.payload[:table_name],
25
25
  event.duration,
26
- event.payload[:deleted]
26
+ event.payload[:deleted] || 0
27
27
  )
28
28
  end
29
29
 
@@ -31,7 +31,7 @@ module DBPurger
31
31
  self.class.metrics.update_delete_records_stats(
32
32
  event.payload[:table_name],
33
33
  event.duration,
34
- event.payload[:records_deleted],
34
+ event.payload[:records_deleted] || 0,
35
35
  event.payload[:num_records]
36
36
  )
37
37
  end
@@ -40,7 +40,7 @@ module DBPurger
40
40
  self.class.metrics.update_lookup_stats(
41
41
  event.payload[:table_name],
42
42
  event.duration,
43
- event.payload[:num_records]
43
+ event.payload[:num_records] || 0
44
44
  )
45
45
  end
46
46
 
@@ -48,8 +48,8 @@ module DBPurger
48
48
  self.class.metrics.update_search_filter_stats(
49
49
  event.payload[:table_name],
50
50
  event.duration,
51
- event.payload[:num_records],
52
- event.payload[:num_records_selected]
51
+ event.payload[:num_records] || 0,
52
+ event.payload[:num_records_selected] || 0
53
53
  )
54
54
  end
55
55
  end
@@ -18,12 +18,21 @@ module DBPurger
18
18
  end
19
19
 
20
20
  def purge!(database, purge_value)
21
+ raise('plan has no base_table or top-level parent_table') if root_tables.empty?
22
+ raise('top-level child_tables require a base_table') unless @base_table || @child_tables.empty?
23
+
21
24
  MetricSubscriber.reset!
22
- num_deleted = PurgeTable.new(database, @base_table, @base_table.field, purge_value).purge!
25
+ num_deleted = purge_root_tables(database, purge_value)
26
+ purge_search_tables(database)
23
27
  MetricSubscriber.finished!
24
28
  num_deleted
25
29
  end
26
30
 
31
+ # tables that receive the purge value directly: the base_table (if any) and top-level parent_tables
32
+ def root_tables
33
+ (@base_table ? [@base_table] : []) + @parent_tables
34
+ end
35
+
27
36
  def tables
28
37
  all_tables = @base_table ? [@base_table] + @base_table.tables : []
29
38
  all_tables += @parent_tables + @parent_tables.map(&:tables) +
@@ -34,11 +43,9 @@ module DBPurger
34
43
  all_tables
35
44
  end
36
45
 
46
+ # the tables of a nested plan (nested plans never have a base_table)
37
47
  def foreign_tables
38
- (@base_table ? [@base_table] : []) +
39
- @parent_tables +
40
- @child_tables +
41
- @search_tables
48
+ @parent_tables + @child_tables + @search_tables
42
49
  end
43
50
 
44
51
  def table_names
@@ -61,5 +68,20 @@ module DBPurger
61
68
  end
62
69
  end
63
70
  end
71
+
72
+ private
73
+
74
+ def purge_root_tables(database, purge_value)
75
+ root_tables.sum do |table|
76
+ PurgeTable.new(database, table, table.field, purge_value).purge!
77
+ end
78
+ end
79
+
80
+ # with a base_table these live in its nested plan and are purged by it
81
+ def purge_search_tables(database)
82
+ @search_tables.each do |table|
83
+ PurgeTableScanner.new(database, table).purge!
84
+ end
85
+ end
64
86
  end
65
87
  end
@@ -7,6 +7,7 @@ module DBPurger
7
7
  class PlanValidator
8
8
  include ActiveModel::Validations
9
9
 
10
+ validate :validate_base_table
10
11
  validate :validate_no_missing_tables
11
12
  validate :validate_no_unknown_tables
12
13
  validate :validate_tables
@@ -28,6 +29,16 @@ module DBPurger
28
29
 
29
30
  private
30
31
 
32
+ # a plan is rooted either by a base_table or by one or more top-level parent_tables
33
+ def validate_base_table
34
+ if @plan.root_tables.empty?
35
+ errors.add(:base_table, 'or a top-level parent_table is required')
36
+ elsif !@plan.child_tables.empty?
37
+ # without a base_table there are no ids to propagate; declared before one, they are never reached
38
+ errors.add(:base_table, 'must be declared before top-level child_tables')
39
+ end
40
+ end
41
+
31
42
  def validate_no_missing_tables
32
43
  errors.add(:missing_tables, missing_tables.sort.join(',')) unless missing_tables.empty?
33
44
  end
@@ -51,6 +62,26 @@ module DBPurger
51
62
  errors.add(:table, "#{table.name}.#{field} is missing in the database")
52
63
  end
53
64
  end
65
+
66
+ validate_mark_deleted_field(table, model)
67
+ validate_batch_size(table)
68
+ validate_nested_tables_have_primary_key(table, model)
69
+ end
70
+
71
+ def validate_mark_deleted_field(table, model)
72
+ return if table.mark_deleted_field.nil? || model.column_names.include?(table.mark_deleted_field.to_s)
73
+
74
+ errors.add(:table, "#{table.name}.#{table.mark_deleted_field} (mark_deleted_field) is missing in the database")
75
+ end
76
+
77
+ def validate_batch_size(table)
78
+ errors.add(:table, "#{table.name} batch_size must be positive") unless table.batch_size.to_i.positive?
79
+ end
80
+
81
+ def validate_nested_tables_have_primary_key(table, model)
82
+ return if model.primary_key || !table.nested_key_tables?
83
+
84
+ errors.add(:table, "#{table.name} has no primary key and cannot have nested child or parent tables")
54
85
  end
55
86
 
56
87
  def find_model_for_table(table)
@@ -0,0 +1,56 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::PlanWriter renders plan DSL source and tracks which tables it has written
5
+ class PlanWriter
6
+ INDENT = ' '
7
+
8
+ attr_reader :output,
9
+ :table_names
10
+
11
+ def initialize
12
+ @output = ''.dup
13
+ @indent_depth = 0
14
+ @table_names = []
15
+ end
16
+
17
+ def table(table_type, table_name, field)
18
+ @table_names << table_name
19
+ write(table_call(table_type, table_name, field))
20
+ end
21
+
22
+ def table_block(table_type, table_name, field)
23
+ @table_names << table_name
24
+ write("#{table_call(table_type, table_name, field)} do")
25
+ @indent_depth += 1
26
+ yield
27
+ @indent_depth -= 1
28
+ write('end')
29
+ end
30
+
31
+ def comment(str)
32
+ write("# #{str}")
33
+ end
34
+
35
+ def ignore_tables(table_names)
36
+ return if table_names.empty?
37
+
38
+ line_break
39
+ table_names.each { |table_name| write("ignore_table #{table_name.to_sym.inspect}") }
40
+ end
41
+
42
+ def line_break
43
+ @output << "\n"
44
+ end
45
+
46
+ private
47
+
48
+ def write(str)
49
+ @output << "#{INDENT * @indent_depth}#{str}\n"
50
+ end
51
+
52
+ def table_call(table_type, table_name, field)
53
+ "#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect})"
54
+ end
55
+ end
56
+ end
@@ -13,10 +13,6 @@ module DBPurger
13
13
  @num_deleted = 0
14
14
  end
15
15
 
16
- def model
17
- @model ||= @database.models.detect { |m| m.table_name == @table.name.to_s }
18
- end
19
-
20
16
  def purge!
21
17
  ActiveSupport::Notifications.instrument('purge.db_purger',
22
18
  table_name: @table.name,
@@ -25,6 +21,7 @@ module DBPurger
25
21
  if model.primary_key
26
22
  purge_in_batches!
27
23
  else
24
+ ensure_no_nested_key_tables!
28
25
  purge_all!
29
26
  end
30
27
  purge_search_tables
@@ -35,6 +32,13 @@ module DBPurger
35
32
 
36
33
  private
37
34
 
35
+ # without a primary key there are no batch ids to propagate, so nested tables would be silently skipped
36
+ def ensure_no_nested_key_tables!
37
+ return unless @table.nested_key_tables?
38
+
39
+ raise("#{@table.name} has no primary key and cannot have nested child or parent tables")
40
+ end
41
+
38
42
  def purge_all!
39
43
  scope = model.where(@purge_field => @purge_value)
40
44
  scope = scope.where(@table.conditions) if @table.conditions
@@ -42,13 +46,27 @@ module DBPurger
42
46
  end
43
47
 
44
48
  def purge_in_batches!
49
+ unless @table.parent_tables?
50
+ each_batch do |batch|
51
+ purge_nested_tables(batch) if @table.nested_tables?
52
+ delete_records(batch)
53
+ end
54
+ return
55
+ end
56
+
57
+ # Parent tables may reference this table's rows and be referenced by its child tables,
58
+ # so purge them after the children but before this table's rows.
59
+ each_batch { |batch| purge_nested_tables(batch) }
60
+ purge_parent_tables
61
+ each_batch { |batch| delete_records(batch) }
62
+ end
63
+
64
+ def each_batch
45
65
  start_id = nil
46
66
  until (batch = next_batch(start_id)).empty?
47
- start_id = batch.last.send(model.primary_key)
48
- purge_nested_tables(batch) if @table.nested_tables?
49
- delete_records(batch)
67
+ start_id = batch.last[model.primary_key]
68
+ yield batch
50
69
  end
51
- purge_parent_tables
52
70
  end
53
71
 
54
72
  def next_batch(start_id)
@@ -38,8 +38,9 @@ module DBPurger
38
38
  end
39
39
  end
40
40
 
41
+ # record[] reads the column; send would call a same-named method instead (e.g. a column called "reload")
41
42
  def batch_values(batch, field)
42
- batch.map { |record| record.send(field) }.compact
43
+ batch.map { |record| record[field] }.compact
43
44
  end
44
45
 
45
46
  def foreign_tables?
@@ -72,7 +73,7 @@ module DBPurger
72
73
  else
73
74
  scope.to_sql.sub(/SELECT .*?FROM/, 'DELETE FROM')
74
75
  end
75
- ::DBPurger.config.explain_file.puts(sql + ';')
76
+ ::DBPurger.config.explain_file.puts("#{sql};")
76
77
  scope.count
77
78
  end
78
79
 
@@ -11,10 +11,6 @@ module DBPurger
11
11
  @num_deleted = 0
12
12
  end
13
13
 
14
- def model
15
- @model ||= @database.models.detect { |m| m.table_name == @table.name.to_s }
16
- end
17
-
18
14
  def purge!
19
15
  ActiveSupport::Notifications.instrument('purge.db_purger',
20
16
  table_name: @table.name) do |payload|
@@ -32,6 +32,15 @@ module DBPurger
32
32
  @nested_plan != nil
33
33
  end
34
34
 
35
+ def parent_tables?
36
+ nested_tables? && !@nested_plan.parent_tables.empty?
37
+ end
38
+
39
+ # nested tables that depend on this table's rows (search tables scan independently)
40
+ def nested_key_tables?
41
+ nested_tables? && !(@nested_plan.child_tables.empty? && @nested_plan.parent_tables.empty?)
42
+ end
43
+
35
44
  def tables
36
45
  @nested_plan ? @nested_plan.tables : []
37
46
  end
@@ -45,7 +54,7 @@ module DBPurger
45
54
  end
46
55
 
47
56
  def mark_deleted_value
48
- @mark_deleted_value || 1
57
+ @mark_deleted_value.nil? ? 1 : @mark_deleted_value
49
58
  end
50
59
  end
51
60
  end
data/lib/db-purger.rb CHANGED
@@ -2,6 +2,7 @@
2
2
 
3
3
  # DBPurger is a tool to delete data from tables based on a initial purge value
4
4
  module DBPurger
5
+ autoload :AssociationGraph, 'db-purger/association_graph'
5
6
  autoload :Config, 'db-purger/config'
6
7
  autoload :DynamicPlanBuilder, 'db-purger/dynamic_plan_builder'
7
8
  autoload :Executor, 'db-purger/executor'
@@ -13,9 +14,25 @@ module DBPurger
13
14
  autoload :Plan, 'db-purger/plan'
14
15
  autoload :PlanBuilder, 'db-purger/plan_builder'
15
16
  autoload :PlanValidator, 'db-purger/plan_validator'
17
+ autoload :PlanWriter, 'db-purger/plan_writer'
16
18
  autoload :Table, 'db-purger/table'
17
19
 
20
+ # The config in effect for the current thread: the one set by with_config, else the global default
18
21
  def self.config
19
- @config ||= Config.new
22
+ Thread.current[:db_purger_config] || default_config
23
+ end
24
+
25
+ def self.default_config
26
+ @default_config ||= Config.new
27
+ end
28
+
29
+ # Runs the block with config in effect for this thread only, so concurrent or
30
+ # later executors can't change explain mode out from under a running purge
31
+ def self.with_config(config)
32
+ previous = Thread.current[:db_purger_config]
33
+ Thread.current[:db_purger_config] = config
34
+ yield
35
+ ensure
36
+ Thread.current[:db_purger_config] = previous
20
37
  end
21
38
  end
metadata CHANGED
@@ -1,49 +1,63 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: db-purger
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.4.1
4
+ version: 0.6.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Doug Youch
8
8
  bindir: bin
9
9
  cert_chain: []
10
- date: 2025-07-31 00:00:00.000000000 Z
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
11
  dependencies:
12
12
  - !ruby/object:Gem::Dependency
13
- name: dynamic-active-model
13
+ name: activerecord
14
14
  requirement: !ruby/object:Gem::Requirement
15
15
  requirements:
16
16
  - - ">="
17
17
  - !ruby/object:Gem::Version
18
- version: '0'
18
+ version: '7.0'
19
19
  type: :runtime
20
20
  prerelease: false
21
21
  version_requirements: !ruby/object:Gem::Requirement
22
22
  requirements:
23
23
  - - ">="
24
24
  - !ruby/object:Gem::Version
25
- version: '0'
25
+ version: '7.0'
26
26
  - !ruby/object:Gem::Dependency
27
- name: inheritance-helper
27
+ name: dynamic-active-model
28
28
  requirement: !ruby/object:Gem::Requirement
29
29
  requirements:
30
+ - - "~>"
31
+ - !ruby/object:Gem::Version
32
+ version: '0.9'
30
33
  - - ">="
31
34
  - !ruby/object:Gem::Version
32
- version: '0'
35
+ version: 0.9.1
33
36
  type: :runtime
34
37
  prerelease: false
35
38
  version_requirements: !ruby/object:Gem::Requirement
36
39
  requirements:
40
+ - - "~>"
41
+ - !ruby/object:Gem::Version
42
+ version: '0.9'
37
43
  - - ">="
38
44
  - !ruby/object:Gem::Version
39
- version: '0'
40
- description: Purge database tables by top level id in batches
45
+ version: 0.9.1
46
+ description: DB Purger deletes (or soft-deletes) every row related to a single top-level
47
+ record (e.g. a company or account) using a declarative Ruby purge plan. Tables are
48
+ purged in primary-key batches, child tables before their parents, with plan validation
49
+ against the live schema, an explain (dry-run) mode that emits SQL, and ActiveSupport::Notifications
50
+ instrumentation for metrics.
41
51
  email: dougyouch@gmail.com
42
52
  executables: []
43
53
  extensions: []
44
54
  extra_rdoc_files: []
45
55
  files:
56
+ - ARCHITECTURE.md
57
+ - LICENSE.txt
58
+ - README.md
46
59
  - lib/db-purger.rb
60
+ - lib/db-purger/association_graph.rb
47
61
  - lib/db-purger/config.rb
48
62
  - lib/db-purger/dynamic_plan_builder.rb
49
63
  - lib/db-purger/executor.rb
@@ -52,6 +66,7 @@ files:
52
66
  - lib/db-purger/plan.rb
53
67
  - lib/db-purger/plan_builder.rb
54
68
  - lib/db-purger/plan_validator.rb
69
+ - lib/db-purger/plan_writer.rb
55
70
  - lib/db-purger/purge_table.rb
56
71
  - lib/db-purger/purge_table_helper.rb
57
72
  - lib/db-purger/purge_table_scanner.rb
@@ -59,7 +74,10 @@ files:
59
74
  homepage: https://github.com/dougyouch/db-purger
60
75
  licenses:
61
76
  - MIT
62
- metadata: {}
77
+ metadata:
78
+ source_code_uri: https://github.com/dougyouch/db-purger
79
+ bug_tracker_uri: https://github.com/dougyouch/db-purger/issues
80
+ rubygems_mfa_required: 'true'
63
81
  rdoc_options: []
64
82
  require_paths:
65
83
  - lib
@@ -67,14 +85,14 @@ required_ruby_version: !ruby/object:Gem::Requirement
67
85
  requirements:
68
86
  - - ">="
69
87
  - !ruby/object:Gem::Version
70
- version: '0'
88
+ version: '3.2'
71
89
  required_rubygems_version: !ruby/object:Gem::Requirement
72
90
  requirements:
73
91
  - - ">="
74
92
  - !ruby/object:Gem::Version
75
93
  version: '0'
76
94
  requirements: []
77
- rubygems_version: 3.6.2
95
+ rubygems_version: 4.0.20
78
96
  specification_version: 4
79
- summary: DB Purger by top level id
97
+ summary: Purge all data tied to a top-level id across related tables, in batches
80
98
  test_files: []