db-purger 0.5.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 863444350ef64942e0669feacdf4a119514c16369a31eeb99e31a5f5435028ef
4
- data.tar.gz: 66b063187fdaa33232a6fffe76f9496264e8bd74e2a835013139b4e19f3e045e
3
+ metadata.gz: 45ab4b4aa9e6e8c3778061c5deb8a8eb37fecf153bb1bae9cf298f2c83433f22
4
+ data.tar.gz: bbb9aec2aaa9dcd8e7c0609186795bf17e8fce64ffa951ce6dd1877981b7dbcc
5
5
  SHA512:
6
- metadata.gz: dfd50964ce367548606fe758df1d1cfb48d6891d9ac09993079b851a78658807fa4a5ff8633709a49f42b9d67c9367260721cd2d0f1454dbd7da197f3363f54b
7
- data.tar.gz: ccc2bee89229fad63fd2dc7d20682b07b008b29040cb66a92aa1cd4c48bd8f09c5f6b91834ead524be8f3394d54c8af08ce8df5d61bc5067a6229b14187c7ede
6
+ metadata.gz: 41d99640be3f800277142f9ef9a9ec0f0945785282551c757f92ba92de36dfad331b735533713bb98385023378727ca5365cf995cff84580dfe725cbbf10240d
7
+ data.tar.gz: bfe3512b7f289cec1b4ac7f466918c4aaa6a5c12f9deb61e60e825153dd3e1badbbcbb1eb578e80dd453504f9b7b08f70a996e4d9d10b7d1e7a654245d12ec00
data/ARCHITECTURE.md CHANGED
@@ -31,15 +31,18 @@ db-purger is small (~900 lines) and splits cleanly into three layers: **describe
31
31
  | `lib/db-purger.rb` | Autoloads everything; holds the global `DBPurger.config`. |
32
32
  | `config.rb` | Global options: `explain?`, `explain_file`, `datetime_format`. |
33
33
  | `table.rb` | Value object for one table in the plan: name, match field, options, and a lazily created nested `Plan`. |
34
- | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, search and ignored tables. `#purge!` is the run entry point. |
34
+ | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, nullify, search and ignored tables. `#purge!` is the run entry point; `#root_tables` are the tables it starts from. |
35
35
  | `plan_builder.rb` | The DSL. `instance_eval`s a plan file or block against a `Plan`; nested blocks get a new builder bound to that table's nested plan. |
36
36
  | `plan_validator.rb` | `ActiveModel::Validations` over plan vs. schema: missing tables, unknown tables, unknown columns. |
37
37
  | `executor.rb` | Convenience façade: loads a plan file, applies config options, `verify!`, `purge!`. |
38
38
  | `purge_table.rb` | Purges one table by `field = value(s)` in primary-key batches. Recurses into nested tables. |
39
39
  | `purge_table_scanner.rb` | Purges a `purge_table_search` table: full `find_in_batches` scan filtered through the user's `search_proc`. |
40
+ | `nullify_table.rb` | Sets a nullable column to `NULL` on rows pointing at a batch of ids about to be deleted (`ON DELETE SET NULL` at purge time). |
40
41
  | `purge_table_helper.rb` | Shared behaviour for both purgers: nested-table recursion, delete vs. soft delete vs. explain, transactions. |
41
42
  | `metrics.rb` / `metric_subscriber.rb` | Aggregate timing and row counts per table from the notification events. |
42
- | `dynamic_plan_builder.rb` | Generates plan-file source from `has_many` associations; a bootstrap tool, not used at purge time. |
43
+ | `dynamic_plan_builder.rb` | Generates plan-file source (`build` for a base table, `build_for` for several roots); a bootstrap tool, not used at purge time. |
44
+ | `association_graph.rb` | For the generator: which tables reference a model, and by which column, from its `has_many`/`has_one`/HABTM reflections. |
45
+ | `plan_writer.rb` | For the generator: renders plan DSL text and records which tables it wrote. |
43
46
 
44
47
  ## The plan tree
45
48
 
@@ -52,16 +55,30 @@ Plan (root)
52
55
  └── nested Plan
53
56
  ├── parent_tables: [company_tags(:company_id)]
54
57
  ├── child_tables: [employments(:company_id) ─▶ nested Plan ..., websites(:id, fk: website_id) ...]
58
+ ├── nullify_tables: [companies(:acquired_by_company_id)]
55
59
  ├── search_tables: [users(:id)]
56
60
  └── ignore_tables
57
61
  ```
58
62
 
63
+ Without a `base_table`, top-level `parent_table`s stay in the root plan's `parent_tables` and each one is a
64
+ root:
65
+
66
+ ```
67
+ Plan (root)
68
+ ├── parent_tables: [calls(:oid) ─▶ nested Plan ..., emails(:oid) ─▶ nested Plan ..., sms_messages(:oid) ...]
69
+ ├── search_tables
70
+ └── ignore_tables
71
+ ```
72
+
59
73
  `Plan#tables` flattens this tree for validation; `Table#foreign_keys` collects the `foreign_key:` columns of
60
74
  a table's direct children so the purger can `SELECT` them alongside the primary key.
61
75
 
62
76
  ## Purge algorithm
63
77
 
64
- `Plan#purge!` resets metrics and starts a `PurgeTable` on the base table with the purge value. Each
78
+ `Plan#purge!` resets metrics and starts a `PurgeTable` with the purge value on each root table
79
+ (`root_tables` = the `base_table`, if any, followed by top-level `parent_tables`, in declaration order), then runs
80
+ any top-level search tables. With a `base_table`, `root_tables` is just `[base_table]`, which is the original
81
+ single-root algorithm. Each
65
82
  `PurgeTable#purge!` does:
66
83
 
67
84
  ```
@@ -72,6 +89,8 @@ each_batch = loop:
72
89
  break if batch empty
73
90
 
74
91
  purge_children(batch) =
92
+ for each nullify_table: # optional rows pointing at us
93
+ UPDATE nullify SET field = NULL WHERE field IN (batch.pks) [AND conditions]
75
94
  for each child_table without foreign_key: # rows pointing at us
76
95
  PurgeTable(child, child.field, batch.pks).purge! (recursive)
77
96
 
@@ -103,6 +122,9 @@ Key properties:
103
122
 
104
123
  - **Depth-first, children first.** A row is only deleted after everything referencing it, so FK
105
124
  constraints hold without `ON DELETE CASCADE`.
125
+ - **Unlink before delete.** A `nullify_table` is updated for each batch before that batch's children and rows
126
+ are deleted, so a self-referential foreign key (or a reference from a row that must survive) never blocks the
127
+ delete, whatever the chain depth or id order across batches.
106
128
  - **Two passes when there are parent tables.** A parent table can reference this table (e.g.
107
129
  `company_tags.company_id → companies.id`) *and* be referenced by one of its children, so it is purged
108
130
  between the child pass and the delete pass. Tables without parent tables keep the single pass.
@@ -131,7 +153,7 @@ All three go through ActiveRecord's `*_all` methods: no model callbacks or valid
131
153
  ## Instrumentation
132
154
 
133
155
  Every unit of work is wrapped in `ActiveSupport::Notifications.instrument` under the `db_purger` namespace
134
- (`purge`, `next_batch`, `delete_records`, `search_filter`). The purgers never talk to `Metrics` directly;
156
+ (`purge`, `next_batch`, `delete_records`, `nullify_records`, `search_filter`). The purgers never talk to `Metrics` directly;
135
157
  `MetricSubscriber` (an `ActiveSupport::Subscriber`) translates events into `Metrics` counters. This keeps the
136
158
  purge code free of reporting concerns and lets callers attach their own subscribers (StatsD, logs, progress
137
159
  bars) without changes to the library.
@@ -157,4 +179,8 @@ fetch happens inside `find_in_batches` rather than in a block the scanner contro
157
179
  - `spec/integrations/*` — end-to-end purges over the schema in `spec/support/db/schema.rb`, asserting row-count
158
180
  deltas per table and, for explain mode, the exact SQL in `spec/fixtures/delete_plan.sql`.
159
181
  - `spec/support/test_db.rb` builds a fresh SQLite database and dynamic models (`TestDB::*`) for each run.
182
+ - `spec/integrations/multi_root_plan_spec.rb` purges one org from a separate outreach schema (`spec/support/outreach_db.rb`,
183
+ real FOREIGN KEY constraints) seeded with two orgs by `OutreachSeeder`, and asserts the exact surviving rows of
184
+ every table for the generated plan, a hand-written small-batch plan and the equivalent `base_table` plan.
185
+ - `spec/support/throwaway_db.rb` builds standalone SQLite databases from raw SQL for schema-specific specs.
160
186
  - `spec/fixtures/*.plan.rb` — plan files used for loading and validation cases.
data/README.md CHANGED
@@ -2,6 +2,7 @@
2
2
 
3
3
  [![CI](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml/badge.svg?branch=master)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
4
4
  [![Coverage](https://raw.githubusercontent.com/dougyouch/db-purger/badges/coverage.svg)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
5
+ [![Branch coverage](https://raw.githubusercontent.com/dougyouch/db-purger/badges/branch-coverage.svg)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
5
6
  [![Gem Version](https://img.shields.io/gem/v/db-purger)](https://rubygems.org/gems/db-purger)
6
7
 
7
8
  Purge every row tied to a single top-level record — a company, an account, a tenant — across all of the
@@ -112,16 +113,46 @@ executor.verify! # raises 'purge plan failed verification', errors pr
112
113
  deleted = executor.purge!(42)
113
114
  ```
114
115
 
115
- `purge!` returns the number of base-table rows deleted.
116
+ `purge!` returns the number of root-table rows deleted (the base table, plus any top-level `parent_table`s).
117
+
118
+ ### Plans with several top-level tables
119
+
120
+ `base_table` is shorthand for "one root table, with everything after it nested underneath". When several tables
121
+ are equally top-level (an outreach product's `emails`, `sms_messages` and `calls`, all keyed by `oid`), leave
122
+ `base_table` out and declare each root as a top-level `parent_table`:
123
+
124
+ ```ruby
125
+ # config/outreach.plan.rb
126
+ parent_table(:calls, :oid) do # calls.email_id -> emails.id, so calls go first
127
+ child_table(:call_notes, :call_id)
128
+ child_table(:call_recordings, :call_id)
129
+ child_table(:call_tags, :call_id)
130
+ end
131
+
132
+ parent_table(:emails, :oid) do
133
+ child_table(:email_attachments, :email_id)
134
+ end
135
+
136
+ parent_table(:sms_messages, :oid) do
137
+ child_table(:sms_deliveries, :sms_message_id)
138
+ end
139
+
140
+ ignore_table :users
141
+ ```
142
+
143
+ Each root is purged by `oid = purge_value`, children first, **in declaration order**: when one root's rows
144
+ reference another's, declare the referencing root first. Without a `base_table`, top-level `child_table`s are
145
+ an error (there is no enclosing batch to take ids from). Existing `base_table` plans run exactly as before.
116
146
 
117
147
  ## The plan DSL
118
148
 
119
149
  | Method | Meaning |
120
150
  |---|---|
121
- | `base_table(table, field, opts = {}, &block)` | The root of the purge. Rows where `field = purge_value` are purged. Declare it **first** — every subsequent top-level call nests under it. |
151
+ | `base_table(table, field, opts = {}, &block)` | Optional single root. Rows where `field = purge_value` are purged. Declare it **first** — every subsequent top-level call nests under it. |
122
152
  | `child_table(table, field, opts = {}, &block)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch. Purged before that batch is deleted. |
123
153
  | `child_table(table, :id, foreign_key: :col, &block)` | Inverted relationship: the *enclosing* table holds `col` pointing at this table's `id`. Deleted in the same transaction, right after the enclosing batch. |
124
- | `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
154
+ | `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. At the top level of a plan without a `base_table`, each one is a root, purged in declaration order. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
155
+ | `nullify_table(table, field, conditions: nil)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch get `field = NULL` instead of being deleted, before anything in that batch is deleted (`ON DELETE SET NULL` at purge time). Use for optional references that must not take the referencing row down with them: a self-referential `parent_id`/"copied from" column, or a link from another tenant's row. Only `conditions:` is supported. |
125
156
  | `purge_table_search(table, field, opts = {}) { \|batch\| ... }` | Scans the whole table in batches; the block receives each batch and returns the records to purge. For orphans that can't be reached by a key. |
126
157
  | `ignore_table(name_or_regexp)` | Exclude a table from validation. |
127
158
 
@@ -166,14 +197,25 @@ plan.purge!(database, 42)
166
197
 
167
198
  ### Generating a starting plan
168
199
 
169
- `DynamicPlanBuilder` walks the `has_many` associations dynamic-active-model discovered and emits a plan file,
170
- listing every unreachable table as `ignore_table`. Treat the output as a first draft: it only knows about
171
- conventional `<singular_table>_id` foreign keys and cannot infer polymorphic, soft-delete, or search rules.
200
+ `DynamicPlanBuilder` walks the `has_many`, `has_one` and `has_and_belongs_to_many` associations
201
+ dynamic-active-model discovered, using each association's real foreign key, and emits a plan file listing every
202
+ unreachable table as `ignore_table`.
172
203
 
173
204
  ```ruby
174
- puts DBPurger::DynamicPlanBuilder.new(database).build(:companies, :id)
205
+ builder = DBPurger::DynamicPlanBuilder.new(database)
206
+ puts builder.build(:companies, :id) # single base_table plan
207
+ puts builder.build_for(:oid) # one top-level parent_table per table holding oid
175
208
  ```
176
209
 
210
+ `build_for` makes **every** table holding the field a root, so rows with a null foreign key to another root
211
+ (an `email_recipients` row without an email) are still purged, and orders the roots so a root referencing
212
+ another root's rows comes first. HABTM join tables are emitted as leaves, never walking into the shared table on
213
+ the other side, and nothing is nested under a table without a primary key. A foreign-key cycle is written as a
214
+ comment instead of recursing.
215
+
216
+ Treat the output as a first draft: it cannot infer polymorphic (`as:`), soft-delete, `belongs_to`-owned
217
+ (`foreign_key:`) or search rules.
218
+
177
219
  ## Validation
178
220
 
179
221
  `Executor#verify!` (or `DBPurger::PlanValidator.new(database, plan).valid?`) checks that:
@@ -181,8 +223,11 @@ puts DBPurger::DynamicPlanBuilder.new(database).build(:companies, :id)
181
223
  - every table in the database is either in the plan or ignored (`missing_tables`)
182
224
  - every table in the plan exists in the database (`unknown_tables`)
183
225
  - every field, `foreign_key` and `mark_deleted_field` named in the plan is a real column
184
- - the plan has a `base_table`, and every `batch_size` is positive
185
- - tables without a primary key have no nested child or parent tables (there would be no ids to propagate)
226
+ - the plan has a `base_table` or at least one top-level `parent_table`, no top-level `child_table` is left
227
+ unreachable, and every `batch_size` is positive
228
+ - tables without a primary key have no nested child, parent or nullify tables (there would be no ids to propagate)
229
+ - every `nullify_table` field is a nullable column, and no `nullify_table` is left at the top level without a
230
+ `base_table`
186
231
 
187
232
  Run it in CI against your schema so a new table can't ship without a purge decision.
188
233
 
@@ -220,6 +265,7 @@ DBPurger::MetricSubscriber.metrics.as_json
220
265
  # => { took: 12.4, started_at: ..., finished_at: ...,
221
266
  # purge_stats: { employments: { duration:, num_purges:, num_records: } },
222
267
  # delete_stats: { employments: { duration:, num_delete_queries:, num_deleted:, num_expected_to_delete: } },
268
+ # nullify_stats: { cadences: { duration:, num_nullify_queries:, num_nullified: } },
223
269
  # lookup_stats: { ... }, filter_stats: { ... } }
224
270
  ```
225
271
 
@@ -231,12 +277,13 @@ Metrics are reset at the start of each `Plan#purge!`. To feed your own telemetry
231
277
  | `purge.db_purger` | `table_name`, `purge_field`, `deleted` |
232
278
  | `next_batch.db_purger` | `table_name`, `start_id`, `num_records` |
233
279
  | `delete_records.db_purger` | `table_name`, `num_records`, `records_deleted`, `deleted` |
280
+ | `nullify_records.db_purger` | `table_name`, `nullify_field`, `num_records`, `records_nullified` |
234
281
  | `search_filter.db_purger` | `table_name`, `num_records`, `num_records_selected` |
235
282
 
236
283
  ## Caveats
237
284
 
238
- - **Only the base table is the entry point.** Top-level `parent_table`/`child_table` calls made before
239
- `base_table` are ignored by `Plan#purge!`.
285
+ - **Declare `base_table` first.** A `child_table` declared before it is never reached (the validator reports
286
+ it); a `parent_table` declared before it becomes a separate root, purged after the base table.
240
287
  - **Not one big transaction.** Each batch is its own set of statements (foreign-key children share a
241
288
  transaction with their parent batch). An interrupted purge is safe to re-run with the same value.
242
289
  - **Soft-deleted rows still match.** A `mark_deleted_field` table is not filtered on that field; add
@@ -254,14 +301,14 @@ script/console
254
301
 
255
302
  CI (`.github/workflows/ci.yml`) runs RuboCop and the specs on Ruby 4.0 for every push and pull request.
256
303
  The HTML coverage report is attached to each run as the `coverage` artifact, and pushes to `master` refresh
257
- the coverage badge on the `badges` branch.
304
+ the line and branch coverage badges on the `badges` branch.
258
305
 
259
306
  See [ARCHITECTURE.md](ARCHITECTURE.md) for how the pieces fit together.
260
307
 
261
308
  ## Releasing
262
309
 
263
310
  1. Bump `s.version` in `db-purger.gemspec` and merge to `master`.
264
- 2. Tag and push: `git tag v0.5.0 && git push origin v0.5.0`
311
+ 2. Tag and push: `git tag v0.7.0 && git push origin v0.7.0`
265
312
 
266
313
  `.github/workflows/release.yml` re-runs CI, checks the tag matches the gemspec version, publishes to RubyGems
267
314
  via trusted publishing (no API key), and creates a GitHub release with the `.gem` attached.
@@ -0,0 +1,76 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::AssociationGraph answers "which tables reference this model, and by which column" from the
5
+ # has_many, has_one and has_and_belongs_to_many associations dynamic-active-model discovered
6
+ class AssociationGraph
7
+ # a table holding foreign_key that points at the parent model's primary key
8
+ Edge = Struct.new(:model, :foreign_key)
9
+
10
+ def initialize(database)
11
+ @database = database
12
+ @edges = {}
13
+ end
14
+
15
+ # database.models order depends on how the adapter lists tables, which varies by platform;
16
+ # sort so generated plans are deterministic
17
+ def models
18
+ @models ||= @database.models.sort_by(&:table_name)
19
+ end
20
+
21
+ def model_for(table_name)
22
+ models.detect { |model| model.table_name == table_name.to_s }
23
+ end
24
+
25
+ # one edge per (table, foreign key); a habtm join table also reached by a has_many appears once
26
+ def edges(model)
27
+ @edges[model] ||= model.reflect_on_all_associations
28
+ .filter_map { |reflection| edge_for(reflection) }
29
+ .uniq { |edge| edge_key(edge) }
30
+ .sort_by { |edge| edge_key(edge) }
31
+ end
32
+
33
+ # every model reachable from model through edges, excluding model itself unless there is a cycle
34
+ def reachable(model, seen = Set.new)
35
+ edges(model).each do |edge|
36
+ next if seen.include?(edge.model)
37
+
38
+ seen << edge.model
39
+ reachable(edge.model, seen)
40
+ end
41
+ seen
42
+ end
43
+
44
+ def column?(model, field)
45
+ model.column_names.include?(field.to_s)
46
+ end
47
+
48
+ private
49
+
50
+ def edge_for(reflection)
51
+ return if skip_reflection?(reflection)
52
+
53
+ case reflection
54
+ when ActiveRecord::Reflection::HasAndBelongsToManyReflection
55
+ join_table_edge(reflection)
56
+ when ActiveRecord::Reflection::HasManyReflection, ActiveRecord::Reflection::HasOneReflection
57
+ Edge.new(reflection.klass, reflection.foreign_key.to_s)
58
+ end
59
+ end
60
+
61
+ # through associations are reached via their own direct associations; polymorphic (as:) ones need a
62
+ # type condition the generator cannot infer
63
+ def skip_reflection?(reflection)
64
+ reflection.options[:through] || reflection.options[:as]
65
+ end
66
+
67
+ def join_table_edge(reflection)
68
+ join_model = model_for(reflection.join_table)
69
+ Edge.new(join_model, reflection.foreign_key.to_s) if join_model
70
+ end
71
+
72
+ def edge_key(edge)
73
+ [edge.model.table_name, edge.foreign_key]
74
+ end
75
+ end
76
+ end
@@ -3,115 +3,124 @@
3
3
  module DBPurger
4
4
  # DBPurger::DynamicPlanBuilder generates a purge plan based on the database relations
5
5
  class DynamicPlanBuilder
6
- INDENT = ' '
7
-
8
- attr_reader :output
9
-
10
6
  def initialize(database)
11
- @database = database
12
- @output = ''.dup
13
- @indent_depth = 0
14
- @tables = []
7
+ @graph = AssociationGraph.new(database)
8
+ @writer = PlanWriter.new
9
+ end
10
+
11
+ def output
12
+ @writer.output
15
13
  end
16
14
 
15
+ # plan rooted at a single base table
17
16
  def build(base_table_name, field)
18
- write_table('base', base_table_name.to_s, field, [], nil)
19
- line_break
20
- model = find_model_for_table(base_table_name)
21
- foreign_key = foreign_key_name(model)
17
+ model = @graph.model_for(base_table_name)
18
+ @writer.table('base', model.table_name, field)
19
+ @writer.line_break
22
20
  if model.primary_key == field.to_s
23
- add_parent_tables(base_table_name, foreign_key)
21
+ add_referencing_parent_tables(model)
24
22
  else
25
- add_parent_tables(base_table_name, field)
26
- unless (child_models = find_child_models(model, foreign_key)).empty?
27
- line_break unless field == :id
28
- add_child_tables(child_models, foreign_key, 0)
29
- end
23
+ add_sibling_parent_tables(model, field)
24
+ add_base_child_tables(model)
30
25
  end
31
- ignore_missing_tables
32
- @output
26
+ finish
27
+ end
28
+
29
+ # plan with one top-level parent_table per table holding field (e.g. :oid), each purged by field
30
+ # directly so rows with a null foreign key are not missed; the tables referencing each root are nested
31
+ # under it, and roots are ordered so a root referencing another root's rows is purged first
32
+ def build_for(field)
33
+ @root_field = field.to_s
34
+ ordered_root_models.each_with_index do |model, idx|
35
+ @writer.line_break if idx.positive?
36
+ write_table('parent', model, field, [])
37
+ end
38
+ finish
33
39
  end
34
40
 
35
41
  private
36
42
 
37
- def find_model_for_table(base_table_name)
38
- @database.models.detect { |m| m.table_name == base_table_name.to_s }
43
+ def finish
44
+ @writer.ignore_tables(@graph.models.map(&:table_name) - @writer.table_names)
45
+ output
39
46
  end
40
47
 
41
- def write(str)
42
- @output << "#{INDENT * @indent_depth}#{str}\n"
48
+ # base_table(:companies, :id): tables holding companies.id are keyed directly on the purge value
49
+ def add_referencing_parent_tables(model)
50
+ @graph.edges(model).each { |edge| write_edge('parent', edge, [model]) }
43
51
  end
44
52
 
45
- def line_break
46
- @output << "\n"
53
+ # base_table(:employments, :company_id): other tables holding company_id share the purge value
54
+ def add_sibling_parent_tables(model, field)
55
+ @graph.models.each do |sibling|
56
+ next if sibling == model || !@graph.column?(sibling, field)
57
+
58
+ write_table('parent', sibling, field, [model])
59
+ end
47
60
  end
48
61
 
49
- def add_parent_tables(base_table_name, field)
50
- sorted_models.each do |model|
51
- next if model.table_name == base_table_name.to_s
52
- next unless column?(model, field)
62
+ def add_base_child_tables(model)
63
+ edges = nestable_edges(model)
64
+ return if edges.empty?
53
65
 
54
- foreign_key = foreign_key_name(model)
55
- write_table('parent', model.table_name, field, find_child_models(model, foreign_key), foreign_key)
56
- end
66
+ @writer.line_break
67
+ edges.each { |edge| write_edge('child', edge, [model]) }
57
68
  end
58
69
 
59
- def add_child_tables(child_models, field, change_indent_by = 1)
60
- @indent_depth += change_indent_by
61
- child_models.each do |model|
62
- add_child_table(model, field)
70
+ def write_edge(table_type, edge, ancestors)
71
+ if ancestors.include?(edge.model)
72
+ @writer.comment("#{table_type}_table(#{edge.model.table_name.to_sym.inspect}, " \
73
+ "#{edge.foreign_key.to_sym.inspect}) skipped: cycle back to #{edge.model.table_name}")
74
+ else
75
+ write_table(table_type, edge.model, edge.foreign_key, ancestors)
63
76
  end
64
- @indent_depth -= change_indent_by
65
77
  end
66
78
 
67
- def add_child_table(model, field)
68
- foreign_key = foreign_key_name(model)
69
- write_table('child', model.table_name, field, find_child_models(model, foreign_key), foreign_key)
79
+ def write_table(table_type, model, field, ancestors)
80
+ edges = nestable_edges(model)
81
+ if edges.empty?
82
+ @writer.table(table_type, model.table_name, field)
83
+ warn_unnestable(model)
84
+ else
85
+ @writer.table_block(table_type, model.table_name, field) do
86
+ edges.each { |edge| write_edge('child', edge, ancestors + [model]) }
87
+ end
88
+ end
70
89
  end
71
90
 
72
- def find_child_models(model, field)
73
- model_has_many_associations(model).map(&:klass).select { |m| column?(m, field) }.sort_by(&:table_name)
74
- end
91
+ # purging nested tables needs this table's primary keys to propagate
92
+ def nestable_edges(model)
93
+ return [] unless model.primary_key
75
94
 
76
- # database.models order depends on how the adapter lists tables, which varies by platform;
77
- # sort so the generated plan is deterministic
78
- def sorted_models
79
- @sorted_models ||= @database.models.sort_by(&:table_name)
95
+ @graph.edges(model).reject { |edge| root_model?(edge.model) }
80
96
  end
81
97
 
82
- def model_has_many_associations(model)
83
- model.reflect_on_all_associations.select do |assoc|
84
- assoc.is_a?(ActiveRecord::Reflection::HasManyReflection)
85
- end
98
+ def root_model?(model)
99
+ @root_field && @graph.column?(model, @root_field)
86
100
  end
87
101
 
88
- def foreign_key_name(model)
89
- "#{model.table_name.singularize}_id"
90
- end
102
+ def warn_unnestable(model)
103
+ return if model.primary_key || (edges = @graph.edges(model)).empty?
91
104
 
92
- def column?(model, field)
93
- model.columns.detect { |c| c.name == field.to_s } != nil
105
+ @writer.comment("#{model.table_name} has no primary key; cannot nest " \
106
+ "#{edges.map { |edge| edge.model.table_name }.join(', ')}")
94
107
  end
95
108
 
96
- def write_table(table_type, table_name, field, child_models, foreign_key)
97
- @tables << table_name
98
- if child_models.empty?
99
- write("#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect})")
100
- else
101
- write("#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect}) do")
102
- add_child_tables(child_models, foreign_key)
103
- write('end')
109
+ # repeatedly take the first root (by name) whose referencing roots are already written; on a cycle,
110
+ # fall back to name order so no root is dropped
111
+ def ordered_root_models
112
+ remaining = @graph.models.select { |model| root_model?(model) }
113
+ ordered = []
114
+ until remaining.empty?
115
+ model = remaining.detect { |root| (purged_first(root) & remaining).empty? } || remaining.first
116
+ ordered << remaining.delete(model)
104
117
  end
118
+ ordered
105
119
  end
106
120
 
107
- def ignore_missing_tables
108
- missing_tables = sorted_models.map(&:table_name) - @tables
109
- return if missing_tables.empty?
110
-
111
- line_break
112
- missing_tables.each do |table_name|
113
- write("ignore_table #{table_name.to_sym.inspect}")
114
- end
121
+ # roots holding rows that reference root's rows (directly or through nested tables)
122
+ def purged_first(root)
123
+ @graph.reachable(root).select { |model| model != root && root_model?(model) }
115
124
  end
116
125
  end
117
126
  end
@@ -36,6 +36,14 @@ module DBPurger
36
36
  )
37
37
  end
38
38
 
39
+ def nullify_records(event)
40
+ self.class.metrics.update_nullify_records_stats(
41
+ event.payload[:table_name],
42
+ event.duration,
43
+ event.payload[:records_nullified] || 0
44
+ )
45
+ end
46
+
39
47
  def next_batch(event)
40
48
  self.class.metrics.update_lookup_stats(
41
49
  event.payload[:table_name],
@@ -7,6 +7,7 @@ module DBPurger
7
7
  :finished_at,
8
8
  :purge_stats,
9
9
  :delete_stats,
10
+ :nullify_stats,
10
11
  :lookup_stats,
11
12
  :filter_stats
12
13
 
@@ -18,6 +19,7 @@ module DBPurger
18
19
  @started_at = Time.now
19
20
  @purge_stats = {}
20
21
  @delete_stats = {}
22
+ @nullify_stats = {}
21
23
  @lookup_stats = {}
22
24
  @filter_stats = {}
23
25
  @finished_at = nil
@@ -48,6 +50,14 @@ module DBPurger
48
50
  stats
49
51
  end
50
52
 
53
+ def update_nullify_records_stats(table_name, duration, num_nullified)
54
+ stats = (@nullify_stats[table_name] ||= Hash.new(0))
55
+ stats[:duration] += duration
56
+ stats[:num_nullify_queries] += 1
57
+ stats[:num_nullified] += num_nullified
58
+ stats
59
+ end
60
+
51
61
  def update_lookup_stats(table_name, duration, records_found)
52
62
  stats = (@lookup_stats[table_name] ||= Hash.new(0))
53
63
  stats[:duration] += duration
@@ -72,6 +82,7 @@ module DBPurger
72
82
  finished_at: @finished_at,
73
83
  purge_stats: @purge_stats,
74
84
  delete_stats: @delete_stats,
85
+ nullify_stats: @nullify_stats,
75
86
  lookup_stats: @lookup_stats,
76
87
  filter_stats: @filter_stats
77
88
  }
@@ -0,0 +1,45 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::NullifyTable clears a nullable column that points at a batch of rows about to be deleted,
5
+ # the purge-time equivalent of ON DELETE SET NULL. It updates the referencing rows instead of deleting them.
6
+ class NullifyTable
7
+ include PurgeTableHelper
8
+
9
+ def initialize(database, table, purge_values)
10
+ @database = database
11
+ @table = table
12
+ @purge_values = purge_values
13
+ end
14
+
15
+ def nullify!
16
+ ActiveSupport::Notifications.instrument('nullify_records.db_purger',
17
+ table_name: @table.name,
18
+ nullify_field: @table.field,
19
+ num_records: @purge_values.size) do |payload|
20
+ payload[:records_nullified] = ::DBPurger.config.explain? ? explain_nullify : nullify_records
21
+ end
22
+ end
23
+
24
+ private
25
+
26
+ def nullify_records
27
+ scope.update_all(@table.field => nil)
28
+ end
29
+
30
+ def scope
31
+ scope = model.where(@table.field => @purge_values)
32
+ scope = scope.where(@table.conditions) if @table.conditions
33
+ scope
34
+ end
35
+
36
+ def explain_nullify
37
+ ::DBPurger.config.explain_file.puts("#{explain_update_sql(scope, field_quoted, 'NULL')};")
38
+ scope.count
39
+ end
40
+
41
+ def field_quoted
42
+ model.connection.quote_column_name(@table.field)
43
+ end
44
+ end
45
+ end
@@ -8,39 +8,48 @@ module DBPurger
8
8
  attr_reader :parent_tables,
9
9
  :child_tables,
10
10
  :ignore_tables,
11
- :search_tables
11
+ :search_tables,
12
+ :nullify_tables
12
13
 
13
14
  def initialize
14
15
  @parent_tables = []
15
16
  @child_tables = []
16
17
  @ignore_tables = []
17
18
  @search_tables = []
19
+ @nullify_tables = []
18
20
  end
19
21
 
20
22
  def purge!(database, purge_value)
21
- raise('plan has no base_table') unless @base_table
23
+ raise('plan has no base_table or top-level parent_table') if root_tables.empty?
24
+ raise('top-level child_tables require a base_table') unless @base_table || @child_tables.empty?
25
+ raise('top-level nullify_tables require a base_table') unless @base_table || @nullify_tables.empty?
22
26
 
23
27
  MetricSubscriber.reset!
24
- num_deleted = PurgeTable.new(database, @base_table, @base_table.field, purge_value).purge!
28
+ num_deleted = purge_root_tables(database, purge_value)
29
+ purge_search_tables(database)
25
30
  MetricSubscriber.finished!
26
31
  num_deleted
27
32
  end
28
33
 
34
+ # tables that receive the purge value directly: the base_table (if any) and top-level parent_tables
35
+ def root_tables
36
+ (@base_table ? [@base_table] : []) + @parent_tables
37
+ end
38
+
29
39
  def tables
30
40
  all_tables = @base_table ? [@base_table] + @base_table.tables : []
31
41
  all_tables += @parent_tables + @parent_tables.map(&:tables) +
32
42
  @child_tables + @child_tables.map(&:tables) +
33
- @search_tables + @search_tables.map(&:tables)
43
+ @search_tables + @search_tables.map(&:tables) +
44
+ @nullify_tables
34
45
  all_tables.flatten!
35
46
  all_tables.compact!
36
47
  all_tables
37
48
  end
38
49
 
50
+ # the tables of a nested plan (nested plans never have a base_table)
39
51
  def foreign_tables
40
- (@base_table ? [@base_table] : []) +
41
- @parent_tables +
42
- @child_tables +
43
- @search_tables
52
+ @parent_tables + @child_tables + @search_tables
44
53
  end
45
54
 
46
55
  def table_names
@@ -51,7 +60,8 @@ module DBPurger
51
60
  @base_table.nil? &&
52
61
  @parent_tables.empty? &&
53
62
  @child_tables.empty? &&
54
- @search_tables.empty?
63
+ @search_tables.empty? &&
64
+ @nullify_tables.empty?
55
65
  end
56
66
 
57
67
  def ignore_table?(table_name)
@@ -63,5 +73,20 @@ module DBPurger
63
73
  end
64
74
  end
65
75
  end
76
+
77
+ private
78
+
79
+ def purge_root_tables(database, purge_value)
80
+ root_tables.sum do |table|
81
+ PurgeTable.new(database, table, table.field, purge_value).purge!
82
+ end
83
+ end
84
+
85
+ # with a base_table these live in its nested plan and are purged by it
86
+ def purge_search_tables(database)
87
+ @search_tables.each do |table|
88
+ PurgeTableScanner.new(database, table).purge!
89
+ end
90
+ end
66
91
  end
67
92
  end
@@ -3,6 +3,8 @@
3
3
  module DBPurger
4
4
  # DBPurger::PlanBuilder is used to build the relationships between tables in a convenient way
5
5
  class PlanBuilder
6
+ NULLIFY_TABLE_OPTIONS = %i[conditions].freeze
7
+
6
8
  def initialize(plan)
7
9
  @plan = plan
8
10
  end
@@ -31,6 +33,22 @@ module DBPurger
31
33
  table
32
34
  end
33
35
 
36
+ # rows whose field matches the enclosing batch's primary keys get field set to NULL instead of being deleted
37
+ def nullify_table(table_name, field, options = {})
38
+ unsupported_options = options.keys - NULLIFY_TABLE_OPTIONS
39
+ unless unsupported_options.empty?
40
+ raise(ArgumentError, "nullify_table does not support #{unsupported_options.map(&:inspect).join(', ')}")
41
+ end
42
+
43
+ table = create_table(table_name, field, options)
44
+ if @plan.base_table
45
+ @plan.base_table.nested_plan.nullify_tables << table
46
+ else
47
+ @plan.nullify_tables << table
48
+ end
49
+ table
50
+ end
51
+
34
52
  def ignore_table(table_name)
35
53
  @plan.ignore_tables << table_name
36
54
  end
@@ -11,6 +11,7 @@ module DBPurger
11
11
  validate :validate_no_missing_tables
12
12
  validate :validate_no_unknown_tables
13
13
  validate :validate_tables
14
+ validate :validate_nullify_tables
14
15
 
15
16
  def initialize(database, plan)
16
17
  @database = database
@@ -29,8 +30,16 @@ module DBPurger
29
30
 
30
31
  private
31
32
 
33
+ # a plan is rooted either by a base_table or by one or more top-level parent_tables
32
34
  def validate_base_table
33
- errors.add(:base_table, 'is required') unless @plan.base_table
35
+ if @plan.root_tables.empty?
36
+ errors.add(:base_table, 'or a top-level parent_table is required')
37
+ elsif !@plan.child_tables.empty?
38
+ # without a base_table there are no ids to propagate; declared before one, they are never reached
39
+ errors.add(:base_table, 'must be declared before top-level child_tables')
40
+ elsif !@plan.nullify_tables.empty?
41
+ errors.add(:base_table, 'must be declared before top-level nullify_tables')
42
+ end
34
43
  end
35
44
 
36
45
  def validate_no_missing_tables
@@ -45,6 +54,27 @@ module DBPurger
45
54
  @plan.tables.each { |table| validate_table_definition(table) }
46
55
  end
47
56
 
57
+ # a NOT NULL column can't be unlinked, so the purge would fail at runtime instead of here
58
+ def validate_nullify_tables
59
+ nullify_tables.each do |table|
60
+ next unless (model = find_model_for_table(table)) # reported by validate_tables
61
+ next unless not_nullable_column?(model, table.field)
62
+
63
+ errors.add(:table, "#{table.name}.#{table.field} (nullify_table) is not nullable")
64
+ end
65
+ end
66
+
67
+ # a missing column is reported by validate_tables
68
+ def not_nullable_column?(model, field)
69
+ column = model.columns_hash[field.to_s]
70
+ column && !column.null
71
+ end
72
+
73
+ def nullify_tables
74
+ @plan.nullify_tables +
75
+ @plan.tables.select(&:nested_tables?).flat_map { |table| table.nested_plan.nullify_tables }
76
+ end
77
+
48
78
  def validate_table_definition(table)
49
79
  unless (model = find_model_for_table(table))
50
80
  errors.add(:table, "#{table.name} has no model")
@@ -0,0 +1,56 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::PlanWriter renders plan DSL source and tracks which tables it has written
5
+ class PlanWriter
6
+ INDENT = ' '
7
+
8
+ attr_reader :output,
9
+ :table_names
10
+
11
+ def initialize
12
+ @output = ''.dup
13
+ @indent_depth = 0
14
+ @table_names = []
15
+ end
16
+
17
+ def table(table_type, table_name, field)
18
+ @table_names << table_name
19
+ write(table_call(table_type, table_name, field))
20
+ end
21
+
22
+ def table_block(table_type, table_name, field)
23
+ @table_names << table_name
24
+ write("#{table_call(table_type, table_name, field)} do")
25
+ @indent_depth += 1
26
+ yield
27
+ @indent_depth -= 1
28
+ write('end')
29
+ end
30
+
31
+ def comment(str)
32
+ write("# #{str}")
33
+ end
34
+
35
+ def ignore_tables(table_names)
36
+ return if table_names.empty?
37
+
38
+ line_break
39
+ table_names.each { |table_name| write("ignore_table #{table_name.to_sym.inspect}") }
40
+ end
41
+
42
+ def line_break
43
+ @output << "\n"
44
+ end
45
+
46
+ private
47
+
48
+ def write(str)
49
+ @output << "#{INDENT * @indent_depth}#{str}\n"
50
+ end
51
+
52
+ def table_call(table_type, table_name, field)
53
+ "#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect})"
54
+ end
55
+ end
56
+ end
@@ -13,10 +13,6 @@ module DBPurger
13
13
  @num_deleted = 0
14
14
  end
15
15
 
16
- def model
17
- @model ||= @database.models.detect { |m| m.table_name == @table.name.to_s }
18
- end
19
-
20
16
  def purge!
21
17
  ActiveSupport::Notifications.instrument('purge.db_purger',
22
18
  table_name: @table.name,
@@ -10,9 +10,19 @@ module DBPurger
10
10
  private
11
11
 
12
12
  def purge_nested_tables(batch)
13
+ nullify_tables(batch) unless @table.nested_plan.nullify_tables.empty?
13
14
  purge_child_tables(batch) unless @table.nested_plan.child_tables.empty?
14
15
  end
15
16
 
17
+ # unlink rows that point at this batch before anything in it is deleted
18
+ def nullify_tables(batch)
19
+ ids = batch_values(batch, model.primary_key)
20
+
21
+ @table.nested_plan.nullify_tables.each do |table|
22
+ NullifyTable.new(@database, table, ids).nullify!
23
+ end
24
+ end
25
+
16
26
  def purge_child_tables(batch)
17
27
  ids = batch_values(batch, model.primary_key)
18
28
 
@@ -69,7 +79,7 @@ module DBPurger
69
79
  def explain(scope)
70
80
  sql =
71
81
  if @table.mark_deleted_field
72
- explain_update_sql(scope)
82
+ explain_update_sql(scope, mark_deleted_field_quoted, mark_deleted_value_quoted)
73
83
  else
74
84
  scope.to_sql.sub(/SELECT .*?FROM/, 'DELETE FROM')
75
85
  end
@@ -77,10 +87,10 @@ module DBPurger
77
87
  scope.count
78
88
  end
79
89
 
80
- def explain_update_sql(scope)
90
+ def explain_update_sql(scope, field_quoted, value_quoted)
81
91
  sql = scope.to_sql.dup
82
92
  sql.sub!(/SELECT .*?FROM/, 'UPDATE')
83
- sql.sub!('WHERE', "SET #{mark_deleted_field_quoted} = #{mark_deleted_value_quoted} WHERE")
93
+ sql.sub!('WHERE', "SET #{field_quoted} = #{value_quoted} WHERE")
84
94
  sql
85
95
  end
86
96
 
@@ -11,10 +11,6 @@ module DBPurger
11
11
  @num_deleted = 0
12
12
  end
13
13
 
14
- def model
15
- @model ||= @database.models.detect { |m| m.table_name == @table.name.to_s }
16
- end
17
-
18
14
  def purge!
19
15
  ActiveSupport::Notifications.instrument('purge.db_purger',
20
16
  table_name: @table.name) do |payload|
@@ -38,7 +38,8 @@ module DBPurger
38
38
 
39
39
  # nested tables that depend on this table's rows (search tables scan independently)
40
40
  def nested_key_tables?
41
- nested_tables? && !(@nested_plan.child_tables.empty? && @nested_plan.parent_tables.empty?)
41
+ nested_tables? &&
42
+ !(@nested_plan.child_tables.empty? && @nested_plan.parent_tables.empty? && @nested_plan.nullify_tables.empty?)
42
43
  end
43
44
 
44
45
  def tables
data/lib/db-purger.rb CHANGED
@@ -2,17 +2,20 @@
2
2
 
3
3
  # DBPurger is a tool to delete data from tables based on a initial purge value
4
4
  module DBPurger
5
+ autoload :AssociationGraph, 'db-purger/association_graph'
5
6
  autoload :Config, 'db-purger/config'
6
7
  autoload :DynamicPlanBuilder, 'db-purger/dynamic_plan_builder'
7
8
  autoload :Executor, 'db-purger/executor'
8
9
  autoload :Metrics, 'db-purger/metrics'
9
10
  autoload :MetricSubscriber, 'db-purger/metric_subscriber'
11
+ autoload :NullifyTable, 'db-purger/nullify_table'
10
12
  autoload :PurgeTable, 'db-purger/purge_table'
11
13
  autoload :PurgeTableHelper, 'db-purger/purge_table_helper'
12
14
  autoload :PurgeTableScanner, 'db-purger/purge_table_scanner'
13
15
  autoload :Plan, 'db-purger/plan'
14
16
  autoload :PlanBuilder, 'db-purger/plan_builder'
15
17
  autoload :PlanValidator, 'db-purger/plan_validator'
18
+ autoload :PlanWriter, 'db-purger/plan_writer'
16
19
  autoload :Table, 'db-purger/table'
17
20
 
18
21
  # The config in effect for the current thread: the one set by with_config, else the global default
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: db-purger
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.5.0
4
+ version: 0.7.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Doug Youch
@@ -29,14 +29,20 @@ dependencies:
29
29
  requirements:
30
30
  - - "~>"
31
31
  - !ruby/object:Gem::Version
32
- version: '0.7'
32
+ version: '0.9'
33
+ - - ">="
34
+ - !ruby/object:Gem::Version
35
+ version: 0.9.1
33
36
  type: :runtime
34
37
  prerelease: false
35
38
  version_requirements: !ruby/object:Gem::Requirement
36
39
  requirements:
37
40
  - - "~>"
38
41
  - !ruby/object:Gem::Version
39
- version: '0.7'
42
+ version: '0.9'
43
+ - - ">="
44
+ - !ruby/object:Gem::Version
45
+ version: 0.9.1
40
46
  description: DB Purger deletes (or soft-deletes) every row related to a single top-level
41
47
  record (e.g. a company or account) using a declarative Ruby purge plan. Tables are
42
48
  purged in primary-key batches, child tables before their parents, with plan validation
@@ -51,14 +57,17 @@ files:
51
57
  - LICENSE.txt
52
58
  - README.md
53
59
  - lib/db-purger.rb
60
+ - lib/db-purger/association_graph.rb
54
61
  - lib/db-purger/config.rb
55
62
  - lib/db-purger/dynamic_plan_builder.rb
56
63
  - lib/db-purger/executor.rb
57
64
  - lib/db-purger/metric_subscriber.rb
58
65
  - lib/db-purger/metrics.rb
66
+ - lib/db-purger/nullify_table.rb
59
67
  - lib/db-purger/plan.rb
60
68
  - lib/db-purger/plan_builder.rb
61
69
  - lib/db-purger/plan_validator.rb
70
+ - lib/db-purger/plan_writer.rb
62
71
  - lib/db-purger/purge_table.rb
63
72
  - lib/db-purger/purge_table_helper.rb
64
73
  - lib/db-purger/purge_table_scanner.rb