db-purger 0.5.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 863444350ef64942e0669feacdf4a119514c16369a31eeb99e31a5f5435028ef
4
- data.tar.gz: 66b063187fdaa33232a6fffe76f9496264e8bd74e2a835013139b4e19f3e045e
3
+ metadata.gz: 7b756bb20c0bcea96d050a3772cb2606676891d485de6db0bf9058253f0092c3
4
+ data.tar.gz: 9dfe62f982ecd99ef96dadd535a7496d1802d0392083b8f63f9166f354bb9368
5
5
  SHA512:
6
- metadata.gz: dfd50964ce367548606fe758df1d1cfb48d6891d9ac09993079b851a78658807fa4a5ff8633709a49f42b9d67c9367260721cd2d0f1454dbd7da197f3363f54b
7
- data.tar.gz: ccc2bee89229fad63fd2dc7d20682b07b008b29040cb66a92aa1cd4c48bd8f09c5f6b91834ead524be8f3394d54c8af08ce8df5d61bc5067a6229b14187c7ede
6
+ metadata.gz: 2a365e0a0c25b33eb79060930289c2820823e9e25fefbf4edcdf13ccfad6e55595e3d380d5da97a76476262051b9ee36095cdf51e30ebf3248bcf43dfd8f0fa2
7
+ data.tar.gz: 9171105e396d405cacde9debad9b6923d1d318c24ebbf269a9f872acbf6e087a0ea90506e9d24c0ab99407e1874fcd9d46e74afb73a736203b1f1da178791970
data/ARCHITECTURE.md CHANGED
@@ -31,7 +31,7 @@ db-purger is small (~900 lines) and splits cleanly into three layers: **describe
31
31
  | `lib/db-purger.rb` | Autoloads everything; holds the global `DBPurger.config`. |
32
32
  | `config.rb` | Global options: `explain?`, `explain_file`, `datetime_format`. |
33
33
  | `table.rb` | Value object for one table in the plan: name, match field, options, and a lazily created nested `Plan`. |
34
- | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, search and ignored tables. `#purge!` is the run entry point. |
34
+ | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, search and ignored tables. `#purge!` is the run entry point; `#root_tables` are the tables it starts from. |
35
35
  | `plan_builder.rb` | The DSL. `instance_eval`s a plan file or block against a `Plan`; nested blocks get a new builder bound to that table's nested plan. |
36
36
  | `plan_validator.rb` | `ActiveModel::Validations` over plan vs. schema: missing tables, unknown tables, unknown columns. |
37
37
  | `executor.rb` | Convenience façade: loads a plan file, applies config options, `verify!`, `purge!`. |
@@ -39,7 +39,9 @@ db-purger is small (~900 lines) and splits cleanly into three layers: **describe
39
39
  | `purge_table_scanner.rb` | Purges a `purge_table_search` table: full `find_in_batches` scan filtered through the user's `search_proc`. |
40
40
  | `purge_table_helper.rb` | Shared behaviour for both purgers: nested-table recursion, delete vs. soft delete vs. explain, transactions. |
41
41
  | `metrics.rb` / `metric_subscriber.rb` | Aggregate timing and row counts per table from the notification events. |
42
- | `dynamic_plan_builder.rb` | Generates plan-file source from `has_many` associations; a bootstrap tool, not used at purge time. |
42
+ | `dynamic_plan_builder.rb` | Generates plan-file source (`build` for a base table, `build_for` for several roots); a bootstrap tool, not used at purge time. |
43
+ | `association_graph.rb` | For the generator: which tables reference a model, and by which column, from its `has_many`/`has_one`/HABTM reflections. |
44
+ | `plan_writer.rb` | For the generator: renders plan DSL text and records which tables it wrote. |
43
45
 
44
46
  ## The plan tree
45
47
 
@@ -56,12 +58,25 @@ Plan (root)
56
58
  └── ignore_tables
57
59
  ```
58
60
 
61
+ Without a `base_table`, top-level `parent_table`s stay in the root plan's `parent_tables` and each one is a
62
+ root:
63
+
64
+ ```
65
+ Plan (root)
66
+ ├── parent_tables: [calls(:oid) ─▶ nested Plan ..., emails(:oid) ─▶ nested Plan ..., sms_messages(:oid) ...]
67
+ ├── search_tables
68
+ └── ignore_tables
69
+ ```
70
+
59
71
  `Plan#tables` flattens this tree for validation; `Table#foreign_keys` collects the `foreign_key:` columns of
60
72
  a table's direct children so the purger can `SELECT` them alongside the primary key.
61
73
 
62
74
  ## Purge algorithm
63
75
 
64
- `Plan#purge!` resets metrics and starts a `PurgeTable` on the base table with the purge value. Each
76
+ `Plan#purge!` resets metrics and starts a `PurgeTable` with the purge value on each root table
77
+ (`root_tables` = the `base_table`, if any, followed by top-level `parent_tables`, in declaration order), then runs
78
+ any top-level search tables. With a `base_table`, `root_tables` is just `[base_table]`, which is the original
79
+ single-root algorithm. Each
65
80
  `PurgeTable#purge!` does:
66
81
 
67
82
  ```
@@ -157,4 +172,8 @@ fetch happens inside `find_in_batches` rather than in a block the scanner contro
157
172
  - `spec/integrations/*` — end-to-end purges over the schema in `spec/support/db/schema.rb`, asserting row-count
158
173
  deltas per table and, for explain mode, the exact SQL in `spec/fixtures/delete_plan.sql`.
159
174
  - `spec/support/test_db.rb` builds a fresh SQLite database and dynamic models (`TestDB::*`) for each run.
175
+ - `spec/integrations/multi_root_plan_spec.rb` purges one org from a separate outreach schema (`spec/support/outreach_db.rb`,
176
+ real FOREIGN KEY constraints) seeded with two orgs by `OutreachSeeder`, and asserts the exact surviving rows of
177
+ every table for the generated plan, a hand-written small-batch plan and the equivalent `base_table` plan.
178
+ - `spec/support/throwaway_db.rb` builds standalone SQLite databases from raw SQL for schema-specific specs.
160
179
  - `spec/fixtures/*.plan.rb` — plan files used for loading and validation cases.
data/README.md CHANGED
@@ -2,6 +2,7 @@
2
2
 
3
3
  [![CI](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml/badge.svg?branch=master)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
4
4
  [![Coverage](https://raw.githubusercontent.com/dougyouch/db-purger/badges/coverage.svg)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
5
+ [![Branch coverage](https://raw.githubusercontent.com/dougyouch/db-purger/badges/branch-coverage.svg)](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
5
6
  [![Gem Version](https://img.shields.io/gem/v/db-purger)](https://rubygems.org/gems/db-purger)
6
7
 
7
8
  Purge every row tied to a single top-level record — a company, an account, a tenant — across all of the
@@ -112,16 +113,45 @@ executor.verify! # raises 'purge plan failed verification', errors pr
112
113
  deleted = executor.purge!(42)
113
114
  ```
114
115
 
115
- `purge!` returns the number of base-table rows deleted.
116
+ `purge!` returns the number of root-table rows deleted (the base table, plus any top-level `parent_table`s).
117
+
118
+ ### Plans with several top-level tables
119
+
120
+ `base_table` is shorthand for "one root table, with everything after it nested underneath". When several tables
121
+ are equally top-level (an outreach product's `emails`, `sms_messages` and `calls`, all keyed by `oid`), leave
122
+ `base_table` out and declare each root as a top-level `parent_table`:
123
+
124
+ ```ruby
125
+ # config/outreach.plan.rb
126
+ parent_table(:calls, :oid) do # calls.email_id -> emails.id, so calls go first
127
+ child_table(:call_notes, :call_id)
128
+ child_table(:call_recordings, :call_id)
129
+ child_table(:call_tags, :call_id)
130
+ end
131
+
132
+ parent_table(:emails, :oid) do
133
+ child_table(:email_attachments, :email_id)
134
+ end
135
+
136
+ parent_table(:sms_messages, :oid) do
137
+ child_table(:sms_deliveries, :sms_message_id)
138
+ end
139
+
140
+ ignore_table :users
141
+ ```
142
+
143
+ Each root is purged by `oid = purge_value`, children first, **in declaration order**: when one root's rows
144
+ reference another's, declare the referencing root first. Without a `base_table`, top-level `child_table`s are
145
+ an error (there is no enclosing batch to take ids from). Existing `base_table` plans run exactly as before.
116
146
 
117
147
  ## The plan DSL
118
148
 
119
149
  | Method | Meaning |
120
150
  |---|---|
121
- | `base_table(table, field, opts = {}, &block)` | The root of the purge. Rows where `field = purge_value` are purged. Declare it **first** — every subsequent top-level call nests under it. |
151
+ | `base_table(table, field, opts = {}, &block)` | Optional single root. Rows where `field = purge_value` are purged. Declare it **first** — every subsequent top-level call nests under it. |
122
152
  | `child_table(table, field, opts = {}, &block)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch. Purged before that batch is deleted. |
123
153
  | `child_table(table, :id, foreign_key: :col, &block)` | Inverted relationship: the *enclosing* table holds `col` pointing at this table's `id`. Deleted in the same transaction, right after the enclosing batch. |
124
- | `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
154
+ | `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. At the top level of a plan without a `base_table`, each one is a root, purged in declaration order. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
125
155
  | `purge_table_search(table, field, opts = {}) { \|batch\| ... }` | Scans the whole table in batches; the block receives each batch and returns the records to purge. For orphans that can't be reached by a key. |
126
156
  | `ignore_table(name_or_regexp)` | Exclude a table from validation. |
127
157
 
@@ -166,14 +196,25 @@ plan.purge!(database, 42)
166
196
 
167
197
  ### Generating a starting plan
168
198
 
169
- `DynamicPlanBuilder` walks the `has_many` associations dynamic-active-model discovered and emits a plan file,
170
- listing every unreachable table as `ignore_table`. Treat the output as a first draft: it only knows about
171
- conventional `<singular_table>_id` foreign keys and cannot infer polymorphic, soft-delete, or search rules.
199
+ `DynamicPlanBuilder` walks the `has_many`, `has_one` and `has_and_belongs_to_many` associations
200
+ dynamic-active-model discovered, using each association's real foreign key, and emits a plan file listing every
201
+ unreachable table as `ignore_table`.
172
202
 
173
203
  ```ruby
174
- puts DBPurger::DynamicPlanBuilder.new(database).build(:companies, :id)
204
+ builder = DBPurger::DynamicPlanBuilder.new(database)
205
+ puts builder.build(:companies, :id) # single base_table plan
206
+ puts builder.build_for(:oid) # one top-level parent_table per table holding oid
175
207
  ```
176
208
 
209
+ `build_for` makes **every** table holding the field a root, so rows with a null foreign key to another root
210
+ (an `email_recipients` row without an email) are still purged, and orders the roots so a root referencing
211
+ another root's rows comes first. HABTM join tables are emitted as leaves, never walking into the shared table on
212
+ the other side, and nothing is nested under a table without a primary key. A foreign-key cycle is written as a
213
+ comment instead of recursing.
214
+
215
+ Treat the output as a first draft: it cannot infer polymorphic (`as:`), soft-delete, `belongs_to`-owned
216
+ (`foreign_key:`) or search rules.
217
+
177
218
  ## Validation
178
219
 
179
220
  `Executor#verify!` (or `DBPurger::PlanValidator.new(database, plan).valid?`) checks that:
@@ -181,7 +222,8 @@ puts DBPurger::DynamicPlanBuilder.new(database).build(:companies, :id)
181
222
  - every table in the database is either in the plan or ignored (`missing_tables`)
182
223
  - every table in the plan exists in the database (`unknown_tables`)
183
224
  - every field, `foreign_key` and `mark_deleted_field` named in the plan is a real column
184
- - the plan has a `base_table`, and every `batch_size` is positive
225
+ - the plan has a `base_table` or at least one top-level `parent_table`, no top-level `child_table` is left
226
+ unreachable, and every `batch_size` is positive
185
227
  - tables without a primary key have no nested child or parent tables (there would be no ids to propagate)
186
228
 
187
229
  Run it in CI against your schema so a new table can't ship without a purge decision.
@@ -235,8 +277,8 @@ Metrics are reset at the start of each `Plan#purge!`. To feed your own telemetry
235
277
 
236
278
  ## Caveats
237
279
 
238
- - **Only the base table is the entry point.** Top-level `parent_table`/`child_table` calls made before
239
- `base_table` are ignored by `Plan#purge!`.
280
+ - **Declare `base_table` first.** A `child_table` declared before it is never reached (the validator reports
281
+ it); a `parent_table` declared before it becomes a separate root, purged after the base table.
240
282
  - **Not one big transaction.** Each batch is its own set of statements (foreign-key children share a
241
283
  transaction with their parent batch). An interrupted purge is safe to re-run with the same value.
242
284
  - **Soft-deleted rows still match.** A `mark_deleted_field` table is not filtered on that field; add
@@ -254,14 +296,14 @@ script/console
254
296
 
255
297
  CI (`.github/workflows/ci.yml`) runs RuboCop and the specs on Ruby 4.0 for every push and pull request.
256
298
  The HTML coverage report is attached to each run as the `coverage` artifact, and pushes to `master` refresh
257
- the coverage badge on the `badges` branch.
299
+ the line and branch coverage badges on the `badges` branch.
258
300
 
259
301
  See [ARCHITECTURE.md](ARCHITECTURE.md) for how the pieces fit together.
260
302
 
261
303
  ## Releasing
262
304
 
263
305
  1. Bump `s.version` in `db-purger.gemspec` and merge to `master`.
264
- 2. Tag and push: `git tag v0.5.0 && git push origin v0.5.0`
306
+ 2. Tag and push: `git tag v0.6.0 && git push origin v0.6.0`
265
307
 
266
308
  `.github/workflows/release.yml` re-runs CI, checks the tag matches the gemspec version, publishes to RubyGems
267
309
  via trusted publishing (no API key), and creates a GitHub release with the `.gem` attached.
@@ -0,0 +1,76 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::AssociationGraph answers "which tables reference this model, and by which column" from the
5
+ # has_many, has_one and has_and_belongs_to_many associations dynamic-active-model discovered
6
+ class AssociationGraph
7
+ # a table holding foreign_key that points at the parent model's primary key
8
+ Edge = Struct.new(:model, :foreign_key)
9
+
10
+ def initialize(database)
11
+ @database = database
12
+ @edges = {}
13
+ end
14
+
15
+ # database.models order depends on how the adapter lists tables, which varies by platform;
16
+ # sort so generated plans are deterministic
17
+ def models
18
+ @models ||= @database.models.sort_by(&:table_name)
19
+ end
20
+
21
+ def model_for(table_name)
22
+ models.detect { |model| model.table_name == table_name.to_s }
23
+ end
24
+
25
+ # one edge per (table, foreign key); a habtm join table also reached by a has_many appears once
26
+ def edges(model)
27
+ @edges[model] ||= model.reflect_on_all_associations
28
+ .filter_map { |reflection| edge_for(reflection) }
29
+ .uniq { |edge| edge_key(edge) }
30
+ .sort_by { |edge| edge_key(edge) }
31
+ end
32
+
33
+ # every model reachable from model through edges, excluding model itself unless there is a cycle
34
+ def reachable(model, seen = Set.new)
35
+ edges(model).each do |edge|
36
+ next if seen.include?(edge.model)
37
+
38
+ seen << edge.model
39
+ reachable(edge.model, seen)
40
+ end
41
+ seen
42
+ end
43
+
44
+ def column?(model, field)
45
+ model.column_names.include?(field.to_s)
46
+ end
47
+
48
+ private
49
+
50
+ def edge_for(reflection)
51
+ return if skip_reflection?(reflection)
52
+
53
+ case reflection
54
+ when ActiveRecord::Reflection::HasAndBelongsToManyReflection
55
+ join_table_edge(reflection)
56
+ when ActiveRecord::Reflection::HasManyReflection, ActiveRecord::Reflection::HasOneReflection
57
+ Edge.new(reflection.klass, reflection.foreign_key.to_s)
58
+ end
59
+ end
60
+
61
+ # through associations are reached via their own direct associations; polymorphic (as:) ones need a
62
+ # type condition the generator cannot infer
63
+ def skip_reflection?(reflection)
64
+ reflection.options[:through] || reflection.options[:as]
65
+ end
66
+
67
+ def join_table_edge(reflection)
68
+ join_model = model_for(reflection.join_table)
69
+ Edge.new(join_model, reflection.foreign_key.to_s) if join_model
70
+ end
71
+
72
+ def edge_key(edge)
73
+ [edge.model.table_name, edge.foreign_key]
74
+ end
75
+ end
76
+ end
@@ -3,115 +3,124 @@
3
3
  module DBPurger
4
4
  # DBPurger::DynamicPlanBuilder generates a purge plan based on the database relations
5
5
  class DynamicPlanBuilder
6
- INDENT = ' '
7
-
8
- attr_reader :output
9
-
10
6
  def initialize(database)
11
- @database = database
12
- @output = ''.dup
13
- @indent_depth = 0
14
- @tables = []
7
+ @graph = AssociationGraph.new(database)
8
+ @writer = PlanWriter.new
9
+ end
10
+
11
+ def output
12
+ @writer.output
15
13
  end
16
14
 
15
+ # plan rooted at a single base table
17
16
  def build(base_table_name, field)
18
- write_table('base', base_table_name.to_s, field, [], nil)
19
- line_break
20
- model = find_model_for_table(base_table_name)
21
- foreign_key = foreign_key_name(model)
17
+ model = @graph.model_for(base_table_name)
18
+ @writer.table('base', model.table_name, field)
19
+ @writer.line_break
22
20
  if model.primary_key == field.to_s
23
- add_parent_tables(base_table_name, foreign_key)
21
+ add_referencing_parent_tables(model)
24
22
  else
25
- add_parent_tables(base_table_name, field)
26
- unless (child_models = find_child_models(model, foreign_key)).empty?
27
- line_break unless field == :id
28
- add_child_tables(child_models, foreign_key, 0)
29
- end
23
+ add_sibling_parent_tables(model, field)
24
+ add_base_child_tables(model)
30
25
  end
31
- ignore_missing_tables
32
- @output
26
+ finish
27
+ end
28
+
29
+ # plan with one top-level parent_table per table holding field (e.g. :oid), each purged by field
30
+ # directly so rows with a null foreign key are not missed; the tables referencing each root are nested
31
+ # under it, and roots are ordered so a root referencing another root's rows is purged first
32
+ def build_for(field)
33
+ @root_field = field.to_s
34
+ ordered_root_models.each_with_index do |model, idx|
35
+ @writer.line_break if idx.positive?
36
+ write_table('parent', model, field, [])
37
+ end
38
+ finish
33
39
  end
34
40
 
35
41
  private
36
42
 
37
- def find_model_for_table(base_table_name)
38
- @database.models.detect { |m| m.table_name == base_table_name.to_s }
43
+ def finish
44
+ @writer.ignore_tables(@graph.models.map(&:table_name) - @writer.table_names)
45
+ output
39
46
  end
40
47
 
41
- def write(str)
42
- @output << "#{INDENT * @indent_depth}#{str}\n"
48
+ # base_table(:companies, :id): tables holding companies.id are keyed directly on the purge value
49
+ def add_referencing_parent_tables(model)
50
+ @graph.edges(model).each { |edge| write_edge('parent', edge, [model]) }
43
51
  end
44
52
 
45
- def line_break
46
- @output << "\n"
53
+ # base_table(:employments, :company_id): other tables holding company_id share the purge value
54
+ def add_sibling_parent_tables(model, field)
55
+ @graph.models.each do |sibling|
56
+ next if sibling == model || !@graph.column?(sibling, field)
57
+
58
+ write_table('parent', sibling, field, [model])
59
+ end
47
60
  end
48
61
 
49
- def add_parent_tables(base_table_name, field)
50
- sorted_models.each do |model|
51
- next if model.table_name == base_table_name.to_s
52
- next unless column?(model, field)
62
+ def add_base_child_tables(model)
63
+ edges = nestable_edges(model)
64
+ return if edges.empty?
53
65
 
54
- foreign_key = foreign_key_name(model)
55
- write_table('parent', model.table_name, field, find_child_models(model, foreign_key), foreign_key)
56
- end
66
+ @writer.line_break
67
+ edges.each { |edge| write_edge('child', edge, [model]) }
57
68
  end
58
69
 
59
- def add_child_tables(child_models, field, change_indent_by = 1)
60
- @indent_depth += change_indent_by
61
- child_models.each do |model|
62
- add_child_table(model, field)
70
+ def write_edge(table_type, edge, ancestors)
71
+ if ancestors.include?(edge.model)
72
+ @writer.comment("#{table_type}_table(#{edge.model.table_name.to_sym.inspect}, " \
73
+ "#{edge.foreign_key.to_sym.inspect}) skipped: cycle back to #{edge.model.table_name}")
74
+ else
75
+ write_table(table_type, edge.model, edge.foreign_key, ancestors)
63
76
  end
64
- @indent_depth -= change_indent_by
65
77
  end
66
78
 
67
- def add_child_table(model, field)
68
- foreign_key = foreign_key_name(model)
69
- write_table('child', model.table_name, field, find_child_models(model, foreign_key), foreign_key)
79
+ def write_table(table_type, model, field, ancestors)
80
+ edges = nestable_edges(model)
81
+ if edges.empty?
82
+ @writer.table(table_type, model.table_name, field)
83
+ warn_unnestable(model)
84
+ else
85
+ @writer.table_block(table_type, model.table_name, field) do
86
+ edges.each { |edge| write_edge('child', edge, ancestors + [model]) }
87
+ end
88
+ end
70
89
  end
71
90
 
72
- def find_child_models(model, field)
73
- model_has_many_associations(model).map(&:klass).select { |m| column?(m, field) }.sort_by(&:table_name)
74
- end
91
+ # purging nested tables needs this table's primary keys to propagate
92
+ def nestable_edges(model)
93
+ return [] unless model.primary_key
75
94
 
76
- # database.models order depends on how the adapter lists tables, which varies by platform;
77
- # sort so the generated plan is deterministic
78
- def sorted_models
79
- @sorted_models ||= @database.models.sort_by(&:table_name)
95
+ @graph.edges(model).reject { |edge| root_model?(edge.model) }
80
96
  end
81
97
 
82
- def model_has_many_associations(model)
83
- model.reflect_on_all_associations.select do |assoc|
84
- assoc.is_a?(ActiveRecord::Reflection::HasManyReflection)
85
- end
98
+ def root_model?(model)
99
+ @root_field && @graph.column?(model, @root_field)
86
100
  end
87
101
 
88
- def foreign_key_name(model)
89
- "#{model.table_name.singularize}_id"
90
- end
102
+ def warn_unnestable(model)
103
+ return if model.primary_key || (edges = @graph.edges(model)).empty?
91
104
 
92
- def column?(model, field)
93
- model.columns.detect { |c| c.name == field.to_s } != nil
105
+ @writer.comment("#{model.table_name} has no primary key; cannot nest " \
106
+ "#{edges.map { |edge| edge.model.table_name }.join(', ')}")
94
107
  end
95
108
 
96
- def write_table(table_type, table_name, field, child_models, foreign_key)
97
- @tables << table_name
98
- if child_models.empty?
99
- write("#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect})")
100
- else
101
- write("#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect}) do")
102
- add_child_tables(child_models, foreign_key)
103
- write('end')
109
+ # repeatedly take the first root (by name) whose referencing roots are already written; on a cycle,
110
+ # fall back to name order so no root is dropped
111
+ def ordered_root_models
112
+ remaining = @graph.models.select { |model| root_model?(model) }
113
+ ordered = []
114
+ until remaining.empty?
115
+ model = remaining.detect { |root| (purged_first(root) & remaining).empty? } || remaining.first
116
+ ordered << remaining.delete(model)
104
117
  end
118
+ ordered
105
119
  end
106
120
 
107
- def ignore_missing_tables
108
- missing_tables = sorted_models.map(&:table_name) - @tables
109
- return if missing_tables.empty?
110
-
111
- line_break
112
- missing_tables.each do |table_name|
113
- write("ignore_table #{table_name.to_sym.inspect}")
114
- end
121
+ # roots holding rows that reference root's rows (directly or through nested tables)
122
+ def purged_first(root)
123
+ @graph.reachable(root).select { |model| model != root && root_model?(model) }
115
124
  end
116
125
  end
117
126
  end
@@ -18,14 +18,21 @@ module DBPurger
18
18
  end
19
19
 
20
20
  def purge!(database, purge_value)
21
- raise('plan has no base_table') unless @base_table
21
+ raise('plan has no base_table or top-level parent_table') if root_tables.empty?
22
+ raise('top-level child_tables require a base_table') unless @base_table || @child_tables.empty?
22
23
 
23
24
  MetricSubscriber.reset!
24
- num_deleted = PurgeTable.new(database, @base_table, @base_table.field, purge_value).purge!
25
+ num_deleted = purge_root_tables(database, purge_value)
26
+ purge_search_tables(database)
25
27
  MetricSubscriber.finished!
26
28
  num_deleted
27
29
  end
28
30
 
31
+ # tables that receive the purge value directly: the base_table (if any) and top-level parent_tables
32
+ def root_tables
33
+ (@base_table ? [@base_table] : []) + @parent_tables
34
+ end
35
+
29
36
  def tables
30
37
  all_tables = @base_table ? [@base_table] + @base_table.tables : []
31
38
  all_tables += @parent_tables + @parent_tables.map(&:tables) +
@@ -36,11 +43,9 @@ module DBPurger
36
43
  all_tables
37
44
  end
38
45
 
46
+ # the tables of a nested plan (nested plans never have a base_table)
39
47
  def foreign_tables
40
- (@base_table ? [@base_table] : []) +
41
- @parent_tables +
42
- @child_tables +
43
- @search_tables
48
+ @parent_tables + @child_tables + @search_tables
44
49
  end
45
50
 
46
51
  def table_names
@@ -63,5 +68,20 @@ module DBPurger
63
68
  end
64
69
  end
65
70
  end
71
+
72
+ private
73
+
74
+ def purge_root_tables(database, purge_value)
75
+ root_tables.sum do |table|
76
+ PurgeTable.new(database, table, table.field, purge_value).purge!
77
+ end
78
+ end
79
+
80
+ # with a base_table these live in its nested plan and are purged by it
81
+ def purge_search_tables(database)
82
+ @search_tables.each do |table|
83
+ PurgeTableScanner.new(database, table).purge!
84
+ end
85
+ end
66
86
  end
67
87
  end
@@ -29,8 +29,14 @@ module DBPurger
29
29
 
30
30
  private
31
31
 
32
+ # a plan is rooted either by a base_table or by one or more top-level parent_tables
32
33
  def validate_base_table
33
- errors.add(:base_table, 'is required') unless @plan.base_table
34
+ if @plan.root_tables.empty?
35
+ errors.add(:base_table, 'or a top-level parent_table is required')
36
+ elsif !@plan.child_tables.empty?
37
+ # without a base_table there are no ids to propagate; declared before one, they are never reached
38
+ errors.add(:base_table, 'must be declared before top-level child_tables')
39
+ end
34
40
  end
35
41
 
36
42
  def validate_no_missing_tables
@@ -0,0 +1,56 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::PlanWriter renders plan DSL source and tracks which tables it has written
5
+ class PlanWriter
6
+ INDENT = ' '
7
+
8
+ attr_reader :output,
9
+ :table_names
10
+
11
+ def initialize
12
+ @output = ''.dup
13
+ @indent_depth = 0
14
+ @table_names = []
15
+ end
16
+
17
+ def table(table_type, table_name, field)
18
+ @table_names << table_name
19
+ write(table_call(table_type, table_name, field))
20
+ end
21
+
22
+ def table_block(table_type, table_name, field)
23
+ @table_names << table_name
24
+ write("#{table_call(table_type, table_name, field)} do")
25
+ @indent_depth += 1
26
+ yield
27
+ @indent_depth -= 1
28
+ write('end')
29
+ end
30
+
31
+ def comment(str)
32
+ write("# #{str}")
33
+ end
34
+
35
+ def ignore_tables(table_names)
36
+ return if table_names.empty?
37
+
38
+ line_break
39
+ table_names.each { |table_name| write("ignore_table #{table_name.to_sym.inspect}") }
40
+ end
41
+
42
+ def line_break
43
+ @output << "\n"
44
+ end
45
+
46
+ private
47
+
48
+ def write(str)
49
+ @output << "#{INDENT * @indent_depth}#{str}\n"
50
+ end
51
+
52
+ def table_call(table_type, table_name, field)
53
+ "#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect})"
54
+ end
55
+ end
56
+ end
@@ -13,10 +13,6 @@ module DBPurger
13
13
  @num_deleted = 0
14
14
  end
15
15
 
16
- def model
17
- @model ||= @database.models.detect { |m| m.table_name == @table.name.to_s }
18
- end
19
-
20
16
  def purge!
21
17
  ActiveSupport::Notifications.instrument('purge.db_purger',
22
18
  table_name: @table.name,
@@ -11,10 +11,6 @@ module DBPurger
11
11
  @num_deleted = 0
12
12
  end
13
13
 
14
- def model
15
- @model ||= @database.models.detect { |m| m.table_name == @table.name.to_s }
16
- end
17
-
18
14
  def purge!
19
15
  ActiveSupport::Notifications.instrument('purge.db_purger',
20
16
  table_name: @table.name) do |payload|
data/lib/db-purger.rb CHANGED
@@ -2,6 +2,7 @@
2
2
 
3
3
  # DBPurger is a tool to delete data from tables based on a initial purge value
4
4
  module DBPurger
5
+ autoload :AssociationGraph, 'db-purger/association_graph'
5
6
  autoload :Config, 'db-purger/config'
6
7
  autoload :DynamicPlanBuilder, 'db-purger/dynamic_plan_builder'
7
8
  autoload :Executor, 'db-purger/executor'
@@ -13,6 +14,7 @@ module DBPurger
13
14
  autoload :Plan, 'db-purger/plan'
14
15
  autoload :PlanBuilder, 'db-purger/plan_builder'
15
16
  autoload :PlanValidator, 'db-purger/plan_validator'
17
+ autoload :PlanWriter, 'db-purger/plan_writer'
16
18
  autoload :Table, 'db-purger/table'
17
19
 
18
20
  # The config in effect for the current thread: the one set by with_config, else the global default
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: db-purger
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.5.0
4
+ version: 0.6.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Doug Youch
@@ -29,14 +29,20 @@ dependencies:
29
29
  requirements:
30
30
  - - "~>"
31
31
  - !ruby/object:Gem::Version
32
- version: '0.7'
32
+ version: '0.9'
33
+ - - ">="
34
+ - !ruby/object:Gem::Version
35
+ version: 0.9.1
33
36
  type: :runtime
34
37
  prerelease: false
35
38
  version_requirements: !ruby/object:Gem::Requirement
36
39
  requirements:
37
40
  - - "~>"
38
41
  - !ruby/object:Gem::Version
39
- version: '0.7'
42
+ version: '0.9'
43
+ - - ">="
44
+ - !ruby/object:Gem::Version
45
+ version: 0.9.1
40
46
  description: DB Purger deletes (or soft-deletes) every row related to a single top-level
41
47
  record (e.g. a company or account) using a declarative Ruby purge plan. Tables are
42
48
  purged in primary-key batches, child tables before their parents, with plan validation
@@ -51,6 +57,7 @@ files:
51
57
  - LICENSE.txt
52
58
  - README.md
53
59
  - lib/db-purger.rb
60
+ - lib/db-purger/association_graph.rb
54
61
  - lib/db-purger/config.rb
55
62
  - lib/db-purger/dynamic_plan_builder.rb
56
63
  - lib/db-purger/executor.rb
@@ -59,6 +66,7 @@ files:
59
66
  - lib/db-purger/plan.rb
60
67
  - lib/db-purger/plan_builder.rb
61
68
  - lib/db-purger/plan_validator.rb
69
+ - lib/db-purger/plan_writer.rb
62
70
  - lib/db-purger/purge_table.rb
63
71
  - lib/db-purger/purge_table_helper.rb
64
72
  - lib/db-purger/purge_table_scanner.rb