db-purger 0.6.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 7b756bb20c0bcea96d050a3772cb2606676891d485de6db0bf9058253f0092c3
4
- data.tar.gz: 9dfe62f982ecd99ef96dadd535a7496d1802d0392083b8f63f9166f354bb9368
3
+ metadata.gz: 45ab4b4aa9e6e8c3778061c5deb8a8eb37fecf153bb1bae9cf298f2c83433f22
4
+ data.tar.gz: bbb9aec2aaa9dcd8e7c0609186795bf17e8fce64ffa951ce6dd1877981b7dbcc
5
5
  SHA512:
6
- metadata.gz: 2a365e0a0c25b33eb79060930289c2820823e9e25fefbf4edcdf13ccfad6e55595e3d380d5da97a76476262051b9ee36095cdf51e30ebf3248bcf43dfd8f0fa2
7
- data.tar.gz: 9171105e396d405cacde9debad9b6923d1d318c24ebbf269a9f872acbf6e087a0ea90506e9d24c0ab99407e1874fcd9d46e74afb73a736203b1f1da178791970
6
+ metadata.gz: 41d99640be3f800277142f9ef9a9ec0f0945785282551c757f92ba92de36dfad331b735533713bb98385023378727ca5365cf995cff84580dfe725cbbf10240d
7
+ data.tar.gz: bfe3512b7f289cec1b4ac7f466918c4aaa6a5c12f9deb61e60e825153dd3e1badbbcbb1eb578e80dd453504f9b7b08f70a996e4d9d10b7d1e7a654245d12ec00
data/ARCHITECTURE.md CHANGED
@@ -31,12 +31,13 @@ db-purger is small (~900 lines) and splits cleanly into three layers: **describe
31
31
  | `lib/db-purger.rb` | Autoloads everything; holds the global `DBPurger.config`. |
32
32
  | `config.rb` | Global options: `explain?`, `explain_file`, `datetime_format`. |
33
33
  | `table.rb` | Value object for one table in the plan: name, match field, options, and a lazily created nested `Plan`. |
34
- | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, search and ignored tables. `#purge!` is the run entry point; `#root_tables` are the tables it starts from. |
34
+ | `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, nullify, search and ignored tables. `#purge!` is the run entry point; `#root_tables` are the tables it starts from. |
35
35
  | `plan_builder.rb` | The DSL. `instance_eval`s a plan file or block against a `Plan`; nested blocks get a new builder bound to that table's nested plan. |
36
36
  | `plan_validator.rb` | `ActiveModel::Validations` over plan vs. schema: missing tables, unknown tables, unknown columns. |
37
37
  | `executor.rb` | Convenience façade: loads a plan file, applies config options, `verify!`, `purge!`. |
38
38
  | `purge_table.rb` | Purges one table by `field = value(s)` in primary-key batches. Recurses into nested tables. |
39
39
  | `purge_table_scanner.rb` | Purges a `purge_table_search` table: full `find_in_batches` scan filtered through the user's `search_proc`. |
40
+ | `nullify_table.rb` | Sets a nullable column to `NULL` on rows pointing at a batch of ids about to be deleted (`ON DELETE SET NULL` at purge time). |
40
41
  | `purge_table_helper.rb` | Shared behaviour for both purgers: nested-table recursion, delete vs. soft delete vs. explain, transactions. |
41
42
  | `metrics.rb` / `metric_subscriber.rb` | Aggregate timing and row counts per table from the notification events. |
42
43
  | `dynamic_plan_builder.rb` | Generates plan-file source (`build` for a base table, `build_for` for several roots); a bootstrap tool, not used at purge time. |
@@ -54,6 +55,7 @@ Plan (root)
54
55
  └── nested Plan
55
56
  ├── parent_tables: [company_tags(:company_id)]
56
57
  ├── child_tables: [employments(:company_id) ─▶ nested Plan ..., websites(:id, fk: website_id) ...]
58
+ ├── nullify_tables: [companies(:acquired_by_company_id)]
57
59
  ├── search_tables: [users(:id)]
58
60
  └── ignore_tables
59
61
  ```
@@ -87,6 +89,8 @@ each_batch = loop:
87
89
  break if batch empty
88
90
 
89
91
  purge_children(batch) =
92
+ for each nullify_table: # optional rows pointing at us
93
+ UPDATE nullify SET field = NULL WHERE field IN (batch.pks) [AND conditions]
90
94
  for each child_table without foreign_key: # rows pointing at us
91
95
  PurgeTable(child, child.field, batch.pks).purge! (recursive)
92
96
 
@@ -118,6 +122,9 @@ Key properties:
118
122
 
119
123
  - **Depth-first, children first.** A row is only deleted after everything referencing it, so FK
120
124
  constraints hold without `ON DELETE CASCADE`.
125
+ - **Unlink before delete.** A `nullify_table` is updated for each batch before that batch's children and rows
126
+ are deleted, so a self-referential foreign key (or a reference from a row that must survive) never blocks the
127
+ delete, whatever the chain depth or id order across batches.
121
128
  - **Two passes when there are parent tables.** A parent table can reference this table (e.g.
122
129
  `company_tags.company_id → companies.id`) *and* be referenced by one of its children, so it is purged
123
130
  between the child pass and the delete pass. Tables without parent tables keep the single pass.
@@ -146,7 +153,7 @@ All three go through ActiveRecord's `*_all` methods: no model callbacks or valid
146
153
  ## Instrumentation
147
154
 
148
155
  Every unit of work is wrapped in `ActiveSupport::Notifications.instrument` under the `db_purger` namespace
149
- (`purge`, `next_batch`, `delete_records`, `search_filter`). The purgers never talk to `Metrics` directly;
156
+ (`purge`, `next_batch`, `delete_records`, `nullify_records`, `search_filter`). The purgers never talk to `Metrics` directly;
150
157
  `MetricSubscriber` (an `ActiveSupport::Subscriber`) translates events into `Metrics` counters. This keeps the
151
158
  purge code free of reporting concerns and lets callers attach their own subscribers (StatsD, logs, progress
152
159
  bars) without changes to the library.
data/README.md CHANGED
@@ -152,6 +152,7 @@ an error (there is no enclosing batch to take ids from). Existing `base_table` p
152
152
  | `child_table(table, field, opts = {}, &block)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch. Purged before that batch is deleted. |
153
153
  | `child_table(table, :id, foreign_key: :col, &block)` | Inverted relationship: the *enclosing* table holds `col` pointing at this table's `id`. Deleted in the same transaction, right after the enclosing batch. |
154
154
  | `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. At the top level of a plan without a `base_table`, each one is a root, purged in declaration order. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
155
+ | `nullify_table(table, field, conditions: nil)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch get `field = NULL` instead of being deleted, before anything in that batch is deleted (`ON DELETE SET NULL` at purge time). Use for optional references that must not take the referencing row down with them: a self-referential `parent_id`/"copied from" column, or a link from another tenant's row. Only `conditions:` is supported. |
155
156
  | `purge_table_search(table, field, opts = {}) { \|batch\| ... }` | Scans the whole table in batches; the block receives each batch and returns the records to purge. For orphans that can't be reached by a key. |
156
157
  | `ignore_table(name_or_regexp)` | Exclude a table from validation. |
157
158
 
@@ -224,7 +225,9 @@ Treat the output as a first draft: it cannot infer polymorphic (`as:`), soft-del
224
225
  - every field, `foreign_key` and `mark_deleted_field` named in the plan is a real column
225
226
  - the plan has a `base_table` or at least one top-level `parent_table`, no top-level `child_table` is left
226
227
  unreachable, and every `batch_size` is positive
227
- - tables without a primary key have no nested child or parent tables (there would be no ids to propagate)
228
+ - tables without a primary key have no nested child, parent or nullify tables (there would be no ids to propagate)
229
+ - every `nullify_table` field is a nullable column, and no `nullify_table` is left at the top level without a
230
+ `base_table`
228
231
 
229
232
  Run it in CI against your schema so a new table can't ship without a purge decision.
230
233
 
@@ -262,6 +265,7 @@ DBPurger::MetricSubscriber.metrics.as_json
262
265
  # => { took: 12.4, started_at: ..., finished_at: ...,
263
266
  # purge_stats: { employments: { duration:, num_purges:, num_records: } },
264
267
  # delete_stats: { employments: { duration:, num_delete_queries:, num_deleted:, num_expected_to_delete: } },
268
+ # nullify_stats: { cadences: { duration:, num_nullify_queries:, num_nullified: } },
265
269
  # lookup_stats: { ... }, filter_stats: { ... } }
266
270
  ```
267
271
 
@@ -273,6 +277,7 @@ Metrics are reset at the start of each `Plan#purge!`. To feed your own telemetry
273
277
  | `purge.db_purger` | `table_name`, `purge_field`, `deleted` |
274
278
  | `next_batch.db_purger` | `table_name`, `start_id`, `num_records` |
275
279
  | `delete_records.db_purger` | `table_name`, `num_records`, `records_deleted`, `deleted` |
280
+ | `nullify_records.db_purger` | `table_name`, `nullify_field`, `num_records`, `records_nullified` |
276
281
  | `search_filter.db_purger` | `table_name`, `num_records`, `num_records_selected` |
277
282
 
278
283
  ## Caveats
@@ -303,7 +308,7 @@ See [ARCHITECTURE.md](ARCHITECTURE.md) for how the pieces fit together.
303
308
  ## Releasing
304
309
 
305
310
  1. Bump `s.version` in `db-purger.gemspec` and merge to `master`.
306
- 2. Tag and push: `git tag v0.6.0 && git push origin v0.6.0`
311
+ 2. Tag and push: `git tag v0.7.0 && git push origin v0.7.0`
307
312
 
308
313
  `.github/workflows/release.yml` re-runs CI, checks the tag matches the gemspec version, publishes to RubyGems
309
314
  via trusted publishing (no API key), and creates a GitHub release with the `.gem` attached.
@@ -36,6 +36,14 @@ module DBPurger
36
36
  )
37
37
  end
38
38
 
39
+ def nullify_records(event)
40
+ self.class.metrics.update_nullify_records_stats(
41
+ event.payload[:table_name],
42
+ event.duration,
43
+ event.payload[:records_nullified] || 0
44
+ )
45
+ end
46
+
39
47
  def next_batch(event)
40
48
  self.class.metrics.update_lookup_stats(
41
49
  event.payload[:table_name],
@@ -7,6 +7,7 @@ module DBPurger
7
7
  :finished_at,
8
8
  :purge_stats,
9
9
  :delete_stats,
10
+ :nullify_stats,
10
11
  :lookup_stats,
11
12
  :filter_stats
12
13
 
@@ -18,6 +19,7 @@ module DBPurger
18
19
  @started_at = Time.now
19
20
  @purge_stats = {}
20
21
  @delete_stats = {}
22
+ @nullify_stats = {}
21
23
  @lookup_stats = {}
22
24
  @filter_stats = {}
23
25
  @finished_at = nil
@@ -48,6 +50,14 @@ module DBPurger
48
50
  stats
49
51
  end
50
52
 
53
+ def update_nullify_records_stats(table_name, duration, num_nullified)
54
+ stats = (@nullify_stats[table_name] ||= Hash.new(0))
55
+ stats[:duration] += duration
56
+ stats[:num_nullify_queries] += 1
57
+ stats[:num_nullified] += num_nullified
58
+ stats
59
+ end
60
+
51
61
  def update_lookup_stats(table_name, duration, records_found)
52
62
  stats = (@lookup_stats[table_name] ||= Hash.new(0))
53
63
  stats[:duration] += duration
@@ -72,6 +82,7 @@ module DBPurger
72
82
  finished_at: @finished_at,
73
83
  purge_stats: @purge_stats,
74
84
  delete_stats: @delete_stats,
85
+ nullify_stats: @nullify_stats,
75
86
  lookup_stats: @lookup_stats,
76
87
  filter_stats: @filter_stats
77
88
  }
@@ -0,0 +1,45 @@
1
+ # frozen_string_literal: true
2
+
3
+ module DBPurger
4
+ # DBPurger::NullifyTable clears a nullable column that points at a batch of rows about to be deleted,
5
+ # the purge-time equivalent of ON DELETE SET NULL. It updates the referencing rows instead of deleting them.
6
+ class NullifyTable
7
+ include PurgeTableHelper
8
+
9
+ def initialize(database, table, purge_values)
10
+ @database = database
11
+ @table = table
12
+ @purge_values = purge_values
13
+ end
14
+
15
+ def nullify!
16
+ ActiveSupport::Notifications.instrument('nullify_records.db_purger',
17
+ table_name: @table.name,
18
+ nullify_field: @table.field,
19
+ num_records: @purge_values.size) do |payload|
20
+ payload[:records_nullified] = ::DBPurger.config.explain? ? explain_nullify : nullify_records
21
+ end
22
+ end
23
+
24
+ private
25
+
26
+ def nullify_records
27
+ scope.update_all(@table.field => nil)
28
+ end
29
+
30
+ def scope
31
+ scope = model.where(@table.field => @purge_values)
32
+ scope = scope.where(@table.conditions) if @table.conditions
33
+ scope
34
+ end
35
+
36
+ def explain_nullify
37
+ ::DBPurger.config.explain_file.puts("#{explain_update_sql(scope, field_quoted, 'NULL')};")
38
+ scope.count
39
+ end
40
+
41
+ def field_quoted
42
+ model.connection.quote_column_name(@table.field)
43
+ end
44
+ end
45
+ end
@@ -8,18 +8,21 @@ module DBPurger
8
8
  attr_reader :parent_tables,
9
9
  :child_tables,
10
10
  :ignore_tables,
11
- :search_tables
11
+ :search_tables,
12
+ :nullify_tables
12
13
 
13
14
  def initialize
14
15
  @parent_tables = []
15
16
  @child_tables = []
16
17
  @ignore_tables = []
17
18
  @search_tables = []
19
+ @nullify_tables = []
18
20
  end
19
21
 
20
22
  def purge!(database, purge_value)
21
23
  raise('plan has no base_table or top-level parent_table') if root_tables.empty?
22
24
  raise('top-level child_tables require a base_table') unless @base_table || @child_tables.empty?
25
+ raise('top-level nullify_tables require a base_table') unless @base_table || @nullify_tables.empty?
23
26
 
24
27
  MetricSubscriber.reset!
25
28
  num_deleted = purge_root_tables(database, purge_value)
@@ -37,7 +40,8 @@ module DBPurger
37
40
  all_tables = @base_table ? [@base_table] + @base_table.tables : []
38
41
  all_tables += @parent_tables + @parent_tables.map(&:tables) +
39
42
  @child_tables + @child_tables.map(&:tables) +
40
- @search_tables + @search_tables.map(&:tables)
43
+ @search_tables + @search_tables.map(&:tables) +
44
+ @nullify_tables
41
45
  all_tables.flatten!
42
46
  all_tables.compact!
43
47
  all_tables
@@ -56,7 +60,8 @@ module DBPurger
56
60
  @base_table.nil? &&
57
61
  @parent_tables.empty? &&
58
62
  @child_tables.empty? &&
59
- @search_tables.empty?
63
+ @search_tables.empty? &&
64
+ @nullify_tables.empty?
60
65
  end
61
66
 
62
67
  def ignore_table?(table_name)
@@ -3,6 +3,8 @@
3
3
  module DBPurger
4
4
  # DBPurger::PlanBuilder is used to build the relationships between tables in a convenient way
5
5
  class PlanBuilder
6
+ NULLIFY_TABLE_OPTIONS = %i[conditions].freeze
7
+
6
8
  def initialize(plan)
7
9
  @plan = plan
8
10
  end
@@ -31,6 +33,22 @@ module DBPurger
31
33
  table
32
34
  end
33
35
 
36
+ # rows whose field matches the enclosing batch's primary keys get field set to NULL instead of being deleted
37
+ def nullify_table(table_name, field, options = {})
38
+ unsupported_options = options.keys - NULLIFY_TABLE_OPTIONS
39
+ unless unsupported_options.empty?
40
+ raise(ArgumentError, "nullify_table does not support #{unsupported_options.map(&:inspect).join(', ')}")
41
+ end
42
+
43
+ table = create_table(table_name, field, options)
44
+ if @plan.base_table
45
+ @plan.base_table.nested_plan.nullify_tables << table
46
+ else
47
+ @plan.nullify_tables << table
48
+ end
49
+ table
50
+ end
51
+
34
52
  def ignore_table(table_name)
35
53
  @plan.ignore_tables << table_name
36
54
  end
@@ -11,6 +11,7 @@ module DBPurger
11
11
  validate :validate_no_missing_tables
12
12
  validate :validate_no_unknown_tables
13
13
  validate :validate_tables
14
+ validate :validate_nullify_tables
14
15
 
15
16
  def initialize(database, plan)
16
17
  @database = database
@@ -36,6 +37,8 @@ module DBPurger
36
37
  elsif !@plan.child_tables.empty?
37
38
  # without a base_table there are no ids to propagate; declared before one, they are never reached
38
39
  errors.add(:base_table, 'must be declared before top-level child_tables')
40
+ elsif !@plan.nullify_tables.empty?
41
+ errors.add(:base_table, 'must be declared before top-level nullify_tables')
39
42
  end
40
43
  end
41
44
 
@@ -51,6 +54,27 @@ module DBPurger
51
54
  @plan.tables.each { |table| validate_table_definition(table) }
52
55
  end
53
56
 
57
+ # a NOT NULL column can't be unlinked, so the purge would fail at runtime instead of here
58
+ def validate_nullify_tables
59
+ nullify_tables.each do |table|
60
+ next unless (model = find_model_for_table(table)) # reported by validate_tables
61
+ next unless not_nullable_column?(model, table.field)
62
+
63
+ errors.add(:table, "#{table.name}.#{table.field} (nullify_table) is not nullable")
64
+ end
65
+ end
66
+
67
+ # a missing column is reported by validate_tables
68
+ def not_nullable_column?(model, field)
69
+ column = model.columns_hash[field.to_s]
70
+ column && !column.null
71
+ end
72
+
73
+ def nullify_tables
74
+ @plan.nullify_tables +
75
+ @plan.tables.select(&:nested_tables?).flat_map { |table| table.nested_plan.nullify_tables }
76
+ end
77
+
54
78
  def validate_table_definition(table)
55
79
  unless (model = find_model_for_table(table))
56
80
  errors.add(:table, "#{table.name} has no model")
@@ -10,9 +10,19 @@ module DBPurger
10
10
  private
11
11
 
12
12
  def purge_nested_tables(batch)
13
+ nullify_tables(batch) unless @table.nested_plan.nullify_tables.empty?
13
14
  purge_child_tables(batch) unless @table.nested_plan.child_tables.empty?
14
15
  end
15
16
 
17
+ # unlink rows that point at this batch before anything in it is deleted
18
+ def nullify_tables(batch)
19
+ ids = batch_values(batch, model.primary_key)
20
+
21
+ @table.nested_plan.nullify_tables.each do |table|
22
+ NullifyTable.new(@database, table, ids).nullify!
23
+ end
24
+ end
25
+
16
26
  def purge_child_tables(batch)
17
27
  ids = batch_values(batch, model.primary_key)
18
28
 
@@ -69,7 +79,7 @@ module DBPurger
69
79
  def explain(scope)
70
80
  sql =
71
81
  if @table.mark_deleted_field
72
- explain_update_sql(scope)
82
+ explain_update_sql(scope, mark_deleted_field_quoted, mark_deleted_value_quoted)
73
83
  else
74
84
  scope.to_sql.sub(/SELECT .*?FROM/, 'DELETE FROM')
75
85
  end
@@ -77,10 +87,10 @@ module DBPurger
77
87
  scope.count
78
88
  end
79
89
 
80
- def explain_update_sql(scope)
90
+ def explain_update_sql(scope, field_quoted, value_quoted)
81
91
  sql = scope.to_sql.dup
82
92
  sql.sub!(/SELECT .*?FROM/, 'UPDATE')
83
- sql.sub!('WHERE', "SET #{mark_deleted_field_quoted} = #{mark_deleted_value_quoted} WHERE")
93
+ sql.sub!('WHERE', "SET #{field_quoted} = #{value_quoted} WHERE")
84
94
  sql
85
95
  end
86
96
 
@@ -38,7 +38,8 @@ module DBPurger
38
38
 
39
39
  # nested tables that depend on this table's rows (search tables scan independently)
40
40
  def nested_key_tables?
41
- nested_tables? && !(@nested_plan.child_tables.empty? && @nested_plan.parent_tables.empty?)
41
+ nested_tables? &&
42
+ !(@nested_plan.child_tables.empty? && @nested_plan.parent_tables.empty? && @nested_plan.nullify_tables.empty?)
42
43
  end
43
44
 
44
45
  def tables
data/lib/db-purger.rb CHANGED
@@ -8,6 +8,7 @@ module DBPurger
8
8
  autoload :Executor, 'db-purger/executor'
9
9
  autoload :Metrics, 'db-purger/metrics'
10
10
  autoload :MetricSubscriber, 'db-purger/metric_subscriber'
11
+ autoload :NullifyTable, 'db-purger/nullify_table'
11
12
  autoload :PurgeTable, 'db-purger/purge_table'
12
13
  autoload :PurgeTableHelper, 'db-purger/purge_table_helper'
13
14
  autoload :PurgeTableScanner, 'db-purger/purge_table_scanner'
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: db-purger
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.6.0
4
+ version: 0.7.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Doug Youch
@@ -63,6 +63,7 @@ files:
63
63
  - lib/db-purger/executor.rb
64
64
  - lib/db-purger/metric_subscriber.rb
65
65
  - lib/db-purger/metrics.rb
66
+ - lib/db-purger/nullify_table.rb
66
67
  - lib/db-purger/plan.rb
67
68
  - lib/db-purger/plan_builder.rb
68
69
  - lib/db-purger/plan_validator.rb