db-purger 0.5.0 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/ARCHITECTURE.md +30 -4
- data/README.md +60 -13
- data/lib/db-purger/association_graph.rb +76 -0
- data/lib/db-purger/dynamic_plan_builder.rb +82 -73
- data/lib/db-purger/metric_subscriber.rb +8 -0
- data/lib/db-purger/metrics.rb +11 -0
- data/lib/db-purger/nullify_table.rb +45 -0
- data/lib/db-purger/plan.rb +34 -9
- data/lib/db-purger/plan_builder.rb +18 -0
- data/lib/db-purger/plan_validator.rb +31 -1
- data/lib/db-purger/plan_writer.rb +56 -0
- data/lib/db-purger/purge_table.rb +0 -4
- data/lib/db-purger/purge_table_helper.rb +13 -3
- data/lib/db-purger/purge_table_scanner.rb +0 -4
- data/lib/db-purger/table.rb +2 -1
- data/lib/db-purger.rb +3 -0
- metadata +12 -3
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 45ab4b4aa9e6e8c3778061c5deb8a8eb37fecf153bb1bae9cf298f2c83433f22
|
|
4
|
+
data.tar.gz: bbb9aec2aaa9dcd8e7c0609186795bf17e8fce64ffa951ce6dd1877981b7dbcc
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 41d99640be3f800277142f9ef9a9ec0f0945785282551c757f92ba92de36dfad331b735533713bb98385023378727ca5365cf995cff84580dfe725cbbf10240d
|
|
7
|
+
data.tar.gz: bfe3512b7f289cec1b4ac7f466918c4aaa6a5c12f9deb61e60e825153dd3e1badbbcbb1eb578e80dd453504f9b7b08f70a996e4d9d10b7d1e7a654245d12ec00
|
data/ARCHITECTURE.md
CHANGED
|
@@ -31,15 +31,18 @@ db-purger is small (~900 lines) and splits cleanly into three layers: **describe
|
|
|
31
31
|
| `lib/db-purger.rb` | Autoloads everything; holds the global `DBPurger.config`. |
|
|
32
32
|
| `config.rb` | Global options: `explain?`, `explain_file`, `datetime_format`. |
|
|
33
33
|
| `table.rb` | Value object for one table in the plan: name, match field, options, and a lazily created nested `Plan`. |
|
|
34
|
-
| `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, search and ignored tables. `#purge!` is the run entry point. |
|
|
34
|
+
| `plan.rb` | A node in the plan tree: one optional `base_table` plus lists of parent, child, nullify, search and ignored tables. `#purge!` is the run entry point; `#root_tables` are the tables it starts from. |
|
|
35
35
|
| `plan_builder.rb` | The DSL. `instance_eval`s a plan file or block against a `Plan`; nested blocks get a new builder bound to that table's nested plan. |
|
|
36
36
|
| `plan_validator.rb` | `ActiveModel::Validations` over plan vs. schema: missing tables, unknown tables, unknown columns. |
|
|
37
37
|
| `executor.rb` | Convenience façade: loads a plan file, applies config options, `verify!`, `purge!`. |
|
|
38
38
|
| `purge_table.rb` | Purges one table by `field = value(s)` in primary-key batches. Recurses into nested tables. |
|
|
39
39
|
| `purge_table_scanner.rb` | Purges a `purge_table_search` table: full `find_in_batches` scan filtered through the user's `search_proc`. |
|
|
40
|
+
| `nullify_table.rb` | Sets a nullable column to `NULL` on rows pointing at a batch of ids about to be deleted (`ON DELETE SET NULL` at purge time). |
|
|
40
41
|
| `purge_table_helper.rb` | Shared behaviour for both purgers: nested-table recursion, delete vs. soft delete vs. explain, transactions. |
|
|
41
42
|
| `metrics.rb` / `metric_subscriber.rb` | Aggregate timing and row counts per table from the notification events. |
|
|
42
|
-
| `dynamic_plan_builder.rb` | Generates plan-file source
|
|
43
|
+
| `dynamic_plan_builder.rb` | Generates plan-file source (`build` for a base table, `build_for` for several roots); a bootstrap tool, not used at purge time. |
|
|
44
|
+
| `association_graph.rb` | For the generator: which tables reference a model, and by which column, from its `has_many`/`has_one`/HABTM reflections. |
|
|
45
|
+
| `plan_writer.rb` | For the generator: renders plan DSL text and records which tables it wrote. |
|
|
43
46
|
|
|
44
47
|
## The plan tree
|
|
45
48
|
|
|
@@ -52,16 +55,30 @@ Plan (root)
|
|
|
52
55
|
└── nested Plan
|
|
53
56
|
├── parent_tables: [company_tags(:company_id)]
|
|
54
57
|
├── child_tables: [employments(:company_id) ─▶ nested Plan ..., websites(:id, fk: website_id) ...]
|
|
58
|
+
├── nullify_tables: [companies(:acquired_by_company_id)]
|
|
55
59
|
├── search_tables: [users(:id)]
|
|
56
60
|
└── ignore_tables
|
|
57
61
|
```
|
|
58
62
|
|
|
63
|
+
Without a `base_table`, top-level `parent_table`s stay in the root plan's `parent_tables` and each one is a
|
|
64
|
+
root:
|
|
65
|
+
|
|
66
|
+
```
|
|
67
|
+
Plan (root)
|
|
68
|
+
├── parent_tables: [calls(:oid) ─▶ nested Plan ..., emails(:oid) ─▶ nested Plan ..., sms_messages(:oid) ...]
|
|
69
|
+
├── search_tables
|
|
70
|
+
└── ignore_tables
|
|
71
|
+
```
|
|
72
|
+
|
|
59
73
|
`Plan#tables` flattens this tree for validation; `Table#foreign_keys` collects the `foreign_key:` columns of
|
|
60
74
|
a table's direct children so the purger can `SELECT` them alongside the primary key.
|
|
61
75
|
|
|
62
76
|
## Purge algorithm
|
|
63
77
|
|
|
64
|
-
`Plan#purge!` resets metrics and starts a `PurgeTable`
|
|
78
|
+
`Plan#purge!` resets metrics and starts a `PurgeTable` with the purge value on each root table
|
|
79
|
+
(`root_tables` = the `base_table`, if any, followed by top-level `parent_tables`, in declaration order), then runs
|
|
80
|
+
any top-level search tables. With a `base_table`, `root_tables` is just `[base_table]`, which is the original
|
|
81
|
+
single-root algorithm. Each
|
|
65
82
|
`PurgeTable#purge!` does:
|
|
66
83
|
|
|
67
84
|
```
|
|
@@ -72,6 +89,8 @@ each_batch = loop:
|
|
|
72
89
|
break if batch empty
|
|
73
90
|
|
|
74
91
|
purge_children(batch) =
|
|
92
|
+
for each nullify_table: # optional rows pointing at us
|
|
93
|
+
UPDATE nullify SET field = NULL WHERE field IN (batch.pks) [AND conditions]
|
|
75
94
|
for each child_table without foreign_key: # rows pointing at us
|
|
76
95
|
PurgeTable(child, child.field, batch.pks).purge! (recursive)
|
|
77
96
|
|
|
@@ -103,6 +122,9 @@ Key properties:
|
|
|
103
122
|
|
|
104
123
|
- **Depth-first, children first.** A row is only deleted after everything referencing it, so FK
|
|
105
124
|
constraints hold without `ON DELETE CASCADE`.
|
|
125
|
+
- **Unlink before delete.** A `nullify_table` is updated for each batch before that batch's children and rows
|
|
126
|
+
are deleted, so a self-referential foreign key (or a reference from a row that must survive) never blocks the
|
|
127
|
+
delete, whatever the chain depth or id order across batches.
|
|
106
128
|
- **Two passes when there are parent tables.** A parent table can reference this table (e.g.
|
|
107
129
|
`company_tags.company_id → companies.id`) *and* be referenced by one of its children, so it is purged
|
|
108
130
|
between the child pass and the delete pass. Tables without parent tables keep the single pass.
|
|
@@ -131,7 +153,7 @@ All three go through ActiveRecord's `*_all` methods: no model callbacks or valid
|
|
|
131
153
|
## Instrumentation
|
|
132
154
|
|
|
133
155
|
Every unit of work is wrapped in `ActiveSupport::Notifications.instrument` under the `db_purger` namespace
|
|
134
|
-
(`purge`, `next_batch`, `delete_records`, `search_filter`). The purgers never talk to `Metrics` directly;
|
|
156
|
+
(`purge`, `next_batch`, `delete_records`, `nullify_records`, `search_filter`). The purgers never talk to `Metrics` directly;
|
|
135
157
|
`MetricSubscriber` (an `ActiveSupport::Subscriber`) translates events into `Metrics` counters. This keeps the
|
|
136
158
|
purge code free of reporting concerns and lets callers attach their own subscribers (StatsD, logs, progress
|
|
137
159
|
bars) without changes to the library.
|
|
@@ -157,4 +179,8 @@ fetch happens inside `find_in_batches` rather than in a block the scanner contro
|
|
|
157
179
|
- `spec/integrations/*` — end-to-end purges over the schema in `spec/support/db/schema.rb`, asserting row-count
|
|
158
180
|
deltas per table and, for explain mode, the exact SQL in `spec/fixtures/delete_plan.sql`.
|
|
159
181
|
- `spec/support/test_db.rb` builds a fresh SQLite database and dynamic models (`TestDB::*`) for each run.
|
|
182
|
+
- `spec/integrations/multi_root_plan_spec.rb` purges one org from a separate outreach schema (`spec/support/outreach_db.rb`,
|
|
183
|
+
real FOREIGN KEY constraints) seeded with two orgs by `OutreachSeeder`, and asserts the exact surviving rows of
|
|
184
|
+
every table for the generated plan, a hand-written small-batch plan and the equivalent `base_table` plan.
|
|
185
|
+
- `spec/support/throwaway_db.rb` builds standalone SQLite databases from raw SQL for schema-specific specs.
|
|
160
186
|
- `spec/fixtures/*.plan.rb` — plan files used for loading and validation cases.
|
data/README.md
CHANGED
|
@@ -2,6 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
|
|
4
4
|
[](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
|
|
5
|
+
[](https://github.com/dougyouch/db-purger/actions/workflows/ci.yml)
|
|
5
6
|
[](https://rubygems.org/gems/db-purger)
|
|
6
7
|
|
|
7
8
|
Purge every row tied to a single top-level record — a company, an account, a tenant — across all of the
|
|
@@ -112,16 +113,46 @@ executor.verify! # raises 'purge plan failed verification', errors pr
|
|
|
112
113
|
deleted = executor.purge!(42)
|
|
113
114
|
```
|
|
114
115
|
|
|
115
|
-
`purge!` returns the number of
|
|
116
|
+
`purge!` returns the number of root-table rows deleted (the base table, plus any top-level `parent_table`s).
|
|
117
|
+
|
|
118
|
+
### Plans with several top-level tables
|
|
119
|
+
|
|
120
|
+
`base_table` is shorthand for "one root table, with everything after it nested underneath". When several tables
|
|
121
|
+
are equally top-level (an outreach product's `emails`, `sms_messages` and `calls`, all keyed by `oid`), leave
|
|
122
|
+
`base_table` out and declare each root as a top-level `parent_table`:
|
|
123
|
+
|
|
124
|
+
```ruby
|
|
125
|
+
# config/outreach.plan.rb
|
|
126
|
+
parent_table(:calls, :oid) do # calls.email_id -> emails.id, so calls go first
|
|
127
|
+
child_table(:call_notes, :call_id)
|
|
128
|
+
child_table(:call_recordings, :call_id)
|
|
129
|
+
child_table(:call_tags, :call_id)
|
|
130
|
+
end
|
|
131
|
+
|
|
132
|
+
parent_table(:emails, :oid) do
|
|
133
|
+
child_table(:email_attachments, :email_id)
|
|
134
|
+
end
|
|
135
|
+
|
|
136
|
+
parent_table(:sms_messages, :oid) do
|
|
137
|
+
child_table(:sms_deliveries, :sms_message_id)
|
|
138
|
+
end
|
|
139
|
+
|
|
140
|
+
ignore_table :users
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
Each root is purged by `oid = purge_value`, children first, **in declaration order**: when one root's rows
|
|
144
|
+
reference another's, declare the referencing root first. Without a `base_table`, top-level `child_table`s are
|
|
145
|
+
an error (there is no enclosing batch to take ids from). Existing `base_table` plans run exactly as before.
|
|
116
146
|
|
|
117
147
|
## The plan DSL
|
|
118
148
|
|
|
119
149
|
| Method | Meaning |
|
|
120
150
|
|---|---|
|
|
121
|
-
| `base_table(table, field, opts = {}, &block)` |
|
|
151
|
+
| `base_table(table, field, opts = {}, &block)` | Optional single root. Rows where `field = purge_value` are purged. Declare it **first** — every subsequent top-level call nests under it. |
|
|
122
152
|
| `child_table(table, field, opts = {}, &block)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch. Purged before that batch is deleted. |
|
|
123
153
|
| `child_table(table, :id, foreign_key: :col, &block)` | Inverted relationship: the *enclosing* table holds `col` pointing at this table's `id`. Deleted in the same transaction, right after the enclosing batch. |
|
|
124
|
-
| `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
|
|
154
|
+
| `parent_table(table, field, opts = {}, &block)` | Rows whose `field` matches the original **purge value**. At the top level of a plan without a `base_table`, each one is a root, purged in declaration order. Purged after the enclosing table's child tables but before the enclosing table's own rows, so it may both reference the base (`company_tags.company_id → companies.id`) and be referenced by a child table. Use for sibling tables that share the same key (e.g. `company_id`). |
|
|
155
|
+
| `nullify_table(table, field, conditions: nil)` | Rows whose `field` matches the **primary key** of the enclosing table's current batch get `field = NULL` instead of being deleted, before anything in that batch is deleted (`ON DELETE SET NULL` at purge time). Use for optional references that must not take the referencing row down with them: a self-referential `parent_id`/"copied from" column, or a link from another tenant's row. Only `conditions:` is supported. |
|
|
125
156
|
| `purge_table_search(table, field, opts = {}) { \|batch\| ... }` | Scans the whole table in batches; the block receives each batch and returns the records to purge. For orphans that can't be reached by a key. |
|
|
126
157
|
| `ignore_table(name_or_regexp)` | Exclude a table from validation. |
|
|
127
158
|
|
|
@@ -166,14 +197,25 @@ plan.purge!(database, 42)
|
|
|
166
197
|
|
|
167
198
|
### Generating a starting plan
|
|
168
199
|
|
|
169
|
-
`DynamicPlanBuilder` walks the `has_many`
|
|
170
|
-
|
|
171
|
-
|
|
200
|
+
`DynamicPlanBuilder` walks the `has_many`, `has_one` and `has_and_belongs_to_many` associations
|
|
201
|
+
dynamic-active-model discovered, using each association's real foreign key, and emits a plan file listing every
|
|
202
|
+
unreachable table as `ignore_table`.
|
|
172
203
|
|
|
173
204
|
```ruby
|
|
174
|
-
|
|
205
|
+
builder = DBPurger::DynamicPlanBuilder.new(database)
|
|
206
|
+
puts builder.build(:companies, :id) # single base_table plan
|
|
207
|
+
puts builder.build_for(:oid) # one top-level parent_table per table holding oid
|
|
175
208
|
```
|
|
176
209
|
|
|
210
|
+
`build_for` makes **every** table holding the field a root, so rows with a null foreign key to another root
|
|
211
|
+
(an `email_recipients` row without an email) are still purged, and orders the roots so a root referencing
|
|
212
|
+
another root's rows comes first. HABTM join tables are emitted as leaves, never walking into the shared table on
|
|
213
|
+
the other side, and nothing is nested under a table without a primary key. A foreign-key cycle is written as a
|
|
214
|
+
comment instead of recursing.
|
|
215
|
+
|
|
216
|
+
Treat the output as a first draft: it cannot infer polymorphic (`as:`), soft-delete, `belongs_to`-owned
|
|
217
|
+
(`foreign_key:`) or search rules.
|
|
218
|
+
|
|
177
219
|
## Validation
|
|
178
220
|
|
|
179
221
|
`Executor#verify!` (or `DBPurger::PlanValidator.new(database, plan).valid?`) checks that:
|
|
@@ -181,8 +223,11 @@ puts DBPurger::DynamicPlanBuilder.new(database).build(:companies, :id)
|
|
|
181
223
|
- every table in the database is either in the plan or ignored (`missing_tables`)
|
|
182
224
|
- every table in the plan exists in the database (`unknown_tables`)
|
|
183
225
|
- every field, `foreign_key` and `mark_deleted_field` named in the plan is a real column
|
|
184
|
-
- the plan has a `base_table`,
|
|
185
|
-
|
|
226
|
+
- the plan has a `base_table` or at least one top-level `parent_table`, no top-level `child_table` is left
|
|
227
|
+
unreachable, and every `batch_size` is positive
|
|
228
|
+
- tables without a primary key have no nested child, parent or nullify tables (there would be no ids to propagate)
|
|
229
|
+
- every `nullify_table` field is a nullable column, and no `nullify_table` is left at the top level without a
|
|
230
|
+
`base_table`
|
|
186
231
|
|
|
187
232
|
Run it in CI against your schema so a new table can't ship without a purge decision.
|
|
188
233
|
|
|
@@ -220,6 +265,7 @@ DBPurger::MetricSubscriber.metrics.as_json
|
|
|
220
265
|
# => { took: 12.4, started_at: ..., finished_at: ...,
|
|
221
266
|
# purge_stats: { employments: { duration:, num_purges:, num_records: } },
|
|
222
267
|
# delete_stats: { employments: { duration:, num_delete_queries:, num_deleted:, num_expected_to_delete: } },
|
|
268
|
+
# nullify_stats: { cadences: { duration:, num_nullify_queries:, num_nullified: } },
|
|
223
269
|
# lookup_stats: { ... }, filter_stats: { ... } }
|
|
224
270
|
```
|
|
225
271
|
|
|
@@ -231,12 +277,13 @@ Metrics are reset at the start of each `Plan#purge!`. To feed your own telemetry
|
|
|
231
277
|
| `purge.db_purger` | `table_name`, `purge_field`, `deleted` |
|
|
232
278
|
| `next_batch.db_purger` | `table_name`, `start_id`, `num_records` |
|
|
233
279
|
| `delete_records.db_purger` | `table_name`, `num_records`, `records_deleted`, `deleted` |
|
|
280
|
+
| `nullify_records.db_purger` | `table_name`, `nullify_field`, `num_records`, `records_nullified` |
|
|
234
281
|
| `search_filter.db_purger` | `table_name`, `num_records`, `num_records_selected` |
|
|
235
282
|
|
|
236
283
|
## Caveats
|
|
237
284
|
|
|
238
|
-
- **
|
|
239
|
-
`
|
|
285
|
+
- **Declare `base_table` first.** A `child_table` declared before it is never reached (the validator reports
|
|
286
|
+
it); a `parent_table` declared before it becomes a separate root, purged after the base table.
|
|
240
287
|
- **Not one big transaction.** Each batch is its own set of statements (foreign-key children share a
|
|
241
288
|
transaction with their parent batch). An interrupted purge is safe to re-run with the same value.
|
|
242
289
|
- **Soft-deleted rows still match.** A `mark_deleted_field` table is not filtered on that field; add
|
|
@@ -254,14 +301,14 @@ script/console
|
|
|
254
301
|
|
|
255
302
|
CI (`.github/workflows/ci.yml`) runs RuboCop and the specs on Ruby 4.0 for every push and pull request.
|
|
256
303
|
The HTML coverage report is attached to each run as the `coverage` artifact, and pushes to `master` refresh
|
|
257
|
-
the coverage
|
|
304
|
+
the line and branch coverage badges on the `badges` branch.
|
|
258
305
|
|
|
259
306
|
See [ARCHITECTURE.md](ARCHITECTURE.md) for how the pieces fit together.
|
|
260
307
|
|
|
261
308
|
## Releasing
|
|
262
309
|
|
|
263
310
|
1. Bump `s.version` in `db-purger.gemspec` and merge to `master`.
|
|
264
|
-
2. Tag and push: `git tag v0.
|
|
311
|
+
2. Tag and push: `git tag v0.7.0 && git push origin v0.7.0`
|
|
265
312
|
|
|
266
313
|
`.github/workflows/release.yml` re-runs CI, checks the tag matches the gemspec version, publishes to RubyGems
|
|
267
314
|
via trusted publishing (no API key), and creates a GitHub release with the `.gem` attached.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module DBPurger
|
|
4
|
+
# DBPurger::AssociationGraph answers "which tables reference this model, and by which column" from the
|
|
5
|
+
# has_many, has_one and has_and_belongs_to_many associations dynamic-active-model discovered
|
|
6
|
+
class AssociationGraph
|
|
7
|
+
# a table holding foreign_key that points at the parent model's primary key
|
|
8
|
+
Edge = Struct.new(:model, :foreign_key)
|
|
9
|
+
|
|
10
|
+
def initialize(database)
|
|
11
|
+
@database = database
|
|
12
|
+
@edges = {}
|
|
13
|
+
end
|
|
14
|
+
|
|
15
|
+
# database.models order depends on how the adapter lists tables, which varies by platform;
|
|
16
|
+
# sort so generated plans are deterministic
|
|
17
|
+
def models
|
|
18
|
+
@models ||= @database.models.sort_by(&:table_name)
|
|
19
|
+
end
|
|
20
|
+
|
|
21
|
+
def model_for(table_name)
|
|
22
|
+
models.detect { |model| model.table_name == table_name.to_s }
|
|
23
|
+
end
|
|
24
|
+
|
|
25
|
+
# one edge per (table, foreign key); a habtm join table also reached by a has_many appears once
|
|
26
|
+
def edges(model)
|
|
27
|
+
@edges[model] ||= model.reflect_on_all_associations
|
|
28
|
+
.filter_map { |reflection| edge_for(reflection) }
|
|
29
|
+
.uniq { |edge| edge_key(edge) }
|
|
30
|
+
.sort_by { |edge| edge_key(edge) }
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
# every model reachable from model through edges, excluding model itself unless there is a cycle
|
|
34
|
+
def reachable(model, seen = Set.new)
|
|
35
|
+
edges(model).each do |edge|
|
|
36
|
+
next if seen.include?(edge.model)
|
|
37
|
+
|
|
38
|
+
seen << edge.model
|
|
39
|
+
reachable(edge.model, seen)
|
|
40
|
+
end
|
|
41
|
+
seen
|
|
42
|
+
end
|
|
43
|
+
|
|
44
|
+
def column?(model, field)
|
|
45
|
+
model.column_names.include?(field.to_s)
|
|
46
|
+
end
|
|
47
|
+
|
|
48
|
+
private
|
|
49
|
+
|
|
50
|
+
def edge_for(reflection)
|
|
51
|
+
return if skip_reflection?(reflection)
|
|
52
|
+
|
|
53
|
+
case reflection
|
|
54
|
+
when ActiveRecord::Reflection::HasAndBelongsToManyReflection
|
|
55
|
+
join_table_edge(reflection)
|
|
56
|
+
when ActiveRecord::Reflection::HasManyReflection, ActiveRecord::Reflection::HasOneReflection
|
|
57
|
+
Edge.new(reflection.klass, reflection.foreign_key.to_s)
|
|
58
|
+
end
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
# through associations are reached via their own direct associations; polymorphic (as:) ones need a
|
|
62
|
+
# type condition the generator cannot infer
|
|
63
|
+
def skip_reflection?(reflection)
|
|
64
|
+
reflection.options[:through] || reflection.options[:as]
|
|
65
|
+
end
|
|
66
|
+
|
|
67
|
+
def join_table_edge(reflection)
|
|
68
|
+
join_model = model_for(reflection.join_table)
|
|
69
|
+
Edge.new(join_model, reflection.foreign_key.to_s) if join_model
|
|
70
|
+
end
|
|
71
|
+
|
|
72
|
+
def edge_key(edge)
|
|
73
|
+
[edge.model.table_name, edge.foreign_key]
|
|
74
|
+
end
|
|
75
|
+
end
|
|
76
|
+
end
|
|
@@ -3,115 +3,124 @@
|
|
|
3
3
|
module DBPurger
|
|
4
4
|
# DBPurger::DynamicPlanBuilder generates a purge plan based on the database relations
|
|
5
5
|
class DynamicPlanBuilder
|
|
6
|
-
INDENT = ' '
|
|
7
|
-
|
|
8
|
-
attr_reader :output
|
|
9
|
-
|
|
10
6
|
def initialize(database)
|
|
11
|
-
@
|
|
12
|
-
@
|
|
13
|
-
|
|
14
|
-
|
|
7
|
+
@graph = AssociationGraph.new(database)
|
|
8
|
+
@writer = PlanWriter.new
|
|
9
|
+
end
|
|
10
|
+
|
|
11
|
+
def output
|
|
12
|
+
@writer.output
|
|
15
13
|
end
|
|
16
14
|
|
|
15
|
+
# plan rooted at a single base table
|
|
17
16
|
def build(base_table_name, field)
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
foreign_key = foreign_key_name(model)
|
|
17
|
+
model = @graph.model_for(base_table_name)
|
|
18
|
+
@writer.table('base', model.table_name, field)
|
|
19
|
+
@writer.line_break
|
|
22
20
|
if model.primary_key == field.to_s
|
|
23
|
-
|
|
21
|
+
add_referencing_parent_tables(model)
|
|
24
22
|
else
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
line_break unless field == :id
|
|
28
|
-
add_child_tables(child_models, foreign_key, 0)
|
|
29
|
-
end
|
|
23
|
+
add_sibling_parent_tables(model, field)
|
|
24
|
+
add_base_child_tables(model)
|
|
30
25
|
end
|
|
31
|
-
|
|
32
|
-
|
|
26
|
+
finish
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
# plan with one top-level parent_table per table holding field (e.g. :oid), each purged by field
|
|
30
|
+
# directly so rows with a null foreign key are not missed; the tables referencing each root are nested
|
|
31
|
+
# under it, and roots are ordered so a root referencing another root's rows is purged first
|
|
32
|
+
def build_for(field)
|
|
33
|
+
@root_field = field.to_s
|
|
34
|
+
ordered_root_models.each_with_index do |model, idx|
|
|
35
|
+
@writer.line_break if idx.positive?
|
|
36
|
+
write_table('parent', model, field, [])
|
|
37
|
+
end
|
|
38
|
+
finish
|
|
33
39
|
end
|
|
34
40
|
|
|
35
41
|
private
|
|
36
42
|
|
|
37
|
-
def
|
|
38
|
-
@
|
|
43
|
+
def finish
|
|
44
|
+
@writer.ignore_tables(@graph.models.map(&:table_name) - @writer.table_names)
|
|
45
|
+
output
|
|
39
46
|
end
|
|
40
47
|
|
|
41
|
-
|
|
42
|
-
|
|
48
|
+
# base_table(:companies, :id): tables holding companies.id are keyed directly on the purge value
|
|
49
|
+
def add_referencing_parent_tables(model)
|
|
50
|
+
@graph.edges(model).each { |edge| write_edge('parent', edge, [model]) }
|
|
43
51
|
end
|
|
44
52
|
|
|
45
|
-
|
|
46
|
-
|
|
53
|
+
# base_table(:employments, :company_id): other tables holding company_id share the purge value
|
|
54
|
+
def add_sibling_parent_tables(model, field)
|
|
55
|
+
@graph.models.each do |sibling|
|
|
56
|
+
next if sibling == model || !@graph.column?(sibling, field)
|
|
57
|
+
|
|
58
|
+
write_table('parent', sibling, field, [model])
|
|
59
|
+
end
|
|
47
60
|
end
|
|
48
61
|
|
|
49
|
-
def
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
next unless column?(model, field)
|
|
62
|
+
def add_base_child_tables(model)
|
|
63
|
+
edges = nestable_edges(model)
|
|
64
|
+
return if edges.empty?
|
|
53
65
|
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
end
|
|
66
|
+
@writer.line_break
|
|
67
|
+
edges.each { |edge| write_edge('child', edge, [model]) }
|
|
57
68
|
end
|
|
58
69
|
|
|
59
|
-
def
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
70
|
+
def write_edge(table_type, edge, ancestors)
|
|
71
|
+
if ancestors.include?(edge.model)
|
|
72
|
+
@writer.comment("#{table_type}_table(#{edge.model.table_name.to_sym.inspect}, " \
|
|
73
|
+
"#{edge.foreign_key.to_sym.inspect}) skipped: cycle back to #{edge.model.table_name}")
|
|
74
|
+
else
|
|
75
|
+
write_table(table_type, edge.model, edge.foreign_key, ancestors)
|
|
63
76
|
end
|
|
64
|
-
@indent_depth -= change_indent_by
|
|
65
77
|
end
|
|
66
78
|
|
|
67
|
-
def
|
|
68
|
-
|
|
69
|
-
|
|
79
|
+
def write_table(table_type, model, field, ancestors)
|
|
80
|
+
edges = nestable_edges(model)
|
|
81
|
+
if edges.empty?
|
|
82
|
+
@writer.table(table_type, model.table_name, field)
|
|
83
|
+
warn_unnestable(model)
|
|
84
|
+
else
|
|
85
|
+
@writer.table_block(table_type, model.table_name, field) do
|
|
86
|
+
edges.each { |edge| write_edge('child', edge, ancestors + [model]) }
|
|
87
|
+
end
|
|
88
|
+
end
|
|
70
89
|
end
|
|
71
90
|
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
91
|
+
# purging nested tables needs this table's primary keys to propagate
|
|
92
|
+
def nestable_edges(model)
|
|
93
|
+
return [] unless model.primary_key
|
|
75
94
|
|
|
76
|
-
|
|
77
|
-
# sort so the generated plan is deterministic
|
|
78
|
-
def sorted_models
|
|
79
|
-
@sorted_models ||= @database.models.sort_by(&:table_name)
|
|
95
|
+
@graph.edges(model).reject { |edge| root_model?(edge.model) }
|
|
80
96
|
end
|
|
81
97
|
|
|
82
|
-
def
|
|
83
|
-
model
|
|
84
|
-
assoc.is_a?(ActiveRecord::Reflection::HasManyReflection)
|
|
85
|
-
end
|
|
98
|
+
def root_model?(model)
|
|
99
|
+
@root_field && @graph.column?(model, @root_field)
|
|
86
100
|
end
|
|
87
101
|
|
|
88
|
-
def
|
|
89
|
-
|
|
90
|
-
end
|
|
102
|
+
def warn_unnestable(model)
|
|
103
|
+
return if model.primary_key || (edges = @graph.edges(model)).empty?
|
|
91
104
|
|
|
92
|
-
|
|
93
|
-
|
|
105
|
+
@writer.comment("#{model.table_name} has no primary key; cannot nest " \
|
|
106
|
+
"#{edges.map { |edge| edge.model.table_name }.join(', ')}")
|
|
94
107
|
end
|
|
95
108
|
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
109
|
+
# repeatedly take the first root (by name) whose referencing roots are already written; on a cycle,
|
|
110
|
+
# fall back to name order so no root is dropped
|
|
111
|
+
def ordered_root_models
|
|
112
|
+
remaining = @graph.models.select { |model| root_model?(model) }
|
|
113
|
+
ordered = []
|
|
114
|
+
until remaining.empty?
|
|
115
|
+
model = remaining.detect { |root| (purged_first(root) & remaining).empty? } || remaining.first
|
|
116
|
+
ordered << remaining.delete(model)
|
|
104
117
|
end
|
|
118
|
+
ordered
|
|
105
119
|
end
|
|
106
120
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
line_break
|
|
112
|
-
missing_tables.each do |table_name|
|
|
113
|
-
write("ignore_table #{table_name.to_sym.inspect}")
|
|
114
|
-
end
|
|
121
|
+
# roots holding rows that reference root's rows (directly or through nested tables)
|
|
122
|
+
def purged_first(root)
|
|
123
|
+
@graph.reachable(root).select { |model| model != root && root_model?(model) }
|
|
115
124
|
end
|
|
116
125
|
end
|
|
117
126
|
end
|
|
@@ -36,6 +36,14 @@ module DBPurger
|
|
|
36
36
|
)
|
|
37
37
|
end
|
|
38
38
|
|
|
39
|
+
def nullify_records(event)
|
|
40
|
+
self.class.metrics.update_nullify_records_stats(
|
|
41
|
+
event.payload[:table_name],
|
|
42
|
+
event.duration,
|
|
43
|
+
event.payload[:records_nullified] || 0
|
|
44
|
+
)
|
|
45
|
+
end
|
|
46
|
+
|
|
39
47
|
def next_batch(event)
|
|
40
48
|
self.class.metrics.update_lookup_stats(
|
|
41
49
|
event.payload[:table_name],
|
data/lib/db-purger/metrics.rb
CHANGED
|
@@ -7,6 +7,7 @@ module DBPurger
|
|
|
7
7
|
:finished_at,
|
|
8
8
|
:purge_stats,
|
|
9
9
|
:delete_stats,
|
|
10
|
+
:nullify_stats,
|
|
10
11
|
:lookup_stats,
|
|
11
12
|
:filter_stats
|
|
12
13
|
|
|
@@ -18,6 +19,7 @@ module DBPurger
|
|
|
18
19
|
@started_at = Time.now
|
|
19
20
|
@purge_stats = {}
|
|
20
21
|
@delete_stats = {}
|
|
22
|
+
@nullify_stats = {}
|
|
21
23
|
@lookup_stats = {}
|
|
22
24
|
@filter_stats = {}
|
|
23
25
|
@finished_at = nil
|
|
@@ -48,6 +50,14 @@ module DBPurger
|
|
|
48
50
|
stats
|
|
49
51
|
end
|
|
50
52
|
|
|
53
|
+
def update_nullify_records_stats(table_name, duration, num_nullified)
|
|
54
|
+
stats = (@nullify_stats[table_name] ||= Hash.new(0))
|
|
55
|
+
stats[:duration] += duration
|
|
56
|
+
stats[:num_nullify_queries] += 1
|
|
57
|
+
stats[:num_nullified] += num_nullified
|
|
58
|
+
stats
|
|
59
|
+
end
|
|
60
|
+
|
|
51
61
|
def update_lookup_stats(table_name, duration, records_found)
|
|
52
62
|
stats = (@lookup_stats[table_name] ||= Hash.new(0))
|
|
53
63
|
stats[:duration] += duration
|
|
@@ -72,6 +82,7 @@ module DBPurger
|
|
|
72
82
|
finished_at: @finished_at,
|
|
73
83
|
purge_stats: @purge_stats,
|
|
74
84
|
delete_stats: @delete_stats,
|
|
85
|
+
nullify_stats: @nullify_stats,
|
|
75
86
|
lookup_stats: @lookup_stats,
|
|
76
87
|
filter_stats: @filter_stats
|
|
77
88
|
}
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module DBPurger
|
|
4
|
+
# DBPurger::NullifyTable clears a nullable column that points at a batch of rows about to be deleted,
|
|
5
|
+
# the purge-time equivalent of ON DELETE SET NULL. It updates the referencing rows instead of deleting them.
|
|
6
|
+
class NullifyTable
|
|
7
|
+
include PurgeTableHelper
|
|
8
|
+
|
|
9
|
+
def initialize(database, table, purge_values)
|
|
10
|
+
@database = database
|
|
11
|
+
@table = table
|
|
12
|
+
@purge_values = purge_values
|
|
13
|
+
end
|
|
14
|
+
|
|
15
|
+
def nullify!
|
|
16
|
+
ActiveSupport::Notifications.instrument('nullify_records.db_purger',
|
|
17
|
+
table_name: @table.name,
|
|
18
|
+
nullify_field: @table.field,
|
|
19
|
+
num_records: @purge_values.size) do |payload|
|
|
20
|
+
payload[:records_nullified] = ::DBPurger.config.explain? ? explain_nullify : nullify_records
|
|
21
|
+
end
|
|
22
|
+
end
|
|
23
|
+
|
|
24
|
+
private
|
|
25
|
+
|
|
26
|
+
def nullify_records
|
|
27
|
+
scope.update_all(@table.field => nil)
|
|
28
|
+
end
|
|
29
|
+
|
|
30
|
+
def scope
|
|
31
|
+
scope = model.where(@table.field => @purge_values)
|
|
32
|
+
scope = scope.where(@table.conditions) if @table.conditions
|
|
33
|
+
scope
|
|
34
|
+
end
|
|
35
|
+
|
|
36
|
+
def explain_nullify
|
|
37
|
+
::DBPurger.config.explain_file.puts("#{explain_update_sql(scope, field_quoted, 'NULL')};")
|
|
38
|
+
scope.count
|
|
39
|
+
end
|
|
40
|
+
|
|
41
|
+
def field_quoted
|
|
42
|
+
model.connection.quote_column_name(@table.field)
|
|
43
|
+
end
|
|
44
|
+
end
|
|
45
|
+
end
|
data/lib/db-purger/plan.rb
CHANGED
|
@@ -8,39 +8,48 @@ module DBPurger
|
|
|
8
8
|
attr_reader :parent_tables,
|
|
9
9
|
:child_tables,
|
|
10
10
|
:ignore_tables,
|
|
11
|
-
:search_tables
|
|
11
|
+
:search_tables,
|
|
12
|
+
:nullify_tables
|
|
12
13
|
|
|
13
14
|
def initialize
|
|
14
15
|
@parent_tables = []
|
|
15
16
|
@child_tables = []
|
|
16
17
|
@ignore_tables = []
|
|
17
18
|
@search_tables = []
|
|
19
|
+
@nullify_tables = []
|
|
18
20
|
end
|
|
19
21
|
|
|
20
22
|
def purge!(database, purge_value)
|
|
21
|
-
raise('plan has no base_table')
|
|
23
|
+
raise('plan has no base_table or top-level parent_table') if root_tables.empty?
|
|
24
|
+
raise('top-level child_tables require a base_table') unless @base_table || @child_tables.empty?
|
|
25
|
+
raise('top-level nullify_tables require a base_table') unless @base_table || @nullify_tables.empty?
|
|
22
26
|
|
|
23
27
|
MetricSubscriber.reset!
|
|
24
|
-
num_deleted =
|
|
28
|
+
num_deleted = purge_root_tables(database, purge_value)
|
|
29
|
+
purge_search_tables(database)
|
|
25
30
|
MetricSubscriber.finished!
|
|
26
31
|
num_deleted
|
|
27
32
|
end
|
|
28
33
|
|
|
34
|
+
# tables that receive the purge value directly: the base_table (if any) and top-level parent_tables
|
|
35
|
+
def root_tables
|
|
36
|
+
(@base_table ? [@base_table] : []) + @parent_tables
|
|
37
|
+
end
|
|
38
|
+
|
|
29
39
|
def tables
|
|
30
40
|
all_tables = @base_table ? [@base_table] + @base_table.tables : []
|
|
31
41
|
all_tables += @parent_tables + @parent_tables.map(&:tables) +
|
|
32
42
|
@child_tables + @child_tables.map(&:tables) +
|
|
33
|
-
@search_tables + @search_tables.map(&:tables)
|
|
43
|
+
@search_tables + @search_tables.map(&:tables) +
|
|
44
|
+
@nullify_tables
|
|
34
45
|
all_tables.flatten!
|
|
35
46
|
all_tables.compact!
|
|
36
47
|
all_tables
|
|
37
48
|
end
|
|
38
49
|
|
|
50
|
+
# the tables of a nested plan (nested plans never have a base_table)
|
|
39
51
|
def foreign_tables
|
|
40
|
-
|
|
41
|
-
@parent_tables +
|
|
42
|
-
@child_tables +
|
|
43
|
-
@search_tables
|
|
52
|
+
@parent_tables + @child_tables + @search_tables
|
|
44
53
|
end
|
|
45
54
|
|
|
46
55
|
def table_names
|
|
@@ -51,7 +60,8 @@ module DBPurger
|
|
|
51
60
|
@base_table.nil? &&
|
|
52
61
|
@parent_tables.empty? &&
|
|
53
62
|
@child_tables.empty? &&
|
|
54
|
-
@search_tables.empty?
|
|
63
|
+
@search_tables.empty? &&
|
|
64
|
+
@nullify_tables.empty?
|
|
55
65
|
end
|
|
56
66
|
|
|
57
67
|
def ignore_table?(table_name)
|
|
@@ -63,5 +73,20 @@ module DBPurger
|
|
|
63
73
|
end
|
|
64
74
|
end
|
|
65
75
|
end
|
|
76
|
+
|
|
77
|
+
private
|
|
78
|
+
|
|
79
|
+
def purge_root_tables(database, purge_value)
|
|
80
|
+
root_tables.sum do |table|
|
|
81
|
+
PurgeTable.new(database, table, table.field, purge_value).purge!
|
|
82
|
+
end
|
|
83
|
+
end
|
|
84
|
+
|
|
85
|
+
# with a base_table these live in its nested plan and are purged by it
|
|
86
|
+
def purge_search_tables(database)
|
|
87
|
+
@search_tables.each do |table|
|
|
88
|
+
PurgeTableScanner.new(database, table).purge!
|
|
89
|
+
end
|
|
90
|
+
end
|
|
66
91
|
end
|
|
67
92
|
end
|
|
@@ -3,6 +3,8 @@
|
|
|
3
3
|
module DBPurger
|
|
4
4
|
# DBPurger::PlanBuilder is used to build the relationships between tables in a convenient way
|
|
5
5
|
class PlanBuilder
|
|
6
|
+
NULLIFY_TABLE_OPTIONS = %i[conditions].freeze
|
|
7
|
+
|
|
6
8
|
def initialize(plan)
|
|
7
9
|
@plan = plan
|
|
8
10
|
end
|
|
@@ -31,6 +33,22 @@ module DBPurger
|
|
|
31
33
|
table
|
|
32
34
|
end
|
|
33
35
|
|
|
36
|
+
# rows whose field matches the enclosing batch's primary keys get field set to NULL instead of being deleted
|
|
37
|
+
def nullify_table(table_name, field, options = {})
|
|
38
|
+
unsupported_options = options.keys - NULLIFY_TABLE_OPTIONS
|
|
39
|
+
unless unsupported_options.empty?
|
|
40
|
+
raise(ArgumentError, "nullify_table does not support #{unsupported_options.map(&:inspect).join(', ')}")
|
|
41
|
+
end
|
|
42
|
+
|
|
43
|
+
table = create_table(table_name, field, options)
|
|
44
|
+
if @plan.base_table
|
|
45
|
+
@plan.base_table.nested_plan.nullify_tables << table
|
|
46
|
+
else
|
|
47
|
+
@plan.nullify_tables << table
|
|
48
|
+
end
|
|
49
|
+
table
|
|
50
|
+
end
|
|
51
|
+
|
|
34
52
|
def ignore_table(table_name)
|
|
35
53
|
@plan.ignore_tables << table_name
|
|
36
54
|
end
|
|
@@ -11,6 +11,7 @@ module DBPurger
|
|
|
11
11
|
validate :validate_no_missing_tables
|
|
12
12
|
validate :validate_no_unknown_tables
|
|
13
13
|
validate :validate_tables
|
|
14
|
+
validate :validate_nullify_tables
|
|
14
15
|
|
|
15
16
|
def initialize(database, plan)
|
|
16
17
|
@database = database
|
|
@@ -29,8 +30,16 @@ module DBPurger
|
|
|
29
30
|
|
|
30
31
|
private
|
|
31
32
|
|
|
33
|
+
# a plan is rooted either by a base_table or by one or more top-level parent_tables
|
|
32
34
|
def validate_base_table
|
|
33
|
-
|
|
35
|
+
if @plan.root_tables.empty?
|
|
36
|
+
errors.add(:base_table, 'or a top-level parent_table is required')
|
|
37
|
+
elsif !@plan.child_tables.empty?
|
|
38
|
+
# without a base_table there are no ids to propagate; declared before one, they are never reached
|
|
39
|
+
errors.add(:base_table, 'must be declared before top-level child_tables')
|
|
40
|
+
elsif !@plan.nullify_tables.empty?
|
|
41
|
+
errors.add(:base_table, 'must be declared before top-level nullify_tables')
|
|
42
|
+
end
|
|
34
43
|
end
|
|
35
44
|
|
|
36
45
|
def validate_no_missing_tables
|
|
@@ -45,6 +54,27 @@ module DBPurger
|
|
|
45
54
|
@plan.tables.each { |table| validate_table_definition(table) }
|
|
46
55
|
end
|
|
47
56
|
|
|
57
|
+
# a NOT NULL column can't be unlinked, so the purge would fail at runtime instead of here
|
|
58
|
+
def validate_nullify_tables
|
|
59
|
+
nullify_tables.each do |table|
|
|
60
|
+
next unless (model = find_model_for_table(table)) # reported by validate_tables
|
|
61
|
+
next unless not_nullable_column?(model, table.field)
|
|
62
|
+
|
|
63
|
+
errors.add(:table, "#{table.name}.#{table.field} (nullify_table) is not nullable")
|
|
64
|
+
end
|
|
65
|
+
end
|
|
66
|
+
|
|
67
|
+
# a missing column is reported by validate_tables
|
|
68
|
+
def not_nullable_column?(model, field)
|
|
69
|
+
column = model.columns_hash[field.to_s]
|
|
70
|
+
column && !column.null
|
|
71
|
+
end
|
|
72
|
+
|
|
73
|
+
def nullify_tables
|
|
74
|
+
@plan.nullify_tables +
|
|
75
|
+
@plan.tables.select(&:nested_tables?).flat_map { |table| table.nested_plan.nullify_tables }
|
|
76
|
+
end
|
|
77
|
+
|
|
48
78
|
def validate_table_definition(table)
|
|
49
79
|
unless (model = find_model_for_table(table))
|
|
50
80
|
errors.add(:table, "#{table.name} has no model")
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module DBPurger
|
|
4
|
+
# DBPurger::PlanWriter renders plan DSL source and tracks which tables it has written
|
|
5
|
+
class PlanWriter
|
|
6
|
+
INDENT = ' '
|
|
7
|
+
|
|
8
|
+
attr_reader :output,
|
|
9
|
+
:table_names
|
|
10
|
+
|
|
11
|
+
def initialize
|
|
12
|
+
@output = ''.dup
|
|
13
|
+
@indent_depth = 0
|
|
14
|
+
@table_names = []
|
|
15
|
+
end
|
|
16
|
+
|
|
17
|
+
def table(table_type, table_name, field)
|
|
18
|
+
@table_names << table_name
|
|
19
|
+
write(table_call(table_type, table_name, field))
|
|
20
|
+
end
|
|
21
|
+
|
|
22
|
+
def table_block(table_type, table_name, field)
|
|
23
|
+
@table_names << table_name
|
|
24
|
+
write("#{table_call(table_type, table_name, field)} do")
|
|
25
|
+
@indent_depth += 1
|
|
26
|
+
yield
|
|
27
|
+
@indent_depth -= 1
|
|
28
|
+
write('end')
|
|
29
|
+
end
|
|
30
|
+
|
|
31
|
+
def comment(str)
|
|
32
|
+
write("# #{str}")
|
|
33
|
+
end
|
|
34
|
+
|
|
35
|
+
def ignore_tables(table_names)
|
|
36
|
+
return if table_names.empty?
|
|
37
|
+
|
|
38
|
+
line_break
|
|
39
|
+
table_names.each { |table_name| write("ignore_table #{table_name.to_sym.inspect}") }
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
def line_break
|
|
43
|
+
@output << "\n"
|
|
44
|
+
end
|
|
45
|
+
|
|
46
|
+
private
|
|
47
|
+
|
|
48
|
+
def write(str)
|
|
49
|
+
@output << "#{INDENT * @indent_depth}#{str}\n"
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
def table_call(table_type, table_name, field)
|
|
53
|
+
"#{table_type}_table(#{table_name.to_sym.inspect}, #{field.to_sym.inspect})"
|
|
54
|
+
end
|
|
55
|
+
end
|
|
56
|
+
end
|
|
@@ -10,9 +10,19 @@ module DBPurger
|
|
|
10
10
|
private
|
|
11
11
|
|
|
12
12
|
def purge_nested_tables(batch)
|
|
13
|
+
nullify_tables(batch) unless @table.nested_plan.nullify_tables.empty?
|
|
13
14
|
purge_child_tables(batch) unless @table.nested_plan.child_tables.empty?
|
|
14
15
|
end
|
|
15
16
|
|
|
17
|
+
# unlink rows that point at this batch before anything in it is deleted
|
|
18
|
+
def nullify_tables(batch)
|
|
19
|
+
ids = batch_values(batch, model.primary_key)
|
|
20
|
+
|
|
21
|
+
@table.nested_plan.nullify_tables.each do |table|
|
|
22
|
+
NullifyTable.new(@database, table, ids).nullify!
|
|
23
|
+
end
|
|
24
|
+
end
|
|
25
|
+
|
|
16
26
|
def purge_child_tables(batch)
|
|
17
27
|
ids = batch_values(batch, model.primary_key)
|
|
18
28
|
|
|
@@ -69,7 +79,7 @@ module DBPurger
|
|
|
69
79
|
def explain(scope)
|
|
70
80
|
sql =
|
|
71
81
|
if @table.mark_deleted_field
|
|
72
|
-
explain_update_sql(scope)
|
|
82
|
+
explain_update_sql(scope, mark_deleted_field_quoted, mark_deleted_value_quoted)
|
|
73
83
|
else
|
|
74
84
|
scope.to_sql.sub(/SELECT .*?FROM/, 'DELETE FROM')
|
|
75
85
|
end
|
|
@@ -77,10 +87,10 @@ module DBPurger
|
|
|
77
87
|
scope.count
|
|
78
88
|
end
|
|
79
89
|
|
|
80
|
-
def explain_update_sql(scope)
|
|
90
|
+
def explain_update_sql(scope, field_quoted, value_quoted)
|
|
81
91
|
sql = scope.to_sql.dup
|
|
82
92
|
sql.sub!(/SELECT .*?FROM/, 'UPDATE')
|
|
83
|
-
sql.sub!('WHERE', "SET #{
|
|
93
|
+
sql.sub!('WHERE', "SET #{field_quoted} = #{value_quoted} WHERE")
|
|
84
94
|
sql
|
|
85
95
|
end
|
|
86
96
|
|
|
@@ -11,10 +11,6 @@ module DBPurger
|
|
|
11
11
|
@num_deleted = 0
|
|
12
12
|
end
|
|
13
13
|
|
|
14
|
-
def model
|
|
15
|
-
@model ||= @database.models.detect { |m| m.table_name == @table.name.to_s }
|
|
16
|
-
end
|
|
17
|
-
|
|
18
14
|
def purge!
|
|
19
15
|
ActiveSupport::Notifications.instrument('purge.db_purger',
|
|
20
16
|
table_name: @table.name) do |payload|
|
data/lib/db-purger/table.rb
CHANGED
|
@@ -38,7 +38,8 @@ module DBPurger
|
|
|
38
38
|
|
|
39
39
|
# nested tables that depend on this table's rows (search tables scan independently)
|
|
40
40
|
def nested_key_tables?
|
|
41
|
-
nested_tables? &&
|
|
41
|
+
nested_tables? &&
|
|
42
|
+
!(@nested_plan.child_tables.empty? && @nested_plan.parent_tables.empty? && @nested_plan.nullify_tables.empty?)
|
|
42
43
|
end
|
|
43
44
|
|
|
44
45
|
def tables
|
data/lib/db-purger.rb
CHANGED
|
@@ -2,17 +2,20 @@
|
|
|
2
2
|
|
|
3
3
|
# DBPurger is a tool to delete data from tables based on a initial purge value
|
|
4
4
|
module DBPurger
|
|
5
|
+
autoload :AssociationGraph, 'db-purger/association_graph'
|
|
5
6
|
autoload :Config, 'db-purger/config'
|
|
6
7
|
autoload :DynamicPlanBuilder, 'db-purger/dynamic_plan_builder'
|
|
7
8
|
autoload :Executor, 'db-purger/executor'
|
|
8
9
|
autoload :Metrics, 'db-purger/metrics'
|
|
9
10
|
autoload :MetricSubscriber, 'db-purger/metric_subscriber'
|
|
11
|
+
autoload :NullifyTable, 'db-purger/nullify_table'
|
|
10
12
|
autoload :PurgeTable, 'db-purger/purge_table'
|
|
11
13
|
autoload :PurgeTableHelper, 'db-purger/purge_table_helper'
|
|
12
14
|
autoload :PurgeTableScanner, 'db-purger/purge_table_scanner'
|
|
13
15
|
autoload :Plan, 'db-purger/plan'
|
|
14
16
|
autoload :PlanBuilder, 'db-purger/plan_builder'
|
|
15
17
|
autoload :PlanValidator, 'db-purger/plan_validator'
|
|
18
|
+
autoload :PlanWriter, 'db-purger/plan_writer'
|
|
16
19
|
autoload :Table, 'db-purger/table'
|
|
17
20
|
|
|
18
21
|
# The config in effect for the current thread: the one set by with_config, else the global default
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: db-purger
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.7.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Doug Youch
|
|
@@ -29,14 +29,20 @@ dependencies:
|
|
|
29
29
|
requirements:
|
|
30
30
|
- - "~>"
|
|
31
31
|
- !ruby/object:Gem::Version
|
|
32
|
-
version: '0.
|
|
32
|
+
version: '0.9'
|
|
33
|
+
- - ">="
|
|
34
|
+
- !ruby/object:Gem::Version
|
|
35
|
+
version: 0.9.1
|
|
33
36
|
type: :runtime
|
|
34
37
|
prerelease: false
|
|
35
38
|
version_requirements: !ruby/object:Gem::Requirement
|
|
36
39
|
requirements:
|
|
37
40
|
- - "~>"
|
|
38
41
|
- !ruby/object:Gem::Version
|
|
39
|
-
version: '0.
|
|
42
|
+
version: '0.9'
|
|
43
|
+
- - ">="
|
|
44
|
+
- !ruby/object:Gem::Version
|
|
45
|
+
version: 0.9.1
|
|
40
46
|
description: DB Purger deletes (or soft-deletes) every row related to a single top-level
|
|
41
47
|
record (e.g. a company or account) using a declarative Ruby purge plan. Tables are
|
|
42
48
|
purged in primary-key batches, child tables before their parents, with plan validation
|
|
@@ -51,14 +57,17 @@ files:
|
|
|
51
57
|
- LICENSE.txt
|
|
52
58
|
- README.md
|
|
53
59
|
- lib/db-purger.rb
|
|
60
|
+
- lib/db-purger/association_graph.rb
|
|
54
61
|
- lib/db-purger/config.rb
|
|
55
62
|
- lib/db-purger/dynamic_plan_builder.rb
|
|
56
63
|
- lib/db-purger/executor.rb
|
|
57
64
|
- lib/db-purger/metric_subscriber.rb
|
|
58
65
|
- lib/db-purger/metrics.rb
|
|
66
|
+
- lib/db-purger/nullify_table.rb
|
|
59
67
|
- lib/db-purger/plan.rb
|
|
60
68
|
- lib/db-purger/plan_builder.rb
|
|
61
69
|
- lib/db-purger/plan_validator.rb
|
|
70
|
+
- lib/db-purger/plan_writer.rb
|
|
62
71
|
- lib/db-purger/purge_table.rb
|
|
63
72
|
- lib/db-purger/purge_table_helper.rb
|
|
64
73
|
- lib/db-purger/purge_table_scanner.rb
|