sidekiq-batch-jobs 0.1.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +185 -0
  3. data/LICENSE.txt +21 -0
  4. data/README.md +480 -68
  5. data/Rakefile +13 -0
  6. data/app/models/sidekiq_batch/abandoned_enrollment_error.rb +9 -0
  7. data/app/models/sidekiq_batch/abandoned_enrollment_reaper.rb +76 -0
  8. data/app/models/sidekiq_batch/announcement.rb +106 -0
  9. data/app/models/sidekiq_batch/batch_enrollment_context.rb +165 -40
  10. data/app/models/sidekiq_batch/client_middleware.rb +5 -7
  11. data/app/models/sidekiq_batch/completion_query.rb +110 -0
  12. data/app/models/sidekiq_batch/groomer_worker.rb +15 -0
  13. data/app/models/sidekiq_batch/jid_index.rb +49 -0
  14. data/app/models/sidekiq_batch/middleware.rb +58 -38
  15. data/app/models/sidekiq_batch/orphaned_callback_reaper.rb +37 -0
  16. data/app/models/sidekiq_batch/orphaned_job_error.rb +9 -0
  17. data/app/models/sidekiq_batch/reaper_worker.rb +58 -0
  18. data/app/models/sidekiq_batch/record_groomer.rb +31 -0
  19. data/app/models/sidekiq_batch/stuck_job_reaper.rb +67 -0
  20. data/app/models/sidekiq_batch.rb +145 -59
  21. data/app/models/sidekiq_batch_job.rb +38 -18
  22. data/lib/generators/sidekiq/batch/jobs/install/install_generator.rb +11 -1
  23. data/lib/generators/sidekiq/batch/jobs/install/templates/create_sidekiq_batch_tables.rb.tt +14 -2
  24. data/lib/generators/sidekiq/batch/jobs/upgrade/templates/upgrade_sidekiq_batch_tables.rb.tt +5 -0
  25. data/lib/generators/sidekiq/batch/jobs/upgrade/upgrade_generator.rb +129 -0
  26. data/lib/sidekiq/batch/jobs/configuration.rb +100 -0
  27. data/lib/sidekiq/batch/jobs/engine.rb +6 -2
  28. data/lib/sidekiq/batch/jobs/enum_compat.rb +29 -0
  29. data/lib/sidekiq/batch/jobs/failure_policy.rb +129 -0
  30. data/lib/sidekiq/batch/jobs/schema.rb +74 -0
  31. data/lib/sidekiq/batch/jobs/version.rb +1 -1
  32. data/lib/sidekiq/batch/jobs.rb +65 -21
  33. metadata +41 -124
  34. data/docker-compose.yml +0 -26
  35. data/sig/sidekiq/batch/jobs.rbs +0 -8
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: e99dbc78b86da74277380f2ba6fe62c5433ab3815d23b0f756ff6173267fb168
4
- data.tar.gz: 96f45bcea22e9cd2359190946e62b54d92c8c192ddfbc87dd4148ccb9abe53e1
3
+ metadata.gz: 148b48f6e660baae58a41a709b7b43010e8c7a1c29585368d9d624d23ce7bfa7
4
+ data.tar.gz: 5c606573cc3b223bd29b6843dc041cb8296c157b2d063b5e6d5055a736dca4b7
5
5
  SHA512:
6
- metadata.gz: 7d665aba6579df375e91325f433cecf574abb0310ad9efd142358232890b5ddabccdac3f6d3a3a9773080f7facd7617fc4d4c0ac4116285e938e6ec805a4c15e
7
- data.tar.gz: dd13456bd144b5d0a104a78b7ddf8eee1469c06a28ac497add4b78a6f4c666fa2c2151fcb2824513292b65f13af16e237bcc9e7713dec07d1d14e79ab54a509e
6
+ metadata.gz: e6b4be14970f3567194ab9f8c81fe8dde1493308c83b73165fdcfb7291f2b3d67644d0ca97168f90816955a9de78479d81111abfdf4e42c1522fb4bc6d6be8a3
7
+ data.tar.gz: 4c910d40ef9a6a183541a18c8fb5d8ff2f8483ea26ad5b1cc22b2d611f0728cc3f26c36d494715a36d64b0bce24cb51683ba672e967e01dd692b750ee004e96e
data/CHANGELOG.md ADDED
@@ -0,0 +1,185 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
+ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
+
8
+ ## [0.3.0] - 2026-09-07
9
+
10
+ ### Breaking
11
+
12
+ - **A migration is required, and the new code does not work without it.** The completion
13
+ statement writes both new columns, so against an unmigrated database every completion check
14
+ raises `PG::UndefinedColumn` and `#progress` raises `NameError` on a finished batch. See
15
+ *Upgrading from 0.2.0* below.
16
+
17
+ ### Added
18
+
19
+ - `complete_count` and `failed_count` on `sidekiq_batches`, stamped by the completion statement
20
+ in the same `UPDATE` that transitions the batch. `#progress` reads them back, so a finished
21
+ batch reports its counts without touching a job row, and `#percentage_progress` now answers
22
+ 100.0 from the status alone: a batch only transitions once nothing is pending, so a terminal
23
+ one has nothing left to count. A running batch counts rows exactly as before, so polling a
24
+ live batch is unchanged.
25
+
26
+ Written once, by the statement that already runs once per batch, rather than bumped from
27
+ `complete!` and `fail!`. A per-job counter would serialise every worker in a batch behind a
28
+ single row lock and churn a tuple per job, to speed up reads that were never on the hot path.
29
+ The completion statement takes that lock anyway, and a terminal batch transitions no further
30
+ jobs, so the tally cannot drift from the rows.
31
+
32
+ - `rails g sidekiq:batch:jobs:upgrade`, which replaces the copy-paste migration these notes
33
+ used to hand you. It reads your tables and emits only what they are missing, so upgrading
34
+ from 0.1.0 and upgrading from 0.2.0 are the same command and each writes the migration that
35
+ applies. Nothing to do means no file, so it is safe to run against a current database, and
36
+ safe to run twice. `Sidekiq::Batch::Jobs::Schema` is the schema it diffs against.
37
+
38
+ ### Upgrading from 0.2.0
39
+
40
+ ```bash
41
+ bin/rails g sidekiq:batch:jobs:upgrade
42
+ bin/rails db:migrate
43
+ ```
44
+
45
+ Two nullable columns, which on Postgres 11+ is a metadata-only change however large the table.
46
+
47
+ Run it before the workers pick up the new code, not after. The completion statement names both
48
+ columns, so until they exist every completion check fails.
49
+
50
+ Nothing is lost if that order slips. The failure is caught where every completion check is
51
+ caught: the job itself still succeeds, the batch keeps its rows and stays `running`, and the
52
+ reason goes to `config.on_alert`. Once the columns exist, the next job to finish transitions
53
+ its batch normally, and `SidekiqBatch::ReaperWorker` completes any batch whose last job already
54
+ ran, once it has been quiet for `config.stuck_after` (two hours by default).
55
+
56
+ No backfill. NULL means "no tally was stamped", which is every running batch and every batch
57
+ that finished under 0.2.0; both fall back to counting job rows the way 0.2.0 always did.
58
+ Batches that finish after the migration carry their own counts.
59
+
60
+ ## [0.2.0] - 2026-08-31
61
+
62
+ 0.1.0 could enrol a batch, detect completion and fire a callback. This release keeps that core
63
+ and builds the rest of the gem around it: a maintenance layer that repairs batches nothing else
64
+ can reach, a failure policy you control, and a third callback event. It also renames the
65
+ callback vocabulary to match Sidekiq Pro's, which is the change most likely to affect you.
66
+
67
+ ### Breaking
68
+
69
+ - **`:complete` now fires whatever the outcome.** In 0.1.0 it meant "all finished *and* all
70
+ succeeded"; that meaning is now `:success`. If you registered `on(:complete, MyWorker)` as a
71
+ success hook, it will start running on failed batches too, with nothing to warn you. Change
72
+ those registrations to `on(:success, MyWorker)` to keep 0.1.0's behaviour.
73
+ - Batch status `complete` is now `succeeded`, so each terminal status names the event that
74
+ fires with it. The stored integer is unchanged, so no data migration is needed, but code
75
+ comparing `batch.status == "complete"` or calling `complete_status?` must be updated.
76
+ `SidekiqBatchJob` statuses are untouched: a *job* that finished is still `complete`.
77
+ - `SidekiqBatch#fire_callback(event)` is gone, replaced by `#fire_callbacks`. A batch now
78
+ announces two events at once, so firing one by hand would spend an announcement the batch has
79
+ not finished making.
80
+ - `Sidekiq::Batch::Jobs.auto_install = false` and `.disable_auto_install!` are gone, along with
81
+ `.reset_installed!`. Configure the gem through `Sidekiq::Batch::Jobs.configure` instead.
82
+ - A batch can only be enrolled once. A second `jobs { … }` block on the same batch raises
83
+ `AlreadyStartedError` rather than enrolling into a batch whose `total_jobs` is already
84
+ stamped. There is no reopening a batch the way Sidekiq Pro allows.
85
+ - **A migration is required.** See *Upgrading from 0.1.0* below.
86
+
87
+ ### Added
88
+
89
+ - A configurable failure policy: `:any_failure` (the default, and what 0.1.0 always did),
90
+ `:all_failed`, `{ tolerate: 10 }` or `{ tolerate: "5%" }`, per batch or globally through
91
+ `config.failure_policy`. All four are the same rule with a different threshold, evaluated
92
+ inside the same atomic statement that transitions the batch, so choosing one never costs a
93
+ race. Nothing forgives a `jobs {}` block that died partway through enrolling: that batch
94
+ fails whatever the policy, with the cause recorded in `enrollment_error`.
95
+ - A third callback event. `:complete` fires however the batch went, `:success` and `:failure`
96
+ name the outcome and are mutually exclusive, so "always do X, and separately tell me when it
97
+ went badly" is two registrations rather than one worker doing both. Each event's enqueue is
98
+ claimed atomically and independently, so a callback that cannot be delivered never causes a
99
+ redelivery of one that already went out. Delivery is at-least-once: the push shares a
100
+ transaction with the claim, so a crash between the two rolls that claim back and the reaper
101
+ re-enqueues. Write callbacks to be idempotent.
102
+ - `SidekiqBatch::ReaperWorker`, three recoveries in one pass: batches whose `jobs {}` block died
103
+ mid-enrollment and left them `pending` forever, batches stalled by a job that vanished from
104
+ Redis (SIGKILL, OOM, pod eviction, or **Kill All** on Sidekiq's Retries page, which moves jobs
105
+ to the dead set without running death handlers), and callbacks orphaned by a crash between the
106
+ completion `UPDATE` and the enqueue.
107
+ - `SidekiqBatch::GroomerWorker`, which deletes terminal batches past the retention window and
108
+ long-abandoned `pending` ones as a backstop for hosts that do not run the reaper.
109
+ - `Sidekiq::Batch::Jobs.configure` for `base_class_name`, `stuck_after`, `retention`,
110
+ `failure_policy`, `on_alert`, `maintenance_queue`, `error_message_max` and `auto_install`.
111
+ `on_alert` is the gem's only unprompted signal; point it somewhere you read.
112
+ - `#percentage_progress`, `#eta` and `#terminal?` alongside the existing `#progress`.
113
+ - Rails 6.1 and Sidekiq 7 support. 0.1.0 required Ruby >= 3.2 and only worked on Rails 7+;
114
+ 0.2.0 runs on Ruby >= 3.0 and is tested against Rails 6.1, 7.1 and 8.0 with Sidekiq 7 and 8.
115
+
116
+ ### Fixed
117
+
118
+ Found by a full audit of 0.1.0. Each fix has a spec that was confirmed to fail against the
119
+ old code.
120
+
121
+ - **A job's row could be marked failed on its first attempt.** `final_attempt?` compared
122
+ `retry_count` against the wrong bound and read the worker class instead of the payload, so a
123
+ `retry: false` worker pushed with `set(retry: 5)` failed its batch while Sidekiq went on to
124
+ retry and succeed. Unrecoverable once it happened. The same bug made the middleware's
125
+ failure path dead code for every retrying worker.
126
+ - **An exception or a crash inside `jobs { … }` stranded the batch in `pending` forever.**
127
+ Its queued jobs ran, nothing ever announced, and neither the batch nor its rows were ever
128
+ cleaned up. Ordinary exceptions are now recorded and the batch finished; a killed process is
129
+ recovered by the reaper.
130
+ - **A batch's own callback could be enrolled into the batch it announces**, freezing
131
+ `total_jobs` below the row count and leaving a `pending` row nothing would look at.
132
+ - **Reopening a batch deleted it.** A second `jobs { … }` block that enrolled nothing, or
133
+ raised before its first push, destroyed the batch and every row tracking a job that was
134
+ already running.
135
+ - **Every Sidekiq job in the host application ran a Postgres query**, batch-tracked or not.
136
+ Untracked jobs now cost zero queries and have no Postgres coupling at all.
137
+ - **Enrollment cost five round trips per job**, two of which bought nothing. Now three, which
138
+ is the floor.
139
+ - Enrolling inside an open transaction was only detected at the start of the block, and only
140
+ on the default connection.
141
+ - `fire_callback` was public and consumed the batch's one announcement, so calling it early
142
+ permanently suppressed the real callback.
143
+ - A bookkeeping failure could reach the caller as if their enqueue had failed, inviting a retry
144
+ that enqueued the whole batch a second time; another could replace a job's own exception on
145
+ its way to Sidekiq, misreporting why the job died.
146
+ - `batch.destroy` loaded and destroyed children one row at a time despite the FK already
147
+ cascading.
148
+
149
+ ### Upgrading from 0.1.0
150
+
151
+ Four columns and two indexes:
152
+
153
+ ```ruby
154
+ class UpgradeSidekiqBatchTables < ActiveRecord::Migration[7.1]
155
+ def change
156
+ add_column :sidekiq_batches, :callbacks_fired, :jsonb, null: false, default: {}
157
+ add_column :sidekiq_batches, :failure_policy, :string
158
+ add_column :sidekiq_batches, :failure_tolerance, :integer
159
+ add_column :sidekiq_batches, :enrollment_error, :jsonb
160
+
161
+ add_index :sidekiq_batches, [:status, :created_at]
162
+ add_index :sidekiq_batch_jobs, :updated_at
163
+ remove_index :sidekiq_batches, :status
164
+ end
165
+ end
166
+ ```
167
+
168
+ The bare `:status` index goes because both composites lead with `status`, and Postgres serves a
169
+ status-only lookup from either. `sidekiq_batch_jobs.updated_at` is what stops the reaper
170
+ sequentially scanning the whole jobs table.
171
+
172
+ No data backfill. The status integers are unchanged, and existing rows keep a NULL
173
+ `failure_policy`, which the completion statement reads as `:any_failure`: exactly what 0.1.0
174
+ did. Batches that already announced under 0.1.0 have `callback_fired_at` set, so the new
175
+ orphaned-callback reaper leaves them alone.
176
+
177
+ Then re-read the `:complete` note under **Breaking** above. It is the one change that alters
178
+ behaviour without raising anything.
179
+
180
+ ## [0.1.0] - 2026-05-13
181
+
182
+ First release. `batch.jobs { … }` enrollment with rows committed before their jobs reach Redis,
183
+ atomic completion detection, two callback events (`:complete` and `:failure`), `#progress`, a
184
+ `context` column, and client middleware, server middleware and a death handler wired up by the
185
+ Rails engine.
data/LICENSE.txt ADDED
@@ -0,0 +1,21 @@
1
+ The MIT License (MIT)
2
+
3
+ Copyright (c) 2026 Douglas Greyling
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.