sidekiq-batch-jobs 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 9e70220429bbc1911259964dccf5cc7a67ac5ecdbb0d2ecafe98fb783d70c043
4
- data.tar.gz: fb4d826947c4acf259e3641022031b656a6b4bd879ba748b1e0cc1f9342e691b
3
+ metadata.gz: 148b48f6e660baae58a41a709b7b43010e8c7a1c29585368d9d624d23ce7bfa7
4
+ data.tar.gz: 5c606573cc3b223bd29b6843dc041cb8296c157b2d063b5e6d5055a736dca4b7
5
5
  SHA512:
6
- metadata.gz: 9b1bbb5b23cae02b4d51ab59c8a471f75cb46371236a6b9dd1494979cc9523c4a46c00e10e4402a713d6c9d426087e6f8ab7fd57bad48d0998b6457d90eb966a
7
- data.tar.gz: d514d78ae12916da738b6d421a46eb4401d3e8ee3664dc8af67c84c21d26accb7ac2a8609450fe255ebb6ea38957d12a505c602402782fc69bb70c5976b59844
6
+ metadata.gz: e6b4be14970f3567194ab9f8c81fe8dde1493308c83b73165fdcfb7291f2b3d67644d0ca97168f90816955a9de78479d81111abfdf4e42c1522fb4bc6d6be8a3
7
+ data.tar.gz: 4c910d40ef9a6a183541a18c8fb5d8ff2f8483ea26ad5b1cc22b2d611f0728cc3f26c36d494715a36d64b0bce24cb51683ba672e967e01dd692b750ee004e96e
data/CHANGELOG.md CHANGED
@@ -5,6 +5,58 @@ All notable changes to this project are documented here.
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
 
8
+ ## [0.3.0] - 2026-09-07
9
+
10
+ ### Breaking
11
+
12
+ - **A migration is required, and the new code does not work without it.** The completion
13
+ statement writes both new columns, so against an unmigrated database every completion check
14
+ raises `PG::UndefinedColumn` and `#progress` raises `NameError` on a finished batch. See
15
+ *Upgrading from 0.2.0* below.
16
+
17
+ ### Added
18
+
19
+ - `complete_count` and `failed_count` on `sidekiq_batches`, stamped by the completion statement
20
+ in the same `UPDATE` that transitions the batch. `#progress` reads them back, so a finished
21
+ batch reports its counts without touching a job row, and `#percentage_progress` now answers
22
+ 100.0 from the status alone: a batch only transitions once nothing is pending, so a terminal
23
+ one has nothing left to count. A running batch counts rows exactly as before, so polling a
24
+ live batch is unchanged.
25
+
26
+ Written once, by the statement that already runs once per batch, rather than bumped from
27
+ `complete!` and `fail!`. A per-job counter would serialise every worker in a batch behind a
28
+ single row lock and churn a tuple per job, to speed up reads that were never on the hot path.
29
+ The completion statement takes that lock anyway, and a terminal batch transitions no further
30
+ jobs, so the tally cannot drift from the rows.
31
+
32
+ - `rails g sidekiq:batch:jobs:upgrade`, which replaces the copy-paste migration these notes
33
+ used to hand you. It reads your tables and emits only what they are missing, so upgrading
34
+ from 0.1.0 and upgrading from 0.2.0 are the same command and each writes the migration that
35
+ applies. Nothing to do means no file, so it is safe to run against a current database, and
36
+ safe to run twice. `Sidekiq::Batch::Jobs::Schema` is the schema it diffs against.
37
+
38
+ ### Upgrading from 0.2.0
39
+
40
+ ```bash
41
+ bin/rails g sidekiq:batch:jobs:upgrade
42
+ bin/rails db:migrate
43
+ ```
44
+
45
+ Two nullable columns, which on Postgres 11+ is a metadata-only change however large the table.
46
+
47
+ Run it before the workers pick up the new code, not after. The completion statement names both
48
+ columns, so until they exist every completion check fails.
49
+
50
+ Nothing is lost if that order slips. The failure is caught where every completion check is
51
+ caught: the job itself still succeeds, the batch keeps its rows and stays `running`, and the
52
+ reason goes to `config.on_alert`. Once the columns exist, the next job to finish transitions
53
+ its batch normally, and `SidekiqBatch::ReaperWorker` completes any batch whose last job already
54
+ ran, once it has been quiet for `config.stuck_after` (two hours by default).
55
+
56
+ No backfill. NULL means "no tally was stamped", which is every running batch and every batch
57
+ that finished under 0.2.0; both fall back to counting job rows the way 0.2.0 always did.
58
+ Batches that finish after the migration carry their own counts.
59
+
8
60
  ## [0.2.0] - 2026-08-31
9
61
 
10
62
  0.1.0 could enrol a batch, detect completion and fire a callback. This release keeps that core
data/README.md CHANGED
@@ -9,6 +9,7 @@ Tested against Ruby 3.0–3.4, Rails 6.1–8.0, Sidekiq 7–8.
9
9
  ## Contents
10
10
 
11
11
  - [Installation](#installation)
12
+ - [Upgrading an existing install](#upgrading-an-existing-install)
12
13
  - [Configuration](#configuration)
13
14
  - [Usage](#usage)
14
15
  - [The four steps](#the-four-steps)
@@ -53,6 +54,26 @@ death handler for you at boot, and again after every code reload.
53
54
  (`Sidekiq.configure_embed`) never runs this gem's server middleware, so its jobs stay
54
55
  `pending` even when they succeed, and the reaper eventually marks them failed.
55
56
 
57
+ ### Upgrading an existing install
58
+
59
+ Some releases add columns. Rather than a migration per release for you to apply in order,
60
+ there is one command that reads your tables and writes only what they are missing:
61
+
62
+ ```bash
63
+ bin/rails g sidekiq:batch:jobs:upgrade
64
+ bin/rails db:migrate
65
+ ```
66
+
67
+ It works from any earlier version, and writes no file at all when your schema is already
68
+ current, so it is safe to run whenever a release mentions the schema, and safe to run twice.
69
+ Read the migration before you run it, as you would any generated one. Do not use the install
70
+ generator for this: that one builds the tables from scratch.
71
+
72
+ Run it **before** the new code reaches your workers. The gem's SQL names every column it
73
+ expects, so between deploying and migrating, completion checks fail. Nothing is lost if that
74
+ order slips: the failure goes to `config.on_alert`, jobs still run, batches keep their rows,
75
+ and everything stalled completes once the migration lands.
76
+
56
77
  ### Configuration
57
78
 
58
79
  Every setting is optional and has a working default. To change one, do it in an initializer:
@@ -316,6 +337,8 @@ batch.status # "pending" | "running" | "succeeded" | "failed"
316
337
  batch.terminal? # true once the batch has finished, either way
317
338
  batch.completed_at # nil until the batch finishes
318
339
  batch.total_jobs # stamped when the jobs {} block returns; 0 before that
340
+ batch.complete_count # the final tally, stamped when the batch finishes;
341
+ batch.failed_count # both nil while it is still running
319
342
  batch.context # the jsonb hash you stashed when creating the batch
320
343
 
321
344
  batch.failure_policy # "any_failure" | "all_failed" | "tolerate_jobs" | "tolerate_percent"
@@ -326,6 +349,19 @@ batch.enrollment_error # nil normally; {"class" =>, "message" =>} if the
326
349
  batch.callbacks_fired # event => when that callback went out
327
350
  ```
328
351
 
352
+ **Reading a finished batch costs nothing.** The completion statement stamps `complete_count`
353
+ and `failed_count` in the same `UPDATE` that transitions the batch, so `progress` on a terminal
354
+ batch reads two attributes instead of counting job rows, and `percentage_progress` answers
355
+ 100.0 from the status alone: a batch only transitions once nothing is pending, so there is
356
+ nothing left to count. This matters most where you are most likely to look, which is from
357
+ inside a callback on a batch that just finished.
358
+
359
+ Those are the only two places the numbers are written. Bumping a counter from `complete!` and
360
+ `fail!` would put every worker in the batch behind one row lock and write a new version of the
361
+ batch row per job, all to speed up a read that is not on the hot path. The completion statement
362
+ takes that lock once per batch anyway, and a terminal batch transitions no further jobs, so its
363
+ tally cannot drift. A batch that is still running counts rows, exactly as before.
364
+
329
365
  Each enrolled job has a row of its own:
330
366
 
331
367
  ```ruby
@@ -9,9 +9,7 @@ class SidekiqBatch
9
9
  def sql
10
10
  <<~SQL.squish
11
11
  UPDATE sidekiq_batches
12
- SET status = #{outcome},
13
- completed_at = NOW(),
14
- updated_at = NOW()
12
+ SET #{assignments}
15
13
  WHERE id = ?
16
14
  AND status = #{batch_status("running")}
17
15
  AND NOT EXISTS (#{jobs_with(job_status("pending"))})
@@ -21,6 +19,15 @@ class SidekiqBatch
21
19
 
22
20
  private
23
21
 
22
+ def assignments
23
+ <<~SQL
24
+ status = #{outcome},
25
+ (complete_count, failed_count) = (#{final_counts}),
26
+ completed_at = NOW(),
27
+ updated_at = NOW()
28
+ SQL
29
+ end
30
+
24
31
  def outcome
25
32
  <<~SQL
26
33
  CASE
@@ -62,6 +69,25 @@ class SidekiqBatch
62
69
  "GREATEST(COALESCE(failure_tolerance, 0), 0)"
63
70
  end
64
71
 
72
+ # The tally the batch keeps once it is finished, so reading its progress
73
+ # later costs no job rows at all. Free of the contention a per-job counter
74
+ # would buy: this statement's WHERE matches nothing until the last job
75
+ # lands, so it runs once per batch and takes the row lock once. Nothing
76
+ # can drift either, since a terminal batch transitions no further jobs.
77
+ #
78
+ # One multi-column assignment rather than two scalar subqueries, so the
79
+ # batch's rows are walked once for both numbers. The CASE in `status`
80
+ # cannot read them: every SET expression sees the pre-UPDATE row, which is
81
+ # why the tolerating branch below keeps a count of its own.
82
+ def final_counts
83
+ <<~SQL
84
+ SELECT COUNT(*) FILTER (WHERE status = #{job_status("complete")}),
85
+ COUNT(*) FILTER (WHERE status = #{job_status("failed")})
86
+ FROM sidekiq_batch_jobs
87
+ WHERE sidekiq_batch_id = sidekiq_batches.id
88
+ SQL
89
+ end
90
+
65
91
  def failed_count
66
92
  "SELECT COUNT(*) FROM sidekiq_batch_jobs " \
67
93
  "WHERE sidekiq_batch_id = sidekiq_batches.id AND status = #{job_status("failed")}"
@@ -8,10 +8,12 @@
8
8
  # callback_fired_at :datetime
9
9
  # callbacks :jsonb not null
10
10
  # callbacks_fired :jsonb not null
11
+ # complete_count :integer
11
12
  # completed_at :datetime
12
13
  # context :jsonb not null
13
14
  # description :string
14
15
  # enrollment_error :jsonb
16
+ # failed_count :integer
15
17
  # failure_policy :string
16
18
  # failure_tolerance :integer
17
19
  # status :integer default("pending"), not null
@@ -112,6 +114,8 @@ class SidekiqBatch < ::Sidekiq::Batch::Jobs.base_class
112
114
  def percentage_progress
113
115
  return 0.0 if total_jobs.zero?
114
116
 
117
+ return 100.0 if terminal?
118
+
115
119
  ((finished_jobs_count.to_f / total_jobs) * 100).round(2)
116
120
  end
117
121
 
@@ -131,14 +135,7 @@ class SidekiqBatch < ::Sidekiq::Batch::Jobs.base_class
131
135
  end
132
136
 
133
137
  def progress
134
- counts = sidekiq_batch_jobs.group(:status).count
135
-
136
- {
137
- total: total_jobs,
138
- complete: counts.fetch("complete", 0),
139
- failed: counts.fetch("failed", 0),
140
- pending: counts.fetch("pending", 0)
141
- }
138
+ stamped_counts || count_by_status
142
139
  end
143
140
 
144
141
  def pending_jobs
@@ -192,6 +189,23 @@ class SidekiqBatch < ::Sidekiq::Batch::Jobs.base_class
192
189
  end
193
190
  end
194
191
 
192
+ def stamped_counts
193
+ return nil unless terminal? && complete_count && failed_count
194
+
195
+ { total: total_jobs, complete: complete_count, failed: failed_count, pending: 0 }
196
+ end
197
+
198
+ def count_by_status
199
+ counts = sidekiq_batch_jobs.group(:status).count
200
+
201
+ {
202
+ total: total_jobs,
203
+ complete: counts.fetch("complete", 0),
204
+ failed: counts.fetch("failed", 0),
205
+ pending: counts.fetch("pending", 0)
206
+ }
207
+ end
208
+
195
209
  def finished_jobs_count
196
210
  sidekiq_batch_jobs.where(status: %w[complete failed]).count
197
211
  end
@@ -4,6 +4,8 @@ class CreateSidekiqBatchTables < ActiveRecord::Migration<%= migration_version %>
4
4
  t.string :description
5
5
  t.integer :status, null: false, default: 0
6
6
  t.integer :total_jobs, null: false, default: 0
7
+ t.integer :complete_count
8
+ t.integer :failed_count
7
9
  t.jsonb :callbacks, null: false, default: {}
8
10
  t.jsonb :callbacks_fired, null: false, default: {}
9
11
  t.jsonb :context, null: false, default: {}
@@ -0,0 +1,5 @@
1
+ class <%= migration_class_name %> < ActiveRecord::Migration<%= migration_version %>
2
+ def change
3
+ <%= migration_body %>
4
+ end
5
+ end
@@ -0,0 +1,129 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "rails/generators"
4
+ require "rails/generators/active_record"
5
+
6
+ module Sidekiq
7
+ module Batch
8
+ module Jobs
9
+ module Generators
10
+ # Brings an existing installation up to the schema this version of the
11
+ # gem expects.
12
+ #
13
+ # It reads the host's real tables and emits only what they are missing,
14
+ # rather than shipping one migration per release. A user upgrading from
15
+ # 0.1.0 and a user upgrading from the release before this one run the
16
+ # same command and each get exactly the difference that applies to
17
+ # them, in one migration. Running it on a current database writes
18
+ # nothing at all, so it is safe to run whenever a release mentions the
19
+ # schema, and safe to run twice.
20
+ class UpgradeGenerator < ::Rails::Generators::Base
21
+ include ::Rails::Generators::Migration
22
+
23
+ source_root File.expand_path("templates", __dir__)
24
+
25
+ desc "Adds whatever the sidekiq-batch-jobs tables are missing for this version of the gem."
26
+
27
+ def self.next_migration_number(dirname)
28
+ ::ActiveRecord::Generators::Base.next_migration_number(dirname)
29
+ end
30
+
31
+ def copy_migration
32
+ return report_missing_tables if missing_tables.any?
33
+ return report_up_to_date if changes.empty?
34
+
35
+ migration_template("upgrade_sidekiq_batch_tables.rb.tt", "db/migrate/#{migration_basename}.rb")
36
+ end
37
+
38
+ # NOTE: must stay private. Thor turns every public instance method on
39
+ # a generator into a step it invokes.
40
+ private
41
+
42
+ # Named for the version it upgrades to, so a later release's migration
43
+ # cannot collide with this one: two files called
44
+ # `upgrade_sidekiq_batch_tables` would be a duplicate migration name.
45
+ def migration_basename
46
+ "upgrade_sidekiq_batch_tables_to_v#{::Sidekiq::Batch::Jobs::VERSION.tr(".", "_")}"
47
+ end
48
+
49
+ # The template cannot hardcode a version: `ActiveRecord::Migration[7.1]`
50
+ # is unknown on Rails 6.1. Track whatever the host app runs instead.
51
+ def migration_version
52
+ "[#{::ActiveRecord::Migration.current_version}]"
53
+ end
54
+
55
+ # Read by the template.
56
+ def migration_body
57
+ changes.map { |line| " #{line}" }.join("\n")
58
+ end
59
+
60
+ def changes
61
+ @changes ||= added_columns + added_indexes + dropped_indexes
62
+ end
63
+
64
+ def added_columns
65
+ Schema::COLUMNS.flat_map do |table, columns|
66
+ columns
67
+ .reject { |name, _type, _options| connection.column_exists?(table, name) }
68
+ .map { |name, type, options| "add_column :#{table}, :#{name}, :#{type}#{arguments(options)}" }
69
+ end
70
+ end
71
+
72
+ def added_indexes
73
+ Schema::INDEXES.flat_map do |table, indexes|
74
+ indexes
75
+ .reject { |columns, _options| index?(table, columns) }
76
+ .map { |columns, options| "add_index :#{table}, #{columns.inspect}#{arguments(options)}" }
77
+ end
78
+ end
79
+
80
+ def dropped_indexes
81
+ Schema::SUPERSEDED_INDEXES.flat_map do |table, indexes|
82
+ indexes
83
+ .select { |columns| index?(table, columns) }
84
+ .map { |columns| "remove_index :#{table}, #{columns.inspect}" }
85
+ end
86
+ end
87
+
88
+ # Deliberately unqualified by the index's options. An index on the
89
+ # right columns is what the gem's queries need; rebuilding one a host
90
+ # has already tuned differently is not this generator's business.
91
+ def index?(table, columns)
92
+ connection.index_exists?(table, columns)
93
+ end
94
+
95
+ def arguments(options)
96
+ return "" if options.empty?
97
+
98
+ ", #{options.map { |key, value| "#{key}: #{value.inspect}" }.join(", ")}"
99
+ end
100
+
101
+ def missing_tables
102
+ @missing_tables ||= Schema::TABLES.reject { |table| connection.table_exists?(table) }
103
+ end
104
+
105
+ def report_missing_tables
106
+ say_status :skip,
107
+ "#{missing_tables.join(" and ")} not found. This upgrades an existing installation; " \
108
+ "run `rails g sidekiq:batch:jobs:install` to create the tables.",
109
+ :yellow
110
+ end
111
+
112
+ def report_up_to_date
113
+ say_status :identical, "the sidekiq-batch-jobs tables are already current", :blue
114
+ end
115
+
116
+ # Reported rather than raised as a connection error, since the reason
117
+ # this generator needs a database at all is not obvious from one.
118
+ def connection
119
+ @connection ||= ::ActiveRecord::Base.connection
120
+ rescue ::ActiveRecord::ActiveRecordError => e
121
+ raise ::Rails::Generators::Error,
122
+ "sidekiq-batch-jobs: this generator compares your schema against the one this version " \
123
+ "expects, and the database could not be read (#{e.class}: #{e.message})."
124
+ end
125
+ end
126
+ end
127
+ end
128
+ end
129
+ end
@@ -0,0 +1,74 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Sidekiq
4
+ module Batch
5
+ module Jobs
6
+ # The tables this version of the gem expects, as data rather than as a
7
+ # `create_table` block, so the upgrade generator can diff them against a
8
+ # host application's real schema and emit only what is missing. That is
9
+ # what lets one command upgrade a database from any earlier version,
10
+ # rather than a per-release migration the user has to apply in order.
11
+ #
12
+ # The install generator's template is this same schema written out for a
13
+ # fresh database. Two files that have to agree, and the upgrade generator
14
+ # spec catches it when they stop: run against a database built from the
15
+ # current schema, this diff has to come out empty.
16
+ module Schema
17
+ BATCHES = :sidekiq_batches
18
+ JOBS = :sidekiq_batch_jobs
19
+
20
+ TABLES = [BATCHES, JOBS].freeze
21
+
22
+ # [name, type, options]. `id` and the timestamps are left out
23
+ # deliberately: both have been on both tables since 0.1.0, and neither
24
+ # is something an upgrade could sensibly add to a populated table.
25
+ COLUMNS = {
26
+ BATCHES => [
27
+ [:description, :string, {}],
28
+ [:status, :integer, { null: false, default: 0 }],
29
+ [:total_jobs, :integer, { null: false, default: 0 }],
30
+ [:complete_count, :integer, {}],
31
+ [:failed_count, :integer, {}],
32
+ [:callbacks, :jsonb, { null: false, default: {} }],
33
+ [:callbacks_fired, :jsonb, { null: false, default: {} }],
34
+ [:context, :jsonb, { null: false, default: {} }],
35
+ [:failure_policy, :string, {}],
36
+ [:failure_tolerance, :integer, {}],
37
+ [:enrollment_error, :jsonb, {}],
38
+ [:completed_at, :datetime, {}],
39
+ [:callback_fired_at, :datetime, {}]
40
+ ].freeze,
41
+ JOBS => [
42
+ [:sidekiq_batch_id, :bigint, { null: false }],
43
+ [:jid, :string, { null: false }],
44
+ [:worker_class, :string, { null: false }],
45
+ [:args, :jsonb, { null: false, default: [] }],
46
+ [:status, :integer, { null: false, default: 0 }],
47
+ [:error_class, :string, {}],
48
+ [:error_message, :text, {}]
49
+ ].freeze
50
+ }.freeze
51
+
52
+ # [columns, options].
53
+ INDEXES = {
54
+ BATCHES => [
55
+ [%i[status callback_fired_at], {}],
56
+ [%i[status created_at], {}]
57
+ ].freeze,
58
+ JOBS => [
59
+ [%i[jid], { unique: true }],
60
+ [%i[sidekiq_batch_id status], {}],
61
+ [%i[updated_at], {}]
62
+ ].freeze
63
+ }.freeze
64
+
65
+ # Indexes an older version created that a current one should not keep.
66
+ # The bare `status` index goes because both composites above lead with
67
+ # `status`, and Postgres serves a status-only lookup from either.
68
+ SUPERSEDED_INDEXES = {
69
+ BATCHES => [%i[status]].freeze
70
+ }.freeze
71
+ end
72
+ end
73
+ end
74
+ end
@@ -3,7 +3,7 @@
3
3
  module Sidekiq
4
4
  module Batch
5
5
  module Jobs
6
- VERSION = "0.2.0"
6
+ VERSION = "0.3.0"
7
7
  end
8
8
  end
9
9
  end
@@ -2,6 +2,7 @@
2
2
 
3
3
  require "sidekiq"
4
4
  require_relative "jobs/version"
5
+ require_relative "jobs/schema"
5
6
  require_relative "jobs/enum_compat"
6
7
  require_relative "jobs/failure_policy"
7
8
  require_relative "jobs/configuration"
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: sidekiq-batch-jobs
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.2.0
4
+ version: 0.3.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Douglas Greyling
@@ -101,11 +101,14 @@ files:
101
101
  - app/models/sidekiq_batch_job.rb
102
102
  - lib/generators/sidekiq/batch/jobs/install/install_generator.rb
103
103
  - lib/generators/sidekiq/batch/jobs/install/templates/create_sidekiq_batch_tables.rb.tt
104
+ - lib/generators/sidekiq/batch/jobs/upgrade/templates/upgrade_sidekiq_batch_tables.rb.tt
105
+ - lib/generators/sidekiq/batch/jobs/upgrade/upgrade_generator.rb
104
106
  - lib/sidekiq/batch/jobs.rb
105
107
  - lib/sidekiq/batch/jobs/configuration.rb
106
108
  - lib/sidekiq/batch/jobs/engine.rb
107
109
  - lib/sidekiq/batch/jobs/enum_compat.rb
108
110
  - lib/sidekiq/batch/jobs/failure_policy.rb
111
+ - lib/sidekiq/batch/jobs/schema.rb
109
112
  - lib/sidekiq/batch/jobs/version.rb
110
113
  homepage: https://github.com/douglasgreyling/sidekiq-batch-jobs
111
114
  licenses: