workhorse 1.5.2 → 2.0.0.rc0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +102 -0
  3. data/README.md +303 -74
  4. data/Rakefile +1 -0
  5. data/VERSION +1 -1
  6. data/bin/rubocop +5 -1
  7. data/lib/generators/workhorse/install_generator.rb +10 -1
  8. data/lib/generators/workhorse/templates/config/initializers/workhorse.rb +55 -0
  9. data/lib/generators/workhorse/templates/create_table_jobs.rb +11 -13
  10. data/lib/generators/workhorse/templates/create_table_workhorse_schedules.rb +29 -0
  11. data/lib/workhorse/daemon.rb +52 -6
  12. data/lib/workhorse/db_job.rb +58 -7
  13. data/lib/workhorse/enqueuer.rb +51 -8
  14. data/lib/workhorse/jobs/cleanup_succeeded_jobs.rb +26 -8
  15. data/lib/workhorse/jobs/detect_late_schedules_job.rb +59 -0
  16. data/lib/workhorse/notifiers/base.rb +55 -0
  17. data/lib/workhorse/notifiers/file_system.rb +66 -0
  18. data/lib/workhorse/notifiers/none.rb +8 -0
  19. data/lib/workhorse/notifiers/redis.rb +227 -0
  20. data/lib/workhorse/performer.rb +28 -0
  21. data/lib/workhorse/poller.rb +289 -47
  22. data/lib/workhorse/pool.rb +12 -6
  23. data/lib/workhorse/schedule.rb +288 -0
  24. data/lib/workhorse/schedules.rb +197 -0
  25. data/lib/workhorse/worker.rb +77 -27
  26. data/lib/workhorse.rb +136 -0
  27. data/test/lib/db_schema.rb +21 -1
  28. data/test/lib/jobs.rb +29 -0
  29. data/test/lib/test_helper.rb +9 -14
  30. data/test/workhorse/daemon_test.rb +33 -0
  31. data/test/workhorse/db_job_test.rb +1 -1
  32. data/test/workhorse/notifier_test.rb +500 -0
  33. data/test/workhorse/poller_test.rb +8 -4
  34. data/test/workhorse/schedule_test.rb +967 -0
  35. data/test/workhorse/worker_test.rb +92 -0
  36. data/workhorse.gemspec +6 -5
  37. metadata +29 -3
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 9ffbb2c24919d7b745fd87bd6a8675fdae5448a6c0e11c3af486d34baf5a9343
4
- data.tar.gz: 9cf2270f0c3a65d244d679baf56367033e314e4a8b7ed11cc17f4d07e4a6fe63
3
+ metadata.gz: 69859406636c3463da70d9e274f023e5b0838720e428f77b8f0babad2e15baf9
4
+ data.tar.gz: 6af9c55d5f9066be6801898d3786311d3a0f5c33723baf3a5e4d464c107ace85
5
5
  SHA512:
6
- metadata.gz: 22b02451a64ae383b75616b9556cc05927f479f47db7d7993347df0f520d36ebfda8c7f3f4de49574c84499dfd603285bab83295df72c6100b0f6f438730d576
7
- data.tar.gz: '0549ab8a2e2683dd13b60bc85a5d2d85a260dee938a13714c45fe1233af217a433af2e71f899a7056c3afd8f503f6f94e1a5b9cee5e97bbf5fb93fed9ae75a03'
6
+ metadata.gz: b674430cc63891334e49f4ce52b1cc407dea1fdd5575beab8792b82c4f12729596bfb16829ebd859bfefc2881a701e75ff85d8bca8e1e479a081b50e1bc7d0e0
7
+ data.tar.gz: 3db6d62ef9a9dbd950dcb40ac6399a6d28ab166bc75a8b592153de315672423f76b2a6518e58d219de789edf1cc2faa81368f568c637585e5d2a05247b626e14
data/CHANGELOG.md CHANGED
@@ -1,5 +1,107 @@
1
1
  # Workhorse Changelog
2
2
 
3
+ ## 2.0.0.rc0 - 2026-09-29
4
+
5
+ Sitrox reference: #154443.
6
+
7
+ ### Breaking changes
8
+
9
+ * **Support for Oracle is dropped.** Workhorse supports MySQL and MariaDB
10
+ only. The Oracle branches of the global lock, the row limiting and the
11
+ generated migrations are gone, along with
12
+ `Workhorse::Poller::ORACLE_LOCK_MODE` and `ORACLE_LOCK_HANDLE`. It was never
13
+ covered by CI, so it only ever had manual verification. An Oracle
14
+ installation has no upgrade path and should stay on 1.x. See
15
+ [Database support](README.md#database-support).
16
+
17
+ * A forked daemon worker no longer runs the `at_exit` handlers registered by
18
+ the process that started it. They belong to that process, and one that waits
19
+ on threads the fork did not inherit hangs a worker which has already
20
+ finished - which then ignores `TERM`. This skips interpreter finalisation as
21
+ a whole, so buffered output is dropped too; a worker that dies of an
22
+ unhandled exception now reports it through `Workhorse.on_exception` and
23
+ exits non-zero.
24
+
25
+ ### Added
26
+
27
+ * *Scheduling*: workhorse runs jobs on a cron schedule itself, without an
28
+ external scheduler process. Each schedule owns a row in the new
29
+ `workhorse_schedules` table holding the occurrence it is waiting for, so an
30
+ occurrence whose time passes while nothing is running is not lost, and what
31
+ happens to it is a per-schedule `catch_up` policy. Timezones and daylight
32
+ saving are handled. See [Scheduling](README.md#scheduling).
33
+
34
+ * *Notifications*: enqueuing a job announces it and a waiting worker polls
35
+ straight away, rather than sleeping out its polling interval. Polling stays
36
+ the floor. `:file` suits workers sharing a filesystem with the application
37
+ and `:redis` those that do not; both are off by default. See
38
+ [Notifications](README.md#notifications).
39
+
40
+ * Job deadlines and lateness reporting: `expires_at`, `max_lateness`, the new
41
+ `expired` state, `Workhorse.on_job_expired`, `Workhorse.on_job_late` and
42
+ `Workhorse::DbJob#lateness`. A job past its deadline is expired rather than
43
+ performed. See [Lateness and deadlines](README.md#lateness-and-deadlines).
44
+
45
+ * `Workhorse::Jobs::DetectLateSchedulesJob`, which reports schedules whose
46
+ next occurrence lies well in the past. Neither callback above can fire for a
47
+ job that was never created, so this is what catches materialization having
48
+ stopped. See
49
+ [Detecting schedules that stopped](README.md#detecting-schedules-that-stopped).
50
+
51
+ * `Workhorse.shutdown_timeout`, the seconds the daemon's `stop` waits for a
52
+ worker before killing it. Defaults to 300, `nil` restores the previous
53
+ behaviour. A worker that ignores `TERM` used to leave `stop` - and whatever
54
+ waits on it, usually a deployment - looping forever.
55
+
56
+ * `Workhorse.enqueue_job_class`, the keyword arguments `expires_at:` and
57
+ `max_lateness:` on `Workhorse.enqueue` and `Workhorse.enqueue_active_job`,
58
+ `priority:` on the latter, and the scope `Workhorse::DbJob.expired`.
59
+
60
+ * The composite indexes `[state, perform_at]`, `[state, priority, created_at]`
61
+ and `[state, expires_at]` in the generated `jobs` migration, replacing the
62
+ single-column index on `state`. `rails generate workhorse:install` now emits
63
+ two migrations rather than one.
64
+
65
+ * `fugit` as a runtime dependency, for parsing cron expressions.
66
+
67
+ ### Changed
68
+
69
+ * `Workhorse::Jobs::CleanupSucceededJobs` also deletes jobs in the new
70
+ `expired` state, with a `states` argument to opt out. A schedule using
71
+ `expires_after` that regularly misses its window would otherwise grow the
72
+ jobs table without bound.
73
+
74
+ * Failures to obtain the global lock on a poll that a notification or an
75
+ instant repoll brought forward no longer count towards
76
+ `max_global_lock_fails`, and are logged at `debug`. Several workers woken by
77
+ one announcement race for the lock and all but one lose, which says nothing
78
+ about a crashed worker.
79
+
80
+ * `Workhorse::DbJob#reset!` accepts `expired` as the terminal state it is.
81
+
82
+ ### Fixed
83
+
84
+ * A deadlock between shutting a worker down and the poller posting a job.
85
+ `Worker#shutdown` held the worker's mutex while waiting for the poller
86
+ thread, which could be waiting for that same mutex in `Worker#perform`. The
87
+ worker then ignored `TERM`. A job that was locked but cannot be performed
88
+ because the worker is shutting down is now reset to `waiting` rather than
89
+ left locked, where it would have blocked its queue.
90
+
91
+ * `Worker#shutdown` raising when called concurrently, which the daemon does by
92
+ sending both `TERM` and `INT`.
93
+
94
+ ### Documentation
95
+
96
+ * PostgreSQL is documented as unsupported. Workers emit `GET_LOCK` on every
97
+ poll, which PostgreSQL does not provide, so a worker fails on its first one.
98
+ The requirements previously listed it as supported, which it never was.
99
+
100
+ ### Upgrading
101
+
102
+ Nothing breaks without migrating, but the new features need one. See
103
+ [Upgrading from 1.x](README.md#upgrading-from-1x).
104
+
3
105
  ## 1.5.2 - 2026-08-04
4
106
 
5
107
  * Fix `Poller#valid_queues` raising `NoMethodError` on the Oracle adapter. The
data/README.md CHANGED
@@ -35,8 +35,8 @@ What it does not do:
35
35
 
36
36
  * Ruby `>= 3.0` (may work with earlier versions but is untested)
37
37
  * Rails `>= 7.0`
38
- * A database and table handler that properly supports row-level locking (such as
39
- MySQL with InnoDB, PostgreSQL, or Oracle).
38
+ * MySQL or MariaDB with InnoDB. No other database is supported, see
39
+ [Database support](#database-support).
40
40
  * If you are planning on using the daemons handler:
41
41
  * An operating system and file system that supports file locking.
42
42
  * MRI Ruby (aka "CRuby") as jRuby does not support `fork`. See the
@@ -60,21 +60,80 @@ What it does not do:
60
60
 
61
61
  This generates:
62
62
 
63
- * A database migration for creating a table named `jobs`
63
+ * Two database migrations, creating the tables `jobs` and
64
+ `workhorse_schedules`
64
65
  * The initializer `config/initializers/workhorse.rb` for global configuration
65
66
  * This can be skipped using the `--skip-initializer` flag
66
67
  * The daemon worker script `bin/workhorse.rb`
67
68
 
68
69
  Please customize the initializer and worker script to your liking.
69
70
 
70
- ### Oracle
71
+ ### Database support
71
72
 
72
- When using Oracle databases, make sure your schema has access to the package
73
- `DBMS_LOCK`:
73
+ **MySQL and MariaDB are the only supported databases.** Workhorse serialises
74
+ job pickup with `GET_LOCK`, a MySQL advisory lock, which is emitted on every
75
+ poll; a database that does not provide it fails on the first poll. InnoDB is
76
+ required, as MyISAM supports neither transactions nor row-level locking. Both
77
+ the `mysql2` and the `trilogy` adapter are covered by CI.
74
78
 
79
+ Oracle was supported until 2.0.0 and is not any more — see the changelog entry
80
+ for that release. PostgreSQL has never been supported, despite the
81
+ requirements once listing it; supporting it would mean an advisory-lock
82
+ dialect of its own (`pg_advisory_lock`) and is not currently planned.
83
+
84
+ ## Upgrading from 1.x
85
+
86
+ Workhorse 2.0 drops support for Oracle, see
87
+ [Database support](#database-support). An Oracle installation has no upgrade
88
+ path and should stay on 1.x.
89
+
90
+ Otherwise nothing breaks without migrating: [scheduling](#scheduling) is inert
91
+ until a schedule is declared, [notifications](#notifications) are off until a
92
+ notifier is selected, and the new job columns are only written when used. To
93
+ take the new features up, add this migration:
94
+
95
+ ```ruby
96
+ class UpgradeWorkhorseToV2 < ActiveRecord::Migration[7.1]
97
+ def change
98
+ add_column :jobs, :expires_at, :datetime, null: true
99
+ add_column :jobs, :max_lateness, :integer, null: true
100
+
101
+ create_table :workhorse_schedules do |t|
102
+ t.string :key, null: false
103
+ t.string :cron, null: false
104
+ t.string :timezone, null: true
105
+ t.boolean :enabled, null: false, default: true
106
+ t.datetime :next_at, null: false
107
+ t.datetime :last_enqueued_at, null: true
108
+ t.datetime :last_occurrence, null: true
109
+ t.integer :last_job_id, null: true
110
+ t.timestamps null: false
111
+ end
112
+
113
+ # The index names are given explicitly because the ones Rails would
114
+ # derive are longer than some tools accept.
115
+ add_index :workhorse_schedules, :key,
116
+ unique: true, length: 191, name: 'idx_wh_schedules_key'
117
+ add_index :workhorse_schedules, %i[enabled next_at],
118
+ name: 'idx_wh_schedules_due'
119
+
120
+ add_index :jobs, %i[state perform_at],
121
+ length: { state: 191 }, name: 'idx_jobs_state_perform_at'
122
+ add_index :jobs, %i[state priority created_at],
123
+ length: { state: 191 }, name: 'idx_jobs_state_prio_created'
124
+ add_index :jobs, %i[state expires_at],
125
+ length: { state: 191 }, name: 'idx_jobs_state_expires_at'
126
+
127
+ # Now redundant, as `state` leads both indexes above
128
+ remove_index :jobs, :state
129
+ end
130
+ end
75
131
  ```
76
- GRANT execute ON DBMS_LOCK TO <schema-name>;
77
- ```
132
+
133
+ Two behaviour changes worth knowing about, both described in the changelog: a
134
+ forked daemon worker no longer runs `at_exit` handlers registered by the
135
+ process that started it, and `Workhorse::Jobs::CleanupSucceededJobs` now also
136
+ deletes jobs in the new `expired` state.
78
137
 
79
138
  ## Queuing jobs
80
139
 
@@ -122,86 +181,155 @@ If you do not want to pass any parameters to the operation, just omit the third
122
181
  Workhorse.enqueue_op Operations::Jobs::CleanUpDatabase, queue: :maintenance, priority: 2
123
182
  ```
124
183
 
125
- ### Scheduling
184
+ ## Scheduling
126
185
 
127
- Workhorse has no out-of-the-box functionality to support scheduling of regular
128
- jobs, such as maintenance or backup jobs. There are two primary ways of
129
- achieving regular execution:
186
+ Workhorse runs jobs on a schedule itself, without an external scheduler
187
+ process. Schedules are declared in code and their state is kept in the
188
+ database:
130
189
 
131
- 1. Rescheduling by the same job after successful execution and setting
132
- `perform_at`
190
+ ```ruby
191
+ # config/initializers/workhorse.rb
192
+ Workhorse.schedules do
193
+ schedule 'cleanup_jobs',
194
+ job: 'Workhorse::Jobs::CleanupSucceededJobs',
195
+ cron: '10 0 * * *'
196
+
197
+ schedule 'morning_digest',
198
+ job: 'Jobs::MorningDigest',
199
+ cron: '0 8 * * 1-5',
200
+ timezone: 'Europe/Zurich',
201
+ queue: :reports,
202
+ priority: -10,
203
+ catch_up: :skip,
204
+ grace: 15.minutes,
205
+ max_lateness: 60.seconds
206
+ end
207
+ ```
133
208
 
134
- This is simple to set up and requires no additional dependencies. However,
135
- the time taken to execute a job and the time delay caused by the polling
136
- interval cannot easily be factored into the calculation of the interval,
137
- leading to a slight shift in effective execution date. (This can be mitigated
138
- by scheduling the job before knowing whether the current run will succeed.
139
- Proceed down this path at your own peril!)
209
+ Each schedule owns a row in `workhorse_schedules` holding the next occurrence
210
+ that has not been materialised yet. Workers reconcile those rows against the
211
+ declarations above on startup, and materialise the occurrences that have come
212
+ due during their regular poll. There is no scheduler process to keep alive and
213
+ no single point of failure: any worker will do.
140
214
 
141
- *Example:* A job that takes 5 seconds to run and reschedules itself every
142
- 10 minutes. If started at 12:00 sharp, after one hour it will execute at
143
- 13:00:30 at the earliest due to cumulative execution time.
215
+ ### Why the occurrence is a row
144
216
 
145
- In its most basic form, the `perform` method of a job would look as follows:
217
+ An in-memory scheduler computes the next occurrence from *now*, so an
218
+ occurrence whose time passes while it is not running never happens and leaves
219
+ no trace — a deployment, a restart or a crash at the wrong minute silently
220
+ skips a nightly job. Because the next occurrence is persisted here, a worker
221
+ coming back at any later point still sees that it is due, and the schedule
222
+ decides what to do about it.
146
223
 
147
- ```ruby
148
- class MyJob
149
- def perform
150
- # Do all the work
224
+ ### Catch-up
151
225
 
152
- # Perform again after 10 minutes (600 seconds)
153
- Workhorse.enqueue MyJob.new, perform_at: Time.now + 600
154
- end
155
- end
156
- ```
226
+ What should happen to an occurrence whose time has passed depends on the job,
227
+ so it is stated per schedule:
157
228
 
158
- 2. Using an external scheduler
229
+ | `catch_up` | Behaviour | Suits |
230
+ |-------------|----------------------------------------------------------|-----------------------------------------|
231
+ | `:run_once` | Collapse all missed occurrences into one (**default**) | Cleanup, maintenance, idempotent work |
232
+ | `:run` | Materialise each, up to `max_catch_up` (default 10) | Per-period reports that must all exist |
233
+ | `:skip` | Drop those older than `grace` | "Send the 08:00 digest" |
159
234
 
160
- A more elaborate setup requires an external scheduler, but which can still be
161
- called from Ruby. One such scheduler is
162
- [rufus-scheduler](https://github.com/jmettraux/rufus-scheduler). A small
163
- example of an adapted `bin/workhorse.rb` to accommodate for the additional
164
- cog in the mechanism is given below:
235
+ `:run_once` is the default deliberately: after a long outage it is the safe
236
+ behaviour. A schedule running every minute that was down for a day would
237
+ otherwise enqueue 1440 jobs at once, which `max_catch_up` also guards against.
165
238
 
166
- ```ruby
167
- #!/usr/bin/env ruby
239
+ `:skip` requires `grace`, as without one it has no way to tell an occurrence
240
+ that is a moment late from one that is a day late.
168
241
 
169
- require './config/environment'
242
+ ### Lateness and deadlines
170
243
 
171
- Workhorse::Daemon::ShellHandler.run do |daemon|
172
- # Start scheduler process
173
- daemon.worker 'Scheduler' do
174
- scheduler = Rufus::Scheduler.new
244
+ A materialised job's `perform_at` is the occurrence's own time, not the moment
245
+ it was enqueued. The difference between it and `started_at` is therefore the
246
+ lateness of that occurrence, available on every job as
247
+ `Workhorse::DbJob#lateness`, and derivable in SQL from the `started_at` and
248
+ `perform_at` columns.
175
249
 
176
- scheduler.cron '0/10 * * * *' do
177
- Workhorse.enqueue Workhorse::Jobs::CleanupSucceededJobs.new
178
- end
250
+ Two options act on it, both settable on a schedule:
179
251
 
180
- Signal.trap 'TERM' do
181
- scheduler.shutdown
182
- end
252
+ * **`max_lateness`** — seconds the job may start late before
253
+ `Workhorse.on_job_late` is called. The job still runs; it was just late.
254
+ * **`expires_after`** — seconds after the occurrence at which the job is no
255
+ longer worth running. It is then set to state `expired` and
256
+ `Workhorse.on_job_expired` is called instead of it being performed. For
257
+ "send the 08:00 reminder", running it at 11:40 is often worse than not
258
+ running it at all.
183
259
 
184
- scheduler.join
185
- end
260
+ Hand-enqueued jobs take `max_lateness:` under the same name, and the deadline
261
+ as an absolute time rather than an offset:
186
262
 
187
- # Start 5 worker processes with 3 threads each
188
- 5.times do
189
- daemon.worker do
190
- Workhorse::Worker.start_and_wait(pool_size: 3, polling_interval: 10, logger: Rails.logger)
191
- end
192
- end
193
- end
194
- ```
263
+ ```ruby
264
+ Workhorse.enqueue(job, expires_at: 1.hour.from_now, max_lateness: 60)
265
+ ```
266
+
267
+ ```ruby
268
+ Workhorse.setup do |config|
269
+ config.on_job_expired = proc do |db_job|
270
+ ExceptionNotifier.notify_exception(
271
+ StandardError.new("Job #{db_job.id} (#{db_job.description}) expired")
272
+ )
273
+ end
195
274
 
196
- This allows starting and stopping the daemon with the usual interface.
197
- Note that the scheduler is handled like a Workhorse worker, the consequence
198
- of which is that only one 'worker' should be started by the ShellHandler.
199
- Otherwise there would be multiple jobs scheduled at the same time.
275
+ config.on_job_late = proc do |db_job, lateness|
276
+ ExceptionNotifier.notify_exception(
277
+ StandardError.new("Job #{db_job.id} started #{lateness.round}s late")
278
+ )
279
+ end
280
+ end
281
+ ```
282
+
283
+ Both callbacks are best-effort: anything they raise goes to
284
+ `Workhorse.on_exception` and never affects the worker or the job. An expiry is
285
+ logged at `warn` whether or not a callback is configured.
286
+
287
+ Expired jobs stay in the table like any other finished job.
288
+ `Workhorse::Jobs::CleanupSucceededJobs` removes them along with succeeded ones;
289
+ pass `states: [Workhorse::DbJob::STATE_SUCCEEDED]` to keep them.
200
290
 
201
- Please refer to the documentation for
202
- [rufus-scheduler](https://github.com/jmettraux/rufus-scheduler) (or the
203
- scheduler of your choice) for further options concerning the timing of the
204
- jobs.
291
+ ### Detecting schedules that stopped
292
+
293
+ Neither callback can fire for a job that was never created, so if no worker is
294
+ polling or the global lock is stuck, occurrences simply stop being
295
+ materialised and nothing says so. `Workhorse::Jobs::DetectLateSchedulesJob`
296
+ covers that case by reporting schedules whose next occurrence lies well in the
297
+ past:
298
+
299
+ ```ruby
300
+ Workhorse.schedules do
301
+ schedule 'detect_late_schedules',
302
+ job: 'Workhorse::Jobs::DetectLateSchedulesJob',
303
+ cron: '*/30 * * * *'
304
+ end
305
+ ```
306
+
307
+ It is itself performed by a worker, so it reports a stall only while at least
308
+ one worker is still running — use it alongside external monitoring rather than
309
+ instead of it.
310
+
311
+ ### Timezones
312
+
313
+ Without `timezone`, a cron expression is read in the process's local time.
314
+ Given one, occurrences are computed in that zone, including across daylight
315
+ saving changes. A `30 2 * * *` schedule in `Europe/Zurich` has no occurrence
316
+ on the day the clocks go forward, because 02:30 does not exist that day — and
317
+ exactly one on the day they go back, although 02:30 happens twice.
318
+
319
+ ### Disabling a schedule
320
+
321
+ Setting `enabled` to `false` on the row stops its occurrences from being
322
+ materialised, without a deployment:
323
+
324
+ ```ruby
325
+ Workhorse::Schedule.find_by(key: 'morning_digest').update!(enabled: false)
326
+ ```
327
+
328
+ Removing a schedule from the declarations deletes its row: the next worker to
329
+ start after a day has passed without any worker declaring it removes it. The delay matters during a rolling deployment, where
330
+ the old and the new version run at once: were rows removed immediately, each
331
+ version would delete the other's schedules and reset the occurrences they were
332
+ waiting for.
205
333
 
206
334
  ## Configuring and starting workers
207
335
 
@@ -278,6 +406,100 @@ polling interval.
278
406
  This setting is recommended for all setups and may eventually be enabled by
279
407
  default.
280
408
 
409
+ ### Notifications
410
+
411
+ Instant repolling only helps *after* a job has been performed. A worker that is
412
+ idle and waiting for new work still sleeps out its entire polling interval, so
413
+ a job enqueued just after a poll waits almost a full interval before it starts.
414
+
415
+ Shortening the polling interval is the obvious remedy and a poor one: every
416
+ poll acquires a global database lock, so more frequent polling across several
417
+ workers increases contention, and a worker that fails to acquire the lock skips
418
+ its poll and waits another whole interval. Polling costs the same whether or
419
+ not anything is happening.
420
+
421
+ *Notifications* turn the question around. Enqueuing a job announces it, and a
422
+ waiting worker polls straight away instead of sleeping out its interval:
423
+
424
+ ```ruby
425
+ # config/initializers/workhorse.rb
426
+ Workhorse.setup do |config|
427
+ config.notifier = :file
428
+ end
429
+ ```
430
+
431
+ Polling remains the floor. A notification that is never delivered — a worker
432
+ that was restarting, an enqueue from a host that cannot reach the others —
433
+ costs latency and nothing else, as the regular poll still finds the job. For
434
+ the same reason, raise `polling_interval` only as far as you are willing to
435
+ wait when a notification *is* missed.
436
+
437
+ Note what this does and does not do to database load. A notification only
438
+ brings the next poll forward; it never replaces one, so the saving comes from
439
+ raising `polling_interval`, not from notifications by themselves. An
440
+ announcement also wakes *every* worker with a free thread, and all but the one
441
+ that wins the job take the global lock for nothing. So a workload that enqueues
442
+ in bursts while many workers sit idle can take the lock more often than plain
443
+ polling would, rather than less. A worker whose threads are all busy does not
444
+ react, so the effect is bounded by how much spare capacity there is.
445
+
446
+ Two notifiers ship with workhorse:
447
+
448
+ #### `:file`
449
+
450
+ Touches a single file, which waiting workers stat once per 0.1 seconds on a
451
+ tick the poller performs anyway. It costs no database work at all and around a
452
+ microsecond of CPU per check, and a job starts within roughly 100 milliseconds.
453
+
454
+ It requires that the processes enqueueing jobs and the workers share a
455
+ filesystem, which in practice means the same host. Where they do not, the
456
+ touch never reaches those workers and they fall back to polling.
457
+
458
+ The file defaults to `tmp/pids/workhorse.wake` below the Rails root and can be
459
+ moved with `config.notification_path`. Every process involved must agree on it.
460
+
461
+ #### `:redis`
462
+
463
+ Publishes on a Redis pub/sub channel, for deployments whose workers do not
464
+ share a filesystem with the application. Redis is a soft dependency: it is not
465
+ declared as a dependency of this gem and workhorse never requires it — supply
466
+ a client through `config.notification_redis`.
467
+
468
+ ```ruby
469
+ Workhorse.setup do |config|
470
+ config.notifier = :redis
471
+ config.notification_redis = -> { Redis.new(url: ENV['REDIS_URL']) }
472
+ config.notification_channel = 'workhorse:jobs' # optional
473
+ end
474
+ ```
475
+
476
+ `notification_redis` is given something callable above because a subscribed
477
+ Redis connection cannot be used for anything else: `subscribe` occupies it
478
+ until it returns. The notifier calls it once for publishing and once for the
479
+ subscriber, so the two get separate connections.
480
+
481
+ Passing a client rather than a callable also works, but the subscriber then
482
+ falls back to `dup` — in redis-rb a shallow copy that may share the
483
+ connection, in which case publishing can block behind the subscription.
484
+
485
+ The channel can be changed at any point before a worker starts; the
486
+ subscriber thread captures the one it was started with.
487
+
488
+ #### Writing your own
489
+
490
+ Subclass `Workhorse::Notifiers::Base` and assign an instance to
491
+ `config.notifier`. A notifier announces jobs with `notify` and exposes a
492
+ `token` that changes whenever a notification has arrived; workers compare it
493
+ against the last value they saw, so several workers in one process stay
494
+ independent of one another. `notify(queue: nil)` must never raise — a job must
495
+ still be enqueued when it cannot be announced. `start` and `stop` are optional
496
+ hooks, called as a worker starts and shuts down, for anything that needs a
497
+ thread or a connection of its own.
498
+
499
+ Note that jobs with a future `perform_at` are not announced, as a woken worker
500
+ would find nothing to do; they are picked up by the regular poll once they are
501
+ due.
502
+
281
503
  ## Transactions
282
504
 
283
505
  By default, each job is run in an individual database transaction. An exception
@@ -412,6 +634,7 @@ DbJob.locked
412
634
  DbJob.started
413
635
  DbJob.succeeded
414
636
  DbJob.failed
637
+ DbJob.expired
415
638
  ```
416
639
  ### Resetting jobs
417
640
 
@@ -419,8 +642,8 @@ Jobs in a state other than `waiting` are either being processed or else already
419
642
  in a final state such as `succeeded` and won't be performed again. Workhorse
420
643
  provides an API method for resetting jobs in the following cases:
421
644
 
422
- * A job has succeeded or failed (states `succeeded` and `failed`) and needs to
423
- re-run. In these cases, perform a non-forced reset:
645
+ * A job has succeeded, failed or expired (states `succeeded`, `failed` and
646
+ `expired`) and needs to re-run. In these cases, perform a non-forced reset:
424
647
 
425
648
  ```ruby
426
649
  db_job.reset!
@@ -461,8 +684,9 @@ configuration or else using `self.queue_adapter` in a job class inheriting from
461
684
  Per default, jobs remain in the database, no matter in which state. This can
462
685
  eventually lead to a very large jobs database. You are advised to clean your
463
686
  jobs database on a regular interval. Workhorse provides the job
464
- `Workhorse::Jobs::CleanupSucceededJobs` for this purpose that cleans up all
465
- succeeded jobs. You can run this using your scheduler in a specific interval.
687
+ `Workhorse::Jobs::CleanupSucceededJobs` for this purpose, which cleans up
688
+ succeeded and expired jobs — pass `states:` to narrow that. Schedule it with
689
+ `Workhorse.schedules`, see [Scheduling](#scheduling).
466
690
 
467
691
  ## Memory handling
468
692
 
@@ -587,6 +811,11 @@ In the event that this still happens, Workhorse takes the following steps:
587
811
  - Retries acquiring the lock on the next poll.
588
812
  - Calls the `on_exception` callback (if configured) after a configurable number of consecutive failures to obtain the lock.
589
813
 
814
+ Failures on a poll that a notification or an instant repoll brought forward
815
+ are expected — several workers race for the lock and all but one lose — so
816
+ they are logged at `debug` only and do not count towards
817
+ `max_global_lock_fails`.
818
+
590
819
  The maximum number of consecutive failures can be configured using
591
820
  `config.max_global_lock_fails`, which defaults to 10.
592
821
 
data/Rakefile CHANGED
@@ -16,6 +16,7 @@ task :gemspec do
16
16
  spec.add_dependency 'activesupport', '>= 7.0.0'
17
17
  spec.add_dependency 'activerecord', '>= 7.0.0'
18
18
  spec.add_dependency 'concurrent-ruby'
19
+ spec.add_dependency 'fugit'
19
20
  end
20
21
 
21
22
  File.write('workhorse.gemspec', gemspec.to_ruby.strip)
data/VERSION CHANGED
@@ -1 +1 @@
1
- 1.5.2
1
+ 2.0.0.rc0
data/bin/rubocop CHANGED
@@ -1 +1,5 @@
1
- bundle exec rubocop "$@"
1
+ #!/bin/bash
2
+
3
+ BASE_PATH=`(cd $(dirname $0)/.. && pwd -P)`
4
+
5
+ BUNDLE_GEMFILE=$BASE_PATH/Gemfile $BASE_PATH/bin/ruby -e "require 'rubygems';require 'bundler/setup'; load Gem.activate_bin_path('rubocop', 'rubocop', '>= 0.a')" -- $@
@@ -7,12 +7,21 @@ module Workhorse
7
7
 
8
8
  source_root File.expand_path('templates', __dir__)
9
9
 
10
+ # Returns the version for the next generated migration. Counts up rather
11
+ # than returning the current time, as several migrations are generated
12
+ # within the same second and would otherwise collide.
10
13
  def self.next_migration_number(_dir)
11
- Time.now.utc.strftime('%Y%m%d%H%M%S')
14
+ @next_migration_number = [
15
+ Time.now.utc.strftime('%Y%m%d%H%M%S').to_i,
16
+ (@next_migration_number || 0) + 1
17
+ ].max
18
+
19
+ return @next_migration_number.to_s
12
20
  end
13
21
 
14
22
  def install_migration
15
23
  migration_template 'create_table_jobs.rb', 'db/migrate/create_table_jobs.rb'
24
+ migration_template 'create_table_workhorse_schedules.rb', 'db/migrate/create_table_workhorse_schedules.rb'
16
25
  end
17
26
 
18
27
  def install_daemon_script