workhorse 1.5.2 → 2.0.0.rc1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. checksums.yaml +4 -4
  2. data/.github/workflows/ruby.yml +137 -1
  3. data/CHANGELOG.md +150 -0
  4. data/Gemfile +16 -1
  5. data/README.md +316 -72
  6. data/Rakefile +1 -0
  7. data/VERSION +1 -1
  8. data/bin/rubocop +5 -1
  9. data/lib/generators/workhorse/install_generator.rb +10 -1
  10. data/lib/generators/workhorse/templates/config/initializers/workhorse.rb +55 -0
  11. data/lib/generators/workhorse/templates/create_table_jobs.rb +15 -2
  12. data/lib/generators/workhorse/templates/create_table_workhorse_schedules.rb +42 -0
  13. data/lib/workhorse/daemon/shell_handler.rb +4 -1
  14. data/lib/workhorse/daemon.rb +57 -7
  15. data/lib/workhorse/db_job.rb +98 -8
  16. data/lib/workhorse/enqueuer.rb +51 -8
  17. data/lib/workhorse/jobs/cleanup_succeeded_jobs.rb +26 -8
  18. data/lib/workhorse/jobs/detect_late_schedules_job.rb +59 -0
  19. data/lib/workhorse/notifiers/base.rb +55 -0
  20. data/lib/workhorse/notifiers/file_system.rb +66 -0
  21. data/lib/workhorse/notifiers/none.rb +8 -0
  22. data/lib/workhorse/notifiers/redis.rb +227 -0
  23. data/lib/workhorse/performer.rb +29 -2
  24. data/lib/workhorse/poller.rb +303 -21
  25. data/lib/workhorse/pool.rb +12 -6
  26. data/lib/workhorse/schedule.rb +288 -0
  27. data/lib/workhorse/schedules.rb +197 -0
  28. data/lib/workhorse/worker.rb +102 -31
  29. data/lib/workhorse.rb +136 -0
  30. data/test/lib/db_schema.rb +36 -3
  31. data/test/lib/jobs.rb +29 -0
  32. data/test/lib/test_helper.rb +113 -20
  33. data/test/workhorse/daemon_test.rb +33 -0
  34. data/test/workhorse/db_job_test.rb +2 -4
  35. data/test/workhorse/notifier_test.rb +487 -0
  36. data/test/workhorse/performer_test.rb +7 -9
  37. data/test/workhorse/poller_test.rb +97 -23
  38. data/test/workhorse/schedule_test.rb +967 -0
  39. data/test/workhorse/worker_test.rb +201 -76
  40. data/workhorse.gemspec +6 -5
  41. metadata +29 -3
data/README.md CHANGED
@@ -35,8 +35,8 @@ What it does not do:
35
35
 
36
36
  * Ruby `>= 3.0` (may work with earlier versions but is untested)
37
37
  * Rails `>= 7.0`
38
- * A database and table handler that properly supports row-level locking (such as
39
- MySQL with InnoDB, PostgreSQL, or Oracle).
38
+ * One of the supported databases (see [Database support](#database-support)):
39
+ MySQL / MariaDB with InnoDB, or Oracle. **PostgreSQL is not supported.**
40
40
  * If you are planning on using the daemons handler:
41
41
  * An operating system and file system that supports file locking.
42
42
  * MRI Ruby (aka "CRuby") as jRuby does not support `fork`. See the
@@ -60,22 +60,96 @@ What it does not do:
60
60
 
61
61
  This generates:
62
62
 
63
- * A database migration for creating a table named `jobs`
63
+ * Two database migrations, creating the tables `jobs` and
64
+ `workhorse_schedules`
64
65
  * The initializer `config/initializers/workhorse.rb` for global configuration
65
66
  * This can be skipped using the `--skip-initializer` flag
66
67
  * The daemon worker script `bin/workhorse.rb`
67
68
 
68
69
  Please customize the initializer and worker script to your liking.
69
70
 
70
- ### Oracle
71
+ ### Database support
71
72
 
72
- When using Oracle databases, make sure your schema has access to the package
73
- `DBMS_LOCK`:
73
+ Workhorse serialises job pickup using a database-level lock, which is
74
+ necessarily written against a specific database's dialect. Two families are
75
+ implemented:
76
+
77
+ | Database | Supported | Lock used | Covered by CI |
78
+ |-------------------|-----------|----------------------|---------------|
79
+ | MySQL / MariaDB | Yes | `GET_LOCK` | Yes, against both the `mysql2` and the `trilogy` adapter |
80
+ | Oracle 12c+ | Yes | `DBMS_LOCK` | Yes, against `activerecord-oracle_enhanced-adapter` |
81
+ | PostgreSQL | **No** | — | — |
82
+ | Everything else | **No** | — | — |
83
+
84
+ There is no PostgreSQL implementation: workers emit `GET_LOCK` on every poll,
85
+ which PostgreSQL does not provide, so a worker fails on its first poll.
86
+ Supporting it would mean an advisory-lock dialect of its own
87
+ (`pg_advisory_lock`) and is not currently planned. Note that InnoDB is required
88
+ on MySQL / MariaDB, as MyISAM supports neither transactions nor row-level
89
+ locking.
90
+
91
+ Oracle 12c is the minimum, as job selection limits its rows with
92
+ `FETCH FIRST … ROWS ONLY`. When using Oracle, make sure your schema has access
93
+ to the package `DBMS_LOCK`:
74
94
 
75
95
  ```
76
96
  GRANT execute ON DBMS_LOCK TO <schema-name>;
77
97
  ```
78
98
 
99
+ ## Upgrading from 1.x
100
+
101
+ Nothing breaks without migrating: [scheduling](#scheduling) is inert
102
+ until a schedule is declared, [notifications](#notifications) are off until a
103
+ notifier is selected, and the new job columns are only written when used. To
104
+ take the new features up, add this migration:
105
+
106
+ ```ruby
107
+ class UpgradeWorkhorseToV2 < ActiveRecord::Migration[7.1]
108
+ def change
109
+ add_column :jobs, :expires_at, :datetime, null: true
110
+ add_column :jobs, :max_lateness, :integer, null: true
111
+
112
+ create_table :workhorse_schedules do |t|
113
+ t.string :key, null: false
114
+ t.string :cron, null: false
115
+ t.string :timezone, null: true
116
+ t.boolean :enabled, null: false, default: true
117
+ t.datetime :next_at, null: false
118
+ t.datetime :last_enqueued_at, null: true
119
+ t.datetime :last_occurrence, null: true
120
+ t.integer :last_job_id, null: true
121
+ t.timestamps null: false
122
+ end
123
+
124
+ # The index names are given explicitly because the ones Rails would
125
+ # derive are longer than some tools accept.
126
+ add_index :workhorse_schedules, :key,
127
+ unique: true, length: 191, name: 'idx_wh_schedules_key'
128
+ add_index :workhorse_schedules, %i[enabled next_at],
129
+ name: 'idx_wh_schedules_due'
130
+
131
+ add_index :jobs, %i[state perform_at],
132
+ length: { state: 191 }, name: 'idx_jobs_state_perform_at'
133
+ add_index :jobs, %i[state priority created_at],
134
+ length: { state: 191 }, name: 'idx_jobs_state_prio_created'
135
+ add_index :jobs, %i[state expires_at],
136
+ length: { state: 191 }, name: 'idx_jobs_state_expires_at'
137
+
138
+ # Now redundant, as `state` leads both indexes above
139
+ remove_index :jobs, :state
140
+ end
141
+ end
142
+ ```
143
+
144
+ On Oracle, drop every `length:` option above — it indexes the whole column and
145
+ rejects a prefix length. The generated migrations do this for you; this one is
146
+ written out by hand.
147
+
148
+ Two behaviour changes worth knowing about, both described in the changelog: a
149
+ forked daemon worker no longer runs `at_exit` handlers registered by the
150
+ process that started it, and `Workhorse::Jobs::CleanupSucceededJobs` now also
151
+ deletes jobs in the new `expired` state.
152
+
79
153
  ## Queuing jobs
80
154
 
81
155
  ### Basic jobs
@@ -122,86 +196,155 @@ If you do not want to pass any parameters to the operation, just omit the third
122
196
  Workhorse.enqueue_op Operations::Jobs::CleanUpDatabase, queue: :maintenance, priority: 2
123
197
  ```
124
198
 
125
- ### Scheduling
199
+ ## Scheduling
126
200
 
127
- Workhorse has no out-of-the-box functionality to support scheduling of regular
128
- jobs, such as maintenance or backup jobs. There are two primary ways of
129
- achieving regular execution:
201
+ Workhorse runs jobs on a schedule itself, without an external scheduler
202
+ process. Schedules are declared in code and their state is kept in the
203
+ database:
130
204
 
131
- 1. Rescheduling by the same job after successful execution and setting
132
- `perform_at`
205
+ ```ruby
206
+ # config/initializers/workhorse.rb
207
+ Workhorse.schedules do
208
+ schedule 'cleanup_jobs',
209
+ job: 'Workhorse::Jobs::CleanupSucceededJobs',
210
+ cron: '10 0 * * *'
211
+
212
+ schedule 'morning_digest',
213
+ job: 'Jobs::MorningDigest',
214
+ cron: '0 8 * * 1-5',
215
+ timezone: 'Europe/Zurich',
216
+ queue: :reports,
217
+ priority: -10,
218
+ catch_up: :skip,
219
+ grace: 15.minutes,
220
+ max_lateness: 60.seconds
221
+ end
222
+ ```
133
223
 
134
- This is simple to set up and requires no additional dependencies. However,
135
- the time taken to execute a job and the time delay caused by the polling
136
- interval cannot easily be factored into the calculation of the interval,
137
- leading to a slight shift in effective execution date. (This can be mitigated
138
- by scheduling the job before knowing whether the current run will succeed.
139
- Proceed down this path at your own peril!)
224
+ Each schedule owns a row in `workhorse_schedules` holding the next occurrence
225
+ that has not been materialised yet. Workers reconcile those rows against the
226
+ declarations above on startup, and materialise the occurrences that have come
227
+ due during their regular poll. There is no scheduler process to keep alive and
228
+ no single point of failure: any worker will do.
140
229
 
141
- *Example:* A job that takes 5 seconds to run and reschedules itself every
142
- 10 minutes. If started at 12:00 sharp, after one hour it will execute at
143
- 13:00:30 at the earliest due to cumulative execution time.
230
+ ### Why the occurrence is a row
144
231
 
145
- In its most basic form, the `perform` method of a job would look as follows:
232
+ An in-memory scheduler computes the next occurrence from *now*, so an
233
+ occurrence whose time passes while it is not running never happens and leaves
234
+ no trace — a deployment, a restart or a crash at the wrong minute silently
235
+ skips a nightly job. Because the next occurrence is persisted here, a worker
236
+ coming back at any later point still sees that it is due, and the schedule
237
+ decides what to do about it.
146
238
 
147
- ```ruby
148
- class MyJob
149
- def perform
150
- # Do all the work
239
+ ### Catch-up
151
240
 
152
- # Perform again after 10 minutes (600 seconds)
153
- Workhorse.enqueue MyJob.new, perform_at: Time.now + 600
154
- end
155
- end
156
- ```
241
+ What should happen to an occurrence whose time has passed depends on the job,
242
+ so it is stated per schedule:
157
243
 
158
- 2. Using an external scheduler
244
+ | `catch_up` | Behaviour | Suits |
245
+ |-------------|----------------------------------------------------------|-----------------------------------------|
246
+ | `:run_once` | Collapse all missed occurrences into one (**default**) | Cleanup, maintenance, idempotent work |
247
+ | `:run` | Materialise each, up to `max_catch_up` (default 10) | Per-period reports that must all exist |
248
+ | `:skip` | Drop those older than `grace` | "Send the 08:00 digest" |
159
249
 
160
- A more elaborate setup requires an external scheduler, but which can still be
161
- called from Ruby. One such scheduler is
162
- [rufus-scheduler](https://github.com/jmettraux/rufus-scheduler). A small
163
- example of an adapted `bin/workhorse.rb` to accommodate for the additional
164
- cog in the mechanism is given below:
250
+ `:run_once` is the default deliberately: after a long outage it is the safe
251
+ behaviour. A schedule running every minute that was down for a day would
252
+ otherwise enqueue 1440 jobs at once, which `max_catch_up` also guards against.
165
253
 
166
- ```ruby
167
- #!/usr/bin/env ruby
254
+ `:skip` requires `grace`, as without one it has no way to tell an occurrence
255
+ that is a moment late from one that is a day late.
168
256
 
169
- require './config/environment'
257
+ ### Lateness and deadlines
170
258
 
171
- Workhorse::Daemon::ShellHandler.run do |daemon|
172
- # Start scheduler process
173
- daemon.worker 'Scheduler' do
174
- scheduler = Rufus::Scheduler.new
259
+ A materialised job's `perform_at` is the occurrence's own time, not the moment
260
+ it was enqueued. The difference between it and `started_at` is therefore the
261
+ lateness of that occurrence, available on every job as
262
+ `Workhorse::DbJob#lateness`, and derivable in SQL from the `started_at` and
263
+ `perform_at` columns.
175
264
 
176
- scheduler.cron '0/10 * * * *' do
177
- Workhorse.enqueue Workhorse::Jobs::CleanupSucceededJobs.new
178
- end
265
+ Two options act on it, both settable on a schedule:
179
266
 
180
- Signal.trap 'TERM' do
181
- scheduler.shutdown
182
- end
267
+ * **`max_lateness`** — seconds the job may start late before
268
+ `Workhorse.on_job_late` is called. The job still runs; it was just late.
269
+ * **`expires_after`** — seconds after the occurrence at which the job is no
270
+ longer worth running. It is then set to state `expired` and
271
+ `Workhorse.on_job_expired` is called instead of it being performed. For
272
+ "send the 08:00 reminder", running it at 11:40 is often worse than not
273
+ running it at all.
183
274
 
184
- scheduler.join
185
- end
275
+ Hand-enqueued jobs take `max_lateness:` under the same name, and the deadline
276
+ as an absolute time rather than an offset:
186
277
 
187
- # Start 5 worker processes with 3 threads each
188
- 5.times do
189
- daemon.worker do
190
- Workhorse::Worker.start_and_wait(pool_size: 3, polling_interval: 10, logger: Rails.logger)
191
- end
192
- end
193
- end
194
- ```
278
+ ```ruby
279
+ Workhorse.enqueue(job, expires_at: 1.hour.from_now, max_lateness: 60)
280
+ ```
281
+
282
+ ```ruby
283
+ Workhorse.setup do |config|
284
+ config.on_job_expired = proc do |db_job|
285
+ ExceptionNotifier.notify_exception(
286
+ StandardError.new("Job #{db_job.id} (#{db_job.description}) expired")
287
+ )
288
+ end
289
+
290
+ config.on_job_late = proc do |db_job, lateness|
291
+ ExceptionNotifier.notify_exception(
292
+ StandardError.new("Job #{db_job.id} started #{lateness.round}s late")
293
+ )
294
+ end
295
+ end
296
+ ```
297
+
298
+ Both callbacks are best-effort: anything they raise goes to
299
+ `Workhorse.on_exception` and never affects the worker or the job. An expiry is
300
+ logged at `warn` whether or not a callback is configured.
301
+
302
+ Expired jobs stay in the table like any other finished job.
303
+ `Workhorse::Jobs::CleanupSucceededJobs` removes them along with succeeded ones;
304
+ pass `states: [Workhorse::DbJob::STATE_SUCCEEDED]` to keep them.
305
+
306
+ ### Detecting schedules that stopped
307
+
308
+ Neither callback can fire for a job that was never created, so if no worker is
309
+ polling or the global lock is stuck, occurrences simply stop being
310
+ materialised and nothing says so. `Workhorse::Jobs::DetectLateSchedulesJob`
311
+ covers that case by reporting schedules whose next occurrence lies well in the
312
+ past:
195
313
 
196
- This allows starting and stopping the daemon with the usual interface.
197
- Note that the scheduler is handled like a Workhorse worker, the consequence
198
- of which is that only one 'worker' should be started by the ShellHandler.
199
- Otherwise there would be multiple jobs scheduled at the same time.
314
+ ```ruby
315
+ Workhorse.schedules do
316
+ schedule 'detect_late_schedules',
317
+ job: 'Workhorse::Jobs::DetectLateSchedulesJob',
318
+ cron: '*/30 * * * *'
319
+ end
320
+ ```
321
+
322
+ It is itself performed by a worker, so it reports a stall only while at least
323
+ one worker is still running — use it alongside external monitoring rather than
324
+ instead of it.
325
+
326
+ ### Timezones
327
+
328
+ Without `timezone`, a cron expression is read in the process's local time.
329
+ Given one, occurrences are computed in that zone, including across daylight
330
+ saving changes. A `30 2 * * *` schedule in `Europe/Zurich` has no occurrence
331
+ on the day the clocks go forward, because 02:30 does not exist that day — and
332
+ exactly one on the day they go back, although 02:30 happens twice.
333
+
334
+ ### Disabling a schedule
335
+
336
+ Setting `enabled` to `false` on the row stops its occurrences from being
337
+ materialised, without a deployment:
338
+
339
+ ```ruby
340
+ Workhorse::Schedule.find_by(key: 'morning_digest').update!(enabled: false)
341
+ ```
200
342
 
201
- Please refer to the documentation for
202
- [rufus-scheduler](https://github.com/jmettraux/rufus-scheduler) (or the
203
- scheduler of your choice) for further options concerning the timing of the
204
- jobs.
343
+ Removing a schedule from the declarations deletes its row: the next worker to
344
+ start after a day has passed without any worker declaring it removes it. The delay matters during a rolling deployment, where
345
+ the old and the new version run at once: were rows removed immediately, each
346
+ version would delete the other's schedules and reset the occurrences they were
347
+ waiting for.
205
348
 
206
349
  ## Configuring and starting workers
207
350
 
@@ -278,6 +421,100 @@ polling interval.
278
421
  This setting is recommended for all setups and may eventually be enabled by
279
422
  default.
280
423
 
424
+ ### Notifications
425
+
426
+ Instant repolling only helps *after* a job has been performed. A worker that is
427
+ idle and waiting for new work still sleeps out its entire polling interval, so
428
+ a job enqueued just after a poll waits almost a full interval before it starts.
429
+
430
+ Shortening the polling interval is the obvious remedy and a poor one: every
431
+ poll acquires a global database lock, so more frequent polling across several
432
+ workers increases contention, and a worker that fails to acquire the lock skips
433
+ its poll and waits another whole interval. Polling costs the same whether or
434
+ not anything is happening.
435
+
436
+ *Notifications* turn the question around. Enqueuing a job announces it, and a
437
+ waiting worker polls straight away instead of sleeping out its interval:
438
+
439
+ ```ruby
440
+ # config/initializers/workhorse.rb
441
+ Workhorse.setup do |config|
442
+ config.notifier = :file
443
+ end
444
+ ```
445
+
446
+ Polling remains the floor. A notification that is never delivered — a worker
447
+ that was restarting, an enqueue from a host that cannot reach the others —
448
+ costs latency and nothing else, as the regular poll still finds the job. For
449
+ the same reason, raise `polling_interval` only as far as you are willing to
450
+ wait when a notification *is* missed.
451
+
452
+ Note what this does and does not do to database load. A notification only
453
+ brings the next poll forward; it never replaces one, so the saving comes from
454
+ raising `polling_interval`, not from notifications by themselves. An
455
+ announcement also wakes *every* worker with a free thread, and all but the one
456
+ that wins the job take the global lock for nothing. So a workload that enqueues
457
+ in bursts while many workers sit idle can take the lock more often than plain
458
+ polling would, rather than less. A worker whose threads are all busy does not
459
+ react, so the effect is bounded by how much spare capacity there is.
460
+
461
+ Two notifiers ship with workhorse:
462
+
463
+ #### `:file`
464
+
465
+ Touches a single file, which waiting workers stat once per 0.1 seconds on a
466
+ tick the poller performs anyway. It costs no database work at all and around a
467
+ microsecond of CPU per check, and a job starts within roughly 100 milliseconds.
468
+
469
+ It requires that the processes enqueueing jobs and the workers share a
470
+ filesystem, which in practice means the same host. Where they do not, the
471
+ touch never reaches those workers and they fall back to polling.
472
+
473
+ The file defaults to `tmp/pids/workhorse.wake` below the Rails root and can be
474
+ moved with `config.notification_path`. Every process involved must agree on it.
475
+
476
+ #### `:redis`
477
+
478
+ Publishes on a Redis pub/sub channel, for deployments whose workers do not
479
+ share a filesystem with the application. Redis is a soft dependency: it is not
480
+ declared as a dependency of this gem and workhorse never requires it — supply
481
+ a client through `config.notification_redis`.
482
+
483
+ ```ruby
484
+ Workhorse.setup do |config|
485
+ config.notifier = :redis
486
+ config.notification_redis = -> { Redis.new(url: ENV['REDIS_URL']) }
487
+ config.notification_channel = 'workhorse:jobs' # optional
488
+ end
489
+ ```
490
+
491
+ `notification_redis` is given something callable above because a subscribed
492
+ Redis connection cannot be used for anything else: `subscribe` occupies it
493
+ until it returns. The notifier calls it once for publishing and once for the
494
+ subscriber, so the two get separate connections.
495
+
496
+ Passing a client rather than a callable also works, but the subscriber then
497
+ falls back to `dup` — in redis-rb a shallow copy that may share the
498
+ connection, in which case publishing can block behind the subscription.
499
+
500
+ The channel can be changed at any point before a worker starts; the
501
+ subscriber thread captures the one it was started with.
502
+
503
+ #### Writing your own
504
+
505
+ Subclass `Workhorse::Notifiers::Base` and assign an instance to
506
+ `config.notifier`. A notifier announces jobs with `notify` and exposes a
507
+ `token` that changes whenever a notification has arrived; workers compare it
508
+ against the last value they saw, so several workers in one process stay
509
+ independent of one another. `notify(queue: nil)` must never raise — a job must
510
+ still be enqueued when it cannot be announced. `start` and `stop` are optional
511
+ hooks, called as a worker starts and shuts down, for anything that needs a
512
+ thread or a connection of its own.
513
+
514
+ Note that jobs with a future `perform_at` are not announced, as a woken worker
515
+ would find nothing to do; they are picked up by the regular poll once they are
516
+ due.
517
+
281
518
  ## Transactions
282
519
 
283
520
  By default, each job is run in an individual database transaction. An exception
@@ -412,6 +649,7 @@ DbJob.locked
412
649
  DbJob.started
413
650
  DbJob.succeeded
414
651
  DbJob.failed
652
+ DbJob.expired
415
653
  ```
416
654
  ### Resetting jobs
417
655
 
@@ -419,8 +657,8 @@ Jobs in a state other than `waiting` are either being processed or else already
419
657
  in a final state such as `succeeded` and won't be performed again. Workhorse
420
658
  provides an API method for resetting jobs in the following cases:
421
659
 
422
- * A job has succeeded or failed (states `succeeded` and `failed`) and needs to
423
- re-run. In these cases, perform a non-forced reset:
660
+ * A job has succeeded, failed or expired (states `succeeded`, `failed` and
661
+ `expired`) and needs to re-run. In these cases, perform a non-forced reset:
424
662
 
425
663
  ```ruby
426
664
  db_job.reset!
@@ -461,8 +699,9 @@ configuration or else using `self.queue_adapter` in a job class inheriting from
461
699
  Per default, jobs remain in the database, no matter in which state. This can
462
700
  eventually lead to a very large jobs database. You are advised to clean your
463
701
  jobs database on a regular interval. Workhorse provides the job
464
- `Workhorse::Jobs::CleanupSucceededJobs` for this purpose that cleans up all
465
- succeeded jobs. You can run this using your scheduler in a specific interval.
702
+ `Workhorse::Jobs::CleanupSucceededJobs` for this purpose, which cleans up
703
+ succeeded and expired jobs — pass `states:` to narrow that. Schedule it with
704
+ `Workhorse.schedules`, see [Scheduling](#scheduling).
466
705
 
467
706
  ## Memory handling
468
707
 
@@ -587,6 +826,11 @@ In the event that this still happens, Workhorse takes the following steps:
587
826
  - Retries acquiring the lock on the next poll.
588
827
  - Calls the `on_exception` callback (if configured) after a configurable number of consecutive failures to obtain the lock.
589
828
 
829
+ Failures on a poll that a notification or an instant repoll brought forward
830
+ are expected — several workers race for the lock and all but one lose — so
831
+ they are logged at `debug` only and do not count towards
832
+ `max_global_lock_fails`.
833
+
590
834
  The maximum number of consecutive failures can be configured using
591
835
  `config.max_global_lock_fails`, which defaults to 10.
592
836
 
data/Rakefile CHANGED
@@ -16,6 +16,7 @@ task :gemspec do
16
16
  spec.add_dependency 'activesupport', '>= 7.0.0'
17
17
  spec.add_dependency 'activerecord', '>= 7.0.0'
18
18
  spec.add_dependency 'concurrent-ruby'
19
+ spec.add_dependency 'fugit'
19
20
  end
20
21
 
21
22
  File.write('workhorse.gemspec', gemspec.to_ruby.strip)
data/VERSION CHANGED
@@ -1 +1 @@
1
- 1.5.2
1
+ 2.0.0.rc1
data/bin/rubocop CHANGED
@@ -1 +1,5 @@
1
- bundle exec rubocop "$@"
1
+ #!/bin/bash
2
+
3
+ BASE_PATH=`(cd $(dirname $0)/.. && pwd -P)`
4
+
5
+ BUNDLE_GEMFILE=$BASE_PATH/Gemfile $BASE_PATH/bin/ruby -e "require 'rubygems';require 'bundler/setup'; load Gem.activate_bin_path('rubocop', 'rubocop', '>= 0.a')" -- $@
@@ -7,12 +7,21 @@ module Workhorse
7
7
 
8
8
  source_root File.expand_path('templates', __dir__)
9
9
 
10
+ # Returns the version for the next generated migration. Counts up rather
11
+ # than returning the current time, as several migrations are generated
12
+ # within the same second and would otherwise collide.
10
13
  def self.next_migration_number(_dir)
11
- Time.now.utc.strftime('%Y%m%d%H%M%S')
14
+ @next_migration_number = [
15
+ Time.now.utc.strftime('%Y%m%d%H%M%S').to_i,
16
+ (@next_migration_number || 0) + 1
17
+ ].max
18
+
19
+ return @next_migration_number.to_s
12
20
  end
13
21
 
14
22
  def install_migration
15
23
  migration_template 'create_table_jobs.rb', 'db/migrate/create_table_jobs.rb'
24
+ migration_template 'create_table_workhorse_schedules.rb', 'db/migrate/create_table_workhorse_schedules.rb'
16
25
  end
17
26
 
18
27
  def install_daemon_script
@@ -23,4 +23,59 @@ Workhorse.setup do |config|
23
23
  # # Do something with exception, i.e.
24
24
  # # ExceptionNotifier.notify_exception(exception)
25
25
  # end
26
+
27
+ # Seconds the daemon's `stop` waits for a worker to finish what it is doing
28
+ # before killing it. Set to nil to wait indefinitely.
29
+ #
30
+ # config.shutdown_timeout = 300
31
+
32
+ # Enable this to let an enqueued job start without waiting for the next
33
+ # poll. Use :file where the workers share a filesystem with the application
34
+ # and :redis where they do not. Polling stays the floor either way, so raise
35
+ # the polling interval only as far as you are willing to wait when a
36
+ # notification is missed.
37
+ #
38
+ # config.notifier = :file
39
+ # config.notification_path = Rails.root.join('tmp', 'pids', 'workhorse.wake')
40
+ #
41
+ # config.notifier = :redis
42
+ # config.notification_redis = -> { Redis.new(url: ENV['REDIS_URL']) }
43
+ # config.notification_channel = 'workhorse:jobs'
44
+
45
+ # Enable and configure these to be told about a job that passed its
46
+ # `expires_at` before any worker got to it, and about one that started later
47
+ # than its `max_lateness` allows. Neither can affect the worker or the job.
48
+ #
49
+ # config.on_job_expired = proc do |db_job|
50
+ # # Do something with the job, i.e.
51
+ # # ExceptionNotifier.notify_exception(
52
+ # # StandardError.new("Job #{db_job.id} (#{db_job.description}) expired")
53
+ # # )
54
+ # end
55
+ #
56
+ # config.on_job_late = proc do |db_job, lateness|
57
+ # # Do something with the job, i.e.
58
+ # # ExceptionNotifier.notify_exception(
59
+ # # StandardError.new("Job #{db_job.id} started #{lateness.round}s late")
60
+ # # )
61
+ # end
26
62
  end
63
+
64
+ # Jobs that run on a schedule. Each of these owns a row in the
65
+ # `workhorse_schedules` table holding the next occurrence that has not been
66
+ # materialized yet, so an occurrence whose time passes while nothing is
67
+ # running is not lost. See the README for the catch-up policies.
68
+ #
69
+ # Workhorse.schedules do
70
+ # schedule 'cleanup_jobs',
71
+ # job: 'Workhorse::Jobs::CleanupSucceededJobs',
72
+ # cron: '10 0 * * *'
73
+ #
74
+ # schedule 'detect_stale_jobs',
75
+ # job: 'Workhorse::Jobs::DetectStaleJobsJob',
76
+ # cron: '30 * * * *'
77
+ #
78
+ # schedule 'detect_late_schedules',
79
+ # job: 'Workhorse::Jobs::DetectLateSchedulesJob',
80
+ # cron: '*/30 * * * *'
81
+ # end
@@ -17,17 +17,30 @@ class CreateTableJobs < ActiveRecord::Migration[7.1]
17
17
  t.integer :priority, null: false
18
18
  t.datetime :perform_at, null: true
19
19
 
20
+ # Deadline; the job is then set to state 'expired' rather than performed.
21
+ t.datetime :expires_at, null: true
22
+
23
+ # Seconds the job may start later than its perform_at before
24
+ # Workhorse.on_job_late is called.
25
+ t.integer :max_lateness, null: true
26
+
20
27
  t.string :description, null: true
21
28
 
22
29
  t.timestamps null: false
23
30
  end
24
31
 
32
+ # The index names are given explicitly because the ones Rails would derive
33
+ # exceed the 30 characters Oracle allows before 12.2.
25
34
  if oracle?
26
35
  add_index :jobs, :queue
27
- add_index :jobs, :state
36
+ add_index :jobs, %i[state perform_at], name: 'idx_jobs_state_perform_at'
37
+ add_index :jobs, %i[state priority created_at], name: 'idx_jobs_state_prio_created'
38
+ add_index :jobs, %i[state expires_at], name: 'idx_jobs_state_expires_at'
28
39
  else
29
40
  add_index :jobs, :queue, length: 191
30
- add_index :jobs, :state, length: 191
41
+ add_index :jobs, %i[state perform_at], length: { state: 191 }, name: 'idx_jobs_state_perform_at'
42
+ add_index :jobs, %i[state priority created_at], length: { state: 191 }, name: 'idx_jobs_state_prio_created'
43
+ add_index :jobs, %i[state expires_at], length: { state: 191 }, name: 'idx_jobs_state_expires_at'
31
44
  end
32
45
  add_index :jobs, :perform_at
33
46
  end
@@ -0,0 +1,42 @@
1
+ class CreateTableWorkhorseSchedules < ActiveRecord::Migration[7.1]
2
+ def change
3
+ # Schedules are addressed by key; the job class and its options live in
4
+ # the Workhorse.schedules definition rather than here.
5
+ create_table :workhorse_schedules, force: true do |t|
6
+ t.string :key, null: false
7
+
8
+ # The cron expression and timezone are kept here so that a change to
9
+ # either is recognised on reconciliation.
10
+ t.string :cron, null: false
11
+ t.string :timezone, null: true
12
+
13
+ # Lets a schedule be switched off without a deployment.
14
+ t.boolean :enabled, null: false, default: true
15
+
16
+ # The next occurrence that has not been materialised yet.
17
+ t.datetime :next_at, null: false
18
+
19
+ t.datetime :last_enqueued_at, null: true
20
+ t.datetime :last_occurrence, null: true
21
+ t.integer :last_job_id, null: true
22
+
23
+ t.timestamps null: false
24
+ end
25
+
26
+ # The index names are given explicitly because the ones Rails would derive
27
+ # exceed the 30 characters Oracle allows before 12.2.
28
+ if oracle?
29
+ add_index :workhorse_schedules, :key, unique: true, name: 'idx_wh_schedules_key'
30
+ else
31
+ add_index :workhorse_schedules, :key, unique: true, length: 191, name: 'idx_wh_schedules_key'
32
+ end
33
+
34
+ add_index :workhorse_schedules, %i[enabled next_at], name: 'idx_wh_schedules_due'
35
+ end
36
+
37
+ private
38
+
39
+ def oracle?
40
+ ActiveRecord::Base.connection.adapter_name == 'OracleEnhanced'
41
+ end
42
+ end
@@ -146,7 +146,10 @@ module Workhorse
146
146
 
147
147
  def self.acquire_lock(lockfile_path, flags)
148
148
  if Workhorse.lock_shell_commands
149
- lockfile = File.open(lockfile_path, 'a') # rubocop:disable Style/FileOpen
149
+ # Not the block form: the lockfile is returned to the caller, which
150
+ # holds the flock for as long as the command runs.
151
+ # rubocop:disable-next Style/FileOpen
152
+ lockfile = File.open(lockfile_path, 'a')
150
153
  result = lockfile.flock(flags)
151
154
 
152
155
  if result == false