workhorse 1.5.2 → 2.0.0.rc1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.github/workflows/ruby.yml +137 -1
- data/CHANGELOG.md +150 -0
- data/Gemfile +16 -1
- data/README.md +316 -72
- data/Rakefile +1 -0
- data/VERSION +1 -1
- data/bin/rubocop +5 -1
- data/lib/generators/workhorse/install_generator.rb +10 -1
- data/lib/generators/workhorse/templates/config/initializers/workhorse.rb +55 -0
- data/lib/generators/workhorse/templates/create_table_jobs.rb +15 -2
- data/lib/generators/workhorse/templates/create_table_workhorse_schedules.rb +42 -0
- data/lib/workhorse/daemon/shell_handler.rb +4 -1
- data/lib/workhorse/daemon.rb +57 -7
- data/lib/workhorse/db_job.rb +98 -8
- data/lib/workhorse/enqueuer.rb +51 -8
- data/lib/workhorse/jobs/cleanup_succeeded_jobs.rb +26 -8
- data/lib/workhorse/jobs/detect_late_schedules_job.rb +59 -0
- data/lib/workhorse/notifiers/base.rb +55 -0
- data/lib/workhorse/notifiers/file_system.rb +66 -0
- data/lib/workhorse/notifiers/none.rb +8 -0
- data/lib/workhorse/notifiers/redis.rb +227 -0
- data/lib/workhorse/performer.rb +29 -2
- data/lib/workhorse/poller.rb +303 -21
- data/lib/workhorse/pool.rb +12 -6
- data/lib/workhorse/schedule.rb +288 -0
- data/lib/workhorse/schedules.rb +197 -0
- data/lib/workhorse/worker.rb +102 -31
- data/lib/workhorse.rb +136 -0
- data/test/lib/db_schema.rb +36 -3
- data/test/lib/jobs.rb +29 -0
- data/test/lib/test_helper.rb +113 -20
- data/test/workhorse/daemon_test.rb +33 -0
- data/test/workhorse/db_job_test.rb +2 -4
- data/test/workhorse/notifier_test.rb +487 -0
- data/test/workhorse/performer_test.rb +7 -9
- data/test/workhorse/poller_test.rb +97 -23
- data/test/workhorse/schedule_test.rb +967 -0
- data/test/workhorse/worker_test.rb +201 -76
- data/workhorse.gemspec +6 -5
- metadata +29 -3
data/README.md
CHANGED
|
@@ -35,8 +35,8 @@ What it does not do:
|
|
|
35
35
|
|
|
36
36
|
* Ruby `>= 3.0` (may work with earlier versions but is untested)
|
|
37
37
|
* Rails `>= 7.0`
|
|
38
|
-
*
|
|
39
|
-
MySQL with InnoDB,
|
|
38
|
+
* One of the supported databases (see [Database support](#database-support)):
|
|
39
|
+
MySQL / MariaDB with InnoDB, or Oracle. **PostgreSQL is not supported.**
|
|
40
40
|
* If you are planning on using the daemons handler:
|
|
41
41
|
* An operating system and file system that supports file locking.
|
|
42
42
|
* MRI Ruby (aka "CRuby") as jRuby does not support `fork`. See the
|
|
@@ -60,22 +60,96 @@ What it does not do:
|
|
|
60
60
|
|
|
61
61
|
This generates:
|
|
62
62
|
|
|
63
|
-
*
|
|
63
|
+
* Two database migrations, creating the tables `jobs` and
|
|
64
|
+
`workhorse_schedules`
|
|
64
65
|
* The initializer `config/initializers/workhorse.rb` for global configuration
|
|
65
66
|
* This can be skipped using the `--skip-initializer` flag
|
|
66
67
|
* The daemon worker script `bin/workhorse.rb`
|
|
67
68
|
|
|
68
69
|
Please customize the initializer and worker script to your liking.
|
|
69
70
|
|
|
70
|
-
###
|
|
71
|
+
### Database support
|
|
71
72
|
|
|
72
|
-
|
|
73
|
-
|
|
73
|
+
Workhorse serialises job pickup using a database-level lock, which is
|
|
74
|
+
necessarily written against a specific database's dialect. Two families are
|
|
75
|
+
implemented:
|
|
76
|
+
|
|
77
|
+
| Database | Supported | Lock used | Covered by CI |
|
|
78
|
+
|-------------------|-----------|----------------------|---------------|
|
|
79
|
+
| MySQL / MariaDB | Yes | `GET_LOCK` | Yes, against both the `mysql2` and the `trilogy` adapter |
|
|
80
|
+
| Oracle 12c+ | Yes | `DBMS_LOCK` | Yes, against `activerecord-oracle_enhanced-adapter` |
|
|
81
|
+
| PostgreSQL | **No** | — | — |
|
|
82
|
+
| Everything else | **No** | — | — |
|
|
83
|
+
|
|
84
|
+
There is no PostgreSQL implementation: workers emit `GET_LOCK` on every poll,
|
|
85
|
+
which PostgreSQL does not provide, so a worker fails on its first poll.
|
|
86
|
+
Supporting it would mean an advisory-lock dialect of its own
|
|
87
|
+
(`pg_advisory_lock`) and is not currently planned. Note that InnoDB is required
|
|
88
|
+
on MySQL / MariaDB, as MyISAM supports neither transactions nor row-level
|
|
89
|
+
locking.
|
|
90
|
+
|
|
91
|
+
Oracle 12c is the minimum, as job selection limits its rows with
|
|
92
|
+
`FETCH FIRST … ROWS ONLY`. When using Oracle, make sure your schema has access
|
|
93
|
+
to the package `DBMS_LOCK`:
|
|
74
94
|
|
|
75
95
|
```
|
|
76
96
|
GRANT execute ON DBMS_LOCK TO <schema-name>;
|
|
77
97
|
```
|
|
78
98
|
|
|
99
|
+
## Upgrading from 1.x
|
|
100
|
+
|
|
101
|
+
Nothing breaks without migrating: [scheduling](#scheduling) is inert
|
|
102
|
+
until a schedule is declared, [notifications](#notifications) are off until a
|
|
103
|
+
notifier is selected, and the new job columns are only written when used. To
|
|
104
|
+
take the new features up, add this migration:
|
|
105
|
+
|
|
106
|
+
```ruby
|
|
107
|
+
class UpgradeWorkhorseToV2 < ActiveRecord::Migration[7.1]
|
|
108
|
+
def change
|
|
109
|
+
add_column :jobs, :expires_at, :datetime, null: true
|
|
110
|
+
add_column :jobs, :max_lateness, :integer, null: true
|
|
111
|
+
|
|
112
|
+
create_table :workhorse_schedules do |t|
|
|
113
|
+
t.string :key, null: false
|
|
114
|
+
t.string :cron, null: false
|
|
115
|
+
t.string :timezone, null: true
|
|
116
|
+
t.boolean :enabled, null: false, default: true
|
|
117
|
+
t.datetime :next_at, null: false
|
|
118
|
+
t.datetime :last_enqueued_at, null: true
|
|
119
|
+
t.datetime :last_occurrence, null: true
|
|
120
|
+
t.integer :last_job_id, null: true
|
|
121
|
+
t.timestamps null: false
|
|
122
|
+
end
|
|
123
|
+
|
|
124
|
+
# The index names are given explicitly because the ones Rails would
|
|
125
|
+
# derive are longer than some tools accept.
|
|
126
|
+
add_index :workhorse_schedules, :key,
|
|
127
|
+
unique: true, length: 191, name: 'idx_wh_schedules_key'
|
|
128
|
+
add_index :workhorse_schedules, %i[enabled next_at],
|
|
129
|
+
name: 'idx_wh_schedules_due'
|
|
130
|
+
|
|
131
|
+
add_index :jobs, %i[state perform_at],
|
|
132
|
+
length: { state: 191 }, name: 'idx_jobs_state_perform_at'
|
|
133
|
+
add_index :jobs, %i[state priority created_at],
|
|
134
|
+
length: { state: 191 }, name: 'idx_jobs_state_prio_created'
|
|
135
|
+
add_index :jobs, %i[state expires_at],
|
|
136
|
+
length: { state: 191 }, name: 'idx_jobs_state_expires_at'
|
|
137
|
+
|
|
138
|
+
# Now redundant, as `state` leads both indexes above
|
|
139
|
+
remove_index :jobs, :state
|
|
140
|
+
end
|
|
141
|
+
end
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
On Oracle, drop every `length:` option above — it indexes the whole column and
|
|
145
|
+
rejects a prefix length. The generated migrations do this for you; this one is
|
|
146
|
+
written out by hand.
|
|
147
|
+
|
|
148
|
+
Two behaviour changes worth knowing about, both described in the changelog: a
|
|
149
|
+
forked daemon worker no longer runs `at_exit` handlers registered by the
|
|
150
|
+
process that started it, and `Workhorse::Jobs::CleanupSucceededJobs` now also
|
|
151
|
+
deletes jobs in the new `expired` state.
|
|
152
|
+
|
|
79
153
|
## Queuing jobs
|
|
80
154
|
|
|
81
155
|
### Basic jobs
|
|
@@ -122,86 +196,155 @@ If you do not want to pass any parameters to the operation, just omit the third
|
|
|
122
196
|
Workhorse.enqueue_op Operations::Jobs::CleanUpDatabase, queue: :maintenance, priority: 2
|
|
123
197
|
```
|
|
124
198
|
|
|
125
|
-
|
|
199
|
+
## Scheduling
|
|
126
200
|
|
|
127
|
-
Workhorse
|
|
128
|
-
|
|
129
|
-
|
|
201
|
+
Workhorse runs jobs on a schedule itself, without an external scheduler
|
|
202
|
+
process. Schedules are declared in code and their state is kept in the
|
|
203
|
+
database:
|
|
130
204
|
|
|
131
|
-
|
|
132
|
-
|
|
205
|
+
```ruby
|
|
206
|
+
# config/initializers/workhorse.rb
|
|
207
|
+
Workhorse.schedules do
|
|
208
|
+
schedule 'cleanup_jobs',
|
|
209
|
+
job: 'Workhorse::Jobs::CleanupSucceededJobs',
|
|
210
|
+
cron: '10 0 * * *'
|
|
211
|
+
|
|
212
|
+
schedule 'morning_digest',
|
|
213
|
+
job: 'Jobs::MorningDigest',
|
|
214
|
+
cron: '0 8 * * 1-5',
|
|
215
|
+
timezone: 'Europe/Zurich',
|
|
216
|
+
queue: :reports,
|
|
217
|
+
priority: -10,
|
|
218
|
+
catch_up: :skip,
|
|
219
|
+
grace: 15.minutes,
|
|
220
|
+
max_lateness: 60.seconds
|
|
221
|
+
end
|
|
222
|
+
```
|
|
133
223
|
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
Proceed down this path at your own peril!)
|
|
224
|
+
Each schedule owns a row in `workhorse_schedules` holding the next occurrence
|
|
225
|
+
that has not been materialised yet. Workers reconcile those rows against the
|
|
226
|
+
declarations above on startup, and materialise the occurrences that have come
|
|
227
|
+
due during their regular poll. There is no scheduler process to keep alive and
|
|
228
|
+
no single point of failure: any worker will do.
|
|
140
229
|
|
|
141
|
-
|
|
142
|
-
10 minutes. If started at 12:00 sharp, after one hour it will execute at
|
|
143
|
-
13:00:30 at the earliest due to cumulative execution time.
|
|
230
|
+
### Why the occurrence is a row
|
|
144
231
|
|
|
145
|
-
|
|
232
|
+
An in-memory scheduler computes the next occurrence from *now*, so an
|
|
233
|
+
occurrence whose time passes while it is not running never happens and leaves
|
|
234
|
+
no trace — a deployment, a restart or a crash at the wrong minute silently
|
|
235
|
+
skips a nightly job. Because the next occurrence is persisted here, a worker
|
|
236
|
+
coming back at any later point still sees that it is due, and the schedule
|
|
237
|
+
decides what to do about it.
|
|
146
238
|
|
|
147
|
-
|
|
148
|
-
class MyJob
|
|
149
|
-
def perform
|
|
150
|
-
# Do all the work
|
|
239
|
+
### Catch-up
|
|
151
240
|
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
end
|
|
155
|
-
end
|
|
156
|
-
```
|
|
241
|
+
What should happen to an occurrence whose time has passed depends on the job,
|
|
242
|
+
so it is stated per schedule:
|
|
157
243
|
|
|
158
|
-
|
|
244
|
+
| `catch_up` | Behaviour | Suits |
|
|
245
|
+
|-------------|----------------------------------------------------------|-----------------------------------------|
|
|
246
|
+
| `:run_once` | Collapse all missed occurrences into one (**default**) | Cleanup, maintenance, idempotent work |
|
|
247
|
+
| `:run` | Materialise each, up to `max_catch_up` (default 10) | Per-period reports that must all exist |
|
|
248
|
+
| `:skip` | Drop those older than `grace` | "Send the 08:00 digest" |
|
|
159
249
|
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
example of an adapted `bin/workhorse.rb` to accommodate for the additional
|
|
164
|
-
cog in the mechanism is given below:
|
|
250
|
+
`:run_once` is the default deliberately: after a long outage it is the safe
|
|
251
|
+
behaviour. A schedule running every minute that was down for a day would
|
|
252
|
+
otherwise enqueue 1440 jobs at once, which `max_catch_up` also guards against.
|
|
165
253
|
|
|
166
|
-
|
|
167
|
-
|
|
254
|
+
`:skip` requires `grace`, as without one it has no way to tell an occurrence
|
|
255
|
+
that is a moment late from one that is a day late.
|
|
168
256
|
|
|
169
|
-
|
|
257
|
+
### Lateness and deadlines
|
|
170
258
|
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
259
|
+
A materialised job's `perform_at` is the occurrence's own time, not the moment
|
|
260
|
+
it was enqueued. The difference between it and `started_at` is therefore the
|
|
261
|
+
lateness of that occurrence, available on every job as
|
|
262
|
+
`Workhorse::DbJob#lateness`, and derivable in SQL from the `started_at` and
|
|
263
|
+
`perform_at` columns.
|
|
175
264
|
|
|
176
|
-
|
|
177
|
-
Workhorse.enqueue Workhorse::Jobs::CleanupSucceededJobs.new
|
|
178
|
-
end
|
|
265
|
+
Two options act on it, both settable on a schedule:
|
|
179
266
|
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
267
|
+
* **`max_lateness`** — seconds the job may start late before
|
|
268
|
+
`Workhorse.on_job_late` is called. The job still runs; it was just late.
|
|
269
|
+
* **`expires_after`** — seconds after the occurrence at which the job is no
|
|
270
|
+
longer worth running. It is then set to state `expired` and
|
|
271
|
+
`Workhorse.on_job_expired` is called instead of it being performed. For
|
|
272
|
+
"send the 08:00 reminder", running it at 11:40 is often worse than not
|
|
273
|
+
running it at all.
|
|
183
274
|
|
|
184
|
-
|
|
185
|
-
|
|
275
|
+
Hand-enqueued jobs take `max_lateness:` under the same name, and the deadline
|
|
276
|
+
as an absolute time rather than an offset:
|
|
186
277
|
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
278
|
+
```ruby
|
|
279
|
+
Workhorse.enqueue(job, expires_at: 1.hour.from_now, max_lateness: 60)
|
|
280
|
+
```
|
|
281
|
+
|
|
282
|
+
```ruby
|
|
283
|
+
Workhorse.setup do |config|
|
|
284
|
+
config.on_job_expired = proc do |db_job|
|
|
285
|
+
ExceptionNotifier.notify_exception(
|
|
286
|
+
StandardError.new("Job #{db_job.id} (#{db_job.description}) expired")
|
|
287
|
+
)
|
|
288
|
+
end
|
|
289
|
+
|
|
290
|
+
config.on_job_late = proc do |db_job, lateness|
|
|
291
|
+
ExceptionNotifier.notify_exception(
|
|
292
|
+
StandardError.new("Job #{db_job.id} started #{lateness.round}s late")
|
|
293
|
+
)
|
|
294
|
+
end
|
|
295
|
+
end
|
|
296
|
+
```
|
|
297
|
+
|
|
298
|
+
Both callbacks are best-effort: anything they raise goes to
|
|
299
|
+
`Workhorse.on_exception` and never affects the worker or the job. An expiry is
|
|
300
|
+
logged at `warn` whether or not a callback is configured.
|
|
301
|
+
|
|
302
|
+
Expired jobs stay in the table like any other finished job.
|
|
303
|
+
`Workhorse::Jobs::CleanupSucceededJobs` removes them along with succeeded ones;
|
|
304
|
+
pass `states: [Workhorse::DbJob::STATE_SUCCEEDED]` to keep them.
|
|
305
|
+
|
|
306
|
+
### Detecting schedules that stopped
|
|
307
|
+
|
|
308
|
+
Neither callback can fire for a job that was never created, so if no worker is
|
|
309
|
+
polling or the global lock is stuck, occurrences simply stop being
|
|
310
|
+
materialised and nothing says so. `Workhorse::Jobs::DetectLateSchedulesJob`
|
|
311
|
+
covers that case by reporting schedules whose next occurrence lies well in the
|
|
312
|
+
past:
|
|
195
313
|
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
314
|
+
```ruby
|
|
315
|
+
Workhorse.schedules do
|
|
316
|
+
schedule 'detect_late_schedules',
|
|
317
|
+
job: 'Workhorse::Jobs::DetectLateSchedulesJob',
|
|
318
|
+
cron: '*/30 * * * *'
|
|
319
|
+
end
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
It is itself performed by a worker, so it reports a stall only while at least
|
|
323
|
+
one worker is still running — use it alongside external monitoring rather than
|
|
324
|
+
instead of it.
|
|
325
|
+
|
|
326
|
+
### Timezones
|
|
327
|
+
|
|
328
|
+
Without `timezone`, a cron expression is read in the process's local time.
|
|
329
|
+
Given one, occurrences are computed in that zone, including across daylight
|
|
330
|
+
saving changes. A `30 2 * * *` schedule in `Europe/Zurich` has no occurrence
|
|
331
|
+
on the day the clocks go forward, because 02:30 does not exist that day — and
|
|
332
|
+
exactly one on the day they go back, although 02:30 happens twice.
|
|
333
|
+
|
|
334
|
+
### Disabling a schedule
|
|
335
|
+
|
|
336
|
+
Setting `enabled` to `false` on the row stops its occurrences from being
|
|
337
|
+
materialised, without a deployment:
|
|
338
|
+
|
|
339
|
+
```ruby
|
|
340
|
+
Workhorse::Schedule.find_by(key: 'morning_digest').update!(enabled: false)
|
|
341
|
+
```
|
|
200
342
|
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
343
|
+
Removing a schedule from the declarations deletes its row: the next worker to
|
|
344
|
+
start after a day has passed without any worker declaring it removes it. The delay matters during a rolling deployment, where
|
|
345
|
+
the old and the new version run at once: were rows removed immediately, each
|
|
346
|
+
version would delete the other's schedules and reset the occurrences they were
|
|
347
|
+
waiting for.
|
|
205
348
|
|
|
206
349
|
## Configuring and starting workers
|
|
207
350
|
|
|
@@ -278,6 +421,100 @@ polling interval.
|
|
|
278
421
|
This setting is recommended for all setups and may eventually be enabled by
|
|
279
422
|
default.
|
|
280
423
|
|
|
424
|
+
### Notifications
|
|
425
|
+
|
|
426
|
+
Instant repolling only helps *after* a job has been performed. A worker that is
|
|
427
|
+
idle and waiting for new work still sleeps out its entire polling interval, so
|
|
428
|
+
a job enqueued just after a poll waits almost a full interval before it starts.
|
|
429
|
+
|
|
430
|
+
Shortening the polling interval is the obvious remedy and a poor one: every
|
|
431
|
+
poll acquires a global database lock, so more frequent polling across several
|
|
432
|
+
workers increases contention, and a worker that fails to acquire the lock skips
|
|
433
|
+
its poll and waits another whole interval. Polling costs the same whether or
|
|
434
|
+
not anything is happening.
|
|
435
|
+
|
|
436
|
+
*Notifications* turn the question around. Enqueuing a job announces it, and a
|
|
437
|
+
waiting worker polls straight away instead of sleeping out its interval:
|
|
438
|
+
|
|
439
|
+
```ruby
|
|
440
|
+
# config/initializers/workhorse.rb
|
|
441
|
+
Workhorse.setup do |config|
|
|
442
|
+
config.notifier = :file
|
|
443
|
+
end
|
|
444
|
+
```
|
|
445
|
+
|
|
446
|
+
Polling remains the floor. A notification that is never delivered — a worker
|
|
447
|
+
that was restarting, an enqueue from a host that cannot reach the others —
|
|
448
|
+
costs latency and nothing else, as the regular poll still finds the job. For
|
|
449
|
+
the same reason, raise `polling_interval` only as far as you are willing to
|
|
450
|
+
wait when a notification *is* missed.
|
|
451
|
+
|
|
452
|
+
Note what this does and does not do to database load. A notification only
|
|
453
|
+
brings the next poll forward; it never replaces one, so the saving comes from
|
|
454
|
+
raising `polling_interval`, not from notifications by themselves. An
|
|
455
|
+
announcement also wakes *every* worker with a free thread, and all but the one
|
|
456
|
+
that wins the job take the global lock for nothing. So a workload that enqueues
|
|
457
|
+
in bursts while many workers sit idle can take the lock more often than plain
|
|
458
|
+
polling would, rather than less. A worker whose threads are all busy does not
|
|
459
|
+
react, so the effect is bounded by how much spare capacity there is.
|
|
460
|
+
|
|
461
|
+
Two notifiers ship with workhorse:
|
|
462
|
+
|
|
463
|
+
#### `:file`
|
|
464
|
+
|
|
465
|
+
Touches a single file, which waiting workers stat once per 0.1 seconds on a
|
|
466
|
+
tick the poller performs anyway. It costs no database work at all and around a
|
|
467
|
+
microsecond of CPU per check, and a job starts within roughly 100 milliseconds.
|
|
468
|
+
|
|
469
|
+
It requires that the processes enqueueing jobs and the workers share a
|
|
470
|
+
filesystem, which in practice means the same host. Where they do not, the
|
|
471
|
+
touch never reaches those workers and they fall back to polling.
|
|
472
|
+
|
|
473
|
+
The file defaults to `tmp/pids/workhorse.wake` below the Rails root and can be
|
|
474
|
+
moved with `config.notification_path`. Every process involved must agree on it.
|
|
475
|
+
|
|
476
|
+
#### `:redis`
|
|
477
|
+
|
|
478
|
+
Publishes on a Redis pub/sub channel, for deployments whose workers do not
|
|
479
|
+
share a filesystem with the application. Redis is a soft dependency: it is not
|
|
480
|
+
declared as a dependency of this gem and workhorse never requires it — supply
|
|
481
|
+
a client through `config.notification_redis`.
|
|
482
|
+
|
|
483
|
+
```ruby
|
|
484
|
+
Workhorse.setup do |config|
|
|
485
|
+
config.notifier = :redis
|
|
486
|
+
config.notification_redis = -> { Redis.new(url: ENV['REDIS_URL']) }
|
|
487
|
+
config.notification_channel = 'workhorse:jobs' # optional
|
|
488
|
+
end
|
|
489
|
+
```
|
|
490
|
+
|
|
491
|
+
`notification_redis` is given something callable above because a subscribed
|
|
492
|
+
Redis connection cannot be used for anything else: `subscribe` occupies it
|
|
493
|
+
until it returns. The notifier calls it once for publishing and once for the
|
|
494
|
+
subscriber, so the two get separate connections.
|
|
495
|
+
|
|
496
|
+
Passing a client rather than a callable also works, but the subscriber then
|
|
497
|
+
falls back to `dup` — in redis-rb a shallow copy that may share the
|
|
498
|
+
connection, in which case publishing can block behind the subscription.
|
|
499
|
+
|
|
500
|
+
The channel can be changed at any point before a worker starts; the
|
|
501
|
+
subscriber thread captures the one it was started with.
|
|
502
|
+
|
|
503
|
+
#### Writing your own
|
|
504
|
+
|
|
505
|
+
Subclass `Workhorse::Notifiers::Base` and assign an instance to
|
|
506
|
+
`config.notifier`. A notifier announces jobs with `notify` and exposes a
|
|
507
|
+
`token` that changes whenever a notification has arrived; workers compare it
|
|
508
|
+
against the last value they saw, so several workers in one process stay
|
|
509
|
+
independent of one another. `notify(queue: nil)` must never raise — a job must
|
|
510
|
+
still be enqueued when it cannot be announced. `start` and `stop` are optional
|
|
511
|
+
hooks, called as a worker starts and shuts down, for anything that needs a
|
|
512
|
+
thread or a connection of its own.
|
|
513
|
+
|
|
514
|
+
Note that jobs with a future `perform_at` are not announced, as a woken worker
|
|
515
|
+
would find nothing to do; they are picked up by the regular poll once they are
|
|
516
|
+
due.
|
|
517
|
+
|
|
281
518
|
## Transactions
|
|
282
519
|
|
|
283
520
|
By default, each job is run in an individual database transaction. An exception
|
|
@@ -412,6 +649,7 @@ DbJob.locked
|
|
|
412
649
|
DbJob.started
|
|
413
650
|
DbJob.succeeded
|
|
414
651
|
DbJob.failed
|
|
652
|
+
DbJob.expired
|
|
415
653
|
```
|
|
416
654
|
### Resetting jobs
|
|
417
655
|
|
|
@@ -419,8 +657,8 @@ Jobs in a state other than `waiting` are either being processed or else already
|
|
|
419
657
|
in a final state such as `succeeded` and won't be performed again. Workhorse
|
|
420
658
|
provides an API method for resetting jobs in the following cases:
|
|
421
659
|
|
|
422
|
-
* A job has succeeded or
|
|
423
|
-
re-run. In these cases, perform a non-forced reset:
|
|
660
|
+
* A job has succeeded, failed or expired (states `succeeded`, `failed` and
|
|
661
|
+
`expired`) and needs to re-run. In these cases, perform a non-forced reset:
|
|
424
662
|
|
|
425
663
|
```ruby
|
|
426
664
|
db_job.reset!
|
|
@@ -461,8 +699,9 @@ configuration or else using `self.queue_adapter` in a job class inheriting from
|
|
|
461
699
|
Per default, jobs remain in the database, no matter in which state. This can
|
|
462
700
|
eventually lead to a very large jobs database. You are advised to clean your
|
|
463
701
|
jobs database on a regular interval. Workhorse provides the job
|
|
464
|
-
`Workhorse::Jobs::CleanupSucceededJobs` for this purpose
|
|
465
|
-
succeeded
|
|
702
|
+
`Workhorse::Jobs::CleanupSucceededJobs` for this purpose, which cleans up
|
|
703
|
+
succeeded and expired jobs — pass `states:` to narrow that. Schedule it with
|
|
704
|
+
`Workhorse.schedules`, see [Scheduling](#scheduling).
|
|
466
705
|
|
|
467
706
|
## Memory handling
|
|
468
707
|
|
|
@@ -587,6 +826,11 @@ In the event that this still happens, Workhorse takes the following steps:
|
|
|
587
826
|
- Retries acquiring the lock on the next poll.
|
|
588
827
|
- Calls the `on_exception` callback (if configured) after a configurable number of consecutive failures to obtain the lock.
|
|
589
828
|
|
|
829
|
+
Failures on a poll that a notification or an instant repoll brought forward
|
|
830
|
+
are expected — several workers race for the lock and all but one lose — so
|
|
831
|
+
they are logged at `debug` only and do not count towards
|
|
832
|
+
`max_global_lock_fails`.
|
|
833
|
+
|
|
590
834
|
The maximum number of consecutive failures can be configured using
|
|
591
835
|
`config.max_global_lock_fails`, which defaults to 10.
|
|
592
836
|
|
data/Rakefile
CHANGED
data/VERSION
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
|
|
1
|
+
2.0.0.rc1
|
data/bin/rubocop
CHANGED
|
@@ -7,12 +7,21 @@ module Workhorse
|
|
|
7
7
|
|
|
8
8
|
source_root File.expand_path('templates', __dir__)
|
|
9
9
|
|
|
10
|
+
# Returns the version for the next generated migration. Counts up rather
|
|
11
|
+
# than returning the current time, as several migrations are generated
|
|
12
|
+
# within the same second and would otherwise collide.
|
|
10
13
|
def self.next_migration_number(_dir)
|
|
11
|
-
|
|
14
|
+
@next_migration_number = [
|
|
15
|
+
Time.now.utc.strftime('%Y%m%d%H%M%S').to_i,
|
|
16
|
+
(@next_migration_number || 0) + 1
|
|
17
|
+
].max
|
|
18
|
+
|
|
19
|
+
return @next_migration_number.to_s
|
|
12
20
|
end
|
|
13
21
|
|
|
14
22
|
def install_migration
|
|
15
23
|
migration_template 'create_table_jobs.rb', 'db/migrate/create_table_jobs.rb'
|
|
24
|
+
migration_template 'create_table_workhorse_schedules.rb', 'db/migrate/create_table_workhorse_schedules.rb'
|
|
16
25
|
end
|
|
17
26
|
|
|
18
27
|
def install_daemon_script
|
|
@@ -23,4 +23,59 @@ Workhorse.setup do |config|
|
|
|
23
23
|
# # Do something with exception, i.e.
|
|
24
24
|
# # ExceptionNotifier.notify_exception(exception)
|
|
25
25
|
# end
|
|
26
|
+
|
|
27
|
+
# Seconds the daemon's `stop` waits for a worker to finish what it is doing
|
|
28
|
+
# before killing it. Set to nil to wait indefinitely.
|
|
29
|
+
#
|
|
30
|
+
# config.shutdown_timeout = 300
|
|
31
|
+
|
|
32
|
+
# Enable this to let an enqueued job start without waiting for the next
|
|
33
|
+
# poll. Use :file where the workers share a filesystem with the application
|
|
34
|
+
# and :redis where they do not. Polling stays the floor either way, so raise
|
|
35
|
+
# the polling interval only as far as you are willing to wait when a
|
|
36
|
+
# notification is missed.
|
|
37
|
+
#
|
|
38
|
+
# config.notifier = :file
|
|
39
|
+
# config.notification_path = Rails.root.join('tmp', 'pids', 'workhorse.wake')
|
|
40
|
+
#
|
|
41
|
+
# config.notifier = :redis
|
|
42
|
+
# config.notification_redis = -> { Redis.new(url: ENV['REDIS_URL']) }
|
|
43
|
+
# config.notification_channel = 'workhorse:jobs'
|
|
44
|
+
|
|
45
|
+
# Enable and configure these to be told about a job that passed its
|
|
46
|
+
# `expires_at` before any worker got to it, and about one that started later
|
|
47
|
+
# than its `max_lateness` allows. Neither can affect the worker or the job.
|
|
48
|
+
#
|
|
49
|
+
# config.on_job_expired = proc do |db_job|
|
|
50
|
+
# # Do something with the job, i.e.
|
|
51
|
+
# # ExceptionNotifier.notify_exception(
|
|
52
|
+
# # StandardError.new("Job #{db_job.id} (#{db_job.description}) expired")
|
|
53
|
+
# # )
|
|
54
|
+
# end
|
|
55
|
+
#
|
|
56
|
+
# config.on_job_late = proc do |db_job, lateness|
|
|
57
|
+
# # Do something with the job, i.e.
|
|
58
|
+
# # ExceptionNotifier.notify_exception(
|
|
59
|
+
# # StandardError.new("Job #{db_job.id} started #{lateness.round}s late")
|
|
60
|
+
# # )
|
|
61
|
+
# end
|
|
26
62
|
end
|
|
63
|
+
|
|
64
|
+
# Jobs that run on a schedule. Each of these owns a row in the
|
|
65
|
+
# `workhorse_schedules` table holding the next occurrence that has not been
|
|
66
|
+
# materialized yet, so an occurrence whose time passes while nothing is
|
|
67
|
+
# running is not lost. See the README for the catch-up policies.
|
|
68
|
+
#
|
|
69
|
+
# Workhorse.schedules do
|
|
70
|
+
# schedule 'cleanup_jobs',
|
|
71
|
+
# job: 'Workhorse::Jobs::CleanupSucceededJobs',
|
|
72
|
+
# cron: '10 0 * * *'
|
|
73
|
+
#
|
|
74
|
+
# schedule 'detect_stale_jobs',
|
|
75
|
+
# job: 'Workhorse::Jobs::DetectStaleJobsJob',
|
|
76
|
+
# cron: '30 * * * *'
|
|
77
|
+
#
|
|
78
|
+
# schedule 'detect_late_schedules',
|
|
79
|
+
# job: 'Workhorse::Jobs::DetectLateSchedulesJob',
|
|
80
|
+
# cron: '*/30 * * * *'
|
|
81
|
+
# end
|
|
@@ -17,17 +17,30 @@ class CreateTableJobs < ActiveRecord::Migration[7.1]
|
|
|
17
17
|
t.integer :priority, null: false
|
|
18
18
|
t.datetime :perform_at, null: true
|
|
19
19
|
|
|
20
|
+
# Deadline; the job is then set to state 'expired' rather than performed.
|
|
21
|
+
t.datetime :expires_at, null: true
|
|
22
|
+
|
|
23
|
+
# Seconds the job may start later than its perform_at before
|
|
24
|
+
# Workhorse.on_job_late is called.
|
|
25
|
+
t.integer :max_lateness, null: true
|
|
26
|
+
|
|
20
27
|
t.string :description, null: true
|
|
21
28
|
|
|
22
29
|
t.timestamps null: false
|
|
23
30
|
end
|
|
24
31
|
|
|
32
|
+
# The index names are given explicitly because the ones Rails would derive
|
|
33
|
+
# exceed the 30 characters Oracle allows before 12.2.
|
|
25
34
|
if oracle?
|
|
26
35
|
add_index :jobs, :queue
|
|
27
|
-
add_index :jobs, :
|
|
36
|
+
add_index :jobs, %i[state perform_at], name: 'idx_jobs_state_perform_at'
|
|
37
|
+
add_index :jobs, %i[state priority created_at], name: 'idx_jobs_state_prio_created'
|
|
38
|
+
add_index :jobs, %i[state expires_at], name: 'idx_jobs_state_expires_at'
|
|
28
39
|
else
|
|
29
40
|
add_index :jobs, :queue, length: 191
|
|
30
|
-
add_index :jobs,
|
|
41
|
+
add_index :jobs, %i[state perform_at], length: { state: 191 }, name: 'idx_jobs_state_perform_at'
|
|
42
|
+
add_index :jobs, %i[state priority created_at], length: { state: 191 }, name: 'idx_jobs_state_prio_created'
|
|
43
|
+
add_index :jobs, %i[state expires_at], length: { state: 191 }, name: 'idx_jobs_state_expires_at'
|
|
31
44
|
end
|
|
32
45
|
add_index :jobs, :perform_at
|
|
33
46
|
end
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
class CreateTableWorkhorseSchedules < ActiveRecord::Migration[7.1]
|
|
2
|
+
def change
|
|
3
|
+
# Schedules are addressed by key; the job class and its options live in
|
|
4
|
+
# the Workhorse.schedules definition rather than here.
|
|
5
|
+
create_table :workhorse_schedules, force: true do |t|
|
|
6
|
+
t.string :key, null: false
|
|
7
|
+
|
|
8
|
+
# The cron expression and timezone are kept here so that a change to
|
|
9
|
+
# either is recognised on reconciliation.
|
|
10
|
+
t.string :cron, null: false
|
|
11
|
+
t.string :timezone, null: true
|
|
12
|
+
|
|
13
|
+
# Lets a schedule be switched off without a deployment.
|
|
14
|
+
t.boolean :enabled, null: false, default: true
|
|
15
|
+
|
|
16
|
+
# The next occurrence that has not been materialised yet.
|
|
17
|
+
t.datetime :next_at, null: false
|
|
18
|
+
|
|
19
|
+
t.datetime :last_enqueued_at, null: true
|
|
20
|
+
t.datetime :last_occurrence, null: true
|
|
21
|
+
t.integer :last_job_id, null: true
|
|
22
|
+
|
|
23
|
+
t.timestamps null: false
|
|
24
|
+
end
|
|
25
|
+
|
|
26
|
+
# The index names are given explicitly because the ones Rails would derive
|
|
27
|
+
# exceed the 30 characters Oracle allows before 12.2.
|
|
28
|
+
if oracle?
|
|
29
|
+
add_index :workhorse_schedules, :key, unique: true, name: 'idx_wh_schedules_key'
|
|
30
|
+
else
|
|
31
|
+
add_index :workhorse_schedules, :key, unique: true, length: 191, name: 'idx_wh_schedules_key'
|
|
32
|
+
end
|
|
33
|
+
|
|
34
|
+
add_index :workhorse_schedules, %i[enabled next_at], name: 'idx_wh_schedules_due'
|
|
35
|
+
end
|
|
36
|
+
|
|
37
|
+
private
|
|
38
|
+
|
|
39
|
+
def oracle?
|
|
40
|
+
ActiveRecord::Base.connection.adapter_name == 'OracleEnhanced'
|
|
41
|
+
end
|
|
42
|
+
end
|
|
@@ -146,7 +146,10 @@ module Workhorse
|
|
|
146
146
|
|
|
147
147
|
def self.acquire_lock(lockfile_path, flags)
|
|
148
148
|
if Workhorse.lock_shell_commands
|
|
149
|
-
lockfile
|
|
149
|
+
# Not the block form: the lockfile is returned to the caller, which
|
|
150
|
+
# holds the flock for as long as the command runs.
|
|
151
|
+
# rubocop:disable-next Style/FileOpen
|
|
152
|
+
lockfile = File.open(lockfile_path, 'a')
|
|
150
153
|
result = lockfile.flock(flags)
|
|
151
154
|
|
|
152
155
|
if result == false
|