workhorse 1.5.2 → 2.0.0.rc0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +102 -0
- data/README.md +303 -74
- data/Rakefile +1 -0
- data/VERSION +1 -1
- data/bin/rubocop +5 -1
- data/lib/generators/workhorse/install_generator.rb +10 -1
- data/lib/generators/workhorse/templates/config/initializers/workhorse.rb +55 -0
- data/lib/generators/workhorse/templates/create_table_jobs.rb +11 -13
- data/lib/generators/workhorse/templates/create_table_workhorse_schedules.rb +29 -0
- data/lib/workhorse/daemon.rb +52 -6
- data/lib/workhorse/db_job.rb +58 -7
- data/lib/workhorse/enqueuer.rb +51 -8
- data/lib/workhorse/jobs/cleanup_succeeded_jobs.rb +26 -8
- data/lib/workhorse/jobs/detect_late_schedules_job.rb +59 -0
- data/lib/workhorse/notifiers/base.rb +55 -0
- data/lib/workhorse/notifiers/file_system.rb +66 -0
- data/lib/workhorse/notifiers/none.rb +8 -0
- data/lib/workhorse/notifiers/redis.rb +227 -0
- data/lib/workhorse/performer.rb +28 -0
- data/lib/workhorse/poller.rb +289 -47
- data/lib/workhorse/pool.rb +12 -6
- data/lib/workhorse/schedule.rb +288 -0
- data/lib/workhorse/schedules.rb +197 -0
- data/lib/workhorse/worker.rb +77 -27
- data/lib/workhorse.rb +136 -0
- data/test/lib/db_schema.rb +21 -1
- data/test/lib/jobs.rb +29 -0
- data/test/lib/test_helper.rb +9 -14
- data/test/workhorse/daemon_test.rb +33 -0
- data/test/workhorse/db_job_test.rb +1 -1
- data/test/workhorse/notifier_test.rb +500 -0
- data/test/workhorse/poller_test.rb +8 -4
- data/test/workhorse/schedule_test.rb +967 -0
- data/test/workhorse/worker_test.rb +92 -0
- data/workhorse.gemspec +6 -5
- metadata +29 -3
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 69859406636c3463da70d9e274f023e5b0838720e428f77b8f0babad2e15baf9
|
|
4
|
+
data.tar.gz: 6af9c55d5f9066be6801898d3786311d3a0f5c33723baf3a5e4d464c107ace85
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: b674430cc63891334e49f4ce52b1cc407dea1fdd5575beab8792b82c4f12729596bfb16829ebd859bfefc2881a701e75ff85d8bca8e1e479a081b50e1bc7d0e0
|
|
7
|
+
data.tar.gz: 3db6d62ef9a9dbd950dcb40ac6399a6d28ab166bc75a8b592153de315672423f76b2a6518e58d219de789edf1cc2faa81368f568c637585e5d2a05247b626e14
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,107 @@
|
|
|
1
1
|
# Workhorse Changelog
|
|
2
2
|
|
|
3
|
+
## 2.0.0.rc0 - 2026-09-29
|
|
4
|
+
|
|
5
|
+
Sitrox reference: #154443.
|
|
6
|
+
|
|
7
|
+
### Breaking changes
|
|
8
|
+
|
|
9
|
+
* **Support for Oracle is dropped.** Workhorse supports MySQL and MariaDB
|
|
10
|
+
only. The Oracle branches of the global lock, the row limiting and the
|
|
11
|
+
generated migrations are gone, along with
|
|
12
|
+
`Workhorse::Poller::ORACLE_LOCK_MODE` and `ORACLE_LOCK_HANDLE`. It was never
|
|
13
|
+
covered by CI, so it only ever had manual verification. An Oracle
|
|
14
|
+
installation has no upgrade path and should stay on 1.x. See
|
|
15
|
+
[Database support](README.md#database-support).
|
|
16
|
+
|
|
17
|
+
* A forked daemon worker no longer runs the `at_exit` handlers registered by
|
|
18
|
+
the process that started it. They belong to that process, and one that waits
|
|
19
|
+
on threads the fork did not inherit hangs a worker which has already
|
|
20
|
+
finished - which then ignores `TERM`. This skips interpreter finalisation as
|
|
21
|
+
a whole, so buffered output is dropped too; a worker that dies of an
|
|
22
|
+
unhandled exception now reports it through `Workhorse.on_exception` and
|
|
23
|
+
exits non-zero.
|
|
24
|
+
|
|
25
|
+
### Added
|
|
26
|
+
|
|
27
|
+
* *Scheduling*: workhorse runs jobs on a cron schedule itself, without an
|
|
28
|
+
external scheduler process. Each schedule owns a row in the new
|
|
29
|
+
`workhorse_schedules` table holding the occurrence it is waiting for, so an
|
|
30
|
+
occurrence whose time passes while nothing is running is not lost, and what
|
|
31
|
+
happens to it is a per-schedule `catch_up` policy. Timezones and daylight
|
|
32
|
+
saving are handled. See [Scheduling](README.md#scheduling).
|
|
33
|
+
|
|
34
|
+
* *Notifications*: enqueuing a job announces it and a waiting worker polls
|
|
35
|
+
straight away, rather than sleeping out its polling interval. Polling stays
|
|
36
|
+
the floor. `:file` suits workers sharing a filesystem with the application
|
|
37
|
+
and `:redis` those that do not; both are off by default. See
|
|
38
|
+
[Notifications](README.md#notifications).
|
|
39
|
+
|
|
40
|
+
* Job deadlines and lateness reporting: `expires_at`, `max_lateness`, the new
|
|
41
|
+
`expired` state, `Workhorse.on_job_expired`, `Workhorse.on_job_late` and
|
|
42
|
+
`Workhorse::DbJob#lateness`. A job past its deadline is expired rather than
|
|
43
|
+
performed. See [Lateness and deadlines](README.md#lateness-and-deadlines).
|
|
44
|
+
|
|
45
|
+
* `Workhorse::Jobs::DetectLateSchedulesJob`, which reports schedules whose
|
|
46
|
+
next occurrence lies well in the past. Neither callback above can fire for a
|
|
47
|
+
job that was never created, so this is what catches materialization having
|
|
48
|
+
stopped. See
|
|
49
|
+
[Detecting schedules that stopped](README.md#detecting-schedules-that-stopped).
|
|
50
|
+
|
|
51
|
+
* `Workhorse.shutdown_timeout`, the seconds the daemon's `stop` waits for a
|
|
52
|
+
worker before killing it. Defaults to 300, `nil` restores the previous
|
|
53
|
+
behaviour. A worker that ignores `TERM` used to leave `stop` - and whatever
|
|
54
|
+
waits on it, usually a deployment - looping forever.
|
|
55
|
+
|
|
56
|
+
* `Workhorse.enqueue_job_class`, the keyword arguments `expires_at:` and
|
|
57
|
+
`max_lateness:` on `Workhorse.enqueue` and `Workhorse.enqueue_active_job`,
|
|
58
|
+
`priority:` on the latter, and the scope `Workhorse::DbJob.expired`.
|
|
59
|
+
|
|
60
|
+
* The composite indexes `[state, perform_at]`, `[state, priority, created_at]`
|
|
61
|
+
and `[state, expires_at]` in the generated `jobs` migration, replacing the
|
|
62
|
+
single-column index on `state`. `rails generate workhorse:install` now emits
|
|
63
|
+
two migrations rather than one.
|
|
64
|
+
|
|
65
|
+
* `fugit` as a runtime dependency, for parsing cron expressions.
|
|
66
|
+
|
|
67
|
+
### Changed
|
|
68
|
+
|
|
69
|
+
* `Workhorse::Jobs::CleanupSucceededJobs` also deletes jobs in the new
|
|
70
|
+
`expired` state, with a `states` argument to opt out. A schedule using
|
|
71
|
+
`expires_after` that regularly misses its window would otherwise grow the
|
|
72
|
+
jobs table without bound.
|
|
73
|
+
|
|
74
|
+
* Failures to obtain the global lock on a poll that a notification or an
|
|
75
|
+
instant repoll brought forward no longer count towards
|
|
76
|
+
`max_global_lock_fails`, and are logged at `debug`. Several workers woken by
|
|
77
|
+
one announcement race for the lock and all but one lose, which says nothing
|
|
78
|
+
about a crashed worker.
|
|
79
|
+
|
|
80
|
+
* `Workhorse::DbJob#reset!` accepts `expired` as the terminal state it is.
|
|
81
|
+
|
|
82
|
+
### Fixed
|
|
83
|
+
|
|
84
|
+
* A deadlock between shutting a worker down and the poller posting a job.
|
|
85
|
+
`Worker#shutdown` held the worker's mutex while waiting for the poller
|
|
86
|
+
thread, which could be waiting for that same mutex in `Worker#perform`. The
|
|
87
|
+
worker then ignored `TERM`. A job that was locked but cannot be performed
|
|
88
|
+
because the worker is shutting down is now reset to `waiting` rather than
|
|
89
|
+
left locked, where it would have blocked its queue.
|
|
90
|
+
|
|
91
|
+
* `Worker#shutdown` raising when called concurrently, which the daemon does by
|
|
92
|
+
sending both `TERM` and `INT`.
|
|
93
|
+
|
|
94
|
+
### Documentation
|
|
95
|
+
|
|
96
|
+
* PostgreSQL is documented as unsupported. Workers emit `GET_LOCK` on every
|
|
97
|
+
poll, which PostgreSQL does not provide, so a worker fails on its first one.
|
|
98
|
+
The requirements previously listed it as supported, which it never was.
|
|
99
|
+
|
|
100
|
+
### Upgrading
|
|
101
|
+
|
|
102
|
+
Nothing breaks without migrating, but the new features need one. See
|
|
103
|
+
[Upgrading from 1.x](README.md#upgrading-from-1x).
|
|
104
|
+
|
|
3
105
|
## 1.5.2 - 2026-08-04
|
|
4
106
|
|
|
5
107
|
* Fix `Poller#valid_queues` raising `NoMethodError` on the Oracle adapter. The
|
data/README.md
CHANGED
|
@@ -35,8 +35,8 @@ What it does not do:
|
|
|
35
35
|
|
|
36
36
|
* Ruby `>= 3.0` (may work with earlier versions but is untested)
|
|
37
37
|
* Rails `>= 7.0`
|
|
38
|
-
*
|
|
39
|
-
|
|
38
|
+
* MySQL or MariaDB with InnoDB. No other database is supported, see
|
|
39
|
+
[Database support](#database-support).
|
|
40
40
|
* If you are planning on using the daemons handler:
|
|
41
41
|
* An operating system and file system that supports file locking.
|
|
42
42
|
* MRI Ruby (aka "CRuby") as jRuby does not support `fork`. See the
|
|
@@ -60,21 +60,80 @@ What it does not do:
|
|
|
60
60
|
|
|
61
61
|
This generates:
|
|
62
62
|
|
|
63
|
-
*
|
|
63
|
+
* Two database migrations, creating the tables `jobs` and
|
|
64
|
+
`workhorse_schedules`
|
|
64
65
|
* The initializer `config/initializers/workhorse.rb` for global configuration
|
|
65
66
|
* This can be skipped using the `--skip-initializer` flag
|
|
66
67
|
* The daemon worker script `bin/workhorse.rb`
|
|
67
68
|
|
|
68
69
|
Please customize the initializer and worker script to your liking.
|
|
69
70
|
|
|
70
|
-
###
|
|
71
|
+
### Database support
|
|
71
72
|
|
|
72
|
-
|
|
73
|
-
`
|
|
73
|
+
**MySQL and MariaDB are the only supported databases.** Workhorse serialises
|
|
74
|
+
job pickup with `GET_LOCK`, a MySQL advisory lock, which is emitted on every
|
|
75
|
+
poll; a database that does not provide it fails on the first poll. InnoDB is
|
|
76
|
+
required, as MyISAM supports neither transactions nor row-level locking. Both
|
|
77
|
+
the `mysql2` and the `trilogy` adapter are covered by CI.
|
|
74
78
|
|
|
79
|
+
Oracle was supported until 2.0.0 and is not any more — see the changelog entry
|
|
80
|
+
for that release. PostgreSQL has never been supported, despite the
|
|
81
|
+
requirements once listing it; supporting it would mean an advisory-lock
|
|
82
|
+
dialect of its own (`pg_advisory_lock`) and is not currently planned.
|
|
83
|
+
|
|
84
|
+
## Upgrading from 1.x
|
|
85
|
+
|
|
86
|
+
Workhorse 2.0 drops support for Oracle, see
|
|
87
|
+
[Database support](#database-support). An Oracle installation has no upgrade
|
|
88
|
+
path and should stay on 1.x.
|
|
89
|
+
|
|
90
|
+
Otherwise nothing breaks without migrating: [scheduling](#scheduling) is inert
|
|
91
|
+
until a schedule is declared, [notifications](#notifications) are off until a
|
|
92
|
+
notifier is selected, and the new job columns are only written when used. To
|
|
93
|
+
take the new features up, add this migration:
|
|
94
|
+
|
|
95
|
+
```ruby
|
|
96
|
+
class UpgradeWorkhorseToV2 < ActiveRecord::Migration[7.1]
|
|
97
|
+
def change
|
|
98
|
+
add_column :jobs, :expires_at, :datetime, null: true
|
|
99
|
+
add_column :jobs, :max_lateness, :integer, null: true
|
|
100
|
+
|
|
101
|
+
create_table :workhorse_schedules do |t|
|
|
102
|
+
t.string :key, null: false
|
|
103
|
+
t.string :cron, null: false
|
|
104
|
+
t.string :timezone, null: true
|
|
105
|
+
t.boolean :enabled, null: false, default: true
|
|
106
|
+
t.datetime :next_at, null: false
|
|
107
|
+
t.datetime :last_enqueued_at, null: true
|
|
108
|
+
t.datetime :last_occurrence, null: true
|
|
109
|
+
t.integer :last_job_id, null: true
|
|
110
|
+
t.timestamps null: false
|
|
111
|
+
end
|
|
112
|
+
|
|
113
|
+
# The index names are given explicitly because the ones Rails would
|
|
114
|
+
# derive are longer than some tools accept.
|
|
115
|
+
add_index :workhorse_schedules, :key,
|
|
116
|
+
unique: true, length: 191, name: 'idx_wh_schedules_key'
|
|
117
|
+
add_index :workhorse_schedules, %i[enabled next_at],
|
|
118
|
+
name: 'idx_wh_schedules_due'
|
|
119
|
+
|
|
120
|
+
add_index :jobs, %i[state perform_at],
|
|
121
|
+
length: { state: 191 }, name: 'idx_jobs_state_perform_at'
|
|
122
|
+
add_index :jobs, %i[state priority created_at],
|
|
123
|
+
length: { state: 191 }, name: 'idx_jobs_state_prio_created'
|
|
124
|
+
add_index :jobs, %i[state expires_at],
|
|
125
|
+
length: { state: 191 }, name: 'idx_jobs_state_expires_at'
|
|
126
|
+
|
|
127
|
+
# Now redundant, as `state` leads both indexes above
|
|
128
|
+
remove_index :jobs, :state
|
|
129
|
+
end
|
|
130
|
+
end
|
|
75
131
|
```
|
|
76
|
-
|
|
77
|
-
|
|
132
|
+
|
|
133
|
+
Two behaviour changes worth knowing about, both described in the changelog: a
|
|
134
|
+
forked daemon worker no longer runs `at_exit` handlers registered by the
|
|
135
|
+
process that started it, and `Workhorse::Jobs::CleanupSucceededJobs` now also
|
|
136
|
+
deletes jobs in the new `expired` state.
|
|
78
137
|
|
|
79
138
|
## Queuing jobs
|
|
80
139
|
|
|
@@ -122,86 +181,155 @@ If you do not want to pass any parameters to the operation, just omit the third
|
|
|
122
181
|
Workhorse.enqueue_op Operations::Jobs::CleanUpDatabase, queue: :maintenance, priority: 2
|
|
123
182
|
```
|
|
124
183
|
|
|
125
|
-
|
|
184
|
+
## Scheduling
|
|
126
185
|
|
|
127
|
-
Workhorse
|
|
128
|
-
|
|
129
|
-
|
|
186
|
+
Workhorse runs jobs on a schedule itself, without an external scheduler
|
|
187
|
+
process. Schedules are declared in code and their state is kept in the
|
|
188
|
+
database:
|
|
130
189
|
|
|
131
|
-
|
|
132
|
-
|
|
190
|
+
```ruby
|
|
191
|
+
# config/initializers/workhorse.rb
|
|
192
|
+
Workhorse.schedules do
|
|
193
|
+
schedule 'cleanup_jobs',
|
|
194
|
+
job: 'Workhorse::Jobs::CleanupSucceededJobs',
|
|
195
|
+
cron: '10 0 * * *'
|
|
196
|
+
|
|
197
|
+
schedule 'morning_digest',
|
|
198
|
+
job: 'Jobs::MorningDigest',
|
|
199
|
+
cron: '0 8 * * 1-5',
|
|
200
|
+
timezone: 'Europe/Zurich',
|
|
201
|
+
queue: :reports,
|
|
202
|
+
priority: -10,
|
|
203
|
+
catch_up: :skip,
|
|
204
|
+
grace: 15.minutes,
|
|
205
|
+
max_lateness: 60.seconds
|
|
206
|
+
end
|
|
207
|
+
```
|
|
133
208
|
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
Proceed down this path at your own peril!)
|
|
209
|
+
Each schedule owns a row in `workhorse_schedules` holding the next occurrence
|
|
210
|
+
that has not been materialised yet. Workers reconcile those rows against the
|
|
211
|
+
declarations above on startup, and materialise the occurrences that have come
|
|
212
|
+
due during their regular poll. There is no scheduler process to keep alive and
|
|
213
|
+
no single point of failure: any worker will do.
|
|
140
214
|
|
|
141
|
-
|
|
142
|
-
10 minutes. If started at 12:00 sharp, after one hour it will execute at
|
|
143
|
-
13:00:30 at the earliest due to cumulative execution time.
|
|
215
|
+
### Why the occurrence is a row
|
|
144
216
|
|
|
145
|
-
|
|
217
|
+
An in-memory scheduler computes the next occurrence from *now*, so an
|
|
218
|
+
occurrence whose time passes while it is not running never happens and leaves
|
|
219
|
+
no trace — a deployment, a restart or a crash at the wrong minute silently
|
|
220
|
+
skips a nightly job. Because the next occurrence is persisted here, a worker
|
|
221
|
+
coming back at any later point still sees that it is due, and the schedule
|
|
222
|
+
decides what to do about it.
|
|
146
223
|
|
|
147
|
-
|
|
148
|
-
class MyJob
|
|
149
|
-
def perform
|
|
150
|
-
# Do all the work
|
|
224
|
+
### Catch-up
|
|
151
225
|
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
end
|
|
155
|
-
end
|
|
156
|
-
```
|
|
226
|
+
What should happen to an occurrence whose time has passed depends on the job,
|
|
227
|
+
so it is stated per schedule:
|
|
157
228
|
|
|
158
|
-
|
|
229
|
+
| `catch_up` | Behaviour | Suits |
|
|
230
|
+
|-------------|----------------------------------------------------------|-----------------------------------------|
|
|
231
|
+
| `:run_once` | Collapse all missed occurrences into one (**default**) | Cleanup, maintenance, idempotent work |
|
|
232
|
+
| `:run` | Materialise each, up to `max_catch_up` (default 10) | Per-period reports that must all exist |
|
|
233
|
+
| `:skip` | Drop those older than `grace` | "Send the 08:00 digest" |
|
|
159
234
|
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
example of an adapted `bin/workhorse.rb` to accommodate for the additional
|
|
164
|
-
cog in the mechanism is given below:
|
|
235
|
+
`:run_once` is the default deliberately: after a long outage it is the safe
|
|
236
|
+
behaviour. A schedule running every minute that was down for a day would
|
|
237
|
+
otherwise enqueue 1440 jobs at once, which `max_catch_up` also guards against.
|
|
165
238
|
|
|
166
|
-
|
|
167
|
-
|
|
239
|
+
`:skip` requires `grace`, as without one it has no way to tell an occurrence
|
|
240
|
+
that is a moment late from one that is a day late.
|
|
168
241
|
|
|
169
|
-
|
|
242
|
+
### Lateness and deadlines
|
|
170
243
|
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
244
|
+
A materialised job's `perform_at` is the occurrence's own time, not the moment
|
|
245
|
+
it was enqueued. The difference between it and `started_at` is therefore the
|
|
246
|
+
lateness of that occurrence, available on every job as
|
|
247
|
+
`Workhorse::DbJob#lateness`, and derivable in SQL from the `started_at` and
|
|
248
|
+
`perform_at` columns.
|
|
175
249
|
|
|
176
|
-
|
|
177
|
-
Workhorse.enqueue Workhorse::Jobs::CleanupSucceededJobs.new
|
|
178
|
-
end
|
|
250
|
+
Two options act on it, both settable on a schedule:
|
|
179
251
|
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
252
|
+
* **`max_lateness`** — seconds the job may start late before
|
|
253
|
+
`Workhorse.on_job_late` is called. The job still runs; it was just late.
|
|
254
|
+
* **`expires_after`** — seconds after the occurrence at which the job is no
|
|
255
|
+
longer worth running. It is then set to state `expired` and
|
|
256
|
+
`Workhorse.on_job_expired` is called instead of it being performed. For
|
|
257
|
+
"send the 08:00 reminder", running it at 11:40 is often worse than not
|
|
258
|
+
running it at all.
|
|
183
259
|
|
|
184
|
-
|
|
185
|
-
|
|
260
|
+
Hand-enqueued jobs take `max_lateness:` under the same name, and the deadline
|
|
261
|
+
as an absolute time rather than an offset:
|
|
186
262
|
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
263
|
+
```ruby
|
|
264
|
+
Workhorse.enqueue(job, expires_at: 1.hour.from_now, max_lateness: 60)
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
```ruby
|
|
268
|
+
Workhorse.setup do |config|
|
|
269
|
+
config.on_job_expired = proc do |db_job|
|
|
270
|
+
ExceptionNotifier.notify_exception(
|
|
271
|
+
StandardError.new("Job #{db_job.id} (#{db_job.description}) expired")
|
|
272
|
+
)
|
|
273
|
+
end
|
|
195
274
|
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
275
|
+
config.on_job_late = proc do |db_job, lateness|
|
|
276
|
+
ExceptionNotifier.notify_exception(
|
|
277
|
+
StandardError.new("Job #{db_job.id} started #{lateness.round}s late")
|
|
278
|
+
)
|
|
279
|
+
end
|
|
280
|
+
end
|
|
281
|
+
```
|
|
282
|
+
|
|
283
|
+
Both callbacks are best-effort: anything they raise goes to
|
|
284
|
+
`Workhorse.on_exception` and never affects the worker or the job. An expiry is
|
|
285
|
+
logged at `warn` whether or not a callback is configured.
|
|
286
|
+
|
|
287
|
+
Expired jobs stay in the table like any other finished job.
|
|
288
|
+
`Workhorse::Jobs::CleanupSucceededJobs` removes them along with succeeded ones;
|
|
289
|
+
pass `states: [Workhorse::DbJob::STATE_SUCCEEDED]` to keep them.
|
|
200
290
|
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
291
|
+
### Detecting schedules that stopped
|
|
292
|
+
|
|
293
|
+
Neither callback can fire for a job that was never created, so if no worker is
|
|
294
|
+
polling or the global lock is stuck, occurrences simply stop being
|
|
295
|
+
materialised and nothing says so. `Workhorse::Jobs::DetectLateSchedulesJob`
|
|
296
|
+
covers that case by reporting schedules whose next occurrence lies well in the
|
|
297
|
+
past:
|
|
298
|
+
|
|
299
|
+
```ruby
|
|
300
|
+
Workhorse.schedules do
|
|
301
|
+
schedule 'detect_late_schedules',
|
|
302
|
+
job: 'Workhorse::Jobs::DetectLateSchedulesJob',
|
|
303
|
+
cron: '*/30 * * * *'
|
|
304
|
+
end
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
It is itself performed by a worker, so it reports a stall only while at least
|
|
308
|
+
one worker is still running — use it alongside external monitoring rather than
|
|
309
|
+
instead of it.
|
|
310
|
+
|
|
311
|
+
### Timezones
|
|
312
|
+
|
|
313
|
+
Without `timezone`, a cron expression is read in the process's local time.
|
|
314
|
+
Given one, occurrences are computed in that zone, including across daylight
|
|
315
|
+
saving changes. A `30 2 * * *` schedule in `Europe/Zurich` has no occurrence
|
|
316
|
+
on the day the clocks go forward, because 02:30 does not exist that day — and
|
|
317
|
+
exactly one on the day they go back, although 02:30 happens twice.
|
|
318
|
+
|
|
319
|
+
### Disabling a schedule
|
|
320
|
+
|
|
321
|
+
Setting `enabled` to `false` on the row stops its occurrences from being
|
|
322
|
+
materialised, without a deployment:
|
|
323
|
+
|
|
324
|
+
```ruby
|
|
325
|
+
Workhorse::Schedule.find_by(key: 'morning_digest').update!(enabled: false)
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
Removing a schedule from the declarations deletes its row: the next worker to
|
|
329
|
+
start after a day has passed without any worker declaring it removes it. The delay matters during a rolling deployment, where
|
|
330
|
+
the old and the new version run at once: were rows removed immediately, each
|
|
331
|
+
version would delete the other's schedules and reset the occurrences they were
|
|
332
|
+
waiting for.
|
|
205
333
|
|
|
206
334
|
## Configuring and starting workers
|
|
207
335
|
|
|
@@ -278,6 +406,100 @@ polling interval.
|
|
|
278
406
|
This setting is recommended for all setups and may eventually be enabled by
|
|
279
407
|
default.
|
|
280
408
|
|
|
409
|
+
### Notifications
|
|
410
|
+
|
|
411
|
+
Instant repolling only helps *after* a job has been performed. A worker that is
|
|
412
|
+
idle and waiting for new work still sleeps out its entire polling interval, so
|
|
413
|
+
a job enqueued just after a poll waits almost a full interval before it starts.
|
|
414
|
+
|
|
415
|
+
Shortening the polling interval is the obvious remedy and a poor one: every
|
|
416
|
+
poll acquires a global database lock, so more frequent polling across several
|
|
417
|
+
workers increases contention, and a worker that fails to acquire the lock skips
|
|
418
|
+
its poll and waits another whole interval. Polling costs the same whether or
|
|
419
|
+
not anything is happening.
|
|
420
|
+
|
|
421
|
+
*Notifications* turn the question around. Enqueuing a job announces it, and a
|
|
422
|
+
waiting worker polls straight away instead of sleeping out its interval:
|
|
423
|
+
|
|
424
|
+
```ruby
|
|
425
|
+
# config/initializers/workhorse.rb
|
|
426
|
+
Workhorse.setup do |config|
|
|
427
|
+
config.notifier = :file
|
|
428
|
+
end
|
|
429
|
+
```
|
|
430
|
+
|
|
431
|
+
Polling remains the floor. A notification that is never delivered — a worker
|
|
432
|
+
that was restarting, an enqueue from a host that cannot reach the others —
|
|
433
|
+
costs latency and nothing else, as the regular poll still finds the job. For
|
|
434
|
+
the same reason, raise `polling_interval` only as far as you are willing to
|
|
435
|
+
wait when a notification *is* missed.
|
|
436
|
+
|
|
437
|
+
Note what this does and does not do to database load. A notification only
|
|
438
|
+
brings the next poll forward; it never replaces one, so the saving comes from
|
|
439
|
+
raising `polling_interval`, not from notifications by themselves. An
|
|
440
|
+
announcement also wakes *every* worker with a free thread, and all but the one
|
|
441
|
+
that wins the job take the global lock for nothing. So a workload that enqueues
|
|
442
|
+
in bursts while many workers sit idle can take the lock more often than plain
|
|
443
|
+
polling would, rather than less. A worker whose threads are all busy does not
|
|
444
|
+
react, so the effect is bounded by how much spare capacity there is.
|
|
445
|
+
|
|
446
|
+
Two notifiers ship with workhorse:
|
|
447
|
+
|
|
448
|
+
#### `:file`
|
|
449
|
+
|
|
450
|
+
Touches a single file, which waiting workers stat once per 0.1 seconds on a
|
|
451
|
+
tick the poller performs anyway. It costs no database work at all and around a
|
|
452
|
+
microsecond of CPU per check, and a job starts within roughly 100 milliseconds.
|
|
453
|
+
|
|
454
|
+
It requires that the processes enqueueing jobs and the workers share a
|
|
455
|
+
filesystem, which in practice means the same host. Where they do not, the
|
|
456
|
+
touch never reaches those workers and they fall back to polling.
|
|
457
|
+
|
|
458
|
+
The file defaults to `tmp/pids/workhorse.wake` below the Rails root and can be
|
|
459
|
+
moved with `config.notification_path`. Every process involved must agree on it.
|
|
460
|
+
|
|
461
|
+
#### `:redis`
|
|
462
|
+
|
|
463
|
+
Publishes on a Redis pub/sub channel, for deployments whose workers do not
|
|
464
|
+
share a filesystem with the application. Redis is a soft dependency: it is not
|
|
465
|
+
declared as a dependency of this gem and workhorse never requires it — supply
|
|
466
|
+
a client through `config.notification_redis`.
|
|
467
|
+
|
|
468
|
+
```ruby
|
|
469
|
+
Workhorse.setup do |config|
|
|
470
|
+
config.notifier = :redis
|
|
471
|
+
config.notification_redis = -> { Redis.new(url: ENV['REDIS_URL']) }
|
|
472
|
+
config.notification_channel = 'workhorse:jobs' # optional
|
|
473
|
+
end
|
|
474
|
+
```
|
|
475
|
+
|
|
476
|
+
`notification_redis` is given something callable above because a subscribed
|
|
477
|
+
Redis connection cannot be used for anything else: `subscribe` occupies it
|
|
478
|
+
until it returns. The notifier calls it once for publishing and once for the
|
|
479
|
+
subscriber, so the two get separate connections.
|
|
480
|
+
|
|
481
|
+
Passing a client rather than a callable also works, but the subscriber then
|
|
482
|
+
falls back to `dup` — in redis-rb a shallow copy that may share the
|
|
483
|
+
connection, in which case publishing can block behind the subscription.
|
|
484
|
+
|
|
485
|
+
The channel can be changed at any point before a worker starts; the
|
|
486
|
+
subscriber thread captures the one it was started with.
|
|
487
|
+
|
|
488
|
+
#### Writing your own
|
|
489
|
+
|
|
490
|
+
Subclass `Workhorse::Notifiers::Base` and assign an instance to
|
|
491
|
+
`config.notifier`. A notifier announces jobs with `notify` and exposes a
|
|
492
|
+
`token` that changes whenever a notification has arrived; workers compare it
|
|
493
|
+
against the last value they saw, so several workers in one process stay
|
|
494
|
+
independent of one another. `notify(queue: nil)` must never raise — a job must
|
|
495
|
+
still be enqueued when it cannot be announced. `start` and `stop` are optional
|
|
496
|
+
hooks, called as a worker starts and shuts down, for anything that needs a
|
|
497
|
+
thread or a connection of its own.
|
|
498
|
+
|
|
499
|
+
Note that jobs with a future `perform_at` are not announced, as a woken worker
|
|
500
|
+
would find nothing to do; they are picked up by the regular poll once they are
|
|
501
|
+
due.
|
|
502
|
+
|
|
281
503
|
## Transactions
|
|
282
504
|
|
|
283
505
|
By default, each job is run in an individual database transaction. An exception
|
|
@@ -412,6 +634,7 @@ DbJob.locked
|
|
|
412
634
|
DbJob.started
|
|
413
635
|
DbJob.succeeded
|
|
414
636
|
DbJob.failed
|
|
637
|
+
DbJob.expired
|
|
415
638
|
```
|
|
416
639
|
### Resetting jobs
|
|
417
640
|
|
|
@@ -419,8 +642,8 @@ Jobs in a state other than `waiting` are either being processed or else already
|
|
|
419
642
|
in a final state such as `succeeded` and won't be performed again. Workhorse
|
|
420
643
|
provides an API method for resetting jobs in the following cases:
|
|
421
644
|
|
|
422
|
-
* A job has succeeded or
|
|
423
|
-
re-run. In these cases, perform a non-forced reset:
|
|
645
|
+
* A job has succeeded, failed or expired (states `succeeded`, `failed` and
|
|
646
|
+
`expired`) and needs to re-run. In these cases, perform a non-forced reset:
|
|
424
647
|
|
|
425
648
|
```ruby
|
|
426
649
|
db_job.reset!
|
|
@@ -461,8 +684,9 @@ configuration or else using `self.queue_adapter` in a job class inheriting from
|
|
|
461
684
|
Per default, jobs remain in the database, no matter in which state. This can
|
|
462
685
|
eventually lead to a very large jobs database. You are advised to clean your
|
|
463
686
|
jobs database on a regular interval. Workhorse provides the job
|
|
464
|
-
`Workhorse::Jobs::CleanupSucceededJobs` for this purpose
|
|
465
|
-
succeeded
|
|
687
|
+
`Workhorse::Jobs::CleanupSucceededJobs` for this purpose, which cleans up
|
|
688
|
+
succeeded and expired jobs — pass `states:` to narrow that. Schedule it with
|
|
689
|
+
`Workhorse.schedules`, see [Scheduling](#scheduling).
|
|
466
690
|
|
|
467
691
|
## Memory handling
|
|
468
692
|
|
|
@@ -587,6 +811,11 @@ In the event that this still happens, Workhorse takes the following steps:
|
|
|
587
811
|
- Retries acquiring the lock on the next poll.
|
|
588
812
|
- Calls the `on_exception` callback (if configured) after a configurable number of consecutive failures to obtain the lock.
|
|
589
813
|
|
|
814
|
+
Failures on a poll that a notification or an instant repoll brought forward
|
|
815
|
+
are expected — several workers race for the lock and all but one lose — so
|
|
816
|
+
they are logged at `debug` only and do not count towards
|
|
817
|
+
`max_global_lock_fails`.
|
|
818
|
+
|
|
590
819
|
The maximum number of consecutive failures can be configured using
|
|
591
820
|
`config.max_global_lock_fails`, which defaults to 10.
|
|
592
821
|
|
data/Rakefile
CHANGED
data/VERSION
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
|
|
1
|
+
2.0.0.rc0
|
data/bin/rubocop
CHANGED
|
@@ -7,12 +7,21 @@ module Workhorse
|
|
|
7
7
|
|
|
8
8
|
source_root File.expand_path('templates', __dir__)
|
|
9
9
|
|
|
10
|
+
# Returns the version for the next generated migration. Counts up rather
|
|
11
|
+
# than returning the current time, as several migrations are generated
|
|
12
|
+
# within the same second and would otherwise collide.
|
|
10
13
|
def self.next_migration_number(_dir)
|
|
11
|
-
|
|
14
|
+
@next_migration_number = [
|
|
15
|
+
Time.now.utc.strftime('%Y%m%d%H%M%S').to_i,
|
|
16
|
+
(@next_migration_number || 0) + 1
|
|
17
|
+
].max
|
|
18
|
+
|
|
19
|
+
return @next_migration_number.to_s
|
|
12
20
|
end
|
|
13
21
|
|
|
14
22
|
def install_migration
|
|
15
23
|
migration_template 'create_table_jobs.rb', 'db/migrate/create_table_jobs.rb'
|
|
24
|
+
migration_template 'create_table_workhorse_schedules.rb', 'db/migrate/create_table_workhorse_schedules.rb'
|
|
16
25
|
end
|
|
17
26
|
|
|
18
27
|
def install_daemon_script
|