solid-jobs 0.1.3 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +73 -4
- data/README.md +149 -118
- data/docs/reliability.md +73 -91
- data/exe/solid-jobs +2 -3
- data/lib/active_job/queue_adapters/solid_jobs_adapter.rb +7 -8
- data/lib/solid_jobs/{config.rb → blueprint.rb} +51 -51
- data/lib/solid_jobs/catalog.rb +427 -0
- data/lib/solid_jobs/claim.rb +315 -0
- data/lib/solid_jobs/{server.rb → conductor.rb} +10 -10
- data/lib/solid_jobs/{cli.rb → console.rb} +12 -9
- data/lib/solid_jobs/{processor.rb → engine.rb} +33 -33
- data/lib/solid_jobs/errors.rb +1 -1
- data/lib/solid_jobs/executor.rb +34 -0
- data/lib/solid_jobs/failure_policy.rb +98 -0
- data/lib/solid_jobs/heartbeat.rb +15 -13
- data/lib/solid_jobs/integrity_check.rb +39 -36
- data/lib/solid_jobs/interceptor_registry.rb +57 -0
- data/lib/solid_jobs/keyspace.rb +42 -0
- data/lib/solid_jobs/{testing.rb → lab.rb} +19 -19
- data/lib/solid_jobs/publisher.rb +186 -0
- data/lib/solid_jobs/rails_adapter.rb +13 -0
- data/lib/solid_jobs/{rails.rb → railtie.rb} +1 -2
- data/lib/solid_jobs/recovery.rb +8 -7
- data/lib/solid_jobs/task.rb +147 -0
- data/lib/solid_jobs/{scheduler.rb → timer.rb} +8 -8
- data/lib/solid_jobs/utilities.rb +5 -11
- data/lib/solid_jobs/version.rb +1 -1
- data/lib/solid_jobs.rb +28 -27
- metadata +19 -18
- data/lib/solid_jobs/active_job.rb +0 -16
- data/lib/solid_jobs/api.rb +0 -457
- data/lib/solid_jobs/client.rb +0 -202
- data/lib/solid_jobs/fetch.rb +0 -310
- data/lib/solid_jobs/job.rb +0 -156
- data/lib/solid_jobs/middleware/chain.rb +0 -104
- data/lib/solid_jobs/retry_service.rb +0 -100
- data/lib/solid_jobs/worker.rb +0 -35
data/docs/reliability.md
CHANGED
|
@@ -1,63 +1,61 @@
|
|
|
1
1
|
# Reliability model
|
|
2
2
|
|
|
3
|
-
SolidJobs uses a Redis-backed at-least-once state machine:
|
|
3
|
+
SolidJobs uses a namespaced Redis-backed at-least-once state machine:
|
|
4
4
|
|
|
5
5
|
```text
|
|
6
6
|
READY
|
|
7
7
|
|
|
|
8
|
-
|
|
|
8
|
+
| atomic claim
|
|
9
9
|
v
|
|
10
|
-
|
|
10
|
+
CLAIMED ---- fenced completion ----> removed
|
|
11
11
|
|
|
|
12
|
-
+---- application error ---->
|
|
12
|
+
+---- application error ----> RETRYING or DISCARDED, then completion
|
|
13
13
|
+---- graceful timeout -----> READY
|
|
14
|
-
+----
|
|
14
|
+
+---- node crash -----------> claim retained
|
|
15
15
|
|
|
|
16
16
|
+---- recovery ----> READY
|
|
17
17
|
```
|
|
18
18
|
|
|
19
|
-
|
|
20
|
-
|
|
19
|
+
Ready tasks live in `solid_jobs:channel:<name>`. Each executor owns at most one
|
|
20
|
+
claimed-task list and one claim journal entry:
|
|
21
21
|
|
|
22
22
|
```text
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
23
|
+
task_id stable across replay
|
|
24
|
+
claim_token unique per execution attempt
|
|
25
|
+
node_id owning server identity
|
|
26
|
+
executor_id owning Executor Ractor
|
|
27
|
+
channel destination used by recovery
|
|
28
|
+
attempt monotonically increasing execution count
|
|
29
|
+
claimed_at wall-clock claim time
|
|
29
30
|
```
|
|
30
31
|
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
or revived
|
|
34
|
-
supervisor reuses the same processor slot.
|
|
32
|
+
Completion and requeue are fenced by `claim_token`. Their Lua scripts verify
|
|
33
|
+
that the executor still owns the current journal generation before changing
|
|
34
|
+
state. A delayed or revived executor cannot complete or requeue a newer claim.
|
|
35
35
|
|
|
36
|
-
The
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
is not stolen while its owner remains alive.
|
|
36
|
+
The envelope retains its canonical `channel`, allowing another node to restore
|
|
37
|
+
it after its owner dies. Recovery uses node liveness and heartbeat, never claim
|
|
38
|
+
age alone, so a legitimate long-running task is not stolen.
|
|
40
39
|
|
|
41
40
|
## Failure boundaries
|
|
42
41
|
|
|
43
|
-
- Before
|
|
44
|
-
- After
|
|
45
|
-
- After the application effect but before
|
|
46
|
-
- Redis unavailable during
|
|
47
|
-
- Graceful shutdown: the active
|
|
42
|
+
- Before claim: the task remains in `solid_jobs:channel:<name>`.
|
|
43
|
+
- After claim or during `execute_task`: the task remains claimed.
|
|
44
|
+
- After the application effect but before completion: recovery replays it.
|
|
45
|
+
- Redis unavailable during completion: the task remains claimed and is replayed.
|
|
46
|
+
- Graceful shutdown: the active task may finish within the configured timeout;
|
|
48
47
|
otherwise it is interrupted and requeued.
|
|
49
|
-
- `SIGKILL`:
|
|
48
|
+
- `SIGKILL`: recovery relies exclusively on Redis state.
|
|
50
49
|
|
|
51
|
-
This design
|
|
52
|
-
|
|
50
|
+
This design favors no task loss over duplicate suppression. Exactly-once side
|
|
51
|
+
effects require application-level idempotency.
|
|
53
52
|
|
|
54
53
|
## Startup isolation
|
|
55
54
|
|
|
56
|
-
The server does not
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
`SolidJobs::StartupBarrier
|
|
60
|
-
ready:
|
|
55
|
+
The server does not claim work while components are booting. Heartbeat,
|
|
56
|
+
Engine, and Timer Ractors initialize their local configuration and Redis
|
|
57
|
+
pool, report `READY` exactly once, and wait behind
|
|
58
|
+
`SolidJobs::StartupBarrier`:
|
|
61
59
|
|
|
62
60
|
```text
|
|
63
61
|
BOOTING -> ALL_READY -> RUNNING
|
|
@@ -65,85 +63,69 @@ BOOTING -> ALL_READY -> RUNNING
|
|
|
65
63
|
```
|
|
66
64
|
|
|
67
65
|
Boot failure is terminal. Already-ready components receive `:abort`, close
|
|
68
|
-
their local resources, and never enter their
|
|
69
|
-
idempotent.
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
Ruby
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
`
|
|
80
|
-
`
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
Ruby 3.4 use the client/API side freely, and prefer one process per
|
|
90
|
-
Processor (`concurrency: 1`) for the server. Server-based stress tests are
|
|
91
|
-
skipped on Ruby < 4 for this reason, and `Server#start` logs a warning when
|
|
92
|
-
it detects `concurrency > 1` on Ruby < 4.
|
|
93
|
-
|
|
94
|
-
### Bounded shutdown
|
|
95
|
-
|
|
96
|
-
`Server#stop` never waits forever for a component. Each Ractor gets
|
|
97
|
-
`shutdown_timeout + Server::STOP_GRACE` to return after `:stop`; the signal
|
|
98
|
-
is re-sent up to `STOP_RESENDS` times, then the component is abandoned with
|
|
99
|
-
an error log so the process can proceed with shutdown.
|
|
66
|
+
their local resources, and never enter their claim loops. Cleanup is
|
|
67
|
+
idempotent.
|
|
68
|
+
|
|
69
|
+
## Ruby 3.4 Ractor caveat
|
|
70
|
+
|
|
71
|
+
Ruby 3.4's Ractor scheduler can deadlock the VM when a GC-triggered scheduler
|
|
72
|
+
barrier runs while several Ractors exchange moved messages.
|
|
73
|
+
`test/support/ractor_barrier_repro.rb` reproduces this without SolidJobs or
|
|
74
|
+
Redis. Ruby 4.0.1 completes the equivalent reproducer.
|
|
75
|
+
|
|
76
|
+
Run multi-Ractor SolidJobs servers on Ruby >= 4.0. On Ruby 3.4, prefer one
|
|
77
|
+
process per Engine (`concurrency: 1`). Conductor stress tests are skipped on
|
|
78
|
+
Ruby versions affected by this runtime issue, and `Conductor#start` warns when it
|
|
79
|
+
detects a multi-Ractor configuration there.
|
|
80
|
+
|
|
81
|
+
## Bounded shutdown
|
|
82
|
+
|
|
83
|
+
`Conductor#stop` never waits forever for a component. Each Ractor gets
|
|
84
|
+
`shutdown_timeout + Conductor::STOP_GRACE` to return after `:stop`; the signal is
|
|
85
|
+
re-sent up to `STOP_RESENDS` times before the component is abandoned with an
|
|
86
|
+
error log.
|
|
100
87
|
|
|
101
88
|
## Configuration scope
|
|
102
89
|
|
|
103
90
|
`SolidJobs.config` and the testing mode are Ractor-local, not thread-local.
|
|
104
|
-
Every thread and fiber inside a Ractor
|
|
105
|
-
|
|
106
|
-
Ractor keeps its own isolated configuration.
|
|
91
|
+
Every thread and fiber inside a Ractor shares its configuration, while each
|
|
92
|
+
Ractor owns isolated mutable state and Redis connections.
|
|
107
93
|
|
|
108
94
|
## Integrity auditing
|
|
109
95
|
|
|
110
|
-
Fault
|
|
111
|
-
Redis-backed state:
|
|
96
|
+
Fault tests can reconcile known task IDs against every Redis-backed state:
|
|
112
97
|
|
|
113
98
|
```ruby
|
|
114
99
|
report = SolidJobs::IntegrityCheck.call(
|
|
115
|
-
expected_job_ids:
|
|
100
|
+
expected_job_ids: submitted_task_ids,
|
|
116
101
|
acked_key: "test:completed",
|
|
117
102
|
)
|
|
118
103
|
|
|
119
104
|
raise report.inspect unless report.ok?
|
|
120
105
|
```
|
|
121
106
|
|
|
122
|
-
The report separates lost
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
sequence.
|
|
107
|
+
The report separates lost tasks, unexpected tasks, inconsistent claim
|
|
108
|
+
journals, dangling attempt indexes, duplicate active states, and malformed
|
|
109
|
+
envelopes. Completion removes the task's attempt index atomically; requeue
|
|
110
|
+
retains it so recovery increments the same attempt sequence.
|
|
127
111
|
|
|
128
|
-
## Backpressure and
|
|
112
|
+
## Backpressure and channels
|
|
129
113
|
|
|
130
|
-
Each
|
|
131
|
-
|
|
132
|
-
|
|
114
|
+
Each Engine owns at most one claim and claims only immediately before
|
|
115
|
+
execution. A node with concurrency `N` therefore holds at most `N` active
|
|
116
|
+
claims, regardless of channel depth.
|
|
133
117
|
|
|
134
|
-
|
|
118
|
+
Channel order defaults to `:weighted`:
|
|
135
119
|
|
|
136
|
-
- `:weighted` uses bounded weighted round-robin
|
|
137
|
-
|
|
138
|
-
- `:
|
|
139
|
-
starve lower-priority queues while a higher-priority queue remains busy;
|
|
140
|
-
- `:random` samples the configured weighted queue list on each reservation.
|
|
120
|
+
- `:weighted` uses bounded weighted round-robin;
|
|
121
|
+
- `:priority` checks channels in declaration order;
|
|
122
|
+
- `:shuffle` samples the configured weighted channel list.
|
|
141
123
|
|
|
142
|
-
Paused
|
|
124
|
+
Paused channels are excluded before claiming.
|
|
143
125
|
|
|
144
126
|
## Retry storms
|
|
145
127
|
|
|
146
|
-
Retries use capped exponential backoff with equal jitter. For
|
|
147
|
-
ceiling is `min(retry_base_delay * 2**n, retry_max_delay)` and the
|
|
128
|
+
Retries use capped exponential backoff with equal jitter. For failure `n`, the
|
|
129
|
+
ceiling is `min(retry_base_delay * 2**(n - 1), retry_max_delay)` and the delay
|
|
148
130
|
is distributed between half and all of that ceiling. Defaults are 15 seconds
|
|
149
|
-
and one hour.
|
|
131
|
+
and one hour.
|
data/exe/solid-jobs
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
require "solid_jobs"
|
|
4
|
-
require "solid_jobs/
|
|
4
|
+
require "solid_jobs/rails_adapter"
|
|
5
5
|
|
|
6
6
|
module ActiveJob
|
|
7
7
|
module QueueAdapters
|
|
@@ -27,16 +27,15 @@ module ActiveJob
|
|
|
27
27
|
private
|
|
28
28
|
|
|
29
29
|
def push(job, at: nil)
|
|
30
|
-
|
|
31
|
-
"
|
|
30
|
+
envelope = {
|
|
31
|
+
"task" => SolidJobs::RailsAdapter::AdapterTask,
|
|
32
32
|
"wrapped" => job.class.name,
|
|
33
|
-
"
|
|
34
|
-
"
|
|
33
|
+
"channel" => job.queue_name,
|
|
34
|
+
"arguments" => [job.serialize],
|
|
35
35
|
}
|
|
36
|
-
|
|
37
|
-
SolidJobs::
|
|
36
|
+
envelope["run_at"] = at if at
|
|
37
|
+
SolidJobs::Publisher.publish(envelope)
|
|
38
38
|
end
|
|
39
39
|
end
|
|
40
40
|
end
|
|
41
41
|
end
|
|
42
|
-
|
|
@@ -3,36 +3,36 @@
|
|
|
3
3
|
require "logger"
|
|
4
4
|
|
|
5
5
|
module SolidJobs
|
|
6
|
-
class
|
|
7
|
-
|
|
8
|
-
"
|
|
9
|
-
"
|
|
6
|
+
class Blueprint
|
|
7
|
+
DEFAULT_TASK_OPTIONS = {
|
|
8
|
+
"channel" => "default",
|
|
9
|
+
"max_failures" => 25,
|
|
10
10
|
}.freeze
|
|
11
11
|
|
|
12
|
-
attr_accessor :concurrency, :
|
|
12
|
+
attr_accessor :concurrency, :discarded_limit, :discarded_retention, :logger,
|
|
13
13
|
:on_complex_arguments, :poll_interval_average, :shutdown_timeout,
|
|
14
14
|
:reliable_fetch, :retry_base_delay, :retry_max_delay
|
|
15
|
-
attr_reader :
|
|
16
|
-
:
|
|
15
|
+
attr_reader :default_task_options, :error_handlers, :execute_interceptors,
|
|
16
|
+
:publish_interceptors, :redis_config
|
|
17
17
|
|
|
18
|
-
def initialize(redis: nil, concurrency: 5,
|
|
18
|
+
def initialize(redis: nil, concurrency: 5, channels: ["default"])
|
|
19
19
|
@redis_config = redis || SolidRedis.config(url: ENV.fetch("REDIS_URL", "redis://127.0.0.1:6379/0"))
|
|
20
20
|
@concurrency = Integer(concurrency)
|
|
21
|
-
@
|
|
22
|
-
@
|
|
23
|
-
@
|
|
24
|
-
@
|
|
21
|
+
@channels = normalize_channels(channels)
|
|
22
|
+
@default_task_options = DEFAULT_TASK_OPTIONS
|
|
23
|
+
@publish_interceptors = InterceptorRegistry.new
|
|
24
|
+
@execute_interceptors = InterceptorRegistry.new
|
|
25
25
|
@error_handlers = []
|
|
26
26
|
@lifecycle_callbacks = Hash.new { |hash, event| hash[event] = [] }
|
|
27
27
|
@on_complex_arguments = :raise
|
|
28
28
|
@poll_interval_average = 5.0
|
|
29
29
|
@shutdown_timeout = 25.0
|
|
30
|
-
@
|
|
30
|
+
@channel_order = :weighted
|
|
31
31
|
@reliable_fetch = true
|
|
32
32
|
@retry_base_delay = 15.0
|
|
33
33
|
@retry_max_delay = 3_600.0
|
|
34
|
-
@
|
|
35
|
-
@
|
|
34
|
+
@discarded_limit = 10_000
|
|
35
|
+
@discarded_retention = 180 * 24 * 60 * 60
|
|
36
36
|
@logger = Logger.new($stdout)
|
|
37
37
|
@redis_pool = nil
|
|
38
38
|
end
|
|
@@ -50,42 +50,42 @@ module SolidJobs
|
|
|
50
50
|
redis_pool.with { |connection| yield connection }
|
|
51
51
|
end
|
|
52
52
|
|
|
53
|
-
def
|
|
54
|
-
@
|
|
53
|
+
def channels
|
|
54
|
+
@channels.map(&:first)
|
|
55
55
|
end
|
|
56
56
|
|
|
57
|
-
def
|
|
58
|
-
@
|
|
57
|
+
def channels=(values)
|
|
58
|
+
@channels = normalize_channels(values)
|
|
59
59
|
end
|
|
60
60
|
|
|
61
|
-
def
|
|
62
|
-
@
|
|
61
|
+
def channel_entries
|
|
62
|
+
@channels.dup
|
|
63
63
|
end
|
|
64
64
|
|
|
65
|
-
def
|
|
66
|
-
@
|
|
65
|
+
def channel_order
|
|
66
|
+
@channel_order
|
|
67
67
|
end
|
|
68
68
|
|
|
69
|
-
def
|
|
69
|
+
def channel_order=(mode)
|
|
70
70
|
mode = mode.to_sym
|
|
71
|
-
unless %i[
|
|
72
|
-
raise ArgumentError, "
|
|
71
|
+
unless %i[priority weighted shuffle].include?(mode)
|
|
72
|
+
raise ArgumentError, "channel_order must be :priority, :weighted, or :shuffle"
|
|
73
73
|
end
|
|
74
74
|
|
|
75
|
-
@
|
|
75
|
+
@channel_order = mode
|
|
76
76
|
end
|
|
77
77
|
|
|
78
|
-
def
|
|
79
|
-
|
|
78
|
+
def prioritized?
|
|
79
|
+
channel_order == :priority
|
|
80
80
|
end
|
|
81
81
|
|
|
82
|
-
def
|
|
83
|
-
self.
|
|
82
|
+
def prioritized=(value)
|
|
83
|
+
self.channel_order = value ? :priority : :weighted
|
|
84
84
|
end
|
|
85
85
|
|
|
86
|
-
def
|
|
87
|
-
@
|
|
88
|
-
|
|
86
|
+
def default_task_options=(options)
|
|
87
|
+
@default_task_options = Utilities.shareable_copy(
|
|
88
|
+
DEFAULT_TASK_OPTIONS.merge(Utilities.stringify_keys(options)),
|
|
89
89
|
)
|
|
90
90
|
end
|
|
91
91
|
|
|
@@ -111,26 +111,26 @@ module SolidJobs
|
|
|
111
111
|
end
|
|
112
112
|
|
|
113
113
|
def inspect
|
|
114
|
-
"#<#{self.class.name} concurrency=#{concurrency}
|
|
114
|
+
"#<#{self.class.name} concurrency=#{concurrency} channels=#{channels.inspect}>"
|
|
115
115
|
end
|
|
116
116
|
|
|
117
117
|
def ractor_snapshot
|
|
118
118
|
Utilities.shareable_copy(
|
|
119
119
|
redis_config: redis_config,
|
|
120
120
|
concurrency: concurrency,
|
|
121
|
-
|
|
122
|
-
|
|
121
|
+
channels: channel_entries,
|
|
122
|
+
channel_order: channel_order,
|
|
123
123
|
reliable_fetch: reliable_fetch,
|
|
124
|
-
|
|
124
|
+
default_task_options: default_task_options,
|
|
125
125
|
on_complex_arguments: on_complex_arguments,
|
|
126
126
|
poll_interval_average: poll_interval_average,
|
|
127
127
|
shutdown_timeout: shutdown_timeout,
|
|
128
|
-
|
|
129
|
-
|
|
128
|
+
discarded_limit: discarded_limit,
|
|
129
|
+
discarded_retention: discarded_retention,
|
|
130
130
|
retry_base_delay: retry_base_delay,
|
|
131
131
|
retry_max_delay: retry_max_delay,
|
|
132
|
-
|
|
133
|
-
|
|
132
|
+
publish_interceptors: publish_interceptors.export,
|
|
133
|
+
execute_interceptors: execute_interceptors.export,
|
|
134
134
|
)
|
|
135
135
|
end
|
|
136
136
|
|
|
@@ -138,31 +138,31 @@ module SolidJobs
|
|
|
138
138
|
config = new(
|
|
139
139
|
redis: snapshot.fetch(:redis_config),
|
|
140
140
|
concurrency: snapshot.fetch(:concurrency),
|
|
141
|
-
|
|
141
|
+
channels: snapshot.fetch(:channels),
|
|
142
142
|
)
|
|
143
|
-
config.
|
|
143
|
+
config.channel_order = snapshot.fetch(:channel_order)
|
|
144
144
|
config.reliable_fetch = snapshot.fetch(:reliable_fetch)
|
|
145
|
-
config.
|
|
145
|
+
config.default_task_options = snapshot.fetch(:default_task_options)
|
|
146
146
|
config.on_complex_arguments = snapshot.fetch(:on_complex_arguments)
|
|
147
147
|
config.poll_interval_average = snapshot.fetch(:poll_interval_average)
|
|
148
148
|
config.shutdown_timeout = snapshot.fetch(:shutdown_timeout)
|
|
149
|
-
config.
|
|
150
|
-
config.
|
|
149
|
+
config.discarded_limit = snapshot.fetch(:discarded_limit)
|
|
150
|
+
config.discarded_retention = snapshot.fetch(:discarded_retention)
|
|
151
151
|
config.retry_base_delay = snapshot.fetch(:retry_base_delay)
|
|
152
152
|
config.retry_max_delay = snapshot.fetch(:retry_max_delay)
|
|
153
|
-
config.
|
|
154
|
-
config.
|
|
153
|
+
config.publish_interceptors.import(snapshot.fetch(:publish_interceptors))
|
|
154
|
+
config.execute_interceptors.import(snapshot.fetch(:execute_interceptors))
|
|
155
155
|
config
|
|
156
156
|
end
|
|
157
157
|
|
|
158
158
|
private
|
|
159
159
|
|
|
160
|
-
def
|
|
160
|
+
def normalize_channels(values)
|
|
161
161
|
Array(values).map do |value|
|
|
162
162
|
name, weight = value.is_a?(Array) ? value : [value, 1]
|
|
163
163
|
name = String(name)
|
|
164
164
|
weight = Integer(weight)
|
|
165
|
-
raise ArgumentError, "
|
|
165
|
+
raise ArgumentError, "Channel weight must be positive" unless weight.positive?
|
|
166
166
|
|
|
167
167
|
[name.freeze, weight].freeze
|
|
168
168
|
end
|