solid-jobs 0.1.2 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +38 -2
- data/README.md +137 -117
- data/docs/reliability.md +71 -89
- data/lib/active_job/queue_adapters/solid_jobs_adapter.rb +6 -7
- data/lib/solid_jobs/active_job.rb +2 -5
- data/lib/solid_jobs/api.rb +133 -156
- data/lib/solid_jobs/claim.rb +315 -0
- data/lib/solid_jobs/cli.rb +9 -6
- data/lib/solid_jobs/config.rb +50 -50
- data/lib/solid_jobs/executor.rb +34 -0
- data/lib/solid_jobs/failure_policy.rb +98 -0
- data/lib/solid_jobs/heartbeat.rb +15 -13
- data/lib/solid_jobs/integrity_check.rb +39 -36
- data/lib/solid_jobs/interceptor_registry.rb +57 -0
- data/lib/solid_jobs/keyspace.rb +42 -0
- data/lib/solid_jobs/processor.rb +30 -30
- data/lib/solid_jobs/publisher.rb +186 -0
- data/lib/solid_jobs/recovery.rb +8 -7
- data/lib/solid_jobs/server.rb +2 -2
- data/lib/solid_jobs/task.rb +147 -0
- data/lib/solid_jobs/testing.rb +18 -18
- data/lib/solid_jobs/{scheduler.rb → timer.rb} +8 -8
- data/lib/solid_jobs/version.rb +1 -2
- data/lib/solid_jobs.rb +18 -18
- metadata +13 -32
- data/lib/solid_jobs/client.rb +0 -202
- data/lib/solid_jobs/fetch.rb +0 -310
- data/lib/solid_jobs/job.rb +0 -156
- data/lib/solid_jobs/middleware/chain.rb +0 -104
- data/lib/solid_jobs/retry_service.rb +0 -100
- data/lib/solid_jobs/worker.rb +0 -35
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 8c178515fb3c56d80f0380da9f25114cd44679135b4d3d39fb468c45760748a7
|
|
4
|
+
data.tar.gz: 94080a99ea85f351714722d4f4f073ab99f01b2e5096a8e304670138171cf1ab
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: fc05837b46c971dfd517cedac8bd46960867160cce776384eced29d41fac3a787b1f45a1a8f2d21ebeddbdb828766087a01396f245a255e5e77443bfa5366a2f
|
|
7
|
+
data.tar.gz: af56380ec6a6e82470d4df8c477b0e87eec5dc9aaddb9d04fcad32f6d630f695254c77008b2c4cb4400b496a2a5a566a4b1e58fec2b93d536ad3447868146099
|
data/CHANGELOG.md
CHANGED
|
@@ -2,7 +2,40 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to this project will be documented in this file.
|
|
4
4
|
|
|
5
|
-
## [
|
|
5
|
+
## [0.2.0] - 2026-10-03
|
|
6
|
+
|
|
7
|
+
### Breaking
|
|
8
|
+
|
|
9
|
+
- Replace the previous job API with the independent `SolidJobs::Task` API:
|
|
10
|
+
`enqueue`, `enqueue_after`, `enqueue_at`, `enqueue_many`, `execute`, and
|
|
11
|
+
`with_options`.
|
|
12
|
+
- Replace the previous payload with a SolidJobs envelope using `id`, `task`,
|
|
13
|
+
`arguments`, `channel`, `run_at`, `created_ms`, and `queued_ms`.
|
|
14
|
+
- Move all Redis data into the `solid_jobs:` keyspace. Existing queued data is
|
|
15
|
+
not read or migrated automatically.
|
|
16
|
+
- Replace middleware chains with `publish_interceptors` and
|
|
17
|
+
`execute_interceptors`, whose interceptors implement `around(context)`.
|
|
18
|
+
- Replace the administrative API with `Metrics`, `Channel`, `StoredTask`,
|
|
19
|
+
`PlannedTasks`, `RetryingTasks`, `DiscardedTasks`, `Node`, `Nodes`,
|
|
20
|
+
`Execution`, and `Claims`.
|
|
21
|
+
- Replace queue configuration with channels and the `:weighted`, `:priority`,
|
|
22
|
+
and `:shuffle` ordering modes.
|
|
23
|
+
- Remove all Sidekiq compatibility aliases and wire-format compatibility.
|
|
24
|
+
|
|
25
|
+
### Changed
|
|
26
|
+
|
|
27
|
+
- Claims use UUID task IDs and claim tokens, with generation-fenced completion
|
|
28
|
+
and requeue operations.
|
|
29
|
+
- Failure handling uses `max_failures`, `retry_within`, `retry_delay`, and
|
|
30
|
+
`after_final_failure`.
|
|
31
|
+
- Testing modes are now `:capture` and `:execute`.
|
|
32
|
+
|
|
33
|
+
## [0.1.3] - 2026-10-01
|
|
34
|
+
|
|
35
|
+
### Changed
|
|
36
|
+
|
|
37
|
+
- Require `solid-redis` 1.0.11 and rely on its `solid-resp-ractor` dependency
|
|
38
|
+
instead of declaring and loading the RESP codec directly.
|
|
6
39
|
|
|
7
40
|
## [0.1.2] - 2026-10-01
|
|
8
41
|
|
|
@@ -56,5 +89,8 @@ All notable changes to this project will be documented in this file.
|
|
|
56
89
|
- Reliable-hot-path and CPU-scaling profilers plus the independent Sidekiq
|
|
57
90
|
comparison benchmark.
|
|
58
91
|
|
|
59
|
-
[
|
|
92
|
+
[0.2.0]: https://github.com/nicolasva/solid-jobs/compare/v0.1.3...v0.2.0
|
|
93
|
+
[0.1.3]: https://github.com/nicolasva/solid-jobs/compare/v0.1.2...v0.1.3
|
|
94
|
+
[0.1.2]: https://github.com/nicolasva/solid-jobs/compare/v0.1.1...v0.1.2
|
|
95
|
+
[0.1.1]: https://github.com/nicolasva/solid-jobs/compare/v0.1.0...v0.1.1
|
|
60
96
|
[0.1.0]: https://github.com/nicolasva/solid-jobs/releases/tag/v0.1.0
|
data/README.md
CHANGED
|
@@ -1,163 +1,183 @@
|
|
|
1
1
|
# SolidJobs
|
|
2
2
|
|
|
3
3
|
[](https://github.com/nicolasva/solid-jobs/actions/workflows/ci.yml)
|
|
4
|
-
[](https://codeclimate.com/github/nicolasva/solid-jobs)
|
|
5
4
|
[](https://rubygems.org/gems/solid-jobs)
|
|
6
|
-
[](https://www.rubydoc.info/gems/solid-jobs)
|
|
7
|
-
[](https://rubygems.org/gems/solid-jobs)
|
|
8
5
|
|
|
9
|
-
SolidJobs is a Ractor-oriented Redis
|
|
10
|
-
|
|
11
|
-
|
|
6
|
+
SolidJobs is a Ractor-oriented Redis task runner for Ruby. It provides its own
|
|
7
|
+
task API, Redis envelope, keyspace, interceptor model, failure policy, and
|
|
8
|
+
inspection API.
|
|
12
9
|
|
|
13
|
-
SolidJobs is
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
SolidJobs and its required gems use pure Ruby and do not require native
|
|
18
|
-
extensions. Hot paths are designed around bounded buffers, reusable immutable
|
|
19
|
-
configuration, and low-allocation command batches.
|
|
10
|
+
SolidJobs is **not compatible with Sidekiq**. It does not read Sidekiq queues,
|
|
11
|
+
does not write Sidekiq payloads, and does not expose Sidekiq's job API. Existing
|
|
12
|
+
applications and queued work must be migrated explicitly.
|
|
20
13
|
|
|
21
14
|
## Installation
|
|
22
15
|
|
|
23
|
-
Add SolidJobs 0.1 to your bundle:
|
|
24
|
-
|
|
25
16
|
```ruby
|
|
26
|
-
gem "solid-jobs"
|
|
17
|
+
gem "solid-jobs"
|
|
27
18
|
```
|
|
28
19
|
|
|
29
|
-
Then run
|
|
30
|
-
|
|
31
|
-
```sh
|
|
32
|
-
bundle install
|
|
33
|
-
```
|
|
34
|
-
|
|
35
|
-
## Delivery semantics
|
|
20
|
+
Then run `bundle install`.
|
|
36
21
|
|
|
37
|
-
|
|
38
|
-
job from `queue:<name>` to a process reservation list before execution and
|
|
39
|
-
removes it only after a successful ACK. Graceful shutdown requeues unfinished
|
|
40
|
-
jobs; reservations owned by a crashed process are recovered into their
|
|
41
|
-
original queues.
|
|
42
|
-
|
|
43
|
-
A process can still crash after the application side effect and before the
|
|
44
|
-
ACK. The recovered job will then run again. Jobs must therefore be idempotent
|
|
45
|
-
or implement an application-level deduplication key when duplicate side
|
|
46
|
-
effects are unsafe. SolidJobs does not claim exactly-once execution.
|
|
22
|
+
## Defining and submitting tasks
|
|
47
23
|
|
|
48
24
|
```ruby
|
|
49
|
-
class
|
|
50
|
-
include SolidJobs::
|
|
25
|
+
class RecalculateAccount
|
|
26
|
+
include SolidJobs::Task
|
|
51
27
|
|
|
52
|
-
|
|
28
|
+
task_options channel: "critical", max_failures: 10
|
|
53
29
|
|
|
54
30
|
def perform(account_id)
|
|
55
31
|
Account.find(account_id).recalculate!
|
|
56
32
|
end
|
|
57
33
|
end
|
|
58
34
|
|
|
59
|
-
|
|
60
|
-
|
|
35
|
+
RecalculateAccount.enqueue(42)
|
|
36
|
+
RecalculateAccount.enqueue_after(30, 42)
|
|
37
|
+
RecalculateAccount.enqueue_at(Time.now + 300, 42)
|
|
38
|
+
RecalculateAccount.enqueue_many([[42], [43], [44]])
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
The persisted envelope is specific to SolidJobs:
|
|
42
|
+
|
|
43
|
+
```json
|
|
44
|
+
{
|
|
45
|
+
"id": "7ec33ed5-1b77-4bde-8960-bd77231f10ef",
|
|
46
|
+
"task": "RecalculateAccount",
|
|
47
|
+
"arguments": [42],
|
|
48
|
+
"channel": "critical",
|
|
49
|
+
"max_failures": 10,
|
|
50
|
+
"created_ms": 1790981000000,
|
|
51
|
+
"queued_ms": 1790981000001
|
|
52
|
+
}
|
|
61
53
|
```
|
|
62
54
|
|
|
63
|
-
|
|
55
|
+
All internal Redis keys use the `solid_jobs:` namespace. Ready work is stored
|
|
56
|
+
under `solid_jobs:channel:<name>`; planned, retrying, discarded, node, claim,
|
|
57
|
+
attempt, and metric data use separate namespaced keys.
|
|
58
|
+
|
|
59
|
+
Run executors:
|
|
64
60
|
|
|
65
61
|
```sh
|
|
66
62
|
bundle exec solid-jobs --require ./config/environment \
|
|
67
|
-
--concurrency 8 --
|
|
63
|
+
--concurrency 8 --channel critical,3 --channel default
|
|
68
64
|
```
|
|
69
65
|
|
|
70
|
-
##
|
|
71
|
-
|
|
72
|
-
SolidJobs uses Minitest exclusively:
|
|
66
|
+
## Configuration
|
|
73
67
|
|
|
74
|
-
```
|
|
75
|
-
|
|
76
|
-
|
|
68
|
+
```ruby
|
|
69
|
+
SolidJobs.configure do |config|
|
|
70
|
+
config.channels = [["critical", 3], "default"]
|
|
71
|
+
config.channel_order = :weighted
|
|
72
|
+
config.concurrency = 8
|
|
73
|
+
config.shutdown_timeout = 25
|
|
74
|
+
config.default_task_options = {
|
|
75
|
+
channel: "default",
|
|
76
|
+
max_failures: 25,
|
|
77
|
+
}
|
|
78
|
+
end
|
|
79
|
+
```
|
|
77
80
|
|
|
78
|
-
|
|
79
|
-
|
|
81
|
+
`channel_order` accepts `:weighted` for bounded weighted round-robin,
|
|
82
|
+
`:priority` for declaration-order priority, and `:shuffle` for random
|
|
83
|
+
selection from the weighted channel list.
|
|
80
84
|
|
|
81
|
-
|
|
82
|
-
SOLID_JOBS_TORTURE=1 STRESS_JOBS=100000 bundle exec rake torture
|
|
85
|
+
## Interceptors
|
|
83
86
|
|
|
84
|
-
|
|
85
|
-
|
|
87
|
+
Publication and execution use SolidJobs interceptors. An interceptor implements
|
|
88
|
+
`around(context)` and yields to continue:
|
|
86
89
|
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
90
|
+
```ruby
|
|
91
|
+
class TraceExecution
|
|
92
|
+
def around(context)
|
|
93
|
+
Telemetry.start(context.envelope.fetch("id"))
|
|
94
|
+
yield
|
|
95
|
+
ensure
|
|
96
|
+
Telemetry.finish
|
|
97
|
+
end
|
|
98
|
+
end
|
|
91
99
|
|
|
92
|
-
|
|
93
|
-
|
|
100
|
+
SolidJobs.configure do |config|
|
|
101
|
+
config.execute_interceptors.use(TraceExecution)
|
|
102
|
+
end
|
|
94
103
|
```
|
|
95
104
|
|
|
96
|
-
|
|
97
|
-
|
|
105
|
+
Publication interceptors receive `SolidJobs::Publisher::Publication`; execution
|
|
106
|
+
interceptors receive `SolidJobs::Executor::Execution`.
|
|
98
107
|
|
|
99
|
-
|
|
108
|
+
## Failure handling
|
|
100
109
|
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
110
|
+
Tasks default to 25 failures. A task can customize the delay or final action:
|
|
111
|
+
|
|
112
|
+
```ruby
|
|
113
|
+
class ImportCatalog
|
|
114
|
+
include SolidJobs::Task
|
|
115
|
+
|
|
116
|
+
task_options max_failures: 8, retry_within: 3_600
|
|
117
|
+
|
|
118
|
+
retry_delay do |failure_count, error, envelope|
|
|
119
|
+
:drop if error.is_a?(InvalidCatalog)
|
|
120
|
+
end
|
|
121
|
+
|
|
122
|
+
after_final_failure do |envelope, error|
|
|
123
|
+
Alerts.catalog_import_failed(envelope.fetch("id"), error)
|
|
124
|
+
end
|
|
125
|
+
end
|
|
105
126
|
```
|
|
106
127
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
128
|
+
`retry_delay` may return a delay in seconds, `:drop`, or `:archive`. Without an
|
|
129
|
+
override, SolidJobs uses capped exponential backoff with equal jitter.
|
|
130
|
+
|
|
131
|
+
## Delivery semantics
|
|
132
|
+
|
|
133
|
+
SolidJobs provides **at-least-once** delivery. An executor atomically claims a
|
|
134
|
+
task before execution and completes the claim only after `perform` returns.
|
|
135
|
+
Graceful shutdown requeues unfinished tasks. Claims owned by a crashed node are
|
|
136
|
+
recovered into their original channels.
|
|
137
|
+
|
|
138
|
+
A process can still crash after an application side effect and before claim
|
|
139
|
+
completion. The recovered task then runs again. Tasks must be idempotent or use
|
|
140
|
+
application-level deduplication when duplicate effects are unsafe.
|
|
141
|
+
|
|
142
|
+
See [docs/reliability.md](docs/reliability.md) for the state machine, claim
|
|
143
|
+
fencing, recovery rules, and Ruby Ractor caveats.
|
|
144
|
+
|
|
145
|
+
## Testing
|
|
111
146
|
|
|
112
|
-
|
|
147
|
+
```ruby
|
|
148
|
+
SolidJobs.testing!(:capture) do
|
|
149
|
+
RecalculateAccount.enqueue(42)
|
|
150
|
+
RecalculateAccount.captured
|
|
151
|
+
end
|
|
152
|
+
|
|
153
|
+
SolidJobs.testing!(:execute) do
|
|
154
|
+
RecalculateAccount.enqueue(42)
|
|
155
|
+
end
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
Run the project suites:
|
|
113
159
|
|
|
114
160
|
```sh
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
bundle exec rake
|
|
161
|
+
bundle exec rake test
|
|
162
|
+
STRESS_JOBS=10000 bundle exec rake stress
|
|
163
|
+
SOLID_JOBS_TORTURE=1 STRESS_JOBS=100000 bundle exec rake torture
|
|
164
|
+
SOLID_JOBS_SOAK=1 SOLID_JOBS_SOAK_SECONDS=86400 bundle exec rake soak
|
|
118
165
|
```
|
|
119
166
|
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
|
138
|
-
| 1 | 142 | 138 | 100.0% | 6.95 | 7.11 | 41.5 MiB | 72.0 MiB | 129.8 | 89.9 |
|
|
139
|
-
| 2 | 144 | 272 | 98.7% | 6.94 | 7.16 | 41.8 MiB | 77.6 MiB | 117.5 | 81.3 |
|
|
140
|
-
| 4 | 143 | 539 | 97.8% | 6.99 | 7.18 | 41.9 MiB | 71.4 MiB | 111.9 | 77.2 |
|
|
141
|
-
| 8 | 144 | 957 | 86.7% | 6.94 | 8.01 | 42.2 MiB | 69.5 MiB | 108.7 | 75.4 |
|
|
142
|
-
|
|
143
|
-
At eight concurrency units, SolidJobs reaches 957 jobs/s versus 144 jobs/s
|
|
144
|
-
for one Sidekiq process. This is Ractor parallelism rather than equal CPU
|
|
145
|
-
efficiency: SolidJobs consumes 766% CPU and 8.01 CPU-seconds per 1,000 jobs,
|
|
146
|
-
while Sidekiq consumes 100% CPU and 6.94 CPU-seconds per 1,000 jobs. SolidJobs
|
|
147
|
-
also uses more RSS, but reaches 13.77 jobs/s/MiB versus 3.41 for Sidekiq.
|
|
148
|
-
|
|
149
|
-
These are local synthetic measurements, not application-capacity claims.
|
|
150
|
-
Queue p95/p99 values in this run use sparse sampling and are excluded from the
|
|
151
|
-
summary until the final latency campaign increases the sample count. Ruby
|
|
152
|
-
3.4.4 eight-Ractor results are also excluded: Ruby 3.4's Ractor scheduler
|
|
153
|
-
crashed (concurrent TCP/RESP initialization, reproduced on macOS arm64 and
|
|
154
|
-
Linux x86_64) or deadlocked VM-wide on a GC barrier under cross-Ractor
|
|
155
|
-
message traffic (`test/support/ractor_barrier_repro.rb` reproduces it without
|
|
156
|
-
SolidJobs).
|
|
157
|
-
Ruby 4.0.1 passed the equivalent reproducers, and `StartupBarrier` serializes
|
|
158
|
-
component initialization before releasing normal parallel processing.
|
|
159
|
-
**Multi-Ractor servers are recommended on Ruby ≥ 4.0**; see
|
|
160
|
-
`docs/reliability.md`.
|
|
161
|
-
|
|
162
|
-
The project is under active development. The Web UI and commercial Sidekiq
|
|
163
|
-
features are not part of the initial scope.
|
|
167
|
+
## Migrating from Sidekiq
|
|
168
|
+
|
|
169
|
+
There is no transparent migration path because compatibility is intentionally
|
|
170
|
+
absent:
|
|
171
|
+
|
|
172
|
+
1. Replace `include Sidekiq::Job` with `include SolidJobs::Task`.
|
|
173
|
+
2. Replace `sidekiq_options` with `task_options`.
|
|
174
|
+
3. Replace `perform_async`, `perform_in`, and `perform_bulk` with `enqueue`,
|
|
175
|
+
`enqueue_after`, and `enqueue_many`.
|
|
176
|
+
4. Replace middleware with SolidJobs interceptors.
|
|
177
|
+
5. Drain or export existing Sidekiq queues before switching. SolidJobs will not
|
|
178
|
+
consume them.
|
|
179
|
+
6. Start SolidJobs with `--channel`; `--queue` is not accepted.
|
|
180
|
+
|
|
181
|
+
The Active Job adapter remains available as
|
|
182
|
+
`ActiveJob::QueueAdapters::SolidJobsAdapter`, but it writes only SolidJobs
|
|
183
|
+
envelopes and keys.
|
data/docs/reliability.md
CHANGED
|
@@ -1,63 +1,61 @@
|
|
|
1
1
|
# Reliability model
|
|
2
2
|
|
|
3
|
-
SolidJobs uses a Redis-backed at-least-once state machine:
|
|
3
|
+
SolidJobs uses a namespaced Redis-backed at-least-once state machine:
|
|
4
4
|
|
|
5
5
|
```text
|
|
6
6
|
READY
|
|
7
7
|
|
|
|
8
|
-
|
|
|
8
|
+
| atomic claim
|
|
9
9
|
v
|
|
10
|
-
|
|
10
|
+
CLAIMED ---- fenced completion ----> removed
|
|
11
11
|
|
|
|
12
|
-
+---- application error ---->
|
|
12
|
+
+---- application error ----> RETRYING or DISCARDED, then completion
|
|
13
13
|
+---- graceful timeout -----> READY
|
|
14
|
-
+----
|
|
14
|
+
+---- node crash -----------> claim retained
|
|
15
15
|
|
|
|
16
16
|
+---- recovery ----> READY
|
|
17
17
|
```
|
|
18
18
|
|
|
19
|
-
|
|
20
|
-
|
|
19
|
+
Ready tasks live in `solid_jobs:channel:<name>`. Each executor owns at most one
|
|
20
|
+
claimed-task list and one claim journal entry:
|
|
21
21
|
|
|
22
22
|
```text
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
23
|
+
task_id stable across replay
|
|
24
|
+
claim_token unique per execution attempt
|
|
25
|
+
node_id owning server identity
|
|
26
|
+
executor_id owning Executor Ractor
|
|
27
|
+
channel destination used by recovery
|
|
28
|
+
attempt monotonically increasing execution count
|
|
29
|
+
claimed_at wall-clock claim time
|
|
29
30
|
```
|
|
30
31
|
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
or revived
|
|
34
|
-
supervisor reuses the same processor slot.
|
|
32
|
+
Completion and requeue are fenced by `claim_token`. Their Lua scripts verify
|
|
33
|
+
that the executor still owns the current journal generation before changing
|
|
34
|
+
state. A delayed or revived executor cannot complete or requeue a newer claim.
|
|
35
35
|
|
|
36
|
-
The
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
is not stolen while its owner remains alive.
|
|
36
|
+
The envelope retains its canonical `channel`, allowing another node to restore
|
|
37
|
+
it after its owner dies. Recovery uses node liveness and heartbeat, never claim
|
|
38
|
+
age alone, so a legitimate long-running task is not stolen.
|
|
40
39
|
|
|
41
40
|
## Failure boundaries
|
|
42
41
|
|
|
43
|
-
- Before
|
|
44
|
-
- After
|
|
45
|
-
- After the application effect but before
|
|
46
|
-
- Redis unavailable during
|
|
47
|
-
- Graceful shutdown: the active
|
|
42
|
+
- Before claim: the task remains in `solid_jobs:channel:<name>`.
|
|
43
|
+
- After claim or during `perform`: the task remains claimed.
|
|
44
|
+
- After the application effect but before completion: recovery replays it.
|
|
45
|
+
- Redis unavailable during completion: the task remains claimed and is replayed.
|
|
46
|
+
- Graceful shutdown: the active task may finish within the configured timeout;
|
|
48
47
|
otherwise it is interrupted and requeued.
|
|
49
|
-
- `SIGKILL`:
|
|
48
|
+
- `SIGKILL`: recovery relies exclusively on Redis state.
|
|
50
49
|
|
|
51
|
-
This design
|
|
52
|
-
|
|
50
|
+
This design favors no task loss over duplicate suppression. Exactly-once side
|
|
51
|
+
effects require application-level idempotency.
|
|
53
52
|
|
|
54
53
|
## Startup isolation
|
|
55
54
|
|
|
56
|
-
The server does not
|
|
57
|
-
Processor, and
|
|
58
|
-
|
|
59
|
-
`SolidJobs::StartupBarrier
|
|
60
|
-
ready:
|
|
55
|
+
The server does not claim work while components are booting. Heartbeat,
|
|
56
|
+
Processor, and Timer Ractors initialize their local configuration and Redis
|
|
57
|
+
pool, report `READY` exactly once, and wait behind
|
|
58
|
+
`SolidJobs::StartupBarrier`:
|
|
61
59
|
|
|
62
60
|
```text
|
|
63
61
|
BOOTING -> ALL_READY -> RUNNING
|
|
@@ -65,85 +63,69 @@ BOOTING -> ALL_READY -> RUNNING
|
|
|
65
63
|
```
|
|
66
64
|
|
|
67
65
|
Boot failure is terminal. Already-ready components receive `:abort`, close
|
|
68
|
-
their local resources, and never enter their
|
|
69
|
-
idempotent.
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
Ruby
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
`
|
|
80
|
-
`
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
`:work`/`:done`/`:stats` channel, so any multi-Processor server on Ruby 3.4
|
|
85
|
-
is exposed; once frozen, neither `Timeout` nor process exit
|
|
86
|
-
(`rb_ractor_terminate_all`) can recover.
|
|
87
|
-
|
|
88
|
-
Recommendation: **run multi-Ractor SolidJobs servers on Ruby ≥ 4.0**. On
|
|
89
|
-
Ruby 3.4 use the client/API side freely, and prefer one process per
|
|
90
|
-
Processor (`concurrency: 1`) for the server. Server-based stress tests are
|
|
91
|
-
skipped on Ruby < 4 for this reason, and `Server#start` logs a warning when
|
|
92
|
-
it detects `concurrency > 1` on Ruby < 4.
|
|
93
|
-
|
|
94
|
-
### Bounded shutdown
|
|
66
|
+
their local resources, and never enter their claim loops. Cleanup is
|
|
67
|
+
idempotent.
|
|
68
|
+
|
|
69
|
+
## Ruby 3.4 Ractor caveat
|
|
70
|
+
|
|
71
|
+
Ruby 3.4's Ractor scheduler can deadlock the VM when a GC-triggered scheduler
|
|
72
|
+
barrier runs while several Ractors exchange moved messages.
|
|
73
|
+
`test/support/ractor_barrier_repro.rb` reproduces this without SolidJobs or
|
|
74
|
+
Redis. Ruby 4.0.1 completes the equivalent reproducer.
|
|
75
|
+
|
|
76
|
+
Run multi-Ractor SolidJobs servers on Ruby >= 4.0. On Ruby 3.4, prefer one
|
|
77
|
+
process per Processor (`concurrency: 1`). Server stress tests are skipped on
|
|
78
|
+
Ruby versions affected by this runtime issue, and `Server#start` warns when it
|
|
79
|
+
detects a multi-Ractor configuration there.
|
|
80
|
+
|
|
81
|
+
## Bounded shutdown
|
|
95
82
|
|
|
96
83
|
`Server#stop` never waits forever for a component. Each Ractor gets
|
|
97
|
-
`shutdown_timeout + Server::STOP_GRACE` to return after `:stop`; the signal
|
|
98
|
-
|
|
99
|
-
|
|
84
|
+
`shutdown_timeout + Server::STOP_GRACE` to return after `:stop`; the signal is
|
|
85
|
+
re-sent up to `STOP_RESENDS` times before the component is abandoned with an
|
|
86
|
+
error log.
|
|
100
87
|
|
|
101
88
|
## Configuration scope
|
|
102
89
|
|
|
103
90
|
`SolidJobs.config` and the testing mode are Ractor-local, not thread-local.
|
|
104
|
-
Every thread and fiber inside a Ractor
|
|
105
|
-
|
|
106
|
-
Ractor keeps its own isolated configuration.
|
|
91
|
+
Every thread and fiber inside a Ractor shares its configuration, while each
|
|
92
|
+
Ractor owns isolated mutable state and Redis connections.
|
|
107
93
|
|
|
108
94
|
## Integrity auditing
|
|
109
95
|
|
|
110
|
-
Fault
|
|
111
|
-
Redis-backed state:
|
|
96
|
+
Fault tests can reconcile known task IDs against every Redis-backed state:
|
|
112
97
|
|
|
113
98
|
```ruby
|
|
114
99
|
report = SolidJobs::IntegrityCheck.call(
|
|
115
|
-
expected_job_ids:
|
|
100
|
+
expected_job_ids: submitted_task_ids,
|
|
116
101
|
acked_key: "test:completed",
|
|
117
102
|
)
|
|
118
103
|
|
|
119
104
|
raise report.inspect unless report.ok?
|
|
120
105
|
```
|
|
121
106
|
|
|
122
|
-
The report separates lost
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
sequence.
|
|
107
|
+
The report separates lost tasks, unexpected tasks, inconsistent claim
|
|
108
|
+
journals, dangling attempt indexes, duplicate active states, and malformed
|
|
109
|
+
envelopes. Completion removes the task's attempt index atomically; requeue
|
|
110
|
+
retains it so recovery increments the same attempt sequence.
|
|
127
111
|
|
|
128
|
-
## Backpressure and
|
|
112
|
+
## Backpressure and channels
|
|
129
113
|
|
|
130
|
-
Each Processor owns at most one
|
|
131
|
-
|
|
132
|
-
|
|
114
|
+
Each Processor owns at most one claim and claims only immediately before
|
|
115
|
+
execution. A node with concurrency `N` therefore holds at most `N` active
|
|
116
|
+
claims, regardless of channel depth.
|
|
133
117
|
|
|
134
|
-
|
|
118
|
+
Channel order defaults to `:weighted`:
|
|
135
119
|
|
|
136
|
-
- `:weighted` uses bounded weighted round-robin
|
|
137
|
-
|
|
138
|
-
- `:
|
|
139
|
-
starve lower-priority queues while a higher-priority queue remains busy;
|
|
140
|
-
- `:random` samples the configured weighted queue list on each reservation.
|
|
120
|
+
- `:weighted` uses bounded weighted round-robin;
|
|
121
|
+
- `:priority` checks channels in declaration order;
|
|
122
|
+
- `:shuffle` samples the configured weighted channel list.
|
|
141
123
|
|
|
142
|
-
Paused
|
|
124
|
+
Paused channels are excluded before claiming.
|
|
143
125
|
|
|
144
126
|
## Retry storms
|
|
145
127
|
|
|
146
|
-
Retries use capped exponential backoff with equal jitter. For
|
|
147
|
-
ceiling is `min(retry_base_delay * 2**n, retry_max_delay)` and the
|
|
128
|
+
Retries use capped exponential backoff with equal jitter. For failure `n`, the
|
|
129
|
+
ceiling is `min(retry_base_delay * 2**(n - 1), retry_max_delay)` and the delay
|
|
148
130
|
is distributed between half and all of that ceiling. Defaults are 15 seconds
|
|
149
|
-
and one hour.
|
|
131
|
+
and one hour.
|
|
@@ -27,16 +27,15 @@ module ActiveJob
|
|
|
27
27
|
private
|
|
28
28
|
|
|
29
29
|
def push(job, at: nil)
|
|
30
|
-
|
|
31
|
-
"
|
|
30
|
+
envelope = {
|
|
31
|
+
"task" => SolidJobs::ActiveJob::Wrapper,
|
|
32
32
|
"wrapped" => job.class.name,
|
|
33
|
-
"
|
|
34
|
-
"
|
|
33
|
+
"channel" => job.queue_name,
|
|
34
|
+
"arguments" => [job.serialize],
|
|
35
35
|
}
|
|
36
|
-
|
|
37
|
-
SolidJobs::
|
|
36
|
+
envelope["run_at"] = at if at
|
|
37
|
+
SolidJobs::Publisher.publish(envelope)
|
|
38
38
|
end
|
|
39
39
|
end
|
|
40
40
|
end
|
|
41
41
|
end
|
|
42
|
-
|
|
@@ -3,14 +3,11 @@
|
|
|
3
3
|
module SolidJobs
|
|
4
4
|
module ActiveJob
|
|
5
5
|
class Wrapper
|
|
6
|
-
include SolidJobs::
|
|
7
|
-
solid_jobs_options retry: true
|
|
6
|
+
include SolidJobs::Task
|
|
8
7
|
|
|
9
8
|
def perform(job_data)
|
|
10
|
-
::ActiveJob::Base.execute(job_data.merge("provider_job_id" =>
|
|
9
|
+
::ActiveJob::Base.execute(job_data.merge("provider_job_id" => task_id))
|
|
11
10
|
end
|
|
12
11
|
end
|
|
13
|
-
JobWrapper = Wrapper
|
|
14
12
|
end
|
|
15
13
|
end
|
|
16
|
-
|