solid-jobs 0.1.1 → 0.1.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: ce7c73e437bf4a78efbf9422e4a5cb7cf05635b5f3c9cbca4cd44a9baa0ee2d9
4
- data.tar.gz: f0b4513d2c4d75bc9f871ba2c64c7848789747cf797e32044d56eefdab961647
3
+ metadata.gz: 1a9ddd8000ed07b176d4e291ba7e6b5c1b6a3d909efe282ea6ace542d676e142
4
+ data.tar.gz: 242f4734ef254188dee64f1ce73df43cbb83f6d39840d7e8f59c6017a499a5f3
5
5
  SHA512:
6
- metadata.gz: e6d245f323c4cdf117f2bb53584eed2e2c697a72fdc170792954d25feccb15b32d2cae7f3d8a337e7b29c3fedb071ab4ad9b773ac414fe893e389754532ca5af
7
- data.tar.gz: 9abce08d1c4fb800a5f39e56fcebc266a7cc1dd52acd891fdaccf9fd1f2840ead4a3c4ce0a5649acc6c27a8f26ad5554d0599a5eb55f8e742afc568a307ddc09
6
+ metadata.gz: 8d2bb684bb107820e189f12d4ac709ce8ad43e546bdf2da5933fcd781627655a8897030ed53010861a74619d65afb33fd89466a1c12993d8ab51e4de85f9c2f8
7
+ data.tar.gz: 22fb65c94e80dc84dc94f5e6ad3fa9a5047378f45fe29263c5a98817c6a0aeda71f9d54f597a69b0407332e8052f1a5d5da097b97fbd22be912543c56015f78d
data/CHANGELOG.md CHANGED
@@ -4,6 +4,25 @@ All notable changes to this project will be documented in this file.
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.1.2] - 2026-10-01
8
+
9
+ ### Added
10
+
11
+ - `Server#start` logs a warning on Ruby < 4 when `concurrency > 1`, pointing
12
+ to the Ruby 3.4 Ractor GC-barrier deadlock and the recommended setups.
13
+ - `Server#stop` is bounded: components get `shutdown_timeout + STOP_GRACE`
14
+ to return, `:stop` is re-sent up to `STOP_RESENDS` times, then the
15
+ component is abandoned with an error log instead of hanging shutdown.
16
+ - `test/support/ractor_barrier_repro.rb`: standalone reproducer of the Ruby
17
+ 3.4 Ractor GC-barrier deadlock (no SolidJobs, no Redis).
18
+
19
+ ### Changed
20
+
21
+ - Server-based stress tests are skipped on Ruby < 4: Ruby 3.4's Ractor
22
+ scheduler deadlocks the VM on a GC barrier under cross-Ractor `move:`
23
+ traffic. Multi-Ractor servers are recommended on Ruby ≥ 4.0 (see
24
+ `docs/reliability.md`). CI jobs now time out after 20 minutes.
25
+
7
26
  ## [0.1.1] - 2026-10-01
8
27
 
9
28
  ### Fixed
data/README.md CHANGED
@@ -149,11 +149,15 @@ also uses more RSS, but reaches 13.77 jobs/s/MiB versus 3.41 for Sidekiq.
149
149
  These are local synthetic measurements, not application-capacity claims.
150
150
  Queue p95/p99 values in this run use sparse sampling and are excluded from the
151
151
  summary until the final latency campaign increases the sample count. Ruby
152
- 3.4.4 eight-Ractor results are also excluded: concurrent TCP/RESP
153
- initialization triggered a reproducible native crash on the tested Apple
154
- Silicon environment. Ruby 4.0.1 passed the equivalent reproducer 100/100
155
- times, and `StartupBarrier` serializes component initialization before
156
- releasing normal parallel processing.
152
+ 3.4.4 eight-Ractor results are also excluded: Ruby 3.4's Ractor scheduler
153
+ crashed (concurrent TCP/RESP initialization, reproduced on macOS arm64 and
154
+ Linux x86_64) or deadlocked VM-wide on a GC barrier under cross-Ractor
155
+ message traffic (`test/support/ractor_barrier_repro.rb` reproduces it without
156
+ SolidJobs).
157
+ Ruby 4.0.1 passed the equivalent reproducers, and `StartupBarrier` serializes
158
+ component initialization before releasing normal parallel processing.
159
+ **Multi-Ractor servers are recommended on Ruby ≥ 4.0**; see
160
+ `docs/reliability.md`.
157
161
 
158
162
  The project is under active development. The Web UI and commercial Sidekiq
159
163
  features are not part of the initial scope.
data/docs/reliability.md CHANGED
@@ -71,6 +71,33 @@ initialization paths known to crash Ruby 3.4 (reproduced on 3.4.4 macOS arm64
71
71
  and 3.4.11 Linux x86_64; `rake startup_torture` is the reproducer); normal
72
72
  processing remains parallel after `RUNNING`.
73
73
 
74
+ ### Ruby 3.4 Ractor caveat
75
+
76
+ Ruby 3.4's Ractor scheduler can deadlock the whole VM (main thread included)
77
+ when a GC-triggered `rb_ractor_sched_barrier_start` runs while several
78
+ Ractors exchange `move: true` messages: every thread parks in
79
+ `ractor_sched_barrier_join_wait_locked` and the barrier never completes.
80
+ `test/support/ractor_barrier_repro.rb` reproduces it **without SolidJobs or
81
+ Redis** (one receiver, four senders, 24k moved messages per iteration):
82
+ Ruby 3.4.4 freezes within the first iterations, Ruby 4.0.1 completes 30/30.
83
+ Inside SolidJobs the same traffic pattern is the Processor → Heartbeat
84
+ `:work`/`:done`/`:stats` channel, so any multi-Processor server on Ruby 3.4
85
+ is exposed; once frozen, neither `Timeout` nor process exit
86
+ (`rb_ractor_terminate_all`) can recover.
87
+
88
+ Recommendation: **run multi-Ractor SolidJobs servers on Ruby ≥ 4.0**. On
89
+ Ruby 3.4 use the client/API side freely, and prefer one process per
90
+ Processor (`concurrency: 1`) for the server. Server-based stress tests are
91
+ skipped on Ruby < 4 for this reason, and `Server#start` logs a warning when
92
+ it detects `concurrency > 1` on Ruby < 4.
93
+
94
+ ### Bounded shutdown
95
+
96
+ `Server#stop` never waits forever for a component. Each Ractor gets
97
+ `shutdown_timeout + Server::STOP_GRACE` to return after `:stop`; the signal
98
+ is re-sent up to `STOP_RESENDS` times, then the component is abandoned with
99
+ an error log so the process can proceed with shutdown.
100
+
74
101
  ## Configuration scope
75
102
 
76
103
  `SolidJobs.config` and the testing mode are Ractor-local, not thread-local.
@@ -2,13 +2,20 @@
2
2
 
3
3
  require "securerandom"
4
4
  require "socket"
5
+ require "timeout"
5
6
 
6
7
  module SolidJobs
7
8
  class Server
9
+ # Extra time granted beyond shutdown_timeout before a component is
10
+ # considered stuck, and how many times :stop is re-sent before giving up.
11
+ STOP_GRACE = 5.0
12
+ STOP_RESENDS = 3
13
+
8
14
  attr_reader :config, :identity, :ractors
9
15
 
10
- def initialize(config: SolidJobs.config)
16
+ def initialize(config: SolidJobs.config, stop_grace: STOP_GRACE)
11
17
  @config = config
18
+ @stop_grace = stop_grace
12
19
  @ractors = []
13
20
  @scheduler = nil
14
21
  @heartbeat = nil
@@ -20,6 +27,7 @@ module SolidJobs
20
27
  def start
21
28
  raise Error, "Server is already running" if @started
22
29
 
30
+ warn_ruby34_multi_ractor
23
31
  snapshot = Utilities.shareable_copy(
24
32
  config.ractor_snapshot.merge(identity: identity, started_at: @started_at),
25
33
  )
@@ -197,10 +205,12 @@ module SolidJobs
197
205
 
198
206
  @ractors.each { |ractor| ractor.send(:stop) }
199
207
  @scheduler&.send(:stop)
200
- results = @ractors.map { |ractor| RactorSupport.value(ractor) }
201
- RactorSupport.value(@scheduler) if @scheduler
208
+ results = @ractors.each_with_index.map do |ractor, index|
209
+ await_termination(ractor, "processor #{index}")
210
+ end
211
+ await_termination(@scheduler, "scheduler") if @scheduler
202
212
  @heartbeat&.send(:stop)
203
- RactorSupport.value(@heartbeat) if @heartbeat
213
+ await_termination(@heartbeat, "heartbeat") if @heartbeat
204
214
  @ractors = []
205
215
  @scheduler = nil
206
216
  @heartbeat = nil
@@ -213,6 +223,47 @@ module SolidJobs
213
223
  @started
214
224
  end
215
225
 
226
+ private
227
+
228
+ RUBY34_MULTI_RACTOR_WARNING =
229
+ "SolidJobs: Ruby %s can deadlock the whole VM on a Ractor GC barrier " \
230
+ "under multi-Processor load (see docs/reliability.md). Run multi-Ractor " \
231
+ "servers on Ruby >= 4.0, or use concurrency: 1 (one process per Processor)."
232
+
233
+ def warn_ruby34_multi_ractor
234
+ return if RUBY_VERSION >= "4" || config.concurrency <= 1
235
+
236
+ config.logger.warn(format(RUBY34_MULTI_RACTOR_WARNING, RUBY_VERSION))
237
+ end
238
+
239
+ # Waits for a component Ractor to return after :stop. Ruby 3.4 can lose
240
+ # the wakeup of a Ractor.receive running in a secondary thread while the
241
+ # inbox is busy; a fresh :stop message re-triggers it. If the component
242
+ # still does not return, it is abandoned so shutdown never hangs forever.
243
+ def await_termination(ractor, name)
244
+ deadline = config.shutdown_timeout + @stop_grace
245
+ resends = 0
246
+ begin
247
+ Timeout.timeout(deadline) { RactorSupport.value(ractor) }
248
+ rescue Timeout::Error
249
+ if resends < STOP_RESENDS
250
+ resends += 1
251
+ config.logger.warn(
252
+ "SolidJobs #{name} did not stop within #{deadline}s, " \
253
+ "re-sending :stop (#{resends}/#{STOP_RESENDS})",
254
+ )
255
+ begin
256
+ ractor.send(:stop)
257
+ rescue Ractor::ClosedError
258
+ nil
259
+ end
260
+ retry
261
+ end
262
+ config.logger.error("SolidJobs #{name} did not stop; abandoning it")
263
+ nil
264
+ end
265
+ end
266
+
216
267
  def remote_signal
217
268
  config.redis_pool.call("RPOP", "#{identity}-signals")
218
269
  end
@@ -1,6 +1,6 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module SolidJobs
4
- VERSION = "0.1.1"
4
+ VERSION = "0.1.2"
5
5
  end
6
6
 
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: solid-jobs
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.1.1
4
+ version: 0.1.2
5
5
  platform: ruby
6
6
  authors:
7
7
  - Nicolas Vandenbogaerde