solid-jobs 0.1.2 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 1a9ddd8000ed07b176d4e291ba7e6b5c1b6a3d909efe282ea6ace542d676e142
4
- data.tar.gz: 242f4734ef254188dee64f1ce73df43cbb83f6d39840d7e8f59c6017a499a5f3
3
+ metadata.gz: 8c178515fb3c56d80f0380da9f25114cd44679135b4d3d39fb468c45760748a7
4
+ data.tar.gz: 94080a99ea85f351714722d4f4f073ab99f01b2e5096a8e304670138171cf1ab
5
5
  SHA512:
6
- metadata.gz: 8d2bb684bb107820e189f12d4ac709ce8ad43e546bdf2da5933fcd781627655a8897030ed53010861a74619d65afb33fd89466a1c12993d8ab51e4de85f9c2f8
7
- data.tar.gz: 22fb65c94e80dc84dc94f5e6ad3fa9a5047378f45fe29263c5a98817c6a0aeda71f9d54f597a69b0407332e8052f1a5d5da097b97fbd22be912543c56015f78d
6
+ metadata.gz: fc05837b46c971dfd517cedac8bd46960867160cce776384eced29d41fac3a787b1f45a1a8f2d21ebeddbdb828766087a01396f245a255e5e77443bfa5366a2f
7
+ data.tar.gz: af56380ec6a6e82470d4df8c477b0e87eec5dc9aaddb9d04fcad32f6d630f695254c77008b2c4cb4400b496a2a5a566a4b1e58fec2b93d536ad3447868146099
data/CHANGELOG.md CHANGED
@@ -2,7 +2,40 @@
2
2
 
3
3
  All notable changes to this project will be documented in this file.
4
4
 
5
- ## [Unreleased]
5
+ ## [0.2.0] - 2026-10-03
6
+
7
+ ### Breaking
8
+
9
+ - Replace the previous job API with the independent `SolidJobs::Task` API:
10
+ `enqueue`, `enqueue_after`, `enqueue_at`, `enqueue_many`, `execute`, and
11
+ `with_options`.
12
+ - Replace the previous payload with a SolidJobs envelope using `id`, `task`,
13
+ `arguments`, `channel`, `run_at`, `created_ms`, and `queued_ms`.
14
+ - Move all Redis data into the `solid_jobs:` keyspace. Existing queued data is
15
+ not read or migrated automatically.
16
+ - Replace middleware chains with `publish_interceptors` and
17
+ `execute_interceptors`, whose interceptors implement `around(context)`.
18
+ - Replace the administrative API with `Metrics`, `Channel`, `StoredTask`,
19
+ `PlannedTasks`, `RetryingTasks`, `DiscardedTasks`, `Node`, `Nodes`,
20
+ `Execution`, and `Claims`.
21
+ - Replace queue configuration with channels and the `:weighted`, `:priority`,
22
+ and `:shuffle` ordering modes.
23
+ - Remove all Sidekiq compatibility aliases and wire-format compatibility.
24
+
25
+ ### Changed
26
+
27
+ - Claims use UUID task IDs and claim tokens, with generation-fenced completion
28
+ and requeue operations.
29
+ - Failure handling uses `max_failures`, `retry_within`, `retry_delay`, and
30
+ `after_final_failure`.
31
+ - Testing modes are now `:capture` and `:execute`.
32
+
33
+ ## [0.1.3] - 2026-10-01
34
+
35
+ ### Changed
36
+
37
+ - Require `solid-redis` 1.0.11 and rely on its `solid-resp-ractor` dependency
38
+ instead of declaring and loading the RESP codec directly.
6
39
 
7
40
  ## [0.1.2] - 2026-10-01
8
41
 
@@ -56,5 +89,8 @@ All notable changes to this project will be documented in this file.
56
89
  - Reliable-hot-path and CPU-scaling profilers plus the independent Sidekiq
57
90
  comparison benchmark.
58
91
 
59
- [Unreleased]: https://github.com/nicolasva/solid-jobs/compare/v0.1.0...HEAD
92
+ [0.2.0]: https://github.com/nicolasva/solid-jobs/compare/v0.1.3...v0.2.0
93
+ [0.1.3]: https://github.com/nicolasva/solid-jobs/compare/v0.1.2...v0.1.3
94
+ [0.1.2]: https://github.com/nicolasva/solid-jobs/compare/v0.1.1...v0.1.2
95
+ [0.1.1]: https://github.com/nicolasva/solid-jobs/compare/v0.1.0...v0.1.1
60
96
  [0.1.0]: https://github.com/nicolasva/solid-jobs/releases/tag/v0.1.0
data/README.md CHANGED
@@ -1,163 +1,183 @@
1
1
  # SolidJobs
2
2
 
3
3
  [![Build Status](https://github.com/nicolasva/solid-jobs/actions/workflows/ci.yml/badge.svg)](https://github.com/nicolasva/solid-jobs/actions/workflows/ci.yml)
4
- [![Code Climate](https://codeclimate.com/github/nicolasva/solid-jobs/badges/gpa.svg)](https://codeclimate.com/github/nicolasva/solid-jobs)
5
4
  [![Gem Version](https://badge.fury.io/rb/solid-jobs.svg)](https://rubygems.org/gems/solid-jobs)
6
- [![Documentation Status](https://img.shields.io/badge/docs-rubydoc.info-blue.svg)](https://www.rubydoc.info/gems/solid-jobs)
7
- [![Downloads](https://img.shields.io/gem/dt/solid-jobs.svg)](https://rubygems.org/gems/solid-jobs)
8
5
 
9
- SolidJobs is a Ractor-oriented Redis background job system for Ruby. It uses
10
- `solid-redis` for Redis access and keeps mutable clients, pools, middleware,
11
- and runtime state local to their owning Ractor.
6
+ SolidJobs is a Ractor-oriented Redis task runner for Ruby. It provides its own
7
+ task API, Redis envelope, keyspace, interceptor model, failure policy, and
8
+ inspection API.
12
9
 
13
- SolidJobs is an independent implementation. It does not depend on or load the
14
- Sidekiq gem. Its Redis job payloads and queue keys are designed to be
15
- compatible with the Sidekiq 8 open-source data format.
16
-
17
- SolidJobs and its required gems use pure Ruby and do not require native
18
- extensions. Hot paths are designed around bounded buffers, reusable immutable
19
- configuration, and low-allocation command batches.
10
+ SolidJobs is **not compatible with Sidekiq**. It does not read Sidekiq queues,
11
+ does not write Sidekiq payloads, and does not expose Sidekiq's job API. Existing
12
+ applications and queued work must be migrated explicitly.
20
13
 
21
14
  ## Installation
22
15
 
23
- Add SolidJobs 0.1 to your bundle:
24
-
25
16
  ```ruby
26
- gem "solid-jobs", "~> 0.1.0"
17
+ gem "solid-jobs"
27
18
  ```
28
19
 
29
- Then run:
30
-
31
- ```sh
32
- bundle install
33
- ```
34
-
35
- ## Delivery semantics
20
+ Then run `bundle install`.
36
21
 
37
- SolidJobs provides **at-least-once** job delivery. A worker atomically moves a
38
- job from `queue:<name>` to a process reservation list before execution and
39
- removes it only after a successful ACK. Graceful shutdown requeues unfinished
40
- jobs; reservations owned by a crashed process are recovered into their
41
- original queues.
42
-
43
- A process can still crash after the application side effect and before the
44
- ACK. The recovered job will then run again. Jobs must therefore be idempotent
45
- or implement an application-level deduplication key when duplicate side
46
- effects are unsafe. SolidJobs does not claim exactly-once execution.
22
+ ## Defining and submitting tasks
47
23
 
48
24
  ```ruby
49
- class HardJob
50
- include SolidJobs::Job
25
+ class RecalculateAccount
26
+ include SolidJobs::Task
51
27
 
52
- solid_jobs_options queue: "critical", retry: 10
28
+ task_options channel: "critical", max_failures: 10
53
29
 
54
30
  def perform(account_id)
55
31
  Account.find(account_id).recalculate!
56
32
  end
57
33
  end
58
34
 
59
- HardJob.perform_async(42)
60
- HardJob.perform_in(30, 42)
35
+ RecalculateAccount.enqueue(42)
36
+ RecalculateAccount.enqueue_after(30, 42)
37
+ RecalculateAccount.enqueue_at(Time.now + 300, 42)
38
+ RecalculateAccount.enqueue_many([[42], [43], [44]])
39
+ ```
40
+
41
+ The persisted envelope is specific to SolidJobs:
42
+
43
+ ```json
44
+ {
45
+ "id": "7ec33ed5-1b77-4bde-8960-bd77231f10ef",
46
+ "task": "RecalculateAccount",
47
+ "arguments": [42],
48
+ "channel": "critical",
49
+ "max_failures": 10,
50
+ "created_ms": 1790981000000,
51
+ "queued_ms": 1790981000001
52
+ }
61
53
  ```
62
54
 
63
- Run workers:
55
+ All internal Redis keys use the `solid_jobs:` namespace. Ready work is stored
56
+ under `solid_jobs:channel:<name>`; planned, retrying, discarded, node, claim,
57
+ attempt, and metric data use separate namespaced keys.
58
+
59
+ Run executors:
64
60
 
65
61
  ```sh
66
62
  bundle exec solid-jobs --require ./config/environment \
67
- --concurrency 8 --queue critical,3 --queue default
63
+ --concurrency 8 --channel critical,3 --channel default
68
64
  ```
69
65
 
70
- ## Tests
71
-
72
- SolidJobs uses Minitest exclusively:
66
+ ## Configuration
73
67
 
74
- ```sh
75
- # Unit and bounded Redis integration tests
76
- bundle exec rake test
68
+ ```ruby
69
+ SolidJobs.configure do |config|
70
+ config.channels = [["critical", 3], "default"]
71
+ config.channel_order = :weighted
72
+ config.concurrency = 8
73
+ config.shutdown_timeout = 25
74
+ config.default_task_options = {
75
+ channel: "default",
76
+ max_failures: 25,
77
+ }
78
+ end
79
+ ```
77
80
 
78
- # Bounded concurrency and multi-process recovery stress
79
- STRESS_JOBS=10000 bundle exec rake stress
81
+ `channel_order` accepts `:weighted` for bounded weighted round-robin,
82
+ `:priority` for declaration-order priority, and `:shuffle` for random
83
+ selection from the weighted channel list.
80
84
 
81
- # Reproducible random-fault campaign
82
- SOLID_JOBS_TORTURE=1 STRESS_JOBS=100000 bundle exec rake torture
85
+ ## Interceptors
83
86
 
84
- # Re-run an exact failure sequence
85
- SOLID_JOBS_TORTURE=1 STRESS_JOBS=100000 STRESS_SEED=123456 bundle exec rake torture
87
+ Publication and execution use SolidJobs interceptors. An interceptor implements
88
+ `around(context)` and yields to continue:
86
89
 
87
- # Repeated fresh-process RESP/Ractor startup torture
88
- STARTUP_TORTURE_CYCLES=1000 \
89
- STARTUP_TORTURE_READERS=100 \
90
- bundle exec rake startup_torture
90
+ ```ruby
91
+ class TraceExecution
92
+ def around(context)
93
+ Telemetry.start(context.envelope.fetch("id"))
94
+ yield
95
+ ensure
96
+ Telemetry.finish
97
+ end
98
+ end
91
99
 
92
- # Long-running stability; defaults to 24 hours
93
- SOLID_JOBS_SOAK=1 SOLID_JOBS_SOAK_SECONDS=86400 bundle exec rake soak
100
+ SolidJobs.configure do |config|
101
+ config.execute_interceptors.use(TraceExecution)
102
+ end
94
103
  ```
95
104
 
96
- The torture report reconciles enqueued, uniquely completed, duplicate,
97
- dead, queued, and reserved jobs. Any non-zero `LOST` value fails the test.
105
+ Publication interceptors receive `SolidJobs::Publisher::Publication`; execution
106
+ interceptors receive `SolidJobs::Executor::Execution`.
98
107
 
99
- Profile the reliable execution path independently from application work:
108
+ ## Failure handling
100
109
 
101
- ```sh
102
- REDIS_URL=redis://127.0.0.1:6379/0 \
103
- HOT_PATH_JOBS=10000 \
104
- bundle exec rake benchmark:hot_path
110
+ Tasks default to 25 failures. A task can customize the delay or final action:
111
+
112
+ ```ruby
113
+ class ImportCatalog
114
+ include SolidJobs::Task
115
+
116
+ task_options max_failures: 8, retry_within: 3_600
117
+
118
+ retry_delay do |failure_count, error, envelope|
119
+ :drop if error.is_a?(InvalidCatalog)
120
+ end
121
+
122
+ after_final_failure do |envelope, error|
123
+ Alerts.catalog_import_failed(envelope.fetch("id"), error)
124
+ end
125
+ end
105
126
  ```
106
127
 
107
- The report separates time and allocations for reservation, payload reuse,
108
- observability registration, dispatch/perform wrapping, observability cleanup,
109
- and fenced ACK. Fetch decodes the payload once for both reservation metadata
110
- and dispatch; its job body is intentionally empty.
128
+ `retry_delay` may return a delay in seconds, `:drop`, or `:archive`. Without an
129
+ override, SolidJobs uses capped exponential backoff with equal jitter.
130
+
131
+ ## Delivery semantics
132
+
133
+ SolidJobs provides **at-least-once** delivery. An executor atomically claims a
134
+ task before execution and completes the claim only after `perform` returns.
135
+ Graceful shutdown requeues unfinished tasks. Claims owned by a crashed node are
136
+ recovered into their original channels.
137
+
138
+ A process can still crash after an application side effect and before claim
139
+ completion. The recovered task then runs again. Tasks must be idempotent or use
140
+ application-level deduplication when duplicate effects are unsafe.
141
+
142
+ See [docs/reliability.md](docs/reliability.md) for the state machine, claim
143
+ fencing, recovery rules, and Ruby Ractor caveats.
144
+
145
+ ## Testing
111
146
 
112
- Diagnose CPU scaling independently from Redis:
147
+ ```ruby
148
+ SolidJobs.testing!(:capture) do
149
+ RecalculateAccount.enqueue(42)
150
+ RecalculateAccount.captured
151
+ end
152
+
153
+ SolidJobs.testing!(:execute) do
154
+ RecalculateAccount.enqueue(42)
155
+ end
156
+ ```
157
+
158
+ Run the project suites:
113
159
 
114
160
  ```sh
115
- CPU_SCALING_JOBS=1000 \
116
- CPU_SCALING_ITERATIONS=210000 \
117
- bundle exec rake benchmark:cpu_scaling
161
+ bundle exec rake test
162
+ STRESS_JOBS=10000 bundle exec rake stress
163
+ SOLID_JOBS_TORTURE=1 STRESS_JOBS=100000 bundle exec rake torture
164
+ SOLID_JOBS_SOAK=1 SOLID_JOBS_SOAK_SECONDS=86400 bundle exec rake soak
118
165
  ```
119
166
 
120
- This runs the identical CPU loop through pure Ractors and through the
121
- SolidJobs in-memory dispatch path. Each 1/2/4/8 case uses fresh processes and
122
- reports execution latency, scaling efficiency, CPU-seconds per 1,000 jobs,
123
- RSS, allocations, GC time, heap slots, and malloc growth.
124
-
125
- ## Sidekiq comparison
126
-
127
- The separate `benchmark_sidekiq_solid-jobs` bundle runs Sidekiq and SolidJobs
128
- against the same isolated Redis server. It covers enqueue, bulk enqueue,
129
- CPU-bound processing, I/O-bound processing, mixed processing, and long-running
130
- stability. Every measurement runs in a fresh Ruby process; client order
131
- alternates and the default report uses the median of six repetitions.
132
-
133
- The current homogeneous CPU-processing reference uses Ruby 4.0.1,
134
- Sidekiq 8.1.7, a 210,000-iteration integer workload, and 1,000 jobs per case:
135
-
136
- | Concurrency | Sidekiq jobs/s | SolidJobs jobs/s | SolidJobs scaling | Sidekiq CPU-s/1k | SolidJobs CPU-s/1k | Sidekiq RSS | SolidJobs RSS | Sidekiq alloc/job | SolidJobs alloc/job |
137
- |---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
138
- | 1 | 142 | 138 | 100.0% | 6.95 | 7.11 | 41.5 MiB | 72.0 MiB | 129.8 | 89.9 |
139
- | 2 | 144 | 272 | 98.7% | 6.94 | 7.16 | 41.8 MiB | 77.6 MiB | 117.5 | 81.3 |
140
- | 4 | 143 | 539 | 97.8% | 6.99 | 7.18 | 41.9 MiB | 71.4 MiB | 111.9 | 77.2 |
141
- | 8 | 144 | 957 | 86.7% | 6.94 | 8.01 | 42.2 MiB | 69.5 MiB | 108.7 | 75.4 |
142
-
143
- At eight concurrency units, SolidJobs reaches 957 jobs/s versus 144 jobs/s
144
- for one Sidekiq process. This is Ractor parallelism rather than equal CPU
145
- efficiency: SolidJobs consumes 766% CPU and 8.01 CPU-seconds per 1,000 jobs,
146
- while Sidekiq consumes 100% CPU and 6.94 CPU-seconds per 1,000 jobs. SolidJobs
147
- also uses more RSS, but reaches 13.77 jobs/s/MiB versus 3.41 for Sidekiq.
148
-
149
- These are local synthetic measurements, not application-capacity claims.
150
- Queue p95/p99 values in this run use sparse sampling and are excluded from the
151
- summary until the final latency campaign increases the sample count. Ruby
152
- 3.4.4 eight-Ractor results are also excluded: Ruby 3.4's Ractor scheduler
153
- crashed (concurrent TCP/RESP initialization, reproduced on macOS arm64 and
154
- Linux x86_64) or deadlocked VM-wide on a GC barrier under cross-Ractor
155
- message traffic (`test/support/ractor_barrier_repro.rb` reproduces it without
156
- SolidJobs).
157
- Ruby 4.0.1 passed the equivalent reproducers, and `StartupBarrier` serializes
158
- component initialization before releasing normal parallel processing.
159
- **Multi-Ractor servers are recommended on Ruby ≥ 4.0**; see
160
- `docs/reliability.md`.
161
-
162
- The project is under active development. The Web UI and commercial Sidekiq
163
- features are not part of the initial scope.
167
+ ## Migrating from Sidekiq
168
+
169
+ There is no transparent migration path because compatibility is intentionally
170
+ absent:
171
+
172
+ 1. Replace `include Sidekiq::Job` with `include SolidJobs::Task`.
173
+ 2. Replace `sidekiq_options` with `task_options`.
174
+ 3. Replace `perform_async`, `perform_in`, and `perform_bulk` with `enqueue`,
175
+ `enqueue_after`, and `enqueue_many`.
176
+ 4. Replace middleware with SolidJobs interceptors.
177
+ 5. Drain or export existing Sidekiq queues before switching. SolidJobs will not
178
+ consume them.
179
+ 6. Start SolidJobs with `--channel`; `--queue` is not accepted.
180
+
181
+ The Active Job adapter remains available as
182
+ `ActiveJob::QueueAdapters::SolidJobsAdapter`, but it writes only SolidJobs
183
+ envelopes and keys.
data/docs/reliability.md CHANGED
@@ -1,63 +1,61 @@
1
1
  # Reliability model
2
2
 
3
- SolidJobs uses a Redis-backed at-least-once state machine:
3
+ SolidJobs uses a namespaced Redis-backed at-least-once state machine:
4
4
 
5
5
  ```text
6
6
  READY
7
7
  |
8
- | BLMOVE reservation
8
+ | atomic claim
9
9
  v
10
- IN_PROGRESS ---- ACK/LREM ----> removed
10
+ CLAIMED ---- fenced completion ----> removed
11
11
  |
12
- +---- application error ----> RETRY or DEAD, then ACK
12
+ +---- application error ----> RETRYING or DISCARDED, then completion
13
13
  +---- graceful timeout -----> READY
14
- +---- process crash --------> reservation retained
14
+ +---- node crash -----------> claim retained
15
15
  |
16
16
  +---- recovery ----> READY
17
17
  ```
18
18
 
19
- The reservation list is named `<process-identity>:reserved:<processor-id>`.
20
- Each execution receives a distinct reservation journal entry:
19
+ Ready tasks live in `solid_jobs:channel:<name>`. Each executor owns at most one
20
+ claimed-task list and one claim journal entry:
21
21
 
22
22
  ```text
23
- job_id stable across replay
24
- reservation_id unique per execution attempt
25
- process_id owning process identity
26
- worker_id owning Processor Ractor
27
- attempt monotonically increasing execution count
28
- reserved_at wall-clock reservation time
23
+ task_id stable across replay
24
+ claim_token unique per execution attempt
25
+ node_id owning server identity
26
+ executor_id owning Executor Ractor
27
+ channel destination used by recovery
28
+ attempt monotonically increasing execution count
29
+ claimed_at wall-clock claim time
29
30
  ```
30
31
 
31
- ACK and requeue are fenced by `reservation_id`. Their Lua scripts first
32
- verify that the worker still owns the current journal generation. A delayed
33
- or revived worker cannot remove or requeue a newer reservation, even when a
34
- supervisor reuses the same processor slot.
32
+ Completion and requeue are fenced by `claim_token`. Their Lua scripts verify
33
+ that the executor still owns the current journal generation before changing
34
+ state. A delayed or revived executor cannot complete or requeue a newer claim.
35
35
 
36
- The payload retains its canonical `queue` field, allowing another process to
37
- restore it after the owner is no longer alive. Recovery uses process liveness
38
- and heartbeat, never reservation age alone, so a legitimate long-running job
39
- is not stolen while its owner remains alive.
36
+ The envelope retains its canonical `channel`, allowing another node to restore
37
+ it after its owner dies. Recovery uses node liveness and heartbeat, never claim
38
+ age alone, so a legitimate long-running task is not stolen.
40
39
 
41
40
  ## Failure boundaries
42
41
 
43
- - Before reservation: the job remains in `queue:<name>`.
44
- - After reservation or during `perform`: the job remains reserved.
45
- - After the application effect but before ACK: recovery replays the job.
46
- - Redis unavailable during ACK: the job remains reserved and is replayed.
47
- - Graceful shutdown: the active job may finish within the configured timeout;
42
+ - Before claim: the task remains in `solid_jobs:channel:<name>`.
43
+ - After claim or during `perform`: the task remains claimed.
44
+ - After the application effect but before completion: recovery replays it.
45
+ - Redis unavailable during completion: the task remains claimed and is replayed.
46
+ - Graceful shutdown: the active task may finish within the configured timeout;
48
47
  otherwise it is interrupted and requeued.
49
- - `SIGKILL`: no handler runs; recovery relies exclusively on Redis state.
48
+ - `SIGKILL`: recovery relies exclusively on Redis state.
50
49
 
51
- This design intentionally favors no job loss over duplicate suppression.
52
- Exactly-once side effects require application-level idempotency.
50
+ This design favors no task loss over duplicate suppression. Exactly-once side
51
+ effects require application-level idempotency.
53
52
 
54
53
  ## Startup isolation
55
54
 
56
- The server does not reserve work while components are booting. Heartbeat,
57
- Processor, and Scheduler Ractors initialize their local configuration and
58
- Redis pool, report `READY` exactly once, and wait behind
59
- `SolidJobs::StartupBarrier`. Processing starts only after every component is
60
- ready:
55
+ The server does not claim work while components are booting. Heartbeat,
56
+ Processor, and Timer Ractors initialize their local configuration and Redis
57
+ pool, report `READY` exactly once, and wait behind
58
+ `SolidJobs::StartupBarrier`:
61
59
 
62
60
  ```text
63
61
  BOOTING -> ALL_READY -> RUNNING
@@ -65,85 +63,69 @@ BOOTING -> ALL_READY -> RUNNING
65
63
  ```
66
64
 
67
65
  Boot failure is terminal. Already-ready components receive `:abort`, close
68
- their local resources, and never enter their fetch loops. Cleanup is
69
- idempotent. Component startup is serialized, avoiding concurrent TCP/RESP
70
- initialization paths known to crash Ruby 3.4 (reproduced on 3.4.4 macOS arm64
71
- and 3.4.11 Linux x86_64; `rake startup_torture` is the reproducer); normal
72
- processing remains parallel after `RUNNING`.
73
-
74
- ### Ruby 3.4 Ractor caveat
75
-
76
- Ruby 3.4's Ractor scheduler can deadlock the whole VM (main thread included)
77
- when a GC-triggered `rb_ractor_sched_barrier_start` runs while several
78
- Ractors exchange `move: true` messages: every thread parks in
79
- `ractor_sched_barrier_join_wait_locked` and the barrier never completes.
80
- `test/support/ractor_barrier_repro.rb` reproduces it **without SolidJobs or
81
- Redis** (one receiver, four senders, 24k moved messages per iteration):
82
- Ruby 3.4.4 freezes within the first iterations, Ruby 4.0.1 completes 30/30.
83
- Inside SolidJobs the same traffic pattern is the Processor → Heartbeat
84
- `:work`/`:done`/`:stats` channel, so any multi-Processor server on Ruby 3.4
85
- is exposed; once frozen, neither `Timeout` nor process exit
86
- (`rb_ractor_terminate_all`) can recover.
87
-
88
- Recommendation: **run multi-Ractor SolidJobs servers on Ruby ≥ 4.0**. On
89
- Ruby 3.4 use the client/API side freely, and prefer one process per
90
- Processor (`concurrency: 1`) for the server. Server-based stress tests are
91
- skipped on Ruby < 4 for this reason, and `Server#start` logs a warning when
92
- it detects `concurrency > 1` on Ruby < 4.
93
-
94
- ### Bounded shutdown
66
+ their local resources, and never enter their claim loops. Cleanup is
67
+ idempotent.
68
+
69
+ ## Ruby 3.4 Ractor caveat
70
+
71
+ Ruby 3.4's Ractor scheduler can deadlock the VM when a GC-triggered scheduler
72
+ barrier runs while several Ractors exchange moved messages.
73
+ `test/support/ractor_barrier_repro.rb` reproduces this without SolidJobs or
74
+ Redis. Ruby 4.0.1 completes the equivalent reproducer.
75
+
76
+ Run multi-Ractor SolidJobs servers on Ruby >= 4.0. On Ruby 3.4, prefer one
77
+ process per Processor (`concurrency: 1`). Server stress tests are skipped on
78
+ Ruby versions affected by this runtime issue, and `Server#start` warns when it
79
+ detects a multi-Ractor configuration there.
80
+
81
+ ## Bounded shutdown
95
82
 
96
83
  `Server#stop` never waits forever for a component. Each Ractor gets
97
- `shutdown_timeout + Server::STOP_GRACE` to return after `:stop`; the signal
98
- is re-sent up to `STOP_RESENDS` times, then the component is abandoned with
99
- an error log so the process can proceed with shutdown.
84
+ `shutdown_timeout + Server::STOP_GRACE` to return after `:stop`; the signal is
85
+ re-sent up to `STOP_RESENDS` times before the component is abandoned with an
86
+ error log.
100
87
 
101
88
  ## Configuration scope
102
89
 
103
90
  `SolidJobs.config` and the testing mode are Ractor-local, not thread-local.
104
- Every thread and fiber inside a Ractor (for example Puma workers or Rails
105
- request threads) shares the configuration set on that Ractor, while each
106
- Ractor keeps its own isolated configuration.
91
+ Every thread and fiber inside a Ractor shares its configuration, while each
92
+ Ractor owns isolated mutable state and Redis connections.
107
93
 
108
94
  ## Integrity auditing
109
95
 
110
- Fault and chaos tests can reconcile a known set of job IDs against every
111
- Redis-backed state:
96
+ Fault tests can reconcile known task IDs against every Redis-backed state:
112
97
 
113
98
  ```ruby
114
99
  report = SolidJobs::IntegrityCheck.call(
115
- expected_job_ids: submitted_job_ids,
100
+ expected_job_ids: submitted_task_ids,
116
101
  acked_key: "test:completed",
117
102
  )
118
103
 
119
104
  raise report.inspect unless report.ok?
120
105
  ```
121
106
 
122
- The report separates lost jobs, unexpected/orphaned jobs, inconsistent
123
- reservation journals, dangling attempt indexes, duplicate active states, and
124
- malformed payloads. ACK removes the completed job's attempt index atomically;
125
- requeue retains it so a recovered reservation increments the same attempt
126
- sequence.
107
+ The report separates lost tasks, unexpected tasks, inconsistent claim
108
+ journals, dangling attempt indexes, duplicate active states, and malformed
109
+ envelopes. Completion removes the task's attempt index atomically; requeue
110
+ retains it so recovery increments the same attempt sequence.
127
111
 
128
- ## Backpressure and queues
112
+ ## Backpressure and channels
129
113
 
130
- Each Processor owns at most one reservation and reserves only immediately
131
- before execution. A process with concurrency `N` therefore holds at most `N`
132
- active reservations, regardless of Redis queue depth.
114
+ Each Processor owns at most one claim and claims only immediately before
115
+ execution. A node with concurrency `N` therefore holds at most `N` active
116
+ claims, regardless of channel depth.
133
117
 
134
- Queue mode defaults to `:weighted`:
118
+ Channel order defaults to `:weighted`:
135
119
 
136
- - `:weighted` uses bounded weighted round-robin and does not starve configured
137
- queues;
138
- - `:strict` always checks queues in declaration order and may intentionally
139
- starve lower-priority queues while a higher-priority queue remains busy;
140
- - `:random` samples the configured weighted queue list on each reservation.
120
+ - `:weighted` uses bounded weighted round-robin;
121
+ - `:priority` checks channels in declaration order;
122
+ - `:shuffle` samples the configured weighted channel list.
141
123
 
142
- Paused queues are excluded before reservation.
124
+ Paused channels are excluded before claiming.
143
125
 
144
126
  ## Retry storms
145
127
 
146
- Retries use capped exponential backoff with equal jitter. For attempt `n`, the
147
- ceiling is `min(retry_base_delay * 2**n, retry_max_delay)` and the actual delay
128
+ Retries use capped exponential backoff with equal jitter. For failure `n`, the
129
+ ceiling is `min(retry_base_delay * 2**(n - 1), retry_max_delay)` and the delay
148
130
  is distributed between half and all of that ceiling. Defaults are 15 seconds
149
- and one hour. This spreads recovery traffic after a shared external outage.
131
+ and one hour.
@@ -27,16 +27,15 @@ module ActiveJob
27
27
  private
28
28
 
29
29
  def push(job, at: nil)
30
- payload = {
31
- "class" => SolidJobs::ActiveJob::Wrapper,
30
+ envelope = {
31
+ "task" => SolidJobs::ActiveJob::Wrapper,
32
32
  "wrapped" => job.class.name,
33
- "queue" => job.queue_name,
34
- "args" => [job.serialize],
33
+ "channel" => job.queue_name,
34
+ "arguments" => [job.serialize],
35
35
  }
36
- payload["at"] = at if at
37
- SolidJobs::Client.push(payload)
36
+ envelope["run_at"] = at if at
37
+ SolidJobs::Publisher.publish(envelope)
38
38
  end
39
39
  end
40
40
  end
41
41
  end
42
-
@@ -3,14 +3,11 @@
3
3
  module SolidJobs
4
4
  module ActiveJob
5
5
  class Wrapper
6
- include SolidJobs::Job
7
- solid_jobs_options retry: true
6
+ include SolidJobs::Task
8
7
 
9
8
  def perform(job_data)
10
- ::ActiveJob::Base.execute(job_data.merge("provider_job_id" => jid))
9
+ ::ActiveJob::Base.execute(job_data.merge("provider_job_id" => task_id))
11
10
  end
12
11
  end
13
- JobWrapper = Wrapper
14
12
  end
15
13
  end
16
-