solid-jobs 0.1.3 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +73 -4
  3. data/README.md +149 -118
  4. data/docs/reliability.md +73 -91
  5. data/exe/solid-jobs +2 -3
  6. data/lib/active_job/queue_adapters/solid_jobs_adapter.rb +7 -8
  7. data/lib/solid_jobs/{config.rb → blueprint.rb} +51 -51
  8. data/lib/solid_jobs/catalog.rb +427 -0
  9. data/lib/solid_jobs/claim.rb +315 -0
  10. data/lib/solid_jobs/{server.rb → conductor.rb} +10 -10
  11. data/lib/solid_jobs/{cli.rb → console.rb} +12 -9
  12. data/lib/solid_jobs/{processor.rb → engine.rb} +33 -33
  13. data/lib/solid_jobs/errors.rb +1 -1
  14. data/lib/solid_jobs/executor.rb +34 -0
  15. data/lib/solid_jobs/failure_policy.rb +98 -0
  16. data/lib/solid_jobs/heartbeat.rb +15 -13
  17. data/lib/solid_jobs/integrity_check.rb +39 -36
  18. data/lib/solid_jobs/interceptor_registry.rb +57 -0
  19. data/lib/solid_jobs/keyspace.rb +42 -0
  20. data/lib/solid_jobs/{testing.rb → lab.rb} +19 -19
  21. data/lib/solid_jobs/publisher.rb +186 -0
  22. data/lib/solid_jobs/rails_adapter.rb +13 -0
  23. data/lib/solid_jobs/{rails.rb → railtie.rb} +1 -2
  24. data/lib/solid_jobs/recovery.rb +8 -7
  25. data/lib/solid_jobs/task.rb +147 -0
  26. data/lib/solid_jobs/{scheduler.rb → timer.rb} +8 -8
  27. data/lib/solid_jobs/utilities.rb +5 -11
  28. data/lib/solid_jobs/version.rb +1 -1
  29. data/lib/solid_jobs.rb +28 -27
  30. metadata +19 -18
  31. data/lib/solid_jobs/active_job.rb +0 -16
  32. data/lib/solid_jobs/api.rb +0 -457
  33. data/lib/solid_jobs/client.rb +0 -202
  34. data/lib/solid_jobs/fetch.rb +0 -310
  35. data/lib/solid_jobs/job.rb +0 -156
  36. data/lib/solid_jobs/middleware/chain.rb +0 -104
  37. data/lib/solid_jobs/retry_service.rb +0 -100
  38. data/lib/solid_jobs/worker.rb +0 -35
data/docs/reliability.md CHANGED
@@ -1,63 +1,61 @@
1
1
  # Reliability model
2
2
 
3
- SolidJobs uses a Redis-backed at-least-once state machine:
3
+ SolidJobs uses a namespaced Redis-backed at-least-once state machine:
4
4
 
5
5
  ```text
6
6
  READY
7
7
  |
8
- | BLMOVE reservation
8
+ | atomic claim
9
9
  v
10
- IN_PROGRESS ---- ACK/LREM ----> removed
10
+ CLAIMED ---- fenced completion ----> removed
11
11
  |
12
- +---- application error ----> RETRY or DEAD, then ACK
12
+ +---- application error ----> RETRYING or DISCARDED, then completion
13
13
  +---- graceful timeout -----> READY
14
- +---- process crash --------> reservation retained
14
+ +---- node crash -----------> claim retained
15
15
  |
16
16
  +---- recovery ----> READY
17
17
  ```
18
18
 
19
- The reservation list is named `<process-identity>:reserved:<processor-id>`.
20
- Each execution receives a distinct reservation journal entry:
19
+ Ready tasks live in `solid_jobs:channel:<name>`. Each executor owns at most one
20
+ claimed-task list and one claim journal entry:
21
21
 
22
22
  ```text
23
- job_id stable across replay
24
- reservation_id unique per execution attempt
25
- process_id owning process identity
26
- worker_id owning Processor Ractor
27
- attempt monotonically increasing execution count
28
- reserved_at wall-clock reservation time
23
+ task_id stable across replay
24
+ claim_token unique per execution attempt
25
+ node_id owning server identity
26
+ executor_id owning Executor Ractor
27
+ channel destination used by recovery
28
+ attempt monotonically increasing execution count
29
+ claimed_at wall-clock claim time
29
30
  ```
30
31
 
31
- ACK and requeue are fenced by `reservation_id`. Their Lua scripts first
32
- verify that the worker still owns the current journal generation. A delayed
33
- or revived worker cannot remove or requeue a newer reservation, even when a
34
- supervisor reuses the same processor slot.
32
+ Completion and requeue are fenced by `claim_token`. Their Lua scripts verify
33
+ that the executor still owns the current journal generation before changing
34
+ state. A delayed or revived executor cannot complete or requeue a newer claim.
35
35
 
36
- The payload retains its canonical `queue` field, allowing another process to
37
- restore it after the owner is no longer alive. Recovery uses process liveness
38
- and heartbeat, never reservation age alone, so a legitimate long-running job
39
- is not stolen while its owner remains alive.
36
+ The envelope retains its canonical `channel`, allowing another node to restore
37
+ it after its owner dies. Recovery uses node liveness and heartbeat, never claim
38
+ age alone, so a legitimate long-running task is not stolen.
40
39
 
41
40
  ## Failure boundaries
42
41
 
43
- - Before reservation: the job remains in `queue:<name>`.
44
- - After reservation or during `perform`: the job remains reserved.
45
- - After the application effect but before ACK: recovery replays the job.
46
- - Redis unavailable during ACK: the job remains reserved and is replayed.
47
- - Graceful shutdown: the active job may finish within the configured timeout;
42
+ - Before claim: the task remains in `solid_jobs:channel:<name>`.
43
+ - After claim or during `execute_task`: the task remains claimed.
44
+ - After the application effect but before completion: recovery replays it.
45
+ - Redis unavailable during completion: the task remains claimed and is replayed.
46
+ - Graceful shutdown: the active task may finish within the configured timeout;
48
47
  otherwise it is interrupted and requeued.
49
- - `SIGKILL`: no handler runs; recovery relies exclusively on Redis state.
48
+ - `SIGKILL`: recovery relies exclusively on Redis state.
50
49
 
51
- This design intentionally favors no job loss over duplicate suppression.
52
- Exactly-once side effects require application-level idempotency.
50
+ This design favors no task loss over duplicate suppression. Exactly-once side
51
+ effects require application-level idempotency.
53
52
 
54
53
  ## Startup isolation
55
54
 
56
- The server does not reserve work while components are booting. Heartbeat,
57
- Processor, and Scheduler Ractors initialize their local configuration and
58
- Redis pool, report `READY` exactly once, and wait behind
59
- `SolidJobs::StartupBarrier`. Processing starts only after every component is
60
- ready:
55
+ The server does not claim work while components are booting. Heartbeat,
56
+ Engine, and Timer Ractors initialize their local configuration and Redis
57
+ pool, report `READY` exactly once, and wait behind
58
+ `SolidJobs::StartupBarrier`:
61
59
 
62
60
  ```text
63
61
  BOOTING -> ALL_READY -> RUNNING
@@ -65,85 +63,69 @@ BOOTING -> ALL_READY -> RUNNING
65
63
  ```
66
64
 
67
65
  Boot failure is terminal. Already-ready components receive `:abort`, close
68
- their local resources, and never enter their fetch loops. Cleanup is
69
- idempotent. Component startup is serialized, avoiding concurrent TCP/RESP
70
- initialization paths known to crash Ruby 3.4 (reproduced on 3.4.4 macOS arm64
71
- and 3.4.11 Linux x86_64; `rake startup_torture` is the reproducer); normal
72
- processing remains parallel after `RUNNING`.
73
-
74
- ### Ruby 3.4 Ractor caveat
75
-
76
- Ruby 3.4's Ractor scheduler can deadlock the whole VM (main thread included)
77
- when a GC-triggered `rb_ractor_sched_barrier_start` runs while several
78
- Ractors exchange `move: true` messages: every thread parks in
79
- `ractor_sched_barrier_join_wait_locked` and the barrier never completes.
80
- `test/support/ractor_barrier_repro.rb` reproduces it **without SolidJobs or
81
- Redis** (one receiver, four senders, 24k moved messages per iteration):
82
- Ruby 3.4.4 freezes within the first iterations, Ruby 4.0.1 completes 30/30.
83
- Inside SolidJobs the same traffic pattern is the Processor → Heartbeat
84
- `:work`/`:done`/`:stats` channel, so any multi-Processor server on Ruby 3.4
85
- is exposed; once frozen, neither `Timeout` nor process exit
86
- (`rb_ractor_terminate_all`) can recover.
87
-
88
- Recommendation: **run multi-Ractor SolidJobs servers on Ruby ≥ 4.0**. On
89
- Ruby 3.4 use the client/API side freely, and prefer one process per
90
- Processor (`concurrency: 1`) for the server. Server-based stress tests are
91
- skipped on Ruby < 4 for this reason, and `Server#start` logs a warning when
92
- it detects `concurrency > 1` on Ruby < 4.
93
-
94
- ### Bounded shutdown
95
-
96
- `Server#stop` never waits forever for a component. Each Ractor gets
97
- `shutdown_timeout + Server::STOP_GRACE` to return after `:stop`; the signal
98
- is re-sent up to `STOP_RESENDS` times, then the component is abandoned with
99
- an error log so the process can proceed with shutdown.
66
+ their local resources, and never enter their claim loops. Cleanup is
67
+ idempotent.
68
+
69
+ ## Ruby 3.4 Ractor caveat
70
+
71
+ Ruby 3.4's Ractor scheduler can deadlock the VM when a GC-triggered scheduler
72
+ barrier runs while several Ractors exchange moved messages.
73
+ `test/support/ractor_barrier_repro.rb` reproduces this without SolidJobs or
74
+ Redis. Ruby 4.0.1 completes the equivalent reproducer.
75
+
76
+ Run multi-Ractor SolidJobs servers on Ruby >= 4.0. On Ruby 3.4, prefer one
77
+ process per Engine (`concurrency: 1`). Conductor stress tests are skipped on
78
+ Ruby versions affected by this runtime issue, and `Conductor#start` warns when it
79
+ detects a multi-Ractor configuration there.
80
+
81
+ ## Bounded shutdown
82
+
83
+ `Conductor#stop` never waits forever for a component. Each Ractor gets
84
+ `shutdown_timeout + Conductor::STOP_GRACE` to return after `:stop`; the signal is
85
+ re-sent up to `STOP_RESENDS` times before the component is abandoned with an
86
+ error log.
100
87
 
101
88
  ## Configuration scope
102
89
 
103
90
  `SolidJobs.config` and the testing mode are Ractor-local, not thread-local.
104
- Every thread and fiber inside a Ractor (for example Puma workers or Rails
105
- request threads) shares the configuration set on that Ractor, while each
106
- Ractor keeps its own isolated configuration.
91
+ Every thread and fiber inside a Ractor shares its configuration, while each
92
+ Ractor owns isolated mutable state and Redis connections.
107
93
 
108
94
  ## Integrity auditing
109
95
 
110
- Fault and chaos tests can reconcile a known set of job IDs against every
111
- Redis-backed state:
96
+ Fault tests can reconcile known task IDs against every Redis-backed state:
112
97
 
113
98
  ```ruby
114
99
  report = SolidJobs::IntegrityCheck.call(
115
- expected_job_ids: submitted_job_ids,
100
+ expected_job_ids: submitted_task_ids,
116
101
  acked_key: "test:completed",
117
102
  )
118
103
 
119
104
  raise report.inspect unless report.ok?
120
105
  ```
121
106
 
122
- The report separates lost jobs, unexpected/orphaned jobs, inconsistent
123
- reservation journals, dangling attempt indexes, duplicate active states, and
124
- malformed payloads. ACK removes the completed job's attempt index atomically;
125
- requeue retains it so a recovered reservation increments the same attempt
126
- sequence.
107
+ The report separates lost tasks, unexpected tasks, inconsistent claim
108
+ journals, dangling attempt indexes, duplicate active states, and malformed
109
+ envelopes. Completion removes the task's attempt index atomically; requeue
110
+ retains it so recovery increments the same attempt sequence.
127
111
 
128
- ## Backpressure and queues
112
+ ## Backpressure and channels
129
113
 
130
- Each Processor owns at most one reservation and reserves only immediately
131
- before execution. A process with concurrency `N` therefore holds at most `N`
132
- active reservations, regardless of Redis queue depth.
114
+ Each Engine owns at most one claim and claims only immediately before
115
+ execution. A node with concurrency `N` therefore holds at most `N` active
116
+ claims, regardless of channel depth.
133
117
 
134
- Queue mode defaults to `:weighted`:
118
+ Channel order defaults to `:weighted`:
135
119
 
136
- - `:weighted` uses bounded weighted round-robin and does not starve configured
137
- queues;
138
- - `:strict` always checks queues in declaration order and may intentionally
139
- starve lower-priority queues while a higher-priority queue remains busy;
140
- - `:random` samples the configured weighted queue list on each reservation.
120
+ - `:weighted` uses bounded weighted round-robin;
121
+ - `:priority` checks channels in declaration order;
122
+ - `:shuffle` samples the configured weighted channel list.
141
123
 
142
- Paused queues are excluded before reservation.
124
+ Paused channels are excluded before claiming.
143
125
 
144
126
  ## Retry storms
145
127
 
146
- Retries use capped exponential backoff with equal jitter. For attempt `n`, the
147
- ceiling is `min(retry_base_delay * 2**n, retry_max_delay)` and the actual delay
128
+ Retries use capped exponential backoff with equal jitter. For failure `n`, the
129
+ ceiling is `min(retry_base_delay * 2**(n - 1), retry_max_delay)` and the delay
148
130
  is distributed between half and all of that ceiling. Defaults are 15 seconds
149
- and one hour. This spreads recovery traffic after a shared external outage.
131
+ and one hour.
data/exe/solid-jobs CHANGED
@@ -2,7 +2,6 @@
2
2
  # frozen_string_literal: true
3
3
 
4
4
  require "solid_jobs"
5
- require "solid_jobs/cli"
6
-
7
- exit SolidJobs::CLI.new.run
5
+ require "solid_jobs/console"
8
6
 
7
+ exit SolidJobs::Console.new.run
@@ -1,7 +1,7 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  require "solid_jobs"
4
- require "solid_jobs/active_job"
4
+ require "solid_jobs/rails_adapter"
5
5
 
6
6
  module ActiveJob
7
7
  module QueueAdapters
@@ -27,16 +27,15 @@ module ActiveJob
27
27
  private
28
28
 
29
29
  def push(job, at: nil)
30
- payload = {
31
- "class" => SolidJobs::ActiveJob::Wrapper,
30
+ envelope = {
31
+ "task" => SolidJobs::RailsAdapter::AdapterTask,
32
32
  "wrapped" => job.class.name,
33
- "queue" => job.queue_name,
34
- "args" => [job.serialize],
33
+ "channel" => job.queue_name,
34
+ "arguments" => [job.serialize],
35
35
  }
36
- payload["at"] = at if at
37
- SolidJobs::Client.push(payload)
36
+ envelope["run_at"] = at if at
37
+ SolidJobs::Publisher.publish(envelope)
38
38
  end
39
39
  end
40
40
  end
41
41
  end
42
-
@@ -3,36 +3,36 @@
3
3
  require "logger"
4
4
 
5
5
  module SolidJobs
6
- class Config
7
- DEFAULT_JOB_OPTIONS = {
8
- "queue" => "default",
9
- "retry" => true,
6
+ class Blueprint
7
+ DEFAULT_TASK_OPTIONS = {
8
+ "channel" => "default",
9
+ "max_failures" => 25,
10
10
  }.freeze
11
11
 
12
- attr_accessor :concurrency, :dead_max_jobs, :dead_timeout, :logger,
12
+ attr_accessor :concurrency, :discarded_limit, :discarded_retention, :logger,
13
13
  :on_complex_arguments, :poll_interval_average, :shutdown_timeout,
14
14
  :reliable_fetch, :retry_base_delay, :retry_max_delay
15
- attr_reader :client_middleware, :default_job_options, :error_handlers,
16
- :redis_config, :server_middleware
15
+ attr_reader :default_task_options, :error_handlers, :execute_interceptors,
16
+ :publish_interceptors, :redis_config
17
17
 
18
- def initialize(redis: nil, concurrency: 5, queues: ["default"])
18
+ def initialize(redis: nil, concurrency: 5, channels: ["default"])
19
19
  @redis_config = redis || SolidRedis.config(url: ENV.fetch("REDIS_URL", "redis://127.0.0.1:6379/0"))
20
20
  @concurrency = Integer(concurrency)
21
- @queues = normalize_queues(queues)
22
- @default_job_options = DEFAULT_JOB_OPTIONS
23
- @client_middleware = Middleware::Chain.new
24
- @server_middleware = Middleware::Chain.new
21
+ @channels = normalize_channels(channels)
22
+ @default_task_options = DEFAULT_TASK_OPTIONS
23
+ @publish_interceptors = InterceptorRegistry.new
24
+ @execute_interceptors = InterceptorRegistry.new
25
25
  @error_handlers = []
26
26
  @lifecycle_callbacks = Hash.new { |hash, event| hash[event] = [] }
27
27
  @on_complex_arguments = :raise
28
28
  @poll_interval_average = 5.0
29
29
  @shutdown_timeout = 25.0
30
- @queue_mode = :weighted
30
+ @channel_order = :weighted
31
31
  @reliable_fetch = true
32
32
  @retry_base_delay = 15.0
33
33
  @retry_max_delay = 3_600.0
34
- @dead_max_jobs = 10_000
35
- @dead_timeout = 180 * 24 * 60 * 60
34
+ @discarded_limit = 10_000
35
+ @discarded_retention = 180 * 24 * 60 * 60
36
36
  @logger = Logger.new($stdout)
37
37
  @redis_pool = nil
38
38
  end
@@ -50,42 +50,42 @@ module SolidJobs
50
50
  redis_pool.with { |connection| yield connection }
51
51
  end
52
52
 
53
- def queues
54
- @queues.map(&:first)
53
+ def channels
54
+ @channels.map(&:first)
55
55
  end
56
56
 
57
- def queues=(values)
58
- @queues = normalize_queues(values)
57
+ def channels=(values)
58
+ @channels = normalize_channels(values)
59
59
  end
60
60
 
61
- def queue_entries
62
- @queues.dup
61
+ def channel_entries
62
+ @channels.dup
63
63
  end
64
64
 
65
- def queue_mode
66
- @queue_mode
65
+ def channel_order
66
+ @channel_order
67
67
  end
68
68
 
69
- def queue_mode=(mode)
69
+ def channel_order=(mode)
70
70
  mode = mode.to_sym
71
- unless %i[strict weighted random].include?(mode)
72
- raise ArgumentError, "queue_mode must be :strict, :weighted, or :random"
71
+ unless %i[priority weighted shuffle].include?(mode)
72
+ raise ArgumentError, "channel_order must be :priority, :weighted, or :shuffle"
73
73
  end
74
74
 
75
- @queue_mode = mode
75
+ @channel_order = mode
76
76
  end
77
77
 
78
- def strict
79
- queue_mode == :strict
78
+ def prioritized?
79
+ channel_order == :priority
80
80
  end
81
81
 
82
- def strict=(value)
83
- self.queue_mode = value ? :strict : :weighted
82
+ def prioritized=(value)
83
+ self.channel_order = value ? :priority : :weighted
84
84
  end
85
85
 
86
- def default_job_options=(options)
87
- @default_job_options = Utilities.shareable_copy(
88
- DEFAULT_JOB_OPTIONS.merge(Utilities.stringify_keys(options)),
86
+ def default_task_options=(options)
87
+ @default_task_options = Utilities.shareable_copy(
88
+ DEFAULT_TASK_OPTIONS.merge(Utilities.stringify_keys(options)),
89
89
  )
90
90
  end
91
91
 
@@ -111,26 +111,26 @@ module SolidJobs
111
111
  end
112
112
 
113
113
  def inspect
114
- "#<#{self.class.name} concurrency=#{concurrency} queues=#{queues.inspect}>"
114
+ "#<#{self.class.name} concurrency=#{concurrency} channels=#{channels.inspect}>"
115
115
  end
116
116
 
117
117
  def ractor_snapshot
118
118
  Utilities.shareable_copy(
119
119
  redis_config: redis_config,
120
120
  concurrency: concurrency,
121
- queues: queue_entries,
122
- queue_mode: queue_mode,
121
+ channels: channel_entries,
122
+ channel_order: channel_order,
123
123
  reliable_fetch: reliable_fetch,
124
- default_job_options: default_job_options,
124
+ default_task_options: default_task_options,
125
125
  on_complex_arguments: on_complex_arguments,
126
126
  poll_interval_average: poll_interval_average,
127
127
  shutdown_timeout: shutdown_timeout,
128
- dead_max_jobs: dead_max_jobs,
129
- dead_timeout: dead_timeout,
128
+ discarded_limit: discarded_limit,
129
+ discarded_retention: discarded_retention,
130
130
  retry_base_delay: retry_base_delay,
131
131
  retry_max_delay: retry_max_delay,
132
- client_middleware: client_middleware.snapshot,
133
- server_middleware: server_middleware.snapshot,
132
+ publish_interceptors: publish_interceptors.export,
133
+ execute_interceptors: execute_interceptors.export,
134
134
  )
135
135
  end
136
136
 
@@ -138,31 +138,31 @@ module SolidJobs
138
138
  config = new(
139
139
  redis: snapshot.fetch(:redis_config),
140
140
  concurrency: snapshot.fetch(:concurrency),
141
- queues: snapshot.fetch(:queues),
141
+ channels: snapshot.fetch(:channels),
142
142
  )
143
- config.queue_mode = snapshot.fetch(:queue_mode)
143
+ config.channel_order = snapshot.fetch(:channel_order)
144
144
  config.reliable_fetch = snapshot.fetch(:reliable_fetch)
145
- config.default_job_options = snapshot.fetch(:default_job_options)
145
+ config.default_task_options = snapshot.fetch(:default_task_options)
146
146
  config.on_complex_arguments = snapshot.fetch(:on_complex_arguments)
147
147
  config.poll_interval_average = snapshot.fetch(:poll_interval_average)
148
148
  config.shutdown_timeout = snapshot.fetch(:shutdown_timeout)
149
- config.dead_max_jobs = snapshot.fetch(:dead_max_jobs)
150
- config.dead_timeout = snapshot.fetch(:dead_timeout)
149
+ config.discarded_limit = snapshot.fetch(:discarded_limit)
150
+ config.discarded_retention = snapshot.fetch(:discarded_retention)
151
151
  config.retry_base_delay = snapshot.fetch(:retry_base_delay)
152
152
  config.retry_max_delay = snapshot.fetch(:retry_max_delay)
153
- config.client_middleware.restore(snapshot.fetch(:client_middleware))
154
- config.server_middleware.restore(snapshot.fetch(:server_middleware))
153
+ config.publish_interceptors.import(snapshot.fetch(:publish_interceptors))
154
+ config.execute_interceptors.import(snapshot.fetch(:execute_interceptors))
155
155
  config
156
156
  end
157
157
 
158
158
  private
159
159
 
160
- def normalize_queues(values)
160
+ def normalize_channels(values)
161
161
  Array(values).map do |value|
162
162
  name, weight = value.is_a?(Array) ? value : [value, 1]
163
163
  name = String(name)
164
164
  weight = Integer(weight)
165
- raise ArgumentError, "Queue weight must be positive" unless weight.positive?
165
+ raise ArgumentError, "Channel weight must be positive" unless weight.positive?
166
166
 
167
167
  [name.freeze, weight].freeze
168
168
  end