ruby_reactor 0.8.2 → 0.8.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (77) hide show
  1. checksums.yaml +4 -4
  2. data/.claude/skills/speckit-review/SKILL.md +324 -0
  3. data/.release-please-manifest.json +1 -1
  4. data/.specify/extensions.yml +10 -0
  5. data/.specify/feature.json +1 -1
  6. data/.specify/workflows/speckit/workflow.yml +13 -1
  7. data/.specify/workflows/workflow-registry.json +2 -2
  8. data/CHANGELOG.md +82 -0
  9. data/CLAUDE.md +2 -2
  10. data/README.md +35 -2
  11. data/lib/ruby_reactor/adapters/active_job/router.rb +19 -0
  12. data/lib/ruby_reactor/adapters/sidekiq/router.rb +21 -0
  13. data/lib/ruby_reactor/context.rb +26 -0
  14. data/lib/ruby_reactor/context_serializer.rb +4 -2
  15. data/lib/ruby_reactor/dsl/interrupt_builder.rb +14 -0
  16. data/lib/ruby_reactor/dsl/lockable.rb +76 -21
  17. data/lib/ruby_reactor/dsl/step_builder.rb +112 -1
  18. data/lib/ruby_reactor/error/async_result_pending.rb +1 -1
  19. data/lib/ruby_reactor/error/execution_parked.rb +16 -0
  20. data/lib/ruby_reactor/error/reactor_contention_park.rb +26 -0
  21. data/lib/ruby_reactor/error/step_contention_park.rb +26 -0
  22. data/lib/ruby_reactor/executor/async_step_dispatch.rb +109 -3
  23. data/lib/ruby_reactor/executor/compensation_manager.rb +99 -17
  24. data/lib/ruby_reactor/executor/ordered_lock_support.rb +76 -44
  25. data/lib/ruby_reactor/executor/result_handler.rb +31 -11
  26. data/lib/ruby_reactor/executor/retry_manager.rb +9 -1
  27. data/lib/ruby_reactor/executor/step_coordination.rb +788 -0
  28. data/lib/ruby_reactor/executor/step_executor.rb +115 -11
  29. data/lib/ruby_reactor/executor.rb +90 -20
  30. data/lib/ruby_reactor/map/element_executor.rb +24 -2
  31. data/lib/ruby_reactor/map/helpers.rb +35 -11
  32. data/lib/ruby_reactor/max_retries_exhausted_failure.rb +2 -2
  33. data/lib/ruby_reactor/open_telemetry.rb +61 -24
  34. data/lib/ruby_reactor/retry_context.rb +31 -2
  35. data/lib/ruby_reactor/rspec/helpers.rb +15 -0
  36. data/lib/ruby_reactor/rspec/matchers.rb +92 -0
  37. data/lib/ruby_reactor/rspec/test_subject.rb +7 -1
  38. data/lib/ruby_reactor/step/async_reactor_step.rb +40 -24
  39. data/lib/ruby_reactor/step/compose_step.rb +14 -3
  40. data/lib/ruby_reactor/step.rb +49 -7
  41. data/lib/ruby_reactor/step_sweeper.rb +29 -1
  42. data/lib/ruby_reactor/step_worker.rb +260 -37
  43. data/lib/ruby_reactor/version.rb +1 -1
  44. data/lib/ruby_reactor/web/api.rb +72 -7
  45. data/lib/ruby_reactor/web/coordination_serializer.rb +120 -2
  46. data/lib/ruby_reactor/web/public/assets/{index-Dw4KV4QY.js → index-CeZU-ESu.js} +9 -9
  47. data/lib/ruby_reactor/web/public/index.html +1 -1
  48. data/lib/ruby_reactor/worker.rb +56 -30
  49. data/lib/ruby_reactor.rb +27 -5
  50. data/specs/future_improvements.md +250 -0
  51. metadata +8 -28
  52. data/specs/002-step-input-contracts/checklists/requirements.md +0 -49
  53. data/specs/002-step-input-contracts/contracts/dsl-surface.md +0 -193
  54. data/specs/002-step-input-contracts/data-model.md +0 -115
  55. data/specs/002-step-input-contracts/plan.md +0 -165
  56. data/specs/002-step-input-contracts/quickstart.md +0 -170
  57. data/specs/002-step-input-contracts/research.md +0 -233
  58. data/specs/002-step-input-contracts/spec.md +0 -359
  59. data/specs/002-step-input-contracts/tasks.md +0 -367
  60. data/specs/004-inheritable-step-class/checklists/requirements.md +0 -40
  61. data/specs/004-inheritable-step-class/contracts/step-lifecycle.md +0 -85
  62. data/specs/004-inheritable-step-class/data-model.md +0 -116
  63. data/specs/004-inheritable-step-class/plan.md +0 -174
  64. data/specs/004-inheritable-step-class/quickstart.md +0 -112
  65. data/specs/004-inheritable-step-class/research.md +0 -308
  66. data/specs/004-inheritable-step-class/spec.md +0 -316
  67. data/specs/004-inheritable-step-class/tasks.md +0 -258
  68. data/specs/active_job.md +0 -259
  69. data/specs/deferred-003-step-lock-declarations/checklists/requirements.md +0 -51
  70. data/specs/deferred-003-step-lock-declarations/contracts/dsl-surface.md +0 -154
  71. data/specs/deferred-003-step-lock-declarations/data-model.md +0 -131
  72. data/specs/deferred-003-step-lock-declarations/plan.md +0 -166
  73. data/specs/deferred-003-step-lock-declarations/quickstart.md +0 -169
  74. data/specs/deferred-003-step-lock-declarations/research.md +0 -196
  75. data/specs/deferred-003-step-lock-declarations/spec.md +0 -447
  76. data/specs/deferred-003-step-lock-declarations/tasks.md +0 -572
  77. data/specs/possible_feature.md +0 -22
@@ -5,7 +5,7 @@
5
5
  <link rel="icon" type="image/svg+xml" href="./vite.svg" />
6
6
  <meta name="viewport" content="width=device-width, initial-scale=1.0" />
7
7
  <title>ui</title>
8
- <script type="module" crossorigin src="./assets/index-Dw4KV4QY.js"></script>
8
+ <script type="module" crossorigin src="./assets/index-CeZU-ESu.js"></script>
9
9
  <link rel="stylesheet" crossorigin href="./assets/index-BQvIWPdx.css">
10
10
  </head>
11
11
  <body>
@@ -10,6 +10,46 @@ module RubyReactor
10
10
  module Worker
11
11
  TERMINAL_STATUSES = %w[completed failed cancelled skipped].freeze
12
12
 
13
+ # Use the error's `retry_after_seconds` hint when available
14
+ # (RateLimit::ExceededError carries the time until the bucket rolls);
15
+ # otherwise fall back to the configured base + jitter for lock/semaphore
16
+ # contention which has no precise hint. Module-level (not just an
17
+ # instance method) so `StepCoordination::Contended` — which wraps a
18
+ # contention error but is not itself a snooze-worthy reactor-level
19
+ # error — can reuse the identical hint logic from the async_step worker
20
+ # (T036) without including this whole module.
21
+ #
22
+ # OrderedLock::WaitError is deliberately excluded from the hint path: its
23
+ # `retry_after_seconds` is the poison-pill window (the upper bound before
24
+ # a *dead* blocker is force-advanced), NOT how long the *live* blocker
25
+ # will take — which is usually milliseconds. Snoozing for the full window
26
+ # would make every out-of-order nonce sleep up to poison_pill_timeout even
27
+ # though its blocker finishes immediately, collapsing throughput. Re-poll
28
+ # at the base delay instead; poison auto-advance still clears a genuinely
29
+ # dead blocker on a later gate.
30
+ def self.snooze_delay(config, error)
31
+ jitter = config.lock_snooze_jitter.to_f
32
+ jitter_amount = jitter.positive? ? rand(0.0..jitter) : 0.0
33
+
34
+ if hinted_retry?(error)
35
+ [error.retry_after_seconds.to_f, 0.1].max + jitter_amount
36
+ else
37
+ config.lock_snooze_base_delay.to_f + jitter_amount
38
+ end
39
+ end
40
+
41
+ def self.hinted_retry?(error)
42
+ # `StepCoordination::Contended` wraps the WaitError as `.original` —
43
+ # unwrap so the same exclusion applies whether the caller is the
44
+ # reactor-level ordered lock (raises WaitError directly) or a step's
45
+ # (raises Contended, whose OWN `retry_after_seconds` just forwards the
46
+ # wrapped error's hint unchanged).
47
+ original = error.respond_to?(:original) ? error.original : error
48
+ return false if original.is_a?(RubyReactor::OrderedLock::WaitError)
49
+
50
+ error.respond_to?(:retry_after_seconds) && error.retry_after_seconds
51
+ end
52
+
13
53
  # Last line of observability when a job burns its whole retry budget on an
14
54
  # infrastructure failure and the backend then discards it (Sidekiq runs
15
55
  # with `dead: false`): without this, the context would stay "running"
@@ -93,8 +133,11 @@ module RubyReactor
93
133
  RubyReactor::Semaphore::AcquisitionError,
94
134
  RubyReactor::RateLimit::ExceededError,
95
135
  RubyReactor::OrderedLock::WaitError,
96
- RubyReactor::Error::AsyncResultPending => e
97
- # Snooze on expected concurrency, rate, or ordering contention.
136
+ RubyReactor::Error::ExecutionParked => e
137
+ # Snooze on expected concurrency, rate, or ordering contention, or on
138
+ # a park signal (a step's contention, or an awaited background result)
139
+ # — raised at any nesting depth, after every executor on the stack has
140
+ # parked its own holds and saved.
98
141
  # OrderedLock::WaitError carries a poison-pill-derived retry hint,
99
142
  # consumed by compute_snooze_delay below. We avoid the framework's native
100
143
  # retry path so this doesn't burn the job's retry budget or appear
@@ -147,12 +190,14 @@ module RubyReactor
147
190
  # duplicate of the *same* execution may wait arbitrarily long for the
148
191
  # live original to finish (e.g. a sweeper re-enqueue racing a slow but
149
192
  # alive worker). Capping it would fail a legitimately-waiting duplicate.
150
- # A parked async wait is likewise uncapped HERE: its bound is
151
- # `async_park_timeout`, enforced against `dispatched_at` at the wait
152
- # site — counting snoozes would double-bound it with the wrong unit.
193
+ # A park signal is likewise uncapped HERE, bounded where it is raised:
194
+ # an async wait by `async_park_timeout` against `dispatched_at`, a
195
+ # step's contention by `lock_snooze_max_attempts` on that step's own
196
+ # contention counter (`StepExecutor#handle_contention`) — counting
197
+ # snoozes too would double-bound either with the wrong unit.
153
198
  capped = !(error.is_a?(RubyReactor::OrderedLock::WaitError) ||
154
199
  error.is_a?(RubyReactor::Lock::ContextLockContention) ||
155
- error.is_a?(RubyReactor::Error::AsyncResultPending))
200
+ error.is_a?(RubyReactor::Error::ExecutionParked))
156
201
 
157
202
  if capped && max != :infinity && snooze_count >= max
158
203
  escalate_snooze(context, snooze_count, error)
@@ -165,34 +210,15 @@ module RubyReactor
165
210
  self.class.perform_in(delay, context_id, reactor_class_name, snooze_count + 1)
166
211
  end
167
212
 
168
- # Use the error's `retry_after_seconds` hint when available
169
- # (RateLimit::ExceededError carries the time until the bucket rolls);
170
- # otherwise fall back to the configured base + jitter for lock/semaphore
171
- # contention which has no precise hint.
172
- #
173
- # OrderedLock::WaitError is deliberately excluded from the hint path: its
174
- # `retry_after_seconds` is the poison-pill window (the upper bound before
175
- # a *dead* blocker is force-advanced), NOT how long the *live* blocker
176
- # will take — which is usually milliseconds. Snoozing for the full window
177
- # would make every out-of-order nonce sleep up to poison_pill_timeout even
178
- # though its blocker finishes immediately, collapsing throughput. Re-poll
179
- # at the base delay instead; poison auto-advance still clears a genuinely
180
- # dead blocker on a later gate.
213
+ # Instance methods delegate to the module functions above — worker
214
+ # behavior is unchanged, just relocated so other callers (StepWorker,
215
+ # T036) can reuse the same logic without a Worker instance.
181
216
  def compute_snooze_delay(config, error)
182
- jitter = config.lock_snooze_jitter.to_f
183
- jitter_amount = jitter.positive? ? rand(0.0..jitter) : 0.0
184
-
185
- if hinted_retry?(error)
186
- [error.retry_after_seconds.to_f, 0.1].max + jitter_amount
187
- else
188
- config.lock_snooze_base_delay.to_f + jitter_amount
189
- end
217
+ Worker.snooze_delay(config, error)
190
218
  end
191
219
 
192
220
  def hinted_retry?(error)
193
- return false if error.is_a?(RubyReactor::OrderedLock::WaitError)
194
-
195
- error.respond_to?(:retry_after_seconds) && error.retry_after_seconds
221
+ Worker.hinted_retry?(error)
196
222
  end
197
223
 
198
224
  def escalate_snooze(context, snooze_count, error)
data/lib/ruby_reactor.rb CHANGED
@@ -133,14 +133,14 @@ module RubyReactor
133
133
 
134
134
  class Failure
135
135
  attr_reader :error, :retryable, :step_name, :inputs, :backtrace, :reactor_name, :step_arguments, :exception_class,
136
- :file_path, :line_number, :code_snippet, :validation_errors
136
+ :file_path, :line_number, :code_snippet, :validation_errors, :rollback_failures
137
137
 
138
- # rubocop:disable Metrics/ParameterLists, Metrics/CyclomaticComplexity, Metrics/PerceivedComplexity
138
+ # rubocop:disable Metrics/ParameterLists, Metrics/CyclomaticComplexity, Metrics/PerceivedComplexity, Metrics/MethodLength
139
139
  def initialize(error, retryable: nil, step_name: nil, inputs: {}, backtrace: nil, redact_inputs: [],
140
140
  reactor_name: nil, step_arguments: {}, exception_class: nil,
141
141
  file_path: nil, line_number: nil, code_snippet: nil, invalid_payload: false, validation_errors: nil,
142
- **opts)
143
- # rubocop:enable Metrics/ParameterLists, Metrics/CyclomaticComplexity, Metrics/PerceivedComplexity
142
+ rollback_failures: nil, **opts)
143
+ # rubocop:enable Metrics/ParameterLists, Metrics/CyclomaticComplexity, Metrics/PerceivedComplexity, Metrics/MethodLength
144
144
  retryable = opts[:retry] if opts.key?(:retry) # `retry:` wins over `retryable:` when both are given
145
145
  @error = error
146
146
 
@@ -159,6 +159,7 @@ module RubyReactor
159
159
  line_number ||= attributes[:line_number]
160
160
  code_snippet ||= attributes[:code_snippet]
161
161
  validation_errors ||= attributes[:validation_errors]
162
+ rollback_failures ||= attributes[:rollback_failures]
162
163
  end
163
164
 
164
165
  @retryable = if retryable.nil?
@@ -179,6 +180,9 @@ module RubyReactor
179
180
  @code_snippet = code_snippet
180
181
  @invalid_payload = invalid_payload
181
182
  @validation_errors = validation_errors
183
+ # Every undo/compensation that did not complete during this failure's
184
+ # rollback (005 FR-004): `{ step:, kind:, key:, reason:, message: }`.
185
+ @rollback_failures = normalize_rollback_failures(rollback_failures)
182
186
  end
183
187
 
184
188
  def success?
@@ -236,6 +240,7 @@ module RubyReactor
236
240
  line_number: @line_number,
237
241
  code_snippet: @code_snippet,
238
242
  validation_errors: @validation_errors,
243
+ rollback_failures: @rollback_failures,
239
244
  backtrace: @backtrace
240
245
  }
241
246
  end
@@ -334,9 +339,26 @@ module RubyReactor
334
339
  file_path: err[:file_path],
335
340
  line_number: err[:line_number],
336
341
  code_snippet: err[:code_snippet],
337
- validation_errors: err[:validation_errors]
342
+ validation_errors: err[:validation_errors],
343
+ rollback_failures: err[:rollback_failures]
338
344
  }
339
345
  end
346
+
347
+ ROLLBACK_FAILURE_SYMBOLS = %i[step kind reason].freeze
348
+ private_constant :ROLLBACK_FAILURE_SYMBOLS
349
+
350
+ # A stored failure comes back with string keys (and, through plain JSON,
351
+ # string values); `step`, `kind` and `reason` are Symbols on a live one.
352
+ def normalize_rollback_failures(entries)
353
+ return [] unless entries.is_a?(Array)
354
+
355
+ entries.map do |entry|
356
+ entry.to_h do |k, v|
357
+ key = k.to_sym
358
+ [key, ROLLBACK_FAILURE_SYMBOLS.include?(key) && v ? v.to_sym : v]
359
+ end
360
+ end
361
+ end
340
362
  end
341
363
 
342
364
  # Sentinel returned when a step's work is handed off to a worker job and is
@@ -0,0 +1,250 @@
1
+ # Future improvements for review
2
+
3
+ ## Guard
4
+
5
+ I want to add another hook feature to the step: `guard`
6
+ A guard concept is to execute after validations and before run.
7
+ example:
8
+
9
+ ```ruby
10
+ class SendEmail < RubyReactor::Step
11
+ input :email, :string, format?: /\A[^@\s]+@[^@\s]+\z/
12
+
13
+ def guard
14
+ fail!("Prevent spamming") if EmailService.sent_today?(inputs[:email])
15
+ success! # optional
16
+ end
17
+
18
+ def run
19
+ # do the work
20
+ end
21
+ end
22
+ ```
23
+ ## Fenced context writes
24
+
25
+ **Status:** proposal, not scheduled. Follows from 005 (step coordination remediation), R-18 and
26
+ R-19 in `specs/005-step-coordination-remediation/research.md`.
27
+
28
+ ### The rule
29
+
30
+ A context is written by one process only: the controlled execution that owns it.
31
+
32
+ - Async children (`async_step`, `async_reactor`, map elements) never write their parent. The
33
+ parent holds only the link written at dispatch; each child keeps its own record, and the
34
+ dashboard rebuilds its view from the links. 005 made this true for `async_step` (R-18).
35
+ - Today the rule is a convention: any code with a storage adapter can `SET` any context. This
36
+ proposal makes it a guarantee that the storage enforces.
37
+
38
+ ### Why a convention is not enough
39
+
40
+ Holding the lock at some point is not the same as owning the data when the write lands:
41
+
42
+ | Failure | What happens today |
43
+ |---|---|
44
+ | **Paused holder**: a worker holding the `async:<root id>` lock stalls (GC, network, a long rollback wait) past `context_lock_ttl`. The lock expires, a redelivery takes it and makes progress, then the first worker wakes. | The first worker's next `store_context` is a plain `SET`. It overwrites the newer state. Its auto-extender fails silently. Nothing stops the write. |
45
+ | **Stale read**: `Worker#perform` loads the context *before* `resume_execution` takes the context lock. | A previous holder can write between that read and the lock. The new holder then runs on the older snapshot and saves over the newer one. |
46
+ | **Writer outside any execution**: a path writes with no lock at all. | It overwrites whatever the live execution saved. 005 fixed two of these (`StepWorker#save_root`, the map collector's post-resume save); others remain (below). |
47
+
48
+ The first two are the textbook reasons distributed locks need fencing. A lock that only *excludes*
49
+ is not enough. The storage must also *reject writes from anyone who no longer holds the lock*.
50
+
51
+ ### Remaining writers outside a controlled execution
52
+
53
+ Found while auditing for 005 R-19:
54
+
55
+ | Writer | Where | Today |
56
+ |---|---|---|
57
+ | `Worker.record_retries_exhausted` | backend's retries-exhausted hook | Plain read-modify-write, no lock. Marks the context failed even if a redelivery is live. |
58
+ | `Worker#escalate_snooze` | after the executor returned | Runs after the executor released the context lock. |
59
+ | `Worker#handle_deserialization_failure` | before any executor | Writes a failed payload with no lock. |
60
+ | Map collector, failure branch | `Map::Helpers#resume_parent_execution` | Holds `map_collect:<map_id>`, never the parent's `async:` lock. |
61
+ | `Reactor.cancel`, `Reactor.undo` | any process (app code, console) | `find` → mutate → `save_context`: a blind write racing a live worker. |
62
+ | `Reactor#continue` (interrupt resume) | web request / app code | Writes the payload before `resume_execution` takes the lock. |
63
+ | Synchronous `Reactor.run` / `Executor#execute` | caller's process | Never takes the context lock, so its saves are unfenced. |
64
+
65
+ Writers that already belong to the owning execution:
66
+ - `Executor#save_context` and `#checkpoint!` on the worker path;
67
+ - `RetryManager#requeue_job`;
68
+ - `StepExecutor#checkpoint_root!`;
69
+ - `MapStep#prepare_async_execution`.
70
+
71
+ Two writers only *create* a row that nothing has read yet:
72
+ - `Reactor#save_context` at enqueue;
73
+ - `AsyncReactorStep#save`, the child's first row.
74
+
75
+ ### Proposal
76
+
77
+ Four parts. Parts 1 and 2 are the mechanism; parts 3 and 4 route the remaining writers through it.
78
+
79
+ #### 1. Ownership-checked, versioned writes (the fence)
80
+
81
+ - Every context row gets a version: a small key next to the blob,
82
+ `reactor:<Class>:context:<id>:v`, with the same TTL as the blob.
83
+ - Every write goes through one Lua script. The script takes:
84
+ - the lock key it claims to hold (`lock:async:<root id>`, or the map element's
85
+ `lock:map_element:<map>:<index>`);
86
+ - the owner token of that lock: the per-execution UUID `Executor#acquire_context_lock`
87
+ already generates;
88
+ - the version the writer loaded.
89
+ - The script writes only if **the lock is held by that token** *and* **the stored version equals
90
+ the loaded one**. It then bumps the version.
91
+ - It returns `ok`, `lost_lock` or `stale`:
92
+
93
+ ```lua
94
+ -- KEYS: context_key, version_key, lock_key ARGV: blob, ttl, owner, expected_version
95
+ if redis.call('hget', KEYS[3], 'owner') ~= ARGV[3] then return 'lost_lock' end
96
+ local current = tonumber(redis.call('get', KEYS[2]) or '0')
97
+ if current ~= tonumber(ARGV[4]) then return 'stale' end
98
+ redis.call('set', KEYS[1], ARGV[1], 'EX', ARGV[2])
99
+ redis.call('set', KEYS[2], current + 1, 'EX', ARGV[2])
100
+ return 'ok'
101
+ ```
102
+
103
+ Why both checks:
104
+ - The **owner check** stops a paused holder from writing at all, even before anyone else has
105
+ written.
106
+ - The **version check** catches a stale read by the current holder, and any path this proposal
107
+ missed.
108
+
109
+ The lock and the data live in the same Redis, and the check runs atomically inside it, so the
110
+ owner comparison does the job of a fencing token. There are no clock or ordering assumptions.
111
+ If locks and contexts ever move to different stores, switch to classic monotonic fencing tokens
112
+ (`INCR` on acquire, highest-token-wins at the store).
113
+
114
+ **Creating a row** uses the same script with `expected_version = 0` and no lock check: a
115
+ create-only `SET NX` equivalent. It covers `Reactor#save_context` at enqueue and an
116
+ `async_reactor` child's first row.
117
+
118
+ #### 2. Lock, then load
119
+
120
+ - A worker takes the context lock **before** it reads the context. The restructure:
121
+ `Worker#perform` acquires `async:<root id>`, then retrieves and deserializes, then hands
122
+ the token and the loaded version to `Executor`. Today `resume_execution` acquires the lock
123
+ mid-way.
124
+ - A synchronous `Reactor.run` takes the same lock for its whole run. It is one `SET NX` and a
125
+ release; today it takes none.
126
+ - Composed children already run inside the root's lock and write through the root. They reuse
127
+ the root's token.
128
+
129
+ #### 3. External actors send requests, not writes
130
+
131
+ Operator actions must not write a live context. Make them requests the owning execution
132
+ consumes, the same pattern as async children writing their own records:
133
+
134
+ - **`Reactor.cancel` / `.undo`**:
135
+ - write a `cancel_requested` key, `reactor:<Class>:context:<id>:cancel`;
136
+ - the owning execution checks it at each step boundary and cancels or undoes itself, under its
137
+ own lock;
138
+ - if no execution is live (the context lock is free), `cancel` takes the lock itself and
139
+ applies the request directly: lock, load, fenced write.
140
+ - **`Reactor#continue`**: take the lock first (bounded wait, because it is a user action), load,
141
+ store the payload, resume. If the lock is held, return "busy — retry" instead of racing.
142
+ - **Worker bookkeeping** (`record_retries_exhausted`, `escalate_snooze`, deserialization
143
+ failure):
144
+ - take the context lock with `wait: 0`;
145
+ - if another process holds it, that process owns the outcome, so log and skip;
146
+ - move `escalate_snooze` inside the executor's lock scope instead of after it.
147
+ - **Map collector**: take the parent's `async:` lock for both branches. The success branch passes
148
+ the token into `resume_execution`, which must accept an already-held lock (re-entrant by owner).
149
+
150
+ #### 4. What a rejected write does
151
+
152
+ A write that comes back `lost_lock` or `stale` raises `Error::ContextOwnershipLost`:
153
+ - The executor stops at once. It saves nothing else, runs no further steps, and releases only
154
+ holds it still owns (every release is already owner-checked).
155
+ - It logs one structured line: `event="ruby_reactor.context.ownership_lost"`, with the context
156
+ id, the lock key and the expected and actual version.
157
+ - The worker does **not** snooze or retry the job. The current owner is responsible for the run.
158
+ - A rejection is a correctness signal, not an error to hide. Expose it in the log and in the
159
+ `:failed_reactor` middleware event, but never mark the context failed: its owner decides the
160
+ outcome.
161
+
162
+ ### Constraints and open questions
163
+
164
+ - **Redis Cluster**: the script touches three keys, which must share a hash slot. Tag them with
165
+ the root id: `lock:async:{<root>}`, `reactor:<Class>:context:{<root>}`, and the version key. The
166
+ ordered lock already does this. Changing the key format needs a migration or a dual-read window.
167
+ - **Composed children's own rows**: `Executor#save_context` of a composed child also stores the
168
+ child under its own id. That is a second key in a different slot. Either tag it with the root
169
+ id too, or drop these rows. They are an observability path: the root blob already embeds the
170
+ child.
171
+ - **Inline test mode**: `Sidekiq::Testing.inline!` re-enters the worker inside a frame that holds
172
+ the lock, which is why `acquire_context_lock` is skipped there today. The fence needs the same
173
+ exemption, or re-entrancy by owner token.
174
+ - **Redis failover**: replication is asynchronous, so a failover can lose an acknowledged lock
175
+ or write. This proposal gives single-primary guarantees, the same as today's locks. Stronger
176
+ guarantees mean `WAIT`/`WAITAOF` on the lock write, or a different store. Document; don't solve
177
+ here.
178
+ - **Other records**: Step Result Records and map element results have the same shape:
179
+ - the dispatcher creates the record;
180
+ - the unit's job updates it under its liveness lock;
181
+ - it could be fenced the same way.
182
+
183
+ Nothing races there today, because the liveness lock drops concurrent duplicates. Decide
184
+ whether to fence them now or later.
185
+
186
+ ### Rollout
187
+
188
+ - **SemVer**: MINOR, behind `config.fenced_context_writes` (default `false`) for one release.
189
+ Existing contexts have no version key, which is read as `0`.
190
+ - A rolling deploy mixes old workers (plain `SET`, no version bump) with new ones, so the new
191
+ workers would see spurious `stale` results. Enable the flag only after every worker runs the new
192
+ version. Flip the default in the next MAJOR.
193
+ - **Cost**: one `EVAL` in place of one `SET` per save; checkpoints already save once per step.
194
+ One extra `SET NX` and release per synchronous run.
195
+
196
+ ### Test plan
197
+
198
+ All against real Redis, per the constitution.
199
+
200
+ 1. **Paused holder**:
201
+ - take the lock with a short TTL and let it expire;
202
+ - a second owner takes it and writes;
203
+ - the first owner's write returns `lost_lock`, and the stored blob is the second owner's.
204
+ 2. **Stale read**:
205
+ - load at version N;
206
+ - another writer (holding the lock through a handover) bumps the version to N+1;
207
+ - the write returns `stale`.
208
+ 3. **Lock, then load**: a redelivery that races a finishing holder always runs on the holder's
209
+ final state. This is the same repro shape as 005 P4.
210
+ 4. **One spec per writer** in the tables above, showing it goes through the fence, or takes the
211
+ lock and skips when it is held.
212
+ 5. **Cancel of a live run**: the request is consumed at the next step boundary. The running
213
+ execution's saves are never overwritten by the canceller.
214
+ 6. **Cluster slot co-location**, if Redis Cluster is supported: every key a script touches
215
+ hashes to one slot.
216
+
217
+ ### Size
218
+
219
+ - Storage adapter: one new script and a versioned `retrieve_context`.
220
+ - Executor and `Worker`: lock, then load, and pass the token.
221
+ - Six writer paths rerouted, plus the cancel/undo request key.
222
+ - Docs: `locks_and_semaphores.md`, `background_and_async.md`, and the README's "Durability &
223
+ Recovery" section.
224
+
225
+ ## Fan-out map inside a composed child
226
+
227
+ **Status:** bug, already present on main (0.8.1). Raised by the PR #56 review.
228
+
229
+ A root that composes a child, where the child runs a `fan_out` map, never finishes. The root
230
+ stays `running`.
231
+
232
+ - `MapStep#prepare_async_execution` stores the child under its own id.
233
+ - The map collector loads that child blob. Its `root_context` is nil, so the collector resumes
234
+ the child as a standalone execution. On a park, it requeues the child the same way.
235
+ - The child completes in its own blob. Nothing resumes the root, and the root's embedded copy of
236
+ the child stays at the map step.
237
+
238
+ To reproduce: `Root` composes `Child`, and `Child` runs a `map ... fan_out batch_size: 1`. Drain
239
+ the element, collector and `Worker` jobs. The root is still `running`.
240
+
241
+ A literal patch would have the collector write the root's blob. That breaks the single-writer
242
+ rule, because the collector holds only `map_collect:`, never the root's `async:` lock.
243
+
244
+ Direction:
245
+ - The collector records the map's completion on the map's own records, which it already does.
246
+ - The collector requeues the ROOT worker and does not resume the child itself.
247
+ - On re-entry, `MapStep` adopts the finished map: collect the results, don't dispatch again.
248
+
249
+ Spec: compose → fan-out map, with and without a park after the map. The root completes either
250
+ way.
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: ruby_reactor
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.8.2
4
+ version: 0.8.3
5
5
  platform: ruby
6
6
  authors:
7
7
  - Artur
@@ -98,6 +98,7 @@ files:
98
98
  - ".claude/skills/speckit-demo-tests/SKILL.md"
99
99
  - ".claude/skills/speckit-implement/SKILL.md"
100
100
  - ".claude/skills/speckit-plan/SKILL.md"
101
+ - ".claude/skills/speckit-review/SKILL.md"
101
102
  - ".claude/skills/speckit-specify/SKILL.md"
102
103
  - ".claude/skills/speckit-tasks/SKILL.md"
103
104
  - ".claude/skills/speckit-taskstoissues/SKILL.md"
@@ -175,8 +176,11 @@ files:
175
176
  - lib/ruby_reactor/error/dependency_error.rb
176
177
  - lib/ruby_reactor/error/deprecated_dsl_error.rb
177
178
  - lib/ruby_reactor/error/deserialization_error.rb
179
+ - lib/ruby_reactor/error/execution_parked.rb
178
180
  - lib/ruby_reactor/error/input_validation_error.rb
181
+ - lib/ruby_reactor/error/reactor_contention_park.rb
179
182
  - lib/ruby_reactor/error/schema_version_error.rb
183
+ - lib/ruby_reactor/error/step_contention_park.rb
180
184
  - lib/ruby_reactor/error/step_failure_error.rb
181
185
  - lib/ruby_reactor/error/undo_error.rb
182
186
  - lib/ruby_reactor/error/validation_error.rb
@@ -188,6 +192,7 @@ files:
188
192
  - lib/ruby_reactor/executor/ordered_lock_support.rb
189
193
  - lib/ruby_reactor/executor/result_handler.rb
190
194
  - lib/ruby_reactor/executor/retry_manager.rb
195
+ - lib/ruby_reactor/executor/step_coordination.rb
191
196
  - lib/ruby_reactor/executor/step_executor.rb
192
197
  - lib/ruby_reactor/interrupt_result.rb
193
198
  - lib/ruby_reactor/lock.rb
@@ -256,39 +261,14 @@ files:
256
261
  - lib/ruby_reactor/web/config.ru
257
262
  - lib/ruby_reactor/web/coordination_serializer.rb
258
263
  - lib/ruby_reactor/web/public/assets/index-BQvIWPdx.css
259
- - lib/ruby_reactor/web/public/assets/index-Dw4KV4QY.js
264
+ - lib/ruby_reactor/web/public/assets/index-CeZU-ESu.js
260
265
  - lib/ruby_reactor/web/public/index.html
261
266
  - lib/ruby_reactor/web/public/vite.svg
262
267
  - lib/ruby_reactor/worker.rb
263
268
  - llms-full.txt
264
269
  - llms.txt
265
270
  - sig/ruby_reactor.rbs
266
- - specs/002-step-input-contracts/checklists/requirements.md
267
- - specs/002-step-input-contracts/contracts/dsl-surface.md
268
- - specs/002-step-input-contracts/data-model.md
269
- - specs/002-step-input-contracts/plan.md
270
- - specs/002-step-input-contracts/quickstart.md
271
- - specs/002-step-input-contracts/research.md
272
- - specs/002-step-input-contracts/spec.md
273
- - specs/002-step-input-contracts/tasks.md
274
- - specs/004-inheritable-step-class/checklists/requirements.md
275
- - specs/004-inheritable-step-class/contracts/step-lifecycle.md
276
- - specs/004-inheritable-step-class/data-model.md
277
- - specs/004-inheritable-step-class/plan.md
278
- - specs/004-inheritable-step-class/quickstart.md
279
- - specs/004-inheritable-step-class/research.md
280
- - specs/004-inheritable-step-class/spec.md
281
- - specs/004-inheritable-step-class/tasks.md
282
- - specs/active_job.md
283
- - specs/deferred-003-step-lock-declarations/checklists/requirements.md
284
- - specs/deferred-003-step-lock-declarations/contracts/dsl-surface.md
285
- - specs/deferred-003-step-lock-declarations/data-model.md
286
- - specs/deferred-003-step-lock-declarations/plan.md
287
- - specs/deferred-003-step-lock-declarations/quickstart.md
288
- - specs/deferred-003-step-lock-declarations/research.md
289
- - specs/deferred-003-step-lock-declarations/spec.md
290
- - specs/deferred-003-step-lock-declarations/tasks.md
291
- - specs/possible_feature.md
271
+ - specs/future_improvements.md
292
272
  - specs/scoped_rspec_matchers.md
293
273
  - teley/Dockerfile
294
274
  homepage: https://github.com/arturictus/ruby_reactor
@@ -1,49 +0,0 @@
1
- # Specification Quality Checklist: Step Input Contracts
2
-
3
- **Purpose**: Validate specification completeness and quality before proceeding to planning
4
- **Created**: 2026-09-10
5
- **Feature**: [spec.md](../spec.md)
6
-
7
- ## Content Quality
8
-
9
- - [x] No implementation details (languages, frameworks, APIs)
10
- - [x] Focused on user value and business needs
11
- - [x] Written for non-technical stakeholders
12
- - [x] All mandatory sections completed
13
-
14
- ## Requirement Completeness
15
-
16
- - [x] No [NEEDS CLARIFICATION] markers remain
17
- - [x] Requirements are testable and unambiguous
18
- - [x] Success criteria are measurable
19
- - [x] Success criteria are technology-agnostic (no implementation details)
20
- - [x] All acceptance scenarios are defined
21
- - [x] Edge cases are identified
22
- - [x] Scope is clearly bounded
23
- - [x] Dependencies and assumptions identified
24
-
25
- ## Feature Readiness
26
-
27
- - [x] All functional requirements have clear acceptance criteria
28
- - [x] User scenarios cover primary flows
29
- - [x] Feature meets measurable outcomes defined in Success Criteria
30
- - [x] No implementation details leak into specification
31
-
32
- ## Notes
33
-
34
- - All items pass. 0 [NEEDS CLARIFICATION] markers remain.
35
- - Resolved with the user on 2026-09-10:
36
- - **Naming**: a unit of work declares `input`; the reactor keeps `argument` for wiring only
37
- (FR-012).
38
- - **Undeclared arguments**: reactor-load error for contract-owning steps (FR-018); steps
39
- with no contract and no wiring keep today's pass-all-inputs behavior (FR-019).
40
- - **Implicit wiring**: an unwired step input resolves from a same-named reactor input only,
41
- checked at reactor-load time; never from another step's result (FR-020, FR-021).
42
- - **Direct invocation**: the contract is enforced on every entry point, not only via a
43
- reactor (FR-022).
44
- - **Falsey values**: presence means "a value was supplied", never "the value is truthy";
45
- the pre-existing loss of `false` during argument resolution is corrected as part of this
46
- feature (FR-023, SC-011).
47
- - "Non-technical stakeholders" is read as *developers who are not this library's
48
- maintainers*: the spec names no Ruby constructs, gems, or file paths.
49
- - Ready for `/speckit-plan`.