redis_queued_locks 1.16.2 → 1.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. checksums.yaml +4 -4
  2. data/.claude/project-overview.md +205 -0
  3. data/.claude/rules/acquirer.md +61 -0
  4. data/.claude/rules/arguments.md +59 -0
  5. data/.claude/rules/logic.md +57 -0
  6. data/.claude/rules/swarm.md +119 -0
  7. data/.claude/rules/tests.md +42 -0
  8. data/.claude/rules/type-checking.md +42 -0
  9. data/.claude/rules/visitors.md +64 -0
  10. data/.rubocop.rbs.yml +31 -0
  11. data/.rubocop.yml +3 -3
  12. data/.ruby-version +1 -1
  13. data/CHANGELOG.md +18 -0
  14. data/CLAUDE.md +31 -0
  15. data/README.md +19 -3
  16. data/Rakefile +31 -2
  17. data/github_ci/.keep +0 -0
  18. data/lib/redis_queued_locks/acquirer/acquire_lock/delay_execution.rb +2 -2
  19. data/lib/redis_queued_locks/acquirer/acquire_lock/try_to_lock.rb +1 -4
  20. data/lib/redis_queued_locks/acquirer/acquire_lock/yield_expire.rb +1 -1
  21. data/lib/redis_queued_locks/acquirer/acquire_lock.rb +0 -4
  22. data/lib/redis_queued_locks/acquirer/lock_series_poc/instr_visitor.rb +0 -2
  23. data/lib/redis_queued_locks/acquirer/lock_series_poc/log_visitor.rb +0 -1
  24. data/lib/redis_queued_locks/acquirer/lock_series_poc.rb +0 -2
  25. data/lib/redis_queued_locks/client.rb +145 -144
  26. data/lib/redis_queued_locks/config/dsl.rb +4 -10
  27. data/lib/redis_queued_locks/swarm/flush_zombies.rb +8 -6
  28. data/lib/redis_queued_locks/swarm/supervisor.rb +1 -1
  29. data/lib/redis_queued_locks/swarm/swarm_element/isolated.rb +166 -44
  30. data/lib/redis_queued_locks/swarm/swarm_element/threaded.rb +45 -47
  31. data/lib/redis_queued_locks/version.rb +2 -2
  32. data/redis_queued_locks.gemspec +1 -1
  33. data/sig/redis_queued_locks/swarm/flush_zombies.rbs +1 -1
  34. data/sig/redis_queued_locks/swarm/swarm_element/isolated.rbs +13 -5
  35. data/sig/redis_queued_locks/swarm/swarm_element/threaded.rbs +6 -6
  36. metadata +14 -5
  37. data/github_ci/ruby3.3.gemfile +0 -17
  38. data/github_ci/ruby3.3.gemfile.lock +0 -216
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: d49a1559431c569e92cfc1a9b56bd04d9942002aa6d79efcb533370f58b0c184
4
- data.tar.gz: bc8fd235df24386a92275273de91d34804452eb085bbb5a2b7118eb4508319ab
3
+ metadata.gz: de70554e8788c12aaa830c8eca99f7b31f936836a3f664498299d213ba2b6036
4
+ data.tar.gz: fa8bc015bb930d9ccef934ab6884dd1ef930cf2b540a98ab221f78b1a06afac8
5
5
  SHA512:
6
- metadata.gz: 0e7d0944406f13a7a6bac3e59153998fca9731e184ef5da83d8d415af7185f61a4b260166895d3858c501f4ae93cf160fb6b4c5c33e248d2aa82c9bac2265216
7
- data.tar.gz: 8aff1a1fb28cbc510edc30efbab0eed49bbb088feff6235959b314b24d705e0058c6d1a347356874fa2a4b930f7aaba62e7f6b48bebf2622e11600998055f719
6
+ metadata.gz: e7ca083df52b3e53878512da62a8a515fe0bdf94f58d14e4d62db6facf1684b481597d5618f3d46294051774f0a2f076d7d518c0b848eedaefef0fca22158bd9
7
+ data.tar.gz: 1d3e8c5660334ef5925ad951408d262922a4786dce98d539ad9318edce240b9ffa47093d1a1fffa2fb8dc431f7be8d2ed6d0fe565d53743e5f5177b5750c42b6
@@ -0,0 +1,205 @@
1
+ # redis_queued_locks: Project Overview
2
+
3
+ Ruby gem providing **distributed locks on Redis with a prioritized lock-acquisition queue**.
4
+ Each lock has its own request queue; requests are processed FIFO (`:queued`) or first-come
5
+ (`:random`). Requests have TTLs (with requeue), so queues never get stuck. Includes reentrancy
6
+ (conflict) strategies, timed locks, retries/fail-fast, logging, instrumentation with sampling,
7
+ and a background "swarm" that detects dead hosts and flushes their locks.
8
+
9
+ - Repo: https://github.com/0exp/redis_queued_locks (MIT); main branch `master`
10
+ - Runtime dependency: `redis-client ~> 0.20` (locked 0.30.1), nothing else
11
+ - Ruby `>= 4.0` (`.ruby-version` = 4.0.7)
12
+
13
+ ---
14
+
15
+ ## 1. Architecture
16
+
17
+ ```
18
+ RedisQueuedLocks (lib/redis_queued_locks.rb) entry point; requires all parts; extends Debugger::Interface
19
+ │
20
+ ▼
21
+ Client (client.rb) public API facade
22
+ ├── Config (config.rb + config/dsl.rb) declarative settings + validators
23
+ ├── Swarm (swarm.rb + swarm/*) supervised background workers
24
+ └── Acquirer::* modules one stateless module per operation
25
+ │
26
+ ▼
27
+ RedisClient (redis-client gem) WATCH/MULTI, Lua (EVAL), ZSET, HASH, SCAN
28
+ ```
29
+
30
+ `Client.new(redis_client) { |config| ... }`:
31
+ 1. builds `Config` from the block,
32
+ 2. computes `uniq_identity` via `config['uniq_identifier']`,
33
+ 3. creates `Swarm` and starts it when `config['swarm.auto_swarm']` is true.
34
+
35
+ Every other public method is a thin wrapper: it fills option defaults from `config[...]` and calls
36
+ an `Acquirer::*` module function with `redis_client`.
37
+
38
+ ### Client → implementation map
39
+
40
+ | Client method | Implementation |
41
+ |---|---|
42
+ | `lock` / `lock!` | `acquirer/acquire_lock.rb` + `acquirer/acquire_lock/*` |
43
+ | `lock_series` / `lock_series!` | `acquirer/lock_series_poc.rb` (proof of concept) |
44
+ | `unlock` | `acquirer/release_lock.rb` |
45
+ | `extend_lock_ttl` | `acquirer/extend_lock_ttl.rb` |
46
+ | `locked?` / `queued?` | `acquirer/is_locked.rb` / `acquirer/is_queued.rb` |
47
+ | `lock_info` / `queue_info` | `acquirer/lock_info.rb` / `acquirer/queue_info.rb` |
48
+ | `locks`, `locks_info`, `queues`, `queues_info`, `keys` | `acquirer/locks.rb`, `queues.rb`, `keys.rb` (SCAN-based) |
49
+ | `clear_locks` | `acquirer/release_all_locks.rb` |
50
+ | `clear_locks_of`, `clear_current_locks` | `acquirer/release_locks_of.rb` |
51
+ | `clear_dead_requests` | `acquirer/clear_dead_requests.rb` |
52
+ | `current_acquirer_id`, `current_host_id`, `possible_host_ids` | `resource.rb` |
53
+ | `swarmize!`, `deswarmize!`, `swarm_status`, `swarm_info`, `probe_hosts`, `flush_zombies`, `zombie_locks`, `zombie_acquirers`, `zombie_hosts`, `zombies_info` | `swarm.rb` + `swarm/*` |
54
+
55
+ ### Lock acquisition flow (`acquirer/acquire_lock.rb` + `acquire_lock/`)
56
+
57
+ `AcquireLock` is assembled from mixins (`extend`):
58
+
59
+ | Mixin | Role |
60
+ |---|---|
61
+ | `TryToLock` (`try_to_lock.rb`) | One attempt: `ZADD NX` acquirer into the lock queue (timestamp score), then `multi(watch: [lock_key])` and take the lock if allowed (head of queue for `:queued`, any position for `:random`). Handles reentrant conflicts; TTL extension via inline Lua (`PTTL` + `PEXPIRE`). |
62
+ | `WithAcqTimeout` | Global acquisition timeout (`timeout`). |
63
+ | `DelayExecution` | Retry delay + jitter between attempts. |
64
+ | `DequeueFromLockQueue` | Removes the acquirer from the queue on timeout/failure. |
65
+ | `YieldExpire` | Runs the user block (optionally `timed`), then releases/expires the lock. |
66
+ | `LogVisitor` / `InstrVisitor` | One method per lifecycle event for logs and instrumentation. |
67
+
68
+ ### Main `lock` options (defaults from config)
69
+
70
+ `ttl`, `queue_ttl`, `timeout`, `timed`, `retry_count`, `retry_delay`, `retry_jitter`,
71
+ `raise_errors`, `fail_fast`, `conflict_strategy`, `access_strategy`, `read_write_mode`
72
+ (default `:write`; its doc is unfinished), `identity`, `meta`, `logger`, `log_lock_try`,
73
+ `instrumenter`, `instrument`, log/instr sampling options, `log_sample_this`, `instr_sample_this`.
74
+
75
+ - `access_strategy`: `:queued` (FIFO, default) or `:random`.
76
+ - `conflict_strategy` (same process re-acquires its own lock): `:wait_for_lock` (default),
77
+ `:work_through`, `:extendable_work_through`, `:dead_locking`.
78
+
79
+ ### Redis data layout (`resource.rb`)
80
+
81
+ | Key | Type | Purpose |
82
+ |---|---|---|
83
+ | `rql:lock:<name>` | HASH | lock owner + metadata (e.g. `l_spc_ts`, `l_spc_ext_ts` for reentrant cases) |
84
+ | `rql:lock_queue:<name>` | ZSET | acquirer queue scored by enqueue time |
85
+ | `rql:lock_queue:<name>:read` / `:write` | ZSET | read/write mode queues |
86
+ | `rql:swarm:hsts` | HASH | swarm host heartbeats |
87
+
88
+ - Acquirer ID: `rql:acq:<pid>/<thread>/<fiber>/<ractor>/<identity>`
89
+ - Host ID: `rql:hst:<pid>/<thread>/<ractor>/<identity>` (no fiber)
90
+
91
+ ### Instrumentation events
92
+
93
+ `redis_queued_locks.` + `lock_obtained`, `reentrant_lock_obtained`,
94
+ `extendable_reentrant_lock_obtained`, `lock_hold_and_release`, `reentrant_lock_hold_completes`,
95
+ `lock_series_obtained`, `lock_series_hold_and_release`, `explicit_lock_release`,
96
+ `explicit_all_locks_release`, `release_locks_of`.
97
+
98
+ ### Errors (`errors.rb`)
99
+
100
+ Base `RedisQueuedLocks::Error < StandardError`, plus `ArgumentError`, `LockAlreadyObtainedError`,
101
+ `LockAcquirementTimeoutError`, `LockAcquirementRetryLimitError`, `TimedLockTimeoutError`,
102
+ `ConflictLockObtainError`, `SwarmError`, `SwarmArgumentError`, `ConfigError`
103
+ (`ConfigNotFoundError`, `ConfigValidationError`). Internal `*IntermediateTimeoutError`s
104
+ inherit from `Timeout::Error`.
105
+
106
+ ---
107
+
108
+ ## 2. Architecture patterns by module
109
+
110
+ | Module | Patterns |
111
+ |---|---|
112
+ | `Client` | Facade; constructor injection (caller-supplied `RedisClient`); `x` returns a result hash, `x!` raises |
113
+ | `Acquirer::*` | One module per command; stateless `class << self` functions; uniform `{ ok:, result: }` returns (public API contract; swarm element internals use bare scalars/primitives) |
114
+ | `AcquireLock` | Composition via `extend` mixins; optimistic concurrency (WATCH/MULTI) + Lua for atomic updates; strategy options (`access_strategy`, `conflict_strategy`) |
115
+ | Log/Instr visitors | Visitor-style event hooks keep observability out of the algorithm; percent sampling via `sampling_happened?(percent)` |
116
+ | `Logging` / `Instrument` | Null Object defaults (`VoidLogger`, `VoidNotifier`); Adapter (`instrument/active_support.rb`); duck typing (`::Logger` API, `#notify(event, payload)`) |
117
+ | `Config` | Declarative DSL: `setting(key, default)` and `validate(key) { }` registries; access as `config['a.b']`; runtime `Client#configure` |
118
+ | `Swarm` | Template base classes `SwarmElement::Threaded` (Thread) and `SwarmElement::Isolated` (Ractor; commands/replies via a per-element pair of `Ractor::Port`s, termination via `Ractor#monitor`). `ProbeHosts < Threaded`; `FlushZombies < Isolated` with its own connection from `RedisClientBuilder` (plain, pooled or sentinel). `Supervisor` watchdog thread restarts dead elements. |
119
+ | `Resource` | Single source of truth for key names and identity strings |
120
+ | Misc | `Data < Hash` result object; `Utilities::Lock` mutex wrapper; global `Debugger` toggle |
121
+
122
+ ---
123
+
124
+ ## 3. Configuration (`config.rb`)
125
+
126
+ | Group | Keys (default) |
127
+ |---|---|
128
+ | Retry/timeouts | `retry_count` (3), `retry_delay` (200 ms), `retry_jitter` (25 ms), `try_to_lock_timeout` (10 s), `is_timed_by_default` (false) |
129
+ | TTLs | `default_lock_ttl` (5000 ms), `default_queue_ttl` (15 s), `dead_request_ttl` (1 day, ms) |
130
+ | Strategies | `default_conflict_strategy` (`:wait_for_lock`), `default_access_strategy` (`:queued`) |
131
+ | Batching | `lock_release_batch_size` (100), `clear_locks_of__lock_scan_size` / `__queue_scan_size` (300), `key_extraction_batch_size` (500) |
132
+ | Identity | `uniq_identifier` (lambda → `Resource.calc_uniq_identity`) |
133
+ | Logging | `logger` (VoidLogger), `log_lock_try`, `log_sampling_enabled`, `log_sampling_percent` (15), `log_sampler` |
134
+ | Instrumentation | `instrumenter` (VoidNotifier), `instr_sampling_enabled`, `instr_sampling_percent` (15), `instr_sampler` |
135
+ | Errors | `detailed_acq_timeout_error` (false) |
136
+ | Swarm | `swarm.auto_swarm` (false), `swarm.supervisor.liveness_probing_period` (2 s), `swarm.probe_hosts.*` (enabled, `probe_period` 2 s, `redis_config.*`), `swarm.flush_zombies.*` (enabled, `zombie_ttl` 15000 ms, scan sizes 500, `zombie_flush_period` 10, `redis_config.*`) |
137
+
138
+ ---
139
+
140
+ ## 4. Directory map
141
+
142
+ ```
143
+ lib/redis_queued_locks.rb entry point
144
+ lib/redis_queued_locks/
145
+ client.rb public API
146
+ config.rb, config/dsl.rb settings DSL + defaults
147
+ acquirer/*.rb one operation per file
148
+ acquirer/acquire_lock/* lock algorithm mixins + visitors
149
+ swarm.rb, swarm/* supervisor, swarm elements, zombie logic, redis client builder
150
+ logging/, instrument/ void defaults, samplers, ActiveSupport adapter
151
+ resource.rb keys and identities
152
+ data.rb, errors.rb, utilities.rb, utilities/lock.rb, debugger/, version.rb
153
+ sig/ RBS mirror of lib/ + sig/vendor stubs (redis_client, active_support, semantic_logger)
154
+ spec/ redis_queued_locks_spec.rb (~2.6k lines, integration), spec_helper.rb, setup_simplecov.rb
155
+ .github/workflows/ tests, lint, typecheck-static, typecheck-runtime
156
+ bin/console, bin/setup dev scripts
157
+ Rakefile, Steepfile, rbs_collection.yaml
158
+ ```
159
+
160
+ Known quirk: `sig/redis_queued_locks/acquier.rbs` is misspelled (should be `acquirer.rbs`).
161
+
162
+ ---
163
+
164
+ ## 5. Technology stack
165
+
166
+ ### Runtime
167
+ - Ruby >= 4.0; stdlib `timeout`, `securerandom`, `logger`; Thread, Fiber, Ractor, Mutex
168
+ - Redis server: WATCH/MULTI, EVAL (Lua), ZSET, HASH, PTTL/PEXPIRE, SCAN
169
+ - `redis-client ~> 0.20`
170
+ - Optional, duck-typed: any `::Logger`-compatible logger; any `#notify(event, payload)` instrumenter (ActiveSupport adapter included; SemanticLogger has only an RBS stub)
171
+
172
+ ### Test / development
173
+ - RSpec 3.13: random order, no monkey patching, `expect` syntax, `Thread.abort_on_exception = true`
174
+ - rspec-retry 0.6.2: 5 retries per example (marked temporary)
175
+ - A real Redis is required (CI: `supercharge/redis-github-action@1.8.1`)
176
+ - SimpleCov 1.3.2: line + branch coverage, HTML report; 100% minimum disabled (TODO)
177
+ - RBS 4.2 + `rbs collection` + tsort; Steep 2.1 (static); RBS runtime testing via `rbs/test/setup`
178
+ - RuboCop 1.91 through `armitage-rubocop`, in two separate runs so Ruby cops (and their rubydex project index, `AllCops/UseProjectIndex`) never see RBS files:
179
+ - `.rubocop.yml`: Ruby sources (general, rake, rspec presets; plugins rubocop-rspec, -performance, -rake, -thread_safety); several Metrics cops disabled
180
+ - `.rubocop.rbs.yml`: RBS cops only (rbs preset, plugin rubocop-on-rbs, no project index): `RBS/*` on `sig/**/*.rbs` and `RBSInline/*` on inline annotations in `lib/**/*.rb`; Ruby cop departments are excluded explicitly because inline `# rubocop:enable` directives would re-enable them despite `DisabledByDefault`
181
+ - rake, bundler, pry, pry-doc, reline, activesupport 8.1
182
+
183
+ ### Rake tasks
184
+ `rspec` (default), `rubocop` (runs `rubocop:ruby` and `rubocop:rbs`, fails if either fails), `steep:check`, plus bundler gem tasks (`build`, `release`, ...).
185
+
186
+ ### CI (GitHub Actions, ubuntu-latest, Ruby 4.0, on every push)
187
+
188
+ | Workflow | Command |
189
+ |---|---|
190
+ | tests | `bundle exec rake rspec` (with a Redis service) |
191
+ | lint | `bundle exec rake rubocop` |
192
+ | typecheck-static | `bundle exec rbs collection install && bundle exec rake steep:check -j 10` |
193
+ | typecheck-runtime | `RBS_TEST_RAISE=true RUBYOPT='-rrbs/test/setup' RBS_TEST_OPT='-I sig' RBS_TEST_TARGET='RedisQueuedLocks::*' bundle exec rspec --failure-exit-code=0` (never fails the build) |
194
+
195
+ ---
196
+
197
+ ## 6. Conventions
198
+
199
+ - `# frozen_string_literal: true` in every Ruby file.
200
+ - YARD tags on every class/method: `@api public|private`, `@since`, `@version` (bump `@version` when changing behavior).
201
+ - Every `lib/` change gets a matching `sig/` RBS update; use `# steep:ignore` only where Steep can't infer types.
202
+ - New operation: new `Acquirer::*` module, thin `Client` method (+ `!` variant if it should raise), RBS, spec.
203
+ - New config option: `setting` (+ `validate`) in `config.rb`, read via `config['key']`, document defaults.
204
+ - Development gems go in `Gemfile` (`Gemspec/DevelopmentDependencies: Gemfile`), not the gemspec.
205
+ - Temporary debt: rspec-retry, disabled coverage minimum, runtime type-check job that can't fail.
@@ -0,0 +1,61 @@
1
+ ---
2
+ paths:
3
+ - "lib/redis_queued_locks/acquirer.rb"
4
+ - "lib/redis_queued_locks/acquirer/**/*.rb"
5
+ ---
6
+
7
+ # Acquirer modules (`lib/redis_queued_locks/acquirer/**`)
8
+
9
+ ## Observed style and patterns
10
+ - **Module shape**: `module RedisQueuedLocks::Acquirer::<CamelName>` in `acquirer/<snake_name>.rb`, `# @api private`, a single public entry function inside `class << self` named after the file (`ReleaseLock.release_lock`, `IsLocked.locked?`), helpers under `private` in the same `class << self`.
11
+ - **Utilities**: modules that time or instrument `extend RedisQueuedLocks::Utilities` (gives `clock_gettime`, `run_non_critical`).
12
+ - **Signatures** (see `arguments.md` for the design rationale):
13
+ - Read-only queries: few positional args `(redis_client, lock_name)`; collection queries use keywords `(redis_client, scan_size:, with_info:)`.
14
+ - Mutating operations: long positional lists ending with the fixed observability tail
15
+ `logger, instrumenter, instrument, log_sampling_enabled, log_sampling_percent, log_sampler, log_sample_this, instr_sampling_enabled, instr_sampling_percent, instr_sampler, instr_sample_this`
16
+ (order of `logger`/`instrumenter` varies between modules; check the existing signature).
17
+ - Only `AcquireLock.acquire_lock` uses keyword args (`process_id:`, `thread_id:`, ...).
18
+ - **Two-layer mutating operations** (`release_lock`, `release_all_locks`, `release_locks_of`):
19
+ 1. `rel_start_time = clock_gettime`
20
+ 2. call a private `fully_*` helper returning `{ ok:, result: }` and destructure it: `fully_x(...) => { ok:, result: }`
21
+ 3. `time_at = Time.now.to_f`; `rel_time = ((rel_end_time - rel_start_time) / 1_000.0).ceil(2)` (microseconds → ms)
22
+ 4. `instr_sampled = RedisQueuedLocks::Instrument.should_instrument?(...)`
23
+ 5. `run_non_critical { instrumenter.notify('redis_queued_locks.<event>', { at:, rel_time:, ... }) } if instr_sampled`
24
+ 6. return `{ ok: true, result: { ..., rel_time: } }`
25
+ - **Results**: always `{ ok: Boolean, result: ... }` for operations; `result` is a Symbol status (`:ttl_extended`, `:async_expire_or_no_lock`, `:released`, `:nothing_to_release`) or a Symbol-keyed Hash with abbreviated keys (`rel_key_cnt`, `tch_queue_cnt`, `rel_time`). Info queries return a String-keyed Hash / Set or `nil` when absent.
26
+ - **Redis access**:
27
+ - Keys only via `Resource.prepare_lock_key` / `prepare_lock_queue` and `Resource::*_PATTERN`.
28
+ - Raw commands: `redis.call('CMD', ...)` with uppercase string command names and string args (`'0'`, `'-inf'`, `'+inf'`).
29
+ - Pooled connection: wrap multi-command work in `redis.with do |rconn| ... end`.
30
+ - Atomic writes: `rconn.multi do |transact| ... end` (or `multi(watch: [lock_key])` in `TryToLock`); batch reads: `pipelined do |pipeline| ... end` then index `result[0]`, `result[1]` into named vars (`hget_cmd_res`, `pttl_cmd_res`).
31
+ - Iteration: `scan('MATCH', PATTERN, count: scan_size) { |key| ... }`, collecting into `Set.new.tap { |set| ... }`; deletes are batched by scan size.
32
+ - Lua: frozen heredoc constant (`<<~LUA_SCRIPT.strip.tr("\n", '').freeze`) + `call('EVAL', SCRIPT, 1, key, arg)`.
33
+ - Release = `EXPIRE key 0` / `ZREMRANGEBYSCORE queue -inf +inf`; Redis TTL sentinels handled explicitly (`PTTL` `-2` = missing, `-1` = no expiry → `Float::INFINITY`).
34
+ - **Data normalization**: Redis hash strings are converted with `Float(...)` / `Integer(...)` inside `hget_cmd_res.tap do |lock_data| ... end`, optional fields guarded with `if lock_data['x']`.
35
+ - **AcquireLock**:
36
+ - Main module `require_relative`s its parts, then `extend`s the mixins (`TryToLock`, `DelayExecution`, `YieldExpire`, `WithAcqTimeout`, `DequeueFromLockQueue`); mixins are plain modules with instance methods (`def try_to_lock(...)`), visitors are `class << self` modules called explicitly (see `visitors.md`).
37
+ - The algorithm is a numbered step script (`# Step 0`, `# Step 2.1`, `# Step 2.2.a`) driven by a mutable `acq_process` hash (`:should_try`, `:tries`, `:acquired`, `:result`, `:lock_info`, timings).
38
+ - Failure modes are Symbols (`:fail_fast_no_try`, `:fail_fast_after_try`, `:conflict_dead_lock`, ...); exceptions are raised only when `raise_errors` is true, with a message naming the lock key / acquirer id.
39
+ - Timing via monotonic `clock_gettime` (microseconds); wall time only for `at:`/`ts` fields.
40
+ - **Inline typing**: `# @type var result: Symbol`, `{} #: Hash[String,String|Float|Integer]`, and `# steep:ignore` on pattern-matching destructures and splats (`rconn.call('DEL', *keys) # steep:ignore`).
41
+
42
+ ## Claude rules
43
+ 1. New operation: create `acquirer/<snake_name>.rb` with `module RedisQueuedLocks::Acquirer::<CamelName>`, one public `class << self` entry function named after the file, private helpers below `private`; add `require_relative` to `acquirer.rb` and a mirrored `sig/redis_queued_locks/acquirer/<snake_name>.rbs`.
44
+ 2. Read-only query → small positional args or `scan_size:` / `with_info:` keywords, return data or `nil`. Mutating operation → follow the two-layer pattern (public timing/instrumentation wrapper + private `fully_*` helper).
45
+ 3. Keep the observability tail parameters in the established order of the closest sibling module; pass them through unchanged from `Client`.
46
+ 4. Return `{ ok:, result: }`; use Symbol statuses for simple outcomes and abbreviated Symbol keys (`rel_*`, `*_cnt`, `*_time`) consistent with existing results. Don't raise from acquirer modules except through the `raise_errors` path in `AcquireLock`.
47
+ 5. Use `redis.with { |rconn| ... }` for any multi-command work; use `multi` for writes that must be atomic, `pipelined` for independent reads, `multi(watch:)` or Lua when a write depends on a read.
48
+ 6. Use uppercase string Redis commands and string numeric args; handle `PTTL`/`TTL` sentinel values explicitly.
49
+ 7. Iterate keys with `scan('MATCH', Resource::*_PATTERN, count:)`, never `KEYS`; batch deletes by scan size.
50
+ 8. Measure durations with `clock_gettime` and report ms with `/ 1_000.0).ceil(2)`; use `Time.now.to_f` only for event timestamps.
51
+ 9. Wrap every `instrumenter.notify` / logger call in `run_non_critical` (or a visitor) and gate it with `Instrument.should_instrument?` / `Logging.should_log?`; observability must never break locking.
52
+ 10. In `AcquireLock`, add behavior as a new step or mixin rather than growing `acquire_lock`; keep the `# Step N.x` comment numbering, update `acq_process` keys consistently, and add a matching `LogVisitor`/`InstrVisitor` method for each new lifecycle event.
53
+ 11. When normalizing lock hash fields, follow the `Float()` / `Integer()` conversion pattern; if a new lock field is added, update both `lock_info.rb` and `locks.rb` (they duplicate the formatting).
54
+
55
+ ## Recommendations (proposed, not yet project policy)
56
+ Apply to new or touched code; don't refactor existing code for these unless asked.
57
+ 1. Load Lua scripts once (`SCRIPT LOAD` + `EVALSHA`, falling back to `EVAL` on `NOSCRIPT`), as the TODO in `extend_lock_ttl.rb` asks.
58
+ 2. Prefer indexed lookups over full `SCAN` loops for new features (see TODOs in `release_locks_of.rb`, `locks.rb`, `queues.rb`).
59
+ 3. Treat `lock_series_poc.rb` as experimental: don't build new features on it without asking.
60
+ 4. Extract the duplicated lock-hash normalization in `lock_info.rb` / `locks.rb` and queue formatting in `queue_info.rb` / `queues.rb` into shared helpers (their TODOs ask for this).
61
+ 5. Unify the observability tail order (`logger, instrumenter` vs `instrumenter, logger` in `release_lock`) when those signatures are next changed.
@@ -0,0 +1,59 @@
1
+ ---
2
+ paths:
3
+ - "lib/redis_queued_locks/client.rb"
4
+ - "lib/redis_queued_locks/acquirer/**/*.rb"
5
+ - "lib/redis_queued_locks/swarm.rb"
6
+ - "lib/redis_queued_locks/config.rb"
7
+ - "sig/redis_queued_locks/client.rbs"
8
+ - "sig/redis_queued_locks/acquirer/**/*.rbs"
9
+ ---
10
+
11
+ # Long explicit keyword lists (core API design principle)
12
+
13
+ ## The pattern
14
+ The public API and the lock pipeline use **long, flat, explicit argument lists**: every tunable is a named keyword with its default declared directly in the signature (usually `config['...']`), and every value is forwarded explicitly, name by name, through each layer:
15
+
16
+ ```ruby
17
+ # Client (public DSL surface)
18
+ def lock(lock_name, ttl: config['default_lock_ttl'], queue_ttl: config['default_queue_ttl'],
19
+ ..., instr_sample_this: false, &block)
20
+ RedisQueuedLocks::Acquirer::AcquireLock.acquire_lock(
21
+ redis_client, lock_name,
22
+ process_id: RedisQueuedLocks::Resource.get_process_id, ...,
23
+ ttl:, queue_ttl:, ..., instr_sample_this:, &block
24
+ )
25
+ end
26
+ ```
27
+
28
+ This is a deliberate framework-level decision, **not** something to refactor into parameter objects, option hashes, context structs or builders.
29
+
30
+ ## Why (framework-engineering rationale)
31
+ - **Minimal object allocation on the hot path.** Lock acquisition runs in tight retry loops under contention. Keywords forwarded as `key:` shorthand and positional args into stateless `class << self` functions cost no extra objects; a parameter/context object, `**opts` hash, `Data`/`Struct` or builder allocates on every call (and per retry, per event). Fewer allocations = less GC pressure and more predictable latency in a concurrency primitive.
32
+ - **The signature is the DSL.** `client.lock('name', ttl: 1_000, fail_fast: true, conflict_strategy: :work_through) { ... }` reads as a declarative per-call configuration. Users override exactly what they need; everything else falls back to `config[...]`. There is one place to look for what a call accepts.
33
+ - **Layered defaults without merging.** Global config (`Config` DSL) → per-call keyword override → explicit forwarding. No hash merging, no hidden `opts.fetch`, no precedence rules to learn.
34
+ - **Fail-fast on typos.** Ruby rejects unknown keywords (`ArgumentError: unknown keyword`) at the call boundary; an options hash would silently accept `tll:`.
35
+ - **Typed and documented contract.** Every option has its own YARD `@option` line and RBS signature entry, so Steep and RBS runtime checks validate each value individually; opaque bags of options can't be typed this precisely.
36
+ - **Stateless, thread/Ractor-friendly internals.** No intermediate object carries mutable state between layers; each function receives everything it needs as plain values, which keeps `Acquirer::*` modules pure and shareable.
37
+ - **Grep-ability and traceability.** Searching for `log_sample_this:` shows every hop of a value from the public API down to the visitor that uses it.
38
+
39
+ ## How it is applied (layer by layer)
40
+ | Layer | Style |
41
+ |---|---|
42
+ | `Client` public methods | Positional subject (`lock_name`) + long keyword list with defaults from `config['...']` or literals; trailing `&block`. Paired `x` / `x!` methods duplicate the full keyword list (no `**kwargs` delegation). |
43
+ | `Client` → `Acquirer` | Explicit forwarding with Ruby 3.1 shorthand (`ttl:`, `queue_ttl:`); runtime identity values injected here (`process_id:`, `thread_id:`, `fiber_id:`, `ractor_id:`). |
44
+ | `AcquireLock.acquire_lock` | Required keywords without defaults (`ttl:`, `timeout:`, ...): defaults live only in `Client`. |
45
+ | Other `Acquirer::*`, mixins, visitors | Positional parameters in a fixed, documented order (subject, data, observability tail) for the cheapest possible internal calls. |
46
+ | Results | Public API (`Client`, `Acquirer::*`, `Swarm` facade actions): small literal Hashes (`{ ok:, result: }`), not result classes. Internal swarm element APIs: bare scalars/primitives (see `swarm.md`). |
47
+
48
+ ## Claude rules
49
+ 1. **Do not introduce** parameter objects, context/options structs (`Data.define`, `Struct`, `OpenStruct`), builders, or `**opts` / `options = {}` hashes in `Client`, `Acquirer::*`, mixins or visitors. Long explicit lists are the intended design.
50
+ 2. New public option: add it as an explicit keyword to **every** affected `Client` method (`x` and `x!`, e.g. `lock` and `lock!`), with the default taken from `config['...']` (add the `setting` + `validate` first) or a literal for per-call-only flags (`raise_errors: false`, `meta: nil`).
51
+ 3. Forward it explicitly by name through every layer using shorthand (`new_option:`); never forward via `**kwargs`, `...`, or `method(__method__).parameters` tricks.
52
+ 4. Keep defaults only at the public boundary (`Client`); internal functions take required keywords/positionals with no defaults so a missed forward fails loudly.
53
+ 5. Internal positional lists: append new params in the established order (subject → domain data → observability tail `logger, instrumenter, instrument, log_sampling_*, log_sample_this, instr_sampling_*, instr_sample_this`) and update every call site and the RBS signature in the same change.
54
+ 6. Document each new option with its own `@option`/`@param` YARD line (type, meaning, default source) and add it to the RBS method signature as a named param/keyword.
55
+ 7. Keep hot-path code allocation-light: no per-call wrapper objects, no hash merging, no splats to build args, no `tap`/closures purely for argument plumbing; prefer passing existing locals.
56
+ 8. Silence length cops locally (`# rubocop:disable Metrics/MethodLength`) rather than shortening signatures; `Metrics/ParameterLists` is disabled project-wide on purpose.
57
+
58
+ ## Recommendations (proposed, not yet project policy)
59
+ 1. For new internal functions with many same-typed params (e.g. several Integers/Booleans in a row), consider required keywords instead of positionals to prevent misordering; keep the list flat and explicit.
@@ -0,0 +1,57 @@
1
+ ---
2
+ paths:
3
+ - "lib/**/*.rb"
4
+ ---
5
+
6
+ # Logic rules: general (`lib/`)
7
+
8
+ Related rule files (loaded for narrower paths):
9
+ - `acquirer.md`: acquirer operations, Redis access, `AcquireLock` algorithm
10
+ - `visitors.md`: log & instrumentation visitors
11
+ - `arguments.md`: long explicit keyword/parameter lists (core API design principle)
12
+ - `swarm.md`: swarm architecture, element lifecycle, Ractor/Thread coding rules
13
+
14
+ ## Swarm overview (details in `swarm.md`)
15
+ - Purpose: zombie-lock elimination. `ProbeHosts` periodically `HSET`s every host id (`rql:hst:<pid>/<thread>/<ractor>/<identity>`) with `Time.now.to_f` into `rql:swarm:hsts`; `FlushZombies` treats hosts older than `zombie_ttl` (ms) as zombies and deletes their locks, their queue entries and the hosts themselves.
16
+ - `Client#swarm` → `Swarm` facade (one per client) owns a `Supervisor` and the swarm elements; started by `swarmize!` / `swarm.auto_swarm`, stopped by `deswarmize!`.
17
+ - Elements are independent background units: a control unit (`SwarmElement::Threaded` = Thread + `SizedQueue` command channel; `SwarmElement::Isolated` = Ractor driven via its own pair of `Ractor::Port`s: results port created in the main Ractor, where the Swarm and Supervisor live; command port created inside the element Ractor) that spawns, stops and reports on a main-loop Thread with its own Redis connection (`Swarm::RedisClientBuilder`).
18
+ - `ProbeHosts` is Threaded (it must see the client's ractor threads); `FlushZombies` is Isolated (copied config values only).
19
+ - `Supervisor` is one Thread that calls `reswarm_if_dead!` on every element each `liveness_probing_period`, restarting dead control units or stopped main loops; `swarm_status` aggregates `{ running:, state: }` of all of them.
20
+ - Redis work lives in stateless class-level functions (`ProbeHosts.probe_hosts`, `FlushZombies.flush_zombies`, `ZombieInfo.*`, `Acquirers.acquirers`), shared by main loops and the manual public API.
21
+
22
+ ## Observed conventions
23
+ - Every file starts with `# frozen_string_literal: true`.
24
+ - Constants are defined compactly: `class RedisQueuedLocks::Acquirer::IsLocked` / `module ...`, never nested `module A; module B`.
25
+ - Namespace files (`acquirer.rb`, `swarm.rb`, `logging.rb`, `swarm/swarm_element.rb`) only `require_relative` their children; the root `lib/redis_queued_locks.rb` requires the namespaces in dependency order.
26
+ - YARD doc block on every class, module, constant, attr and method:
27
+ `@param name [Type]`, `@option`, `@return [Type]`, blank `#` line, then `@api public|private`, `@since X.Y.Z`, optional `@version X.Y.Z` (latest behavior change).
28
+ - Operations are stateless modules with `class << self` functions; dependencies (`redis_client`, logger, instrumenter, sampling options) are passed as explicit args, never read from globals.
29
+ - Public API results are hashes `{ ok: Boolean, result: Symbol|Hash }` (`Client`, `Acquirer::*`, `Swarm` facade actions); `Client` `!` methods raise `RedisQueuedLocks::*Error`. Internal swarm element APIs return bare scalars/primitives (see `swarm.md`).
30
+ - `Client` methods only fill defaults from `config['...']` and delegate to `Acquirer::*` / `Swarm`.
31
+ - Redis keys come only from `RedisQueuedLocks::Resource.prepare_*` helpers and its `*_PATTERN` / `SWARM_KEY` constants.
32
+ - Config: `setting('key', default)` + `validate('key') { |val| ... }` in `config.rb`; dotted keys for nested groups (`swarm.flush_zombies.zombie_ttl`); units in a trailing `# NOTE: in milliseconds` comment.
33
+ - Errors: subclasses in `errors.rb`, written as `class XError < Error; end` (not `Class.new`) so RBS/Steep can see the superclass.
34
+ - Comments use `# NOTE:`, `# TODO:` and `# @type var x: T` / `#: T` for inline type hints.
35
+ - Rubocop is suppressed locally with `# rubocop:disable Metrics/MethodLength` etc. (paired `enable`), mostly for large algorithm methods.
36
+ - Shared mutable state is guarded by `RedisQueuedLocks::Utilities::Lock#synchronize`; Ractor code must not capture non-shareable objects (build a fresh Redis client via `Swarm::RedisClientBuilder`).
37
+
38
+ ## Claude rules
39
+ 1. Start new files with `# frozen_string_literal: true` and use compact constant paths.
40
+ 2. Add a full YARD block to every new public/private method; new code gets `@since <next version>`; when changing existing behavior, add/bump `@version`.
41
+ 3. New operation = new `Acquirer::*` module + thin `Client` method delegating to it (+ `!` variant only if it must raise); details in `acquirer.md`.
42
+ 4. Return `{ ok:, result: }` hashes from public API operations (`Client`, `Acquirer::*`, `Swarm` facade actions); raise only in `!` methods or on invalid arguments (`RedisQueuedLocks::ArgumentError`). Internal APIs, and swarm element internals in particular, use bare scalars/primitives (`true`/`false`, String, Symbol, Integer, `nil`, a flat Hash of scalars) without the wrapper.
43
+ 5. Never hardcode `rql:` strings; add a `Resource.prepare_*` helper or constant instead.
44
+ 6. Use WATCH/MULTI or a Lua constant for any read-modify-write on lock/queue keys; never do check-then-set in Ruby without a transaction.
45
+ 7. New config option: `setting` with default + `validate` with a type check, both in `config.rb`; read it in `Client` as `config['key']`.
46
+ 8. Logs and instrumentation events go through visitor modules (see `visitors.md`); event names follow `redis_queued_locks.<snake_case>`.
47
+ 9. Keep thread/Ractor safety: guard shared state with `Utilities::Lock`, never share a `RedisClient` across Ractors.
48
+ 10. Disable rubocop cops only locally with a matching `rubocop:enable`, and only for Metrics/Layout on large methods.
49
+ 11. After any change here, update the mirrored `sig/` file (see `type-checking.md`) and add/adjust a spec (see `tests.md`).
50
+
51
+ ## Recommendations (proposed, not yet project policy)
52
+ Apply to new or touched code; don't refactor existing code for these unless asked.
53
+ 1. Avoid `# rubocop:disable all` (used in `client.rb` lock_series and `lock_series_poc.rb`); disable only the specific cops.
54
+ 2. Document every option fully in YARD; `read_write_mode` is currently documented as `?` in `acquire_lock.rb`.
55
+ 3. Spell new identifiers correctly and don't copy existing typos (`swarm_element__termiante` method, `Acquier` in comments and the `acquier.rbs` filename); fix them only in a dedicated change.
56
+ 4. Include context in raised errors (lock name, acquirer id, timeout) so failures are diagnosable.
57
+ 5. Prefer `then` (the modern alias) over `yield_self` in new code.
@@ -0,0 +1,119 @@
1
+ ---
2
+ paths:
3
+ - "lib/redis_queued_locks/swarm.rb"
4
+ - "lib/redis_queued_locks/swarm/**/*.rb"
5
+ - "sig/redis_queued_locks/swarm.rbs"
6
+ - "sig/redis_queued_locks/swarm/**/*.rbs"
7
+ ---
8
+
9
+ # Swarm rules (`lib/redis_queued_locks/swarm/**`)
10
+
11
+ ## Purpose
12
+ The swarm removes **zombie locks**: locks and queue entries left by dead workers.
13
+ - **Host**: a `process/thread/ractor/identity` worker, with the id `rql:hst:<pid>/<thread_id>/<ractor_id>/<identity>` (`Resource.host_identifier`). Fibers are not included because `ObjectSpace` can't see Fibers or Threads once a Ractor exists. Hosts are enumerated with `Thread.list` (`Resource.possible_host_identifiers`).
14
+ - **Liveness**: `ProbeHosts` runs `HSET rql:swarm:hsts <host_id> <Time.now.to_f>` for every possible host of the client's ractor (`Resource::SWARM_KEY`).
15
+ - **Zombie**: a host whose last probe score is `< Resource.calc_zombie_score(zombie_ttl / 1_000.0)` (`now - ttl`). `zombie_ttl` is in milliseconds. Zombie locks are `rql:lock:*` whose `hst_id` field is a zombie host. Zombie acquirers are their `acq_id`s.
16
+ - **Flush** (`FlushZombies.flush_zombies`) runs these steps: `HGETALL` hosts → zombie hosts (return early if none) → `SCAN MATCH rql:lock:*` + `HMGET acq_id hst_id` → `DEL` zombie locks → `SCAN MATCH rql:lock_queue:*` + `ZREM` zombie acquirers → `HDEL` zombie hosts. It is best-effort and non-transactional, with full keyspace scans (`TODO: indexing`).
17
+
18
+ ## Components
19
+ | Object | File | Kind | Role |
20
+ |---|---|---|---|
21
+ | `Swarm` | `swarm.rb` | facade (one per `Client`, `client.swarm`) | owns the supervisor and elements. Public API: `swarm!`/`deswarm!`, `swarm_status`, `swarm_info`, `probe_hosts`, `flush_zombies`, `zombie_*`/`zombies_info`. Guards everything with its own `sync` |
22
+ | `Swarm::Supervisor` | `swarm/supervisor.rb` | plain `Thread` (`visor`) | every `swarm.supervisor.liveness_probing_period` seconds it runs the `observable` block, which calls `reswarm_if_dead!` on every element |
23
+ | `SwarmElement::Threaded` | `swarm/swarm_element/threaded.rb` | abstract base | control `Thread` plus a main-loop `Thread` |
24
+ | `SwarmElement::Isolated` | `swarm/swarm_element/isolated.rb` | abstract base | control `Ractor` plus a main-loop `Thread` inside it, driven through a per-element pair of `Ractor::Port`s |
25
+ | `ProbeHosts` | `swarm/probe_hosts.rb` | `< Threaded` | periodic host liveness probes |
26
+ | `FlushZombies` | `swarm/flush_zombies.rb` | `< Isolated` | periodic zombie flushing |
27
+ | `Acquirers`, `ZombieInfo` | `swarm/acquirers.rb`, `swarm/zombie_info.rb` | stateless `class << self` modules | read-only queries (`HGETALL` swarm hash, lock `SCAN`s) |
28
+ | `RedisClientBuilder` | `swarm/redis_client_builder.rb` | stateless module | `build(pooled:, sentinel:, config:, pool_config:)` returns a fresh `RedisClient` or `RedisClient::Pooled` |
29
+
30
+ `Client` methods `swarmize!`, `deswarmize!`, `swarm_status`/`swarm_state`, `swarm_info`, `probe_hosts`, `flush_zombies`, `zombie_locks`, `zombie_acquirers`, `zombie_hosts`, `zombies_info`/`zombies` only fill defaults from `config['swarm.*']` and delegate to `Swarm`. `Client#initialize` calls `swarm.swarm!` when `swarm.auto_swarm` is set.
31
+
32
+ ## Why Threaded vs Isolated
33
+ - `ProbeHosts` **must** be Threaded. Host ids come from `Thread.list` and `Ractor.current` of the client's ractor, so a probe from another Ractor would announce the wrong hosts. It may read `rql_client` (same ractor).
34
+ - `FlushZombies` is Isolated. The heavy `SCAN`/`DEL` work runs in its own Ractor with its own Redis connection, isolated from the app's ractor and objects. Everything it needs is passed into `Ractor.new(...)` as copied plain values (`config.slice('swarm.flush_zombies.redis_config')`, which `dup`s values, plus Integers). It never touches `rql_client` inside the ractor.
35
+
36
+ ## Element lifecycle (both bases)
37
+ Each element has two layers, with different code for each base:
38
+ 1. **Control unit** (`swarm_element`). Threaded: a `Thread` reading `swarm_element_commands` (`Thread::SizedQueue.new(1)`) and replying on `swarm_element_results` (`SizedQueue(1)`). Isolated: a `Ractor` running `self.swarm_loop(swarm_element_results_port)` that reads `swarm_element_commands_port` and replies on `swarm_element_results_port`, both `Ractor::Port`s (see "Isolated ports" below).
39
+ 2. **Main loop**: a `Thread` that does the actual periodic work (`loop { op(redis_client, ...); sleep(period) }`), with a Redis client built inside the thread by `RedisClientBuilder` from `swarm.<element>.redis_config.*`.
40
+
41
+ Phases: **init** `swarm!` (create control unit, loop not started) → **start** `swarm_loop__start` (kill old main loop, spawn new) → **stop** `swarm_loop__stop` kills only the main loop → **kill** `swarm_element__termiante` (Threaded: kill both threads, close and clear queues, nil them) / `swarm_loop__kill` (Isolated: `:kill` makes the ractor kill and join all its threads, ack, and leave its loop; the host then `join`s the ractor).
42
+
43
+ Command protocol: the same bare-value replies in both bases (see "Internal API: scalars and primitives" below).
44
+ | Command | Reply |
45
+ |---|---|
46
+ | `:status` | `{ alive: Boolean, state: String }` |
47
+ | `:is_active` | `true` / `false` |
48
+ | `:start` / `:stop` | `true` (ack) |
49
+ | `:kill` | Isolated only: `true` (ack), then the ractor finishes. Threaded is terminated from outside. |
50
+
51
+ Every request/reply pair is wrapped in `sync.synchronize` so concurrent callers can't interleave on the channel, and each reply belongs to the command just sent. Both bases send every command through one helper, `swarm_loop__send_command(command)`, which returns the reply, or `nil` when the element is dead (or, in Threaded, terminating). Abstract hooks (`spawn_main_loop!`, `spawn_swarm_element!`) raise `NotImplementedError` in the bases.
52
+
53
+ ## Internal API: scalars and primitives
54
+ `{ ok:, result: }` is the **public API** contract only: `Client` methods, `Acquirer::*` operations, and the `Swarm` facade's public methods (`swarm!`/`deswarm!`, plus the stateless operations behind `client.probe_hosts` / `client.flush_zombies`). Status reports (`swarm_status`, element `#status`, `Supervisor#status`) are plain nested Hashes with no wrapper.
55
+
56
+ Everything inside swarm elements is internal API and uses bare scalars and primitives, with no `{ ok:, result: }` wrapper. That covers control-unit commands and replies, the Isolated handshake, port and queue messages, `swarm_loop__*` helpers, `swarmed__*` predicates and main-loop plumbing.
57
+ - **Commands**: Symbols (`:status`, `:is_active`, `:start`, `:stop`, `:kill`).
58
+ - **Replies**: the value itself. Use `true`/`false` for questions and `true` as the ack for actions. Use a String/Symbol/Integer/Float for a single datum, and a flat Hash of scalars (`{ alive:, state: }`) only when a reply carries several related values. Don't nest.
59
+ - **Handshake**: the commands port itself, sent as the first message.
60
+ - **"No answer" / dead element**: `nil`. On the results port, a Symbol (`:exited`/`:aborted`) is reserved for the `Ractor#monitor` death notice, so never reply with a Symbol there. If a datum is a Symbol, send it as a String.
61
+ - **Errors**: don't wrap them in replies. A failing main loop just dies, and a dead control unit is detected through liveness, the monitor notice or `nil`.
62
+
63
+ ## Isolated ports (`Ractor::Port`)
64
+ Only the Ractor that created a port can `receive` from it; any Ractor can send to it. So each Isolated element owns two ports:
65
+ - **`swarm_element_results_port`**: created in the **main Ractor** by `swarm!`. The Swarm and its Supervisor live there, and so do all callers of the element API. It receives command replies and also the Ractor's termination notice: `swarm!` calls `swarm_element.monitor(results)`, which sends `:exited`/`:aborted`, immediately if the Ractor has already finished.
66
+ - **`swarm_element_commands_port`**: created **inside the element Ractor** by `.swarm_loop`. It's handed to the main Ractor as the first message on the results port (the handshake: the port itself; any other first message means the Ractor died during startup), which `swarm!` waits for.
67
+
68
+ Ports belong to the element instance, so any number of Isolated elements work side by side without sharing channels. A new `swarm!` creates a fresh pair, and stale messages left in an old results port are dropped with it.
69
+
70
+ Failure handling in `swarm_loop__send_command`:
71
+ - A non-Hash reply (`:exited`/`:aborted`) means the Ractor died, so the request returns `nil`.
72
+ - Sending to a finished Ractor's port raises `Ractor::ClosedError`, which is also rescued to `nil`.
73
+ - No request can block forever, because the monitor notice always arrives.
74
+
75
+ Thread hygiene inside the Ractor: a Ractor stays `running` until **all** its threads have finished. Killed but un-joined threads count, and so do helper threads that `Socket.tcp` leaves behind when the loop is killed mid-connect. `.terminate_thread` therefore kills and joins. `:kill` terminates every thread of the Ractor (`Thread.list - [Thread.current]`) before acking, and `swarm_loop__kill` joins the Ractor, so the status is `terminated` as soon as `try_kill!` returns.
76
+
77
+ State predicates (private, same names in both bases): `idle?` (no control unit), `swarmed?`, `swarmed__alive?`, `swarmed__dead?`, `swarmed__running?` (alive and main loop active), `swarmed__stopped?` (alive, loop inactive). Threaded also has `terminating?` (queues nil/closed). In Isolated, `swarmed__running?`/`swarmed__stopped?` return `false` when the request returned `nil` (the element died mid-request). Liveness: Threaded uses `Thread#alive?`; Isolated uses `Utilities.ractor_alive?`/`ractor_status`, which parse `Ractor#to_s` because Ractor has no status API. Thread state uses `Utilities.thread_state` (`'dead'`/`'failed'`/`Thread#status`).
78
+
79
+ Public element API (called by `Swarm`/`Supervisor` only):
80
+ - `try_swarm!`: no-op unless `enabled?`; terminate, then `swarm!`, then start.
81
+ - `reswarm_if_dead!`: no-op unless `enabled?`; `swarmed__stopped?` → restart the loop; `swarmed__dead? || idle?` → `swarm!` and start. This is how a crashed main loop (`abort_on_exception = false`) or a killed element recovers.
82
+ - `try_kill!`: terminate regardless of `enabled?`.
83
+ - `status`: Threaded `{ enabled:, thread: { running:, state: }, main_loop: { running:, state: } }`; Isolated is the same with `ractor:` instead of `thread:`. Unborn parts report `'non_initialized'`. Isolated gets the main loop part from a single `:status` request, and a dead element reports `'non_initialized'`.
84
+
85
+ ## Swarm orchestration (`Swarm#swarm!` / `#deswarm!`)
86
+ - `swarm!` (under `sync`): `supervisor.stop!` → each `element.try_swarm!` → `supervisor.observe! { each element.reswarm_if_dead! }` unless running → `sleep(0.1)` → `{ ok: true, result: :swarming }`. It can be called again; it re-creates everything.
87
+ - `deswarm!`: `supervisor.stop!` first, so it doesn't resurrect elements, then each `element.try_kill!` → `sleep(0.1)` → `{ ok: true, result: :terminating }`.
88
+ - `swarm_status`: `{ auto_swarm:, supervisor: { running:, state:, observable: }, probe_hosts: <status>, flush_zombies: <status> }`.
89
+ - The supervisor block swallows errors (`yield rescue nil`), so a failing element never kills the visor.
90
+
91
+ ## Config (`config.rb`)
92
+ `swarm.auto_swarm`, `swarm.supervisor.liveness_probing_period` (s). Per element: `swarm.<el>.enabled_for_swarm`, a period (`probe_hosts.probe_period` s, `flush_zombies.zombie_flush_period` s), and the 4-key `swarm.<el>.redis_config.{sentinel,pooled,config,pool_config}` group. Flush extras: `zombie_ttl` (ms), `zombie_lock_scan_size`, `zombie_queue_scan_size`.
93
+
94
+ ## Claude rules
95
+ 1. Swarm operations are stateless class-level functions (`ProbeHosts.probe_hosts`, `FlushZombies.flush_zombies`, `ZombieInfo.*`, `Acquirers.acquirers`) that take `redis_client` plus plain values. The manual public API and the main loop call the same function. Don't put Redis logic in instance methods.
96
+ 2. New element: subclass `Threaded` (it needs the client's ractor: threads, `rql_client`, or non-shareable objects) or `Isolated` (self-contained work on copied values). Implement `enabled?` plus `spawn_main_loop!` (Threaded, returns the `Thread`) or `spawn_swarm_element!(swarm_element_results_port)` (Isolated: return `Ractor.new(swarm_element_results_port, <plain args>) { |r_res_p, ...| <Klass>.swarm_loop(r_res_p) { Thread.new { ... } } }`). Don't override `swarm!` in Isolated subclasses; the base wires the ports, the monitor and the handshake.
97
+ 3. Register a new element everywhere: `attr_reader` + `initialize`, `swarm!` (`try_swarm!`), the supervisor block (`reswarm_if_dead!`), `deswarm!` (`try_kill!`), `swarm_status`. Add the config group (`enabled_for_swarm`, period, `redis_config.*`) with `validate`s, and add `Client` delegators with config defaults.
98
+ 4. Main loops build their own Redis client with `RedisClientBuilder` inside the loop thread. Never reuse `rql_client.redis_client` in a loop, and never pass a client, logger or other non-shareable object into a Ractor. Pass `config.slice(...)` and scalars instead.
99
+ 5. Inside a Ractor, reference only shareable constants and module functions (`RedisQueuedLocks::Swarm::X.op`, `RedisQueuedLocks::Resource`, `RedisQueuedLocks::Utilities`). No instance state, no captured locals.
100
+ 6. Change element state only inside `sync.synchronize`. `Utilities::Lock` is a reentrant `Monitor`, so the public→private nesting is intended. Keep the single-slot request/reply pairing atomic.
101
+ 7. Keep the lifecycle API names and semantics (`try_swarm!`, `reswarm_if_dead!`, `try_kill!`, `status`, the `swarmed__*` predicates, the `:status`/`:is_active`/`:start`/`:stop`/`:kill` commands) the same across both bases. The supervisor and `Swarm` rely on that duck typing.
102
+ 8. Main loop threads must have `abort_on_exception = false` (both control units set this) and let errors kill only the loop. Recovery belongs to the supervisor, so don't add retry loops inside elements.
103
+ 9. Stop the supervisor before killing or re-creating elements, or it will race and resurrect them.
104
+ 10. Report element status as the nested `{ running:, state: }` hashes with `'non_initialized'` for missing parts. Update `swarm_status` RBS types when the shape changes.
105
+ 11. Zombie detection compares probe scores (`Time.now.to_f`) with `Resource.calc_zombie_score`. Keep `zombie_ttl` in ms at the API and convert with `/ 1_000.0`. Use `Resource::SWARM_KEY`/`*_PATTERN` and never hardcode keys.
106
+ 12. Mirror every change in `sig/redis_queued_locks/swarm/**` and keep the existing `# steep:ignore` on nil-narrowed `Thread?`/`Ractor?` calls (Steep doesn't narrow `attr_reader` results).
107
+ 13. Isolated communication goes only through `Ractor::Port`s: create the results port in the main Ractor (`swarm!`) and the command port inside the element Ractor (`.swarm_loop`), and register `Ractor#monitor` on the results port before waiting on it. Every command gets exactly one reply, sent before the Ractor leaves its loop. Treat Symbol replies as death notices. Don't use `Ractor.receive`, `Ractor#send` or the default port, and never share one port between elements. Naming: port variables, attributes and params end in `_port` (`swarm_element_results_port`, `swarm_element_commands_port`). Block (proc) params use short initial-letter names (`r_res_p` for the results port, `r_com_p` for the commands port), like the other abbreviated Ractor block params (`rc`, `z_ttl`).
108
+ 14. Inside an element Ractor, terminate threads with kill + join (`.terminate_thread`) and keep `:kill` terminating every thread of the Ractor. Otherwise its status lags as `running`.
109
+ 15. Don't use `{ ok:, result: }` inside swarm elements (see "Internal API: scalars and primitives"). Commands are Symbols. Replies, helper returns and predicates are bare scalars/primitives, or a flat Hash of scalars, with `nil` for "no answer / dead". Keep the wrapper only at the public boundary (`Client`, `Acquirer::*`, `Swarm#swarm!`/`#deswarm!`, `ProbeHosts.probe_hosts`, `FlushZombies.flush_zombies`). Type internal replies in RBS as literal types/unions (`bool`, `String`, `{ alive: bool, state: String }`) rather than `xxxResult` records.
110
+ 16. Swarm specs (in `describe 'swarm'`) deal with real timing: set short periods via config, kill elements with `try_kill!`, wait longer than `liveness_probing_period`, and match states loosely (`eq('running').or(eq('blocking'))`, `'sleep'`/`'run'`).
111
+
112
+ ## Known quirks (don't copy; fix only in a dedicated change)
113
+ - Typos: `swarm_element__termiante`, `@since 19.0.0` on `Threaded#reswarm_if_dead!`, `@api ppublic` on `Swarm#zombie_acquirers`, "lopp"/"teh" in comments.
114
+ - `swarm_loop__stop` is not used by the swarm itself, only by specs.
115
+ - Killing an element while its main loop is still connecting to Redis makes Ruby's `Socket.tcp` helper threads print `terminated with exception ... IO#write` reports. This is harmless noise from the stdlib.
116
+ - `sleep(0.1)` "give a timespot" waits in `swarm!`/`deswarm!`/`observe!` instead of real readiness signalling.
117
+ - RBS runtime type checking (`rbs/test/setup`, the `typecheck-runtime` CI job) can't run inside Ractors. Its hooks on methods called in the element Ractor (`.swarm_loop`, `.flush_zombies`) read `RBS.logger` and raise `Ractor::IsolationError`, so Isolated elements die at startup and the swarm specs fail only under that job.
118
+ - Swarm has no logging or instrumentation, and supervisor errors are silently dropped (`TODO: (CHECK)`).
119
+ - `FlushZombies` scans the whole keyspace on every run and runs `ZREM` for every zombie acquirer on every queue.
@@ -0,0 +1,42 @@
1
+ ---
2
+ paths:
3
+ - "spec/**/*.rb"
4
+ - ".rspec"
5
+ ---
6
+
7
+ # Test rules (`spec/`)
8
+
9
+ ## Observed conventions
10
+ - All behavior specs live in one integration file, `spec/redis_queued_locks_spec.rb`, under `RSpec.describe RedisQueuedLocks`. The file header says it will be reworked; rspec-retry (5 retries) masks flakiness meanwhile.
11
+ - `.rspec` auto-requires `spec_helper`; `spec_helper.rb` loads SimpleCov first (`setup_simplecov.rb`), then `rspec/retry`, `pry`, the gem.
12
+ - RSpec config: random order (`Kernel.srand config.seed`), `disable_monkey_patching!`, `expect` syntax only, `filter_run_when_matching :focus`, `Thread.abort_on_exception = true`.
13
+ - Tests hit a real Redis (db 0) via `let(:redis) { RedisClient.config(db: 0).new_pool(timeout: 5, size: 50, ...) }` with long timeouts (RBS runtime-check runs are slow).
14
+ - `before`: `FLUSHDB`, `DEL Resource::SWARM_KEY`, `RedisQueuedLocks.enable_debugger!`; `after`: `DEL SWARM_KEY`, `FLUSHDB`.
15
+ - Examples are written with `specify '<feature>'` (few `it`); grouped with `describe` only for big areas (`'Lock Series PoC'`, `'swarm'`).
16
+ - Clients are built inline: `RedisQueuedLocks::Client.new(redis) { |config| ... }`.
17
+ - Logger/instrumenter fakes are anonymous classes (`Class.new { def debug(...); def notify(event, payload = {}) }`) collecting calls into arrays.
18
+ - Concurrency is tested with real `Thread.new` and `sleep` to wait for async swarm elements; assertions use `match(...)`, `eq(...).or(eq(...))`, `include`, `raise_error(RedisQueuedLocks::...Error)`.
19
+ - Every example cleans up its own state (`client.clear_locks`, `deswarmize!`, `redis.close`).
20
+ - Coverage: SimpleCov line + branch, HTML only; `minimum_coverage 100` is a TODO.
21
+
22
+ ## Claude rules
23
+ 1. Run specs with a local Redis available: `bundle exec rake rspec` (single example: `bundle exec rspec spec/redis_queued_locks_spec.rb:<line>`).
24
+ 2. Add new examples to `spec/redis_queued_locks_spec.rb` with `specify '<feature>'`, inside an existing `describe` when one fits; don't create new spec files unless asked (the suite is pending a rework).
25
+ 3. Use only `expect` syntax; no `should`, no monkey-patched DSL, no `focus` left behind.
26
+ 4. Use the shared `redis` pool and rely on the global `before`/`after` FLUSHDB; never use a different DB or flush in the middle of an example without reason.
27
+ 5. Build fakes as anonymous classes implementing the duck-typed interface (`debug`, `notify`, `sampling_happened?`), not with doubles of real loggers.
28
+ 6. Clean up everything an example starts: release locks, `deswarmize!` swarm clients, join/kill threads.
29
+ 7. Keep `sleep`-based waits minimal and comment why (`# give a timespot to ...`); prefer polling with a bounded timeout when adding new async checks.
30
+ 8. Do not add new rspec-retry reliance or lower retry settings; do not enable `minimum_coverage` without being asked.
31
+ 9. Specs are also run under RBS runtime checks, so pass correctly typed arguments to public API calls (type violations are logged in CI).
32
+
33
+ ## Recommendations (proposed, not yet project policy)
34
+ Apply to new tests; don't restructure the existing suite unless asked.
35
+ 1. Move toward one spec file per feature, mirroring `lib/` (`spec/redis_queued_locks/acquirer/extend_lock_ttl_spec.rb`), when the planned rework starts.
36
+ 2. Extract the repeated fake logger/notifier classes into `spec/support/` shared helpers or `shared_context`.
37
+ 3. Add a `wait_until(timeout:) { condition }` helper and use it instead of fixed `sleep` for swarm/thread checks; this reduces flakiness and rspec-retry reliance.
38
+ 4. Add Redis-free unit specs for pure modules (`Resource`, `Config` + validators, samplers, `Logging.should_log?`); they are fast and stable.
39
+ 5. Tag slow swarm/Ractor examples (e.g. `:swarm`) so they can be run or skipped separately.
40
+ 6. Read the Redis connection from `ENV['REDIS_URL']` (default `redis://localhost:6379/0`) so specs can run against non-default Redis setups.
41
+ 7. Once the suite is stable, remove rspec-retry and raise `minimum_coverage` step by step toward 100.
42
+ 8. Cover failure paths explicitly: timeouts, retry exhaustion, `fail_fast`, `raise_errors`, each conflict and access strategy.