pgbus 0.16.6 → 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: cddde0da7d83fc84f54404d10bcdbfccb06f25af74379b560ac4f9bb892f0041
4
- data.tar.gz: 44eba7a5a2bbe87659f388c7fb10c5a7b7577a1f045bb8d48c527cfbf9266df5
3
+ metadata.gz: 11b7ad58d91fe1eb9c0388008907523c77324f286f27a0e97a94353e4a8577b0
4
+ data.tar.gz: '0707593bb7c826f48f9a84db9a666057aaed3407b0d1cb559007f57caf080be2'
5
5
  SHA512:
6
- metadata.gz: 3bb680972106246ea3a16265730eb8b0567944e80323cbca8a9f68f7b7302b8afa2087cfee74e0fa99b695f038fd5e2104141971096ac8d76b236437ce147e60
7
- data.tar.gz: cf822271a70b7ec8bcbbca0591838179678113c2b19849a87b4442cf5e0ccd3b2bac1d7a92e8820596c839ff1005c82d0c0ab45d47ddad2df1bae8094b0a9208
6
+ metadata.gz: 0f1422dea206dbeaac10b1fd7dfa74da2338c007cffb48d7a2b5429bf3a94e6f3e78b5795832d0a2f45346fe4350d4c7355d16b773247420894ea5234c50db85
7
+ data.tar.gz: 6d30c3583b7e6b5fa9dc89cf42115f57b182dea18b2d363f1ec413d6096ba32b76d75f008dcf7a2f25de3baf9ceba7e17f1edbb382b96e2aeb23bf52a8824af2
data/CHANGELOG.md CHANGED
@@ -6,10 +6,16 @@
6
6
 
7
7
  - **Concurrency is now visible: a section on the Locks page, three gauges, a home stat card and an MCP tool.** `limits_concurrency … on_conflict: :block` became durable in #460 — a parked job is never dropped and a key never runs past `to:` — but nothing showed it, so an operator whose per-record pipeline looked "stuck" had no way to see that the key's slot was held, how many jobs were parked behind it, or how long the oldest had waited, short of SQL. `pgbus_semaphores` and `pgbus_blocked_executions` appeared on no page, in no gauge, in no tool. One new read, `Web::DataSource#concurrency_stats`, now feeds four surfaces: the Locks page's **Concurrency** section (parked jobs, oldest wait, slots held, keys at limit, then a row per key with value / limit, lease state, parked count and oldest wait, auto-refreshing like the other dashboard panels); a **Parked jobs** card on the dashboard home; three unlabelled gauges on both `/pgbus/api/metrics` and the AppSignal probe (`pgbus_concurrency_blocked_executions`, `pgbus_concurrency_blocked_oldest_age_seconds`, `pgbus_concurrency_slots_held` — keys are per-record, so a per-key label would be an unbounded series); and a read-only `pgbus_concurrency` MCP tool returning the same data. The key rows come from a FULL OUTER JOIN, not a LEFT one: a key can have parked jobs and no semaphore (its holder died and the sweep removed the row) or a semaphore and nothing parked, and both shapes matter to an operator. Each row carries two guarded escape hatches. **Release** drops the semaphore and promotes through `Concurrency::BlockedExecution.promote_next`, so it takes each slot through the same guarded upsert an enqueue uses and can never reintroduce the over-admission #460 closed; when the lease is still fresh the confirm says so, because a job is then probably still running and releasing lets another start beside it. **Discard parked** drops a single key's parked jobs, naming the exact count in the confirm, and resolves the bookkeeping those jobs will never resolve themselves — a parked batch child is marked failed so its batch stops waiting and `on_failure` fires (#413), and an `:until_executed` uniqueness lock is released rather than orphaned (#423); each discard logs a warning naming the job and emits `pgbus.blocked_execution_discarded`. There is deliberately no cross-key "discard all parked": not losing work is the whole point of `:block`. Nothing is added to the enqueue or execute path. Refs #461.
8
8
 
9
+ - **PGMQ: vendored schema v1.13.0, and the upgrade migration now carries upstream's table fixups.** The 1.12.0 → 1.13.0 delta is entirely partitioned-queue work plus one metrics attribute: `pgmq.create_partitioned` gains a `premake` argument (its three-argument form dropped, so pgmq-ruby's three-argument call binds to the new signature with the default), new partitioned queues use `msg_id … GENERATED BY DEFAULT AS IDENTITY` rather than `GENERATED ALWAYS`, and `pgmq.metrics_result` gains `default_partition_length` (NULL for non-partitioned queues, ignored by pgmq-ruby, which reads rows by name). Nothing pgbus calls through `Pgbus::Client` changes, and pgbus creates only partitioned *archives*, whose `msg_id` is a plain column — so no client or pgmq-ruby change. Fresh embedded installs get 1.13.0; existing *embedded* installs: `rake pgbus:pgmq:status`, then `rails generate pgbus:upgrade_pgmq` + `rails db:migrate` (`rails db:migrate:pgbus` on a separate-database install). An `:extension` install is upgraded by PostgreSQL itself (`ALTER EXTENSION pgmq UPDATE`) once the server has the new version packaged — not by this migration, whose function drop refuses to remove extension-owned objects. **New:** because the upgrade works by dropping every pgmq function and composite type and re-running the target schema, it could never carry an upstream change to an existing *table*. The generated migration now reads the version in `pgbus_pgmq_schema_versions` and applies the fixup vendored for each version between it and the target (`lib/pgbus/pgmq_schema/fixups/pgmq_v<VERSION>.sql`, discovered by the same glob as the schema files, via `PgmqSchema.fixups_sql`). This hop's fixup is upstream's own guarded block moving existing partitioned queues to `GENERATED BY DEFAULT`; it selects on `pgmq.meta.is_partitioned`, so it is a no-op unless partitioned queues were created through PGMQ directly. Fixups must be idempotent — an install with no recorded version gets all of them. Closes #459.
10
+
9
11
  - **The batches list and a batch's progress panel now refresh themselves.** Batches was the only dashboard list whose `turbo-frame` was missing `data-auto-refresh`, and the batch detail page had no frame at all, so watching a batch drain meant reloading by hand. Both now poll on the existing `web_refresh_interval` (5s, pauses while you interact) like every other list — the batch progress panel is served by `GET /batches/:id?frame=progress`. This reuses the dashboard's turbo-frame polling rather than the SSE stack in `lib/pgbus/web/stream_app.rb`: that is the host-app-facing stream feature, and pointing it at the dashboard would open a PG `LISTEN` connection per viewer to move four integers.
10
12
 
11
13
  ### Fixed
12
14
 
15
+ - **A still-running event handler is no longer re-run as if it had crashed, and a slow one is no longer redelivered mid-run.** Two halves of one hole, follow-up to #469. First: `Process::Consumer` had no equivalent of `ActiveJob::Executor#with_visibility_heartbeat`, so an event message's visibility timeout was never re-armed — a handler slower than `visibility_timeout` (30s by default) was redelivered *while it was still running*, `read_ct` climbed on every redelivery, and after `max_retries` the event was dead-lettered without the handler ever raising. Workers have been protected from this since the heartbeat landed; event consumers never were. The consumer now tracks each message through `VisibilityHeartbeat` for exactly as long as its handlers run, releasing the entry before the archive so a beat can never re-arm a message that is already gone. Second: `Handler#claim_idempotency?` read a `pgbus_processed_events` row with `completed_at IS NULL` as proof that the previous holder had been killed mid-handler, and re-ran. That state equally describes a handler that is simply still running — on another thread, another fork or another host — so a second delivery (a genuine redelivery, or a duplicate envelope, whose independent visibility timeouts no heartbeat can serialize) executed the handler *concurrently with* the live one, which is precisely the double-execution `idempotent!` exists to prevent. The claim now carries liveness rather than only a claim instant: `EventBus::ClaimBeat` refreshes every in-flight claim's `processed_at` from the same beat that re-arms the message's visibility timeout, so message visibility and claim liveness go quiet together when a process dies. A pending claim silent for longer than twice the heartbeat interval is abandoned and re-runs, exactly as before; a fresher one is owned and the delivery skips, deferring to the holder — which either completes (nothing is lost) or fails, leaving its own message for visibility-timeout redelivery to recover. A skip is also no longer silent: `pgbus.event_skipped` carries the reason (`:completed`, `:cached` or `:owned`), the claim's age in seconds and the delivery's `read_ct`, and the metrics subscriber counts it as `pgbus_event_count` with `status: "skipped"` and the reason as a tag. No migration, no new column, and no change on a table that has not run `pgbus:add_processed_event_completion` — a single-phase claim has no pending state and therefore no ownership question. Closes #470.
16
+
17
+ - **An event is no longer handed to every handler whose pattern matches it, so a wildcard handler ran once per subscriber on the topic.** Each subscriber gets its own queue (`Subscriber#setup!` creates and binds one), so a topic with N matching subscribers puts N copies of every event on the bus — but `Consumer#handle_message` resolved handlers with `Registry#handlers_for(routing_key)`, a pattern match that ignored which queue the message had just been read from, and ran every match. Each of the N copies therefore fanned out to all N handlers: **every handler ran N times per event**, with four `"#"` subscribers meaning four invocations of each handler for every event on the bus. `idempotent!` was the only thing hiding it, and it hides it imperfectly by design: the two-phase claim (#385) deliberately *re-runs* when the existing claim is still pending, on the theory that the prior holder crashed — but a pending claim is also exactly what a handler still running on another host looks like. So one event could have host A win the claim and start a 700ms handler while host B, reading a different subscriber's copy of the same event a few hundred milliseconds later, dispatched to that same handler, lost the claim insert, read `completed_at IS NULL` as "crashed, re-run", and executed it concurrently. In one host app that meant a "task completed" record written twice and the user emailed twice, with a single row in `pgbus_processed_events`, `read_ct = 1` on every archived copy, and the second execution's side effects timestamped inside the first's window — dispatch, not redelivery. A non-idempotent handler simply ran N times with no guard at all. `handlers_for` now takes a required `queue_name:` and selects on ownership *and* pattern, so a message read from queue Q goes to Q's owner(s) alone and each handler runs exactly once per event; duplicate execution again requires a real redelivery, which is the at-least-once contract the docs describe. The keyword is required rather than defaulted precisely so a keyword-less call cannot silently restore the fan-out — the two callers that genuinely want the pattern view (the `Testing.inline!` publish path and `Testing::EventStore#drain!`, neither of which has a queue, since the event never reaches PGMQ) moved to an explicitly named `Registry#subscribers_matching`, where per-subscriber delivery counts already match what owner-only dispatch now produces in production. Two handlers registered against the same explicit `queue_name:` still both run on that queue's delivery. The pattern check is kept alongside the ownership check because a routing key the owner's pattern does not match means a stale `pgmq.topic_bindings` row, and a stale binding must not run the handler. A message on a queue no subscriber in this process owns is still archived rather than looped through visibility-timeout redelivery into the DLQ — that was already the behavior when `handlers_for` returned `[]` — but it is no longer silent: it logs a warning naming the queue and routing key, rate-limited to once per queue per process so a permanently stale binding cannot flood the log, and emits `pgbus.event_unrouted` on every occurrence so the real rate stays visible in metrics. The pending-claim re-run in `Handler#claim_idempotency?` is deliberately untouched here; it is a separate question now that dispatch no longer manufactures the concurrency it was misreading. Closes #469.
18
+
13
19
  - **A dropped ActiveRecord socket no longer pages the host app for an event that was handled successfully.** EventBus consumers are long-lived threads holding a leased AR connection, and a pooler restart, an admin disconnect or a brief failover can kill that socket at any point. Rails reconnects most statements transparently, but `Relation#update_all` is marked `allow_retry: false` — and that is exactly the phase-2 claim stamp in `Handler#complete_claim!`, the one AR write that happens *after* `handle` has already returned. A drop there surfaced as `ActiveRecord::ConnectionFailed` ("PQconsumeInput() SSL error: unexpected eof while reading") on a message whose work was done, so the host app's exception tracker paged for a false failure and PGMQ redelivered the event for a re-run. `EventBus::StaleConnectionRetry` now reconnects this thread's lease (via `ConnectionPool#active_connection?`, the accessor that exists across the supported Rails range) and repeats the stamp once; a second drop still raises, leaving the message to VT redelivery. Only the leased connections are reconnected — `clear_all_connections!` would yank sockets out from under sibling consumers in the same process, turning one recoverable drop into many. The retryable patterns are deliberately *broader* than `Client::STALE_CONNECTION_PATTERNS` and for the opposite reason: that list excludes mid-flight drops because a half-committed enqueue would duplicate a message, whereas this only ever repeats an idempotent `SET completed_at = <now>`. Phase 1 (`claim_idempotency?`) is deliberately not wrapped — its INSERT may have committed before the socket died, and on a legacy schema the retry's empty `result.rows` would read as "another consumer owns this claim", turning a recoverable drop into a silently skipped event.
14
20
 
15
21
 
data/README.md CHANGED
@@ -221,6 +221,14 @@ Pgbus::EventBus::Registry.instance.subscribe(
221
221
  )
222
222
  ```
223
223
 
224
+ Each subscriber gets its own queue, and a queue is consumed only by the
225
+ handler(s) registered against it: one published event matching N subscribers
226
+ becomes N deliveries, one per handler, so every handler runs exactly once per
227
+ event. (Two handlers sharing an explicit `queue_name:` both run on that queue's
228
+ delivery.) A message on a queue no subscriber in this process owns — a stale
229
+ topic binding left by a renamed or removed handler — is archived, logged once
230
+ per queue, and reported as `pgbus.event_unrouted`.
231
+
224
232
  `idempotent!` uses a **two-phase claim**: a *pending* row in
225
233
  `pgbus_processed_events` is inserted before `handle` runs, and only stamped
226
234
  `completed_at` after `handle` returns. Deduplication applies to **completed**
data/Rakefile CHANGED
@@ -16,7 +16,7 @@ require "rubocop/rake_task"
16
16
  # on the unresolvable gem inheritance. Passing explicit paths stops the
17
17
  # discovery. docs/ lints itself in the Docs site workflow. (Same fix as docs-kit.)
18
18
  RuboCop::RakeTask.new do |task|
19
- task.patterns = %w[app benchmarks config lib spec Gemfile Rakefile pgbus.gemspec]
19
+ task.patterns = %w[app benchmarks config lib rakelib spec Gemfile Rakefile pgbus.gemspec]
20
20
  end
21
21
 
22
22
  namespace :bench do
@@ -164,173 +164,8 @@ task :build do
164
164
  sh("rm -rf /tmp/gem-verify #{gem_file}")
165
165
  end
166
166
 
167
- desc "Release a new version (rake release[1.2.3] or rake release[pre] or rake release[1.2.3,force])"
168
- task :release, %i[version force] do |_t, args|
169
- require_relative "lib/pgbus/version"
170
-
171
- def info(msg) = puts "\e[34m→\e[0m #{msg}"
172
- def success(msg) = puts "\e[32m✓\e[0m #{msg}"
173
- def skip(msg) = puts "\e[33m⊘\e[0m #{msg} \e[33m(skipped)\e[0m"
174
- def warn(msg) = puts "\e[33m⚠\e[0m #{msg}"
175
- def error(msg) = puts "\e[31m✗\e[0m #{msg}"
176
- def header(msg) = puts "\n\e[1;36m#{msg}\e[0m\n#{"─" * msg.length}"
177
-
178
- new_version = args[:version]
179
- abort "\e[31mUsage: rake release[X.Y.Z] or rake release[X.Y.Z,force]\e[0m" unless new_version
180
-
181
- force = args[:force]&.to_s&.downcase == "force"
182
-
183
- dirty = `git status --porcelain`.strip
184
- abort "\e[31mAborting: working directory is not clean.\e[0m\n#{dirty}" unless dirty.empty?
185
-
186
- current = Pgbus::VERSION
187
- prerelease = new_version.match?(/alpha|beta|rc|pre/) || new_version == "pre"
188
-
189
- if new_version == "pre"
190
- new_version = current
191
- prerelease = true
192
- end
193
-
194
- tag = "v#{new_version}"
195
- version_file = "lib/pgbus/version.rb"
196
-
197
- title = "Release #{tag}"
198
- title += " (force)" if force
199
- header title
200
- info "Current version: #{current}"
201
- info "New version: #{new_version}"
202
- info "Pre-release: #{prerelease}"
203
-
204
- # Step 0: Force cleanup — delete existing release and tag
205
- if force
206
- header "Force cleanup"
207
- if system("gh release view #{tag} >/dev/null 2>&1")
208
- sh("gh release delete #{tag} --yes --cleanup-tag")
209
- success "Deleted release and remote tag #{tag}"
210
- else
211
- skip "No release #{tag} to delete"
212
- end
213
-
214
- if system("git rev-parse #{tag} >/dev/null 2>&1")
215
- sh("git tag -d #{tag}")
216
- success "Deleted local tag #{tag}"
217
- else
218
- skip "No local tag #{tag} to delete"
219
- end
220
- end
221
-
222
- # Step 1: Update version file
223
- header "Version"
224
- if new_version == current
225
- skip "Version already #{new_version}"
226
- else
227
- content = File.read(version_file)
228
- content.sub!(/VERSION = ".*"/, "VERSION = \"#{new_version}\"")
229
- File.write(version_file, content)
230
- success "Updated #{version_file}"
231
- end
232
-
233
- # Step 1b: Regenerate the frozen lockfiles that pin the pgbus path gem, so the
234
- # bump ships with them in sync. These are installed with `--frozen`/deployment
235
- # in CI, so if they still name the OLD version they instant-fail (the root
236
- # Gemfile.lock on every main-Gemfile leg AND release.yml's own `bundle install`,
237
- # the Rails 7.1 leg with exit 16, and docs-CI on any docs change). Regenerating
238
- # here keeps the version-pin drift out of the release commit instead of
239
- # surfacing on the next PR — or, worse, in the Release workflow itself.
240
- header "Frozen lockfiles"
241
- # The ONLY thing a version bump changes in these frozen lockfiles is the pgbus
242
- # path-gem pin — so bump exactly that line, in place, with a string edit.
243
- #
244
- # We deliberately do NOT run `bundle lock` here: a full re-resolve trips over
245
- # constraints that have nothing to do with pgbus. Concretely, docs/Gemfile.lock
246
- # carries a broad PLATFORMS list (…-gnu / …-musl / arm-linux) for which a
247
- # platform gem like `thruster` ships no variant, so `bundle lock` fails with
248
- # "Could not find gems matching 'thruster' valid for all resolution platforms"
249
- # on any machine whose cache doesn't already hold those exact gems — aborting
250
- # the release. `bundle lock --local` was even worse (wrong-file write + no
251
- # fetch). A targeted pin edit sidesteps all of it, is deterministic on any
252
- # machine, and produces the minimal 2-line diff (the PATH spec + the
253
- # DEPENDENCIES pin). See #338/#341 and the surgical-bump fix.
254
- frozen_lockfiles = %w[Gemfile.lock gemfiles/rails_7_1.gemfile.lock docs/Gemfile.lock]
255
- regenerated_lockfiles = []
256
- frozen_lockfiles.each do |lockfile|
257
- unless File.exist?(lockfile)
258
- skip "#{lockfile} not present"
259
- next
260
- end
261
-
262
- content = File.read(lockfile)
263
- # Matches both the PATH-source spec (" pgbus (X.Y.Z)") and the
264
- # DEPENDENCIES pin (" pgbus (X.Y.Z)"), leaving everything else untouched.
265
- bumped = content.gsub(/^(\s+pgbus) \([^)]*\)$/, "\\1 (#{new_version})")
266
-
267
- if bumped == content
268
- skip "#{lockfile} — no pgbus pin to bump"
269
- next
270
- end
271
-
272
- File.write(lockfile, bumped)
273
- regenerated_lockfiles << lockfile
274
- success "Bumped pgbus pin in #{lockfile}"
275
- end
276
-
277
- # Step 2: Verify gem builds cleanly
278
- header "Build verification"
279
- sh("gem build pgbus.gemspec --strict")
280
- sh("rm -f pgbus-*.gem")
281
- success "Gem builds cleanly"
282
-
283
- # Step 3: Commit version bump (+ any re-synced lockfiles)
284
- header "Git commit"
285
- paths_to_stage = [version_file, *regenerated_lockfiles]
286
- staged_changes = paths_to_stage.any? do |path|
287
- !`git diff #{path}`.strip.empty? || !`git diff --cached #{path}`.strip.empty?
288
- end
289
- if staged_changes
290
- paths_to_stage.each { |path| sh("git add #{path}") }
291
- sh("git commit -m 'chore: bump version to #{new_version}'")
292
- success "Committed version bump"
293
- else
294
- skip "No version change to commit"
295
- end
296
-
297
- # Step 4: Push to origin
298
- header "Git push"
299
- local_sha = `git rev-parse HEAD`.strip
300
- remote_sha = `git rev-parse origin/main 2>/dev/null`.strip
301
- if local_sha == remote_sha
302
- skip "origin/main already at #{local_sha[0..6]}"
303
- else
304
- sh("git push origin main")
305
- success "Pushed to origin/main"
306
- end
307
-
308
- # Step 5: Create release
309
- header "Release"
310
- tag_exists = system("git rev-parse #{tag} >/dev/null 2>&1")
311
- release_exists = system("gh release view #{tag} >/dev/null 2>&1")
312
-
313
- if release_exists
314
- skip "Release #{tag} already exists (use force to re-create)"
315
- elsif tag_exists
316
- info "Tag #{tag} exists, creating release from it"
317
- pre_flag = prerelease ? "--prerelease" : ""
318
- sh("gh release create #{tag} --generate-notes #{pre_flag}".strip)
319
- success "Release #{tag} created from existing tag"
320
- else
321
- pre_flag = prerelease ? "--prerelease" : ""
322
- sh("gh release create #{tag} --generate-notes --target main #{pre_flag}".strip)
323
- success "Release #{tag} created"
324
- end
325
-
326
- puts ""
327
- success "\e[1mRelease #{tag} complete!\e[0m CI will handle the rest:"
328
- puts " • Run tests"
329
- puts " • Build + verify gem"
330
- puts " • Sign with Sigstore"
331
- puts " • Publish to RubyGems"
332
- puts " • Upload assets to the release"
333
- end
167
+ # `rake release[X.Y.Z]` lives in rakelib/release.rake (shared across the
168
+ # zoolutions gems); `bin/release` is its interactive front door.
334
169
 
335
170
  namespace :frontend do
336
171
  # app/frontend/pgbus/style.css is a committed build artifact — the engine ships
@@ -8,6 +8,16 @@ class UpgradePgmqToV<%= target_version_slug.camelize %> < ActiveRecord::Migratio
8
8
  # This uses the vendored SQL which doesn't require the pgmq extension.
9
9
  execute Pgbus::PgmqSchema.install_sql("<%= target_version %>")
10
10
 
11
+ # Apply the table fixups upstream ships for every hop between the
12
+ # version this database records and the target. Dropping and re-creating
13
+ # the functions cannot carry a change to an existing TABLE (1.13.0 moves
14
+ # partitioned queues' msg_id to GENERATED BY DEFAULT), so those ride
15
+ # here. Runs after install_sql because the fixups call
16
+ # pgmq.format_table_name, which the drop step removed. Each fixup is
17
+ # idempotent, so an install with no recorded version gets them all.
18
+ fixups = Pgbus::PgmqSchema.fixups_sql(after: installed_pgmq_version, upto: "<%= target_version %>")
19
+ execute fixups unless fixups.empty?
20
+
11
21
  # Re-install the per-queue NOTIFY insert triggers the function drop
12
22
  # cascaded away (pgmq.notify_insert_throttle records which queues had
13
23
  # them and at what throttle; enable_notify_insert is idempotent).
@@ -34,4 +44,16 @@ class UpgradePgmqToV<%= target_version_slug.camelize %> < ActiveRecord::Migratio
34
44
  "PGMQ schema downgrade is not supported. " \
35
45
  "Restore from a database backup if needed."
36
46
  end
47
+
48
+ private
49
+
50
+ # The version this database last recorded, or nil when it has never been
51
+ # tracked (a pre-tracking install, or one bootstrapped by hand).
52
+ def installed_pgmq_version
53
+ return nil unless connection.table_exists?("pgbus_pgmq_schema_versions")
54
+
55
+ connection.select_value(
56
+ "SELECT version FROM pgbus_pgmq_schema_versions ORDER BY installed_at DESC LIMIT 1"
57
+ )
58
+ end
37
59
  end
@@ -0,0 +1,80 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Pgbus
4
+ module EventBus
5
+ # Liveness for the idempotency claims in flight for one PGMQ message.
6
+ #
7
+ # A two-phase claim (issue #385) is a `pgbus_processed_events` row with
8
+ # `completed_at IS NULL`. That state says "claimed, not finished" — it does
9
+ # NOT say whether the claimer is dead or simply still running (issue #470).
10
+ # Handler resolves the ambiguity by age: a claim whose `processed_at` has
11
+ # gone quiet for longer than the ownership window is abandoned, anything
12
+ # fresher is owned. For that age to mean "silence" rather than
13
+ # "time since the claim was taken", something has to keep stamping it while
14
+ # the handler runs. This is that something.
15
+ #
16
+ # One beat per message, created by Process::Consumer and handed to every
17
+ # handler dispatched for it. A handler registers its claim for exactly the
18
+ # duration of `handle` and the consumer's VisibilityHeartbeat `on_beat` hook
19
+ # drives #touch! on the same cadence that re-arms the message's visibility
20
+ # timeout — so the claim and the message go quiet together when the process
21
+ # dies, and both stay fresh while it lives.
22
+ #
23
+ # #touch! runs on the heartbeat ticker thread while #register / #release run
24
+ # on a pool thread, hence the mutex.
25
+ class ClaimBeat
26
+ def initialize
27
+ @mutex = Mutex.new
28
+ @claims = []
29
+ end
30
+
31
+ def register(event_id, handler_class)
32
+ claim = [event_id, handler_class]
33
+ @mutex.synchronize { @claims << claim unless @claims.include?(claim) }
34
+ self
35
+ end
36
+
37
+ def release(event_id, handler_class)
38
+ @mutex.synchronize { @claims.delete([event_id, handler_class]) }
39
+ self
40
+ end
41
+
42
+ def size
43
+ @mutex.synchronize { @claims.size }
44
+ end
45
+
46
+ def empty?
47
+ size.zero?
48
+ end
49
+
50
+ # Refresh every registered claim's liveness stamp. Returns the number of
51
+ # claims touched. No-op on a legacy schema: without `completed_at` there
52
+ # are no pending claims to keep alive, and Handler's single-phase
53
+ # fallback never consults the age.
54
+ #
55
+ # A claim that fails to update is logged and skipped rather than raised:
56
+ # this runs inside the visibility heartbeat's beat, and one unwritable
57
+ # row must not cost every other in-flight message its VT extension.
58
+ def touch!
59
+ return 0 unless ProcessedEvent.completion_column?
60
+
61
+ now = Time.now.utc
62
+ @mutex.synchronize { @claims.dup }.count { |event_id, handler_class| touch(event_id, handler_class, now) }
63
+ end
64
+
65
+ private
66
+
67
+ def touch(event_id, handler_class, now)
68
+ ProcessedEvent
69
+ .where(event_id: event_id, handler_class: handler_class, completed_at: nil)
70
+ .update_all(processed_at: now)
71
+ true
72
+ rescue StandardError => e
73
+ Pgbus.logger.warn do
74
+ "[Pgbus] Could not refresh idempotency claim #{handler_class}/#{event_id}: #{e.class}: #{e.message}"
75
+ end
76
+ false
77
+ end
78
+ end
79
+ end
80
+ end
@@ -17,8 +17,20 @@ module Pgbus
17
17
  end
18
18
  end
19
19
 
20
- def process(message)
21
- with_rails_executor { process!(message) }
20
+ # Outcome of the two-phase claim. `age` is how long the losing delivery
21
+ # found the existing claim to have been silent, in seconds (nil unless
22
+ # the claim was pending).
23
+ ClaimResult = Data.define(:status, :age) do
24
+ def granted?
25
+ status == :claimed
26
+ end
27
+ end
28
+
29
+ # @param claim_beat [ClaimBeat, nil] the message's claim-liveness beat,
30
+ # supplied by Process::Consumer. Absent for a hand-rolled caller: the
31
+ # handler still runs, its claim simply ages from the claim instant.
32
+ def process(message, claim_beat: nil)
33
+ with_rails_executor { process!(message, claim_beat) }
22
34
  end
23
35
 
24
36
  def handle(event)
@@ -27,12 +39,18 @@ module Pgbus
27
39
 
28
40
  private
29
41
 
30
- def process!(message)
42
+ def process!(message, claim_beat = nil)
31
43
  raw = JSON.parse(message.message)
32
44
  event = build_event(raw)
33
45
  routing_key = raw.dig("headers", "routing_key") || raw["routing_key"]
34
46
 
35
- return :skipped if self.class.idempotent? && !claim_idempotency?(event.event_id)
47
+ if self.class.idempotent?
48
+ claim = claim_idempotency(event.event_id)
49
+ unless claim.granted?
50
+ instrument_skip(claim, event, message, routing_key)
51
+ return :skipped
52
+ end
53
+ end
36
54
 
37
55
  instrument_payload = {
38
56
  event_id: event.event_id,
@@ -42,11 +60,13 @@ module Pgbus
42
60
  read_ct: message.read_ct.to_i,
43
61
  msg_id: message.msg_id.to_i
44
62
  }
45
- Instrumentation.instrument("pgbus.event_processed", instrument_payload) do
46
- # Publisher's Current attributes (issue #431) are set for the handler
47
- # and reverted after (CurrentAttributes#set semantics); the Rails
48
- # executor wrap above additionally resets at completion.
49
- Pgbus::CurrentAttributes.restore(event.context) { handle(event) }
63
+ with_claim_beat(claim_beat, event.event_id) do
64
+ Instrumentation.instrument("pgbus.event_processed", instrument_payload) do
65
+ # Publisher's Current attributes (issue #431) are set for the handler
66
+ # and reverted after (CurrentAttributes#set semantics); the Rails
67
+ # executor wrap above additionally resets at completion.
68
+ Pgbus::CurrentAttributes.restore(event.context) { handle(event) }
69
+ end
50
70
  end
51
71
  complete_claim!(event.event_id) if self.class.idempotent?
52
72
  :handled
@@ -64,7 +84,7 @@ module Pgbus
64
84
 
65
85
  # Mirrors Pgbus::ActiveJob::Executor#execute_job: wrap the handler
66
86
  # invocation in Rails.application.executor (or the reloader in dev)
67
- # so AR connections leased by `claim_idempotency?` and `handle` are
87
+ # so AR connections leased by `claim_idempotency` and `handle` are
68
88
  # released back to the pool when this method returns. Without the
69
89
  # wrap, every consumed event leaks one AR connection on the consumer
70
90
  # thread — in dev that wedges `clear_reloadable_connections!`,
@@ -113,24 +133,32 @@ module Pgbus
113
133
 
114
134
  # Two-phase idempotency claim (issue #385). Phase 1: atomically claim
115
135
  # via INSERT ... ON CONFLICT DO NOTHING with completed_at NULL — a
116
- # *pending* claim. Returns true when this delivery should run handle:
136
+ # *pending* claim. Returns a ClaimResult whose status is one of:
117
137
  #
118
- # - insert won → fresh claim
119
- # - insert lost, completed_at NULL → a prior attempt claimed but was
120
- # killed before finishing (SIGKILL mid-handler); re-run so the crash
121
- # doesn't silently drop the execution. Safe: PGMQ's VT means the
122
- # prior holder is dead or wedged past its timeout — the same
123
- # at-least-once window every non-idempotent handler has.
138
+ # :claimed — insert won (fresh claim), or the row was purged between
139
+ # the losing insert and the read, or an existing pending
140
+ # claim has gone silent for longer than the ownership
141
+ # window: the holder is dead by the heartbeat's own
142
+ # definition, so re-run rather than silently drop the
143
+ # execution a SIGKILL interrupted.
144
+ # :completed — the execution already finished. Skip.
145
+ # :owned — pending, and its liveness stamp is fresh: the holder is
146
+ # still running (issue #470). Skip — running `handle`
147
+ # concurrently with the holder is exactly the
148
+ # double-execution `idempotent!` promises not to do. The
149
+ # holder either completes (nothing lost) or fails, leaving
150
+ # its own message for VT redelivery to recover.
151
+ # :cached — a completed execution already in this process's memory.
124
152
  #
125
- # Returns false (skip) only for a *completed* execution. Phase 2 is
126
- # complete_claim! after handle returns; only completed executions enter
127
- # the in-memory dedup cache.
153
+ # Phase 2 is complete_claim! after handle returns; only completed
154
+ # executions enter the in-memory dedup cache.
128
155
  #
129
156
  # Legacy fallback: without the completed_at column (upgraded gem,
130
- # not-yet-migrated table) this degrades to the old single-phase claim.
131
- def claim_idempotency?(event_id)
157
+ # not-yet-migrated table) this degrades to the old single-phase claim,
158
+ # which has no pending state and therefore no ownership question.
159
+ def claim_idempotency(event_id)
132
160
  cache_key = dedup_key(event_id)
133
- return false if self.class.dedup_cache.seen?(cache_key)
161
+ return ClaimResult.new(status: :cached, age: nil) if self.class.dedup_cache.seen?(cache_key)
134
162
 
135
163
  result = ProcessedEvent.insert(
136
164
  { event_id: event_id, handler_class: self.class.name, processed_at: Time.now.utc },
@@ -139,18 +167,79 @@ module Pgbus
139
167
 
140
168
  unless ProcessedEvent.completion_column?
141
169
  self.class.dedup_cache.mark!(cache_key)
142
- return result.rows.any?
170
+ return ClaimResult.new(status: result.rows.any? ? :claimed : :completed, age: nil)
143
171
  end
144
172
 
145
- return true if result.rows.any?
173
+ return ClaimResult.new(status: :claimed, age: nil) if result.rows.any?
174
+
175
+ inspect_existing_claim(event_id, cache_key)
176
+ end
177
+
178
+ # The insert lost, so a row exists (or existed). `pick` returns nil for
179
+ # the whole row when it has since been purged — not a pending claim,
180
+ # nothing is running, so claim it.
181
+ def inspect_existing_claim(event_id, cache_key)
182
+ completed_at, processed_at = ProcessedEvent
183
+ .where(event_id: event_id, handler_class: self.class.name)
184
+ .pick(:completed_at, :processed_at)
185
+
186
+ if completed_at
187
+ self.class.dedup_cache.mark!(cache_key)
188
+ return ClaimResult.new(status: :completed, age: nil)
189
+ end
190
+
191
+ return ClaimResult.new(status: :claimed, age: nil) if processed_at.nil?
192
+
193
+ age = Time.now.utc - processed_at.to_time.utc
194
+ return ClaimResult.new(status: :owned, age: age) if age < claim_ownership_window
195
+
196
+ ClaimResult.new(status: :claimed, age: age)
197
+ end
198
+
199
+ # How long a pending claim may stay silent before its holder counts as
200
+ # dead. ClaimBeat refreshes a live claim from the visibility heartbeat,
201
+ # which lands every extension inside [interval, 1.5 * interval] of the
202
+ # previous one — two intervals leaves margin for a late beat without
203
+ # stretching the window past the visibility timeout it rides on.
204
+ #
205
+ # With the heartbeat disabled a claim is never refreshed, so the window
206
+ # degrades to "roughly the first two thirds of one visibility timeout
207
+ # after the claim" — a redelivery, which cannot arrive before the VT has
208
+ # lapsed, still re-runs exactly as it did before issue #470.
209
+ def claim_ownership_window
210
+ Pgbus.configuration.effective_visibility_heartbeat_interval * 2
211
+ end
146
212
 
147
- completed_at = ProcessedEvent
148
- .where(event_id: event_id, handler_class: self.class.name)
149
- .pick(:completed_at)
150
- return true if completed_at.nil? # pending claim (or purged row) → re-run
213
+ # Register this claim with the message's beat for exactly the duration of
214
+ # handle: before it, there is nothing to keep alive; after it,
215
+ # complete_claim! owns the row and a beat touching processed_at would
216
+ # race the completion stamp.
217
+ def with_claim_beat(claim_beat, event_id)
218
+ return yield unless claim_beat && self.class.idempotent? && ProcessedEvent.completion_column?
219
+
220
+ claim_beat.register(event_id, self.class.name)
221
+ begin
222
+ yield
223
+ ensure
224
+ claim_beat.release(event_id, self.class.name)
225
+ end
226
+ end
151
227
 
152
- self.class.dedup_cache.mark!(cache_key)
153
- false
228
+ # A skip used to be silent, which made an over-eager re-run (issue #470)
229
+ # invisible in production: nothing distinguished "deduplicated" from
230
+ # "deferred to a live holder". The claim age and read_ct are what tell
231
+ # an operator which one happened.
232
+ def instrument_skip(claim, event, message, routing_key)
233
+ Instrumentation.instrument(
234
+ "pgbus.event_skipped",
235
+ event_id: event.event_id,
236
+ handler: self.class.name,
237
+ routing_key: routing_key,
238
+ reason: claim.status,
239
+ claim_age: claim.age,
240
+ read_ct: message.read_ct.to_i,
241
+ msg_id: message.msg_id.to_i
242
+ )
154
243
  end
155
244
 
156
245
  # Phase 2: stamp the claim completed and only then admit it to the
@@ -165,11 +254,11 @@ module Pgbus
165
254
  # The stamp is an idempotent `SET completed_at = <now>`, so repeating a
166
255
  # statement that may already have committed is safe.
167
256
  #
168
- # Phase 1 (claim_idempotency?) is deliberately NOT wrapped. Its INSERT
257
+ # Phase 1 (claim_idempotency) is deliberately NOT wrapped. Its INSERT
169
258
  # may have committed before the socket died, and on a legacy schema the
170
259
  # retry's empty `result.rows` would read as "someone else owns this
171
- # claim" and return false — turning a recoverable drop into a silently
172
- # skipped event. VT redelivery is the correct recovery there.
260
+ # claim" and report :completed — turning a recoverable drop into a
261
+ # silently skipped event. VT redelivery is the correct recovery there.
173
262
  def complete_claim!(event_id)
174
263
  return unless ProcessedEvent.completion_column?
175
264
 
@@ -25,7 +25,7 @@ module Pgbus
25
25
  Pgbus::Testing.store.push_event(event)
26
26
 
27
27
  if Pgbus::Testing.inline? && delay.to_i <= 0
28
- Pgbus::EventBus::Registry.instance.handlers_for(routing_key).each do |subscriber|
28
+ Pgbus::EventBus::Registry.instance.subscribers_matching(routing_key).each do |subscriber|
29
29
  Pgbus::CurrentAttributes.restore(event.context) { subscriber.handler_class.new.handle(event) }
30
30
  end
31
31
  end
@@ -62,7 +62,30 @@ module Pgbus
62
62
  end
63
63
  end
64
64
 
65
- def handlers_for(routing_key)
65
+ # Subscribers a message read from +queue_name+ must be dispatched to
66
+ # (issue #469). Every subscriber gets its own queue, so a topic with N
67
+ # matching subscribers produces N queue copies of each event; selecting by
68
+ # pattern alone fanned every copy out to every match, running each handler
69
+ # N times per event — and, across hosts, concurrently. Ownership is the
70
+ # primary filter; the pattern check still applies because a routing key
71
+ # the owner's pattern does not match means a stale pgmq.topic_bindings
72
+ # row, and a stale binding must not run the handler.
73
+ #
74
+ # +queue_name+ is required on purpose: a keyword-less call would silently
75
+ # restore the fan-out this closed. Callers that genuinely want the
76
+ # pattern view (the Testing inline/drain paths, which never touch a
77
+ # queue) use #subscribers_matching.
78
+ def handlers_for(routing_key, queue_name:)
79
+ @subscribers.select do |s|
80
+ s.queue_name == queue_name && matches?(s.pattern, routing_key)
81
+ end
82
+ end
83
+
84
+ # Pattern-only selection, with no queue in play. Used by the Testing
85
+ # inline/drain paths, where the event never reaches PGMQ and each matching
86
+ # subscriber is invoked exactly once — the same per-subscriber delivery
87
+ # count owner-only dispatch produces in production.
88
+ def subscribers_matching(routing_key)
66
89
  @subscribers.select { |s| matches?(s.pattern, routing_key) }
67
90
  end
68
91
 
@@ -20,6 +20,10 @@ module Pgbus
20
20
  # payload: queue, job_class, msg_id, vt, extensions
21
21
  # pgbus.event_processed — event handler succeeded
22
22
  # pgbus.event_failed — event handler raised; carries :exception_object
23
+ # pgbus.event_unrouted — a consumer read an event from a queue no
24
+ # subscriber in this process owns (stale topic
25
+ # binding); the message is archived
26
+ # payload: queue_name, routing_key
23
27
  # pgbus.stream.broadcast — stream broadcast (sync or deferred)
24
28
  # pgbus.outbox.publish — outbox row created
25
29
  # pgbus.recurring.enqueue — scheduler enqueued a due recurring task