pgbus 0.16.7 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 82a6db71dd099a9b4fa24cb0a278184a3aed8bbee7acfb3f3cc6ef6b55064038
4
- data.tar.gz: ad7034ee934aee91afdb4a6278200f8178caf0da609d49e57bc6bbcf27d3b7b5
3
+ metadata.gz: 7b7955f6b4ccdd4d3e4f8d8e372e9283ca0b7f801d7a6909d29f03fde3425ce7
4
+ data.tar.gz: f26d692d0e9b979e9c4b77b3e28cafdab4d3c50e32292e793e1885d94f0d5943
5
5
  SHA512:
6
- metadata.gz: 86977d3f705b8e3732211956582334f6e49a60e426af1082fdf585a175c9a189c9dc9ec83d5ceeaac2f0be7dcee40c1414e55a99225fcc90ac955537d8b4f01c
7
- data.tar.gz: 3af2ae55db88774c89f3082b70e5456201c9a85ffd994d43beb197fe0f52d282209fc5de57db66c44de10afe34b37d37fa5404522eabb97d4eb2e91cbfc80da7
6
+ metadata.gz: 30c539f2494acaeffdd10301bba047d6dca86ab2bf3ae776ac6d8d3239f4b5e5a05d530135932cf15122b37dd02acd50de86026315365ea10c085e0100201ebf
7
+ data.tar.gz: 596d817366be305b4abd3c9f1975e6e004ef2448976f58bfb36c9c9b1cd15d0abbc830a98bfd9900d5fc2cfc07b2f6257c1f158b7f8b09a653e97e55c590cc92
data/CHANGELOG.md CHANGED
@@ -2,12 +2,26 @@
2
2
 
3
3
  ### Added
4
4
 
5
+ - **`rake bench:worker_profile` measures the event consumer too.** A consumer cell drives a real `Process::Consumer` over an event-bus topic queue (plain handler, so the transport is measured, not the dedup table) with the same harness as the worker cell. `WP_BENCH_ROLE` (`worker`, `consumer`, `both`) and `WP_BENCH_READ_AHEAD` (passed as `read_ahead:`) are new knobs. Refs #486.
6
+
5
7
  - **`coalesce:` now works on the `Turbo::StreamsChannel` broadcast path.** `Pgbus::Streams::Stream#broadcast` has accepted `coalesce:` since #171, but the Turbo patch forwarded only four options (`durable:`, `exclude:`, `visible_to:`, `event:`), so nothing that broadcasts through `Turbo::StreamsChannel.broadcast_*_to` — every `Turbo::Broadcastable` model, and phlex-reactive's `Streamable.broadcast_to` — could debounce. A background job settling a 177-row collection therefore pushed 177 idempotent replaces of the same count badge at every peer. Any targeted helper now takes `coalesce:` (`@order.broadcast_replace_to :account, coalesce: 100`), as does a direct channel call (`Turbo::StreamsChannel.broadcast_replace_to(@account, target: "order-count", coalesce: true)`), with the same semantics as the direct API: opt-in, per `(stream, target)`, last-write-wins, `true` = the 50ms default window, an integer = the window in ms. Two things the option's plumbing needed beyond a fifth thread-local: the coalescing key, which dedupes on `(stream, target)` while `broadcast_stream_to` only ever sees the rendered `content:` — it is now resolved from the frame's own `target:`/`targets:` through the same `ActionView::RecordIdentifier.dom_id` conversion Turbo applies, so a record target keys on `order_7` rather than on a per-instance `#to_s`; and extraction of the pgbus options at the channel itself, since a direct `Turbo::StreamsChannel` call never passes through `Turbo::Broadcastable` and so previously leaked `durable: true` (and now `coalesce:`) into Turbo's renderer. `broadcast_refresh_to` and `broadcast_render_to` carry no target and so cannot coalesce — they raise the existing actionable "requires target:" error rather than guessing a key — and the `_later_to` variants remain out of scope, as the thread-local cannot follow into a job. The wire is byte-identical when `coalesce:` is absent. Closes #465.
6
8
 
7
9
  - **Concurrency is now visible: a section on the Locks page, three gauges, a home stat card and an MCP tool.** `limits_concurrency … on_conflict: :block` became durable in #460 — a parked job is never dropped and a key never runs past `to:` — but nothing showed it, so an operator whose per-record pipeline looked "stuck" had no way to see that the key's slot was held, how many jobs were parked behind it, or how long the oldest had waited, short of SQL. `pgbus_semaphores` and `pgbus_blocked_executions` appeared on no page, in no gauge, in no tool. One new read, `Web::DataSource#concurrency_stats`, now feeds four surfaces: the Locks page's **Concurrency** section (parked jobs, oldest wait, slots held, keys at limit, then a row per key with value / limit, lease state, parked count and oldest wait, auto-refreshing like the other dashboard panels); a **Parked jobs** card on the dashboard home; three unlabelled gauges on both `/pgbus/api/metrics` and the AppSignal probe (`pgbus_concurrency_blocked_executions`, `pgbus_concurrency_blocked_oldest_age_seconds`, `pgbus_concurrency_slots_held` — keys are per-record, so a per-key label would be an unbounded series); and a read-only `pgbus_concurrency` MCP tool returning the same data. The key rows come from a FULL OUTER JOIN, not a LEFT one: a key can have parked jobs and no semaphore (its holder died and the sweep removed the row) or a semaphore and nothing parked, and both shapes matter to an operator. Each row carries two guarded escape hatches. **Release** drops the semaphore and promotes through `Concurrency::BlockedExecution.promote_next`, so it takes each slot through the same guarded upsert an enqueue uses and can never reintroduce the over-admission #460 closed; when the lease is still fresh the confirm says so, because a job is then probably still running and releasing lets another start beside it. **Discard parked** drops a single key's parked jobs, naming the exact count in the confirm, and resolves the bookkeeping those jobs will never resolve themselves — a parked batch child is marked failed so its batch stops waiting and `on_failure` fires (#413), and an `:until_executed` uniqueness lock is released rather than orphaned (#423); each discard logs a warning naming the job and emits `pgbus.blocked_execution_discarded`. There is deliberately no cross-key "discard all parked": not losing work is the whole point of `:block`. Nothing is added to the enqueue or execute path. Refs #461.
8
10
 
11
+ - **PGMQ: vendored schema v1.13.0, and the upgrade migration now carries upstream's table fixups.** The 1.12.0 → 1.13.0 delta is entirely partitioned-queue work plus one metrics attribute: `pgmq.create_partitioned` gains a `premake` argument (its three-argument form dropped, so pgmq-ruby's three-argument call binds to the new signature with the default), new partitioned queues use `msg_id … GENERATED BY DEFAULT AS IDENTITY` rather than `GENERATED ALWAYS`, and `pgmq.metrics_result` gains `default_partition_length` (NULL for non-partitioned queues, ignored by pgmq-ruby, which reads rows by name). Nothing pgbus calls through `Pgbus::Client` changes, and pgbus creates only partitioned *archives*, whose `msg_id` is a plain column — so no client or pgmq-ruby change. Fresh embedded installs get 1.13.0; existing *embedded* installs: `rake pgbus:pgmq:status`, then `rails generate pgbus:upgrade_pgmq` + `rails db:migrate` (`rails db:migrate:pgbus` on a separate-database install). An `:extension` install is upgraded by PostgreSQL itself (`ALTER EXTENSION pgmq UPDATE`) once the server has the new version packaged — not by this migration, whose function drop refuses to remove extension-owned objects. **New:** because the upgrade works by dropping every pgmq function and composite type and re-running the target schema, it could never carry an upstream change to an existing *table*. The generated migration now reads the version in `pgbus_pgmq_schema_versions` and applies the fixup vendored for each version between it and the target (`lib/pgbus/pgmq_schema/fixups/pgmq_v<VERSION>.sql`, discovered by the same glob as the schema files, via `PgmqSchema.fixups_sql`). This hop's fixup is upstream's own guarded block moving existing partitioned queues to `GENERATED BY DEFAULT`; it selects on `pgmq.meta.is_partitioned`, so it is a no-op unless partitioned queues were created through PGMQ directly. Fixups must be idempotent — an install with no recorded version gets all of them. Closes #459.
12
+
9
13
  - **The batches list and a batch's progress panel now refresh themselves.** Batches was the only dashboard list whose `turbo-frame` was missing `data-auto-refresh`, and the batch detail page had no frame at all, so watching a batch drain meant reloading by hand. Both now poll on the existing `web_refresh_interval` (5s, pauses while you interact) like every other list — the batch progress panel is served by `GET /batches/:id?frame=progress`. This reuses the dashboard's turbo-frame polling rather than the SSE stack in `lib/pgbus/web/stream_app.rb`: that is the host-app-facing stream feature, and pointing it at the dashboard would open a PG `LISTEN` connection per viewer to move four integers.
10
14
 
15
+ - **`pgbus doctor` reports the Ruby JIT, and the boot banner names it.** A new "Ruby JIT" check is `:ok` when YJIT (or ZJIT) is on, when the Ruby has no JIT, or in a local (development/test) environment, and warns otherwise — never fails, never enables anything; turning YJIT on stays with Rails' `config.yjit`. The supervisor's first boot line now ends `jit=yjit|zjit|none`. Measured reason: with Postgres close by, a worker is CPU-bound under the GVL and YJIT is worth ~14–20 % jobs/s (`docs/performance.md`). Refs #484.
16
+
17
+ - **`rake bench:worker_profile`: jobs/s and where a job's time goes, for a real worker.** Drives a real `Pgbus::Process::Worker` against real PGMQ and splits each thread's time (Postgres wait, pool-checkout wait, waiting for work, GVL wait, and whose code was on-CPU) with vernier, locally and optionally behind toxiproxy-injected latency. `vernier` joins the test group; it is not a runtime dependency. The same change repairs `benchmarks/integration_bench.rb`, which no longer ran against 1.0's `ensures_uniqueness`. Refs #484.
18
+
19
+ ### Changed
20
+
21
+ - **perf(worker, consumer): opt-in read-ahead, so one read feeds many jobs when Postgres is on another host.** A worker read `qty = free threads`; once its pool was busy, threads freed one at a time, so it made one read round trip per job. Under +1 ms of latency that capped a 12-thread worker at ~1.3–1.7 k no-op jobs/s however many threads it had (#484), and the event consumer's loop had the same shape. New `config.read_ahead` (default **0**, off; overridable per capsule and per `event_consumers` entry with `read_ahead:`) lets either loop claim up to N messages beyond its free threads. Claimed messages wait in a new `Process::ClaimBuffer`, kept invisible by a new `VisibilityHeartbeat.hold`, until a thread frees. They are handed back to the queue (`set_vt` 0) the moment the worker drains, recycles, pauses or stops, or the consumer shuts down or recycles, with `read_ct` already counted. Buffered claims count as in flight, so `prefetch_limit` still caps everything claimed. Measured with `rake bench:worker_profile` (12 threads, no-op jobs, +1 ms per round trip, rounds 2–3 of three interleaved rounds on a loaded machine): **worker 1 286–1 731 → 2 300–2 462 jobs/s, consumer 1 326–1 779 → 2 288–2 599** at `read_ahead` 12. Locally both are within noise, because there the worker is CPU-bound under the GVL; `read_ahead` 0 overlaps the baseline's run-to-run range in 7 of 8 cells; consumer / proxied / no JIT sits 2–18 % below, plausibly load noise but unproven. A returned message's next claimant logs it as a zombie redelivery when `zombie_detection` is on, and `validate!` warns when `read_ahead > 0` meets `visibility_heartbeat = false`, since nothing then keeps a buffered message invisible. With read-ahead on, the next ceiling under latency is the connection pool (checkout wait 2–3 % → 14–15 %), so raise `pool_size` with it. `VisibilityHeartbeat`'s unregister now drops only the entry it registered, so a stale hold released after a job's own tracking took over cannot drop that job's heartbeat. Worker heartbeat metadata gains `"buffered"`. Refs #486.
22
+
23
+ - **perf(executor): a successful first delivery no longer deletes from `pgbus_failed_events`.** Every successful job ran a `DELETE … WHERE queue_name = $1 AND msg_id = $2` through ActiveRecord, but only a failed attempt ever writes that table, so a job on its first delivery cannot have a row. The DELETE now runs only on redelivery (`read_ct > 1`), still clearing the row a failed attempt left. Measured with `rake bench:worker_profile` (12 threads, no-op jobs, local Postgres): **4 569 → 5 823 jobs/s with YJIT (+27 %), 3 797 → 5 119 without (+35 %)**; under +1 ms injected latency the change is within noise, because there the worker's single reader is the ceiling. This is a per-job query removed, not faster job code. The executor's debug tag is also built only when debug logging is on (9 → 8 pgbus-owned objects per job, now a hard budget). `Client#drop_queue` now also deletes the dropped queue's `pgbus_failed_events` rows: PGMQ restarts `msg_id`s when a queue is recreated, and a surviving row would point the dashboard's retry/discard at an unrelated new message — previously a first-delivery success happened to sweep such a row, and that sweep is gone. Refs #484.
24
+
11
25
  ### Fixed
12
26
 
13
27
  - **A still-running event handler is no longer re-run as if it had crashed, and a slow one is no longer redelivered mid-run.** Two halves of one hole, follow-up to #469. First: `Process::Consumer` had no equivalent of `ActiveJob::Executor#with_visibility_heartbeat`, so an event message's visibility timeout was never re-armed — a handler slower than `visibility_timeout` (30s by default) was redelivered *while it was still running*, `read_ct` climbed on every redelivery, and after `max_retries` the event was dead-lettered without the handler ever raising. Workers have been protected from this since the heartbeat landed; event consumers never were. The consumer now tracks each message through `VisibilityHeartbeat` for exactly as long as its handlers run, releasing the entry before the archive so a beat can never re-arm a message that is already gone. Second: `Handler#claim_idempotency?` read a `pgbus_processed_events` row with `completed_at IS NULL` as proof that the previous holder had been killed mid-handler, and re-ran. That state equally describes a handler that is simply still running — on another thread, another fork or another host — so a second delivery (a genuine redelivery, or a duplicate envelope, whose independent visibility timeouts no heartbeat can serialize) executed the handler *concurrently with* the live one, which is precisely the double-execution `idempotent!` exists to prevent. The claim now carries liveness rather than only a claim instant: `EventBus::ClaimBeat` refreshes every in-flight claim's `processed_at` from the same beat that re-arms the message's visibility timeout, so message visibility and claim liveness go quiet together when a process dies. A pending claim silent for longer than twice the heartbeat interval is abandoned and re-runs, exactly as before; a fresher one is owned and the delivery skips, deferring to the holder — which either completes (nothing is lost) or fails, leaving its own message for visibility-timeout redelivery to recover. A skip is also no longer silent: `pgbus.event_skipped` carries the reason (`:completed`, `:cached` or `:owned`), the claim's age in seconds and the delivery's `read_ct`, and the metrics subscriber counts it as `pgbus_event_count` with `status: "skipped"` and the reason as a tag. No migration, no new column, and no change on a table that has not run `pgbus:add_processed_event_completion` — a single-phase claim has no pending state and therefore no ownership question. Closes #470.
data/README.md CHANGED
@@ -26,6 +26,7 @@ PostgreSQL-native job processing and event bus for Rails, built on [PGMQ](https:
26
26
  - [Client-level circuit breaker (database-down)](#client-level-circuit-breaker-database-down)
27
27
  - [Read timeouts (libpq-native)](#read-timeouts-libpq-native)
28
28
  - [Prefetch flow control](#prefetch-flow-control)
29
+ - [Read-ahead](#read-ahead)
29
30
  - [Worker recycling](#worker-recycling)
30
31
  - [Retry backoff](#retry-backoff)
31
32
  - [Routing and ordering](#routing-and-ordering)
@@ -541,6 +542,24 @@ end
541
542
 
542
543
  The worker tracks in-flight messages with an atomic counter and only fetches `min(idle_threads, prefetch_available)` messages per cycle. The counter is decremented in an `ensure` block so it never gets stuck.
543
544
 
545
+ ### Read-ahead
546
+
547
+ By default a worker reads as many messages as it has free threads. Once the pool is busy, threads free one at a time, so the worker ends up making one read round trip per job. That costs nothing when Postgres is on the same host. When Postgres is on another host, that round trip becomes the throughput ceiling. `read_ahead` lets a worker or event consumer claim up to N messages beyond its free threads and hold them until a thread frees:
548
+
549
+ ```ruby
550
+ Pgbus.configure do |config|
551
+ config.read_ahead = 12 # global; 0 = off (default)
552
+ config.capsule :api, queues: %w[api], threads: 12, read_ahead: 12 # per capsule
553
+ config.event_consumers = [{ topics: ["orders.#"], threads: 8, read_ahead: 8 }]
554
+ end
555
+ ```
556
+
557
+ - **When to turn it on:** when Postgres is not on the same host as the worker. Start with `read_ahead` equal to `threads`. With a local database it buys little, because the worker is CPU-bound there. Under +1 ms of latency, `read_ahead` 12 took a 12-thread worker from ~1.3–1.7k to ~2.3–2.5k no-op jobs/s (`docs/performance.md`, "Read-ahead"). The next ceiling is the connection pool, so raise `pool_size` with it.
558
+ - **Visibility:** buffered messages are kept invisible by the visibility heartbeat until their job starts, so a long wait never lets another worker claim them. With `visibility_heartbeat = false` nothing holds them, and a message buffered past `visibility_timeout` can run twice; `validate!` warns about that combination.
559
+ - **Shutdown and recycle:** when a worker drains, recycles, pauses or shuts down, buffered messages go straight back to the queue (`set_vt` 0), not after `visibility_timeout`. Their `read_ct` has already been incremented, the same as after a crash, so with `zombie_detection` on, the worker that picks one up logs it as a zombie redelivery.
560
+ - **With `prefetch_limit`:** buffered messages count as in flight, so `prefetch_limit` still caps everything a worker has claimed.
561
+ - **Multi-queue workers:** the plain multi-queue read (no priority, fair share or group mode) can claim and discard rows past its limit. A larger read makes that worse; see `docs/performance.md`.
562
+
544
563
  ### Worker recycling
545
564
 
546
565
  Pgbus workers recycle themselves to prevent memory bloat. This is the main reliability difference vs. solid_queue, which leaves workers alive forever.
@@ -2031,11 +2050,12 @@ pgbus help # Show help
2031
2050
 
2032
2051
  #### pgbus doctor
2033
2052
 
2034
- A single preflight command that answers "is this environment healthy enough to run?" — useful as a deploy or CI gate. It runs 11 checks and never raises; a broken environment turns every probe into a failed/warned check instead of a crash:
2053
+ A single preflight command that answers "is this environment healthy enough to run?" — useful as a deploy or CI gate. It runs 12 checks and never raises; a broken environment turns every probe into a failed/warned check instead of a crash:
2035
2054
 
2036
2055
  | Check | Fails (`:fail`) when | Warns (`:warn`) when |
2037
2056
  |---|---|---|
2038
2057
  | Configuration | `Configuration#validate!` raises | — |
2058
+ | Ruby JIT | — | YJIT is available but off outside a local environment (a CPU-bound worker with Postgres nearby runs ~15% fewer jobs/s without it; no measurable effect when network-bound); reports only, never enables |
2039
2059
  | Database | Unreachable (`SELECT 1` via `Client#ping`) | — |
2040
2060
  | PGMQ schema | Schema not installed | Installed but untracked, or behind the vendored version |
2041
2061
  | Queues | A configured queue has no PGMQ table | — |
data/Rakefile CHANGED
@@ -16,7 +16,7 @@ require "rubocop/rake_task"
16
16
  # on the unresolvable gem inheritance. Passing explicit paths stops the
17
17
  # discovery. docs/ lints itself in the Docs site workflow. (Same fix as docs-kit.)
18
18
  RuboCop::RakeTask.new do |task|
19
- task.patterns = %w[app benchmarks config lib spec Gemfile Rakefile pgbus.gemspec]
19
+ task.patterns = %w[app benchmarks config lib rakelib spec Gemfile Rakefile pgbus.gemspec]
20
20
  end
21
21
 
22
22
  namespace :bench do
@@ -25,7 +25,8 @@ namespace :bench do
25
25
  # no-DB unit suite that bench:all runs in CI.
26
26
  db_benches = %w[connection_pool_bench integration_bench streams_bench streams_read_pool_bench
27
27
  execution_modes_bench pool_swap_bench pool_autoscale_bench job_burst_bench
28
- notify_wake_bench notify_chaos_bench streams_hub_bench fair_read_bench].freeze
28
+ notify_wake_bench notify_chaos_bench streams_hub_bench fair_read_bench
29
+ worker_profile_bench].freeze
29
30
  # The unit suite is every *_bench.rb that doesn't need a database, derived
30
31
  # from the directory so a new unit bench is picked up automatically (kept in
31
32
  # sync with bench:one, which globs the same files).
@@ -60,6 +61,11 @@ namespace :bench do
60
61
  ruby "benchmarks/integration_bench.rb"
61
62
  end
62
63
 
64
+ desc "Run worker profiling bench: jobs/s and where a job's time goes (requires PGBUS_DATABASE_URL)"
65
+ task :worker_profile do
66
+ ruby "benchmarks/worker_profile_bench.rb"
67
+ end
68
+
63
69
  desc "Run fair share read benchmark (requires PGBUS_DATABASE_URL)"
64
70
  task :fair_read do
65
71
  ruby "benchmarks/fair_read_bench.rb"
@@ -164,173 +170,8 @@ task :build do
164
170
  sh("rm -rf /tmp/gem-verify #{gem_file}")
165
171
  end
166
172
 
167
- desc "Release a new version (rake release[1.2.3] or rake release[pre] or rake release[1.2.3,force])"
168
- task :release, %i[version force] do |_t, args|
169
- require_relative "lib/pgbus/version"
170
-
171
- def info(msg) = puts "\e[34m→\e[0m #{msg}"
172
- def success(msg) = puts "\e[32m✓\e[0m #{msg}"
173
- def skip(msg) = puts "\e[33m⊘\e[0m #{msg} \e[33m(skipped)\e[0m"
174
- def warn(msg) = puts "\e[33m⚠\e[0m #{msg}"
175
- def error(msg) = puts "\e[31m✗\e[0m #{msg}"
176
- def header(msg) = puts "\n\e[1;36m#{msg}\e[0m\n#{"─" * msg.length}"
177
-
178
- new_version = args[:version]
179
- abort "\e[31mUsage: rake release[X.Y.Z] or rake release[X.Y.Z,force]\e[0m" unless new_version
180
-
181
- force = args[:force]&.to_s&.downcase == "force"
182
-
183
- dirty = `git status --porcelain`.strip
184
- abort "\e[31mAborting: working directory is not clean.\e[0m\n#{dirty}" unless dirty.empty?
185
-
186
- current = Pgbus::VERSION
187
- prerelease = new_version.match?(/alpha|beta|rc|pre/) || new_version == "pre"
188
-
189
- if new_version == "pre"
190
- new_version = current
191
- prerelease = true
192
- end
193
-
194
- tag = "v#{new_version}"
195
- version_file = "lib/pgbus/version.rb"
196
-
197
- title = "Release #{tag}"
198
- title += " (force)" if force
199
- header title
200
- info "Current version: #{current}"
201
- info "New version: #{new_version}"
202
- info "Pre-release: #{prerelease}"
203
-
204
- # Step 0: Force cleanup — delete existing release and tag
205
- if force
206
- header "Force cleanup"
207
- if system("gh release view #{tag} >/dev/null 2>&1")
208
- sh("gh release delete #{tag} --yes --cleanup-tag")
209
- success "Deleted release and remote tag #{tag}"
210
- else
211
- skip "No release #{tag} to delete"
212
- end
213
-
214
- if system("git rev-parse #{tag} >/dev/null 2>&1")
215
- sh("git tag -d #{tag}")
216
- success "Deleted local tag #{tag}"
217
- else
218
- skip "No local tag #{tag} to delete"
219
- end
220
- end
221
-
222
- # Step 1: Update version file
223
- header "Version"
224
- if new_version == current
225
- skip "Version already #{new_version}"
226
- else
227
- content = File.read(version_file)
228
- content.sub!(/VERSION = ".*"/, "VERSION = \"#{new_version}\"")
229
- File.write(version_file, content)
230
- success "Updated #{version_file}"
231
- end
232
-
233
- # Step 1b: Regenerate the frozen lockfiles that pin the pgbus path gem, so the
234
- # bump ships with them in sync. These are installed with `--frozen`/deployment
235
- # in CI, so if they still name the OLD version they instant-fail (the root
236
- # Gemfile.lock on every main-Gemfile leg AND release.yml's own `bundle install`,
237
- # the Rails 7.1 leg with exit 16, and docs-CI on any docs change). Regenerating
238
- # here keeps the version-pin drift out of the release commit instead of
239
- # surfacing on the next PR — or, worse, in the Release workflow itself.
240
- header "Frozen lockfiles"
241
- # The ONLY thing a version bump changes in these frozen lockfiles is the pgbus
242
- # path-gem pin — so bump exactly that line, in place, with a string edit.
243
- #
244
- # We deliberately do NOT run `bundle lock` here: a full re-resolve trips over
245
- # constraints that have nothing to do with pgbus. Concretely, docs/Gemfile.lock
246
- # carries a broad PLATFORMS list (…-gnu / …-musl / arm-linux) for which a
247
- # platform gem like `thruster` ships no variant, so `bundle lock` fails with
248
- # "Could not find gems matching 'thruster' valid for all resolution platforms"
249
- # on any machine whose cache doesn't already hold those exact gems — aborting
250
- # the release. `bundle lock --local` was even worse (wrong-file write + no
251
- # fetch). A targeted pin edit sidesteps all of it, is deterministic on any
252
- # machine, and produces the minimal 2-line diff (the PATH spec + the
253
- # DEPENDENCIES pin). See #338/#341 and the surgical-bump fix.
254
- frozen_lockfiles = %w[Gemfile.lock gemfiles/rails_7_1.gemfile.lock docs/Gemfile.lock]
255
- regenerated_lockfiles = []
256
- frozen_lockfiles.each do |lockfile|
257
- unless File.exist?(lockfile)
258
- skip "#{lockfile} not present"
259
- next
260
- end
261
-
262
- content = File.read(lockfile)
263
- # Matches both the PATH-source spec (" pgbus (X.Y.Z)") and the
264
- # DEPENDENCIES pin (" pgbus (X.Y.Z)"), leaving everything else untouched.
265
- bumped = content.gsub(/^(\s+pgbus) \([^)]*\)$/, "\\1 (#{new_version})")
266
-
267
- if bumped == content
268
- skip "#{lockfile} — no pgbus pin to bump"
269
- next
270
- end
271
-
272
- File.write(lockfile, bumped)
273
- regenerated_lockfiles << lockfile
274
- success "Bumped pgbus pin in #{lockfile}"
275
- end
276
-
277
- # Step 2: Verify gem builds cleanly
278
- header "Build verification"
279
- sh("gem build pgbus.gemspec --strict")
280
- sh("rm -f pgbus-*.gem")
281
- success "Gem builds cleanly"
282
-
283
- # Step 3: Commit version bump (+ any re-synced lockfiles)
284
- header "Git commit"
285
- paths_to_stage = [version_file, *regenerated_lockfiles]
286
- staged_changes = paths_to_stage.any? do |path|
287
- !`git diff #{path}`.strip.empty? || !`git diff --cached #{path}`.strip.empty?
288
- end
289
- if staged_changes
290
- paths_to_stage.each { |path| sh("git add #{path}") }
291
- sh("git commit -m 'chore: bump version to #{new_version}'")
292
- success "Committed version bump"
293
- else
294
- skip "No version change to commit"
295
- end
296
-
297
- # Step 4: Push to origin
298
- header "Git push"
299
- local_sha = `git rev-parse HEAD`.strip
300
- remote_sha = `git rev-parse origin/main 2>/dev/null`.strip
301
- if local_sha == remote_sha
302
- skip "origin/main already at #{local_sha[0..6]}"
303
- else
304
- sh("git push origin main")
305
- success "Pushed to origin/main"
306
- end
307
-
308
- # Step 5: Create release
309
- header "Release"
310
- tag_exists = system("git rev-parse #{tag} >/dev/null 2>&1")
311
- release_exists = system("gh release view #{tag} >/dev/null 2>&1")
312
-
313
- if release_exists
314
- skip "Release #{tag} already exists (use force to re-create)"
315
- elsif tag_exists
316
- info "Tag #{tag} exists, creating release from it"
317
- pre_flag = prerelease ? "--prerelease" : ""
318
- sh("gh release create #{tag} --generate-notes #{pre_flag}".strip)
319
- success "Release #{tag} created from existing tag"
320
- else
321
- pre_flag = prerelease ? "--prerelease" : ""
322
- sh("gh release create #{tag} --generate-notes --target main #{pre_flag}".strip)
323
- success "Release #{tag} created"
324
- end
325
-
326
- puts ""
327
- success "\e[1mRelease #{tag} complete!\e[0m CI will handle the rest:"
328
- puts " • Run tests"
329
- puts " • Build + verify gem"
330
- puts " • Sign with Sigstore"
331
- puts " • Publish to RubyGems"
332
- puts " • Upload assets to the release"
333
- end
173
+ # `rake release[X.Y.Z]` lives in rakelib/release.rake (shared across the
174
+ # zoolutions gems); `bin/release` is its interactive front door.
334
175
 
335
176
  namespace :frontend do
336
177
  # app/frontend/pgbus/style.css is a committed build artifact — the engine ships
@@ -8,6 +8,16 @@ class UpgradePgmqToV<%= target_version_slug.camelize %> < ActiveRecord::Migratio
8
8
  # This uses the vendored SQL which doesn't require the pgmq extension.
9
9
  execute Pgbus::PgmqSchema.install_sql("<%= target_version %>")
10
10
 
11
+ # Apply the table fixups upstream ships for every hop between the
12
+ # version this database records and the target. Dropping and re-creating
13
+ # the functions cannot carry a change to an existing TABLE (1.13.0 moves
14
+ # partitioned queues' msg_id to GENERATED BY DEFAULT), so those ride
15
+ # here. Runs after install_sql because the fixups call
16
+ # pgmq.format_table_name, which the drop step removed. Each fixup is
17
+ # idempotent, so an install with no recorded version gets them all.
18
+ fixups = Pgbus::PgmqSchema.fixups_sql(after: installed_pgmq_version, upto: "<%= target_version %>")
19
+ execute fixups unless fixups.empty?
20
+
11
21
  # Re-install the per-queue NOTIFY insert triggers the function drop
12
22
  # cascaded away (pgmq.notify_insert_throttle records which queues had
13
23
  # them and at what throttle; enable_notify_insert is idempotent).
@@ -34,4 +44,16 @@ class UpgradePgmqToV<%= target_version_slug.camelize %> < ActiveRecord::Migratio
34
44
  "PGMQ schema downgrade is not supported. " \
35
45
  "Restore from a database backup if needed."
36
46
  end
47
+
48
+ private
49
+
50
+ # The version this database last recorded, or nil when it has never been
51
+ # tracked (a pre-tracking install, or one bootstrapped by hand).
52
+ def installed_pgmq_version
53
+ return nil unless connection.table_exists?("pgbus_pgmq_schema_versions")
54
+
55
+ connection.select_value(
56
+ "SELECT version FROM pgbus_pgmq_schema_versions ORDER BY installed_at DESC LIMIT 1"
57
+ )
58
+ end
37
59
  end
@@ -19,8 +19,7 @@ module Pgbus
19
19
 
20
20
  def execute(message, queue_name, source_queue: nil)
21
21
  execution_start = monotonic_now
22
- tag = "msg_id=#{message.msg_id} queue=#{queue_name} read_ct=#{message.read_ct}"
23
- Pgbus.logger.debug { "[Pgbus::Executor] start #{tag}" }
22
+ Pgbus.logger.debug { "[Pgbus::Executor] start #{log_tag(message, queue_name)}" }
24
23
 
25
24
  payload = JSON.parse(message.message)
26
25
  job_class = payload["job_class"]
@@ -42,7 +41,7 @@ module Pgbus
42
41
  read_ct: read_count,
43
42
  msg_id: message.msg_id.to_i
44
43
  )
45
- Pgbus.logger.debug { "[Pgbus::Executor] dead_lettered #{tag} job_class=#{job_class}" }
44
+ Pgbus.logger.debug { "[Pgbus::Executor] dead_lettered #{log_tag(message, queue_name)} job_class=#{job_class}" }
46
45
  return :dead_lettered
47
46
  end
48
47
  uniqueness_key = Uniqueness.extract_key(payload)
@@ -70,7 +69,7 @@ module Pgbus
70
69
  end
71
70
  end
72
71
 
73
- Pgbus.logger.debug { "[Pgbus::Executor] deserialized #{tag} job_class=#{job_class}" }
72
+ Pgbus.logger.debug { "[Pgbus::Executor] deserialized #{log_tag(message, queue_name)} job_class=#{job_class}" }
74
73
  job_succeeded = false
75
74
  retried = false
76
75
 
@@ -95,13 +94,13 @@ module Pgbus
95
94
  # job data, so hand it to the job before perform — that is what makes
96
95
  # `batch` (and `batch.enqueue` for open batches) work inside a job.
97
96
  assign_batch_id(job, payload)
98
- Pgbus.logger.debug { "[Pgbus::Executor] running #{tag} job_class=#{job_class}" }
97
+ Pgbus.logger.debug { "[Pgbus::Executor] running #{log_tag(message, queue_name)} job_class=#{job_class}" }
99
98
  with_visibility_heartbeat(job, queue_name, msg_id, source_queue, payload) { execute_job(job) }
100
99
  # retry_on re-enqueues from inside perform_now and returns normally:
101
100
  # this attempt is done (archive it) but the job is not — the retry
102
101
  # message carries the batch tag and signals on its own outcome.
103
102
  retried = Batch.retry_reenqueued?(payload["job_id"])
104
- Pgbus.logger.debug { "[Pgbus::Executor] perform_returned #{tag} job_class=#{job_class}" }
103
+ Pgbus.logger.debug { "[Pgbus::Executor] perform_returned #{log_tag(message, queue_name)} job_class=#{job_class}" }
105
104
  # Archiving is the exact-once claim on this execution.
106
105
  # :already_archived means another worker archived the message (our
107
106
  # heartbeat lapsed and it was redelivered): that worker owns the
@@ -109,20 +108,24 @@ module Pgbus
109
108
  # concurrency slot a second time (rails/solid_queue#761).
110
109
  if archive_from(queue_name, msg_id, source_queue: source_queue) == :already_archived
111
110
  Pgbus.logger.warn do
112
- "[Pgbus::Executor] already archived elsewhere, skipping signals #{tag} job_class=#{job_class}"
111
+ "[Pgbus::Executor] already archived elsewhere, skipping signals #{log_tag(message, queue_name)} job_class=#{job_class}"
113
112
  end
114
113
  release_duplicate_execution_lock(uniqueness_key, uniqueness_strategy, queue_name, msg_id)
115
114
  return :duplicate
116
115
  end
117
- Pgbus.logger.debug { "[Pgbus::Executor] archived #{tag} job_class=#{job_class}" }
116
+ Pgbus.logger.debug { "[Pgbus::Executor] archived #{log_tag(message, queue_name)} job_class=#{job_class}" }
118
117
  job_succeeded = true
119
118
  release_uniqueness_lock(uniqueness_key)
120
- FailedEventRecorder.clear!(queue_name: queue_name, msg_id: msg_id)
119
+ # Only FailedEventRecorder.record! writes this table, and only after a
120
+ # failed attempt, so a first delivery has no row: skip the per-job
121
+ # DELETE round trip (issue #484). A stale row from a dropped-and-
122
+ # recreated queue reusing this msg_id is the dispatcher's to sweep.
123
+ FailedEventRecorder.clear!(queue_name: queue_name, msg_id: msg_id) if read_count > 1
121
124
  end
122
125
 
123
126
  instrument("pgbus.job_completed", queue: queue_name, job_class: job_class)
124
127
  record_stat(payload, queue_name, "success", execution_start, message: message)
125
- Pgbus.logger.debug { "[Pgbus::Executor] done #{tag} job_class=#{job_class}" }
128
+ Pgbus.logger.debug { "[Pgbus::Executor] done #{log_tag(message, queue_name)} job_class=#{job_class}" }
126
129
  :success
127
130
  rescue *FATAL_EXCEPTIONS
128
131
  # Process-fatal: propagate so the supervisor/OS can react.
@@ -150,7 +153,7 @@ module Pgbus
150
153
  exception_object: e
151
154
  )
152
155
  record_stat(payload, queue_name, "failed", execution_start, message: message)
153
- Pgbus.logger.debug { "[Pgbus::Executor] failed #{tag} job_class=#{payload&.dig("job_class")} error=#{e.class}" }
156
+ Pgbus.logger.debug { "[Pgbus::Executor] failed #{log_tag(message, queue_name)} job_class=#{payload&.dig("job_class")} error=#{e.class}" }
154
157
  # Don't signal concurrency on transient failure — the job will be retried.
155
158
  # Semaphore is released only on success or dead-lettering.
156
159
  :failed
@@ -170,6 +173,12 @@ module Pgbus
170
173
 
171
174
  private
172
175
 
176
+ # Built inside the logger blocks so a job pays for the string only when
177
+ # debug logging is on (issue #484).
178
+ def log_tag(message, queue_name)
179
+ "msg_id=#{message.msg_id} queue=#{queue_name} read_ct=#{message.read_ct}"
180
+ end
181
+
173
182
  def assign_batch_id(job, payload)
174
183
  batch_id = payload[Batch::METADATA_KEY]
175
184
  return unless batch_id && job.respond_to?(:batch_id=)
data/lib/pgbus/client.rb CHANGED
@@ -699,6 +699,9 @@ module Pgbus
699
699
  synchronized { @pgmq.drop_queue(name) }
700
700
  end
701
701
  @queues_created.delete(name)
702
+ # Failed-event rows are keyed by the name the executor saw (physical or
703
+ # logical); a recreated queue reuses msg_ids, so neither may survive.
704
+ FailedEventRecorder.clear_queue!([name, name.delete_prefix("#{config.queue_prefix}_")].uniq)
702
705
  result
703
706
  end
704
707
 
@@ -12,6 +12,12 @@ module Pgbus
12
12
 
13
13
  # Worker settings
14
14
  attr_accessor :polling_interval, :prefetch_limit, :execution_mode
15
+ # read_ahead (issue #486): how many messages a worker or event consumer
16
+ # claims beyond its free threads and holds, heartbeated, until a thread
17
+ # frees. Under database latency one read then feeds many jobs instead of
18
+ # one. 0 (default) is off. Overridable per capsule / event_consumers entry
19
+ # (`read_ahead:`); prefetch_limit still caps everything claimed.
20
+ attr_reader :read_ahead
15
21
  # visibility_heartbeat / visibility_heartbeat_interval: while a job runs,
16
22
  # its message's visibility timeout is re-armed every interval seconds
17
23
  # (default: a third of visibility_timeout), so a job that outlives the
@@ -277,6 +283,7 @@ module Pgbus
277
283
  @visibility_heartbeat_interval = nil
278
284
 
279
285
  @prefetch_limit = nil
286
+ @read_ahead = 0
280
287
  @execution_mode = :threads
281
288
 
282
289
  @max_jobs_per_worker = nil
@@ -573,6 +580,16 @@ module Pgbus
573
580
  (0...priority_levels).map { |p| priority_queue_name(name, p) }
574
581
  end
575
582
 
583
+ def read_ahead=(value)
584
+ @read_ahead = value.nil? ? 0 : value
585
+ end
586
+
587
+ # Read-ahead for one capsule or event_consumers entry, falling back to
588
+ # the global read_ahead (mirrors execution_mode_for).
589
+ def read_ahead_for(entry)
590
+ entry.fetch(:read_ahead, nil) || read_ahead
591
+ end
592
+
576
593
  # Returns the execution mode for a specific worker config hash,
577
594
  # falling back to the global execution_mode setting.
578
595
  def execution_mode_for(worker_config)
@@ -828,6 +845,8 @@ module Pgbus
828
845
  "prefetch_limit must be > 0"
829
846
  end
830
847
 
848
+ validate_read_ahead!
849
+
831
850
  if priority_levels && !(priority_levels.is_a?(Integer) && priority_levels >= 1 && priority_levels <= 10)
832
851
  raise Pgbus::ConfigurationError, "priority_levels must be an integer between 1 and 10"
833
852
  end
@@ -846,6 +865,33 @@ module Pgbus
846
865
  self
847
866
  end
848
867
 
868
+ # Global, per-capsule and per-consumer read_ahead: a non-negative Integer.
869
+ def validate_read_ahead!
870
+ settings = [["read_ahead", read_ahead]]
871
+ Array(workers).each { |w| settings << ["worker read_ahead", w[:read_ahead]] unless w[:read_ahead].nil? }
872
+ Array(event_consumers).each { |c| settings << ["event consumer read_ahead", c[:read_ahead]] unless c[:read_ahead].nil? }
873
+
874
+ settings.each do |name, value|
875
+ next if value.is_a?(Integer) && !value.negative?
876
+
877
+ raise Pgbus::ConfigurationError, "#{name} must be a non-negative Integer (got #{value.inspect})"
878
+ end
879
+ warn_read_ahead_without_heartbeat(settings)
880
+ end
881
+
882
+ # Buffered claims are kept invisible only by the visibility heartbeat.
883
+ # Without it a message buffered longer than visibility_timeout can be
884
+ # claimed by another worker while it still waits here. Legal, but warned.
885
+ def warn_read_ahead_without_heartbeat(settings)
886
+ return if visibility_heartbeat
887
+ return unless settings.any? { |_name, value| value.positive? }
888
+
889
+ Pgbus.logger.warn do
890
+ "[Pgbus] read_ahead is on but visibility_heartbeat is false — a message buffered longer than " \
891
+ "visibility_timeout (#{visibility_timeout}s) can be claimed and run by another worker"
892
+ end
893
+ end
894
+
849
895
  # An explicit shutdown_timeout must be a positive number; nil keeps the
850
896
  # derived drain_timeout + margin default. A value below drain_timeout is
851
897
  # legal but self-defeating (the supervisor SIGKILLs workers mid-drain), so
data/lib/pgbus/doctor.rb CHANGED
@@ -5,7 +5,7 @@ require "pgbus/mcp/health_analyzer"
5
5
 
6
6
  module Pgbus
7
7
  # Preflight diagnostics for a pgbus deployment — the single command that
8
- # answers "is this environment healthy enough to run?". Runs eleven checks and
8
+ # answers "is this environment healthy enough to run?". Runs twelve checks and
9
9
  # returns a machine-readable result plus a human report, so `pgbus doctor`
10
10
  # and `rake pgbus:doctor` can gate a deploy or CI run (exit 0 on success,
11
11
  # 1 on any failure).
@@ -32,6 +32,7 @@ module Pgbus
32
32
  # stable identity of each check (used in the report and to select subsets).
33
33
  CHECKS = {
34
34
  "Configuration" => :check_configuration,
35
+ "Ruby JIT" => :check_ruby_jit,
35
36
  "Database" => :check_database,
36
37
  "PGMQ schema" => :check_pgmq_schema,
37
38
  "Queues" => :check_queues,
@@ -166,6 +167,33 @@ module Pgbus
166
167
  Check.new(name: "Configuration", status: :fail, detail: "#{e.class}: #{e.message}")
167
168
  end
168
169
 
170
+ # 1b. Ruby JIT (issue #484). With Postgres close by, a worker is GVL-bound
171
+ # and YJIT is worth ~15-20% jobs/s (docs/performance.md). Rails enables it via
172
+ # config.yjit (load_defaults 7.2+, or 8.0+ outside local envs). Reports
173
+ # only — enabling a JIT is the application's decision, never pgbus's.
174
+ def check_ruby_jit
175
+ label = RubyJit.label
176
+ return Check.new(name: "Ruby JIT", status: :ok, detail: "#{label.upcase} enabled") unless label == "none"
177
+ return Check.new(name: "Ruby JIT", status: :ok, detail: "no YJIT available on this Ruby") unless RubyJit.yjit_available?
178
+ return Check.new(name: "Ruby JIT", status: :ok, detail: "YJIT off in a local environment (expected)") if local_env?
179
+
180
+ Check.new(name: "Ruby JIT", status: :warn,
181
+ detail: "YJIT is available but off — a CPU-bound worker (Postgres nearby) runs ~15% fewer jobs/s " \
182
+ "without it; no measurable effect when network-bound. " \
183
+ "Enable it with config.yjit = true (or load_defaults 7.2+) or RUBY_YJIT_ENABLE=1")
184
+ rescue StandardError => e
185
+ Check.new(name: "Ruby JIT", status: :warn, detail: "could not determine (#{e.class}: #{e.message})")
186
+ end
187
+
188
+ # A Rails app outside development/test. Without Rails the environment is
189
+ # unknown, so assume local and stay quiet rather than warn on a guess.
190
+ def local_env?
191
+ return true unless defined?(Rails) && Rails.respond_to?(:env) && Rails.env
192
+
193
+ env = Rails.env
194
+ env.respond_to?(:local?) ? env.local? : %w[development test].include?(env.to_s)
195
+ end
196
+
169
197
  # 2. Database connectivity — SELECT 1 via Client#ping.
170
198
  def check_database
171
199
  @client.ping
@@ -47,6 +47,23 @@ module Pgbus
47
47
  false
48
48
  end
49
49
 
50
+ # Drops every row recorded under any of `queue_names`. Called when a
51
+ # queue is dropped: PGMQ restarts msg_ids on recreate, so a surviving
52
+ # row would aim the dashboard's retry/discard at an unrelated message.
53
+ # Queue names are validated word characters, so the array literal is safe.
54
+ def clear_queue!(queue_names)
55
+ connection.exec_delete(
56
+ "DELETE FROM pgbus_failed_events WHERE queue_name = ANY($1::text[])",
57
+ "FailedEvent Clear Queue",
58
+ ["{#{queue_names.join(",")}}"]
59
+ )
60
+ rescue StandardError => e
61
+ Pgbus.logger.error do
62
+ "[Pgbus] Failed to clear failed events for dropped queue(s) #{queue_names.join(", ")}: " \
63
+ "#{e.class}: #{e.message}"
64
+ end
65
+ end
66
+
50
67
  def clear!(queue_name:, msg_id:)
51
68
  connection.exec_delete(
52
69
  "DELETE FROM pgbus_failed_events WHERE queue_name = $1 AND msg_id = $2",