pgbus 0.17.0 → 0.17.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +12 -0
- data/README.md +21 -1
- data/Rakefile +7 -1
- data/lib/pgbus/active_job/executor.rb +20 -11
- data/lib/pgbus/client.rb +3 -0
- data/lib/pgbus/configuration.rb +46 -0
- data/lib/pgbus/doctor.rb +29 -1
- data/lib/pgbus/failed_event_recorder.rb +17 -0
- data/lib/pgbus/process/claim_buffer.rb +115 -0
- data/lib/pgbus/process/consumer.rb +41 -14
- data/lib/pgbus/process/supervisor.rb +3 -1
- data/lib/pgbus/process/worker.rb +48 -19
- data/lib/pgbus/ruby_jit.rb +21 -0
- data/lib/pgbus/version.rb +1 -1
- data/lib/pgbus/visibility_heartbeat.rb +38 -3
- metadata +3 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 7b7955f6b4ccdd4d3e4f8d8e372e9283ca0b7f801d7a6909d29f03fde3425ce7
|
|
4
|
+
data.tar.gz: f26d692d0e9b979e9c4b77b3e28cafdab4d3c50e32292e793e1885d94f0d5943
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 30c539f2494acaeffdd10301bba047d6dca86ab2bf3ae776ac6d8d3239f4b5e5a05d530135932cf15122b37dd02acd50de86026315365ea10c085e0100201ebf
|
|
7
|
+
data.tar.gz: 596d817366be305b4abd3c9f1975e6e004ef2448976f58bfb36c9c9b1cd15d0abbc830a98bfd9900d5fc2cfc07b2f6257c1f158b7f8b09a653e97e55c590cc92
|
data/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
### Added
|
|
4
4
|
|
|
5
|
+
- **`rake bench:worker_profile` measures the event consumer too.** A consumer cell drives a real `Process::Consumer` over an event-bus topic queue (plain handler, so the transport is measured, not the dedup table) with the same harness as the worker cell. `WP_BENCH_ROLE` (`worker`, `consumer`, `both`) and `WP_BENCH_READ_AHEAD` (passed as `read_ahead:`) are new knobs. Refs #486.
|
|
6
|
+
|
|
5
7
|
- **`coalesce:` now works on the `Turbo::StreamsChannel` broadcast path.** `Pgbus::Streams::Stream#broadcast` has accepted `coalesce:` since #171, but the Turbo patch forwarded only four options (`durable:`, `exclude:`, `visible_to:`, `event:`), so nothing that broadcasts through `Turbo::StreamsChannel.broadcast_*_to` — every `Turbo::Broadcastable` model, and phlex-reactive's `Streamable.broadcast_to` — could debounce. A background job settling a 177-row collection therefore pushed 177 idempotent replaces of the same count badge at every peer. Any targeted helper now takes `coalesce:` (`@order.broadcast_replace_to :account, coalesce: 100`), as does a direct channel call (`Turbo::StreamsChannel.broadcast_replace_to(@account, target: "order-count", coalesce: true)`), with the same semantics as the direct API: opt-in, per `(stream, target)`, last-write-wins, `true` = the 50ms default window, an integer = the window in ms. Two things the option's plumbing needed beyond a fifth thread-local: the coalescing key, which dedupes on `(stream, target)` while `broadcast_stream_to` only ever sees the rendered `content:` — it is now resolved from the frame's own `target:`/`targets:` through the same `ActionView::RecordIdentifier.dom_id` conversion Turbo applies, so a record target keys on `order_7` rather than on a per-instance `#to_s`; and extraction of the pgbus options at the channel itself, since a direct `Turbo::StreamsChannel` call never passes through `Turbo::Broadcastable` and so previously leaked `durable: true` (and now `coalesce:`) into Turbo's renderer. `broadcast_refresh_to` and `broadcast_render_to` carry no target and so cannot coalesce — they raise the existing actionable "requires target:" error rather than guessing a key — and the `_later_to` variants remain out of scope, as the thread-local cannot follow into a job. The wire is byte-identical when `coalesce:` is absent. Closes #465.
|
|
6
8
|
|
|
7
9
|
- **Concurrency is now visible: a section on the Locks page, three gauges, a home stat card and an MCP tool.** `limits_concurrency … on_conflict: :block` became durable in #460 — a parked job is never dropped and a key never runs past `to:` — but nothing showed it, so an operator whose per-record pipeline looked "stuck" had no way to see that the key's slot was held, how many jobs were parked behind it, or how long the oldest had waited, short of SQL. `pgbus_semaphores` and `pgbus_blocked_executions` appeared on no page, in no gauge, in no tool. One new read, `Web::DataSource#concurrency_stats`, now feeds four surfaces: the Locks page's **Concurrency** section (parked jobs, oldest wait, slots held, keys at limit, then a row per key with value / limit, lease state, parked count and oldest wait, auto-refreshing like the other dashboard panels); a **Parked jobs** card on the dashboard home; three unlabelled gauges on both `/pgbus/api/metrics` and the AppSignal probe (`pgbus_concurrency_blocked_executions`, `pgbus_concurrency_blocked_oldest_age_seconds`, `pgbus_concurrency_slots_held` — keys are per-record, so a per-key label would be an unbounded series); and a read-only `pgbus_concurrency` MCP tool returning the same data. The key rows come from a FULL OUTER JOIN, not a LEFT one: a key can have parked jobs and no semaphore (its holder died and the sweep removed the row) or a semaphore and nothing parked, and both shapes matter to an operator. Each row carries two guarded escape hatches. **Release** drops the semaphore and promotes through `Concurrency::BlockedExecution.promote_next`, so it takes each slot through the same guarded upsert an enqueue uses and can never reintroduce the over-admission #460 closed; when the lease is still fresh the confirm says so, because a job is then probably still running and releasing lets another start beside it. **Discard parked** drops a single key's parked jobs, naming the exact count in the confirm, and resolves the bookkeeping those jobs will never resolve themselves — a parked batch child is marked failed so its batch stops waiting and `on_failure` fires (#413), and an `:until_executed` uniqueness lock is released rather than orphaned (#423); each discard logs a warning naming the job and emits `pgbus.blocked_execution_discarded`. There is deliberately no cross-key "discard all parked": not losing work is the whole point of `:block`. Nothing is added to the enqueue or execute path. Refs #461.
|
|
@@ -10,6 +12,16 @@
|
|
|
10
12
|
|
|
11
13
|
- **The batches list and a batch's progress panel now refresh themselves.** Batches was the only dashboard list whose `turbo-frame` was missing `data-auto-refresh`, and the batch detail page had no frame at all, so watching a batch drain meant reloading by hand. Both now poll on the existing `web_refresh_interval` (5s, pauses while you interact) like every other list — the batch progress panel is served by `GET /batches/:id?frame=progress`. This reuses the dashboard's turbo-frame polling rather than the SSE stack in `lib/pgbus/web/stream_app.rb`: that is the host-app-facing stream feature, and pointing it at the dashboard would open a PG `LISTEN` connection per viewer to move four integers.
|
|
12
14
|
|
|
15
|
+
- **`pgbus doctor` reports the Ruby JIT, and the boot banner names it.** A new "Ruby JIT" check is `:ok` when YJIT (or ZJIT) is on, when the Ruby has no JIT, or in a local (development/test) environment, and warns otherwise — never fails, never enables anything; turning YJIT on stays with Rails' `config.yjit`. The supervisor's first boot line now ends `jit=yjit|zjit|none`. Measured reason: with Postgres close by, a worker is CPU-bound under the GVL and YJIT is worth ~14–20 % jobs/s (`docs/performance.md`). Refs #484.
|
|
16
|
+
|
|
17
|
+
- **`rake bench:worker_profile`: jobs/s and where a job's time goes, for a real worker.** Drives a real `Pgbus::Process::Worker` against real PGMQ and splits each thread's time (Postgres wait, pool-checkout wait, waiting for work, GVL wait, and whose code was on-CPU) with vernier, locally and optionally behind toxiproxy-injected latency. `vernier` joins the test group; it is not a runtime dependency. The same change repairs `benchmarks/integration_bench.rb`, which no longer ran against 1.0's `ensures_uniqueness`. Refs #484.
|
|
18
|
+
|
|
19
|
+
### Changed
|
|
20
|
+
|
|
21
|
+
- **perf(worker, consumer): opt-in read-ahead, so one read feeds many jobs when Postgres is on another host.** A worker read `qty = free threads`; once its pool was busy, threads freed one at a time, so it made one read round trip per job. Under +1 ms of latency that capped a 12-thread worker at ~1.3–1.7 k no-op jobs/s however many threads it had (#484), and the event consumer's loop had the same shape. New `config.read_ahead` (default **0**, off; overridable per capsule and per `event_consumers` entry with `read_ahead:`) lets either loop claim up to N messages beyond its free threads. Claimed messages wait in a new `Process::ClaimBuffer`, kept invisible by a new `VisibilityHeartbeat.hold`, until a thread frees. They are handed back to the queue (`set_vt` 0) the moment the worker drains, recycles, pauses or stops, or the consumer shuts down or recycles, with `read_ct` already counted. Buffered claims count as in flight, so `prefetch_limit` still caps everything claimed. Measured with `rake bench:worker_profile` (12 threads, no-op jobs, +1 ms per round trip, rounds 2–3 of three interleaved rounds on a loaded machine): **worker 1 286–1 731 → 2 300–2 462 jobs/s, consumer 1 326–1 779 → 2 288–2 599** at `read_ahead` 12. Locally both are within noise, because there the worker is CPU-bound under the GVL; `read_ahead` 0 overlaps the baseline's run-to-run range in 7 of 8 cells; consumer / proxied / no JIT sits 2–18 % below, plausibly load noise but unproven. A returned message's next claimant logs it as a zombie redelivery when `zombie_detection` is on, and `validate!` warns when `read_ahead > 0` meets `visibility_heartbeat = false`, since nothing then keeps a buffered message invisible. With read-ahead on, the next ceiling under latency is the connection pool (checkout wait 2–3 % → 14–15 %), so raise `pool_size` with it. `VisibilityHeartbeat`'s unregister now drops only the entry it registered, so a stale hold released after a job's own tracking took over cannot drop that job's heartbeat. Worker heartbeat metadata gains `"buffered"`. Refs #486.
|
|
22
|
+
|
|
23
|
+
- **perf(executor): a successful first delivery no longer deletes from `pgbus_failed_events`.** Every successful job ran a `DELETE … WHERE queue_name = $1 AND msg_id = $2` through ActiveRecord, but only a failed attempt ever writes that table, so a job on its first delivery cannot have a row. The DELETE now runs only on redelivery (`read_ct > 1`), still clearing the row a failed attempt left. Measured with `rake bench:worker_profile` (12 threads, no-op jobs, local Postgres): **4 569 → 5 823 jobs/s with YJIT (+27 %), 3 797 → 5 119 without (+35 %)**; under +1 ms injected latency the change is within noise, because there the worker's single reader is the ceiling. This is a per-job query removed, not faster job code. The executor's debug tag is also built only when debug logging is on (9 → 8 pgbus-owned objects per job, now a hard budget). `Client#drop_queue` now also deletes the dropped queue's `pgbus_failed_events` rows: PGMQ restarts `msg_id`s when a queue is recreated, and a surviving row would point the dashboard's retry/discard at an unrelated new message — previously a first-delivery success happened to sweep such a row, and that sweep is gone. Refs #484.
|
|
24
|
+
|
|
13
25
|
### Fixed
|
|
14
26
|
|
|
15
27
|
- **A still-running event handler is no longer re-run as if it had crashed, and a slow one is no longer redelivered mid-run.** Two halves of one hole, follow-up to #469. First: `Process::Consumer` had no equivalent of `ActiveJob::Executor#with_visibility_heartbeat`, so an event message's visibility timeout was never re-armed — a handler slower than `visibility_timeout` (30s by default) was redelivered *while it was still running*, `read_ct` climbed on every redelivery, and after `max_retries` the event was dead-lettered without the handler ever raising. Workers have been protected from this since the heartbeat landed; event consumers never were. The consumer now tracks each message through `VisibilityHeartbeat` for exactly as long as its handlers run, releasing the entry before the archive so a beat can never re-arm a message that is already gone. Second: `Handler#claim_idempotency?` read a `pgbus_processed_events` row with `completed_at IS NULL` as proof that the previous holder had been killed mid-handler, and re-ran. That state equally describes a handler that is simply still running — on another thread, another fork or another host — so a second delivery (a genuine redelivery, or a duplicate envelope, whose independent visibility timeouts no heartbeat can serialize) executed the handler *concurrently with* the live one, which is precisely the double-execution `idempotent!` exists to prevent. The claim now carries liveness rather than only a claim instant: `EventBus::ClaimBeat` refreshes every in-flight claim's `processed_at` from the same beat that re-arms the message's visibility timeout, so message visibility and claim liveness go quiet together when a process dies. A pending claim silent for longer than twice the heartbeat interval is abandoned and re-runs, exactly as before; a fresher one is owned and the delivery skips, deferring to the holder — which either completes (nothing is lost) or fails, leaving its own message for visibility-timeout redelivery to recover. A skip is also no longer silent: `pgbus.event_skipped` carries the reason (`:completed`, `:cached` or `:owned`), the claim's age in seconds and the delivery's `read_ct`, and the metrics subscriber counts it as `pgbus_event_count` with `status: "skipped"` and the reason as a tag. No migration, no new column, and no change on a table that has not run `pgbus:add_processed_event_completion` — a single-phase claim has no pending state and therefore no ownership question. Closes #470.
|
data/README.md
CHANGED
|
@@ -26,6 +26,7 @@ PostgreSQL-native job processing and event bus for Rails, built on [PGMQ](https:
|
|
|
26
26
|
- [Client-level circuit breaker (database-down)](#client-level-circuit-breaker-database-down)
|
|
27
27
|
- [Read timeouts (libpq-native)](#read-timeouts-libpq-native)
|
|
28
28
|
- [Prefetch flow control](#prefetch-flow-control)
|
|
29
|
+
- [Read-ahead](#read-ahead)
|
|
29
30
|
- [Worker recycling](#worker-recycling)
|
|
30
31
|
- [Retry backoff](#retry-backoff)
|
|
31
32
|
- [Routing and ordering](#routing-and-ordering)
|
|
@@ -541,6 +542,24 @@ end
|
|
|
541
542
|
|
|
542
543
|
The worker tracks in-flight messages with an atomic counter and only fetches `min(idle_threads, prefetch_available)` messages per cycle. The counter is decremented in an `ensure` block so it never gets stuck.
|
|
543
544
|
|
|
545
|
+
### Read-ahead
|
|
546
|
+
|
|
547
|
+
By default a worker reads as many messages as it has free threads. Once the pool is busy, threads free one at a time, so the worker ends up making one read round trip per job. That costs nothing when Postgres is on the same host. When Postgres is on another host, that round trip becomes the throughput ceiling. `read_ahead` lets a worker or event consumer claim up to N messages beyond its free threads and hold them until a thread frees:
|
|
548
|
+
|
|
549
|
+
```ruby
|
|
550
|
+
Pgbus.configure do |config|
|
|
551
|
+
config.read_ahead = 12 # global; 0 = off (default)
|
|
552
|
+
config.capsule :api, queues: %w[api], threads: 12, read_ahead: 12 # per capsule
|
|
553
|
+
config.event_consumers = [{ topics: ["orders.#"], threads: 8, read_ahead: 8 }]
|
|
554
|
+
end
|
|
555
|
+
```
|
|
556
|
+
|
|
557
|
+
- **When to turn it on:** when Postgres is not on the same host as the worker. Start with `read_ahead` equal to `threads`. With a local database it buys little, because the worker is CPU-bound there. Under +1 ms of latency, `read_ahead` 12 took a 12-thread worker from ~1.3–1.7k to ~2.3–2.5k no-op jobs/s (`docs/performance.md`, "Read-ahead"). The next ceiling is the connection pool, so raise `pool_size` with it.
|
|
558
|
+
- **Visibility:** buffered messages are kept invisible by the visibility heartbeat until their job starts, so a long wait never lets another worker claim them. With `visibility_heartbeat = false` nothing holds them, and a message buffered past `visibility_timeout` can run twice; `validate!` warns about that combination.
|
|
559
|
+
- **Shutdown and recycle:** when a worker drains, recycles, pauses or shuts down, buffered messages go straight back to the queue (`set_vt` 0), not after `visibility_timeout`. Their `read_ct` has already been incremented, the same as after a crash, so with `zombie_detection` on, the worker that picks one up logs it as a zombie redelivery.
|
|
560
|
+
- **With `prefetch_limit`:** buffered messages count as in flight, so `prefetch_limit` still caps everything a worker has claimed.
|
|
561
|
+
- **Multi-queue workers:** the plain multi-queue read (no priority, fair share or group mode) can claim and discard rows past its limit. A larger read makes that worse; see `docs/performance.md`.
|
|
562
|
+
|
|
544
563
|
### Worker recycling
|
|
545
564
|
|
|
546
565
|
Pgbus workers recycle themselves to prevent memory bloat. This is the main reliability difference vs. solid_queue, which leaves workers alive forever.
|
|
@@ -2031,11 +2050,12 @@ pgbus help # Show help
|
|
|
2031
2050
|
|
|
2032
2051
|
#### pgbus doctor
|
|
2033
2052
|
|
|
2034
|
-
A single preflight command that answers "is this environment healthy enough to run?" — useful as a deploy or CI gate. It runs
|
|
2053
|
+
A single preflight command that answers "is this environment healthy enough to run?" — useful as a deploy or CI gate. It runs 12 checks and never raises; a broken environment turns every probe into a failed/warned check instead of a crash:
|
|
2035
2054
|
|
|
2036
2055
|
| Check | Fails (`:fail`) when | Warns (`:warn`) when |
|
|
2037
2056
|
|---|---|---|
|
|
2038
2057
|
| Configuration | `Configuration#validate!` raises | — |
|
|
2058
|
+
| Ruby JIT | — | YJIT is available but off outside a local environment (a CPU-bound worker with Postgres nearby runs ~15% fewer jobs/s without it; no measurable effect when network-bound); reports only, never enables |
|
|
2039
2059
|
| Database | Unreachable (`SELECT 1` via `Client#ping`) | — |
|
|
2040
2060
|
| PGMQ schema | Schema not installed | Installed but untracked, or behind the vendored version |
|
|
2041
2061
|
| Queues | A configured queue has no PGMQ table | — |
|
data/Rakefile
CHANGED
|
@@ -25,7 +25,8 @@ namespace :bench do
|
|
|
25
25
|
# no-DB unit suite that bench:all runs in CI.
|
|
26
26
|
db_benches = %w[connection_pool_bench integration_bench streams_bench streams_read_pool_bench
|
|
27
27
|
execution_modes_bench pool_swap_bench pool_autoscale_bench job_burst_bench
|
|
28
|
-
notify_wake_bench notify_chaos_bench streams_hub_bench fair_read_bench
|
|
28
|
+
notify_wake_bench notify_chaos_bench streams_hub_bench fair_read_bench
|
|
29
|
+
worker_profile_bench].freeze
|
|
29
30
|
# The unit suite is every *_bench.rb that doesn't need a database, derived
|
|
30
31
|
# from the directory so a new unit bench is picked up automatically (kept in
|
|
31
32
|
# sync with bench:one, which globs the same files).
|
|
@@ -60,6 +61,11 @@ namespace :bench do
|
|
|
60
61
|
ruby "benchmarks/integration_bench.rb"
|
|
61
62
|
end
|
|
62
63
|
|
|
64
|
+
desc "Run worker profiling bench: jobs/s and where a job's time goes (requires PGBUS_DATABASE_URL)"
|
|
65
|
+
task :worker_profile do
|
|
66
|
+
ruby "benchmarks/worker_profile_bench.rb"
|
|
67
|
+
end
|
|
68
|
+
|
|
63
69
|
desc "Run fair share read benchmark (requires PGBUS_DATABASE_URL)"
|
|
64
70
|
task :fair_read do
|
|
65
71
|
ruby "benchmarks/fair_read_bench.rb"
|
|
@@ -19,8 +19,7 @@ module Pgbus
|
|
|
19
19
|
|
|
20
20
|
def execute(message, queue_name, source_queue: nil)
|
|
21
21
|
execution_start = monotonic_now
|
|
22
|
-
|
|
23
|
-
Pgbus.logger.debug { "[Pgbus::Executor] start #{tag}" }
|
|
22
|
+
Pgbus.logger.debug { "[Pgbus::Executor] start #{log_tag(message, queue_name)}" }
|
|
24
23
|
|
|
25
24
|
payload = JSON.parse(message.message)
|
|
26
25
|
job_class = payload["job_class"]
|
|
@@ -42,7 +41,7 @@ module Pgbus
|
|
|
42
41
|
read_ct: read_count,
|
|
43
42
|
msg_id: message.msg_id.to_i
|
|
44
43
|
)
|
|
45
|
-
Pgbus.logger.debug { "[Pgbus::Executor] dead_lettered #{
|
|
44
|
+
Pgbus.logger.debug { "[Pgbus::Executor] dead_lettered #{log_tag(message, queue_name)} job_class=#{job_class}" }
|
|
46
45
|
return :dead_lettered
|
|
47
46
|
end
|
|
48
47
|
uniqueness_key = Uniqueness.extract_key(payload)
|
|
@@ -70,7 +69,7 @@ module Pgbus
|
|
|
70
69
|
end
|
|
71
70
|
end
|
|
72
71
|
|
|
73
|
-
Pgbus.logger.debug { "[Pgbus::Executor] deserialized #{
|
|
72
|
+
Pgbus.logger.debug { "[Pgbus::Executor] deserialized #{log_tag(message, queue_name)} job_class=#{job_class}" }
|
|
74
73
|
job_succeeded = false
|
|
75
74
|
retried = false
|
|
76
75
|
|
|
@@ -95,13 +94,13 @@ module Pgbus
|
|
|
95
94
|
# job data, so hand it to the job before perform — that is what makes
|
|
96
95
|
# `batch` (and `batch.enqueue` for open batches) work inside a job.
|
|
97
96
|
assign_batch_id(job, payload)
|
|
98
|
-
Pgbus.logger.debug { "[Pgbus::Executor] running #{
|
|
97
|
+
Pgbus.logger.debug { "[Pgbus::Executor] running #{log_tag(message, queue_name)} job_class=#{job_class}" }
|
|
99
98
|
with_visibility_heartbeat(job, queue_name, msg_id, source_queue, payload) { execute_job(job) }
|
|
100
99
|
# retry_on re-enqueues from inside perform_now and returns normally:
|
|
101
100
|
# this attempt is done (archive it) but the job is not — the retry
|
|
102
101
|
# message carries the batch tag and signals on its own outcome.
|
|
103
102
|
retried = Batch.retry_reenqueued?(payload["job_id"])
|
|
104
|
-
Pgbus.logger.debug { "[Pgbus::Executor] perform_returned #{
|
|
103
|
+
Pgbus.logger.debug { "[Pgbus::Executor] perform_returned #{log_tag(message, queue_name)} job_class=#{job_class}" }
|
|
105
104
|
# Archiving is the exact-once claim on this execution.
|
|
106
105
|
# :already_archived means another worker archived the message (our
|
|
107
106
|
# heartbeat lapsed and it was redelivered): that worker owns the
|
|
@@ -109,20 +108,24 @@ module Pgbus
|
|
|
109
108
|
# concurrency slot a second time (rails/solid_queue#761).
|
|
110
109
|
if archive_from(queue_name, msg_id, source_queue: source_queue) == :already_archived
|
|
111
110
|
Pgbus.logger.warn do
|
|
112
|
-
"[Pgbus::Executor] already archived elsewhere, skipping signals #{
|
|
111
|
+
"[Pgbus::Executor] already archived elsewhere, skipping signals #{log_tag(message, queue_name)} job_class=#{job_class}"
|
|
113
112
|
end
|
|
114
113
|
release_duplicate_execution_lock(uniqueness_key, uniqueness_strategy, queue_name, msg_id)
|
|
115
114
|
return :duplicate
|
|
116
115
|
end
|
|
117
|
-
Pgbus.logger.debug { "[Pgbus::Executor] archived #{
|
|
116
|
+
Pgbus.logger.debug { "[Pgbus::Executor] archived #{log_tag(message, queue_name)} job_class=#{job_class}" }
|
|
118
117
|
job_succeeded = true
|
|
119
118
|
release_uniqueness_lock(uniqueness_key)
|
|
120
|
-
FailedEventRecorder.
|
|
119
|
+
# Only FailedEventRecorder.record! writes this table, and only after a
|
|
120
|
+
# failed attempt, so a first delivery has no row: skip the per-job
|
|
121
|
+
# DELETE round trip (issue #484). A stale row from a dropped-and-
|
|
122
|
+
# recreated queue reusing this msg_id is the dispatcher's to sweep.
|
|
123
|
+
FailedEventRecorder.clear!(queue_name: queue_name, msg_id: msg_id) if read_count > 1
|
|
121
124
|
end
|
|
122
125
|
|
|
123
126
|
instrument("pgbus.job_completed", queue: queue_name, job_class: job_class)
|
|
124
127
|
record_stat(payload, queue_name, "success", execution_start, message: message)
|
|
125
|
-
Pgbus.logger.debug { "[Pgbus::Executor] done #{
|
|
128
|
+
Pgbus.logger.debug { "[Pgbus::Executor] done #{log_tag(message, queue_name)} job_class=#{job_class}" }
|
|
126
129
|
:success
|
|
127
130
|
rescue *FATAL_EXCEPTIONS
|
|
128
131
|
# Process-fatal: propagate so the supervisor/OS can react.
|
|
@@ -150,7 +153,7 @@ module Pgbus
|
|
|
150
153
|
exception_object: e
|
|
151
154
|
)
|
|
152
155
|
record_stat(payload, queue_name, "failed", execution_start, message: message)
|
|
153
|
-
Pgbus.logger.debug { "[Pgbus::Executor] failed #{
|
|
156
|
+
Pgbus.logger.debug { "[Pgbus::Executor] failed #{log_tag(message, queue_name)} job_class=#{payload&.dig("job_class")} error=#{e.class}" }
|
|
154
157
|
# Don't signal concurrency on transient failure — the job will be retried.
|
|
155
158
|
# Semaphore is released only on success or dead-lettering.
|
|
156
159
|
:failed
|
|
@@ -170,6 +173,12 @@ module Pgbus
|
|
|
170
173
|
|
|
171
174
|
private
|
|
172
175
|
|
|
176
|
+
# Built inside the logger blocks so a job pays for the string only when
|
|
177
|
+
# debug logging is on (issue #484).
|
|
178
|
+
def log_tag(message, queue_name)
|
|
179
|
+
"msg_id=#{message.msg_id} queue=#{queue_name} read_ct=#{message.read_ct}"
|
|
180
|
+
end
|
|
181
|
+
|
|
173
182
|
def assign_batch_id(job, payload)
|
|
174
183
|
batch_id = payload[Batch::METADATA_KEY]
|
|
175
184
|
return unless batch_id && job.respond_to?(:batch_id=)
|
data/lib/pgbus/client.rb
CHANGED
|
@@ -699,6 +699,9 @@ module Pgbus
|
|
|
699
699
|
synchronized { @pgmq.drop_queue(name) }
|
|
700
700
|
end
|
|
701
701
|
@queues_created.delete(name)
|
|
702
|
+
# Failed-event rows are keyed by the name the executor saw (physical or
|
|
703
|
+
# logical); a recreated queue reuses msg_ids, so neither may survive.
|
|
704
|
+
FailedEventRecorder.clear_queue!([name, name.delete_prefix("#{config.queue_prefix}_")].uniq)
|
|
702
705
|
result
|
|
703
706
|
end
|
|
704
707
|
|
data/lib/pgbus/configuration.rb
CHANGED
|
@@ -12,6 +12,12 @@ module Pgbus
|
|
|
12
12
|
|
|
13
13
|
# Worker settings
|
|
14
14
|
attr_accessor :polling_interval, :prefetch_limit, :execution_mode
|
|
15
|
+
# read_ahead (issue #486): how many messages a worker or event consumer
|
|
16
|
+
# claims beyond its free threads and holds, heartbeated, until a thread
|
|
17
|
+
# frees. Under database latency one read then feeds many jobs instead of
|
|
18
|
+
# one. 0 (default) is off. Overridable per capsule / event_consumers entry
|
|
19
|
+
# (`read_ahead:`); prefetch_limit still caps everything claimed.
|
|
20
|
+
attr_reader :read_ahead
|
|
15
21
|
# visibility_heartbeat / visibility_heartbeat_interval: while a job runs,
|
|
16
22
|
# its message's visibility timeout is re-armed every interval seconds
|
|
17
23
|
# (default: a third of visibility_timeout), so a job that outlives the
|
|
@@ -277,6 +283,7 @@ module Pgbus
|
|
|
277
283
|
@visibility_heartbeat_interval = nil
|
|
278
284
|
|
|
279
285
|
@prefetch_limit = nil
|
|
286
|
+
@read_ahead = 0
|
|
280
287
|
@execution_mode = :threads
|
|
281
288
|
|
|
282
289
|
@max_jobs_per_worker = nil
|
|
@@ -573,6 +580,16 @@ module Pgbus
|
|
|
573
580
|
(0...priority_levels).map { |p| priority_queue_name(name, p) }
|
|
574
581
|
end
|
|
575
582
|
|
|
583
|
+
def read_ahead=(value)
|
|
584
|
+
@read_ahead = value.nil? ? 0 : value
|
|
585
|
+
end
|
|
586
|
+
|
|
587
|
+
# Read-ahead for one capsule or event_consumers entry, falling back to
|
|
588
|
+
# the global read_ahead (mirrors execution_mode_for).
|
|
589
|
+
def read_ahead_for(entry)
|
|
590
|
+
entry.fetch(:read_ahead, nil) || read_ahead
|
|
591
|
+
end
|
|
592
|
+
|
|
576
593
|
# Returns the execution mode for a specific worker config hash,
|
|
577
594
|
# falling back to the global execution_mode setting.
|
|
578
595
|
def execution_mode_for(worker_config)
|
|
@@ -828,6 +845,8 @@ module Pgbus
|
|
|
828
845
|
"prefetch_limit must be > 0"
|
|
829
846
|
end
|
|
830
847
|
|
|
848
|
+
validate_read_ahead!
|
|
849
|
+
|
|
831
850
|
if priority_levels && !(priority_levels.is_a?(Integer) && priority_levels >= 1 && priority_levels <= 10)
|
|
832
851
|
raise Pgbus::ConfigurationError, "priority_levels must be an integer between 1 and 10"
|
|
833
852
|
end
|
|
@@ -846,6 +865,33 @@ module Pgbus
|
|
|
846
865
|
self
|
|
847
866
|
end
|
|
848
867
|
|
|
868
|
+
# Global, per-capsule and per-consumer read_ahead: a non-negative Integer.
|
|
869
|
+
def validate_read_ahead!
|
|
870
|
+
settings = [["read_ahead", read_ahead]]
|
|
871
|
+
Array(workers).each { |w| settings << ["worker read_ahead", w[:read_ahead]] unless w[:read_ahead].nil? }
|
|
872
|
+
Array(event_consumers).each { |c| settings << ["event consumer read_ahead", c[:read_ahead]] unless c[:read_ahead].nil? }
|
|
873
|
+
|
|
874
|
+
settings.each do |name, value|
|
|
875
|
+
next if value.is_a?(Integer) && !value.negative?
|
|
876
|
+
|
|
877
|
+
raise Pgbus::ConfigurationError, "#{name} must be a non-negative Integer (got #{value.inspect})"
|
|
878
|
+
end
|
|
879
|
+
warn_read_ahead_without_heartbeat(settings)
|
|
880
|
+
end
|
|
881
|
+
|
|
882
|
+
# Buffered claims are kept invisible only by the visibility heartbeat.
|
|
883
|
+
# Without it a message buffered longer than visibility_timeout can be
|
|
884
|
+
# claimed by another worker while it still waits here. Legal, but warned.
|
|
885
|
+
def warn_read_ahead_without_heartbeat(settings)
|
|
886
|
+
return if visibility_heartbeat
|
|
887
|
+
return unless settings.any? { |_name, value| value.positive? }
|
|
888
|
+
|
|
889
|
+
Pgbus.logger.warn do
|
|
890
|
+
"[Pgbus] read_ahead is on but visibility_heartbeat is false — a message buffered longer than " \
|
|
891
|
+
"visibility_timeout (#{visibility_timeout}s) can be claimed and run by another worker"
|
|
892
|
+
end
|
|
893
|
+
end
|
|
894
|
+
|
|
849
895
|
# An explicit shutdown_timeout must be a positive number; nil keeps the
|
|
850
896
|
# derived drain_timeout + margin default. A value below drain_timeout is
|
|
851
897
|
# legal but self-defeating (the supervisor SIGKILLs workers mid-drain), so
|
data/lib/pgbus/doctor.rb
CHANGED
|
@@ -5,7 +5,7 @@ require "pgbus/mcp/health_analyzer"
|
|
|
5
5
|
|
|
6
6
|
module Pgbus
|
|
7
7
|
# Preflight diagnostics for a pgbus deployment — the single command that
|
|
8
|
-
# answers "is this environment healthy enough to run?". Runs
|
|
8
|
+
# answers "is this environment healthy enough to run?". Runs twelve checks and
|
|
9
9
|
# returns a machine-readable result plus a human report, so `pgbus doctor`
|
|
10
10
|
# and `rake pgbus:doctor` can gate a deploy or CI run (exit 0 on success,
|
|
11
11
|
# 1 on any failure).
|
|
@@ -32,6 +32,7 @@ module Pgbus
|
|
|
32
32
|
# stable identity of each check (used in the report and to select subsets).
|
|
33
33
|
CHECKS = {
|
|
34
34
|
"Configuration" => :check_configuration,
|
|
35
|
+
"Ruby JIT" => :check_ruby_jit,
|
|
35
36
|
"Database" => :check_database,
|
|
36
37
|
"PGMQ schema" => :check_pgmq_schema,
|
|
37
38
|
"Queues" => :check_queues,
|
|
@@ -166,6 +167,33 @@ module Pgbus
|
|
|
166
167
|
Check.new(name: "Configuration", status: :fail, detail: "#{e.class}: #{e.message}")
|
|
167
168
|
end
|
|
168
169
|
|
|
170
|
+
# 1b. Ruby JIT (issue #484). With Postgres close by, a worker is GVL-bound
|
|
171
|
+
# and YJIT is worth ~15-20% jobs/s (docs/performance.md). Rails enables it via
|
|
172
|
+
# config.yjit (load_defaults 7.2+, or 8.0+ outside local envs). Reports
|
|
173
|
+
# only — enabling a JIT is the application's decision, never pgbus's.
|
|
174
|
+
def check_ruby_jit
|
|
175
|
+
label = RubyJit.label
|
|
176
|
+
return Check.new(name: "Ruby JIT", status: :ok, detail: "#{label.upcase} enabled") unless label == "none"
|
|
177
|
+
return Check.new(name: "Ruby JIT", status: :ok, detail: "no YJIT available on this Ruby") unless RubyJit.yjit_available?
|
|
178
|
+
return Check.new(name: "Ruby JIT", status: :ok, detail: "YJIT off in a local environment (expected)") if local_env?
|
|
179
|
+
|
|
180
|
+
Check.new(name: "Ruby JIT", status: :warn,
|
|
181
|
+
detail: "YJIT is available but off — a CPU-bound worker (Postgres nearby) runs ~15% fewer jobs/s " \
|
|
182
|
+
"without it; no measurable effect when network-bound. " \
|
|
183
|
+
"Enable it with config.yjit = true (or load_defaults 7.2+) or RUBY_YJIT_ENABLE=1")
|
|
184
|
+
rescue StandardError => e
|
|
185
|
+
Check.new(name: "Ruby JIT", status: :warn, detail: "could not determine (#{e.class}: #{e.message})")
|
|
186
|
+
end
|
|
187
|
+
|
|
188
|
+
# A Rails app outside development/test. Without Rails the environment is
|
|
189
|
+
# unknown, so assume local and stay quiet rather than warn on a guess.
|
|
190
|
+
def local_env?
|
|
191
|
+
return true unless defined?(Rails) && Rails.respond_to?(:env) && Rails.env
|
|
192
|
+
|
|
193
|
+
env = Rails.env
|
|
194
|
+
env.respond_to?(:local?) ? env.local? : %w[development test].include?(env.to_s)
|
|
195
|
+
end
|
|
196
|
+
|
|
169
197
|
# 2. Database connectivity — SELECT 1 via Client#ping.
|
|
170
198
|
def check_database
|
|
171
199
|
@client.ping
|
|
@@ -47,6 +47,23 @@ module Pgbus
|
|
|
47
47
|
false
|
|
48
48
|
end
|
|
49
49
|
|
|
50
|
+
# Drops every row recorded under any of `queue_names`. Called when a
|
|
51
|
+
# queue is dropped: PGMQ restarts msg_ids on recreate, so a surviving
|
|
52
|
+
# row would aim the dashboard's retry/discard at an unrelated message.
|
|
53
|
+
# Queue names are validated word characters, so the array literal is safe.
|
|
54
|
+
def clear_queue!(queue_names)
|
|
55
|
+
connection.exec_delete(
|
|
56
|
+
"DELETE FROM pgbus_failed_events WHERE queue_name = ANY($1::text[])",
|
|
57
|
+
"FailedEvent Clear Queue",
|
|
58
|
+
["{#{queue_names.join(",")}}"]
|
|
59
|
+
)
|
|
60
|
+
rescue StandardError => e
|
|
61
|
+
Pgbus.logger.error do
|
|
62
|
+
"[Pgbus] Failed to clear failed events for dropped queue(s) #{queue_names.join(", ")}: " \
|
|
63
|
+
"#{e.class}: #{e.message}"
|
|
64
|
+
end
|
|
65
|
+
end
|
|
66
|
+
|
|
50
67
|
def clear!(queue_name:, msg_id:)
|
|
51
68
|
connection.exec_delete(
|
|
52
69
|
"DELETE FROM pgbus_failed_events WHERE queue_name = $1 AND msg_id = $2",
|
|
@@ -0,0 +1,115 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Pgbus
|
|
4
|
+
module Process
|
|
5
|
+
# The read-ahead buffer a Worker or Consumer owns (issue #486).
|
|
6
|
+
#
|
|
7
|
+
# Under database latency the claim loop used to issue one read round trip
|
|
8
|
+
# per job: once the pool is full, slots free one at a time and each loop
|
|
9
|
+
# step read `qty = free slots` = 1. With `read_ahead = N` the loop claims
|
|
10
|
+
# up to N messages beyond its free slots and parks them here, so one read
|
|
11
|
+
# returns a batch and the pool never waits on the reader.
|
|
12
|
+
#
|
|
13
|
+
# Surplus claims cannot live inside the execution pool: AsyncPool#post
|
|
14
|
+
# raises at zero capacity and a queued block could not be handed back on
|
|
15
|
+
# drain. So the buffer is drained into the pool only while it has a free
|
|
16
|
+
# slot, and returned to the queue (vt: 0) on drain, recycle or shutdown.
|
|
17
|
+
#
|
|
18
|
+
# Every buffered message is kept invisible by a VisibilityHeartbeat hold
|
|
19
|
+
# from claim until the run starts, under the executor's key
|
|
20
|
+
# (`source_queue || queue_name`, `prefixed: source_queue.nil?`), so the
|
|
21
|
+
# run's own tracking takes over the same entry with no gap. The hold does
|
|
22
|
+
# not know the job class, so a job that opted out with
|
|
23
|
+
# `pgbus_visibility_heartbeat false` is still held while it waits here:
|
|
24
|
+
# the opt-out covers its run, not the wait. With the heartbeat disabled
|
|
25
|
+
# globally there is no hold at all (Configuration#validate! warns).
|
|
26
|
+
#
|
|
27
|
+
# Only ever touched from the owning process's loop thread, so no mutex.
|
|
28
|
+
class ClaimBuffer
|
|
29
|
+
Claim = Struct.new(:queue_name, :message, :source_queue, :hold)
|
|
30
|
+
|
|
31
|
+
HOLD_LABEL = "(read-ahead)"
|
|
32
|
+
|
|
33
|
+
def initialize(config: Pgbus.configuration)
|
|
34
|
+
@config = config
|
|
35
|
+
@claims = []
|
|
36
|
+
end
|
|
37
|
+
|
|
38
|
+
def size
|
|
39
|
+
@claims.size
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
def empty?
|
|
43
|
+
@claims.empty?
|
|
44
|
+
end
|
|
45
|
+
|
|
46
|
+
def push(queue_name, message, source_queue = nil, client: Pgbus.client)
|
|
47
|
+
hold = VisibilityHeartbeat.hold(
|
|
48
|
+
client: client, queue_name: source_queue || queue_name, prefixed: source_queue.nil?,
|
|
49
|
+
msg_id: message.msg_id, job_class: HOLD_LABEL, config: @config
|
|
50
|
+
)
|
|
51
|
+
@claims << Claim.new(queue_name, message, source_queue, hold)
|
|
52
|
+
self
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
def shift
|
|
56
|
+
@claims.shift
|
|
57
|
+
end
|
|
58
|
+
|
|
59
|
+
# Yields the oldest claims, one per free pool slot, and returns how many.
|
|
60
|
+
# Capacity is re-read before every claim: the block's post takes a slot.
|
|
61
|
+
def drain_into(pool)
|
|
62
|
+
drained = 0
|
|
63
|
+
while !@claims.empty? && pool.available_capacity.positive?
|
|
64
|
+
yield @claims.shift
|
|
65
|
+
drained += 1
|
|
66
|
+
end
|
|
67
|
+
drained
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
# How many messages the next read should claim: enough to fill the free
|
|
71
|
+
# slots and top the buffer back up to `read_ahead`, never more than
|
|
72
|
+
# `prefetch_room` (prefetch_limit minus everything already claimed).
|
|
73
|
+
# With read_ahead 0 and an empty buffer this is `free_slots`: the
|
|
74
|
+
# pre-#486 qty.
|
|
75
|
+
def deficit(free_slots:, read_ahead:, prefetch_room: nil)
|
|
76
|
+
want = free_slots + read_ahead - @claims.size
|
|
77
|
+
want = [want, prefetch_room].min if prefetch_room
|
|
78
|
+
want.clamp(0..)
|
|
79
|
+
end
|
|
80
|
+
|
|
81
|
+
# Hand every buffered message back to the queue right away instead of
|
|
82
|
+
# leaving it invisible until its timeout. Their read_ct has already been
|
|
83
|
+
# bumped, the same cost a crash or a stale claim pays. A failed return
|
|
84
|
+
# is logged and skipped: the hold is already dropped, so the message
|
|
85
|
+
# reappears when its current timeout runs out. Returns the count. One
|
|
86
|
+
# claim at a time, so #size stays true while the returns are written.
|
|
87
|
+
def return_all!(client: Pgbus.client)
|
|
88
|
+
returned = 0
|
|
89
|
+
while (claim = @claims.shift)
|
|
90
|
+
return_claim(claim, client)
|
|
91
|
+
returned += 1
|
|
92
|
+
end
|
|
93
|
+
returned
|
|
94
|
+
end
|
|
95
|
+
|
|
96
|
+
private
|
|
97
|
+
|
|
98
|
+
# The hold goes first: a heartbeat tick that still held the entry could
|
|
99
|
+
# otherwise write its extension after the vt: 0 and hide the message
|
|
100
|
+
# again for a full visibility_timeout. (VisibilityHeartbeat#extend!
|
|
101
|
+
# re-checks registration, which shrinks the remaining window to the
|
|
102
|
+
# tick's own check-then-write.)
|
|
103
|
+
def return_claim(claim, client)
|
|
104
|
+
VisibilityHeartbeat.release(claim.hold)
|
|
105
|
+
queue = claim.source_queue || claim.queue_name
|
|
106
|
+
client.set_visibility_timeout(queue, claim.message.msg_id.to_i, vt: 0, prefixed: claim.source_queue.nil?)
|
|
107
|
+
rescue StandardError => e
|
|
108
|
+
Pgbus.logger.warn do
|
|
109
|
+
"[Pgbus] Could not return read-ahead message msg_id=#{claim.message.msg_id} queue=#{queue}: " \
|
|
110
|
+
"#{e.class}: #{e.message}"
|
|
111
|
+
end
|
|
112
|
+
end
|
|
113
|
+
end
|
|
114
|
+
end
|
|
115
|
+
end
|
|
@@ -7,7 +7,7 @@ module Pgbus
|
|
|
7
7
|
class Consumer
|
|
8
8
|
include SignalHandler
|
|
9
9
|
|
|
10
|
-
attr_reader :topics, :threads, :config, :execution_mode,
|
|
10
|
+
attr_reader :topics, :threads, :config, :execution_mode, :read_ahead,
|
|
11
11
|
:queue_names, :wake_signal, :notify_retry_backoff, :circuit_breaker
|
|
12
12
|
# notify_listener is writable so tests can simulate a start_notify_listener
|
|
13
13
|
# success from inside a stub (production sets it in start_notify_listener).
|
|
@@ -30,6 +30,11 @@ module Pgbus
|
|
|
30
30
|
@jobs_processed.value
|
|
31
31
|
end
|
|
32
32
|
|
|
33
|
+
# Messages claimed by read-ahead and not yet handed to the pool (#486).
|
|
34
|
+
def buffered
|
|
35
|
+
@claim_buffer.size
|
|
36
|
+
end
|
|
37
|
+
|
|
33
38
|
# Seed the processed-job counter. Used by tests to drive the recycle
|
|
34
39
|
# thresholds without running thousands of real jobs; production only ever
|
|
35
40
|
# increments it via the AtomicFixnum during message handling.
|
|
@@ -60,7 +65,7 @@ module Pgbus
|
|
|
60
65
|
queue_names: nil, liveness_pipe: nil, stat_buffer: :default,
|
|
61
66
|
notify_listener: nil, notify_retry_at: 0.0,
|
|
62
67
|
notify_retry_backoff: NOTIFY_RETRY_BASE_SECONDS,
|
|
63
|
-
started_at_monotonic: nil, wake_pipe: nil)
|
|
68
|
+
started_at_monotonic: nil, wake_pipe: nil, read_ahead: nil)
|
|
64
69
|
@topics = Array(topics)
|
|
65
70
|
@threads = threads
|
|
66
71
|
@config = config
|
|
@@ -68,6 +73,8 @@ module Pgbus
|
|
|
68
73
|
@shutting_down = false
|
|
69
74
|
@recycling = false
|
|
70
75
|
@jobs_processed = Concurrent::AtomicFixnum.new(0)
|
|
76
|
+
@read_ahead = read_ahead || config.read_ahead
|
|
77
|
+
@claim_buffer = ClaimBuffer.new(config: config)
|
|
71
78
|
@loop_tick_at = Concurrent::AtomicReference.new(nil)
|
|
72
79
|
@started_at_monotonic = started_at_monotonic || monotonic_now
|
|
73
80
|
@wake_signal = WakeSignal.new
|
|
@@ -134,7 +141,12 @@ module Pgbus
|
|
|
134
141
|
check_recycle
|
|
135
142
|
ensure_notify_listener
|
|
136
143
|
|
|
137
|
-
|
|
144
|
+
if @shutting_down
|
|
145
|
+
# No drain loop here: buffered claims go straight back to the
|
|
146
|
+
# queue; shutdown then waits for the ones already running.
|
|
147
|
+
@claim_buffer.return_all!(client: Pgbus.client) unless @claim_buffer.empty?
|
|
148
|
+
break
|
|
149
|
+
end
|
|
138
150
|
|
|
139
151
|
consume
|
|
140
152
|
@stat_buffer&.flush_if_due
|
|
@@ -171,20 +183,35 @@ module Pgbus
|
|
|
171
183
|
@queue_names = @registry.queue_names_for_topics(topics)
|
|
172
184
|
end
|
|
173
185
|
|
|
186
|
+
# Same loop step as Worker#claim_and_execute (issue #486): feed
|
|
187
|
+
# buffered claims to free slots, read the deficit, feed again, and wait
|
|
188
|
+
# only when the step moved nothing.
|
|
174
189
|
def consume
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
190
|
+
drained = drain_claim_buffer
|
|
191
|
+
claimed = claim_deficit
|
|
192
|
+
drained += drain_claim_buffer
|
|
193
|
+
@wake_signal.wait(timeout: wake_timeout) if drained.zero? && claimed.zero?
|
|
194
|
+
end
|
|
195
|
+
|
|
196
|
+
# The hold is released before handle_message's own heartbeat tracking
|
|
197
|
+
# registers the same key, so the message is never untracked in between.
|
|
198
|
+
def drain_claim_buffer
|
|
199
|
+
@claim_buffer.drain_into(@pool) do |claim|
|
|
200
|
+
@pool.post do
|
|
201
|
+
VisibilityHeartbeat.release(claim.hold)
|
|
202
|
+
handle_message(claim.message, claim.queue_name)
|
|
203
|
+
end
|
|
183
204
|
end
|
|
205
|
+
end
|
|
184
206
|
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
207
|
+
def claim_deficit
|
|
208
|
+
want = @claim_buffer.deficit(free_slots: @pool.available_capacity, read_ahead: @read_ahead)
|
|
209
|
+
return 0 if want.zero?
|
|
210
|
+
|
|
211
|
+
tagged_messages = fetch_messages(want)
|
|
212
|
+
client = Pgbus.client
|
|
213
|
+
tagged_messages.each { |queue_name, message| @claim_buffer.push(queue_name, message, client: client) }
|
|
214
|
+
tagged_messages.size
|
|
188
215
|
end
|
|
189
216
|
|
|
190
217
|
# Returns an array of [queue_name, message] pairs. Queues whose circuit
|
|
@@ -171,7 +171,7 @@ module Pgbus
|
|
|
171
171
|
private_constant :ROLE_FLAGS
|
|
172
172
|
|
|
173
173
|
def log_boot_banner
|
|
174
|
-
Pgbus.logger.info { "[Pgbus] boot: pgbus #{Pgbus::VERSION} pid=#{::Process.pid}" }
|
|
174
|
+
Pgbus.logger.info { "[Pgbus] boot: pgbus #{Pgbus::VERSION} pid=#{::Process.pid} jit=#{RubyJit.label}" }
|
|
175
175
|
Pgbus.logger.info do
|
|
176
176
|
"[Pgbus] boot: connection=#{redacted_connection_target} pool=#{banner_field { config.resolved_pool_size }}"
|
|
177
177
|
end
|
|
@@ -330,6 +330,7 @@ module Pgbus
|
|
|
330
330
|
queues: queues, threads: threads, config: config,
|
|
331
331
|
single_active_consumer: single_active, consumer_priority: priority,
|
|
332
332
|
execution_mode: exec_mode, group_mode: grp_mode,
|
|
333
|
+
read_ahead: config.read_ahead_for(worker_config),
|
|
333
334
|
liveness_pipe: liveness_writer, wake_pipe: wake_reader
|
|
334
335
|
)
|
|
335
336
|
worker.run
|
|
@@ -498,6 +499,7 @@ module Pgbus
|
|
|
498
499
|
setup_child_process
|
|
499
500
|
load_rails_app
|
|
500
501
|
consumer = Consumer.new(topics: topics, threads: threads, config: config,
|
|
502
|
+
read_ahead: config.read_ahead_for(consumer_config),
|
|
501
503
|
liveness_pipe: liveness_writer, wake_pipe: wake_reader)
|
|
502
504
|
consumer.run
|
|
503
505
|
end
|
data/lib/pgbus/process/worker.rb
CHANGED
|
@@ -7,7 +7,7 @@ module Pgbus
|
|
|
7
7
|
class Worker
|
|
8
8
|
include SignalHandler
|
|
9
9
|
|
|
10
|
-
attr_reader :queues, :threads, :config, :execution_mode,
|
|
10
|
+
attr_reader :queues, :threads, :config, :execution_mode, :read_ahead,
|
|
11
11
|
:rate_counter, :wake_signal, :restore_streak, :lifecycle
|
|
12
12
|
# stat_buffer is writable so a test can swap in a buffer double after
|
|
13
13
|
# construction to assert graceful_shutdown / check_recycle flush it. The
|
|
@@ -35,7 +35,7 @@ module Pgbus
|
|
|
35
35
|
rate_counter: nil, wake_signal: nil, stat_buffer: :default,
|
|
36
36
|
notify_listener: nil, notify_retry_at: 0.0,
|
|
37
37
|
notify_retry_backoff: NOTIFY_RETRY_BASE_SECONDS,
|
|
38
|
-
started_at_monotonic: nil, wake_pipe: nil)
|
|
38
|
+
started_at_monotonic: nil, wake_pipe: nil, read_ahead: nil)
|
|
39
39
|
@queues = Array(queues)
|
|
40
40
|
@initial_queues = @queues.dup.freeze
|
|
41
41
|
@wildcard = @queues.include?("*")
|
|
@@ -65,7 +65,10 @@ module Pgbus
|
|
|
65
65
|
@last_wildcard_resolve = nil
|
|
66
66
|
@jobs_processed = Concurrent::AtomicFixnum.new(0)
|
|
67
67
|
@jobs_failed = Concurrent::AtomicFixnum.new(0)
|
|
68
|
+
# in_flight counts every claimed message: buffered and running.
|
|
68
69
|
@in_flight = Concurrent::AtomicFixnum.new(0)
|
|
70
|
+
@read_ahead = read_ahead || config.read_ahead
|
|
71
|
+
@claim_buffer = ClaimBuffer.new(config: config)
|
|
69
72
|
@loop_tick_at = Concurrent::AtomicReference.new(nil)
|
|
70
73
|
@rate_counter = rate_counter || RateCounter.new(:processed, :failed, :dequeued)
|
|
71
74
|
@started_at = Time.current
|
|
@@ -119,6 +122,7 @@ module Pgbus
|
|
|
119
122
|
jobs_processed: @jobs_processed.value,
|
|
120
123
|
jobs_failed: @jobs_failed.value,
|
|
121
124
|
in_flight: @in_flight.value,
|
|
125
|
+
buffered: @claim_buffer.size,
|
|
122
126
|
state: @lifecycle.state,
|
|
123
127
|
execution_mode: @execution_mode,
|
|
124
128
|
consumer_priority: @consumer_priority,
|
|
@@ -184,6 +188,9 @@ module Pgbus
|
|
|
184
188
|
check_recycle
|
|
185
189
|
refresh_wildcard_queues
|
|
186
190
|
ensure_notify_listener
|
|
191
|
+
# Draining, paused or stopped: claims nobody will run here go back
|
|
192
|
+
# to the queue now rather than when their timeout runs out.
|
|
193
|
+
return_claims unless @lifecycle.can_process?
|
|
187
194
|
|
|
188
195
|
break if @lifecycle.stopped?
|
|
189
196
|
# quiesced? (all slots free), not idle? (any slot free) — exiting
|
|
@@ -240,32 +247,53 @@ module Pgbus
|
|
|
240
247
|
|
|
241
248
|
private
|
|
242
249
|
|
|
250
|
+
# One loop step (issue #486): feed buffered claims to free slots, read
|
|
251
|
+
# the deficit (free slots + read_ahead - buffered, capped by
|
|
252
|
+
# prefetch_limit), feed again. With read_ahead 0 the buffer never
|
|
253
|
+
# outlives the step and qty is the free slot count, as before. Waits
|
|
254
|
+
# only when the step moved nothing, so a full pool with a topped-up
|
|
255
|
+
# buffer sleeps until a slot frees (the pool's on_state_change wakes it).
|
|
243
256
|
def claim_and_execute
|
|
244
|
-
|
|
245
|
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
252
|
-
|
|
253
|
-
|
|
257
|
+
drained = drain_claim_buffer
|
|
258
|
+
claimed = claim_deficit
|
|
259
|
+
drained += drain_claim_buffer
|
|
260
|
+
@wake_signal.wait(timeout: wake_timeout) if drained.zero? && claimed.zero?
|
|
261
|
+
end
|
|
262
|
+
|
|
263
|
+
# The hold is released before the executor's own heartbeat tracking
|
|
264
|
+
# registers the same key, so the message is never untracked in between.
|
|
265
|
+
def drain_claim_buffer
|
|
266
|
+
@claim_buffer.drain_into(@pool) do |claim|
|
|
267
|
+
@pool.post do
|
|
268
|
+
VisibilityHeartbeat.release(claim.hold)
|
|
269
|
+
process_message(claim.message, claim.queue_name, source_queue: claim.source_queue)
|
|
270
|
+
end
|
|
254
271
|
end
|
|
272
|
+
end
|
|
255
273
|
|
|
256
|
-
|
|
274
|
+
def claim_deficit
|
|
275
|
+
prefetch_room = config.prefetch_limit && (config.prefetch_limit - @in_flight.value)
|
|
276
|
+
want = @claim_buffer.deficit(free_slots: @pool.available_capacity, read_ahead: @read_ahead,
|
|
277
|
+
prefetch_room: prefetch_room)
|
|
278
|
+
return 0 if want.zero?
|
|
257
279
|
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
return
|
|
261
|
-
end
|
|
280
|
+
tagged_messages = fetch_messages(want)
|
|
281
|
+
return 0 if tagged_messages.empty?
|
|
262
282
|
|
|
263
283
|
@rate_counter.increment(:dequeued, tagged_messages.size)
|
|
284
|
+
client = Pgbus.client
|
|
264
285
|
tagged_messages.each do |queue_name, message, source_queue|
|
|
265
286
|
detect_zombie(queue_name, message)
|
|
266
287
|
@in_flight.increment
|
|
267
|
-
@
|
|
288
|
+
@claim_buffer.push(queue_name, message, source_queue, client: client)
|
|
268
289
|
end
|
|
290
|
+
tagged_messages.size
|
|
291
|
+
end
|
|
292
|
+
|
|
293
|
+
def return_claims
|
|
294
|
+
return if @claim_buffer.empty?
|
|
295
|
+
|
|
296
|
+
@in_flight.decrement(@claim_buffer.return_all!(client: Pgbus.client))
|
|
269
297
|
end
|
|
270
298
|
|
|
271
299
|
# Returns an array of [queue_name, message] pairs so we always know
|
|
@@ -831,7 +859,8 @@ module Pgbus
|
|
|
831
859
|
"rates" => @rate_counter.rates.transform_keys(&:to_s),
|
|
832
860
|
"jobs_processed" => @jobs_processed.value,
|
|
833
861
|
"jobs_failed" => @jobs_failed.value,
|
|
834
|
-
"in_flight" => @in_flight.value
|
|
862
|
+
"in_flight" => @in_flight.value,
|
|
863
|
+
"buffered" => @claim_buffer.size
|
|
835
864
|
}
|
|
836
865
|
end
|
|
837
866
|
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Pgbus
|
|
4
|
+
# Reports which Ruby JIT this process runs under (issue #484). Read-only:
|
|
5
|
+
# Rails owns turning YJIT on (config.yjit), pgbus only makes the state
|
|
6
|
+
# visible in the boot log and `pgbus doctor`.
|
|
7
|
+
module RubyJit
|
|
8
|
+
module_function
|
|
9
|
+
|
|
10
|
+
def label
|
|
11
|
+
return "yjit" if yjit_available? && RubyVM::YJIT.enabled?
|
|
12
|
+
return "zjit" if defined?(RubyVM::ZJIT) && RubyVM::ZJIT.enabled?
|
|
13
|
+
|
|
14
|
+
"none"
|
|
15
|
+
end
|
|
16
|
+
|
|
17
|
+
def yjit_available?
|
|
18
|
+
defined?(RubyVM::YJIT) ? true : false
|
|
19
|
+
end
|
|
20
|
+
end
|
|
21
|
+
end
|
data/lib/pgbus/version.rb
CHANGED
|
@@ -30,8 +30,9 @@ module Pgbus
|
|
|
30
30
|
#
|
|
31
31
|
# The ninth member takes the struct out of the 80-byte slot it used to fit
|
|
32
32
|
# (measured: 80 → 160). That is paid at most once per in-flight message, so
|
|
33
|
-
# the whole table is bounded by the execution pool's capacity
|
|
34
|
-
#
|
|
33
|
+
# the whole table is bounded by the execution pool's capacity plus the
|
|
34
|
+
# read-ahead buffer (threads + read_ahead per process, issue #486) — a
|
|
35
|
+
# handful of entries per process, not one per enqueued job.
|
|
35
36
|
Entry = Struct.new(:client, :queue_name, :prefixed, :msg_id, :job_class, :extended_at, :extensions,
|
|
36
37
|
:concurrency, :on_beat, keyword_init: true)
|
|
37
38
|
|
|
@@ -81,6 +82,25 @@ module Pgbus
|
|
|
81
82
|
end
|
|
82
83
|
end
|
|
83
84
|
|
|
85
|
+
# Keep a claimed-but-not-started message invisible until #release.
|
|
86
|
+
# Used by Process::ClaimBuffer for read-ahead (issue #486): a buffered
|
|
87
|
+
# message waits for a free slot and must not be redelivered meanwhile.
|
|
88
|
+
# Same key as #track, so the run's own tracking takes over after the
|
|
89
|
+
# hold is released. Returns the entry, or nil when the heartbeat is off.
|
|
90
|
+
def hold(client:, queue_name:, msg_id:, prefixed: true, job_class: nil, config: Pgbus.configuration)
|
|
91
|
+
return unless config.visibility_heartbeat
|
|
92
|
+
|
|
93
|
+
entry = Entry.new(client: client, queue_name: queue_name, prefixed: prefixed, msg_id: msg_id.to_i,
|
|
94
|
+
job_class: job_class, extended_at: monotonic_now, extensions: 0)
|
|
95
|
+
register(entry, config)
|
|
96
|
+
entry
|
|
97
|
+
end
|
|
98
|
+
|
|
99
|
+
# Drop a #hold. A nil entry (heartbeat off) is a no-op.
|
|
100
|
+
def release(entry)
|
|
101
|
+
unregister(entry) if entry
|
|
102
|
+
end
|
|
103
|
+
|
|
84
104
|
# Extend every tracked message whose last extension is older than the
|
|
85
105
|
# heartbeat interval. Public so tests and callers without the thread
|
|
86
106
|
# can drive it.
|
|
@@ -127,11 +147,22 @@ module Pgbus
|
|
|
127
147
|
end
|
|
128
148
|
end
|
|
129
149
|
|
|
150
|
+
# Only the entry that is registered under the key: a stale hold released
|
|
151
|
+
# after #track re-registered the same message must not drop the running
|
|
152
|
+
# job's entry (issue #486).
|
|
130
153
|
def unregister(entry)
|
|
131
|
-
synchronize
|
|
154
|
+
synchronize do
|
|
155
|
+
key = key_for(entry)
|
|
156
|
+
entries.delete(key) if entries[key].equal?(entry)
|
|
157
|
+
end
|
|
132
158
|
end
|
|
133
159
|
|
|
134
160
|
def extend!(entry, now:, config:)
|
|
161
|
+
# tick! picked this entry under the mutex but runs here without it: if
|
|
162
|
+
# the entry was released meanwhile (a read-ahead claim handed back with
|
|
163
|
+
# vt: 0, a job that finished), extending it would undo that.
|
|
164
|
+
return unless registered?(entry)
|
|
165
|
+
|
|
135
166
|
vt = config.visibility_timeout
|
|
136
167
|
entry.client.set_visibility_timeout(entry.queue_name, entry.msg_id, vt: vt, prefixed: entry.prefixed)
|
|
137
168
|
entry.extended_at = now
|
|
@@ -218,6 +249,10 @@ module Pgbus
|
|
|
218
249
|
@entries ||= {}
|
|
219
250
|
end
|
|
220
251
|
|
|
252
|
+
def registered?(entry)
|
|
253
|
+
synchronize { entries[key_for(entry)].equal?(entry) }
|
|
254
|
+
end
|
|
255
|
+
|
|
221
256
|
def key_for(entry)
|
|
222
257
|
[entry.queue_name, entry.msg_id]
|
|
223
258
|
end
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: pgbus
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.17.
|
|
4
|
+
version: 0.17.1
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Mikael Henriksson
|
|
@@ -348,6 +348,7 @@ files:
|
|
|
348
348
|
- lib/pgbus/pgmq_schema/pgmq_v1.11.1.sql
|
|
349
349
|
- lib/pgbus/pgmq_schema/pgmq_v1.12.0.sql
|
|
350
350
|
- lib/pgbus/pgmq_schema/pgmq_v1.13.0.sql
|
|
351
|
+
- lib/pgbus/process/claim_buffer.rb
|
|
351
352
|
- lib/pgbus/process/consumer.rb
|
|
352
353
|
- lib/pgbus/process/consumer_priority.rb
|
|
353
354
|
- lib/pgbus/process/dispatcher.rb
|
|
@@ -376,6 +377,7 @@ files:
|
|
|
376
377
|
- lib/pgbus/recurring/scheduler.rb
|
|
377
378
|
- lib/pgbus/recurring/task.rb
|
|
378
379
|
- lib/pgbus/retry_backoff.rb
|
|
380
|
+
- lib/pgbus/ruby_jit.rb
|
|
379
381
|
- lib/pgbus/serializer.rb
|
|
380
382
|
- lib/pgbus/stat_buffer.rb
|
|
381
383
|
- lib/pgbus/streams.rb
|