async-rabbitmq 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: c8e21899276061f5e0ecd749000e4b7bf5b81932a491a1704e0b2e643cf647b1
4
+ data.tar.gz: a7fcd5f434e5e677ad7fce840ca388e323264edd8b2347f8ed92ec73a0fd7a3e
5
+ SHA512:
6
+ metadata.gz: 0a41b23ef4e9d1fdad8403e8dbc99783f53205ed6b0fe5cb3e891b79d9763b9a9c0ec05bceda15f78de1fb47e0fb5b65bdbc5b756c2d19cf9c6609d2ee13fd3c
7
+ data.tar.gz: 4e7795488172face945464fbd3d36c9355b7b6500705238079df4ae38b0c6529c67088c84323354c0772bd31562b9efb7b1927c06872b369addc4d56638c5815
data/CHANGELOG.md ADDED
@@ -0,0 +1,268 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project will be documented in this file.
4
+
5
+ ## [0.4.0] - 2026-10-07
6
+
7
+ Reliability review against a payments workload. The two delivery-correctness items
8
+ (stale acks, empty bodies) are the reason to take this.
9
+
10
+ ### Fixed
11
+
12
+ - **Acks are no longer applied to the wrong message after a reconnect.** Delivery tags are
13
+ scoped to a channel on a connection: the broker restarts numbering at 1 when a channel is
14
+ reopened. A handler still running when the connection dropped would ack that number against
15
+ the new channel, acknowledging a message it never processed (a whole range of them with
16
+ `multiple: true`), or draw a 406 that closed the channel and took its consumers with it.
17
+ Deliveries now carry a `VersionedDeliveryTag`, and `basic_ack` / `basic_nack` / `basic_reject`
18
+ drop a tag from an earlier generation and return `false`.
19
+ - **A message with an empty body no longer hangs the channel.** Neither RabbitMQ nor
20
+ amq-protocol sends a body frame when the body is zero bytes, so the delivery never completed
21
+ and the next method frame on that channel was taken as its content — a handler called with
22
+ garbage and a lost delivery. Zero-length content is now routed on its header.
23
+ - **A half-dead connection is detected in heartbeat time rather than in TCP retransmit time.**
24
+ The heartbeat task wrote before it checked liveness, and that write needs the socket lock, so
25
+ a writer parked in `@socket.write` after a partition blocked the heartbeat too and nothing
26
+ noticed for around 15 minutes. Liveness is checked first, the write is bounded, and
27
+ `TCP_USER_TIMEOUT` is set where the platform has it.
28
+ - **TLS connections no longer strand a TCP socket per reconnect.** `sync_close` was never set,
29
+ so closing the SSL socket left its transport open until the GC ran.
30
+ - **A second connection drop during recovery no longer closes channels permanently.**
31
+ `@recovery_in_progress` was cleared before channels were reopened, so a drop in that window
32
+ started a competing recovery while the first was still failing channels through `mark_closed!`.
33
+ Channels are now left `:recovering` and reopened by the next attempt. A `frame_io` that a later
34
+ attempt has replaced no longer reports its dying socket as the live connection's.
35
+ - **A consumer block that raises no longer stalls the consumer.** The delivery stayed unacked
36
+ for the life of the connection, and once `prefetch` messages were stuck that way nothing more
37
+ arrived. Handlers are wrapped: the delivery is nacked (`requeue: false`, so a poison message
38
+ dead-letters instead of looping) and `Channel#on_handler_error` is called.
39
+ - **An RPC timeout no longer leaves the channel out of step.** AMQP replies carry nothing to
40
+ match them to a request, so a reply that arrived after the caller gave up went to whoever
41
+ asked next. The channel is closed and reopened, the late reply discarded, its unconfirmed
42
+ publishes replayed (which is also what resets the client's confirm tag counter to match the
43
+ broker, since it restarts at 1 on the reopened channel), and its consumers re-registered.
44
+ The channel reports `resyncing?` for the duration and parks other fibers, so nothing publishes
45
+ onto a channel that is closing — the broker would drop it silently — and no other request has
46
+ its reply discarded. Confirms that arrive during the window are still applied.
47
+ - **A handler that acks and then raises no longer closes the channel.** The automatic nack was
48
+ unconditional, so a delivery the handler had already settled was nacked a second time: a 406
49
+ that closed the channel and took its consumers with it. Settled deliveries are tracked and
50
+ skipped.
51
+ - **Background loops survive the task that started them.** Channel dispatch, the heartbeat, the
52
+ channel-0 monitor, the frame reader and writer, the recovery loop and every `Cluster` node
53
+ supervisor were children of whichever task opened the channel or called `connect`. A
54
+ short-lived caller took them down with it while the session still reported itself open. They
55
+ are parented at the reactor instead.
56
+ - **The frame writer starts recovery on a non-IO error** instead of exiting quietly and leaving
57
+ publishers blocked once the write queue filled.
58
+ - **Recovery that has to retry tears down the half-built connection first.** A drop while
59
+ channels were being reopened left the session reporting `:open` (so a `Cluster` would place new
60
+ channels on a dead node), left the socket open, and leaked a channel-0 monitor fiber per flap.
61
+ - **`Cluster#update_secret` cannot leave nodes on different secrets.** One node raising skipped
62
+ every node after it. The secret is stored everywhere first, each node is then attempted, and
63
+ the failures are raised together as `ClusterError`.
64
+ - **OpenTelemetry no longer writes into the caller's headers hash** — a frozen hash raised, and
65
+ one shared between fibers raced. The hash is copied before `traceparent` is injected.
66
+ - **Tracking unconfirmed publishes no longer costs a quarter of the send throughput.** Keeping
67
+ the payload and routing for `unconfirmed_messages` built a keyword struct and copied the
68
+ options hash for every confirmed publish, and put it in a second hash alongside
69
+ `@pending_confirms`. Measured against 0.3.0 that cost 22% of publish throughput under simple
70
+ confirms and 33% with `basic_publish_batch`. The routing is now shared by one publish call
71
+ (a batch builds it once), the record is built when `unconfirmed_messages` is read rather than
72
+ when the message is sent, and it lives in the hash that was already there. Back to 0.3.0
73
+ throughput: batch +0.2%, simple confirms within run-to-run spread.
74
+ - **`Notifier#publish` iterates a snapshot**, so a subscriber that unsubscribes itself no longer
75
+ causes the next subscriber to be skipped.
76
+
77
+ ### Added
78
+
79
+ - `Channel#on_handler_error` — called when a consumer block raises, with the exception, the
80
+ `Basic::Deliver` frame and the queue name.
81
+ - `Channel#wait_for_confirms(timeout:)` — raises `ConfirmTimeoutError` instead of waiting
82
+ forever for a broker that accepts a publish and never confirms it. Unbounded by default.
83
+ - `Channel#unconfirmed_messages` — the publishes the broker has not resolved, as
84
+ `UnconfirmedMessage` records carrying the payload, its routing and the publish options
85
+ (`persistent`, `headers`, `message_id`, `correlation_id`, ...), so they can be republished on
86
+ another connection without silently losing their properties. A `Cluster#on_node_down` block that declares a fourth parameter is
87
+ handed them; under `:drop` that is the only chance to see them.
88
+ - `Session.new(tcp_user_timeout:)` — milliseconds, defaulting to twice the heartbeat.
89
+ - `Channel#delivery_generation` — how many times the channel has been opened on a connection.
90
+ The broker restarts both delivery tags and publisher confirm tags at 1 on a reopened channel,
91
+ so a tag only identifies a message together with its generation. Deliveries carry theirs on
92
+ the tag itself; confirm tags from `basic_publish` are plain integers, so code keeping its own
93
+ confirm bookkeeping across a reconnect reads this.
94
+ - `Channel#resyncing?` — true while the channel is being reopened, after an RPC timeout or
95
+ through `Channel#reopen`. Operations park until it is back, as they do during connection
96
+ recovery, and unconfirmed messages are replayed before it goes `:open` so a new publish
97
+ cannot take a delivery tag that no longer matches its place on the wire.
98
+ - `TopologyRegistry#clear_transient`.
99
+
100
+ ### Changed
101
+
102
+ - **A connection no longer keeps the reactor alive by itself.** The background loops are
103
+ transient tasks now, so a program that calls `connect` and `basic_consume` and then lets its
104
+ main task finish exits immediately instead of running the consumer — silently, with no error.
105
+ This is correct Async behaviour (nothing was waiting), but it is a change from 0.3.0, where
106
+ those tasks were children of the caller and held the reactor open. The caller has to block on
107
+ something it owns:
108
+
109
+ ```ruby
110
+ Async do
111
+ session.connect
112
+ ch = session.open_channel
113
+ ch.basic_qos(prefetch_count: 10)
114
+ ch.each("q") { |delivery, header, body| ... } # blocks until the consumer goes away
115
+ end
116
+ ```
117
+
118
+ `Async::Condition#wait`, `sleep`, or anything else that parks the task will do.
119
+ - **`delivery_tag` is a `VersionedDeliveryTag`, not an `Integer`.** It converts (`to_int`),
120
+ compares, sorts, hashes and prints as the integer it wraps, so `basic_ack(di.delivery_tag)`,
121
+ comparisons and array membership are unaffected. Code calling `.is_a?(Integer)` on it, or
122
+ serialising it directly, needs `.to_i`.
123
+ - **`clear_topology_on_drop:` now clears only what the lost connection owned** — exclusive,
124
+ auto-delete and server-named queues, auto-delete exchanges, and the bindings that referred to
125
+ them. Durable topology is kept, so a node that comes back from an empty data directory still
126
+ has it re-declared. Previously the whole registry for that node was dropped.
127
+ - `basic_ack`, `basic_nack` and `basic_reject` return `true` when sent and `false` when the tag
128
+ was stale, where they previously returned the write's result.
129
+ - `Channel#basic_consume` warns once per channel when called without a prior `basic_qos`: the
130
+ broker sends the whole queue as fast as it can and one task is created per delivery, so memory
131
+ tracks queue depth rather than concurrency.
132
+
133
+ ## [0.3.0] - 2026-09-16
134
+
135
+ ### Added
136
+
137
+ - `AsyncRabbitMQ::Cluster`: one connection to every node of a cluster behind the Session API,
138
+ for a process with a channel per client. Each new channel opens on the node with the fewest;
139
+ a node that is down is retried until it is back and then takes new channels until the counts
140
+ are level. `on_node_down:` decides what a lost node's channels do: `:park` (wait, as a Session
141
+ does), `:drop` (close them at once so their fibers move to another node) or a callable that
142
+ closes the ones to drop. With `:drop` the lost connection's topology is forgotten
143
+ (`clear_topology_on_drop: true`). `on_node_down` / `on_node_up` callbacks.
144
+ - `Session#on_connection_lost`, `Session#channels`, `Session#channel_count`,
145
+ `Session#store_secret`, and a `notifier:` option to share one `Notifier` between sessions.
146
+ - `channel.closed` events for a channel given up during recovery: `reason: :dropped`, or `:user`
147
+ when the application closed it.
148
+
149
+ ## [0.2.0] - 2026-09-14
150
+
151
+ Review against Bunny 3.3 / amq-protocol 2.9 and RabbitMQ 4.3 (issues #21–#42).
152
+
153
+ ### Fixed
154
+
155
+ - Channel ids are allocated from a bitset and released when a channel closes. They were a bare
156
+ counter: a process opening a channel per unit of work walked it to 65535, after which the id
157
+ wrapped to 0 and the broker dropped the connection. `Session#open_channel` also refuses to
158
+ exceed the negotiated `channel_max` with `ChannelLimitError`, where before the broker closed
159
+ the whole connection with a 530.
160
+ - The topology registry tracks consumers and forgets an auto-delete queue, with its bindings,
161
+ when its last consumer goes (cancel, broker-side cancel or channel close), and an auto-delete
162
+ exchange when its last binding goes. Before, a worker using a temporary queue per channel
163
+ accumulated one dead queue per channel and re-declared them all on every reconnect.
164
+ - Publishing while the connection is being recovered no longer silently drops the message:
165
+ channels enter a `:recovering` state and park new publishes and requests until they are
166
+ reopened. Messages published under confirms are kept until acked and re-published after a
167
+ reconnect instead of being cleared and reported as confirmed (#21).
168
+ - `basic_get` returned the previous message's body from the second call on (#42).
169
+ - `AsyncRabbitMQ::Pool.new` raised `NameError`; it now uses the `Async::Pool::Controller`
170
+ constructor API (#22).
171
+ - `Session#connect` cleans up after a broker-refused handshake, tries the next address, and
172
+ raises the broker's error; `connection.close` 403 raises `AuthenticationError`. The client
173
+ advertises the standard capabilities (`authentication_failure_close` among them), so bad
174
+ credentials get a 403 instead of a dropped socket; a connection dropped mid-handshake fails
175
+ immediately; a failed connect no longer disables recovery on a reused session (#23).
176
+ - `Session#close` hung while the broker had `connection.blocked` active: only publishes are
177
+ gated, and parked publishers are released with `ConnectionError` on close (#24).
178
+ - Request/reply operations and publish frames are serialised per channel, so a shared channel
179
+ no longer hands replies to the wrong fiber or interleaves frames under backpressure (#25).
180
+ - Heartbeats are sent every T/2 instead of every T (#26).
181
+ - `Channel#each` returns when the channel is closed, the consumer is cancelled (by the client
182
+ or the broker) or the broker closes the channel (#27).
183
+ - `Session.from_uri` percent-decodes credentials and vhost via `AMQ::URI` (a literal `+` is
184
+ preserved) and honours the standard query parameters (#28).
185
+ - `wait_for_confirms` returns `false` when the broker nacked a message; `nacked_tags` and
186
+ `unconfirmed_tags` expose the details (#29).
187
+ - The reader loop no longer logs a warning and re-enters recovery when the socket was closed
188
+ locally.
189
+ - Closing a session whose write queue has not drained within the 5 s close handshake now logs
190
+ how many queued writes are being discarded, instead of dropping them silently. Publishing
191
+ without confirms is still fire-and-forget; the warning just makes the loss visible
192
+ (`FrameIO#pending_writes` exposes the count).
193
+ - A failed TLS handshake no longer leaks the socket it was wrapping, which cost one file
194
+ descriptor per connect attempt against a broker with a bad certificate.
195
+ - Unconfirmed messages are re-published after the topology has been replayed, not during
196
+ channel reopen. Against a broker that lost the topology, a message re-published first hit a
197
+ missing exchange (404, closing the channel again) or a missing binding (dropped while the
198
+ broker acked it).
199
+ - Prefetch is restored after connection recovery.
200
+ - Test suite: RabbitMQ 4.2+ rejects transient non-exclusive queues; specs declare durable or
201
+ exclusive queues (#41).
202
+
203
+ ### Added
204
+
205
+ - TLS without OpenSSL plumbing: `tls_cert:`, `tls_key:`, `tls_ca_certificates:`, `verify_peer:`
206
+ and `tls_min_version:` on `Session.new`, each taking a path or PEM text, with chains and CA
207
+ bundles split as needed (`AsyncRabbitMQ::TLS`). Certificate material implies `tls: true`.
208
+ `tls_context:` remains for anything the options do not cover.
209
+ - An `async-rabbitmq` command: `publish`, `consume`, `inspect` and `purge`, built on the gem's
210
+ own API. Publishing uses confirms and the mandatory flag and reports a nack or an unroutable
211
+ message with a non-zero exit; `consume --peek` prints without acknowledging.
212
+ - Optional OpenTelemetry tracing in `async_rabbitmq/telemetry/open_telemetry`, with the span
213
+ names, attributes and W3C context propagation of `opentelemetry-instrumentation-bunny`.
214
+ Not required by default and not a runtime dependency.
215
+ - Structured events for metrics and tracing: `Session#on_event(pattern) { |name, payload| }`
216
+ and a `instrumenter:` constructor option. 18 events covering the connection, recovery,
217
+ channels, consumers and messages, with durations on `channel.rpc`, `message.consumed`,
218
+ `connection.open` and `recovery.succeeded`. Nothing is emitted, and no payload is built,
219
+ until something subscribes; a subscriber that raises is logged and skipped.
220
+ - Confirm tracking: `confirm_select(tracking: true, outstanding_limit: 1000)` gives publishers
221
+ backpressure and raises `MessageNacked` on a nack (#32).
222
+ - `Channel#basic_publish_batch` (#33).
223
+ - `Channel#reopen` after a broker-initiated channel close (#34).
224
+ - Topology recovery: exchanges, queues and bindings declared through the session are
225
+ re-declared after a reconnect, with server-named queue renames propagated;
226
+ `recover_topology:` and `topology_recovery_filter:` options; `Session#topology` (#35).
227
+ - `Session#update_secret` (`connection.update-secret`) (#36).
228
+ - Transactions: `tx_select`, `tx_commit`, `tx_rollback`, `using_tx?` (#37).
229
+ - `rpc_timeout:` (default 15 s) bounds synchronous channel operations with `RpcTimeoutError`
230
+ (#38).
231
+ - `Channel#durable_queue`, `Channel#stream`, `Queue::Types`, `Exchange::TYPE_*` constants;
232
+ `quorum_queue` and `stream` reject empty names (#39).
233
+ - `Queue#pop` (alias `get`) (#40).
234
+ - `ChannelError#close_method` with `delivery_ack_timeout?`, `unknown_delivery_tag?`,
235
+ `message_too_large?`; `AuthenticationError#code` / `#text` (#30).
236
+ - `Session.new` options `connect_timeout:`, `channel_max:`, `rpc_timeout:`,
237
+ `recover_topology:`, `topology_recovery_filter:`; `FrameIO#blocked?`.
238
+ - Since 0.1.1, before this review: `Session.from_uri`, multi-host failover (`hosts:`,
239
+ `addresses:`, `hosts_shuffle_strategy:`), `auto_recover:` and configurable recovery with
240
+ `on_recovery_attempt` / `on_recovery` / `on_recovery_exhausted`, `on_blocked` /
241
+ `on_unblocked` / `Channel#on_error`, client properties and `connection_name:`, pluggable
242
+ SASL mechanisms with `connection.secure` handling, per-channel consumer `pool_size`, the
243
+ Bunny convenience API (`direct` / `fanout` / `topic` / `headers` / `default_exchange`,
244
+ `temporary_queue`, `quorum_queue`, `Queue#status`, predicates, `with_channel`,
245
+ `queue_exists?` / `exchange_exists?`, named message properties, `Exchange#on_return`).
246
+
247
+ ### Changed
248
+
249
+ - No dependency on the `logger` gem, which stopped being a default gem in Ruby 4.0 and would
250
+ otherwise have to be installed by every application using this one. `logger:` takes any
251
+ object responding to `debug`, `info`, `warn` and `error`; the default is now
252
+ `AsyncRabbitMQ::Log`, which writes warnings and errors to `$stderr` instead of the previous
253
+ default of everything, including debug, to `$stdout`. `AsyncRabbitMQ::Log.silent` says
254
+ nothing.
255
+ - `AsyncRabbitMQ::Pool` explains that `async-pool` has to be in your Gemfile instead of
256
+ failing with `cannot load such file`.
257
+ - `basic_get` defaults to `manual_ack: true`, as in Bunny (#40).
258
+ - `amq-protocol` requirement raised to `~> 2.9` (#30).
259
+ - Requires Ruby 3.4 or later; CI runs Ruby 3.4 and 4.0.
260
+ - A message's frames are written as one buffer.
261
+
262
+ ## [0.1.1] - 2026-04-02
263
+
264
+ ### Added
265
+
266
+ - Expose `passive:` keyword argument on `Channel#queue` and `Channel#exchange` for defensive
267
+ startup checks that assert a resource exists without creating it (AMQP 0-9-1 passive declare).
268
+ - Integration tests for passive declare (exists and 404 paths) for both queues and exchanges.
data/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Russell Penney
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.