async-rabbitmq 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/README.md ADDED
@@ -0,0 +1,594 @@
1
+ # async-rabbitmq
2
+
3
+ A fiber-native RabbitMQ (AMQP 0-9-1) client for Ruby, built on the
4
+ [async](https://github.com/socketry/async) ecosystem and
5
+ [amq-protocol](https://github.com/ruby-amqp/amq-protocol). No threads: the
6
+ reader, writer, heartbeat and every consumer handler are fibers, so it fits
7
+ Falcon, async-http and anything else running under the fiber scheduler.
8
+
9
+ Requires Ruby 3.4+ (CI runs 3.4 and 4.0) and RabbitMQ 3.13+ (tested against 4.x).
10
+
11
+ ```ruby
12
+ require "async"
13
+ require "async_rabbitmq"
14
+
15
+ Sync do
16
+ session = AsyncRabbitMQ::Session.from_uri("amqp://guest:guest@localhost/%2F")
17
+ session.connect
18
+
19
+ channel = session.open_channel
20
+ exchange = channel.topic("events", durable: true)
21
+ queue = channel.durable_queue("events.audit")
22
+ queue.bind(exchange: exchange.name, routing_key: "audit.#")
23
+
24
+ exchange.publish("user 42 logged in", routing_key: "audit.login", persistent: true)
25
+
26
+ queue.subscribe(manual_ack: true) do |delivery, header, body|
27
+ puts "#{header.properties[:content_type]}: #{body}"
28
+ channel.basic_ack(delivery.delivery_tag)
29
+ end
30
+
31
+ sleep 1
32
+ session.close
33
+ end
34
+ ```
35
+
36
+ ## Connecting
37
+
38
+ ```ruby
39
+ AsyncRabbitMQ::Session.new(
40
+ host: "localhost", port: 5672, vhost: "/", username: "guest", password: "guest",
41
+ hosts: ["rabbit1", "rabbit2"], # or addresses: ["rabbit1:5672", "rabbit2:5673"]
42
+ tls: false, # implied by any of the four below
43
+ tls_cert: nil, tls_key: nil, # client certificate and key: a path or PEM text
44
+ tls_ca_certificates: nil, # CA path(s) or PEM text; default is the system store
45
+ verify_peer: true, tls_min_version: :TLS1_2,
46
+ tls_context: nil, # or an OpenSSL::SSL::SSLContext you built, for anything else
47
+ heartbeat: 60, frame_max: 131_072, channel_max: 2047,
48
+ connect_timeout: 30, rpc_timeout: 15, # seconds; rpc_timeout: nil waits forever
49
+ auth_mechanism: nil, # "PLAIN" / "EXTERNAL", negotiated by default
50
+ connection_name: "orders-worker",
51
+ auto_recover: true, recovery_attempts: nil, recovery_interval: 1.0, recovery_max_interval: 30.0,
52
+ recover_topology: true, topology_recovery_filter: nil,
53
+ instrumenter: nil, # ->(event_name, payload) { ... }
54
+ notifier: nil, # or one Notifier shared between sessions
55
+ logger: AsyncRabbitMQ::Log.new # anything with debug/info/warn/error
56
+ )
57
+ ```
58
+
59
+ `Session.from_uri` accepts one or more `amqp://` / `amqps://` URIs and the
60
+ standard query parameters (`heartbeat`, `connection_timeout`, `channel_max`,
61
+ `auth_mechanism`; on `amqps://` also `verify`, `cacertfile`, `certfile`,
62
+ `keyfile`). Several URIs form a failover list. Keyword arguments override the
63
+ URI.
64
+
65
+ `Session#update_secret(new_secret, reason)` rotates the credential on a live
66
+ connection (for refreshed OAuth 2 tokens); the new value is used for reconnects.
67
+ `Session#store_secret` records it without sending it, for a session that is
68
+ down while the secret rotates.
69
+
70
+ ### TLS
71
+
72
+ The ordinary case needs no OpenSSL: give it the CA that signed the broker's
73
+ certificate, and a client certificate and key if the broker asks for one. Each
74
+ value is a file path or PEM text, so material from a secrets manager works as
75
+ is. A certificate value may carry the leaf followed by its chain, and a CA
76
+ value may hold several certificates.
77
+
78
+ ```ruby
79
+ AsyncRabbitMQ::Session.new(host: "rabbit", port: 5671,
80
+ tls_ca_certificates: "/etc/ssl/rabbit-ca.pem")
81
+ AsyncRabbitMQ::Session.new(host: "rabbit", port: 5671,
82
+ tls_cert: ENV["RABBIT_CERT"], tls_key: ENV["RABBIT_KEY"],
83
+ tls_ca_certificates: [ENV["RABBIT_CA"]])
84
+ ```
85
+
86
+ Peers and their hostnames are verified unless `verify_peer: false`, and TLS 1.2
87
+ is the floor unless `tls_min_version:` says otherwise. `amqps://` URIs take the
88
+ same things as `cacertfile`, `certfile`, `keyfile` and `verify`. For anything
89
+ this does not cover, pass `tls_context:` and it is used as given.
90
+
91
+ ### Logging
92
+
93
+ `logger:` takes anything that responds to `debug`, `info`, `warn` and `error`,
94
+ so pass whatever your application already uses. The gem does not depend on the
95
+ `logger` gem, which stopped being a default gem in Ruby 4.0.
96
+
97
+ ```ruby
98
+ AsyncRabbitMQ::Session.new(logger: Rails.logger)
99
+ AsyncRabbitMQ::Session.new(logger: Logger.new($stdout, level: Logger::INFO))
100
+ AsyncRabbitMQ::Session.new(logger: AsyncRabbitMQ::Log.silent) # say nothing
101
+ AsyncRabbitMQ::Session.new(logger: AsyncRabbitMQ::Log.new($stdout, level: :debug))
102
+ ```
103
+
104
+ The default is `AsyncRabbitMQ::Log`, which writes warnings and errors, one line
105
+ each, to `$stderr`. Structured events (see below) are the better hook for
106
+ metrics and tracing; the log is for the things a human should read.
107
+
108
+ ## Channels
109
+
110
+ Everything synchronous (`queue`, `exchange`, `basic_consume`, ...) is a
111
+ request/reply that blocks only the calling fiber and gives up after
112
+ `rpc_timeout` with `RpcTimeoutError`. A channel may be shared by many fibers:
113
+ requests are serialised per channel and a publish's frames always reach the
114
+ wire contiguously.
115
+
116
+ ```ruby
117
+ ch = session.open_channel(pool_size: 4) # up to 4 consumer handlers at once on this channel
118
+ session.with_channel { |ch| ... } # closed afterwards
119
+
120
+ ch.queue("name", durable: true, exclusive: false, auto_delete: false, arguments: {})
121
+ ch.durable_queue("name") # classic, durable, non-exclusive, non-auto-delete
122
+ ch.quorum_queue("name") # x-queue-type: quorum
123
+ ch.stream("name") # x-queue-type: stream
124
+ ch.temporary_queue # server-named, exclusive, auto-delete
125
+ ch.queue("name", passive: true) # assert it exists (404 ChannelError otherwise)
126
+
127
+ ch.exchange("name", type: :topic, durable: true) # or ch.direct / fanout / topic / headers
128
+ ch.default_exchange
129
+
130
+ ch.basic_publish(body, exchange: "", routing_key: "q", persistent: true, content_type: "text/plain",
131
+ headers: {}, correlation_id: "...", reply_to: "...", expiration: "60000", ...)
132
+ ch.basic_publish_batch([b1, b2, b3], routing_key: "q") # one write, best throughput
133
+
134
+ ch.basic_get("q") # => [delivery_info, header, body] or nil; manual ack by default
135
+ ch.basic_get("q", manual_ack: false)
136
+ ch.basic_ack(tag) / basic_nack(tag, requeue: true) / basic_reject(tag, requeue: true)
137
+ ch.basic_qos(prefetch_count: 10) # also sizes the handler pool
138
+ tag = ch.basic_consume("q", manual_ack: true) { |delivery, header, body| ... }
139
+ ch.basic_cancel(tag)
140
+ ch.each("q") { |delivery, header, body| ... } # blocks until the consumer or channel goes away
141
+ ```
142
+
143
+ Consumer handlers run in their own fibers, at most `pool_size` at a time per
144
+ channel. Set `basic_qos` before `basic_consume`: without a prefetch limit the
145
+ broker sends the whole queue as fast as it can and one fiber is created per
146
+ delivery, so memory tracks queue depth rather than concurrency. The client
147
+ warns once per channel if you don't.
148
+
149
+ ### Acknowledging, and delivery tags across a reconnect
150
+
151
+ `delivery.delivery_tag` is a `VersionedDeliveryTag`: the broker's tag plus the
152
+ generation of the channel it arrived on. It converts (`to_int`), compares,
153
+ sorts, hashes and prints as the integer it wraps, so passing it to
154
+ `basic_ack`, comparing it or putting it in a collection all work unchanged.
155
+ Call `.to_i` if you need a real `Integer` — to serialise it, or where something
156
+ tests `is_a?(Integer)`.
157
+
158
+ The generation is what makes an ack safe across a reconnect. Delivery tags are
159
+ scoped to a channel on a connection, and the broker restarts numbering at 1
160
+ when a channel is reopened, so a handler that was still running when the
161
+ connection dropped would otherwise acknowledge whichever message now holds that
162
+ number:
163
+
164
+ ```ruby
165
+ ch.basic_qos(prefetch_count: 10)
166
+ ch.basic_consume("q", manual_ack: true) do |delivery, _header, body|
167
+ handle(body) # may outlive the connection
168
+ ch.basic_ack(delivery.delivery_tag) # => false if the connection went; nothing is sent
169
+ end
170
+ ```
171
+
172
+ `basic_ack`, `basic_nack` and `basic_reject` return `true` when the frame was
173
+ sent and `false` when the tag belonged to an earlier generation and was
174
+ dropped. A dropped ack is not a lost message: the broker requeued it when the
175
+ channel went, and it is redelivered on the new connection.
176
+
177
+ When a handler raises, the delivery is nacked with `requeue: false` — so a
178
+ message that always fails dead-letters instead of looping — and
179
+ `on_handler_error` is called. A delivery the handler already settled itself is
180
+ left alone, so acking and then raising in whatever follows is safe:
181
+
182
+ ```ruby
183
+ ch.on_handler_error do |error, delivery, queue_name|
184
+ Sentry.capture_exception(error, extra: { queue: queue_name })
185
+ end
186
+ ```
187
+
188
+ ### RabbitMQ 4.2+ and transient queues
189
+
190
+ RabbitMQ 4.2 and later refuse to declare a queue that is both non-durable
191
+ and non-exclusive; the broker answers with a **connection-level** error
192
+ (`541 INTERNAL_ERROR`, "Feature transient_nonexcl_queues is deprecated")
193
+ and this client reconnects. The AMQP default (`durable: false,
194
+ exclusive: false`) is therefore rejected by a current broker. Use
195
+ `durable: true` (add `"x-expires"` to have idle queues clean themselves up),
196
+ or `exclusive: true` / `temporary_queue` for a queue that should live only as
197
+ long as the connection.
198
+
199
+ ## Publisher confirms
200
+
201
+ ```ruby
202
+ ch.confirm_select
203
+ ch.basic_publish(...) # => delivery tag
204
+ ch.wait_for_confirms # true, or false if the broker nacked something
205
+ ch.nacked_tags # tags rejected since confirm_select
206
+ ch.unconfirmed_tags
207
+
208
+ ch.confirm_select(tracking: true, outstanding_limit: 1000)
209
+ # publishes park while 1000 messages are unconfirmed (backpressure);
210
+ # wait_for_confirms raises AsyncRabbitMQ::MessageNacked on a nack.
211
+
212
+ ch.wait_for_confirms(timeout: 5) # ConfirmTimeoutError instead of waiting forever
213
+ ch.unconfirmed_messages # what the broker has not resolved: payload, routing and options
214
+ ```
215
+
216
+ `wait_for_confirms` waits indefinitely by default. Pass `timeout:` if a broker
217
+ that accepts a publish and then never confirms it should raise
218
+ `ConfirmTimeoutError` (which carries `unconfirmed_tags`) rather than park the
219
+ fiber for good.
220
+
221
+ Messages published under confirms are kept until the broker acks them. If
222
+ the connection drops first they are re-published on the recovered channel
223
+ (the usual at-least-once trade-off: a message the broker had already
224
+ accepted may be delivered twice).
225
+
226
+ ### What a returned publish means here
227
+
228
+ Two differences from Bunny on the sending side, both deliberate.
229
+
230
+ **`basic_publish` returns when the frames are queued for the writer fiber,
231
+ not when they are on the socket.** Bunny writes inline on the calling thread,
232
+ so there a returned publish means the bytes reached the kernel. Neither is a
233
+ delivery guarantee: RabbitMQ's own guidance is that a client which has written
234
+ frames to its socket still cannot assume the broker received or processed
235
+ them. Publisher confirms are the only thing that tells you. What the queue
236
+ does change is the size of the window. Closing a session with a backlog
237
+ discards whatever the writer has not reached yet, because the close handshake
238
+ is bounded at five seconds; the client logs a warning naming the number of
239
+ queued writes it dropped, since a publisher without confirms has no other way
240
+ to find out. If it matters, `confirm_select` and `wait_for_confirms` before
241
+ `close`.
242
+
243
+ **A channel may be published to from many fibers at once.** RabbitMQ's
244
+ documentation says concurrent publishing on a shared channel is not supported
245
+ by client libraries, and for most clients that is true. Here publishes and
246
+ request/reply calls are serialised per channel, so a message's frames always
247
+ reach the wire contiguously and confirm tags follow wire order. Sharing a
248
+ channel is still a throughput bottleneck, and synchronous calls queue behind a
249
+ publish backlog, so give a busy publisher its own channel when latency on
250
+ declares matters.
251
+
252
+ ## Transactions
253
+
254
+ `ch.tx_select`, `ch.tx_commit`, `ch.tx_rollback`, `ch.using_tx?`. A channel
255
+ cannot be both transactional and in confirm mode.
256
+
257
+ ### The reactor does not stay alive on its own
258
+
259
+ The client's background work — the frame reader and writer, channel dispatch,
260
+ heartbeats, recovery — runs in transient tasks, so it never holds the reactor
261
+ open by itself. A program that connects, subscribes and then lets its main task
262
+ finish exits immediately, without running the consumer and without an error:
263
+
264
+ ```ruby
265
+ Async do
266
+ session.connect
267
+ ch = session.open_channel
268
+ ch.basic_qos(prefetch_count: 10)
269
+ ch.basic_consume("q") { |delivery, header, body| ... }
270
+ end # returns at once; nothing was waiting
271
+ ```
272
+
273
+ Block on something you own instead:
274
+
275
+ ```ruby
276
+ Async do
277
+ session.connect
278
+ ch = session.open_channel
279
+ ch.basic_qos(prefetch_count: 10)
280
+ ch.each("q") { |delivery, header, body| ... } # blocks until the consumer or channel goes away
281
+ end
282
+ ```
283
+
284
+ `Async::Condition#wait`, a `sleep`, or your own supervisor task work equally
285
+ well. Versions up to 0.3.0 kept the reactor alive by accident, because those
286
+ tasks were children of whichever task called `connect`.
287
+
288
+ ## Connection recovery
289
+
290
+ When the connection is lost the session reconnects with exponential backoff
291
+ (`recovery_interval` doubling up to `recovery_max_interval`, ±25% jitter,
292
+ `recovery_attempts: nil` = forever), trying every address in the list.
293
+ Meanwhile each channel is in the *recovering* state: operations already
294
+ waiting for a reply raise `ConnectionError`, and new publishes and requests
295
+ **park** until the channel is reopened, so nothing issued during the outage
296
+ is lost. After reconnect the session
297
+
298
+ 1. reopens every channel with its prefetch, confirm and tx settings,
299
+ 2. replays the recorded topology (`recover_topology`): exchanges, queues
300
+ (server-named queues come back under a new name, which is propagated to
301
+ their bindings, consumers and `Queue` objects), then bindings,
302
+ 3. re-registers consumers and releases the parked callers,
303
+ 4. calls `on_recovery`.
304
+
305
+ Passive declares are not recorded; deleted or unbound entities are not
306
+ recovered. The registry follows what the broker does rather than which channel
307
+ declared what: a durable or exclusive queue outlives the channel that declared
308
+ it and is still recovered, while an auto-delete queue is forgotten, with its
309
+ bindings, the moment its last consumer goes, whether by `basic_cancel`, a
310
+ broker-side cancel or the channel closing, and an auto-delete exchange goes
311
+ when its last binding does. A worker that opens a channel and a temporary
312
+ queue per unit of work therefore leaves nothing behind for the next reconnect
313
+ to re-declare. `topology_recovery_filter:` takes an object implementing any of
314
+ `filter_exchanges`, `filter_queues`, `filter_queue_bindings`,
315
+ `filter_exchange_bindings`. If recovery is disabled, exhausted
316
+ (`on_recovery_exhausted`) or the broker refuses the credentials, channels are
317
+ closed and parked callers raise.
318
+
319
+ Callbacks: `on_connection_lost` (the connection dropped; fires before recovery
320
+ starts), `on_blocked` / `on_unblocked` (connection.blocked),
321
+ `on_recovery_attempt`, `on_recovery`, `on_recovery_exhausted`;
322
+ `Channel#on_return` (mandatory messages the broker could not route),
323
+ `on_cancel` (broker cancelled a consumer), `on_error` (broker closed the
324
+ channel). A broker-closed channel can be reopened in place with
325
+ `Channel#reopen`.
326
+
327
+ ## Errors
328
+
329
+ All errors derive from `AsyncRabbitMQ::Error`: `ConnectionTimeoutError`
330
+ (could not connect), `AuthenticationError` (credentials or vhost refused,
331
+ `code` 403), `ConnectionError` (`code`, `text`; broker closed the
332
+ connection or it was lost), `ChannelError` (`code`, `text`, `channel_id`,
333
+ `close_method`, plus `delivery_ack_timeout?`, `unknown_delivery_tag?`,
334
+ `message_too_large?`), `NotOpenError`, `RpcTimeoutError`, `MessageNacked`
335
+ (`nacked_tags`), `ConfirmTimeoutError` (`unconfirmed_tags`, `channel_id`),
336
+ `ClusterError` (a cluster-wide operation that did not succeed on every node),
337
+ `HeartbeatTimeoutError`, `ChannelLimitError` (every id up to
338
+ the negotiated `channel_max` is in use; the broker would otherwise have closed
339
+ the connection). Channel ids are reused as channels close, so opening a
340
+ channel per unit of work is fine for the life of the connection.
341
+
342
+ ## Events
343
+
344
+ `Session#on_event` subscribes to structured events for metrics, tracing and
345
+ debugging. Nothing is emitted until something subscribes: the payload is built
346
+ inside a block that only runs when a subscriber is listening.
347
+
348
+ ```ruby
349
+ session.on_event { |name, payload| logger.info("#{name} #{payload}") }
350
+ session.on_event("message.") { |name, _| statsd.increment(name) }
351
+ session.on_event("channel.rpc") { |_, p| histogram.record(p[:duration]) }
352
+ session.on_event(/^recovery\./) { |name, p| pager.note(name, p) }
353
+
354
+ # or hand the whole stream to something at construction
355
+ AsyncRabbitMQ::Session.new(instrumenter: ->(name, payload) { ... })
356
+ ```
357
+
358
+ The pattern is `nil` for everything, a String for one event name or, ending in
359
+ a dot, a prefix, or a Regexp. `on_event` returns a handle that
360
+ `session.notifier.unsubscribe(handle)` takes back. A subscriber that raises is
361
+ logged and skipped; it never breaks the connection.
362
+
363
+ | Event | Payload |
364
+ |---|---|
365
+ | `connection.open` | `host`, `port`, `vhost`, `tls`, `heartbeat`, `frame_max`, `channel_max`, `duration` |
366
+ | `connection.closed` | `reason` (`:user`), `host`, `port` |
367
+ | `connection.lost` | `host`, `port`, `error`, `message`, `recovering` |
368
+ | `connection.blocked` | `reason` |
369
+ | `connection.unblocked` | — |
370
+ | `recovery.attempt` | `attempt`, `delay` |
371
+ | `recovery.succeeded` | `attempts`, `host`, `port`, `channels`, `duration` |
372
+ | `recovery.exhausted` | `attempts`, `reason` (`:attempts_exceeded`, `:authentication_failed`) |
373
+ | `heartbeat.sent` | `interval` |
374
+ | `channel.open` | `channel` |
375
+ | `channel.closed` | `channel`, `reason` (`:user`, `:broker`, `:dropped`), `code`, `text` |
376
+ | `channel.rpc` | `channel`, `method` (`"queue.declare-ok"`), `duration` |
377
+ | `consumer.registered` | `channel`, `queue`, `consumer_tag`, `manual_ack` |
378
+ | `consumer.cancelled` | `channel`, `consumer_tag`, `queue`, `reason` (`:client`, `:broker`) |
379
+ | `message.published` | `channel`, `exchange`, `routing_key`, `count`, `bytes`, `delivery_tag` |
380
+ | `message.confirmed` | `channel`, `delivery_tag`, `multiple`, `acked` |
381
+ | `message.returned` | `channel`, `exchange`, `routing_key`, `code`, `text`, `bytes` |
382
+ | `message.consumed` | `channel`, `queue`, `consumer_tag`, `bytes`, `redelivered`, `duration` |
383
+
384
+ Durations are seconds as a Float. `bytes` on `message.published` is the encoded
385
+ frames, including the header and properties; elsewhere it is the body. A batch
386
+ published with `basic_publish_batch` is one event with `count` set and the last
387
+ delivery tag. Event names are API and do not change without a major version;
388
+ `AsyncRabbitMQ::Notifier::EVENTS` lists them.
389
+
390
+ ### OpenTelemetry
391
+
392
+ Tracing is a separate, optional layer. Add `opentelemetry-api` to your bundle,
393
+ then:
394
+
395
+ ```ruby
396
+ require "async_rabbitmq/telemetry/open_telemetry"
397
+ AsyncRabbitMQ::Telemetry::OpenTelemetry.install
398
+ ```
399
+
400
+ Spans and attributes match `opentelemetry-instrumentation-bunny`, so a service
401
+ moving over from Bunny keeps the traces and dashboards it had. Publishing opens
402
+ a PRODUCER span `"<exchange>.<routing key> publish"` and injects the W3C trace
403
+ context into the message headers; `basic_get` opens a CONSUMER span
404
+ `"<destination> receive"`; and a consumer handler runs inside a CONSUMER span
405
+ `"<destination> process"` whose parent is the context extracted from the
406
+ headers, so one trace spans both sides of the broker. Attributes are
407
+ `messaging.system`, `messaging.destination`, `messaging.destination_kind`,
408
+ `messaging.protocol`, `messaging.protocol_version`,
409
+ `messaging.rabbitmq.routing_key`, `messaging.operation`, `net.peer.name` and
410
+ `net.peer.port`, plus `messaging.batch.message_count` on a batch publish.
411
+
412
+ `install` takes `tracer_provider:`, `tracer_name:` and `tracer_version:`.
413
+ `uninstall` stops tracing. One difference from Bunny: a pushed delivery here
414
+ goes straight to the handler, so there is one process span per delivery rather
415
+ than a receive span with a process span under it.
416
+
417
+ ## Connection pool
418
+
419
+ Add `gem "async-pool"` to your Gemfile: it is not a dependency of this gem, so
420
+ that an application that never pools does not install it.
421
+
422
+ ```ruby
423
+ require "async_rabbitmq/pool"
424
+ pool = AsyncRabbitMQ::Pool.new(max: 5, host: "localhost")
425
+ pool.acquire { |session| session.with_channel { |ch| ch.basic_publish(...) } }
426
+ pool.close
427
+ ```
428
+
429
+ Backed by `Async::Pool`; a session is shared by up to `channel_max` fibers
430
+ before another connection is opened.
431
+
432
+ ## One connection per cluster node
433
+
434
+ A `Session` is on one node. Every channel it opens is there, and so is every
435
+ exclusive queue declared through it, because the broker always places those on
436
+ the connecting node, whatever `queue_leader_locator` says. A process that holds
437
+ a channel and a temporary queue per client therefore puts all of its clients on
438
+ one node, and loses them together when that node goes.
439
+
440
+ `AsyncRabbitMQ::Cluster` takes the same options as `Session` and presents the
441
+ same methods, so it drops in where a `Session` was. It pins one connection to
442
+ each address, opens each new channel on the node with the fewest, and retries a
443
+ node that is down until it is back, after which new channels drift to it until
444
+ the counts are level. Nothing is ever moved: a channel stays on its node for
445
+ life.
446
+
447
+ ```ruby
448
+ cluster = AsyncRabbitMQ::Cluster.new(
449
+ addresses: %w[rabbit1:5672 rabbit2:5672 rabbit3:5672], # one connection each
450
+ username: "app", password: secret,
451
+ on_node_down: :drop, # :park (default), :drop, or ->(session, channels, error) { ... }
452
+ clear_topology_on_drop: true # with :drop, forget the lost connection's transient topology
453
+ )
454
+ cluster.connect # every node, concurrently; the ones that are down are retried in the background
455
+ channel = cluster.open_channel # on the node with the fewest channels
456
+ cluster.on_node_down { |session, channels, error| ... }
457
+ # declare a fourth parameter to also receive the publishes the broker never
458
+ # confirmed, so they can be sent again on another node:
459
+ cluster.on_node_down do |session, channels, error, unconfirmed|
460
+ unconfirmed.each { |m| elsewhere.basic_publish(m.payload, exchange: m.exchange, routing_key: m.routing_key, **m.options) }
461
+ end
462
+ cluster.on_node_up { |session| ... }
463
+ ```
464
+
465
+ `on_node_down:` decides what happens to the channels on a node whose connection
466
+ is lost. `:park` is what a `Session` does: they wait, and resume on the same
467
+ node when it returns. `:drop` closes them at once: calls in flight raise
468
+ `ConnectionError`, consumer loops return, later calls raise `NotOpenError`, and
469
+ the fibers using them can open a new channel, which lands on a node that is up.
470
+ A callable receives the node's session, its channels and the error while the
471
+ channels are parked, and closes the ones it wants dropped; the rest wait. With
472
+ `:drop` the lost connection's transient topology is forgotten as well
473
+ (`clear_topology_on_drop: true`): exclusive, auto-delete and server-named
474
+ queues, auto-delete exchanges, and the bindings that referred to them. Durable
475
+ exchanges, queues and bindings are kept, so a node that comes back from an
476
+ empty data directory still has them re-declared; `false` keeps the usual
477
+ registry rules. Publishing and shared durable
478
+ topology are expected to live on a plain `Session` alongside: the cluster is
479
+ for the mass of per-client channels.
480
+
481
+ `connect` connects every node concurrently and returns once each one is either
482
+ connected or has failed a first attempt, so the channels opened next are spread
483
+ over every node that is reachable. The wait is bounded by one `connect_timeout`
484
+ however many nodes are down. It raises only if none can be reached, or at once
485
+ with `AuthenticationError` if a node refuses the credentials. Nodes that are
486
+ down at startup, or later, are retried with the recovery backoff.
487
+ `open?` is true while any node is; `close` closes them all. `open_channel`,
488
+ `with_channel`, `queue_exists?` and `exchange_exists?` use the least loaded
489
+ node; `update_secret` stores the new secret on every node first and then updates each
490
+ one that is up, so a node that refuses it cannot leave the nodes after it on
491
+ the old credential; failures are raised together as `ClusterError`.
492
+ The `on_*` callbacks fire for every node with that node's `Session`, `on_event`
493
+ sees every node's events (each carries `host` and `port`), `topology` is a
494
+ read-only view over every node's registry, and `host` and `port` are the first
495
+ address. `sessions`, `on_node_down` and `on_node_up` are the additions.
496
+
497
+ ## Command line
498
+
499
+ Installing the gem puts an `async-rabbitmq` command on your path, written on
500
+ this API, for the things you would otherwise open a console for.
501
+
502
+ ```bash
503
+ export RABBITMQ_URL=amqp://guest:guest@localhost:5672 # or pass --url
504
+
505
+ async-rabbitmq publish orders '{"id":1}' --count 10 --persistent
506
+ async-rabbitmq publish orders --file payload.json
507
+ echo '{"id":2}' | async-rabbitmq publish orders
508
+ async-rabbitmq publish events.audit --exchange events # queue name becomes the routing key
509
+
510
+ async-rabbitmq inspect orders # orders: 12 messages, 2 consumers
511
+ async-rabbitmq consume orders --count 5
512
+ async-rabbitmq consume orders --peek # print without acknowledging: nothing is removed
513
+ async-rabbitmq purge orders
514
+ ```
515
+
516
+ `publish` uses confirms and mandatory routing, so it exits non-zero and says so
517
+ when the broker nacks a message or sends it back unroutable. `consume`
518
+ acknowledges what it prints, stops after `--count` or `--timeout` seconds of
519
+ quiet, and leaves anything beyond the count on the queue. `--quiet` prints the
520
+ messages and nothing else, for piping. Options go after the command.
521
+
522
+ ## Performance and integrity harness
523
+
524
+ `examples/` holds a sender and a receiver that load the broker and check what
525
+ comes out the other end. They are meant for the failure modes a throughput
526
+ number hides: a message lost, doubled or delivered out of sequence, a body
527
+ handed over with another message's header, a delivery on the wrong queue, a
528
+ reply given to the wrong caller.
529
+
530
+ ```bash
531
+ ruby -Ilib examples/perf_consumer.rb --streams 4 # start first
532
+ ruby -Ilib examples/perf_publisher.rb --streams 4 --messages 50000
533
+ ```
534
+
535
+ Every message is deterministic in `(run, stream, seq, size)` and states its
536
+ identity twice, in the AMQP properties and inside the body, so the receiver can
537
+ rebuild the bytes it should have been given and compare. One queue per
538
+ publisher stream, one consumer, one handler: that is where AMQP promises order,
539
+ so out-of-sequence delivery there is a real fault, and the report says so.
540
+ `--consumers` or `--handlers` above 1 makes deliveries concurrent and the
541
+ report downgrades ordering to an observation.
542
+
543
+ The sender publishes from several fibers over **one shared channel** by default
544
+ (`--channels per-stream` for one each), which is the case worth stressing: it
545
+ checks that confirm delivery tags are unique and cover every publish, that
546
+ nothing came back unroutable, and, with `--rpc-probe N`, that passive declares
547
+ issued from other fibers while the channel is saturated each get their own
548
+ reply. Useful switches: `--batch`, `--confirms none|simple|tracking`,
549
+ `--persistent`, `--rate`, `--size`. The receiver takes `--mode get` to exercise
550
+ the `basic_get` path instead of a consumer, plus `--prefetch` and
551
+ `--[no-]manual-ack`. Both exit non-zero when anything failed, so they can gate
552
+ a build.
553
+
554
+ `examples/perf_fault_inject.rb` publishes one of each fault on purpose; run it
555
+ against a receiver to see the checks fire rather than trusting them.
556
+
557
+ Two things the numbers do not say. Latency is measured from a monotonic clock,
558
+ so both programs must run on the same host, and it only means anything while
559
+ the receiver keeps up (pace the sender with `--rate`). And with `--confirms
560
+ none`, nothing proves the broker received anything: see *What a returned
561
+ publish means here* above.
562
+
563
+ ## Development
564
+
565
+ The suite runs against a real RabbitMQ (with Toxiproxy for network faults):
566
+
567
+ ```bash
568
+ docker compose -f spec/docker-compose.yml up -d --wait
569
+ bundle exec rspec
570
+ ```
571
+
572
+ `spec_helper` starts the containers itself if nothing listens on the
573
+ configured port, and generates the TLS certificates with
574
+ `spec/docker/gen-certs.sh`. Ports and hosts come from `.env`.
575
+
576
+ ## Releasing
577
+
578
+ The version lives in `lib/async_rabbitmq/version.rb` and nowhere else. To cut a
579
+ release: bump it, date the section in `CHANGELOG.md`, run the suite against a
580
+ real broker, then
581
+
582
+ ```bash
583
+ gem build async-rabbitmq.gemspec # writes async-rabbitmq-<version>.gem
584
+ gem install ./async-rabbitmq-<version>.gem # optional: check it installs and the command runs
585
+ gem push async-rabbitmq-<version>.gem # asks for your RubyGems OTP
586
+ git tag -a v<version> -m "v<version>" && git push origin v<version>
587
+ ```
588
+
589
+ The gemspec sets `rubygems_mfa_required`, so publishing and yanking need
590
+ multi-factor authentication on the RubyGems account.
591
+
592
+ ## License
593
+
594
+ MIT, see [LICENSE](LICENSE).
@@ -0,0 +1,12 @@
1
+ #!/usr/bin/env ruby
2
+ # frozen_string_literal: true
3
+
4
+ $LOAD_PATH.unshift(File.expand_path("../lib", __dir__)) if File.directory?(File.expand_path("../lib", __dir__))
5
+
6
+ # The frame writer uses IO::Buffer, still flagged experimental in Ruby 4.0.
7
+ # A command-line tool should print its own output, not the interpreter's.
8
+ Warning[:experimental] = false if Warning.respond_to?(:[]=)
9
+
10
+ require "async_rabbitmq/cli"
11
+
12
+ exit AsyncRabbitMQ::CLI.start(ARGV)