raptor 0.17.0 → 0.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +4 -0
- data/README.md +34 -17
- data/docs/brisrails-talk.md +119 -88
- data/docs/raptor-vs-puma.md +75 -63
- data/lib/rackup/handler/raptor.rb +7 -5
- data/lib/raptor/cli.rb +21 -0
- data/lib/raptor/cluster.rb +13 -3
- data/lib/raptor/stats.rb +1 -1
- data/lib/raptor/version.rb +1 -1
- data/sig/generated/raptor/cli.rbs +11 -2
- data/sig/generated/raptor/cluster.rbs +3 -0
- data/sig/generated/raptor/stats.rbs +1 -1
- metadata +1 -1
data/docs/raptor-vs-puma.md
CHANGED
|
@@ -1,35 +1,34 @@
|
|
|
1
1
|
# Raptor vs Puma: A Design Comparison
|
|
2
2
|
|
|
3
|
-
Raptor is a Ruby web server
|
|
3
|
+
Raptor is a Ruby web server built around Ractor-parallel protocol pipelines, a lock-free app thread pool, and an opinionated cluster architecture. Against the latest Puma release on the same hardware, it leads clearly on the IO-heavy HTTP/1.1 benchmark and lands within a few percent on CPU-heavy throughput, while Falcon's fibers remain the strongest fit for highly concurrent IO. Raptor also speaks HTTP/2 natively, which Puma does not. This document explains how those designs move requests and where each trade-off shows up.
|
|
4
4
|
|
|
5
5
|
## Why this document exists
|
|
6
6
|
|
|
7
7
|
Raptor began as a curiosity project. Ruby 4.0 was landing with a more polished Ractor implementation, and I wanted to see what a web server would look like if you actually leaned into parallel Ruby instead of pretending the GVL was not there. The initial goal was modest. Build something that could parse HTTP in parallel using Ractors, hook it up to Rack, and see if the numbers moved.
|
|
8
8
|
|
|
9
|
-
They did, but the more interesting result was structural.
|
|
9
|
+
They did, but the more interesting result was structural. Committing to Ractor-based protocol pipelines made every boundary around them visible: which requests are worth sending through a Ractor, how partial requests return to the reactor, how completed requests reach ordinary app threads, and how concurrent HTTP/2 responses share one socket. It also pushed the rest of the server toward cheap timeout updates, a lock-free app work queue, explicit backpressure, and cluster dispatch that accounts for current load. No one of those choices explains the benchmark. Together they define the server.
|
|
10
10
|
|
|
11
11
|
This document walks through both servers at systems-design depth. By the end you should be able to describe how Puma and Raptor each move a request from `accept()` to `app.call(env)` and back, why the two architectures make the design decisions they do, and where the performance delta actually comes from. No source code required.
|
|
12
12
|
|
|
13
13
|
## A note on Falcon
|
|
14
14
|
|
|
15
|
-
Falcon is the other next-generation Ruby web server worth naming. It takes a different bet than either Puma or Raptor. Concurrency comes from fibers via the `async` gem instead of threads, HTTP/2 is native, and each request is a lightweight task rather than a slot in
|
|
15
|
+
Falcon is the other next-generation Ruby web server worth naming. It takes a different bet than either Puma or Raptor. Concurrency comes from fibers via the `async` gem instead of threads, HTTP/2 is native, and each request is a lightweight task rather than a slot in an app thread pool. Its strengths show up on workloads with lots of concurrent long-lived connections (WebSockets, SSE, streaming), applications built end-to-end on the `async` ecosystem, and HTTP/2-heavy traffic. On a traditional Rails workload where most of the request budget is a synchronous DB round-trip through ActiveRecord, its advantages over Puma are less pronounced because the fiber scheduler only helps when the underlying I/O is fiber-aware. If your app fits that sweet spot, Falcon belongs in your evaluation.
|
|
16
16
|
|
|
17
17
|
This document focuses on Puma because Puma is the incumbent that any new Ruby web server has to justify itself against; that is the comparison most readers actually need. Falcon appears in the benchmark table in the README as a third data point, but a full design comparison against Falcon would be its own document.
|
|
18
18
|
|
|
19
19
|
## The shape of the benchmark
|
|
20
20
|
|
|
21
|
-
Raptor is a research project. It hasn't run production traffic. The numbers in the README come from a repeatable microbenchmark, not from a real deployment.
|
|
21
|
+
Raptor is a research project. It hasn't run production traffic. The numbers in the README come from a repeatable microbenchmark, not from a real deployment. They use controlled Rack workloads to measure the full server path: accepting connections, parsing, dispatching, running the app, and writing responses. A real Rails application adds database time, downstream services, middleware, and application-specific contention, so a percentage here will not transfer directly to production.
|
|
22
22
|
|
|
23
|
-
The Raptor README carries the [current head-to-head numbers](../README.md#micro-benchmarks) against the latest Puma and Falcon releases, run on the same hardware with the same Rack app on a recent Ruby with YJIT enabled. All servers run one worker process per available CPU
|
|
23
|
+
The Raptor README carries the [current head-to-head numbers](../README.md#micro-benchmarks) against the latest Puma and Falcon releases, run on the same hardware with the same Rack app on a recent Ruby with YJIT enabled. All servers run one worker process per available CPU. Puma and Raptor's fixed mode use three app threads per worker. Raptor's scaling mode starts at three and has no fixed limit, while Falcon uses unbounded fibers. Load generators use four client connections per baseline app thread, so on the 10-core machine that produced the current numbers that's 120 concurrent HTTP/1.1 client connections and 40 h2 connections × 3 streams each.
|
|
24
24
|
|
|
25
25
|
Two workload profiles are measured. **IO-bound** is a GET endpoint that does 5 to 10 short sleeps interleaved with small CPU work per request, simulating a read path that makes several DB or cache calls throughout its lifetime. **CPU-bound** is a POST endpoint with a small JSON body that builds a JSON response in 3 to 5 chunks interleaved with sub-100µs sleeps, simulating a write path that does most of its work in Ruby with a few near-zero-cost cache hits. The workloads are interleaved rather than a single bulk sleep or single bulk serialise so a fiber-per-connection server like Falcon doesn't look artificially good from one-shot IO, and the CPU-bound workload is heavily CPU-dominated by design (roughly 95% CPU / 5% IO by wall time) so it actually measures CPU work rather than smuggling in enough IO for fibers to multiplex.
|
|
26
26
|
|
|
27
|
-
Each cell in the table reports the median throughput and median p95 latency independently across
|
|
27
|
+
Each cell in the table reports the median throughput and median p95 latency independently across 3 runs, and every run boots a fresh server process so state cannot accumulate across measurements. Both Raptor modes are compared with both servers: fixed mode is the like-for-like comparison with Puma, while scaling mode tests whether adaptive OS threads can close the IO-concurrency gap with Falcon's much cheaper fibers. Rather than pin every number into this document (they drift with Ruby versions and hardware), the shape of the result is what matters.
|
|
28
28
|
|
|
29
|
-
- On IO-bound
|
|
30
|
-
- On CPU-bound HTTP/1.1
|
|
31
|
-
- On
|
|
32
|
-
- On HTTP/2, Raptor and Falcon both implement it; Puma doesn't. On CPU-bound h2 Raptor lands close to Falcon on throughput and holds a meaningful edge on p95; on IO-bound h2 Falcon dominates for the same reason it dominates h1 IO. This matters if you're terminating h2 at the app server, less so if nginx or another proxy in front is already handling it.
|
|
29
|
+
- On IO-bound HTTP/1.1, fixed Raptor delivers a little over twice Puma's throughput with less than half its p95 latency. Scaling lifts Raptor to 8.63k req/s without keep-alive and 9.43k with it: still 29.3% behind Falcon on fresh connections, but 51.3% ahead with keep-alive.
|
|
30
|
+
- On CPU-bound HTTP/1.1, fixed and scaling Raptor are effectively identical. Puma leads by 5.0% without keep-alive and 0.7% with it, while both Raptor modes lead Falcon. This is the intended result: the pool gets its IO gains without trading away CPU throughput.
|
|
31
|
+
- On HTTP/2, Raptor and Falcon both implement it; Puma doesn't. Scaling raises Raptor's IO throughput from 1.14k to 4.46k req/s, though it remains 30.3% behind Falcon. On the CPU profile it narrows Falcon's throughput lead from 26.9% to 14.2% while keeping a lower p95. HTTP/2 also varies substantially more between runs in this benchmark, so those medians deserve less confidence than the stable HTTP/1.1 results.
|
|
33
32
|
|
|
34
33
|
The rest of this doc explains why the shape looks like that.
|
|
35
34
|
|
|
@@ -39,16 +38,16 @@ The rest of this doc explains why the shape looks like that.
|
|
|
39
38
|
| ---------------------- | ------------------------------------------------------- | ---------------------------------------------------------------------------- |
|
|
40
39
|
| Ruby requirement | 3.0 and up | 4.0 and up (needs `Ractor::Port`) |
|
|
41
40
|
| Process model | Single, or cluster (pre-forks by default with 2+ workers); optional refork | Cluster only, always pre-forks |
|
|
42
|
-
| Threading model | `ThreadPool` with `Mutex` + `ConditionVariable`; autoscaling min/max, or fixed when min == max |
|
|
43
|
-
| True parallelism |
|
|
41
|
+
| Threading model | `ThreadPool` with `Mutex` + `ConditionVariable`; autoscaling min/max, or fixed when min == max | `AtomicThreadPool` with a CAS-based queue; fixed by default, with GVL-aware adaptive scaling up to a configured limit |
|
|
42
|
+
| True parallelism | One GVL inside each worker; parser runs on an app thread | Protocol Ractors can parse in parallel with the app on the pipeline path |
|
|
44
43
|
| I/O multiplexing | `nio4r` reactor for keep-alive idle and slow reads | `nio4r` reactor for the same, plus a red-black tree for O(log n) timeouts |
|
|
45
|
-
|
|
|
46
|
-
| Work queue | `
|
|
44
|
+
| Cluster dispatch | Workers race on inherited listeners with a load-proportional accept delay | Two-choice load-aware BPF dispatch for TCP on Linux; shared-listener fallback |
|
|
45
|
+
| Work queue | Ruby `Queue` coordinated under the pool mutex | Lock-free Michael-Scott FIFO queue |
|
|
47
46
|
| HTTP/2 | Not implemented | Native C parser + HPACK, lock-free per-connection frame writer |
|
|
48
47
|
| Keep-alive fast path | Same-thread inline dispatch when spare threads exist | Same-thread inline plus a `wait_readable(1ms)` micro-poll on the app thread |
|
|
49
48
|
| Native extensions | 1 (Ragel HTTP/1 parser + MiniSSL) | 3, all Ractor-safe (Ragel HTTP/1 parser; HTTP/2 parser + HPACK; `writev`, `sched_setaffinity`, `prctl` wrappers) |
|
|
50
49
|
| Shared state (worker↔master) | Pipes and signals | Anonymous shared-memory `mmap` region |
|
|
51
|
-
| Restart primitives | Phased (USR1), hot (USR2 re-exec, inherits FDs via env), refork (SIGURG) | Phased (USR1), hot (USR2 re-exec, inherits FDs via env)
|
|
50
|
+
| Restart primitives | Phased (USR1), hot (USR2 re-exec, inherits FDs via env), refork (SIGURG) | Phased (USR1), hot (USR2 re-exec, inherits FDs via env), refork (SIGURG) |
|
|
52
51
|
| systemd integration | `sd_notify` via plugin, `LISTEN_FDS` via binder | Native `sd_notify` + `LISTEN_FDS` socket activation |
|
|
53
52
|
|
|
54
53
|
The rest of the document expands on each row.
|
|
@@ -67,7 +66,7 @@ Signals: INT and TERM start a graceful shutdown, USR1 does a phased restart (inc
|
|
|
67
66
|
|
|
68
67
|
### Threading model
|
|
69
68
|
|
|
70
|
-
Inside a worker, request work is handled by `Puma::ThreadPool`. The pool has a minimum and maximum thread count, and it autoscales; when work arrives and there are more items queued than there are waiting threads, it spawns a new thread up to the max.
|
|
69
|
+
Inside a worker, request work is handled by `Puma::ThreadPool`. The pool has a minimum and maximum thread count, and it autoscales; when work arrives and there are more items queued than there are waiting threads, it spawns a new thread up to the max. Work sits in Ruby's thread-safe `Queue`, while a pool-level `Mutex` protects enqueue/dequeue coordination, thread counts, autoscaling, trimming, and shutdown. A `ConditionVariable` parks idle threads. Adding work signals the condvar; a waiting thread wakes, dequeues an item, and processes it.
|
|
71
70
|
|
|
72
71
|
This is a textbook thread pool. It works, and it has for a decade. But it has three characteristics worth noting for the comparison:
|
|
73
72
|
|
|
@@ -192,7 +191,7 @@ Two kinds of restart are supported:
|
|
|
192
191
|
|
|
193
192
|
Systemd socket activation is a native feature and slots straight into this model. When the service unit is `Type=notify` and there is a socket unit, systemd passes listener FDs via `LISTEN_FDS`. Raptor detects this exactly the same way it detects a hot restart handoff: `Systemd.listen_fds` returns the FDs, the binder is built from them, and the master sends `READY=1` back to systemd once workers have booted. `STOPPING=1` and `RELOADING=1` fire on the corresponding lifecycle events.
|
|
194
193
|
|
|
195
|
-
|
|
194
|
+
Routine worker monitoring does not use pipes. Every worker writes its stats (pid, request count, backlog, busy and available threads, last checkin timestamp, booted flag) into a fixed-size slot in an anonymous shared-memory region allocated with `mmap-ruby` before the fork. The master reads the region directly. There is no serialisation, pipe drain, or signal to trigger the read; it is 49 bytes per worker of native memory. `bundle exec raptor stats` prints a JSON snapshot. Refork coordination is separate and does use a pair of pipes between the master and seed.
|
|
196
195
|
|
|
197
196
|
On Linux, each worker pins itself to a distinct CPU via `sched_setaffinity` when the worker count fits within the process's allowed CPU set, so it stays on one core and its L1/L2 caches stay warm. When workers outnumber available CPUs the pin is skipped and the kernel scheduler manages placement.
|
|
198
197
|
|
|
@@ -212,37 +211,38 @@ The design is Pitchfork's, adapted for Raptor's process model. [Pitchfork](https
|
|
|
212
211
|
|
|
213
212
|
### Threading model
|
|
214
213
|
|
|
215
|
-
This is where Raptor diverges dramatically. Inside a worker there are
|
|
214
|
+
This is where Raptor diverges dramatically. Inside a worker there are several distinct layers of concurrent activity:
|
|
216
215
|
|
|
217
216
|
1. **One server thread** running the accept loop.
|
|
218
217
|
2. **One reactor thread** running the NIO event loop plus timeout tree.
|
|
219
218
|
3. **A `RactorPool` for HTTP/1.1 parsing** sized by `http1.ractors`, defaulting to `round(cores / workers)` clamped to `[1, 3]`.
|
|
220
219
|
4. **A `RactorPool` for HTTP/2 parsing** sized by `http2.ractors`, defaulting to `round(cores / workers)` clamped to `[1, 2]`.
|
|
221
220
|
5. **A collector thread per Ractor pool** that receives parsed results via a `Ractor::Port`.
|
|
222
|
-
6. **An `AtomicThreadPool` of T app threads** running the Rack app and writing responses.
|
|
223
|
-
7.
|
|
221
|
+
6. **An `AtomicThreadPool` of T baseline app threads** running the Rack app and writing responses, with optional adaptive growth.
|
|
222
|
+
7. **One stats thread** that writes the shared-memory slot every second.
|
|
223
|
+
8. **One load reporter thread when BPF dispatch is active** that publishes backlog to the kernel map.
|
|
224
224
|
|
|
225
225
|
That is a lot of moving parts. Let us go through why.
|
|
226
226
|
|
|
227
|
-
**Why Ractors for parsing.** Ractors are Ruby's answer to true parallelism. Multiple Ractors can execute Ruby code simultaneously on different OS threads, each with its own GVL. But Ractors are heavily restricted.
|
|
227
|
+
**Why Ractors for parsing.** Ractors are Ruby's answer to true parallelism. Multiple Ractors can execute Ruby code simultaneously on different OS threads, each with its own GVL. But Ractors are heavily restricted. Shared objects must satisfy Ractor's shareability rules, most global mutable state is inaccessible, and many existing gems assume shared-state semantics that are incompatible with isolation.
|
|
228
228
|
|
|
229
229
|
For a web server, this restriction turns out to be almost exactly right for HTTP parsing. Parsing a request is CPU-bound (tokenising bytes, uppercasing header names, decoding chunked bodies), it does not need to touch any global state, and it produces a result (a hash) that can be safely frozen and handed off. The native HTTP/1 parser (`raptor_http.c`) is declared `rb_ext_ractor_safe(true)`; it holds no per-parser Ruby state in the extension itself, and it writes only into the caller-supplied env hash. Same for the HTTP/2 parser plus HPACK.
|
|
230
230
|
|
|
231
231
|
The HTTP/1 parser also pre-interns the ~40 most common header keys (`HTTP_HOST`, `HTTP_USER_AGENT`, the `HTTP_ACCEPT_*` family, `CONTENT_LENGTH`, `HTTP_X_FORWARDED_*`, the `HTTP_SEC_FETCH_*` client hints, and so on) once at load time. During parsing, a `memcmp` lookup against that table returns the shared frozen `VALUE` for known keys and falls back to `rb_enc_interned_str` for the rest. Every request's env hash therefore reuses the same String object for its header names, which both skips per-request allocation and lets Ruby's hash lookup use the interned key's cached hash code.
|
|
232
232
|
|
|
233
|
-
|
|
233
|
+
Your Rack app still runs on ordinary threads under one worker GVL, so the app does not need to be Ractor-safe. Requests that need the reactor pipeline can have their protocol work run in parallel in another Ractor. Complete requests found by the eager accept and keep-alive paths parse inline instead, avoiding a Ractor handoff when the bytes are already available.
|
|
234
234
|
|
|
235
235
|
**How the Ractor pools actually work.** Raptor uses the `ractor-pool` gem, which is another one of my libraries. Each pool has one coordinator Ractor and M pipeline Ractors. When a pipeline Ractor is idle, it sends itself back to the coordinator via `coordinator.send(Ractor.current, move: true)`. When work arrives at the coordinator, it either forwards it to a waiting Ractor (if any) or queues it. This coordinator-dispatch pattern guarantees that no Ractor sits idle while there is work. Results flow back through a shared `Ractor::Port` (a many-to-one channel added in recent Ruby versions and stable in 4.0) to a Ruby-side collector thread. If `M == 1` the coordinator is skipped and work goes straight to the single pipeline Ractor.
|
|
236
236
|
|
|
237
237
|
Raptor runs two independent pools per worker, one for HTTP/1.1 parsing and one for HTTP/2 parsing. Both defaults scale with headroom via `round(cores / workers)`, clamped to `[1, 3]` for `http1.ractors` and `[1, 2]` for `http2.ractors`. Splitting the pools means h1 and h2 parsing never share ractor slots, so a burst of small HTTP/1.1 requests cannot delay HTTP/2 frame handling on the same connection, and vice versa.
|
|
238
238
|
|
|
239
|
-
|
|
239
|
+
**Why a custom thread pool.** The `AtomicThreadPool` in `atomic-ruby` (another one of my libraries) is backed by an `AtomicQueue`. The queue is a Michael-Scott multi-producer, multi-consumer FIFO: a singly linked list with a dummy sentinel and atomic head and tail pointers. Producers append nodes at the tail; consumers advance the head. Both operations are O(1) and make progress through compare-and-swap rather than a queue-wide mutex. Separate atoms track queue size and active app threads for backpressure.
|
|
240
240
|
|
|
241
|
-
|
|
241
|
+
The pool is fixed at `threads` by default. Setting `max_threads` above that baseline enables adaptive growth; `Float::INFINITY` removes the fixed limit. A native CRuby thread-event hook measures how much active workers spend running, blocked outside the GVL, and waiting to acquire it. The pool only adds a temporary thread after work has remained queued across several samples, every current worker is active, blocked time is above half of measured worker time, and GVL wait is below two percent. That last check matters: adding threads helps when existing threads are asleep in database or network calls, but hurts when CPU-bound Ruby threads are already fighting over the GVL. Temporary threads retire after the queue has remained empty for a second.
|
|
242
242
|
|
|
243
243
|
The pool still uses an `AtomicConditionVariable` under the hood to park idle threads (idle threads call `Thread.stop` and get woken with `Thread#wakeup`; there is no spinning), because idle spinning would waste CPU. The difference from Puma's pool is not "no locks anywhere" but rather "the hot path (enqueue and dequeue when the queue has items) is lock-free". Once every worker is busy the mechanics look similar; where things diverge is under contention when you have many threads all trying to push and pop.
|
|
244
244
|
|
|
245
|
-
The knock-on effect is that
|
|
245
|
+
The knock-on effect is that the server thread can read `pool.queue_size + pool.active_count` on every accept-loop iteration without acquiring the queue's mutation lock. Those are still synchronised atomic reads, but they do not serialise producers and consumers behind one mutex.
|
|
246
246
|
|
|
247
247
|
### I/O model
|
|
248
248
|
|
|
@@ -259,17 +259,19 @@ if @thread_pool.queue_size > @thread_pool.size
|
|
|
259
259
|
end
|
|
260
260
|
```
|
|
261
261
|
|
|
262
|
-
The first is a hard skip. `@reactor.backlog` is `thread_pool.queue_size + thread_pool.active_count`; when the total load reaches
|
|
262
|
+
The first is a hard skip. `@reactor.backlog` is `thread_pool.queue_size + thread_pool.active_count`; when the total load reaches the threshold, this worker stops accepting until it drains. `MIN_BACKPRESSURE_THRESHOLD` is 8, so a three-thread pool trips at 8 concurrent items rather than 4. The floor avoids overreacting to a few active requests while still bounding the amount of work a worker pulls from the kernel.
|
|
263
263
|
|
|
264
264
|
The second is a softer yield. When the queue alone exceeds the pool size — the app threads are all busy and there's a queue building on top of them — the accept loop yields the GVL via `Thread.pass` and re-checks the queue on the next iteration instead of accepting more work. This lets the app threads make progress before the server thread grabs another connection, and only fires under real pool pressure (queue depth greater than pool size), so IO-bound workloads where threads spend most of their time in `sleep` and rarely queue past the pool size aren't affected.
|
|
265
265
|
|
|
266
|
-
|
|
266
|
+
On Linux, Raptor can replace the shared TCP listener with one `SO_REUSEPORT` listener per worker and attach a small BPF program to the group. A reporter thread publishes each worker's backlog into a kernel map every millisecond. For each new connection, the program hashes to two distinct workers and selects the one with the lower reported load: the power-of-two-choices strategy. It then atomically increments that worker's map slot before routing the connection, reserving capacity immediately rather than waiting for the next reporter tick. The accept path also publishes `reactor.backlog + 1` as soon as Ruby receives the socket. These reservations stop a burst of connections from repeatedly choosing the same stale minimum while retaining hash-based spread across the cluster.
|
|
267
|
+
|
|
268
|
+
This path only wraps plain `tcp://` bindings. TLS listeners and non-TCP bindings remain shared listeners inherited from the master. If `libbpf-ruby` or the compiled BPF object is unavailable, Raptor silently uses those shared listeners for TCP too; if the prerequisites exist but the kernel refuses the program, startup raises. In either mode, a worker under Raptor's explicit backpressure stops competing for new accepts until its current work drains.
|
|
267
269
|
|
|
268
270
|
The BPF-based approach was inspired by [a comment](https://github.com/puma/puma/issues/3934#issuecomment-4356462590) by John Hawthorn ([@jhawthorn](https://github.com/jhawthorn)) on a Puma issue about `EPOLLEXCLUSIVE`, where he floated `SO_ATTACH_REUSEPORT_EBPF` as a way to route each connection to the least-busy worker.
|
|
269
271
|
|
|
270
272
|
The reactor is again an `NIO::Selector` loop. Two things make it different from Puma's:
|
|
271
273
|
|
|
272
|
-
1. **Read strategy.** When a socket is readable, the reactor does one `read_nonblock(64KB)`
|
|
274
|
+
1. **Read strategy.** When a socket is readable, the reactor does one `read_nonblock(64KB)` in the reactor thread, updates the buffered state, makes that state shareable, and sends it to the protocol's Ractor pool. The Ractor decides whether the request or frame batch is complete. The collector then sends incomplete state back to the reactor or dispatches completed requests to the app pool. The reactor itself does I/O, not protocol parsing.
|
|
273
275
|
|
|
274
276
|
2. **Timeout data structure.** Instead of a sorted linked list, timeouts are stored in a red-black tree (`red-black-tree` gem, yes, also one of mine). Each connection is represented by a `TimeoutClient < RedBlackTree::Node` ordered by its `timeout_at` value. Insertion is O(log n), deletion by key (needed when a connection's timeout is updated mid-flight, which happens on every read) is O(log n), and in-order traversal is O(k) where k is the number of expired connections. After every selector poll, the reactor walks the tree in order and breaks on the first non-expired node.
|
|
275
277
|
|
|
@@ -299,7 +301,7 @@ The **pipeline path** fires when the first read returns `WaitReadable` (bytes ha
|
|
|
299
301
|
8. An app thread pops the proc, builds a Rack env, calls `@app.call(env)`, and writes the response.
|
|
300
302
|
9. If the response signals keep-alive (HTTP/1.1 default without `Connection: close`), the app thread enters the **eager keep-alive loop**.
|
|
301
303
|
|
|
302
|
-
|
|
304
|
+
The eager keep-alive loop is one of Raptor's more deliberate latency/occupancy trade-offs. Rather than immediately returning the connection to the reactor after a response, the app thread does:
|
|
303
305
|
|
|
304
306
|
```ruby
|
|
305
307
|
loop do
|
|
@@ -312,9 +314,9 @@ loop do
|
|
|
312
314
|
end
|
|
313
315
|
```
|
|
314
316
|
|
|
315
|
-
The thread waits 1 millisecond for the next request. If bytes arrive in that window, it parses them inline on the same thread and calls the Rack app again.
|
|
317
|
+
The thread waits up to 1 millisecond for the next request. If bytes arrive in that window, it parses them inline on the same thread and calls the Rack app again. If no bytes arrive, the connection returns to the reactor. Pipelined requests therefore avoid a reactor round-trip, while an idle keep-alive connection occupies an app thread for at most that short polling window.
|
|
316
318
|
|
|
317
|
-
Puma has a similar shape but
|
|
319
|
+
Puma has a similar shape but does not wait. It checks buffered back-to-back requests, then eagerly drains bytes already available on the socket. If a complete request is ready and the pool has a waiting thread, the current thread loops inline; otherwise Puma queues the client or returns it to the reactor. Raptor's 1ms poll deliberately widens the window in which the next request can stay on the current app thread.
|
|
318
320
|
|
|
319
321
|
The `reactor.persist` call re-registers the socket with the reactor using `persistent_data_timeout` (65s) as the new deadline. When the next bytes arrive, the reactor treats the socket like any other partially-read connection.
|
|
320
322
|
|
|
@@ -340,7 +342,7 @@ So under contention, only one thread does socket I/O at a time (because a socket
|
|
|
340
342
|
|
|
341
343
|
Once the writer thread has claimed a batch of pending frames, it concatenates them into a single buffer and issues one socket write for the whole batch. For a typical response of a HEADERS frame plus several DATA frames, that is one SSL_write call rather than one per frame.
|
|
342
344
|
|
|
343
|
-
Flow control uses similar CAS-protected atoms. The connection-level window and the per-stream windows live in separate `Atom` cells
|
|
345
|
+
Flow control uses similar CAS-protected atoms. The connection-level window and the per-stream windows live in separate `Atom` cells. `acquire` atomically reserves connection capacity and, where per-stream tracking is needed, deducts the same grant from that stream's window. If either window is exhausted, the caller sleeps 1ms and retries until a `WINDOW_UPDATE` makes progress possible.
|
|
344
346
|
|
|
345
347
|
Frame processing also has an eager loop. After processing one batch of frames, the h2 handler tries to `read_nonblock` one more time to see if the next batch is already available. Up to eight rounds are consumed inline before handing back to the reactor, and the loop bails out early once the app thread pool has more queued work than worker slots so one busy connection cannot starve the collector. This is the same principle as the HTTP/1.1 eager keep-alive: amortise the reactor round-trip when the client is actively sending, but back off under saturation.
|
|
346
348
|
|
|
@@ -421,17 +423,17 @@ flowchart TB
|
|
|
421
423
|
class SHM storage
|
|
422
424
|
```
|
|
423
425
|
|
|
424
|
-
The critical structural difference from Puma is that
|
|
426
|
+
The critical structural difference from Puma is that Raptor has a separate protocol pipeline for connections that need more I/O. The reactor reads, a Ractor parses, a collector routes the result, and an app thread runs Rack and writes the response. Raptor's eager paths collapse that machinery when a complete request is already available. This is a hybrid rather than a rule that every request must cross a Ractor.
|
|
425
427
|
|
|
426
428
|
## Part III: Head to head
|
|
427
429
|
|
|
428
430
|
### Parsing model
|
|
429
431
|
|
|
430
|
-
**Puma.** Parsing happens on
|
|
432
|
+
**Puma.** Parsing happens on an app thread. The C parser callbacks build the env hash. A fresh client first enters the thread pool, where eager reads may complete it immediately; partial and idle keep-alive connections wait in the reactor before returning to the pool. Parsing shares the worker's GVL with the app.
|
|
431
433
|
|
|
432
|
-
**Raptor.**
|
|
434
|
+
**Raptor.** Fresh and immediate keep-alive requests parse inline on the server or app thread. A connection that needs more bytes takes the longer path: reactor (I/O) → protocol Ractor pool (parse) → collector → app thread pool (Rack + write). Between keep-alive requests, the app thread does a 1ms micro-poll before returning an idle connection to the reactor. Parsing in the Ractor pipeline has its own GVL; parsing on an eager path does not.
|
|
433
435
|
|
|
434
|
-
|
|
436
|
+
In practice, Puma has one process-wide GVL per worker. Every Ruby thread inside that worker takes turns holding it. Raptor has the same main-Ractor GVL plus one GVL per protocol Ractor, so a pipeline Ractor can parse one connection while an app thread executes Rack for another. That parallelism is real, but so are the costs of making state shareable and crossing the Ractor and collector boundaries. Which side wins depends on how much work the request gives the protocol pipeline; the current CPU benchmark leaves Raptor and Puma close rather than proving a universal parsing advantage.
|
|
435
437
|
|
|
436
438
|
### Timeout management
|
|
437
439
|
|
|
@@ -439,17 +441,17 @@ The concrete effect is that Puma has one process-wide GVL. Every thread inside a
|
|
|
439
441
|
|
|
440
442
|
**Raptor.** Red-black tree keyed by `timeout_at`. Insert O(log n), remove O(log n), in-order traversal breaks early on first non-expired node. Scales cleanly to thousands of connections.
|
|
441
443
|
|
|
442
|
-
At
|
|
444
|
+
At moderate connection counts this is unlikely to dominate either server. The asymptotic difference becomes more relevant as a worker tracks more idle or partial connections and updates more individual deadlines.
|
|
443
445
|
|
|
444
446
|
### Work queue
|
|
445
447
|
|
|
446
|
-
**Puma.**
|
|
448
|
+
**Puma.** Ruby `Queue` plus a pool-level `Mutex` and `ConditionVariable`. Every enqueue takes the mutex and wakes a waiter; every dequeue happens while holding the mutex. The same critical section coordinates pool bookkeeping and autoscaling.
|
|
447
449
|
|
|
448
|
-
**Raptor.**
|
|
450
|
+
**Raptor.** A Michael-Scott FIFO with atomic head and tail pointers and a dummy sentinel node. Producers link at the tail and consumers advance the head using CAS. Queue size and active-thread counts live in separate atoms.
|
|
449
451
|
|
|
450
452
|
The condvar is still there for parking idle threads (spinning would burn CPU), but the hot path when the queue has items is lock-free.
|
|
451
453
|
|
|
452
|
-
Under moderate load
|
|
454
|
+
Under moderate load, queue mechanics are unlikely to dominate either server. Under contention, Raptor avoids one queue-wide mutex and can read its backpressure counters without locking producers or consumers. That is a narrower claim than saying a lock-free queue is always faster: CAS retries and cache-line traffic still have costs.
|
|
453
455
|
|
|
454
456
|
### Keep-alive fast path
|
|
455
457
|
|
|
@@ -457,23 +459,23 @@ Under moderate load these look equivalent. Under high concurrency (many threads
|
|
|
457
459
|
|
|
458
460
|
**Raptor.** After a response, the app thread does `socket.wait_readable(0.001)`, waiting up to 1ms for bytes. If bytes arrive, it parses the next request inline. If the thread pool queue is at least as deep as the pool, the parsed request is handed back to the pool so other threads share the load; otherwise the same thread dispatches it inline. If no bytes arrive, `reactor.persist` and return.
|
|
459
461
|
|
|
460
|
-
The difference is subtle
|
|
462
|
+
The difference is subtle. Puma's `eagerly_finish` catches bytes already available on the socket; Raptor's `wait_readable(1ms)` catches those plus bytes arriving during the next millisecond. A request caught there parses on the response-writing thread and avoids a reactor round-trip. The cost is that an app thread can spend up to 1ms waiting on an otherwise idle connection.
|
|
461
463
|
|
|
462
|
-
This is
|
|
464
|
+
This fast path is one plausible contributor to Raptor's keep-alive result. Requests arriving inside the polling window are parsed without a reactor round-trip: the response-writing thread either continues serving that connection inline (when the pool queue is shallower than the pool) or hands the parsed request back to the pool (when it is not). Puma's app threads also continue inline when data is already available and capacity permits. The material difference is Raptor's short wait for data that has not arrived yet.
|
|
463
465
|
|
|
464
466
|
### Backpressure
|
|
465
467
|
|
|
466
468
|
**Puma.** Cluster mode uses `accept_loop_delay` (sleep proportional to busy ratio) to prevent thundering herd across workers. Single-worker backpressure is implicit; if all threads are busy and the queue is growing, new accepts pile up in the kernel accept queue. Puma does have `queue_requests` (default true) which pushes partial requests into the reactor, freeing the accept loop, but there is no explicit "stop accepting" signal from the worker.
|
|
467
469
|
|
|
468
|
-
**Raptor.** Explicit backpressure, read every iteration of the accept loop: hard skip when `backlog >= max(pool_size * 1.2, 8)`, and a softer `Thread.pass` yield when the queue alone exceeds the pool size.
|
|
470
|
+
**Raptor.** Explicit backpressure, read every iteration of the accept loop: hard skip when `backlog >= max(pool_size * 1.2, 8)`, and a softer `Thread.pass` yield when the queue alone exceeds the pool size. On supported Linux TCP bindings, the BPF reuseport program samples two workers, routes to the less loaded one, and reserves its map slot. TLS, Unix sockets, unsupported platforms, and installations without the BPF prerequisites use inherited shared listeners instead.
|
|
469
471
|
|
|
470
472
|
### Shared state (worker ↔ master)
|
|
471
473
|
|
|
472
474
|
**Puma.** Pipes. Each worker writes ping messages to a pipe read by the master. Signals push the master to check status. Simple, works everywhere, but every stat update involves a syscall on both ends.
|
|
473
475
|
|
|
474
|
-
**Raptor.** Anonymous mmap region shared across workers via `mmap-ruby`. Each worker writes a 49-byte slot for its own vitals every second. The master reads the region directly
|
|
476
|
+
**Raptor.** Anonymous mmap region shared across workers via `mmap-ruby`. Each worker writes a 49-byte slot for its own vitals every second. The master reads the region directly without a pipe exchange.
|
|
475
477
|
|
|
476
|
-
The performance difference here is negligible
|
|
478
|
+
The performance difference here is negligible because the update happens once per worker per second, outside request processing. The design mainly gives the master a fixed-size snapshot it can inspect without draining per-worker messages.
|
|
477
479
|
|
|
478
480
|
### HTTP/2
|
|
479
481
|
|
|
@@ -481,15 +483,15 @@ The performance difference here is negligible. Stats writing is one syscall per
|
|
|
481
483
|
|
|
482
484
|
**Raptor.** Native C parser plus HPACK, per-stream flow control, lock-free frame writer, stream multiplexing over a single connection. Once a request is complete it takes the same path as HTTP/1.1 and enters the same thread pool. Under HTTP/2, a single client connection can be issuing many concurrent requests, and Raptor services all of them in parallel on the same thread pool.
|
|
483
485
|
|
|
484
|
-
Whether that matters depends on your setup. If you terminate TLS at an edge proxy that already speaks HTTP/2, both servers see HTTP/1.1 and it doesn't matter which of them you pick on this axis. If you're building an all-Ruby stack with no proxy in front,
|
|
486
|
+
Whether that matters depends on your setup. If you terminate TLS at an edge proxy that already speaks HTTP/2, both servers see HTTP/1.1 and it doesn't matter which of them you pick on this axis. If you're building an all-Ruby stack with no proxy in front, serving direct HTTP/2 clients, or measuring the app server itself, HTTP/2 support is where Raptor and Puma stop being comparable.
|
|
485
487
|
|
|
486
|
-
At the throughput numbers the benchmark shows,
|
|
488
|
+
At the throughput numbers the benchmark shows, a small set of concurrent connections multiplex many streams, so responses from several app threads share each socket. The writer's CAS-based handoff keeps one active socket writer without parking the other app threads behind a per-connection mutex.
|
|
487
489
|
|
|
488
490
|
### Response writing
|
|
489
491
|
|
|
490
|
-
Both servers
|
|
492
|
+
Both servers support the same fundamental response shapes: file bodies through `IO.copy_stream`, non-blocking writes with `wait_writable(timeout)` on EAGAIN, and chunked transfer encoding for enumerable bodies without a known length. Puma uses `TCP_CORK` on Linux around HTTP/1.1 responses. Raptor corks responses that will close the connection, but skips the cork/uncork socket options on keep-alive responses where its write batching already supplies the important grouping.
|
|
491
493
|
|
|
492
|
-
On the HTTP/1.1 path, Raptor has a small `writev(2)` wrapper (`Raptor::VectorIO`) that scatter-
|
|
494
|
+
On the HTTP/1.1 path, Raptor has a small `writev(2)` wrapper (`Raptor::VectorIO`) that can scatter-write the status line, headers, and body in one call for non-chunked responses. Puma sends the same content over multiple `write` calls batched by `TCP_CORK` at the kernel; Raptor groups the buffers in userspace and lets `writev` handle partial writes when necessary.
|
|
493
495
|
|
|
494
496
|
HTTP/1.1 responses also reuse a per-thread String buffer for the status line and headers rather than allocating one per response. The buffer grows once to fit the largest response the thread has served and stays that size afterwards, so subsequent responses on that thread skip the allocation entirely.
|
|
495
497
|
|
|
@@ -499,7 +501,7 @@ Around the response boundary, HTTP/1.1 also amortises the common per-request all
|
|
|
499
501
|
|
|
500
502
|
### Keep-alive request by request
|
|
501
503
|
|
|
502
|
-
|
|
504
|
+
To make the keep-alive distinction concrete, here is one possible timing for three requests on the same connection. The third request arrives after Puma's non-blocking eager read but inside Raptor's 1ms polling window. Drawn separately so the participant columns stay wide enough to read.
|
|
503
505
|
|
|
504
506
|
**Puma, three keep-alive requests:**
|
|
505
507
|
|
|
@@ -561,37 +563,39 @@ sequenceDiagram
|
|
|
561
563
|
RP->>Client: response 3
|
|
562
564
|
```
|
|
563
565
|
|
|
564
|
-
|
|
566
|
+
In this timing, Puma returns Request 3 to the reactor while Raptor catches it on the current thread. If the bytes were already available, both could stay inline; if they arrived after 1ms, both would use the reactor. The optimisation trades up to 1ms of app-thread occupancy for a wider inline window.
|
|
565
567
|
|
|
566
568
|
## Part IV: What Raptor's design buys you
|
|
567
569
|
|
|
568
|
-
### IO-bound work, where
|
|
570
|
+
### IO-bound work, where app concurrency matters
|
|
571
|
+
|
|
572
|
+
On the IO-bound benchmark profile, each request does 5 to 10 short sleeps interleaved with small CPU work, simulating a request that makes several DB or cache calls throughout its lifetime. The bottleneck is how many requests a worker can keep in flight while they wait on IO. Fixed Raptor and Puma cap application execution at three threads per worker. Falcon spawns a fiber per connection and cooperatively yields on every sleep, so many more client connections can make progress while others wait. That advantage gives Falcon the clear lead over both fixed-thread servers, especially without keep-alive.
|
|
569
573
|
|
|
570
|
-
|
|
574
|
+
Scaling Raptor starts with the same three threads, then adds temporary threads while work is queued and the active threads are mostly blocked outside the GVL. In the current results that raises throughput from 3.11k to 8.63k req/s without keep-alive, and from 3.21k to 9.43k with it. Falcon still leads the first by 29.3%; scaling Raptor leads the second by 51.3%. The benchmark reports fixed and scaling Raptor separately rather than hiding that difference in one result.
|
|
571
575
|
|
|
572
|
-
Between the thread-based servers, Raptor holds a clear lead over Puma on both throughput and p95.
|
|
576
|
+
Between the thread-based servers, Raptor holds a clear lead over Puma on both throughput and p95. Its eager paths, explicit admission control, response batching, and app pool are all designed to reduce coordination, but the benchmark does not isolate enough variables to assign the result to one of them.
|
|
573
577
|
|
|
574
|
-
Real applications that spend most of their time waiting on a database or an upstream service look like this. If your app is IO-heavy and you're free to adopt the `async` ecosystem, Falcon
|
|
578
|
+
Real applications that spend most of their time waiting on a database or an upstream service look like this. If your app is IO-heavy and you're free to adopt the `async` ecosystem, Falcon and scaling Raptor are the interesting comparison: fibers are cheaper, while Raptor can run ordinary blocking Rack code without requiring a fiber-aware stack.
|
|
575
579
|
|
|
576
|
-
### CPU-bound HTTP/1.1, where
|
|
580
|
+
### CPU-bound HTTP/1.1, where Puma and Raptor converge
|
|
577
581
|
|
|
578
582
|
On the CPU-bound benchmark profile, each POST request accepts a small JSON body and builds a JSON response in 3 to 5 chunks totalling 450 to 1500 items, with sub-100µs sleeps between chunks. It's roughly 95% CPU by wall time, so fibers can't multiplex their way to an advantage. The CPU work happens under a single Ruby VM regardless of concurrency model.
|
|
579
583
|
|
|
580
|
-
**Without keep-alive**, every request opens a fresh TCP connection, gets parsed, dispatched, served, and closes.
|
|
584
|
+
**Without keep-alive**, every request opens a fresh TCP connection, gets parsed, dispatched, served, and closes. Puma leads fixed Raptor by 5.0% on throughput and has the lower p95; Raptor still leads Falcon. Fixed and scaling Raptor both deliver 8.22k req/s. The Puma/Raptor gap is small enough that neither architecture has overwhelmed the CPU cost of the Rack workload.
|
|
581
585
|
|
|
582
|
-
**With keep-alive**,
|
|
586
|
+
**With keep-alive**, Puma leads fixed Raptor by 0.7% on throughput, while Raptor's p95 is 0.4ms lower; Raptor leads Falcon on both. Fixed and scaling Raptor both deliver 8.49k req/s, with scaling p95 matching Puma. The result is more useful as a guardrail than a victory claim: adaptive scaling provides the IO-bound gains without hurting this CPU-bound profile.
|
|
583
587
|
|
|
584
588
|
### HTTP/2, when it matters
|
|
585
589
|
|
|
586
590
|
Puma doesn't implement HTTP/2, and most Rails production terminates HTTP/2 at nginx or a similar edge proxy before it reaches the app server. If that describes your stack, Raptor's HTTP/2 support isn't going to help you. Both servers see HTTP/1.1 from the proxy and the throughput numbers above are what actually matter. Puma's [position](https://github.com/puma/puma/issues/2697) is that this is where h2 belongs, and it's a reasonable one.
|
|
587
591
|
|
|
588
|
-
Where Raptor's HTTP/2 support does matter is the all-Ruby stack
|
|
592
|
+
Where Raptor's HTTP/2 support does matter is the all-Ruby stack: no proxy in front, TLS terminated at the app, and browsers or API clients speaking h2 directly to it. In that setup, Puma negotiates HTTP/1.1 instead, so the app-server connection does not get HTTP/2 multiplexing or HPACK header compression.
|
|
589
593
|
|
|
590
|
-
Falcon also speaks HTTP/2 natively, so it's the interesting comparison there rather than Puma.
|
|
594
|
+
Falcon also speaks HTTP/2 natively, so it's the interesting comparison there rather than Puma. On the CPU profile, scaling narrows Raptor's throughput gap from 26.9% to 14.2% and improves its p95 advantage from 12.3% to 21.3%. On IO, scaling raises Raptor from 1.14k to 4.46k req/s, but remains 30.3% behind Falcon. The h2 samples vary substantially more than the h1 samples, so these results establish broad shape rather than a precise ranking.
|
|
591
595
|
|
|
592
|
-
|
|
596
|
+
The benchmark's h2 listener uses TLS, while Raptor's BPF reuseport path only wraps plain TCP listeners. BPF dispatch therefore cannot explain the h2 variance. With 40 physical connections spread across 10 workers, each carrying three streams, placement and per-connection scheduling have coarse granularity; more instrumentation is needed before assigning the variance to a specific mechanism.
|
|
593
597
|
|
|
594
|
-
|
|
598
|
+
Raptor's HTTP/2 CPU-bound throughput remains in the same broad range as its HTTP/1.1 result while multiplexing streams onto shared sockets. The lock-free `Writer` and flow-control atoms are part of how it coordinates that work, but this benchmark does not provide a mutex-based Raptor control case from which to quantify their individual effect.
|
|
595
599
|
|
|
596
600
|
## Part V: What Raptor gives up
|
|
597
601
|
|
|
@@ -601,4 +605,12 @@ No design comes free. Two disclosures matter most.
|
|
|
601
605
|
|
|
602
606
|
**Ruby version.** Raptor requires Ruby 4.0 because it depends on `Ractor::Port` and on Ractor internals having stabilised. Puma works on 3.0 and up. If you need to support older Ruby, Puma wins by default.
|
|
603
607
|
|
|
604
|
-
A handful of smaller trade-offs are worth naming briefly.
|
|
608
|
+
A handful of smaller trade-offs are worth naming briefly. A request on Raptor's reactor pipeline pays for shareability and Ractor/collector handoffs; eager requests avoid them. Raptor's core dependencies (`ractor-pool`, `atomic-ruby`, `red-black-tree`, `mmap-ruby`, `libbpf-ruby`) are libraries I wrote specifically to make it work, which is either "purpose-built" or "narrower testing surface" depending on how you look at it. Debugging is harder because the slow path crosses Ractor and thread boundaries, so tracing it end-to-end means stitching several stack traces together. Raptor has no single-process mode, and on a single-CPU container that adds coordination without the possibility of parallel execution.
|
|
609
|
+
|
|
610
|
+
## Which server to choose
|
|
611
|
+
|
|
612
|
+
Choose Puma when production history, broad Ruby compatibility, and a mature operational ecosystem matter most. That is still the default answer for most Rails deployments.
|
|
613
|
+
|
|
614
|
+
Try Raptor when Ruby 4 is available and you want to evaluate explicit admission control, Ractor-parallel protocol paths, direct HTTP/2, and a lock-free app pool. Benchmark your own Rack application rather than extrapolating from mine. The current numbers justify the experiment; they do not replace production evidence.
|
|
615
|
+
|
|
616
|
+
Choose Falcon when the application and its dependencies are fiber-aware and the workload benefits from keeping many I/O-bound requests or long-lived connections in flight. These are architectural choices, not a universal leaderboard.
|
|
@@ -40,11 +40,12 @@ module Rackup
|
|
|
40
40
|
# @rbs () -> Hash[String, String]
|
|
41
41
|
def self.valid_options
|
|
42
42
|
{
|
|
43
|
-
"Host=HOST"
|
|
44
|
-
"Port=PORT"
|
|
45
|
-
"Workers=NUM"
|
|
46
|
-
"Threads=NUM"
|
|
47
|
-
"
|
|
43
|
+
"Host=HOST" => "Hostname to listen on (default: #{DEFAULT_OPTIONS[:Host]})",
|
|
44
|
+
"Port=PORT" => "Port to listen on (default: #{DEFAULT_OPTIONS[:Port]})",
|
|
45
|
+
"Workers=NUM" => "Number of worker processes (default: available processor count)",
|
|
46
|
+
"Threads=NUM" => "Number of threads per worker (default: 3)",
|
|
47
|
+
"MaxThreads=NUM" => "Maximum threads per worker (`unlimited` for no limit; default: fixed at Threads)",
|
|
48
|
+
"Config=PATH" => "Load additional configuration from PATH"
|
|
48
49
|
}
|
|
49
50
|
end
|
|
50
51
|
|
|
@@ -81,6 +82,7 @@ module Rackup
|
|
|
81
82
|
drain_accept_queue: config.key?(:drain_accept_queue) ? config[:drain_accept_queue] : cli_defaults[:drain_accept_queue],
|
|
82
83
|
workers: (options[:Workers] || config[:workers] || Concurrent.available_processor_count).to_i,
|
|
83
84
|
threads: (options[:Threads] || config[:threads] || cli_defaults[:threads]).to_i,
|
|
85
|
+
max_threads: ::Raptor::CLI.parse_max_threads(options[:MaxThreads] || config[:max_threads]),
|
|
84
86
|
app: app
|
|
85
87
|
}
|
|
86
88
|
result[:rackup] = config[:rackup] if config.key?(:rackup)
|
data/lib/raptor/cli.rb
CHANGED
|
@@ -34,6 +34,7 @@ module Raptor
|
|
|
34
34
|
drain_accept_queue: false,
|
|
35
35
|
workers: DEFAULT_WORKER_COUNT,
|
|
36
36
|
threads: 3,
|
|
37
|
+
max_threads: nil,
|
|
37
38
|
rackup: "config.ru",
|
|
38
39
|
chdir: nil,
|
|
39
40
|
environment: nil,
|
|
@@ -98,6 +99,21 @@ module Raptor
|
|
|
98
99
|
DEFAULT_CONFIG_PATHS.find { |path| File.exist?(File.join(root, path)) }
|
|
99
100
|
end
|
|
100
101
|
|
|
102
|
+
# Parses a maximum thread count from a CLI, config, or Rack handler value.
|
|
103
|
+
#
|
|
104
|
+
# @param value [Integer, Float, String, nil] maximum thread count
|
|
105
|
+
# @return [Integer, Float, nil] parsed thread count, `Float::INFINITY` for
|
|
106
|
+
# `"unlimited"`, or nil for a fixed-size pool
|
|
107
|
+
#
|
|
108
|
+
# @rbs (Integer | Float | String | nil value) -> (Integer | Float)?
|
|
109
|
+
def self.parse_max_threads(value)
|
|
110
|
+
return unless value
|
|
111
|
+
return Float::INFINITY if value == "unlimited" || value == Float::INFINITY
|
|
112
|
+
return value if value.is_a?(Integer)
|
|
113
|
+
|
|
114
|
+
Integer(value, 10)
|
|
115
|
+
end
|
|
116
|
+
|
|
101
117
|
# @rbs @command: Symbol
|
|
102
118
|
# @rbs @options: Hash[Symbol, untyped]
|
|
103
119
|
# @rbs @parser: OptionParser
|
|
@@ -130,6 +146,7 @@ module Raptor
|
|
|
130
146
|
@parser.parse!(argv)
|
|
131
147
|
|
|
132
148
|
@options[:rackup] = argv.first if @command == :server && argv.first
|
|
149
|
+
@options[:max_threads] = self.class.parse_max_threads(@options[:max_threads])
|
|
133
150
|
end
|
|
134
151
|
|
|
135
152
|
# Runs the requested command.
|
|
@@ -242,6 +259,10 @@ module Raptor
|
|
|
242
259
|
@options[:threads] = num
|
|
243
260
|
end
|
|
244
261
|
|
|
262
|
+
opts.on("--max-threads NUM", "Maximum application threads per worker (`unlimited` for no limit; default: fixed at --threads)") do |num|
|
|
263
|
+
@options[:max_threads] = num
|
|
264
|
+
end
|
|
265
|
+
|
|
245
266
|
opts.on("-C", "--chdir PATH", String, "Change to PATH before loading the Rack application (default: none)") do |path|
|
|
246
267
|
@options[:chdir] = path
|
|
247
268
|
end
|