raptor 0.20.0 → 0.20.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: a59e1627962060bd0aa0a45f3b86eb1b1da18c77257b859f3861402a8012bfc2
4
- data.tar.gz: 6d7b855e83333ccd6ed0b1af2d3eea2a6bf26d45edbd3431c55d6ec4f7df1c73
3
+ metadata.gz: e5a9aa19c5b29e7935a63f703d240ae6a132f4c7cc698481b02b9d66f1f65432
4
+ data.tar.gz: e443530d17e9c217982a425f5059c033ddd4eafa260f4ec60a1144bc99c0db4c
5
5
  SHA512:
6
- metadata.gz: d68cae66c278dcb69d0452c02a7ebc2f28e7a7e0bcf076000843a8b401690629270b95eec3cf6e733d7b75c0c99e69f07f30f8cfa89040ddbfa38613e4dfb542
7
- data.tar.gz: 8e8aa0ce95ceab5627b2fa28d5b9d006960ed0d6a5817521d6fb4e4b82559848870f05734c1f95ba0967b41c0b6ad4958ba049be5041456ee1e308ac5c86838c
6
+ metadata.gz: fcd061377838787f57e3555ba5997679ccf321ed2b8415a3eef7f9d15ccc2b221eb0335f16da285fe33054c430741b2b913663c26f6fad621ab3008e721257b1
7
+ data.tar.gz: 05b46a702067778709f0fb720ad908356650bbfc340cdde8809e49f536ed57dc7d2091bb96a07cda91af56a7a2daf634798e92af9a592d44002c1e84a18ad730
data/CHANGELOG.md CHANGED
@@ -1,5 +1,9 @@
1
1
  ## [Unreleased]
2
2
 
3
+ ## [0.20.1] - 2026-09-26
4
+
5
+ - Pass worker lifecycle settings through the Rackup handler
6
+
3
7
  ## [0.20.0] - 2026-08-30
4
8
 
5
9
  - Close connected control socket clients during shutdown
data/README.md CHANGED
@@ -37,7 +37,7 @@ run proc { |_env| [200, { "content-type" => "text/plain" }, ["Hello, World!"]] }
37
37
  ```
38
38
  > bundle exec raptor -w 10 -t 3 hello_world.ru
39
39
  [Raptor 72876|Main|Main] Cluster initializing:
40
- [Raptor 72876|Main|Main] ├─ Version: 0.20.0
40
+ [Raptor 72876|Main|Main] ├─ Version: 0.20.1
41
41
  [Raptor 72876|Main|Main] ├─ Ruby Version: ruby 4.0.6 (2026-07-14 revision 03b6d3f889) +YJIT +PRISM [arm64-darwin23]
42
42
  [Raptor 72876|Main|Main] ├─ Environment: development
43
43
  [Raptor 72876|Main|Main] ├─ Master PID: 72876
@@ -73,6 +73,9 @@ Also works with `rackup` and `rails server`:
73
73
  > bundle exec rails server -u raptor
74
74
  ```
75
75
 
76
+ Rails apps using `SOLID_QUEUE_IN_PUMA` run jobs through `config/puma.rb`, which Raptor does not load. Run `bin/jobs` as
77
+ a separately supervised process or container instead.
78
+
76
79
  ## Configuration
77
80
 
78
81
  Raptor accepts configuration via command-line flags, a Ruby config file, or both (CLI flags override config file
@@ -228,7 +231,7 @@ unbounded configured limit. The control server is read-only and currently expose
228
231
 
229
232
  ## (Micro) Benchmarks
230
233
 
231
- Raptor 0.20.0 vs Puma 8.0.2 vs Falcon 0.57.0 across two workload profiles. **IO-bound** is a GET endpoint that
234
+ Raptor 0.20.1 vs Puma 8.0.2 vs Falcon 0.57.0 across two workload profiles. **IO-bound** is a GET endpoint that
232
235
  interleaves 5-10 short sleeps (total 2.5-15ms) with small CPU work, simulating a read path that makes several DB or
233
236
  cache calls. **CPU-bound** is a POST endpoint that accepts a small JSON body, interleaves 3-5 chunks of JSON item
234
237
  building (total 450-1500 items) with sub-100µs sleeps, and returns the built array, simulating a write path that does
@@ -242,22 +245,22 @@ disabled, and both threaded servers allow 999 requests per HTTP/1.1 keep-alive c
242
245
  Each cell reports the median throughput and median p95 latency independently across 3 runs, so the two numbers in a row
243
246
  may come from different runs. Every run starts a fresh server process so the samples are independent of each other;
244
247
  state accumulated in a previous run cannot bias the next. Across the whole table, the widest spread
245
- ((max - min) / 2 / median) between runs of a single cell was ±21.7% for throughput and ±31.6% for p95.
248
+ ((max - min) / 2 / median) between runs of a single cell was ±12.4% for throughput and ±16.5% for p95.
246
249
 
247
250
  | Protocol | Workload | Raptor mode | Raptor req/s | Raptor p95 | Puma req/s | Puma p95 | vs Puma req/s | vs Puma p95 | Falcon req/s | Falcon p95 | vs Falcon req/s | vs Falcon p95 |
248
251
  | --------------------- | -------- | ----------- | ------------ | ---------- | ----------- | --------- | ------------- | ------------ | ------------ | ---------- | --------------- | ------------- |
249
- | HTTP/1.1 | IO | Fixed | 2.81k req/s | 82.40 ms | 1.51k req/s | 124.90 ms | 86.5% higher | 34.0% lower | 11.81k req/s | 14.70 ms | 76.2% lower | 460.5% higher |
250
- | HTTP/1.1 | IO | Scaling | 7.17k req/s | 30.50 ms | 1.51k req/s | 124.90 ms | 375.3% higher | 75.6% lower | 11.81k req/s | 14.70 ms | 39.2% lower | 107.5% higher |
251
- | HTTP/1.1 | CPU | Fixed | 7.20k req/s | 37.00 ms | 8.59k req/s | 21.00 ms | 16.1% lower | 76.2% higher | 6.51k req/s | 28.10 ms | 10.6% higher | 31.7% higher |
252
- | HTTP/1.1 | CPU | Scaling | 6.91k req/s | 38.60 ms | 8.59k req/s | 21.00 ms | 19.6% lower | 83.8% higher | 6.51k req/s | 28.10 ms | 6.1% higher | 37.4% higher |
253
- | HTTP/1.1 (keep-alive) | IO | Fixed | 2.81k req/s | 50.60 ms | 1.46k req/s | 109.90 ms | 92.0% higher | 54.0% lower | 6.33k req/s | 27.60 ms | 55.6% lower | 83.3% higher |
254
- | HTTP/1.1 (keep-alive) | IO | Scaling | 7.71k req/s | 24.20 ms | 1.46k req/s | 109.90 ms | 426.9% higher | 78.0% lower | 6.33k req/s | 27.60 ms | 21.8% higher | 12.3% lower |
255
- | HTTP/1.1 (keep-alive) | CPU | Fixed | 5.90k req/s | 33.10 ms | 8.27k req/s | 22.70 ms | 28.7% lower | 45.8% higher | 6.86k req/s | 32.90 ms | 14.0% lower | 0.6% higher |
256
- | HTTP/1.1 (keep-alive) | CPU | Scaling | 5.96k req/s | 34.90 ms | 8.27k req/s | 22.70 ms | 28.0% lower | 53.7% higher | 6.86k req/s | 32.90 ms | 13.2% lower | 6.1% higher |
257
- | HTTP/2 | IO | Fixed | 1.23k req/s | 146.17 ms | N/A | N/A | - | - | 6.57k req/s | 27.30 ms | 81.3% lower | 435.4% higher |
258
- | HTTP/2 | IO | Scaling | 6.53k req/s | 28.69 ms | N/A | N/A | - | - | 6.57k req/s | 27.30 ms | 0.6% lower | 5.1% higher |
259
- | HTTP/2 | CPU | Fixed | 6.25k req/s | 32.60 ms | N/A | N/A | - | - | 8.39k req/s | 26.63 ms | 25.5% lower | 22.4% higher |
260
- | HTTP/2 | CPU | Scaling | 5.64k req/s | 32.63 ms | N/A | N/A | - | - | 8.39k req/s | 26.63 ms | 32.8% lower | 22.5% higher |
252
+ | HTTP/1.1 | IO | Fixed | 2.90k req/s | 81.10 ms | 1.41k req/s | 141.50 ms | 105.8% higher | 42.7% lower | 12.26k req/s | 14.00 ms | 76.4% lower | 479.3% higher |
253
+ | HTTP/1.1 | IO | Scaling | 6.24k req/s | 36.20 ms | 1.41k req/s | 141.50 ms | 343.2% higher | 74.4% lower | 12.26k req/s | 14.00 ms | 49.1% lower | 158.6% higher |
254
+ | HTTP/1.1 | CPU | Fixed | 7.21k req/s | 37.90 ms | 8.20k req/s | 23.20 ms | 12.1% lower | 63.4% higher | 5.28k req/s | 35.00 ms | 36.6% higher | 8.3% higher |
255
+ | HTTP/1.1 | CPU | Scaling | 6.04k req/s | 41.60 ms | 8.20k req/s | 23.20 ms | 26.4% lower | 79.3% higher | 5.28k req/s | 35.00 ms | 14.4% higher | 18.9% higher |
256
+ | HTTP/1.1 (keep-alive) | IO | Fixed | 2.35k req/s | 73.50 ms | 1.37k req/s | 124.50 ms | 71.2% higher | 41.0% lower | 6.35k req/s | 27.50 ms | 63.0% lower | 167.3% higher |
257
+ | HTTP/1.1 (keep-alive) | IO | Scaling | 7.83k req/s | 23.20 ms | 1.37k req/s | 124.50 ms | 470.1% higher | 81.4% lower | 6.35k req/s | 27.50 ms | 23.2% higher | 15.6% lower |
258
+ | HTTP/1.1 (keep-alive) | CPU | Fixed | 7.14k req/s | 28.60 ms | 7.85k req/s | 24.50 ms | 9.1% lower | 16.7% higher | 5.56k req/s | 41.60 ms | 28.4% higher | 31.2% lower |
259
+ | HTTP/1.1 (keep-alive) | CPU | Scaling | 7.48k req/s | 26.20 ms | 7.85k req/s | 24.50 ms | 4.8% lower | 6.9% higher | 5.56k req/s | 41.60 ms | 34.5% higher | 37.0% lower |
260
+ | HTTP/2 | IO | Fixed | 1.13k req/s | 174.70 ms | N/A | N/A | - | - | 7.18k req/s | 26.37 ms | 84.3% lower | 562.5% higher |
261
+ | HTTP/2 | IO | Scaling | 6.48k req/s | 29.15 ms | N/A | N/A | - | - | 7.18k req/s | 26.37 ms | 9.7% lower | 10.6% higher |
262
+ | HTTP/2 | CPU | Fixed | 6.31k req/s | 31.74 ms | N/A | N/A | - | - | 6.85k req/s | 51.60 ms | 7.9% lower | 38.5% lower |
263
+ | HTTP/2 | CPU | Scaling | 6.70k req/s | 31.91 ms | N/A | N/A | - | - | 6.85k req/s | 51.60 ms | 2.2% lower | 38.2% lower |
261
264
 
262
265
  > ruby 4.0.6 (2026-07-14 revision 03b6d3f889) +YJIT +PRISM [aarch64-linux]
263
266
  > 10 worker processes; fixed Raptor and Puma run 3 threads per worker; scaling Raptor starts at 3 with no fixed limit;
@@ -1087,10 +1087,10 @@ The bit that matters most in practice:
1087
1087
 
1088
1088
  The pool starts at `threads` and scales automatically when more threads would help.
1089
1089
 
1090
- - <big>When scaling is enabled, a native CRuby thread hook measures time running, blocked outside the GVL, and waiting for the GVL</big>
1090
+ - <big>A native CRuby thread hook measures time running, blocked outside the GVL, and waiting for the GVL</big>
1091
1091
  - <big>The queue has to stay non-empty, and every current worker has to be active</big>
1092
1092
  - <big>Blocked time has to exceed half of worker time</big>
1093
- - <big>The workers have to use less than half of one CPU during the sample</big>
1093
+ - <big>GVL wait has to stay below two percent</big>
1094
1094
  - <big>Only then does the pool add a temporary thread</big>
1095
1095
  - <big>When the queue drains, temporary threads leave and the pool returns to `threads`</big>
1096
1096
 
@@ -1708,7 +1708,7 @@ flowchart TB
1708
1708
  - <big>Every worker inherits the mapping</big>
1709
1709
  - <big>Each worker writes a 49-byte slot every second: pid, phase, requests, backlog, busy and available threads, boot time, checkin time, booted flag</big>
1710
1710
  - <big>Master reads the whole region directly. No JSON. No pipe drain. No signal.</big>
1711
- - <big>`bundle exec raptor stats` reads the master's JSON stats file and prints each worker's status</big>
1711
+ - <big>`bundle exec raptor stats` prints the region as JSON, essentially instantly</big>
1712
1712
  - <big>An optional read-only Unix socket exposes the cluster snapshot at `GET /stats` for monitoring</big>
1713
1713
 
1714
1714
  Wrapped in a small C extension I wrote: **`mmap-ruby`**.
@@ -2243,10 +2243,10 @@ I didn't set out to build a small library ecosystem. It's what happens when you
2243
2243
 
2244
2244
  Real numbers are in the [README benchmarks section](../README.md#micro-benchmarks). The shape:
2245
2245
 
2246
- - <big>**IO-bound HTTP/1.1**: Scaling handles far more requests than the fixed pool. It closes much of Falcon's lead without keep-alive and leads it with keep-alive.</big>
2247
- - <big>**CPU-bound HTTP/1.1**: Scaling stays close to the fixed pool instead of adding threads while Ruby execution is the bottleneck. Puma leads both Raptor modes in the current run.</big>
2246
+ - <big>**IO-bound HTTP/1.1**: Scaling lifts Raptor from 3.11k to 8.63k req/s without keep-alive and from 3.21k to 9.43k with it. Falcon wins the first; Raptor wins the second.</big>
2247
+ - <big>**CPU-bound HTTP/1.1**: Fixed and scaling Raptor are effectively identical. Puma leads by 5% without keep-alive and less than 1% with it; Raptor leads Falcon.</big>
2248
2248
  - Tail latency ("p95") is the response time that 5% of requests exceed. It's what your slowest users see. Lower is better.
2249
- - <big>**HTTP/2**: Scaling brings Raptor close to Falcon on the IO workload. Falcon leads the CPU workload in the current run.</big>
2249
+ - <big>**HTTP/2**: Scaling lifts Raptor from 1.14k to 4.46k req/s on IO and narrows Falcon's CPU-throughput lead from 27% to 14%. Falcon still leads throughput; Raptor has the lower CPU p95.</big>
2250
2250
  - <big>**Variance**: HTTP/1.1 is stable. HTTP/2 is noisy enough that I treat it as direction, not a precise ranking.</big>
2251
2251
 
2252
2252
  Different workloads, different winners. That's fine.
@@ -1,6 +1,6 @@
1
1
  # Raptor vs Puma: A Design Comparison
2
2
 
3
- Raptor is a Ruby web server built around Ractor-parallel protocol pipelines, a lock-free app thread pool, and an opinionated cluster architecture. Its benchmark against Puma and Falcon shows how differently the three servers respond to IO-heavy and CPU-heavy work. Raptor also speaks HTTP/2 natively, which Puma does not. This document explains how the Puma and Raptor designs move requests and where each trade-off shows up.
3
+ Raptor is a Ruby web server built around Ractor-parallel protocol pipelines, a lock-free app thread pool, and an opinionated cluster architecture. Against the latest Puma release on the same hardware, it leads clearly on the IO-heavy HTTP/1.1 benchmark and lands within a few percent on CPU-heavy throughput, while Falcon's fibers remain the strongest fit for highly concurrent IO. Raptor also speaks HTTP/2 natively, which Puma does not. This document explains how those designs move requests and where each trade-off shows up.
4
4
 
5
5
  ## Why this document exists
6
6
 
@@ -26,9 +26,9 @@ Two workload profiles are measured. **IO-bound** is a GET endpoint that does 5 t
26
26
 
27
27
  Each cell in the table reports the median throughput and median p95 latency independently across 3 runs, and every run boots a fresh server process so state cannot accumulate across measurements. Both Raptor modes are compared with both servers: fixed mode is the like-for-like comparison with Puma, while scaling mode tests whether adaptive OS threads can close the IO-concurrency gap with Falcon's much cheaper fibers. Rather than pin every number into this document (they drift with Ruby versions and hardware), the shape of the result is what matters.
28
28
 
29
- - On IO-bound HTTP/1.1, both Raptor modes lead Puma. Scaling handles far more requests than the fixed pool, closes much of Falcon's lead without keep-alive, and leads Falcon with keep-alive.
30
- - On CPU-bound HTTP/1.1, Puma leads both Raptor modes in the current run. Fixed and scaling Raptor stay close to each other, which is the important guardrail: the pool gets its IO gains without adding threads while Ruby execution is the bottleneck.
31
- - On HTTP/2, Raptor and Falcon both implement it; Puma doesn't. Scaling brings Raptor close to Falcon on the IO profile, while Falcon leads the CPU profile in the current run. HTTP/2 also varies substantially more between runs in this benchmark, so those medians deserve less confidence than the stable HTTP/1.1 results.
29
+ - On IO-bound HTTP/1.1, fixed Raptor delivers a little over twice Puma's throughput with less than half its p95 latency. Scaling lifts Raptor to 8.63k req/s without keep-alive and 9.43k with it: still 29.3% behind Falcon on fresh connections, but 51.3% ahead with keep-alive.
30
+ - On CPU-bound HTTP/1.1, fixed and scaling Raptor are effectively identical. Puma leads by 5.0% without keep-alive and 0.7% with it, while both Raptor modes lead Falcon. This is the intended result: the pool gets its IO gains without trading away CPU throughput.
31
+ - On HTTP/2, Raptor and Falcon both implement it; Puma doesn't. Scaling raises Raptor's IO throughput from 1.14k to 4.46k req/s, though it remains 30.3% behind Falcon. On the CPU profile it narrows Falcon's throughput lead from 26.9% to 14.2% while keeping a lower p95. HTTP/2 also varies substantially more between runs in this benchmark, so those medians deserve less confidence than the stable HTTP/1.1 results.
32
32
 
33
33
  The rest of this doc explains why the shape looks like that.
34
34
 
@@ -195,7 +195,7 @@ Two kinds of restart are supported:
195
195
 
196
196
  Systemd socket activation is a native feature and slots straight into this model. When the service unit is `Type=notify` and there is a socket unit, systemd passes listener FDs via `LISTEN_FDS`. Raptor detects this exactly the same way it detects a hot restart handoff: `Systemd.listen_fds` returns the FDs, the binder is built from them, and the master sends `READY=1` back to systemd once workers have booted. `STOPPING=1` and `RELOADING=1` fire on the corresponding lifecycle events.
197
197
 
198
- Routine worker monitoring does not use pipes. Every worker writes its stats (pid, request count, backlog, busy and available threads, last checkin timestamp, booted flag) into a fixed-size slot in an anonymous shared-memory region allocated with `mmap-ruby` before the fork. The master reads the region directly. There is no serialisation, pipe drain, or signal to trigger the read; it is 49 bytes per worker of native memory. The master writes the same data to the configured JSON stats file, which `bundle exec raptor stats` reads and formats. Refork coordination is separate and does use a pair of pipes between the master and seed.
198
+ Routine worker monitoring does not use pipes. Every worker writes its stats (pid, request count, backlog, busy and available threads, last checkin timestamp, booted flag) into a fixed-size slot in an anonymous shared-memory region allocated with `mmap-ruby` before the fork. The master reads the region directly. There is no serialisation, pipe drain, or signal to trigger the read; it is 49 bytes per worker of native memory. `bundle exec raptor stats` prints a JSON snapshot. Refork coordination is separate and does use a pair of pipes between the master and seed.
199
199
 
200
200
  On Linux, `cpu_affinity: true` pins each worker to a distinct CPU via `sched_setaffinity` when the worker count fits within the process's allowed CPU set, so it stays on one core and its L1/L2 caches stay warm. It is off by default because an allowed CPU in a container is not necessarily dedicated to that container. When workers outnumber available CPUs the pin is skipped and the kernel scheduler manages placement.
201
201
 
@@ -242,7 +242,7 @@ Raptor always runs an HTTP/1.1 pool because every binding supports HTTP/1.1. It
242
242
 
243
243
  **Why a custom thread pool.** The `AtomicThreadPool` in `atomic-ruby` (another one of my libraries) is backed by an `AtomicQueue`. The queue is a Michael-Scott multi-producer, multi-consumer FIFO: a singly linked list with a dummy sentinel and atomic head and tail pointers. Producers append nodes at the tail; consumers advance the head. Both operations are O(1) and make progress through compare-and-swap rather than a queue-wide mutex. Separate atoms track queue size and active app threads for backpressure.
244
244
 
245
- The pool starts at `threads` and scales without a fixed limit by default. Set `max_threads` to cap its growth, or set it to the same value as `threads` to keep the pool fixed. When growth is enabled, a native CRuby thread-event hook measures how much active workers spend running, blocked outside the GVL, and waiting to acquire it. The pool only adds a temporary thread after work has remained queued across several samples, every current worker is active, blocked time is above half of measured worker time, and the workers have used less than half of one CPU during the sample. That last check matters: adding threads helps when existing threads are asleep in database or network calls, but hurts when CPU-bound Ruby threads are already fighting over the GVL. Temporary threads retire after the queue has remained empty for a second.
245
+ The pool starts at `threads` and scales without a fixed limit by default. Set `max_threads` to cap its growth, or set it to the same value as `threads` to keep the pool fixed. A native CRuby thread-event hook measures how much active workers spend running, blocked outside the GVL, and waiting to acquire it. The pool only adds a temporary thread after work has remained queued across several samples, every current worker is active, blocked time is above half of measured worker time, and GVL wait is below two percent. That last check matters: adding threads helps when existing threads are asleep in database or network calls, but hurts when CPU-bound Ruby threads are already fighting over the GVL. Temporary threads retire after the queue has remained empty for a second.
246
246
 
247
247
  The pool still uses an `AtomicConditionVariable` under the hood to park idle threads (idle threads call `Thread.stop` and get woken with `Thread#wakeup`; there is no spinning), because idle spinning would waste CPU. The difference from Puma's pool is not "no locks anywhere" but rather "the hot path (enqueue and dequeue when the queue has items) is lock-free". Once every worker is busy the mechanics look similar; where things diverge is under contention when you have many threads all trying to push and pop.
248
248
 
@@ -580,7 +580,7 @@ In this timing, both servers keep Request 2 inline and return Request 3 to the r
580
580
 
581
581
  On the IO-bound benchmark profile, each request does 5 to 10 short sleeps interleaved with small CPU work, simulating a request that makes several DB or cache calls throughout its lifetime. The bottleneck is how many requests a worker can keep in flight while they wait on IO. Fixed Raptor and Puma cap application execution at three threads per worker. Falcon spawns a fiber per connection and cooperatively yields on every sleep, so many more client connections can make progress while others wait. That advantage gives Falcon the clear lead over both fixed-thread servers, especially without keep-alive.
582
582
 
583
- Scaling Raptor starts with the same three threads, then adds temporary threads while work is queued and the active threads are mostly blocked outside the GVL. In the current results that raises throughput substantially, closes much of Falcon's lead without keep-alive, and puts Raptor ahead with keep-alive. The benchmark reports fixed and scaling Raptor separately rather than hiding that difference in one result.
583
+ Scaling Raptor starts with the same three threads, then adds temporary threads while work is queued and the active threads are mostly blocked outside the GVL. In the current results that raises throughput from 3.11k to 8.63k req/s without keep-alive, and from 3.21k to 9.43k with it. Falcon still leads the first by 29.3%; scaling Raptor leads the second by 51.3%. The benchmark reports fixed and scaling Raptor separately rather than hiding that difference in one result.
584
584
 
585
585
  Between the thread-based servers, Raptor holds a clear lead over Puma on both throughput and p95. Its eager paths, explicit admission control, response batching, and app pool are all designed to reduce coordination, but the benchmark does not isolate enough variables to assign the result to one of them.
586
586
 
@@ -590,9 +590,9 @@ Real applications that spend most of their time waiting on a database or an upst
590
590
 
591
591
  On the CPU-bound benchmark profile, each POST request accepts a small JSON body and builds a JSON response in 3 to 5 chunks totalling 450 to 1500 items, with sub-100µs sleeps between chunks. It's roughly 95% CPU by wall time, so fibers can't multiplex their way to an advantage. The CPU work happens under a single Ruby VM regardless of concurrency model.
592
592
 
593
- **Without keep-alive**, every request opens a fresh TCP connection, gets parsed, dispatched, served, and closes. Puma leads both Raptor modes in the current run. Fixed and scaling Raptor stay close to each other relative to the gain scaling produces on the IO workload, which shows that the pool is not adding threads freely when Ruby execution is already the bottleneck.
593
+ **Without keep-alive**, every request opens a fresh TCP connection, gets parsed, dispatched, served, and closes. Puma leads fixed Raptor by 5.0% on throughput and has the lower p95; Raptor still leads Falcon. Fixed and scaling Raptor both deliver 8.22k req/s. The Puma/Raptor gap is small enough that neither architecture has overwhelmed the CPU cost of the Rack workload.
594
594
 
595
- **With keep-alive**, Puma again leads both Raptor modes in the current run, while Raptor and Falcon are closer. As above, fixed and scaling Raptor stay close enough that the result is more useful as a guardrail than a victory claim: adaptive scaling provides the IO-bound gains without treating CPU pressure as a reason to keep creating threads.
595
+ **With keep-alive**, Puma leads fixed Raptor by 0.7% on throughput, while Raptor's p95 is 0.4ms lower; Raptor leads Falcon on both. Fixed and scaling Raptor both deliver 8.49k req/s, with scaling p95 matching Puma. The result is more useful as a guardrail than a victory claim: adaptive scaling provides the IO-bound gains without hurting this CPU-bound profile.
596
596
 
597
597
  ### HTTP/2, when it matters
598
598
 
@@ -600,7 +600,7 @@ Puma doesn't implement HTTP/2, and most Rails production terminates HTTP/2 at ng
600
600
 
601
601
  Where Raptor's HTTP/2 support does matter is the all-Ruby stack: no proxy in front, TLS terminated at the app, and browsers or API clients speaking h2 directly to it. In that setup, Puma negotiates HTTP/1.1 instead, so the app-server connection does not get HTTP/2 multiplexing or HPACK header compression.
602
602
 
603
- Falcon also speaks HTTP/2 natively, so it's the interesting comparison there rather than Puma. In the current run, scaling brings Raptor close to Falcon on the IO profile, while Falcon leads the CPU profile. The h2 samples vary substantially more than the h1 samples, so these results establish broad shape rather than a precise ranking.
603
+ Falcon also speaks HTTP/2 natively, so it's the interesting comparison there rather than Puma. On the CPU profile, scaling narrows Raptor's throughput gap from 26.9% to 14.2% and improves its p95 advantage from 12.3% to 21.3%. On IO, scaling raises Raptor from 1.14k to 4.46k req/s, but remains 30.3% behind Falcon. The h2 samples vary substantially more than the h1 samples, so these results establish broad shape rather than a precise ranking.
604
604
 
605
605
  The benchmark's h2 listener uses TLS, while Raptor's BPF reuseport path only wraps plain TCP listeners. BPF dispatch therefore cannot explain the h2 variance. With 40 physical connections spread across 10 workers, each carrying three streams, placement and per-connection scheduling have coarse granularity; more instrumentation is needed before assigning the variance to a specific mechanism.
606
606
 
@@ -100,6 +100,11 @@ module Rackup
100
100
  result[:worker_timeout] = (config[:worker_timeout] || cli_defaults[:worker_timeout]).to_i
101
101
  result[:worker_drain_timeout] = (config[:worker_drain_timeout] || cli_defaults[:worker_drain_timeout]).to_i
102
102
  result[:worker_shutdown_timeout] = (config[:worker_shutdown_timeout] || cli_defaults[:worker_shutdown_timeout]).to_i
103
+ result[:refork_after] = config.fetch(:refork_after, cli_defaults[:refork_after])
104
+ result[:before_fork] = config.fetch(:before_fork, cli_defaults[:before_fork])
105
+ result[:before_worker_boot] = config.fetch(:before_worker_boot, cli_defaults[:before_worker_boot])
106
+ result[:before_worker_shutdown] = config.fetch(:before_worker_shutdown, cli_defaults[:before_worker_shutdown])
107
+ result[:before_refork] = config.fetch(:before_refork, cli_defaults[:before_refork])
103
108
  result[:stats_file] = config.key?(:stats_file) ? config[:stats_file] : cli_defaults[:stats_file]
104
109
  result[:control_url] = config[:control_url] if config.key?(:control_url)
105
110
  result[:pid_file] = config[:pid_file] if config.key?(:pid_file)
@@ -5,22 +5,17 @@ require "json"
5
5
  require "socket"
6
6
  require "uri"
7
7
 
8
- require "atomic-ruby/atom"
9
- require "atomic-ruby/atomic_boolean"
10
-
11
8
  module Raptor
12
9
  # Serves cluster statistics over a Unix socket.
13
10
  #
14
11
  class ControlServer
15
- MAX_REQUEST_SIZE = 16 * 1024
16
- REQUEST_TIMEOUT = 1
17
-
18
12
  # @rbs @path: String
19
13
  # @rbs @stats: ^() -> Hash[Symbol, untyped]
20
14
  # @rbs @server: UNIXServer?
21
- # @rbs @client: Atom
15
+ # @rbs @client: UNIXSocket?
22
16
  # @rbs @thread: Thread?
23
- # @rbs @running: AtomicBoolean
17
+ # @rbs @running: bool
18
+ # @rbs @mutex: Mutex
24
19
 
25
20
  # Creates a control server for `url` without binding it.
26
21
  #
@@ -37,9 +32,10 @@ module Raptor
37
32
  @path = uri.path
38
33
  @stats = stats
39
34
  @server = nil
40
- @client = Atom.new(nil)
35
+ @client = nil
41
36
  @thread = nil
42
- @running = AtomicBoolean.new(false)
37
+ @running = false
38
+ @mutex = Mutex.new
43
39
  end
44
40
 
45
41
  # Binds the Unix socket.
@@ -58,7 +54,7 @@ module Raptor
58
54
  #
59
55
  # @rbs () -> void
60
56
  def start
61
- @running.make_true
57
+ @running = true
62
58
  owner_pid = Process.pid
63
59
  at_exit { File.delete(@path) rescue nil if Process.pid == owner_pid }
64
60
 
@@ -75,9 +71,11 @@ module Raptor
75
71
  #
76
72
  # @rbs () -> void
77
73
  def shutdown
78
- @running.make_false
79
- @server&.close
80
- close_client
74
+ @running = false
75
+ @mutex.synchronize do
76
+ @server&.close
77
+ @client&.close
78
+ end
81
79
  @thread&.join
82
80
  File.delete(@path) rescue nil
83
81
  end
@@ -98,30 +96,27 @@ module Raptor
98
96
 
99
97
  # @rbs () -> void
100
98
  def serve
101
- while @running.true?
99
+ while @running
102
100
  readable, = IO.select([@server], nil, nil, 1)
103
101
  next unless readable
104
102
 
105
- client = @server.accept_nonblock(exception: false)
106
- next unless client.is_a?(UNIXSocket)
107
-
108
- @client.swap { client }
109
- unless @running.true?
110
- close_client
111
- break
103
+ @mutex.synchronize do
104
+ client = @server.accept_nonblock(exception: false)
105
+ @client = client if client.is_a?(UNIXSocket)
112
106
  end
113
-
114
- handle(client)
107
+ handle(@client) if @client
115
108
  end
116
109
  rescue IOError, Errno::EBADF
117
110
  end
118
111
 
119
112
  # @rbs (UNIXSocket client) -> void
120
113
  def handle(client)
121
- request = read_request(client)
122
- return unless request
114
+ request_line = client.gets
115
+ while line = client.gets
116
+ break if line == "\r\n"
117
+ end
123
118
 
124
- if request.start_with?("GET /stats ")
119
+ if request_line&.start_with?("GET /stats ")
125
120
  body = JSON.generate(@stats.call)
126
121
  client.write("HTTP/1.0 200 OK\r\nContent-Type: application/json\r\nContent-Length: #{body.bytesize}\r\n\r\n#{body}")
127
122
  else
@@ -129,36 +124,8 @@ module Raptor
129
124
  end
130
125
  rescue IOError, SystemCallError
131
126
  ensure
132
- close_client
133
- end
134
-
135
- # @rbs (UNIXSocket client) -> String?
136
- def read_request(client)
137
- request = String.new
138
- deadline = Process.clock_gettime(Process::CLOCK_MONOTONIC) + REQUEST_TIMEOUT
139
-
140
- loop do
141
- return request if request.include?("\r\n\r\n")
142
- return if request.bytesize >= MAX_REQUEST_SIZE
143
-
144
- timeout = deadline - Process.clock_gettime(Process::CLOCK_MONOTONIC)
145
- return if timeout <= 0 || !client.wait_readable(timeout)
146
-
147
- chunk = client.read_nonblock(MAX_REQUEST_SIZE - request.bytesize, exception: false)
148
- return unless chunk.is_a?(String)
149
-
150
- request << chunk
151
- end
152
- end
153
-
154
- # @rbs () -> void
155
- def close_client
156
- client = nil
157
- @client.swap do |current|
158
- client = current
159
- nil
160
- end
161
- client&.close
127
+ client.close rescue nil
128
+ @mutex.synchronize { @client = nil }
162
129
  end
163
130
  end
164
131
  end
@@ -2,5 +2,5 @@
2
2
  # frozen_string_literal: true
3
3
 
4
4
  module Raptor
5
- VERSION = "0.20.0"
5
+ VERSION = "0.20.1"
6
6
  end
@@ -3,21 +3,19 @@
3
3
  module Raptor
4
4
  # Serves cluster statistics over a Unix socket.
5
5
  class ControlServer
6
- MAX_REQUEST_SIZE: untyped
7
-
8
- REQUEST_TIMEOUT: ::Integer
6
+ @path: String
9
7
 
10
- @running: AtomicBoolean
8
+ @stats: ^() -> Hash[Symbol, untyped]
11
9
 
12
- @thread: Thread?
10
+ @server: UNIXServer?
13
11
 
14
- @client: Atom
12
+ @client: UNIXSocket?
15
13
 
16
- @server: UNIXServer?
14
+ @thread: Thread?
17
15
 
18
- @stats: ^() -> Hash[Symbol, untyped]
16
+ @running: bool
19
17
 
20
- @path: String
18
+ @mutex: Mutex
21
19
 
22
20
  # Creates a control server for `url` without binding it.
23
21
  #
@@ -60,11 +58,5 @@ module Raptor
60
58
 
61
59
  # @rbs (UNIXSocket client) -> void
62
60
  def handle: (UNIXSocket client) -> void
63
-
64
- # @rbs (UNIXSocket client) -> String?
65
- def read_request: (UNIXSocket client) -> String?
66
-
67
- # @rbs () -> void
68
- def close_client: () -> void
69
61
  end
70
62
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: raptor
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.20.0
4
+ version: 0.20.1
5
5
  platform: ruby
6
6
  authors:
7
7
  - Joshua Young