raptor 0.21.0 → 0.22.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 9fa644eb1b4bc0f253823d74dcd806125176ef6940410f6ed332a7ae446be2e6
4
- data.tar.gz: a100f4830e44b1ed3cce30e2549ddfe95c4bc648d4d0e1b0f4ffa9aee9906269
3
+ metadata.gz: bfbcd63c857341a9ad73f6d4d40d03dac757deb1074dde954dc966b9b268822b
4
+ data.tar.gz: 2d2d78611b34df0061f231c18d0c3c0e4bc13f8314b1e61bcfcb9ef7a215297d
5
5
  SHA512:
6
- metadata.gz: 79ac90d569223b7826731aae20441dcec3362cd92b8e03eca3ee243492c0f41d20edab8dc5ddb06a3099b642bb6774b1bd013397d3c33798be9057c57c4f53e4
7
- data.tar.gz: e007cf30032b2a0053c8f7b15fca20a376a7dcd85c5cb3ec57591dd48d84daa69f51d4e294a60cfa6ac8c9df1f15b39d1f30a33ffa6709071efa72dfde4a6a33
6
+ metadata.gz: 6898ded5b172745d15dd411bc6638686cccb07485ca03cf5b6857194e4b8d263e9b02da8fd24f287dd1a60d43a128f7cd2e5524a1e95122233023c58b18fb086
7
+ data.tar.gz: 18081bd442db06fc3d60b2686eda2bfb0b36118f6709b23ed88ac8ede2e0a03a39b4ac4bc3c8d95c7d3b50fff52ec071bc414e9041ba7ae9d747fbb450bc0cf6
data/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  ## [Unreleased]
2
2
 
3
+ ## [0.22.1] - 2026-10-03
4
+
5
+ - Stop waiting for more HTTP/2 frames on the collector thread
6
+ - Run `DetachedBody#on_open` callbacks on the request thread
7
+ - Attach detached response bodies without polling
8
+
9
+ ## [0.22.0] - 2026-10-03
10
+
11
+ - Add detached HTTP/1.1 response bodies
12
+ - Add detached HTTP/2 response bodies
13
+ - Move HTTP/2 writes to the reactor
14
+
3
15
  ## [0.21.0] - 2026-09-27
4
16
 
5
17
  - Add HTTP/2 keepalive
data/README.md CHANGED
@@ -37,7 +37,7 @@ run proc { |_env| [200, { "content-type" => "text/plain" }, ["Hello, World!"]] }
37
37
  ```
38
38
  > bundle exec raptor -w 10 -t 3 hello_world.ru
39
39
  [Raptor 72876|Main|Main] Cluster initializing:
40
- [Raptor 72876|Main|Main] ├─ Version: 0.21.0
40
+ [Raptor 72876|Main|Main] ├─ Version: 0.22.1
41
41
  [Raptor 72876|Main|Main] ├─ Ruby Version: ruby 4.0.6 (2026-07-14 revision 03b6d3f889) +YJIT +PRISM [arm64-darwin23]
42
42
  [Raptor 72876|Main|Main] ├─ Environment: development
43
43
  [Raptor 72876|Main|Main] ├─ Master PID: 72876
@@ -78,21 +78,11 @@ a separately supervised process or container instead.
78
78
 
79
79
  ## Configuration
80
80
 
81
- Raptor accepts configuration via command-line flags, a Ruby config file, or both (CLI flags override config file
82
- values). Run `bundle exec raptor --help` for the full flag list.
81
+ Raptor reads options from a Ruby config file, environment variables, and command line flags, in increasing order of
82
+ precedence. Run `bundle exec raptor --help` for the full flag list.
83
83
 
84
- The config file is a Ruby file that evaluates to a hash of options. By default Raptor loads `raptor.rb` then
85
- `config/raptor.rb` from the working directory; pass `-c PATH` to point at a specific file. Settings are nested under
86
- `connection:` (shared across protocols), `http1:` (HTTP/1.1-specific), and `http2:` (HTTP/2-specific).
87
-
88
- Use an `ssl://` bind for HTTP/2 negotiated with ALPN, or an `h2c://` bind for cleartext HTTP/2 clients using prior
89
- knowledge. Each `h2c://` listener accepts HTTP/2 only.
90
-
91
- HTTP/2 applications can populate `env["raptor.response_trailers"]` with trailing response headers. Values may be
92
- strings or arrays of strings.
93
-
94
- Idle HTTP/2 connections receive a PING after `keepalive_interval` seconds and close when its acknowledgement does not
95
- arrive within `keepalive_timeout`. Set `keepalive_interval` to `0` to disable these probes.
84
+ The config file evaluates to a hash of options. By default Raptor loads `raptor.rb` then `config/raptor.rb` from the
85
+ working directory. Pass `-c PATH` to point at a specific file.
96
86
 
97
87
  ```ruby
98
88
  # raptor.rb
@@ -146,35 +136,20 @@ arrive within `keepalive_timeout`. Set `keepalive_interval` to `0` to disable th
146
136
  }
147
137
  ```
148
138
 
149
- `threads` sets the number of application threads each worker keeps running. By default, Raptor adds temporary threads
150
- without a fixed limit when queued work is held up by blocking operations. It does not add threads when waiting for the
151
- GVL is the bottleneck, and temporary threads leave after the queue drains. Set `max_threads` to cap growth, or set it
152
- to the same value as `threads` for a fixed pool.
153
-
154
- Set `cpu_affinity` to `true` to pin each worker to a distinct CPU when the worker count fits within the process's
155
- allowed CPU set. It is off by default because container runtimes commonly expose CPUs that are shared with other
156
- containers.
139
+ `RAPTOR_WORKERS`, `RAPTOR_THREADS`, and `RAPTOR_MAX_THREADS` set the corresponding options without a config file.
157
140
 
158
- Raptor clears application thread locals after each request by default. Set `clean_thread_locals` to `false` to disable
159
- it. Set `clean_fiber_locals` to `true` to run each request in a fresh Fiber, isolating Fiber-local state as well.
160
-
161
- `RAPTOR_WORKERS`, `RAPTOR_THREADS`, and `RAPTOR_MAX_THREADS` can set the corresponding options without a config file.
162
- Config files override defaults, environment variables override config files, and command-line options override both.
163
- `RAPTOR_MAX_THREADS=unlimited` leaves adaptive growth uncapped.
164
-
165
- `before_worker_boot` and `before_worker_shutdown` hooks receive the worker index.
141
+ By default each worker adds threads beyond `threads` only while queued requests are waiting on blocking work such as
142
+ database or network calls. It stops adding them once Ruby execution becomes the bottleneck, and the extra threads
143
+ retire when the queue is idle. Set `max_threads` to cap this growth, or set it to `threads` for a fixed pool.
166
144
 
167
145
  ## Bindings
168
146
 
169
- Raptor accepts multiple `binds:` URIs across three schemes.
147
+ `binds` accepts any combination of these URIs.
170
148
 
171
- - `tcp://host:port` for TCP. Host can be a specific IP, `0.0.0.0` / `[::]`, or `localhost` (expanded to both IPv4 and
172
- IPv6 loopback addresses).
173
- - `unix:///path/to/socket` for a Unix domain socket. Stale sockets left by crashed processes are cleaned up
174
- automatically.
175
- - `ssl://host:port?cert=/path/to.crt&key=/path/to.key` for TLS. HTTP/1.1 and HTTP/2 are negotiated via ALPN.
176
-
177
- Multiple binds can be combined freely.
149
+ - `tcp://host:port` for TCP. `localhost` binds both IPv4 and IPv6 loopback addresses.
150
+ - `unix:///path/to/socket` for a Unix domain socket.
151
+ - `ssl://host:port?cert=/path/to.crt&key=/path/to.key` for TLS, negotiating HTTP/1.1 or HTTP/2 via ALPN.
152
+ - `h2c://host:port` for cleartext HTTP/2.
178
153
 
179
154
  ## Signals
180
155
 
@@ -186,22 +161,16 @@ Send to the master process.
186
161
  | `TERM` | Graceful shutdown |
187
162
  | `HUP` | Reopen `stdout_file`, `stderr_file`, and `access_log_file` |
188
163
  | `USR1` | Phased restart (rolling worker replacement) |
189
- | `USR2` | Hot restart (re-exec master, inheriting listening sockets) |
190
-
191
- ## Restarts
164
+ | `USR2` | Hot restart (restart master, keeping listening sockets) |
192
165
 
193
- - **Phased restart** (`USR1`) replaces workers one at a time, waiting for each new worker to boot before retiring the
194
- previous one. The master process keeps running, so existing workers continue serving until they are individually
195
- replaced. Use to pick up code changes that don't affect the master's boot path.
196
- - **Hot restart** (`USR2`) re-execs the master process with its original command line, inheriting the listening sockets
197
- so accepted connections continue to be served across the swap. The successor master re-runs initialization from
198
- scratch. Use to pick up changes that affect master-level state (config layout, dependency upgrades, Raptor itself).
166
+ A phased restart replaces workers one at a time while the master keeps running, picking up application code changes.
167
+ A hot restart starts a new master with the original command line and the same listening sockets, picking up changes
168
+ to configuration, dependencies, or Raptor itself.
199
169
 
200
170
  ## systemd
201
171
 
202
- Raptor implements socket activation (`LISTEN_FDS`) and `sd_notify`, so it integrates cleanly with `Type=notify` units.
203
- When the socket unit is active, systemd hands the pre-bound listening file descriptors to Raptor, which serves them in
204
- place of `binds:`. `READY=1`, `STOPPING=1`, and `RELOADING=1` lifecycle messages are emitted automatically.
172
+ Raptor supports socket activation and `sd_notify`, so it works with `Type=notify` units. When the socket unit is
173
+ active, Raptor serves the listening sockets systemd passes in place of `binds`.
205
174
 
206
175
  ```ini
207
176
  # /etc/systemd/system/myapp.socket
@@ -224,8 +193,7 @@ KillMode=mixed
224
193
 
225
194
  ## Stats
226
195
 
227
- Each worker writes per-worker stats (request count, busy and available threads, backlog, last check-in) to shared
228
- memory and to a JSON file (default `tmp/raptor.json`; set via `stats_file`).
196
+ Workers publish their stats to `stats_file`, which defaults to `tmp/raptor.json`.
229
197
 
230
198
  ```
231
199
  > bundle exec raptor stats
@@ -235,14 +203,12 @@ Worker 1 (phase 0): pid=91351, requests=1199, busy=1/3, backlog=0, booted, last_
235
203
  ...
236
204
  ```
237
205
 
238
- Set `control_url` to a Unix socket URL such as `unix:///tmp/raptor-control.sock` to expose cluster stats over `/stats`.
239
- For adaptive pools, `max_threads` in each worker's status is its current thread count, so
240
- `pool_capacity / max_threads` measures the capacity available at that moment rather than comparing against an
241
- unbounded configured limit. The control server is read-only and currently exposes only `/stats`.
206
+ Set `control_url` to a Unix socket URL such as `unix:///tmp/raptor-control.sock` to serve the same stats over
207
+ `GET /stats`.
242
208
 
243
209
  ## (Micro) Benchmarks
244
210
 
245
- Raptor 0.21.0 vs Puma 8.0.2 vs Falcon 0.57.0 across two workload profiles. **IO-bound** is a GET endpoint that
211
+ Raptor 0.22.1 vs Puma 8.0.2 vs Falcon 0.57.0 across two workload profiles. **IO-bound** is a GET endpoint that
246
212
  interleaves 5-10 short sleeps (total 2.5-15ms) with small CPU work, simulating a read path that makes several DB or
247
213
  cache calls. **CPU-bound** is a POST endpoint that accepts a small JSON body, interleaves 3-5 chunks of JSON item
248
214
  building (total 450-1500 items) with sub-100µs sleeps, and returns the built array, simulating a write path that does
@@ -256,22 +222,22 @@ disabled, and both threaded servers allow 999 requests per HTTP/1.1 keep-alive c
256
222
  Each cell reports the median throughput and median p95 latency independently across 3 runs, so the two numbers in a row
257
223
  may come from different runs. Every run starts a fresh server process so the samples are independent of each other;
258
224
  state accumulated in a previous run cannot bias the next. Across the whole table, the widest spread
259
- ((max - min) / 2 / median) between runs of a single cell was ±22.1% for throughput and ±26.6% for p95.
225
+ ((max - min) / 2 / median) between runs of a single cell was ±19.6% for throughput and ±13.6% for p95.
260
226
 
261
227
  | Protocol | Workload | Raptor mode | Raptor req/s | Raptor p95 | Puma req/s | Puma p95 | vs Puma req/s | vs Puma p95 | Falcon req/s | Falcon p95 | vs Falcon req/s | vs Falcon p95 |
262
228
  | --------------------- | -------- | ----------- | ------------ | ---------- | ----------- | --------- | ------------- | ------------ | ------------ | ---------- | --------------- | ------------- |
263
- | HTTP/1.1 | IO | Fixed | 2.83k req/s | 84.00 ms | 1.51k req/s | 126.00 ms | 87.5% higher | 33.3% lower | 12.19k req/s | 14.10 ms | 76.8% lower | 495.7% higher |
264
- | HTTP/1.1 | IO | Scaling | 7.12k req/s | 31.00 ms | 1.51k req/s | 126.00 ms | 371.2% higher | 75.4% lower | 12.19k req/s | 14.10 ms | 41.6% lower | 119.9% higher |
265
- | HTTP/1.1 | CPU | Fixed | 7.42k req/s | 37.30 ms | 8.63k req/s | 21.30 ms | 14.0% lower | 75.1% higher | 6.55k req/s | 28.10 ms | 13.2% higher | 32.7% higher |
266
- | HTTP/1.1 | CPU | Scaling | 6.66k req/s | 37.70 ms | 8.63k req/s | 21.30 ms | 22.8% lower | 77.0% higher | 6.55k req/s | 28.10 ms | 1.6% higher | 34.2% higher |
267
- | HTTP/1.1 (keep-alive) | IO | Fixed | 2.43k req/s | 65.40 ms | 1.47k req/s | 106.30 ms | 65.2% higher | 38.5% lower | 6.22k req/s | 28.10 ms | 61.0% lower | 132.7% higher |
268
- | HTTP/1.1 (keep-alive) | IO | Scaling | 7.77k req/s | 23.40 ms | 1.47k req/s | 106.30 ms | 428.4% higher | 78.0% lower | 6.22k req/s | 28.10 ms | 24.9% higher | 16.7% lower |
269
- | HTTP/1.1 (keep-alive) | CPU | Fixed | 7.46k req/s | 26.60 ms | 8.46k req/s | 22.80 ms | 11.8% lower | 16.7% higher | 6.92k req/s | 33.80 ms | 7.9% higher | 21.3% lower |
270
- | HTTP/1.1 (keep-alive) | CPU | Scaling | 7.53k req/s | 26.50 ms | 8.46k req/s | 22.80 ms | 11.0% lower | 16.2% higher | 6.92k req/s | 33.80 ms | 8.9% higher | 21.6% lower |
271
- | HTTP/2 | IO | Fixed | 1.25k req/s | 138.74 ms | N/A | N/A | - | - | 6.17k req/s | 28.61 ms | 79.7% lower | 384.9% higher |
272
- | HTTP/2 | IO | Scaling | 6.20k req/s | 30.00 ms | N/A | N/A | - | - | 6.17k req/s | 28.61 ms | 0.5% higher | 4.9% higher |
273
- | HTTP/2 | CPU | Fixed | 6.84k req/s | 32.02 ms | N/A | N/A | - | - | 7.90k req/s | 56.51 ms | 13.4% lower | 43.3% lower |
274
- | HTTP/2 | CPU | Scaling | 7.12k req/s | 31.07 ms | N/A | N/A | - | - | 7.90k req/s | 56.51 ms | 9.9% lower | 45.0% lower |
229
+ | HTTP/1.1 | IO | Fixed | 2.95k req/s | 80.20 ms | 1.55k req/s | 122.60 ms | 90.1% higher | 34.6% lower | 12.24k req/s | 14.00 ms | 75.9% lower | 472.9% higher |
230
+ | HTTP/1.1 | IO | Scaling | 7.05k req/s | 30.40 ms | 1.55k req/s | 122.60 ms | 354.1% higher | 75.2% lower | 12.24k req/s | 14.00 ms | 42.4% lower | 117.1% higher |
231
+ | HTTP/1.1 | CPU | Fixed | 7.29k req/s | 35.00 ms | 8.73k req/s | 20.30 ms | 16.4% lower | 72.4% higher | 6.71k req/s | 26.80 ms | 8.7% higher | 30.6% higher |
232
+ | HTTP/1.1 | CPU | Scaling | 6.73k req/s | 36.90 ms | 8.73k req/s | 20.30 ms | 22.8% lower | 81.8% higher | 6.71k req/s | 26.80 ms | 0.4% higher | 37.7% higher |
233
+ | HTTP/1.1 (keep-alive) | IO | Fixed | 2.52k req/s | 70.70 ms | 1.50k req/s | 102.60 ms | 67.8% higher | 31.1% lower | 6.23k req/s | 28.20 ms | 59.5% lower | 150.7% higher |
234
+ | HTTP/1.1 (keep-alive) | IO | Scaling | 8.19k req/s | 21.90 ms | 1.50k req/s | 102.60 ms | 445.8% higher | 78.7% lower | 6.23k req/s | 28.20 ms | 31.6% higher | 22.3% lower |
235
+ | HTTP/1.1 (keep-alive) | CPU | Fixed | 7.04k req/s | 28.30 ms | 8.62k req/s | 21.50 ms | 18.3% lower | 31.6% higher | 7.07k req/s | 32.10 ms | 0.4% lower | 11.8% lower |
236
+ | HTTP/1.1 (keep-alive) | CPU | Scaling | 7.78k req/s | 25.70 ms | 8.62k req/s | 21.50 ms | 9.8% lower | 19.5% higher | 7.07k req/s | 32.10 ms | 9.9% higher | 19.9% lower |
237
+ | HTTP/2 | IO | Fixed | 1.50k req/s | 128.92 ms | N/A | N/A | - | - | 6.57k req/s | 27.29 ms | 77.2% lower | 372.4% higher |
238
+ | HTTP/2 | IO | Scaling | 9.41k req/s | 19.30 ms | N/A | N/A | - | - | 6.57k req/s | 27.29 ms | 43.1% higher | 29.3% lower |
239
+ | HTTP/2 | CPU | Fixed | 7.75k req/s | 26.55 ms | N/A | N/A | - | - | 7.24k req/s | 49.42 ms | 7.1% higher | 46.3% lower |
240
+ | HTTP/2 | CPU | Scaling | 7.54k req/s | 26.97 ms | N/A | N/A | - | - | 7.24k req/s | 49.42 ms | 4.2% higher | 45.4% lower |
275
241
 
276
242
  > ruby 4.0.7 (2026-09-15 revision 229531a6cf) +YJIT +PRISM [aarch64-linux]
277
243
  > 10 worker processes; fixed Raptor and Puma run 3 threads per worker; scaling Raptor starts at 3 with no fixed limit;
@@ -43,7 +43,7 @@ The rest of this doc explains why the shape looks like that.
43
43
  | I/O multiplexing | `nio4r` reactor for keep-alive idle and slow reads | `nio4r` reactor for the same, plus a red-black tree for O(log n) timeouts |
44
44
  | Cluster dispatch | Workers race on inherited listeners with a load-proportional accept delay | Two-choice load-aware BPF dispatch for TCP on Linux; shared-listener fallback |
45
45
  | Work queue | Ruby `Queue` coordinated under the pool mutex | Lock-free Michael-Scott FIFO queue |
46
- | HTTP/2 | Not implemented | Native C parser + HPACK, lock-free per-connection frame writer |
46
+ | HTTP/2 | Not implemented | Native C parser + HPACK, reactor-owned frame scheduler |
47
47
  | Keep-alive fast path | Same-thread inline dispatch when spare threads exist | Same-thread inline dispatch for bytes that are already waiting |
48
48
  | Native extensions | 1 (Ragel HTTP/1 parser + MiniSSL) | 3, all Ractor-safe (Ragel HTTP/1 parser; HTTP/2 parser + HPACK; `writev`, `sched_setaffinity`, `prctl` wrappers) |
49
49
  | Shared state (worker↔master) | Pipes and signals | Anonymous shared-memory `mmap` region |
@@ -326,6 +326,8 @@ Puma has a similar shape. It checks buffered back-to-back requests, then eagerly
326
326
 
327
327
  The `reactor.persist` call re-registers the socket with the reactor using `persistent_data_timeout` (65s) as the new deadline. When the next bytes arrive, the reactor treats the socket like any other partially-read connection.
328
328
 
329
+ A `Raptor::DetachedBody` returns the application thread while keeping the response open, chunked on HTTP/1.1 and ended by closing the connection on HTTP/1.0. The reactor owns the response until it closes, buffers later writes within fixed limits, and pauses the connection's next request without occupying an application thread.
330
+
329
331
  ### HTTP/2 request lifecycle
330
332
 
331
333
  Raptor speaks HTTP/2 on TLS connections where the client negotiates it via ALPN and on `h2c://` listeners for cleartext clients using prior knowledge. The binder sets `alpn_protocols = ["h2", "http/1.1"]` on the SSL context and the ALPN callback picks h2 whenever the client offers it. Puma does not do this. Puma's SSL context does not advertise `h2` in ALPN, so clients transparently fall back to HTTP/1.1.
@@ -337,26 +339,19 @@ From there the shape is similar to HTTP/1.1:
337
339
  1. Reactor reads frames.
338
340
  2. The HTTP/2 parser (native C, with an HPACK decoder using a static Huffman table) parses the frames in the HTTP/2 Ractor pool.
339
341
  3. Completed requests (once `HEADERS` and `DATA` are complete for a stream) go to the thread pool as separate work items. **A single connection can be servicing many streams in parallel across the thread pool.**
340
- 4. Each stream's response is written back through the connection's `Writer`, which serialises frame writes across threads without a mutex.
342
+ 4. Each stream's response frames are queued for the reactor, which owns every socket write after connection setup.
341
343
 
342
344
  Responses to `HEAD` requests and statuses that prohibit a message body end with the response `HEADERS` frame.
343
345
  Early hints and response-finished callbacks follow the same Rack lifecycle as HTTP/1.1.
344
346
  Trailing request `HEADERS` complete an open request. Rack has no standard request-trailer key, so Raptor validates them without adding them to the environment.
345
347
 
346
- The `Writer` is worth a paragraph. Naive per-connection writing would need a mutex around every socket write. Contention grows with concurrent streams. Raptor's `Writer` stores the "pending frames" queue in an `Atom` whose value is either `:idle` (nobody is writing) or an array of frames waiting to go out. A thread that wants to write does a CAS:
347
-
348
- - If current value is `:idle`, the thread claims the writer by CAS-ing to its own array of frames, then loops draining any additional frames other threads have appended.
349
- - If current value is an array (someone is already writing), the thread CAS-appends its frames and returns immediately; the current writer will pick them up and flush them.
350
-
351
- So under contention, only one thread does socket I/O at a time (because a socket can only be written to serially anyway), but no thread ever blocks on a lock. The "loser" of the CAS hands its frames off to the "winner" and returns immediately to whatever it was doing next, whether that is starting another stream, waiting for the next work item, or servicing a different connection.
352
-
353
- Once the writer thread has claimed a batch of pending frames, it concatenates them into a single buffer and issues one socket write for the whole batch. Frames handed off concurrently can share that write, while sequential body chunks reach the socket as the Rack body yields or writes them.
348
+ The `Writer` hands serialized frames to the reactor, which writes as the socket becomes ready and closes clients that stop reading. Application threads never wait for socket writability, and the connection has a single I/O owner without a per-connection mutex.
354
349
 
355
- Flow control uses similar CAS-protected atoms. The connection-level window and the per-stream windows live in separate `Atom` cells. `acquire` atomically reserves connection capacity and, where per-stream tracking is needed, deducts the same grant from that stream's window. If either window is exhausted, the caller parks on an `AtomicConditionVariable`; a `WINDOW_UPDATE`, stream reset, or connection shutdown wakes it.
350
+ Flow control uses similar CAS-protected atoms. Ordinary Rack bodies wait for connection and stream capacity as they yield. A `Raptor::DetachedBody` instead returns its application thread immediately; the reactor schedules its bounded buffer as capacity becomes available and closes it when the client cancels. This keeps long-lived streams from consuming one application thread each.
356
351
 
357
- Frame processing also has an eager loop. After processing one batch of frames, the h2 handler tries to `read_nonblock` one more time to see if the next batch is already available. Up to eight rounds are consumed inline before handing back to the reactor, and the loop bails out early once the app thread pool has more queued work than worker slots so one busy connection cannot starve the collector. This is the same principle as the HTTP/1.1 eager keep-alive: amortise the reactor round-trip when the client is actively sending, but back off under saturation.
352
+ Frame processing also has an eager loop. After processing one batch of frames, the h2 handler checks `wait_readable(0)` and reads the next batch only if it has already arrived. Up to eight rounds are consumed inline before handing back to the reactor, and the loop bails out early once the app thread pool has more queued work than worker slots so one busy connection cannot starve the collector. This is the same principle as the HTTP/1.1 eager keep-alive: amortise the reactor round-trip when the client is actively sending, but back off under saturation.
358
353
 
359
- During worker shutdown, Raptor stops accepting connections and sends GOAWAY with the last stream handed to the Rack application. Later streams are refused while the application pool drains. The reactor remains active during that period so in-flight responses can receive flow-control updates and finish before their connections close.
354
+ During worker shutdown, Raptor stops accepting connections and sends GOAWAY with the last stream handed to the Rack application. Later streams are refused while application work and detached bodies drain. The reactor remains active so in-flight responses can receive flow-control updates; detached bodies still open when the drain period expires are cancelled before their connections close.
360
355
 
361
356
  ### Raptor request flow diagram
362
357
 
@@ -408,7 +403,8 @@ flowchart TB
408
403
  COL --> CHK
409
404
  CHK -->|"no, more bytes needed"| RCT
410
405
  CHK -->|"yes, push proc"| ATP
411
- ATP -->|"app.call + write"| KA
406
+ ATP -->|"HTTP/2 response frames"| RCT
407
+ ATP -->|"HTTP/1.1 response write"| KA
412
408
  KA -->|"no, close"| CLS["close socket"]
413
409
  KA -->|"yes"| EAG
414
410
  EAG -.->|"bytes ready, parse+dispatch on same thread"| ATP
@@ -495,11 +491,11 @@ For external monitoring, `control_url` can expose a read-only `GET /stats` endpo
495
491
 
496
492
  **Puma.** Not implemented. Puma's [position](https://github.com/puma/puma/issues/2697) is that HTTP/2 belongs at the edge (nginx, Caddy, ALB), which terminates it and speaks HTTP/1.1 to the app server. That's a reasonable call for the deployments Puma is aimed at, and it's where most Rails production actually sits.
497
493
 
498
- **Raptor.** Native C parser plus HPACK, per-stream flow control, lock-free frame writer, stream multiplexing over a single connection, configurable PING keepalive, and response trailers exposed through `env["raptor.response_trailers"]`. Once a request is complete it takes the same path as HTTP/1.1 and enters the same thread pool. Under HTTP/2, a single client connection can be issuing many concurrent requests, and Raptor services all of them in parallel on the same thread pool.
494
+ **Raptor.** Native C parser plus HPACK, per-stream flow control, reactor-owned response writes, detached long-lived responses, stream multiplexing over a single connection, configurable PING keepalive, and response trailers exposed through `env["raptor.response_trailers"]`. Once a request is complete it takes the same path as HTTP/1.1 and enters the same thread pool. Under HTTP/2, a single client connection can be issuing many concurrent requests, and Raptor services all of them in parallel on the same thread pool.
499
495
 
500
496
  Whether that matters depends on your setup. If you terminate TLS at an edge proxy that already speaks HTTP/2, both servers see HTTP/1.1 and it doesn't matter which of them you pick on this axis. If you're building an all-Ruby stack with no proxy in front, serving direct HTTP/2 clients, or measuring the app server itself, HTTP/2 support is where Raptor and Puma stop being comparable.
501
497
 
502
- At the throughput numbers the benchmark shows, a small set of concurrent connections multiplex many streams, so responses from several app threads share each socket. The writer's CAS-based handoff keeps one active socket writer without parking the other app threads behind a per-connection mutex.
498
+ At the throughput numbers the benchmark shows, a small set of concurrent connections multiplex many streams, so responses from several app threads share each socket. Those threads queue frames while the reactor owns non-blocking writes for the connection.
503
499
 
504
500
  ### Response writing
505
501
 
@@ -610,7 +606,7 @@ Falcon also speaks HTTP/2 natively, so it's the interesting comparison there rat
610
606
 
611
607
  The benchmark's h2 listener uses TLS, while Raptor's BPF reuseport path only wraps plain TCP listeners. BPF dispatch therefore cannot explain the h2 variance. With 40 physical connections spread across 10 workers, each carrying three streams, placement and per-connection scheduling have coarse granularity; more instrumentation is needed before assigning the variance to a specific mechanism.
612
608
 
613
- Raptor's HTTP/2 CPU-bound throughput remains in the same broad range as its HTTP/1.1 result while multiplexing streams onto shared sockets. The lock-free `Writer` and flow-control atoms are part of how it coordinates that work, but this benchmark does not provide a mutex-based Raptor control case from which to quantify their individual effect.
609
+ Raptor's HTTP/2 CPU-bound throughput remains in the same broad range as its HTTP/1.1 result while multiplexing streams onto shared sockets. The reactor-owned frame scheduler and flow-control atoms are part of how it coordinates that work, but this benchmark does not provide an alternative Raptor control case from which to quantify their individual effect.
614
610
 
615
611
  ## Part V: What Raptor gives up
616
612
 
data/lib/raptor/cli.rb CHANGED
@@ -1,10 +1,11 @@
1
1
  # rbs_inline: enabled
2
2
  # frozen_string_literal: true
3
3
 
4
- require "concurrent/utility/processor_counter"
5
4
  require "json"
6
5
  require "optparse"
7
6
 
7
+ require "concurrent/utility/processor_counter"
8
+
8
9
  require_relative "cluster"
9
10
 
10
11
  module Raptor
@@ -46,18 +47,18 @@ module Raptor
46
47
  chunk_data_timeout: 10,
47
48
  write_timeout: 5,
48
49
  max_body_size: nil,
49
- body_spool_threshold: 1024 * 1024,
50
+ body_spool_threshold: 1024 * 1024
50
51
  },
51
52
  http1: {
52
53
  ractors: nil,
53
54
  persistent_data_timeout: 65,
54
- max_keepalive_requests: 1000,
55
+ max_keepalive_requests: 1000
55
56
  },
56
57
  http2: {
57
58
  ractors: nil,
58
59
  max_concurrent_streams: 100,
59
60
  keepalive_interval: 10,
60
- keepalive_timeout: 5,
61
+ keepalive_timeout: 5
61
62
  },
62
63
  worker_boot_timeout: 60,
63
64
  worker_timeout: 60,
@@ -1,11 +1,11 @@
1
1
  # rbs_inline: enabled
2
2
  # frozen_string_literal: true
3
3
 
4
- require "concurrent/utility/processor_counter"
5
4
  require "json"
6
5
  require "time"
7
6
 
8
7
  require "atomic-ruby/atomic_thread_pool"
8
+ require "concurrent/utility/processor_counter"
9
9
  require "rack/builder"
10
10
  require "ractor-pool"
11
11
 
@@ -339,8 +339,8 @@ module Raptor
339
339
  pool_capacity: [capacity - total_work, 0].max,
340
340
  busy_threads: active,
341
341
  max_threads: capacity,
342
- requests_count: stat.fetch(:requests, 0),
343
- },
342
+ requests_count: stat.fetch(:requests, 0)
343
+ }
344
344
  }
345
345
  end
346
346
 
@@ -350,7 +350,7 @@ module Raptor
350
350
  phase: @phase,
351
351
  booted_workers: worker_status.count { |worker| worker[:booted] },
352
352
  old_workers: worker_status.count { |worker| worker[:phase] != @phase },
353
- worker_status: worker_status,
353
+ worker_status: worker_status
354
354
  }
355
355
  end
356
356
 
@@ -946,6 +946,7 @@ module Raptor
946
946
  server_thread.join
947
947
  http1.shutdown
948
948
  http2.shutdown(reactor)
949
+ reactor.drain_detached_bodies(@worker_drain_timeout)
949
950
  drain_thread_pool(thread_pool)
950
951
  reactor.shutdown
951
952
  reactor_thread.join
@@ -5,17 +5,19 @@ require "json"
5
5
  require "socket"
6
6
  require "uri"
7
7
 
8
+ require "atomic-ruby/atom"
9
+
8
10
  module Raptor
9
11
  # Serves cluster statistics over a Unix socket.
10
12
  #
11
13
  class ControlServer
14
+ SHUTDOWN = :shutdown
15
+
12
16
  # @rbs @path: String
13
17
  # @rbs @stats: ^() -> Hash[Symbol, untyped]
14
18
  # @rbs @server: UNIXServer?
15
- # @rbs @client: UNIXSocket?
19
+ # @rbs @client: Atom
16
20
  # @rbs @thread: Thread?
17
- # @rbs @running: bool
18
- # @rbs @mutex: Mutex
19
21
 
20
22
  # Creates a control server for `url` without binding it.
21
23
  #
@@ -32,10 +34,8 @@ module Raptor
32
34
  @path = uri.path
33
35
  @stats = stats
34
36
  @server = nil
35
- @client = nil
37
+ @client = Atom.new(nil)
36
38
  @thread = nil
37
- @running = false
38
- @mutex = Mutex.new
39
39
  end
40
40
 
41
41
  # Binds the Unix socket.
@@ -54,7 +54,6 @@ module Raptor
54
54
  #
55
55
  # @rbs () -> void
56
56
  def start
57
- @running = true
58
57
  owner_pid = Process.pid
59
58
  at_exit { File.delete(@path) rescue nil if Process.pid == owner_pid }
60
59
 
@@ -71,17 +70,24 @@ module Raptor
71
70
  #
72
71
  # @rbs () -> void
73
72
  def shutdown
74
- @running = false
75
- @mutex.synchronize do
76
- @server&.close
77
- @client&.close
73
+ client = nil
74
+ @client.swap do |current|
75
+ client = current
76
+ SHUTDOWN
78
77
  end
78
+ @server&.close
79
+ client.close if client.is_a?(UNIXSocket)
79
80
  @thread&.join
80
81
  File.delete(@path) rescue nil
81
82
  end
82
83
 
83
84
  private
84
85
 
86
+ # Removes a stale socket while refusing to replace an active server.
87
+ #
88
+ # @return [void]
89
+ # @raise [RuntimeError] if another server is listening on the socket
90
+ #
85
91
  # @rbs () -> void
86
92
  def remove_stale_socket
87
93
  return unless File.exist?(@path)
@@ -94,21 +100,34 @@ module Raptor
94
100
  end
95
101
  end
96
102
 
103
+ # Accepts and handles control requests until shutdown begins.
104
+ #
105
+ # @return [void]
106
+ #
97
107
  # @rbs () -> void
98
108
  def serve
99
- while @running
109
+ until @client.value == SHUTDOWN
100
110
  readable, = IO.select([@server], nil, nil, 1)
101
111
  next unless readable
102
112
 
103
- @mutex.synchronize do
104
- client = @server.accept_nonblock(exception: false)
105
- @client = client if client.is_a?(UNIXSocket)
113
+ client = @server.accept_nonblock(exception: false)
114
+ next unless client.is_a?(UNIXSocket)
115
+
116
+ if @client.swap { |current| current == SHUTDOWN ? current : client } == SHUTDOWN
117
+ client.close
118
+ return
106
119
  end
107
- handle(@client) if @client
120
+
121
+ handle(client)
108
122
  end
109
123
  rescue IOError, Errno::EBADF
110
124
  end
111
125
 
126
+ # Writes the response for one control-socket request.
127
+ #
128
+ # @param client [UNIXSocket] connected control client
129
+ # @return [void]
130
+ #
112
131
  # @rbs (UNIXSocket client) -> void
113
132
  def handle(client)
114
133
  request_line = client.gets
@@ -125,7 +144,7 @@ module Raptor
125
144
  rescue IOError, SystemCallError
126
145
  ensure
127
146
  client.close rescue nil
128
- @mutex.synchronize { @client = nil }
147
+ @client.swap { |current| current.equal?(client) ? nil : current }
129
148
  end
130
149
  end
131
150
  end