kino 0.4.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/doc/architecture.md CHANGED
@@ -11,10 +11,13 @@ tokio (Rust threads) Ruby
11
11
  └──────────────────────────┘ └────────────────────────────┘
12
12
  ```
13
13
 
14
- All network I/O lives in Rust on a tokio multi-threaded runtime; hyper
15
- parses HTTP/1.1 and handles keep-alive; rustls terminates TLS. Ruby never
16
- touches a socket. Each request becomes a Rust-side `RequestCtx` pushed to a
17
- bounded flume MPMC queue; Ruby workers pull from it.
14
+ All network I/O lives in Rust on tokio runtimes; hyper parses HTTP/1.1 and
15
+ handles keep-alive; rustls terminates TLS. The default is Tokio's
16
+ multi-thread runtime. With `io_shards true`, one current-thread runtime
17
+ accepts connections and assigns them to current-thread I/O shards, where
18
+ each connection stays for its lifetime. Ruby never touches a socket. Each
19
+ request becomes a Rust-side `RequestCtx` pushed to a bounded flume MPMC
20
+ queue; Ruby workers pull from it.
18
21
 
19
22
  ## Topology
20
23
 
@@ -106,13 +109,63 @@ plain decrement after an `.await` would never run).
106
109
 
107
110
  ## Graceful shutdown
108
111
 
109
- `stop_accepting` → drain until queue + in-flight reach zero or the
110
- deadline passes → `close_queue` (idle workers see Disconnected and exit) →
112
+ `stop_accepting` → every live connection switches to graceful shutdown
113
+ (finish the in-flight request, answer with `Connection: close` on
114
+ HTTP/1 or GOAWAY on h2, take nothing new—so drains converge instead
115
+ of chasing chatty keep-alive clients) → drain until queue + in-flight
116
+ reach zero or the deadline passes → `close_queue` (idle workers see
117
+ Disconnected and exit) →
111
118
  join workers → past deadline: abort remaining clients (a 500, or a
112
119
  connection abort mid-stream), interrupt blocked workers, reap
113
120
  stragglers → tear down the tokio runtime. Idempotent;
114
121
  a second INT/TERM force-exits.
115
122
 
123
+ ## HTTP/2
124
+
125
+ One connection builder serves every protocol: hyper-util's auto builder
126
+ picks h2 by ALPN on TLS connections, by the 24-byte preface sniff on
127
+ plaintext (prior-knowledge h2c), and HTTP/1.x otherwise; `http2 false`
128
+ pins the HTTP/1 codec and skips the sniff entirely. Streams multiplex
129
+ into the same bounded queue as keep-alive requests, so h2 concurrency
130
+ is bounded by workers × threads exactly as h1's is, and the bounded
131
+ body channels give per-stream backpressure for free: while the
132
+ forwarder blocks, hyper withholds WINDOW_UPDATE and the client stalls.
133
+
134
+ The env bridge fills SERVER_NAME/SERVER_PORT/HTTP_HOST from the request
135
+ URI's `:authority` (h2 requests carry no Host header), through the same
136
+ host LRU the Host-header path uses, keyed by the authority bytes;
137
+ SERVER_PROTOCOL is the interned "HTTP/2". h2 trailer frames are dropped
138
+ (Rack has no trailer surface), and the h2 codec itself rejects
139
+ connection-ish headers before they can reach the env.
140
+
141
+ The upload path needed one h2-shaped fix: `read_body` drains every
142
+ already-queued chunk in a single native call (one GVL round-trip and
143
+ one Ruby string per 64 KB read), because a body arriving as 16 KB DATA
144
+ frames otherwise paid one crossing per frame—that alone took h2
145
+ uploads from half of h1's throughput to parity. A knob sweep over
146
+ hyper's h2 codec (frame size, adaptive windows, window sizes) moved
147
+ nothing after that, so hyper's defaults stay.
148
+
149
+ SETTINGS_MAX_CONCURRENT_STREAMS is derived from slot capacity
150
+ (workers × threads, clamped to [8, 1024]) rather than hyper's flat 200:
151
+ a smart balancer sees the server's true admission, and a hostile client
152
+ cannot multiply one connection into hundreds of queued requests—the
153
+ h2 analogue of h1's one-request-per-connection shape. The codec's own
154
+ abuse bounds ship as hyper/h2 defaults and were reviewed: 16 KB header
155
+ lists, 20 pending remote resets then GOAWAY (rapid reset), per-second
156
+ reset-churn and empty-frame budgets.
157
+
158
+ Low-cardinality header values (UA, accept-*, sec-ch-*, sec-fetch-*)
159
+ are interned in an LRU of frozen strings—the env-side analogue of
160
+ HPACK's wire dedup, and the practical form of it: hyper does not expose
161
+ HPACK table indices, so the cache keys on value bytes and works for
162
+ HTTP/1 too. Cookie and authorization are deliberately excluded
163
+ (per-user cardinality, secret lifetime).
164
+
165
+ Deferred until benchmarks justify it: slot-aware flow-control window
166
+ grants (a memory/abuse lever, not a throughput one—the knob sweep
167
+ showed windows don't gate upload throughput).
168
+
116
169
  ## Timer waits: `Kino.sleep`
117
170
 
118
171
  MRI's `sleep` parks the thread on the VM timer, whose wakeups inside
@@ -125,7 +178,8 @@ at the interrupt tick so `Thread#kill` and shutdown stay responsive.
125
178
 
126
179
  - **tokio + hyper**: the bottleneck is the Ruby dispatch boundary, not raw
127
180
  I/O throughput; what matters is HTTP correctness, keep-alive, TLS, and
128
- h2-later—hyper's territory. Cross-platform out of the box.
181
+ h2 (since shipped, via hyper-util's protocol-auto builder)—hyper's
182
+ territory. Cross-platform out of the box.
129
183
  - **monoio**: thread-per-core io_uring looks great in echo-server
130
184
  benchmarks, but hyper only works through its poll-io compat layer
131
185
  (forfeiting io_uring on the hot path), and the share-nothing advantage
data/doc/benchmarks.md CHANGED
@@ -17,9 +17,21 @@ the deployment most apps run today.
17
17
  9R14 (Genoa), 16 GB RAM, Amazon Linux 2023, kernel 6.18. A realistic
18
18
  app-server size, deliberately: nobody provisions a 32-core box per
19
19
  app process.
20
- - Toolchain built on the box via mise: Ruby 4.0.5 (**YJIT enabled**,
21
- `RUBY_YJIT_ENABLE=1` for every server), Rust 1.96, Kino compiled in
20
+ - Toolchain built on the box via mise: Ruby 4.0.6 (**YJIT enabled**,
21
+ `RUBY_YJIT_ENABLE=1` for every server), Rust stable, Kino compiled in
22
22
  the release profile.
23
+ - **2026-09 full re-measurement** (a fresh c7a.2xlarge): every number in
24
+ this document and the README was re-run on Ruby 4.0.6 / kernel 6.18 /
25
+ Puma 8.0.2, adding the sharded-I/O and HTTP/2 studies. Kino's numbers
26
+ reproduced within 1-3% across the board—including the /io slot
27
+ ceiling and the arena balloon, both intact. The one real shift: Puma
28
+ 8.0.2 is faster than 7.x (142k plaintext, 61k /cpu), narrowing
29
+ ractor mode's /cpu lead to +25% (was +34%). Reversed-boot-order
30
+ re-runs reproduced within ~1%. A methodology trap worth recording:
31
+ a 5-second single-endpoint warmup understates memory badly (81 MB
32
+ where the full battery shows 137 MB) and hides the arena balloon
33
+ entirely—memory is only comparable after the full endpoint battery,
34
+ and /io numbers are only comparable at equal slot counts.
23
35
  - Load generator: wrk 4.2 on the same host, 8-second windows, 64
24
36
  connections (`bench/run.sh 8 64`). Same-host load generation costs
25
37
  both sides CPU equally; we verified the generator was not the
@@ -249,14 +261,17 @@ visible.
249
261
 
250
262
  | config | RSS | PSS |
251
263
  |---|---:|---:|
252
- | Kino :ractor 8×1 (default) | 151 | **148** |
253
- | Kino lanes 8×1 | 137 | **135** |
254
- | Kino :ractor 8×3 | 171 | **169** |
264
+ | Kino :ractor 8×1 (default) | 137 | **135** |
265
+ | Kino lanes 8×1 | 128 | **126** |
266
+ | Kino :ractor 8×3 | 170 | **168** |
255
267
  | Kino :threaded 8×3 (`MALLOC_ARENA_MAX=2`) | 109 | **107** |
256
- | Kino :threaded 8×3 (no arena cap) | 668 | **666**¹ |
257
- | Puma cluster 8×3 | 1,213 | **1,068** |
268
+ | Kino :threaded 8×3 (no arena cap) | 672 | **670**¹ |
269
+ | Puma cluster 8×3 | 1,216 | **1,072** |
258
270
 
259
- The tiny app is ~7× lighter than the cluster in ractor mode, ~10× in
271
+ (2026-09 re-measure; the 2026-06 numbers reproduced within a few MB on
272
+ every row—the arena-capped threaded row to the megabyte.)
273
+
274
+ The tiny app is ~8× lighter than the cluster in ractor mode, ~10× in
260
275
  arena-capped threaded mode. RSS ≈ PSS for every Kino row (one process,
261
276
  nothing to share) and within ~12% for Puma here: a trivial app has almost
262
277
  no shared state, so Puma's footprint is ~1,051 MB of *private* per-worker
@@ -280,18 +295,19 @@ Here copy-on-write **does** matter, which is exactly why PSS is mandatory:
280
295
 
281
296
  | config | RSS | PSS |
282
297
  |---|---:|---:|
283
- | Kino :threaded (one process) | 97 | **92** |
284
- | Puma cluster 8×3 (preload) | 794 | **389** |
298
+ | Kino :threaded (one process) | 97 | **95** |
299
+ | Puma cluster 8×5 (preload) | 813 | **405** |
300
+ | Puma cluster 8×5 (no preload) | 824 | **646** |
285
301
 
286
302
  Puma serves the same Rails framework from 8 forks that share it
287
- copy-on-write; RSS counts that shared framework once per worker (794 MB),
288
- PSS counts it once (389 MB). The fair ratio is **~4×**, not the ~8× a
303
+ copy-on-write; RSS counts that shared framework once per worker (813 MB),
304
+ PSS counts it once (405 MB). The fair ratio is **~4×**, not the ~8× a
289
305
  naive RSS sum reports—this is the correction that prompted the whole
290
- re-measure. Preload barely helps (389 vs 400 MB without): Ruby's GC
291
- dirties most heap pages, breaking copy-on-write, so even a preloaded
292
- cluster keeps a large private heap per worker. That is why "CoW should
293
- make a fork cluster nearly free" is only half true—it shares the code,
294
- not the live object heap.
306
+ re-measure. The 2026-09 run also split out preload: it now saves a
307
+ third (405 vs 646 MB PSS)—worth turning on—but even preloaded, Ruby's
308
+ GC dirties heap pages and breaks copy-on-write, so each worker keeps a
309
+ large private heap. That is why "CoW should make a fork cluster nearly
310
+ free" is only half true—it shares the code, not the live object heap.
295
311
 
296
312
  ## Run-to-run variance (a.k.a. "is this a regression?")
297
313
 
@@ -376,27 +392,46 @@ crash semantics, stealing fairness, and drain behavior have spec
376
392
  coverage but not production mileage. (On loopback-bound macOS, lanes
377
393
  lose a few percent instead; see the secondary table below.)
378
394
 
395
+ ## Sharded I/O (`io_shards true`)
396
+
397
+ `bench/studies.sh 8 64 shards`, ractor 8×3, 2026-09 reference box:
398
+
399
+ | case (/plaintext) | req/s |
400
+ |---|---:|
401
+ | shared tokio runtime (baseline) | 193,461 |
402
+ | io_shards, default shard count | 195,436 |
403
+ | io_shards, `io_threads 2` | 170,509 |
404
+ | io_shards, `io_threads 4` | 196,298 |
405
+ | io_shards, `io_threads 8` | 199,657 |
406
+
407
+ On 8 cores the shards buy +1-3%, best at `io_threads 8` (the
408
+ half-the-cores default is close behind; 2 shards choke on accept
409
+ handoff). The design removes work-stealing and cross-thread wakeups,
410
+ so the win scales with scheduler contention—expect more on boxes with
411
+ more cores and connections, and measure on your own core count.
412
+
379
413
  ## Logging costs
380
414
 
381
415
  Measured at full plaintext saturation (one log line per request—rates
382
416
  that no real deployment logs at; treat these as worst-case ceilings, not
383
417
  typical costs):
384
418
 
385
- | case (8×3, same session) | req/s |
419
+ | case (8×3, same session, 2026-09) | req/s |
386
420
  |---|---:|
387
- | threaded, no logging | 219,168 |
388
- | threaded, `log_requests true` (native access log) | 193,998 (−11%) |
389
- | ractor, access log off / on | 197,596 / 181,050 (−8%) |
390
- | app logs 1 line/req via shared `::Logger` (file) | **62,961** |
391
- | app logs 1 line/req via `Kino::Logger` (file) | **149,519 (2.4×)** |
421
+ | threaded, no logging | 213,166 |
422
+ | threaded, `log_requests true` (native access log) | 173,501 (−19%) |
423
+ | ractor, access log off / on | 191,152 / 163,777 (−14%) |
424
+ | app logs 1 line/req via shared `::Logger` (file) | **90,175** |
425
+ | app logs 1 line/req via `Kino::Logger` (file) | **154,653 (1.7×)** |
392
426
 
393
427
  The shared-`::Logger` cost is the mutex: 24 worker threads serialize
394
428
  through one lock plus a write syscall per line. `Kino::Logger` hands the
395
429
  formatted line to a lock-free channel and returns—the remaining cost vs
396
430
  not logging at all is Ruby-side formatting, which no device can remove.
397
- (In the Docker environment the same comparison showed 8.5×—overlay-fs
398
- write latency punished the synchronous logger far harder than this box's
399
- NVMe does. The ranking is environment-independent; the multiple is not.)
431
+ (The multiple moves with the environment: 2.4× on the 2026-06 box,
432
+ 1.7× on the 2026-09 one, 8.5× under Docker, where overlay-fs write
433
+ latency punished the synchronous logger hardest. The ranking is
434
+ environment-independent; the multiple is not.)
400
435
 
401
436
  One trade-off worth knowing: the sink **never blocks** request threads,
402
437
  so at absurd rates against a slow disk it drops lines once its 8192-line
@@ -408,6 +443,65 @@ Puma comparison note: request logging is opt-in there too (`--quiet` is
408
443
  the default, `-v/--log-requests` enables it)—Kino's default-off
409
444
  `log_requests` matches the ecosystem's standard behavior.
410
445
 
446
+ ## HTTP/2
447
+
448
+ `bench/h2.sh`, Linux only (on macOS run it under Docker; the 2026-09
449
+ numbers below are from the c7a.2xlarge reference box). Every lane is
450
+ measured with h2load so the generator is identical everywhere: h2
451
+ lanes run 8 connections × 8 concurrent streams, h1 lanes 64
452
+ connections—the same total in-flight. Servers without native h2 get
453
+ the standard pattern instead: nginx terminating h2 and proxying
454
+ HTTP/1.1 upstream over keep-alive. One labeled run (kino ractor 8×3,
455
+ falcon `--count 8`, puma `-w 8 -t 3:3`, 5 s/lane); re-run the whole
456
+ script for close calls, per the variance section.
457
+
458
+ | target (h2 unless noted) | /plaintext | /10k | /big-cookie | /upload (64 KB) |
459
+ |---|---:|---:|---:|---:|
460
+ | kino h2c | 207,461 | 150,105 | 195,093 | 26,304 |
461
+ | kino h1 cleartext (same boot) | 116,082 | 97,634 | 110,614 | 24,338 |
462
+ | kino h2 TLS | 164,044 | 116,617 | 154,301 | 19,226 |
463
+ | kino h1 TLS (same boot) | 80,789 | 69,942 | 76,654 | 17,210 |
464
+ | falcon TLS (native h2) | 55,637 | 37,541 | 49,151 | 18,838 |
465
+ | nginx h2 → puma h1 | 81,234 | 57,891 | 51,332 | 1,247 |
466
+ | nginx h2 → kino h1 | 109,219 | 66,634 | 54,996 | 1,217 |
467
+
468
+ What the numbers say:
469
+
470
+ - **Native h2 beats h1 on the same server by +79% cleartext and +103%
471
+ over TLS** on /plaintext: the same 64 in-flight requests ride 8
472
+ connections instead of 64, so frames batch into fewer, larger
473
+ syscalls—and TLS amplifies it, since h1's 64 connections each pay
474
+ crypto per record. The `/big-cookie` lane (a ~2 KB cookie per
475
+ request) shows HPACK on top: the cookie crosses the wire once per
476
+ connection, not once per request, and holds 94% of bare-plaintext
477
+ throughput where h1 loses 5%.
478
+ - **Native h2 beats proxied h2 by +50%** with the *same backend*: the
479
+ nginx→kino-h1 lane is the proxy-cost control, and the extra hop,
480
+ re-parse, and re-serialize cost ~55k req/s on /plaintext.
481
+ - **Uploads run at h1 parity**—but only after a fix this lane
482
+ caught: h2 delivers bodies as 16 KB DATA frames, and `read_body`
483
+ originally crossed the GVL once per chunk, halving upload
484
+ throughput. It now drains every queued chunk per crossing
485
+ (doc/architecture.md), and a knob sweep over hyper's h2 codec
486
+ (frame size, adaptive/bigger windows) moved nothing afterwards.
487
+ nginx's h2 upload collapse (~1.5k) is its default-config
488
+ request-body flow control; tune `http2_body_preread_size`/buffering
489
+ before drawing conclusions there.
490
+ - falcon lands at roughly a third of kino-h2-TLS on fast handlers and
491
+ slightly behind on uploads. Single-run caveat: the nginx lanes
492
+ showed ±20% swings between runs in the Docker environment; on the
493
+ reference box the kino-vs-kino and kino-vs-proxy ratios reproduced
494
+ across runs, the nginx lanes remain the noisiest.
495
+ - **Header-value interning** (user-agent, accept-*, sec-ch-*: one
496
+ frozen string instead of a fresh allocation per request) was
497
+ measured with a realistic 11-header browser set on /plaintext,
498
+ using the within-boot header cost (bare vs with-headers) as the
499
+ drift-resistant metric: over h2 that cost fell from ~16% to ~12-13%
500
+ (~+3-5% throughput on the headers lane); h1's smaller header cost
501
+ stayed within noise. The effect the tiny app understates: 6-8 fewer
502
+ string allocations per request is GC pressure a real app feels more
503
+ than this one does.
504
+
411
505
  ## Hot-path notes
412
506
 
413
507
  For the curious, the dispatch-path work behind the numbers: a try-pop
data/ext/kino/Cargo.toml CHANGED
@@ -1,6 +1,6 @@
1
1
  [package]
2
2
  name = "kino"
3
- version = "0.4.0"
3
+ version = "0.6.0"
4
4
  edition = "2021"
5
5
  authors = ["Yaroslav Markin <yaroslav@markin.net>"]
6
6
  license = "MIT"
@@ -21,8 +21,8 @@ smallvec = "1"
21
21
  lru = "0.18"
22
22
  mimalloc = { version = "0.1", default-features = false }
23
23
  tokio = { version = "1.45", features = ["rt-multi-thread", "net", "time", "sync", "io-util", "macros"] }
24
- hyper = { version = "1.6", features = ["http1", "server"] }
25
- hyper-util = { version = "0.1", features = ["server", "tokio", "http1"] }
24
+ hyper = { version = "1.6", features = ["http1", "http2", "server"] }
25
+ hyper-util = { version = "0.1", features = ["server", "server-auto", "tokio", "http1", "http2"] }
26
26
  http = "1"
27
27
  http-body-util = "0.1"
28
28
  bytes = "1"
@@ -49,3 +49,8 @@ rb-sys-env = "0.2.2"
49
49
  # undefined (the host process provides them at runtime); feature
50
50
  # unification turns this on only for `cargo test`.
51
51
  rb-sys = { version = "0.9", features = ["link-ruby"] }
52
+ # The protocol tests drive the server end of a duplex pipe with hyper's
53
+ # own client, and pause the clock to test timeouts; feature unification
54
+ # enables both only for `cargo test`.
55
+ hyper = { version = "1.6", features = ["client"] }
56
+ tokio = { version = "1.45", features = ["test-util"] }