kino 0.5.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/doc/benchmarks.md CHANGED
@@ -17,9 +17,21 @@ the deployment most apps run today.
17
17
  9R14 (Genoa), 16 GB RAM, Amazon Linux 2023, kernel 6.18. A realistic
18
18
  app-server size, deliberately: nobody provisions a 32-core box per
19
19
  app process.
20
- - Toolchain built on the box via mise: Ruby 4.0.5 (**YJIT enabled**,
21
- `RUBY_YJIT_ENABLE=1` for every server), Rust 1.96, Kino compiled in
20
+ - Toolchain built on the box via mise: Ruby 4.0.6 (**YJIT enabled**,
21
+ `RUBY_YJIT_ENABLE=1` for every server), Rust stable, Kino compiled in
22
22
  the release profile.
23
+ - **2026-09 full re-measurement** (a fresh c7a.2xlarge): every number in
24
+ this document and the README was re-run on Ruby 4.0.6 / kernel 6.18 /
25
+ Puma 8.0.2, adding the sharded-I/O and HTTP/2 studies. Kino's numbers
26
+ reproduced within 1-3% across the board—including the /io slot
27
+ ceiling and the arena balloon, both intact. The one real shift: Puma
28
+ 8.0.2 is faster than 7.x (142k plaintext, 61k /cpu), narrowing
29
+ ractor mode's /cpu lead to +25% (was +34%). Reversed-boot-order
30
+ re-runs reproduced within ~1%. A methodology trap worth recording:
31
+ a 5-second single-endpoint warmup understates memory badly (81 MB
32
+ where the full battery shows 137 MB) and hides the arena balloon
33
+ entirely—memory is only comparable after the full endpoint battery,
34
+ and /io numbers are only comparable at equal slot counts.
23
35
  - Load generator: wrk 4.2 on the same host, 8-second windows, 64
24
36
  connections (`bench/run.sh 8 64`). Same-host load generation costs
25
37
  both sides CPU equally; we verified the generator was not the
@@ -249,14 +261,17 @@ visible.
249
261
 
250
262
  | config | RSS | PSS |
251
263
  |---|---:|---:|
252
- | Kino :ractor 8×1 (default) | 151 | **148** |
253
- | Kino lanes 8×1 | 137 | **135** |
254
- | Kino :ractor 8×3 | 171 | **169** |
264
+ | Kino :ractor 8×1 (default) | 137 | **135** |
265
+ | Kino lanes 8×1 | 128 | **126** |
266
+ | Kino :ractor 8×3 | 170 | **168** |
255
267
  | Kino :threaded 8×3 (`MALLOC_ARENA_MAX=2`) | 109 | **107** |
256
- | Kino :threaded 8×3 (no arena cap) | 668 | **666**¹ |
257
- | Puma cluster 8×3 | 1,213 | **1,068** |
268
+ | Kino :threaded 8×3 (no arena cap) | 672 | **670**¹ |
269
+ | Puma cluster 8×3 | 1,216 | **1,072** |
258
270
 
259
- The tiny app is ~7× lighter than the cluster in ractor mode, ~10× in
271
+ (2026-09 re-measure; the 2026-06 numbers reproduced within a few MB on
272
+ every row—the arena-capped threaded row to the megabyte.)
273
+
274
+ The tiny app is ~8× lighter than the cluster in ractor mode, ~10× in
260
275
  arena-capped threaded mode. RSS ≈ PSS for every Kino row (one process,
261
276
  nothing to share) and within ~12% for Puma here: a trivial app has almost
262
277
  no shared state, so Puma's footprint is ~1,051 MB of *private* per-worker
@@ -280,18 +295,19 @@ Here copy-on-write **does** matter, which is exactly why PSS is mandatory:
280
295
 
281
296
  | config | RSS | PSS |
282
297
  |---|---:|---:|
283
- | Kino :threaded (one process) | 97 | **92** |
284
- | Puma cluster 8×3 (preload) | 794 | **389** |
298
+ | Kino :threaded (one process) | 97 | **95** |
299
+ | Puma cluster 8×5 (preload) | 813 | **405** |
300
+ | Puma cluster 8×5 (no preload) | 824 | **646** |
285
301
 
286
302
  Puma serves the same Rails framework from 8 forks that share it
287
- copy-on-write; RSS counts that shared framework once per worker (794 MB),
288
- PSS counts it once (389 MB). The fair ratio is **~4×**, not the ~8× a
303
+ copy-on-write; RSS counts that shared framework once per worker (813 MB),
304
+ PSS counts it once (405 MB). The fair ratio is **~4×**, not the ~8× a
289
305
  naive RSS sum reports—this is the correction that prompted the whole
290
- re-measure. Preload barely helps (389 vs 400 MB without): Ruby's GC
291
- dirties most heap pages, breaking copy-on-write, so even a preloaded
292
- cluster keeps a large private heap per worker. That is why "CoW should
293
- make a fork cluster nearly free" is only half true—it shares the code,
294
- not the live object heap.
306
+ re-measure. The 2026-09 run also split out preload: it now saves a
307
+ third (405 vs 646 MB PSS)—worth turning on—but even preloaded, Ruby's
308
+ GC dirties heap pages and breaks copy-on-write, so each worker keeps a
309
+ large private heap. That is why "CoW should make a fork cluster nearly
310
+ free" is only half true—it shares the code, not the live object heap.
295
311
 
296
312
  ## Run-to-run variance (a.k.a. "is this a regression?")
297
313
 
@@ -376,27 +392,46 @@ crash semantics, stealing fairness, and drain behavior have spec
376
392
  coverage but not production mileage. (On loopback-bound macOS, lanes
377
393
  lose a few percent instead; see the secondary table below.)
378
394
 
395
+ ## Sharded I/O (`io_shards true`)
396
+
397
+ `bench/studies.sh 8 64 shards`, ractor 8×3, 2026-09 reference box:
398
+
399
+ | case (/plaintext) | req/s |
400
+ |---|---:|
401
+ | shared tokio runtime (baseline) | 193,461 |
402
+ | io_shards, default shard count | 195,436 |
403
+ | io_shards, `io_threads 2` | 170,509 |
404
+ | io_shards, `io_threads 4` | 196,298 |
405
+ | io_shards, `io_threads 8` | 199,657 |
406
+
407
+ On 8 cores the shards buy +1-3%, best at `io_threads 8` (the
408
+ half-the-cores default is close behind; 2 shards choke on accept
409
+ handoff). The design removes work-stealing and cross-thread wakeups,
410
+ so the win scales with scheduler contention—expect more on boxes with
411
+ more cores and connections, and measure on your own core count.
412
+
379
413
  ## Logging costs
380
414
 
381
415
  Measured at full plaintext saturation (one log line per request—rates
382
416
  that no real deployment logs at; treat these as worst-case ceilings, not
383
417
  typical costs):
384
418
 
385
- | case (8×3, same session) | req/s |
419
+ | case (8×3, same session, 2026-09) | req/s |
386
420
  |---|---:|
387
- | threaded, no logging | 219,168 |
388
- | threaded, `log_requests true` (native access log) | 193,998 (−11%) |
389
- | ractor, access log off / on | 197,596 / 181,050 (−8%) |
390
- | app logs 1 line/req via shared `::Logger` (file) | **62,961** |
391
- | app logs 1 line/req via `Kino::Logger` (file) | **149,519 (2.4×)** |
421
+ | threaded, no logging | 213,166 |
422
+ | threaded, `log_requests true` (native access log) | 173,501 (−19%) |
423
+ | ractor, access log off / on | 191,152 / 163,777 (−14%) |
424
+ | app logs 1 line/req via shared `::Logger` (file) | **90,175** |
425
+ | app logs 1 line/req via `Kino::Logger` (file) | **154,653 (1.7×)** |
392
426
 
393
427
  The shared-`::Logger` cost is the mutex: 24 worker threads serialize
394
428
  through one lock plus a write syscall per line. `Kino::Logger` hands the
395
429
  formatted line to a lock-free channel and returns—the remaining cost vs
396
430
  not logging at all is Ruby-side formatting, which no device can remove.
397
- (In the Docker environment the same comparison showed 8.5×—overlay-fs
398
- write latency punished the synchronous logger far harder than this box's
399
- NVMe does. The ranking is environment-independent; the multiple is not.)
431
+ (The multiple moves with the environment: 2.4× on the 2026-06 box,
432
+ 1.7× on the 2026-09 one, 8.5× under Docker, where overlay-fs write
433
+ latency punished the synchronous logger hardest. The ranking is
434
+ environment-independent; the multiple is not.)
400
435
 
401
436
  One trade-off worth knowing: the sink **never blocks** request threads,
402
437
  so at absurd rates against a slow disk it drops lines once its 8192-line
@@ -408,6 +443,65 @@ Puma comparison note: request logging is opt-in there too (`--quiet` is
408
443
  the default, `-v/--log-requests` enables it)—Kino's default-off
409
444
  `log_requests` matches the ecosystem's standard behavior.
410
445
 
446
+ ## HTTP/2
447
+
448
+ `bench/h2.sh`, Linux only (on macOS run it under Docker; the 2026-09
449
+ numbers below are from the c7a.2xlarge reference box). Every lane is
450
+ measured with h2load so the generator is identical everywhere: h2
451
+ lanes run 8 connections × 8 concurrent streams, h1 lanes 64
452
+ connections—the same total in-flight. Servers without native h2 get
453
+ the standard pattern instead: nginx terminating h2 and proxying
454
+ HTTP/1.1 upstream over keep-alive. One labeled run (kino ractor 8×3,
455
+ falcon `--count 8`, puma `-w 8 -t 3:3`, 5 s/lane); re-run the whole
456
+ script for close calls, per the variance section.
457
+
458
+ | target (h2 unless noted) | /plaintext | /10k | /big-cookie | /upload (64 KB) |
459
+ |---|---:|---:|---:|---:|
460
+ | kino h2c | 207,461 | 150,105 | 195,093 | 26,304 |
461
+ | kino h1 cleartext (same boot) | 116,082 | 97,634 | 110,614 | 24,338 |
462
+ | kino h2 TLS | 164,044 | 116,617 | 154,301 | 19,226 |
463
+ | kino h1 TLS (same boot) | 80,789 | 69,942 | 76,654 | 17,210 |
464
+ | falcon TLS (native h2) | 55,637 | 37,541 | 49,151 | 18,838 |
465
+ | nginx h2 → puma h1 | 81,234 | 57,891 | 51,332 | 1,247 |
466
+ | nginx h2 → kino h1 | 109,219 | 66,634 | 54,996 | 1,217 |
467
+
468
+ What the numbers say:
469
+
470
+ - **Native h2 beats h1 on the same server by +79% cleartext and +103%
471
+ over TLS** on /plaintext: the same 64 in-flight requests ride 8
472
+ connections instead of 64, so frames batch into fewer, larger
473
+ syscalls—and TLS amplifies it, since h1's 64 connections each pay
474
+ crypto per record. The `/big-cookie` lane (a ~2 KB cookie per
475
+ request) shows HPACK on top: the cookie crosses the wire once per
476
+ connection, not once per request, and holds 94% of bare-plaintext
477
+ throughput where h1 loses 5%.
478
+ - **Native h2 beats proxied h2 by +50%** with the *same backend*: the
479
+ nginx→kino-h1 lane is the proxy-cost control, and the extra hop,
480
+ re-parse, and re-serialize cost ~55k req/s on /plaintext.
481
+ - **Uploads run at h1 parity**—but only after a fix this lane
482
+ caught: h2 delivers bodies as 16 KB DATA frames, and `read_body`
483
+ originally crossed the GVL once per chunk, halving upload
484
+ throughput. It now drains every queued chunk per crossing
485
+ (doc/architecture.md), and a knob sweep over hyper's h2 codec
486
+ (frame size, adaptive/bigger windows) moved nothing afterwards.
487
+ nginx's h2 upload collapse (~1.5k) is its default-config
488
+ request-body flow control; tune `http2_body_preread_size`/buffering
489
+ before drawing conclusions there.
490
+ - falcon lands at roughly a third of kino-h2-TLS on fast handlers and
491
+ slightly behind on uploads. Single-run caveat: the nginx lanes
492
+ showed ±20% swings between runs in the Docker environment; on the
493
+ reference box the kino-vs-kino and kino-vs-proxy ratios reproduced
494
+ across runs, the nginx lanes remain the noisiest.
495
+ - **Header-value interning** (user-agent, accept-*, sec-ch-*: one
496
+ frozen string instead of a fresh allocation per request) was
497
+ measured with a realistic 11-header browser set on /plaintext,
498
+ using the within-boot header cost (bare vs with-headers) as the
499
+ drift-resistant metric: over h2 that cost fell from ~16% to ~12-13%
500
+ (~+3-5% throughput on the headers lane); h1's smaller header cost
501
+ stayed within noise. The effect the tiny app understates: 6-8 fewer
502
+ string allocations per request is GC pressure a real app feels more
503
+ than this one does.
504
+
411
505
  ## Hot-path notes
412
506
 
413
507
  For the curious, the dispatch-path work behind the numbers: a try-pop
data/ext/kino/Cargo.toml CHANGED
@@ -1,6 +1,6 @@
1
1
  [package]
2
2
  name = "kino"
3
- version = "0.5.0"
3
+ version = "0.7.0"
4
4
  edition = "2021"
5
5
  authors = ["Yaroslav Markin <yaroslav@markin.net>"]
6
6
  license = "MIT"
@@ -21,8 +21,8 @@ smallvec = "1"
21
21
  lru = "0.18"
22
22
  mimalloc = { version = "0.1", default-features = false }
23
23
  tokio = { version = "1.45", features = ["rt-multi-thread", "net", "time", "sync", "io-util", "macros"] }
24
- hyper = { version = "1.6", features = ["http1", "server"] }
25
- hyper-util = { version = "0.1", features = ["server", "tokio", "http1"] }
24
+ hyper = { version = "1.6", features = ["http1", "http2", "server"] }
25
+ hyper-util = { version = "0.1", features = ["server", "server-auto", "tokio", "http1", "http2"] }
26
26
  http = "1"
27
27
  http-body-util = "0.1"
28
28
  bytes = "1"
@@ -49,3 +49,8 @@ rb-sys-env = "0.2.2"
49
49
  # undefined (the host process provides them at runtime); feature
50
50
  # unification turns this on only for `cargo test`.
51
51
  rb-sys = { version = "0.9", features = ["link-ruby"] }
52
+ # The protocol tests drive the server end of a duplex pipe with hyper's
53
+ # own client, and pause the clock to test timeouts; feature unification
54
+ # enables both only for `cargo test`.
55
+ hyper = { version = "1.6", features = ["client"] }
56
+ tokio = { version = "1.45", features = ["test-util"] }
@@ -18,6 +18,8 @@ pub struct WorkerStat {
18
18
  pub in_flight: usize,
19
19
  pub busy_ms: u64,
20
20
  pub quarantined: bool,
21
+ /// Sent home by the pool scaler; leaves at its next idle tick.
22
+ pub retired: bool,
21
23
  }
22
24
 
23
25
  /// Read every slot's per-worker sensors in one pass under the slots read
@@ -46,6 +48,7 @@ pub fn collect_worker_status(server: &ServerInner) -> Vec<WorkerStat> {
46
48
  now.saturating_sub(started)
47
49
  },
48
50
  quarantined,
51
+ retired: slot.retired.load(Ordering::Relaxed),
49
52
  }
50
53
  })
51
54
  .collect()
@@ -70,6 +73,10 @@ pub struct StatsSnapshot {
70
73
  pub worker_status: Vec<WorkerStat>,
71
74
  pub quarantined_count: usize,
72
75
  pub quarantine_replacements: u64,
76
+ pub max_workers: usize,
77
+ pub active_workers: usize,
78
+ pub scale_ups: u64,
79
+ pub scale_downs: u64,
73
80
  pub queue_histogram: crate::registry::QueueHistogramSnapshot,
74
81
  }
75
82
 
@@ -94,6 +101,10 @@ impl StatsSnapshot {
94
101
  worker_status,
95
102
  quarantined_count,
96
103
  quarantine_replacements: server.quarantine_replacements.load(Ordering::Relaxed),
104
+ max_workers: server.topology.max_workers,
105
+ active_workers: server.active_workers.load(Ordering::Relaxed),
106
+ scale_ups: server.scale_ups.load(Ordering::Relaxed),
107
+ scale_downs: server.scale_downs.load(Ordering::Relaxed),
97
108
  queue_histogram: server.queue_histogram.snapshot(),
98
109
  }
99
110
  }
@@ -114,9 +125,10 @@ pub fn stats_json(s: &StatsSnapshot) -> String {
114
125
  let mut out = String::with_capacity(256);
115
126
  write!(
116
127
  out,
117
- r#"{{"mode":"{}","lanes":{},"workers":{},"threads":{},"batch":{},"respawns":{},"queued":{},"in_flight":{},"served":{},"rejected":{},"timeouts":{}"#,
128
+ r#"{{"mode":"{}","lanes":{},"workers":{},"threads":{},"batch":{},"respawns":{},"queued":{},"in_flight":{},"served":{},"rejected":{},"timeouts":{},"max_workers":{},"active_workers":{},"scale_ups":{},"scale_downs":{}"#,
118
129
  s.mode, s.lanes, s.workers, s.threads, s.batch, s.respawns,
119
- s.queued, s.in_flight, s.served, s.rejected, s.timeouts
130
+ s.queued, s.in_flight, s.served, s.rejected, s.timeouts,
131
+ s.max_workers, s.active_workers, s.scale_ups, s.scale_downs
120
132
  )
121
133
  .expect("writing to a String cannot fail");
122
134
  if let Some(depths) = &s.lane_depths {
@@ -134,8 +146,8 @@ pub fn stats_json(s: &StatsSnapshot) -> String {
134
146
  }
135
147
  write!(
136
148
  out,
137
- r#"{{"index":{},"served":{},"in_flight":{},"busy_ms":{},"quarantined":{}}}"#,
138
- w.index, w.served, w.in_flight, w.busy_ms, w.quarantined
149
+ r#"{{"index":{},"served":{},"in_flight":{},"busy_ms":{},"quarantined":{},"retired":{}}}"#,
150
+ w.index, w.served, w.in_flight, w.busy_ms, w.quarantined, w.retired
139
151
  )
140
152
  .expect("writing to a String cannot fail");
141
153
  }
@@ -247,6 +259,34 @@ pub fn metrics_text(s: &StatsSnapshot) -> String {
247
259
  "Configured threads per worker.",
248
260
  s.threads,
249
261
  );
262
+ metric(
263
+ &mut out,
264
+ "kino_max_workers",
265
+ "gauge",
266
+ "Worker pool ceiling (equals kino_workers for a fixed pool).",
267
+ s.max_workers,
268
+ );
269
+ metric(
270
+ &mut out,
271
+ "kino_active_workers",
272
+ "gauge",
273
+ "Workers currently alive and serving.",
274
+ s.active_workers,
275
+ );
276
+ metric(
277
+ &mut out,
278
+ "kino_scale_ups_total",
279
+ "counter",
280
+ "Workers added by the pool scaler.",
281
+ s.scale_ups,
282
+ );
283
+ metric(
284
+ &mut out,
285
+ "kino_scale_downs_total",
286
+ "counter",
287
+ "Idle workers retired by the pool scaler.",
288
+ s.scale_downs,
289
+ );
250
290
  metric(
251
291
  &mut out,
252
292
  "kino_ready",
@@ -641,6 +681,10 @@ mod tests {
641
681
  worker_status: vec![],
642
682
  quarantined_count: 0,
643
683
  quarantine_replacements: 0,
684
+ max_workers: 32,
685
+ active_workers: 12,
686
+ scale_ups: 3,
687
+ scale_downs: 1,
644
688
  queue_histogram: crate::registry::QueueHistogramSnapshot {
645
689
  buckets: [0; crate::registry::QUEUE_BOUNDS_US.len()],
646
690
  overflow: 0,
@@ -665,6 +709,10 @@ mod tests {
665
709
  r#""served":100"#,
666
710
  r#""rejected":5"#,
667
711
  r#""timeouts":6"#,
712
+ r#""max_workers":32"#,
713
+ r#""active_workers":12"#,
714
+ r#""scale_ups":3"#,
715
+ r#""scale_downs":1"#,
668
716
  r#""state":"ready""#,
669
717
  r#""version":""#,
670
718
  ] {
@@ -690,6 +738,49 @@ mod tests {
690
738
  assert!(draining.contains("kino_ready 0"));
691
739
  }
692
740
 
741
+ #[test]
742
+ fn metrics_text_reports_the_elastic_pool() {
743
+ let text = metrics_text(&snapshot(crate::registry::STATE_READY));
744
+ assert!(text.contains("# TYPE kino_max_workers gauge"));
745
+ assert!(text.contains("kino_max_workers 32"));
746
+ assert!(text.contains("# TYPE kino_active_workers gauge"));
747
+ assert!(text.contains("kino_active_workers 12"));
748
+ assert!(text.contains("# TYPE kino_scale_ups_total counter"));
749
+ assert!(text.contains("kino_scale_ups_total 3"));
750
+ assert!(text.contains("# TYPE kino_scale_downs_total counter"));
751
+ assert!(text.contains("kino_scale_downs_total 1"));
752
+ }
753
+
754
+ #[test]
755
+ fn worker_status_reports_retired_slots() {
756
+ let server = crate::registry::test_server(false, 4);
757
+ server.register_worker();
758
+ server.register_worker();
759
+ server.slots.read()[1].retire();
760
+
761
+ let status = collect_worker_status(&server);
762
+
763
+ assert_eq!(
764
+ status.iter().map(|w| w.retired).collect::<Vec<_>>(),
765
+ vec![false, true]
766
+ );
767
+ assert!(status.iter().all(|w| w.busy_ms == 0));
768
+ }
769
+
770
+ #[test]
771
+ fn snapshot_reads_the_pool_counters() {
772
+ let server = crate::registry::test_server(false, 4);
773
+ server.active_workers.store(3, Ordering::Relaxed);
774
+ server.scale_ups.fetch_add(2, Ordering::Relaxed);
775
+ server.scale_downs.fetch_add(1, Ordering::Relaxed);
776
+
777
+ let s = StatsSnapshot::take(&server);
778
+
779
+ assert_eq!(s.active_workers, 3);
780
+ assert_eq!(s.max_workers, server.topology.max_workers);
781
+ assert_eq!((s.scale_ups, s.scale_downs), (2, 1));
782
+ }
783
+
693
784
  #[test]
694
785
  fn metrics_text_includes_lane_depth_samples_when_lanes_are_on() {
695
786
  let mut s = snapshot(crate::registry::STATE_READY);
@@ -757,6 +848,7 @@ mod tests {
757
848
  in_flight: 1,
758
849
  busy_ms: 4,
759
850
  quarantined: false,
851
+ retired: false,
760
852
  },
761
853
  WorkerStat {
762
854
  index: 1,
@@ -764,10 +856,11 @@ mod tests {
764
856
  in_flight: 0,
765
857
  busy_ms: 0,
766
858
  quarantined: false,
859
+ retired: true,
767
860
  },
768
861
  ];
769
862
  let json = stats_json(&s);
770
- assert!(json.contains(r#""worker_status":[{"index":0,"served":10,"in_flight":1,"busy_ms":4,"quarantined":false},{"index":1,"served":7,"in_flight":0,"busy_ms":0,"quarantined":false}]"#), "got {json}");
863
+ assert!(json.contains(r#""worker_status":[{"index":0,"served":10,"in_flight":1,"busy_ms":4,"quarantined":false,"retired":false},{"index":1,"served":7,"in_flight":0,"busy_ms":0,"quarantined":false,"retired":true}]"#), "got {json}");
771
864
  }
772
865
 
773
866
  #[test]
@@ -786,6 +879,7 @@ mod tests {
786
879
  in_flight: 1,
787
880
  busy_ms: 4,
788
881
  quarantined: false,
882
+ retired: false,
789
883
  },
790
884
  WorkerStat {
791
885
  index: 1,
@@ -793,6 +887,7 @@ mod tests {
793
887
  in_flight: 0,
794
888
  busy_ms: 0,
795
889
  quarantined: false,
890
+ retired: true,
796
891
  },
797
892
  ];
798
893
  let text = metrics_text(&s);
@@ -835,6 +930,7 @@ mod tests {
835
930
  in_flight: 1,
836
931
  busy_ms: 0,
837
932
  quarantined: true,
933
+ retired: false,
838
934
  },
839
935
  WorkerStat {
840
936
  index: 1,
@@ -842,6 +938,7 @@ mod tests {
842
938
  in_flight: 1,
843
939
  busy_ms: 5,
844
940
  quarantined: false,
941
+ retired: false,
845
942
  },
846
943
  ];
847
944
  let json = stats_json(&s);
@@ -850,7 +947,7 @@ mod tests {
850
947
  "top-level count: {json}"
851
948
  );
852
949
  assert!(
853
- json.contains(r#"{"index":0,"served":3,"in_flight":1,"busy_ms":0,"quarantined":true}"#),
950
+ json.contains(r#"{"index":0,"served":3,"in_flight":1,"busy_ms":0,"quarantined":true,"retired":false}"#),
854
951
  "{json}"
855
952
  );
856
953
  assert!(json.contains(r#""quarantined":false"#));