loaderx 1.4.2__tar.gz → 1.5.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (102) hide show
  1. {loaderx-1.4.2/loaderx.egg-info → loaderx-1.5.2}/PKG-INFO +153 -141
  2. {loaderx-1.4.2 → loaderx-1.5.2}/README.md +152 -140
  3. {loaderx-1.4.2 → loaderx-1.5.2}/build.zig +1 -0
  4. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/__init__.py +1 -1
  5. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/zrecord.py +3 -3
  6. {loaderx-1.4.2 → loaderx-1.5.2/loaderx.egg-info}/PKG-INFO +153 -141
  7. {loaderx-1.4.2 → loaderx-1.5.2}/scripts/bench.py +14 -8
  8. {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord/format.zig +50 -34
  9. loaderx-1.5.2/src/zrecord/storage.zig +541 -0
  10. {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord.zig +73 -36
  11. {loaderx-1.4.2 → loaderx-1.5.2}/tests/test_loaderx.py +16 -5
  12. loaderx-1.4.2/src/zrecord/storage.zig +0 -396
  13. {loaderx-1.4.2 → loaderx-1.5.2}/LICENSE +0 -0
  14. {loaderx-1.4.2 → loaderx-1.5.2}/MANIFEST.in +0 -0
  15. {loaderx-1.4.2 → loaderx-1.5.2}/build.zig.zon +0 -0
  16. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/_lib.py +0 -0
  17. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/dataloader.py +0 -0
  18. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/dataset.py +0 -0
  19. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/sampler.py +0 -0
  20. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/utils.py +0 -0
  21. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/SOURCES.txt +0 -0
  22. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/dependency_links.txt +0 -0
  23. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/requires.txt +0 -0
  24. {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/top_level.txt +0 -0
  25. {loaderx-1.4.2 → loaderx-1.5.2}/pyproject.toml +0 -0
  26. {loaderx-1.4.2 → loaderx-1.5.2}/scripts/build_wheels.py +0 -0
  27. {loaderx-1.4.2 → loaderx-1.5.2}/setup.cfg +0 -0
  28. {loaderx-1.4.2 → loaderx-1.5.2}/setup.py +0 -0
  29. {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord/codec.zig +0 -0
  30. {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord/exec.zig +0 -0
  31. {loaderx-1.4.2 → loaderx-1.5.2}/src/zsampler.zig +0 -0
  32. {loaderx-1.4.2 → loaderx-1.5.2}/src/zstd/c.zig +0 -0
  33. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/COPYING +0 -0
  34. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/LICENSE +0 -0
  35. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/allocations.h +0 -0
  36. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/bits.h +0 -0
  37. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/bitstream.h +0 -0
  38. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/compiler.h +0 -0
  39. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/cpu.h +0 -0
  40. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/debug.c +0 -0
  41. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/debug.h +0 -0
  42. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/entropy_common.c +0 -0
  43. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/error_private.c +0 -0
  44. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/error_private.h +0 -0
  45. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/fse.h +0 -0
  46. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/fse_decompress.c +0 -0
  47. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/huf.h +0 -0
  48. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/mem.h +0 -0
  49. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/pool.c +0 -0
  50. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/pool.h +0 -0
  51. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/portability_macros.h +0 -0
  52. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/threading.c +0 -0
  53. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/threading.h +0 -0
  54. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/xxhash.c +0 -0
  55. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/xxhash.h +0 -0
  56. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_common.c +0 -0
  57. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_deps.h +0 -0
  58. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_internal.h +0 -0
  59. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_trace.h +0 -0
  60. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/clevels.h +0 -0
  61. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/fse_compress.c +0 -0
  62. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/hist.c +0 -0
  63. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/hist.h +0 -0
  64. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/huf_compress.c +0 -0
  65. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress.c +0 -0
  66. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_internal.h +0 -0
  67. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_literals.c +0 -0
  68. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_literals.h +0 -0
  69. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_sequences.c +0 -0
  70. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_sequences.h +0 -0
  71. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_superblock.c +0 -0
  72. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_superblock.h +0 -0
  73. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_cwksp.h +0 -0
  74. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_double_fast.c +0 -0
  75. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_double_fast.h +0 -0
  76. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_fast.c +0 -0
  77. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_fast.h +0 -0
  78. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_lazy.c +0 -0
  79. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_lazy.h +0 -0
  80. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_ldm.c +0 -0
  81. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_ldm.h +0 -0
  82. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_ldm_geartab.h +0 -0
  83. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_opt.c +0 -0
  84. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_opt.h +0 -0
  85. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstdmt_compress.c +0 -0
  86. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstdmt_compress.h +0 -0
  87. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/huf_decompress.c +0 -0
  88. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_ddict.c +0 -0
  89. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_ddict.h +0 -0
  90. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress.c +0 -0
  91. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress_block.c +0 -0
  92. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress_block.h +0 -0
  93. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress_internal.h +0 -0
  94. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/cover.c +0 -0
  95. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/cover.h +0 -0
  96. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/divsufsort.c +0 -0
  97. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/divsufsort.h +0 -0
  98. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/fastcover.c +0 -0
  99. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/zdict.c +0 -0
  100. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/zdict.h +0 -0
  101. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/zstd.h +0 -0
  102. {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/zstd_errors.h +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: loaderx
3
- Version: 1.4.2
3
+ Version: 1.5.2
4
4
  Summary: A compact, high-performance persistent record store with zero-copy batch gathering and transparent per-record compression, designed for single-machine AI training and serving pipelines
5
5
  Author-email: Ben0i0d <ben0i0d@foxmail.com>
6
6
  License: MIT License
@@ -238,9 +238,11 @@ and how much data trains it — so a caller picks a tier, never a number:
238
238
 
239
239
  The sample is the byte budget the dictionary trains on (a strided subset of the
240
240
  records), so each tier costs the same training time whatever the record size.
241
- On this box's 12 KiB sample, `"balanced"` gathers ~3660 MiB/s at 1.17x against
242
- `"max"`'s ~2400 MiB/s at 1.34x — for a loader the throughput is the hot path,
243
- for archival the ratio is.
241
+ At realistic image scale (768 KiB records) the tiers converge — on the earlier
242
+ measurement box's structured data `"balanced"` and `"max"` both gather
243
+ ~1.6 GiB/s at a 1.74x ratio — because a dictionary is a small fraction of a
244
+ large frame. The tiers still matter at small record sizes, where the dict is
245
+ most of a record and `"max"` trades gather throughput for ratio.
244
246
 
245
247
  `"zstd_dict"` records can only be read from a store that has the dictionary
246
248
  (`dict.zr`). The dictionary is loaded on open and shared, lock-free, across all
@@ -383,59 +385,58 @@ For practical integration examples, please refer to the **[Data2Latent](https://
383
385
  the index sampler, the record store, and the full data loader — each through the
384
386
  binding a client actually uses, so cffi, the GIL and the NumPy allocation are all
385
387
  inside the timings. One run is enough: this is a qualitative horizontal
386
- comparison, and the defaults are sized so `python3 scripts/bench.py all` — which
387
- prints this machine block first — finishes in under half a minute. Warm page
388
- cache.
388
+ comparison, sized to a realistic image workload (625 MiB of 256 KiB records)
389
+ that finishes `python3 scripts/bench.py all` — which prints this machine block
390
+ first — in under half a minute. Warm page cache.
389
391
 
390
- **Machine** — one box, inside a container on a personal laptop (AMD Strix Point
391
- APU, 12 cores / 24 threads):
392
+ **Machine** — one box, an LXC container on a server (AMD EPYC 7303, 16 cores /
393
+ 32 threads):
392
394
 
393
395
  | machine | value |
394
396
  |-----------|-------|
395
- | CPU | AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M, 1 socket, 12 cores / 24 threads |
396
- | frequency | 605–5158 MHz |
397
- | caches | L1d 576 KiB, L1i 384 KiB, L2 12 MiB, L3 24 MiB |
398
- | NUMA | 1 node |
399
- | memory | 31 GiB (not limited by cgroup) |
400
- | OS | Debian GNU/Linux forky/sid, kernel 7.1.3+deb13-amd64 |
401
- | python | 3.14.6, numpy 2.5.1 |
397
+ | CPU | AMD EPYC 7303 16-Core Processor, 1 socket, 16 cores / 32 threads |
398
+ | frequency | 1500–3437 MHz |
399
+ | caches | L1d 512 KiB, L1i 512 KiB, L2 8 MiB, L3 64 MiB |
400
+ | NUMA | 4 nodes |
401
+ | memory | 32 GiB (32 GiB cgroup limit) |
402
+ | OS | Debian GNU/Linux 13 (trixie), kernel 6.12.95+deb13-amd64 |
403
+ | python | 3.14.7, numpy 2.5.1 |
402
404
 
403
- The container sees all 24 threads and has no CPU quota. The benchmark runs on
404
- the ordinary page cache.
405
+ The container sees all 32 threads under a 32 GiB cgroup limit. The benchmark
406
+ runs on the ordinary page cache.
405
407
 
406
408
  **1. Sampler** — index generation on its own, IID (with replacement), 1M index
407
409
  space, against NumPy's modern API. The µs-scale figures fluctuate with box load;
408
- the stable signal is the ~2x margin, widest at the smallest batches.
410
+ the stable signal is the ~1.9x margin, roughly flat across batch sizes.
409
411
 
410
412
  | sampler | batch | per batch | vs default_rng |
411
413
  |-------------------|-------|-----------|----------------|
412
- | numpy default_rng | 256 | 5.7 µs | 1.00x |
413
- | **zsampler** | 256 | 1.7 µs | **3.32x** |
414
- | numpy default_rng | 1024 | 4.4 µs | 1.00x |
415
- | **zsampler** | 1024 | 2.3 µs | **1.89x** |
416
- | numpy default_rng | 8192 | 18.6 µs | 1.00x |
417
- | **zsampler** | 8192 | 9.5 µs | **1.97x** |
418
-
419
- **2. Store** — Zrecord against the alternatives: random batch gather, 10000
420
- records, batch 256, structured (mildly compressible) data. A single 12 KiB
421
- sample is the horizontal comparison; other record sizes are a `--shape` away.
422
- Two backends stay out of the default because they would dominate the runtime
423
- without moving the comparison — they remain reachable through `--only`:
424
- arrayrecord (no usable wheel here, and one record per Python call) and
425
- hdf5-gzip (the same driver as hdf5 plus a codec axis, ~180 MiB/s gather vs
426
- hdf5's ~430).
427
-
428
- Small records — 12 KiB per record, 3 MiB per batch:
414
+ | numpy default_rng | 256 | 5.6 µs | 1.00x |
415
+ | **zsampler** | 256 | 3.0 µs | **1.89x** |
416
+ | numpy default_rng | 1024 | 8.0 µs | 1.00x |
417
+ | **zsampler** | 1024 | 4.3 µs | **1.85x** |
418
+ | numpy default_rng | 8192 | 31.3 µs | 1.00x |
419
+ | **zsampler** | 8192 | 17.0 µs | **1.84x** |
420
+
421
+ **2. Store** — Zrecord against the alternatives: random batch gather, 2500
422
+ records, batch 256, structured (mildly compressible) data. A single 256 KiB
423
+ sample (512×512 single-channel, a realistic image record) is the horizontal
424
+ comparison; other record sizes are a `--shape` away. hdf5-gzip, arrayrecord,
425
+ torch and grain are installed and run in the default tables.
426
+
427
+ Realistic records — 256 KiB per record, 64 MiB per batch:
429
428
 
430
429
  | store | write | gather | on disk | ratio |
431
430
  |-------------------|------------|-------------|---------|-------|
432
- | zrecord-zstd | 1022 MiB/s | 11126 MiB/s | 108 MiB | 1.08x |
433
- | zrecord-zstdict | 33 MiB/s | 4587 MiB/s | 100 MiB | 1.18x |
434
- | zrecord-raw | 1573 MiB/s | 19208 MiB/s | 117 MiB | 1.00x |
435
- | npy-mmap | 2566 MiB/s | 15762 MiB/s | 117 MiB | 1.00x |
436
- | hdf5 | 1052 MiB/s | 380 MiB/s | 118 MiB | 0.99x |
437
- | blosc2 | 32 MiB/s | 439 MiB/s | 108 MiB | 1.09x |
438
- | tensorstore | 463 MiB/s | 104 MiB/s | 107 MiB | 1.09x |
431
+ | zrecord-zstd | 772 MiB/s | 4031 MiB/s | 555 MiB | 1.13x |
432
+ | zrecord-zstdict | 72 MiB/s | 2123 MiB/s | 388 MiB | 1.61x |
433
+ | zrecord-raw | 880 MiB/s | 5214 MiB/s | 625 MiB | 1.00x |
434
+ | npy-mmap | 1398 MiB/s | 1853 MiB/s | 625 MiB | 1.00x |
435
+ | hdf5 | 975 MiB/s | 1046 MiB/s | 625 MiB | 1.00x |
436
+ | hdf5-gzip | 37 MiB/s | 182 MiB/s | 445 MiB | 1.40x |
437
+ | blosc2 | 48 MiB/s | 353 MiB/s | 551 MiB | 1.13x |
438
+ | tensorstore | 720 MiB/s | 109 MiB/s | 462 MiB | 1.35x |
439
+ | arrayrecord | 331 MiB/s | 865 MiB/s | 567 MiB | 1.10x |
439
440
 
440
441
  `write` is page-cache ingestion — the store is built and left open, exactly
441
442
  like `np.save` and the other backends, which defer durability to the kernel.
@@ -448,34 +449,36 @@ lazy ones.
448
449
 
449
450
  `gather` is fully page-cache-warmed — the benchmark sweeps every record once
450
451
  before timing, so it measures the pure access path (memory), not first-touch
451
- page faults or disk. That correction matters: mmap pays a fault per page on
452
- first touch while zrecord's `pread` reads the just-written cache directly, so a
453
- sparse warm-up inflated zrecord's margin over `.npy`. Fully warm, zrecord-raw
454
- leads by ~1.2x on this box (24 cores) — its gather is a fan-out across every
455
- core, numpy's fancy index is single-threaded — and on the 12-core reference box
456
- the margin is wider, because single-threaded NumPy copy is slower there.
452
+ page faults or disk. At 256 KiB records the batch (64 MiB) exceeds L3, so these
453
+ numbers are DRAM-bandwidth-bound for every backend; the 12 KiB table that
454
+ predates the record-size bump was an L3-fit artifact. Fully warm, zrecord-raw
455
+ leads npy-mmap by ~2.8x on this box — its gather fans out across every core
456
+ while numpy's fancy index is single-threaded — and the gap widens where
457
+ single-threaded copy is slower.
457
458
 
458
459
  **3. Loader** — the full input pipeline end to end (sample, fetch, collate, hand
459
- over a batch), same workload, 4 workers. The `peak RSS` column (added with the
460
- uniform-codec rewrite) is the highest resident set size of the whole process
461
- tree while batches are flowing — measured on this box, so the numbers below are
462
- the box's, not the reference machine's. (grain stays out of the default: it
463
- needs JAX, its build writes ArrayRecord one sample at a time — tens of seconds
464
- for this workload — and it lands ~50 batches/s, an order of magnitude below the
465
- rest. `--only grain` reaches it when installed.)
460
+ over a batch), same workload, 4 workers, batch 256 (64 MiB). The `peak RSS`
461
+ column (added with the uniform-codec rewrite) is the highest resident set size
462
+ of the whole process tree while batches are flowing. torch, grain and
463
+ arrayrecord are installed and wired into the defaults, and all four loaders run
464
+ at the full batch.
466
465
 
467
466
  | loader | batches/s | peak RSS |
468
467
  |----------------|-----------|----------|
469
- | **loaderx** | 4115.5 | ~220 MiB |
470
- | loaderx-raw | 4833.8 | ~245 MiB |
471
- | torch | 770.7 | ~3000 MiB |
472
-
473
- The memory gap is the architecture, not an accident: loaderx prefetches in
474
- threads inside one process, so four workers share one interpreter, one numpy and
475
- one set of gather buffers; torch's four workers are four processes, each a full
476
- Python runtime holding its own copy of the worker state — ~13x the resident
477
- memory for the same parallelism, which is a real cost on a training box that is
478
- already feeding the model.
468
+ | **loaderx** | 59.2 | 2477 MiB |
469
+ | loaderx-raw | 80.1 | 2478 MiB |
470
+ | torch | 37.6 | 14525 MiB |
471
+ | grain | 15.8 | 4567 MiB |
472
+
473
+ At 64 MiB per batch the loader is DRAM-bound, not sampler-bound — the per-batch
474
+ gather (zstd ~4.0 GiB/s, raw ~5.2 GiB/s) is the whole story, and the prefetch
475
+ threads keep it at store-gather speed while the Python side collates. The
476
+ memory is the prefetch buffers plus the two dataset handles (raw + label stores
477
+ over the same backing data). loaderx prefetches in threads inside one process,
478
+ so workers share one interpreter, one numpy and one set of gather buffers; torch
479
+ and grain run a worker process per prefetch thread, which is most of their RSS
480
+ (torch's peak also includes the shared-memory collated batches). grain's
481
+ batches/s is the pipeline's floor here, ~5x below loaderx-raw.
479
482
 
480
483
  **Free-threaded Python.** loaderx targets free-threaded builds (no GIL), and
481
484
  the key sections are also measured on 3.14t — the question is whether anything
@@ -485,20 +488,22 @@ and `--transform augment`); they are not part of the default `all`:
485
488
  * **Sampler** — zsampler draws ~1.9–2x faster than free-threaded numpy, the
486
489
  same margin as on the GIL build: a batch is one Zig call either way.
487
490
  * **Store** — the free-threaded numbers track the GIL table within ~10%.
488
- `zrecord-raw` gathers ~19.4–21.4 GiB/s on both (still ~1.7x `npy-mmap`),
489
- `zrecord-zstd` ~9.4–10.7 GiB/s. Notably, the alternatives do not gain from the
490
- build: blosc2's free-threaded wheel re-enables the GIL to load its extension,
491
- so it runs exactly as on the GIL build. Random gather, 12 KiB records:
491
+ `zrecord-raw` gathers 4.9 vs 5.2 GiB/s (still ~2.6x `npy-mmap`),
492
+ `zrecord-zstd` 3.9 vs 4.0 GiB/s. Notably, the alternatives do not gain from
493
+ the build: blosc2's free-threaded wheel re-enables the GIL to load its
494
+ extension, so it runs exactly as on the GIL build. Random gather, 256 KiB
495
+ records:
492
496
 
493
497
  | store | CPython 3.14 (GIL) | free-threaded 3.14t |
494
498
  |--------------|--------------------|---------------------|
495
- | zrecord-raw | 21412 MiB/s | 19374 MiB/s |
496
- | zrecord-zstd | 10668 MiB/s | 9399 MiB/s |
497
- | npy-mmap | 12063 MiB/s | 10399 MiB/s |
498
- | hdf5 | 436 MiB/s | 307 MiB/s |
499
- | blosc2 | 434 MiB/s | 419 MiB/s |
500
- | tensorstore | 87 MiB/s | 86 MiB/s |
501
- * **Loader, identity** — unchanged, ~4100–6000 batches/s on both interpreters.
499
+ | zrecord-raw | 5214 MiB/s | 4886 MiB/s |
500
+ | zrecord-zstd | 4031 MiB/s | 3913 MiB/s |
501
+ | npy-mmap | 1853 MiB/s | 1861 MiB/s |
502
+ | hdf5 | 1046 MiB/s | 1030 MiB/s |
503
+ | blosc2 | 353 MiB/s | 334 MiB/s |
504
+ | tensorstore | 109 MiB/s | 110 MiB/s |
505
+ * **Loader, identity** — unchanged, ~60 batches/s (loaderx) on both
506
+ interpreters at 64 MiB per batch.
502
507
  * **Loader, CPU-heavy transform** — the one place the free-threaded build
503
508
  matters. A transform runs on the prefetch threads, so the GIL serializes it on
504
509
  standard CPython — the case where worker processes win, and the reason for the
@@ -508,13 +513,14 @@ and `--transform augment`); they are not part of the default `all`:
508
513
 
509
514
  | loader | CPython 3.14 (GIL) | free-threaded 3.14t |
510
515
  |-----------------|--------------------|---------------------|
511
- | **loaderx** | 128.9 | 687.9 batches/s |
512
- | loaderx-raw | 127.9 | 727.2 batches/s |
513
- | torch | 408.7 | n/a |
516
+ | **loaderx** | 24.3 | 41.1 batches/s |
517
+ | loaderx-raw | 22.3 | 43.3 batches/s |
514
518
 
515
- Under the GIL, loaderx's prefetch threads serialize the transform and fall to
516
- ~1/3 of torch's worker processes; free-threaded, they parallelize and loaderx
517
- reaches ~1.7x torch while keeping the zero-copy, IPC-free path.
519
+ Free-threaded still wins — the transform parallelizes across the prefetch
520
+ threads — ~1.7x on loaderx and ~1.9x on raw at 256 KiB. The margin is
521
+ narrower than at 12 KiB records (where a batch fit in cache and the transform
522
+ was the whole cost) because the 64 MiB gather is DRAM-bound, but wider than
523
+ at the 768 KiB scale, where the pipeline was even more DRAM-bound.
518
524
 
519
525
  **Conclusion** — why the numbers look like this.
520
526
 
@@ -522,15 +528,15 @@ and `--transform augment`); they are not part of the default `all`:
522
528
  zrecord gathers a whole batch and decompresses it, in a single cffi call into
523
529
  Zig with the GIL released and the batch copied straight into its destination
524
530
  buffer. The contenders do the same work one record at a time from Python. That
525
- one fact runs through all three tables: the sampler's margin is widest at the
526
- smallest batches (a batch costs one call either way, so the per-index work is
531
+ one fact runs through all three tables: the sampler's margin is roughly flat
532
+ across batch sizes (a batch costs one call either way, so the per-index work is
527
533
  what divides them), and the store gather column is where one-record-per-call
528
534
  costs the most. The speedup is not bought with distribution shortcuts: the IID
529
535
  draw is unbiased like NumPy's (Lemire with rejection, so uniformity costs
530
536
  nothing over a real index space).
531
537
 
532
538
  **The layout matches what a training loader does.** zrecord is built for random
533
- record access — the meta table is mmap'd, any record is found in O(1), a dense
539
+ record access — the record table is indexed in O(1), a dense
534
540
  store gathers at a fixed stride. Array stores are built for contiguous scans, so
535
541
  a scattered batch — exactly what a loader reads — fights their layout. A raw
536
542
  gather is a fan-out across every core, where NumPy's fancy index is one thread;
@@ -540,17 +546,17 @@ on the multi-core memory subsystem that closes the gap to an uncompressed
540
546
  **Compression is in the kernel, and there is one codec.** `zrecord-zstd` is not
541
547
  "storage plus a codec": the layout, the multi-core decompress and the GIL-free
542
548
  copy are one path, so turning compression on costs part of a margin, not an
543
- order of magnitude — ~10.7 GiB/s here, one to two orders of magnitude ahead of
544
- the other compressed stores (87–434 MiB/s) at the same ~1.1x ratio. The modest
545
- ratio is the data, not the format: these samples are barely compressible, and on
546
- a smooth image set plain zstd reaches 7.6x and the dictionary (at the `"max"`
547
- tier) 16.5x.
549
+ order of magnitude — ~3.9 GiB/s here, ~11x ahead of the other compressed stores
550
+ (109–353 MiB/s) at the same ~1.1x ratio. The modest ratio is the data, not the
551
+ format: these samples are barely compressible, and on a smooth image set plain
552
+ zstd reaches 7.6x and the dictionary (at the `"max"` tier) 16.5x.
548
553
 
549
554
  **The loader gap is architecture, not storage.** loaderx uses threads and never
550
555
  ends an epoch, so a step pays no IPC and never waits on an epoch boundary; torch
551
- restarts per epoch with worker processes. That is most of the ~5–6x end to end.
552
- Storage also differs per loader — each reads from what it was built for — so the
553
- loader table is a different comparison than the store table, not a rerun of it.
556
+ restarts per epoch with worker processes. That is most of the ~1.6x over torch
557
+ and ~3.8x over grain here. Storage also differs per loader — each reads from
558
+ what it was built for — so the loader table is a different comparison than the
559
+ store table, not a rerun of it.
554
560
  The one crack in the thread model is a CPU-heavy transform, which the GIL
555
561
  serializes — that is exactly what free-threaded Python removes, so loaderx is
556
562
  developed and benchmarked against free-threaded builds first.
@@ -563,7 +569,7 @@ page-cache ingestion with durability deferred, matching the other backends —
563
569
  zrecord's own durability (`sync`/`close`) is a separate, explicit cost that
564
570
  neither this table nor the competition pays. This is a single qualitative pass —
565
571
  the µs-scale sampler timings and the loader `batches/s` fluctuate with box load
566
- (on these 12 cores the compressed loader trails the raw one by ~25–30%), so
572
+ (on these 16 cores the compressed loader trails the raw one by ~25–30%), so
567
573
  treat the absolute numbers as ballpark and the cross-backend margins as the
568
574
  signal.
569
575
 
@@ -621,11 +627,15 @@ by record against the npy ground truth.
621
627
 
622
628
  ## Current Limitations
623
629
  * Single-host only; multi-host training is not supported.
624
- * A single sample must be at most 1 MiB, and a store holds at most 2^24
625
- (16777216) records.
626
- * Metadata is mapped and used in place, so a store carries the host's byte
627
- order and is not portable to a machine of the opposite endianness. Every
628
- published platform is little-endian, so this only matters if you build for
630
+ * A single sample must be at most 2 GiB (2^31 bytes). There is no fixed record
631
+ count: `length` is a u64 and the record table grows on demand, so how many
632
+ records a store holds is bounded by its total chunk capacity (up to 2^64
633
+ bytes) divided by the average record size — e.g. roughly 2^44 records at
634
+ 1 MiB each, 2^33 at 2 GiB each.
635
+ * Metadata is read and written as the host's struct layout, so a store carries
636
+ the host's byte order and is not portable to a machine of the opposite
637
+ endianness. Every published platform is little-endian, so this only matters
638
+ if you build for
629
639
  one yourself.
630
640
 
631
641
  ## Build
@@ -739,49 +749,57 @@ A record-based data runtime, focused on delivering extreme throughput and low la
739
749
  `zstd` store never touches a dictionary even if one is present.
740
750
 
741
751
  ## Persistence format
742
- Zrecord storage is metadata plus chunked data:
752
+ Zrecord storage is metadata plus chunked data. The extension is the type:
753
+ `.zr` files are store-global singletons, `.loc` files are record-table segments,
754
+ `.chunk` files are record data:
743
755
  ```
744
756
  zrecord/
745
- ├── meta.zr
746
- ├── 0.zr
747
- └── 1.zr
757
+ ├── meta.zr header (global state)
758
+ ├── dict.zr zstd dictionary (only in dict stores)
759
+ ├── 0.loc record table segment [0, 2^28)
760
+ ├── 1.loc record table segment [2^28, 2^29)
761
+ ├── 0.chunk record data
762
+ └── 1.chunk
748
763
  ```
749
764
 
750
- ### Metadata (meta.zr)
765
+ ### Metadata (meta.zr + {id}.loc)
751
766
 
752
- meta.zr is **used as the values it holds, not decoded into them**. Both structs
753
- below are `extern`, exactly 16 bytes each, and naturally aligned; an mmap is page
754
- aligned, so the header is a `*Header` and the record table is a `[]RecordLoc`.
755
- Looking up a record is `table[idx]`, and there is no serializer anywhere in the
756
- code.
767
+ Files are read and written **positionally** — pread/pwrite at computed offsets,
768
+ no mmap. The header is a bit-packed struct and a `.loc` segment an array of
769
+ 16-byte `RecordLoc`s, exactly as wide as they declare, so a location is one
770
+ pread/pwrite of 16 bytes at a computed offset and there is no serializer
771
+ anywhere in the code. Every header field is byte-aligned (no bit fields cross a
772
+ byte), so the packed header reads as plain memory.
757
773
 
758
- The fields are deliberately wider than the limits need — `chunk_id` is a `u16`
759
- for 4096 chunks, the lengths are `u32` for a 1 MiB cap. Bit-packing them would
760
- save 5 bytes per record and cost a hand-written codec on the hottest lookup
761
- path. The slack buys a machine-shaped array; it is worth it.
774
+ The record table is partitioned into `{id}.loc` segments so it can **grow by
775
+ appending a segment** instead of reserving the maximum. The id→segment mapping
776
+ is pure arithmetic — `seg = idx >> 28`, `off = (idx & (2^28−1)) × 16` — so a
777
+ segment needs no per-record bookkeeping. A segment is immutable once published,
778
+ so each mapping's base never moves and lock-free readers are safe to index it.
762
779
 
763
- **1. Header** — 16 bytes of global state, the same width as a `RecordLoc`, so
764
- the table starts at a clean 16-byte offset.
780
+ **1. Header** — 32 bytes, the whole of meta.zr.
765
781
  * `magic` (`ZREC`) and `version` are ordinary fields, so opening a directory
766
782
  that is not a Zrecord store fails immediately instead of decoding garbage.
767
783
  * `codec` is the store's one compression method, stamped at creation and
768
784
  immutable — there is no per-record tag anywhere.
769
- * `length` is the total number of records | `tail_chunk`/`tail_offset` mark the
770
- last write position.
785
+ * `length` (u64) is the total number of records | `tail_chunk`/`tail_offset`
786
+ mark the last write position.
771
787
  * There is no chunk count. Chunks are created in order and the frontier is
772
788
  always in the last one, so the store holds exactly chunks `0..=tail_chunk` —
773
789
  a count would be a second copy of that fact to keep in sync.
774
790
 
775
791
  ```zig
776
792
  const Codec = enum(u8) { raw = 0, zstd = 1, zstdict = 2, _ };
777
- const Header = extern struct {
778
- magic: [4]u8, version: u8, codec: Codec,
779
- tail_chunk: u16, tail_offset: u32, length: u32,
793
+ const Header = packed struct {
794
+ magic: u32, version: u8, codec: Codec,
795
+ tail_chunk: u32, tail_offset: u32, length: u64,
796
+ _reserved: u80,
780
797
  };
781
798
  ```
782
799
 
783
- **2. Record table** — 16 bytes per entry, indexed directly. Mapping an index to
784
- a physical address is what makes random access efficient.
800
+ **2. Record table** — a `.loc` segment is 2^28 entries of 16 bytes (4 GiB,
801
+ sparse), indexed directly. Mapping an index to a physical address is what makes
802
+ random access efficient.
785
803
  * `chunk_id` is the containing chunk | `offset` is the position within it |
786
804
  `phys_length`/`logic_length` are the stored and original sizes. The codec is
787
805
  not here: it is the header's, so a record is stored exactly the way the store
@@ -790,26 +808,17 @@ a physical address is what makes random access efficient.
790
808
  ```zig
791
809
  const RecordLoc = extern struct {
792
810
  offset: u32, phys_length: u32, logic_length: u32,
793
- chunk_id: u16, _reserved: [2]u8,
811
+ chunk_id: u32,
794
812
  };
795
813
  ```
796
814
 
797
815
  There is no liveness flag. Every entry below `length` is live, because deletion
798
816
  swaps the tail into the hole rather than tombstoning.
799
817
 
800
- **3. Maximum length.** Running out of space is effectively impossible, so the
801
- metadata is sized statically at the maximum.
802
- * Current design: up to 2^12 (4096) chunks, 2^32 (4 GiB) per chunk, 2^20 (1 MiB)
803
- per record.
804
- * Derivation: N = chunks × 2^32 (chunk size) / 2^20 (max record size) ⇒ a
805
- maximum length of 2^24 (16777216). Equality here is the point: the store can
806
- never run out of chunks before it runs out of record slots, which is what
807
- lets the table be sized statically. It is a compile-time assertion.
808
- * meta.zr is a fixed 16 + 2^24 × 16 bytes (256 MiB). It is sparse — a store
809
- with one record allocates 4 KiB of it — and the mapping does not prefault.
810
-
811
- ### Chunk data (x.zr)
812
- Densely packed record data.
818
+ **3. No maximum length.** The table grows a `.loc` segment at a time and the
819
+ data grows a chunk at a time, so there is no static record-count cap to size
820
+ against. The real bounds are the field widths — 2^32 chunks of 2^32 bytes
821
+ (2^64 bytes total), 2^31 (2 GiB) per record — and disk.
813
822
 
814
823
  ## Executor
815
824
 
@@ -829,9 +838,11 @@ Densely packed record data.
829
838
  side (executed on async threads).
830
839
  * Committed records are immutable, so the read path is lock free; `length` is
831
840
  published to readers through an atomic.
832
- * Every record is read at the offset its table entry records — the meta table is
833
- mmap'd, so random access is one array subscript and one positional read, with
834
- no batching assumptions about layout. Compressed records are read into a
841
+ * Every record is read at the offset its table entry records — the record table
842
+ is addressed by pure arithmetic, so random access is one pread for the
843
+ location and one for the bytes, with no batching assumptions about layout.
844
+ Locations are read once per batch (contiguous runs in one pread). Compressed
845
+ records are read into a
835
846
  per-shard staging buffer and decompressed in place into the destination.
836
847
 
837
848
  **3. Concurrency model.** `Io.Group.async` shards work by CPU count, and shards
@@ -864,7 +875,8 @@ is reclaimed by an offline `compact` that rewrites the live records in place.
864
875
  that compaction cannot remove.
865
876
 
866
877
  **5. File access.**
867
- * Metadata: a 256 MiB file created at init, accessed by mmap thereafter.
878
+ * Metadata: `meta.zr` (32 bytes) plus one file per `.loc` segment, each 4 GiB
879
+ (sparse), opened when the segment is created.
868
880
  * Chunk data: 4 GiB files created at init, accessed concurrently through
869
881
  `readPositionalAll`/`writePositionalAll`.
870
882